Method and system for monitoring leakage of buried gas pipeline under action of external disturbance

By using a dual-distributed optical fiber monitoring system and a multi-model integrated CatBoost regression prediction model, the blind spots and insufficient anti-interference capabilities of buried gas pipeline leakage monitoring under external disturbances have been solved, achieving comprehensive and accurate leakage and disturbance identification, which is suitable for long-distance buried pipeline monitoring.

CN121993744APending Publication Date: 2026-05-08CHINA UNIV OF MINING & TECH +3
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF MINING & TECH
Filing Date
2026-01-28
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies for monitoring leaks in buried gas pipelines under external disturbances suffer from problems such as large monitoring blind spots, lack of multi-parameter fusion mechanisms, insufficient anti-interference capabilities, and inadequate analysis of coupling mechanisms. These issues make it difficult to distinguish between disturbance signals and leakage signals, leading to missed or incorrect detections.

Method used

A dual-distributed optical fiber monitoring system is adopted, combined with a multi-model integrated CatBoost regression prediction model. Through multi-source data processing of temperature signal, disturbance sound wave signal and leakage sound signal, and using the CatBoost algorithm with Min-Max normalization, polynomial feature expansion and Bayesian optimization, accurate identification of leakage and disturbance is achieved.

Benefits of technology

It achieves high sensitivity and low false alarm identification of leaks and disturbances in complex environments, provides all-round monitoring and anti-interference capabilities, ensures leak location and disturbance identification without blind spots, and is suitable for monitoring long-distance buried pipelines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121993744A_ABST
    Figure CN121993744A_ABST
Patent Text Reader

Abstract

The invention discloses a buried gas pipeline leakage monitoring method and system under the action of external disturbance, and the method comprises the following steps: collecting multi-source data of a gas pipeline, and carrying out the preprocessing of the multi-source data, and obtaining an input feature; constructing a data set according to the input features, and dividing the data set into a training set and a test set; establishing three Catboost regression prediction models, namely a first model, a second model and a third model, and training the Catboost regression prediction models by using the training set data; wherein the first model is used for judging whether the pipeline leaks or not, the second model is used for identifying external disturbance, and the third model is used for identifying a pipeline leakage point / external disturbance position; the optimal hyper-parameter combination is obtained, the prediction effect of the CatBoost regression prediction model is checked based on test set data, and visualization and performance evaluation are carried out on the prediction result of the CatBoost regression prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gas pipeline leakage monitoring technology, and in particular to a method and system for monitoring leaks in buried gas pipelines under external disturbances. Background Technology

[0002] During transportation, distribution, and use, natural gas may leak due to pipeline aging, construction problems, or natural disasters. As the "lifeline" of urban energy transmission, the safe and stable operation of buried gas pipelines is directly related to public safety and socio-economic development. In recent years, with the acceleration of urbanization and increased development of underground space, pipeline leakage accidents caused by external disturbances have become frequent. Currently, most research on pipeline leakage monitoring technology focuses on overhead pipelines without external disturbances, with limited research on leakage identification in buried pipelines with complex operating conditions and significant external disturbances. Therefore, research on buried gas pipeline leakage monitoring technology applicable to environments with external disturbances has high practical significance.

[0003] Distributed fiber optic monitoring, as a highly sensitive and accurate monitoring method, can achieve distributed continuous monitoring over distances ranging from several meters to tens of kilometers. Leveraging millimeter- to meter-level spatial resolution, it constructs a blind-spot-free monitoring network to accurately capture subtle signals from localized micro-leaks or small-scale external disturbances. Simultaneously, it utilizes optical signal spatiotemporal analysis technology to achieve high-precision location of anomalies, providing precise guidance for on-site investigation. Patent CN114352947A proposes a method, system, device, and storage medium for detecting gas pipeline leaks, using an LSTM network based on a first sequence and a preset expert scoring model to identify the gas pipeline leak status. Patent CN116246659A proposes a method and system for processing acoustic signals from gas pipeline leaks, using blind source separation technology to extract low-frequency leakage acoustic wave features to identify the gas pipeline leak status. Patent CN118423622A proposes a method for distributed optical fiber monitoring of low-pressure gas pipeline leaks. When a low-pressure gas pipeline leaks, it generates sound waves that propagate forward through the soil, causing soil particles to vibrate. This vibration is then transmitted to the surface of the optical fiber, further altering the internal optical path and transmitting the leak signal back to the computer. After signal processing, the leak location and pipeline operating status can be visually observed. While these methods can fill gaps in gas pipeline leak monitoring to some extent, they still have significant limitations: 1. Large monitoring blind spots: Most existing technologies choose to lay the fiber parallel to the pipeline at a certain distance, which cannot achieve comprehensive monitoring of pipeline leaks. 2. Lack of multi-parameter fusion mechanism: Existing technologies mostly focus on monitoring and analyzing single physical parameters, failing to establish a collaborative correlation model between leak characteristic parameters and multiple external disturbance parameters. This makes it difficult to distinguish between "simple disturbance signals" and "disturbance-induced leak signals," and in complex disturbance environments, parameter coupling can easily lead to missed or false detections. 3. Insufficient anti-interference capability in complex environments: When faced with multi-source interference such as soil noise, surrounding traffic vibration, and industrial electromagnetic interference in buried scenarios, characteristic signals are easily submerged, and there is a lack of targeted anti-interference signal processing strategies. 4. Insufficient analysis of coupling mechanism: Traditional distributed monitoring mostly uses single fiber optic sensing technology, which can only obtain single parameters such as temperature or strain, making it difficult to fully reflect the coupling mechanism between leakage and external disturbance. Summary of the Invention

[0004] This solution addresses the problems and needs raised above by proposing a monitoring method and system for buried gas pipeline leakage under external disturbances. The above-mentioned technical objectives are achieved by adopting the following technical features, and several other technical benefits are also brought about.

[0005] One object of the present invention is to provide a method for monitoring leakage in buried gas pipelines under external disturbance, comprising the following steps: S10: Collect multi-source data from the gas pipeline, including temperature signal, disturbance sound wave signal and leakage sound signal, and preprocess the multi-source data to obtain input features; S20: Construct a dataset based on the input features, and divide the dataset into a training set and a test set; S30: Establish three CatBoost regression prediction models, namely the first model, the second model, and the third model, and train the CatBoost regression prediction models using the training set data; among them, the first model is used to determine whether the pipeline is leaking, the second model is used to identify external disturbances, and the third model is used to identify the pipeline leak point / external disturbance location. S40: Perform Bayesian optimization iteration on the CatBoost regression prediction model to obtain the optimal hyperparameter combination, then verify the prediction effect of the CatBoost regression prediction model based on the test set data, and visualize and evaluate the prediction results of the CatBoost regression prediction model.

[0006] Furthermore, the method for monitoring leaks in buried gas pipelines under external disturbances according to the present invention may also have the following technical features: In one example of the present invention, step S10 involves preprocessing the multi-source data, including the following steps: S11: Apply Min-Max normalization to all input variables, its mathematical definition is as follows: In the formula, The value is the normalized value, which is generally between 0 and 1; This is the maximum value of this feature in the original dataset; This is the minimum value of this feature in the original dataset; During the inference phase, in order to restore the model output to actual unit units, an inverse normalization transformation is performed on the normalized predicted values: S12: Further introduces a second-order polynomial feature expansion and cross-term construction mechanism, assuming the input vector is: Its expanded expression is: In the formula, x represents the original input feature vector. For the i-th feature component, For feature dimension, This represents the second-order interaction term between different features.

[0007] In one example of the present invention, the training process of the CatBoost model in step S40 is as follows: S41: Initialize a weak learner, usually a decision tree, denoted as... The prediction result is the initial predicted value. There is an error between the initial predicted value and the actual value. S42: In regression tasks, calculate the residuals for each sample, i.e., the true values. Compared with the current model predictions The difference ,in, Indicates the number of iterations; in classification tasks, it calculates the negative gradient of the loss function with respect to the current model predictions. S43: Use the calculated residual or negative gradient as the new target value and construct a new decision tree using a symmetric tree structure. ; S44: Update the current model based on the newly trained decision tree, using the following update formula: ,in, It is the learning rate, used to control the degree to which each tree contributes to model updates; S45: Repeat steps S42-S44 to continuously train new decision trees and update the model until the preset number of iterations is reached and the loss function converges to a certain extent.

[0008] In one example of the present invention, prior to step S41, the method further includes: preprocessing the raw data using an improved GTBS method, specifically including: Using weighting coefficients and prior distribution term Smoothing is performed to obtain an improved GTBS method and handle the discrete feature problem of GBDT. The expression of the improved GTBS method is as follows: In the formula, Let be the smoothed feature value of the k-th sample along the i-th feature dimension. For the first The sample at the th The values ​​taken on each feature dimension For the first The target variable corresponding to each sample For smoothing coefficients, These are prior values.

[0009] In one example of the present invention, in step S40, Bayesian optimization iteration is performed on the CatBoost regression prediction model to obtain the optimal hyperparameter combination, including: S401: Defines the search space for CatBoost hyperparameters; S402: Establish a Gaussian process as a surrogate model to approximate the mapping relationship between the objective function and hyperparameters of the CatBoost model; S403: The root mean square error of the CatBoost model under K-fold cross-validation is used as the objective function, and K sets of hyperparameter combinations are randomly selected as the initial evaluation points; S404: Perform L rounds of iterative optimization on the hyperparameters of CatBoost. In each round of iterative optimization, based on the data of all currently evaluated points, update the surrogate model and use the expected improvement acquisition function to calculate the expected improvement acquisition function EI value of all points in the search space. Select the hyperparameter combination with the largest EI value as the point to be evaluated in the next round, and train CatBoost using the hyperparameter combination with the largest EI value. At the same time, calculate the root mean square error of the hyperparameter combination with the largest EI value under five-fold cross-validation. S405: Complete a total of K+L objective function evaluations, including K initial evaluation points and L rounds of iterative optimization, and select the hyperparameter combination with the smallest root mean square error as the optimal hyperparameter combination.

[0010] In one example of this invention, the expression for the acquisition function EI is expected to be improved as follows: In the formula, { In order to improve, For hyperparameter combination, The mean of the predictions from the surrogate model. This is the best performance value observed so far. This represents the current optimal combination of hyperparameters. For adjustable exploration parameters, For intermediate parameters, The standard deviation of the surrogate model. The cumulative distribution function of the standard normal distribution. It is the probability density function of the standard normal distribution.

[0011] Another object of the present invention is to provide a monitoring system for leaks in buried gas pipelines under external disturbance, comprising: The data acquisition module is configured to acquire multi-source data from gas pipelines, including temperature signals, disturbance sound wave signals, and leakage sound signals. The module also preprocesses the multi-source data to obtain input features. The data partitioning module is configured to construct a dataset based on input features and divide the dataset into a training set and a test set. The model training module is configured to build three CatBoost regression prediction models, namely the first model, the second model, and the third model, and to train the CatBoost regression prediction models using training set data. The first model is used to determine whether a pipeline leak has occurred, the second model is used to identify external disturbances, and the third model is used to identify the pipeline leak point / location of external disturbances. The model prediction module is configured to perform Bayesian optimization iteration on the CatBoost regression prediction model to obtain the optimal hyperparameter combination, then verify the prediction effect of the CatBoost regression prediction model based on the test set data, and visualize and evaluate the prediction results of the CatBoost regression prediction model.

[0012] In one example of the present invention, the data partitioning module includes: The normalization unit, configured to apply Min-Max normalization to all input variables, is mathematically defined as follows: In the formula, The value is the normalized value, which is generally between 0 and 1; This is the maximum value of this feature in the original dataset; This is the minimum value of this feature in the original dataset; During the inference phase, in order to restore the model output to actual unit units, an inverse normalization transformation is performed on the normalized predicted values: The data expansion unit is configured to further introduce a second-order polynomial feature expansion and cross-term construction mechanism, assuming the input vector is: Its expanded expression is: In the formula, x represents the original input feature vector. For the i-th feature component, For feature dimension, This represents the second-order interaction term between different features.

[0013] In one example of the present invention, the model prediction module includes: The initial prediction unit is configured to initialize a weak learner, denoted as . The prediction result is the initial predicted value. There is an error between the initial predicted value and the actual value. The loss function unit is configured to calculate the residual, i.e., the true value, for each sample in a regression task. Compared with the current model predictions The difference ,in, Indicates the number of iterations; in classification tasks, it calculates the negative gradient of the loss function with respect to the current model predictions. Decision tree units use the calculated residuals or negative gradients as new target values ​​and construct a new decision tree using a symmetric tree structure. At the same time, regularization techniques are used to limit the depth of the decision tree and control the number of leaf nodes. The model update unit is configured to update the current model based on the newly trained decision tree, and the update formula is as follows: ,in, It is the learning rate, used to control the degree to which each tree contributes to model updates; The iterative unit is configured to continuously train new decision trees and update the model until a preset number of iterations is reached and the loss function converges to a certain extent.

[0014] In one example of the present invention, the model prediction module further includes: Improve the GTBS cell by configuring it to use weighting factors. and prior distribution term Smoothing is performed to obtain an improved GTBS method and handle the discrete feature problem of GBDT. The expression of the improved GTBS method is as follows: In the formula, Let be the smoothed feature value of the k-th sample along the i-th feature dimension. For the first The sample at the th The values ​​taken on each feature dimension For the first The target variable corresponding to each sample For smoothing coefficients, These are prior values.

[0015] Compared with the prior art, the present invention has the following beneficial effects: This invention employs a combined architecture of "dual distributed optical fibers + multi-model integration," specifically addressing the shortcomings of traditional single-fiber or single-sensor technologies. The two optical fibers are directly intertwined and laid, sensing the thermal effect of leakage through temperature signals. Distributed acoustic sensors capture the vibration of leaking airflow and the acoustic signals from external disturbances such as construction work and collisions, as well as the acoustic signals of pipeline leaks. These three sensors complement each other from both thermal and acoustic perspectives. Compared to existing single-fiber monitoring systems, which suffer from insufficient sensitivity to minute leaks and weak anti-interference capabilities, this multi-sensor combination can simultaneously acquire heterogeneous data from multiple sources, significantly improving the comprehensiveness of signal identification in complex buried environments. It avoids misjudgment based on a single parameter and accurately distinguishes between "leakage signals" and "environmental interference signals."

[0016] This invention employs normalization in the data preprocessing stage to effectively eliminate dimensional differences between different sensing devices and the influence of environmental noise, providing a high-quality data foundation for subsequent algorithm analysis. The invention introduces the Bayesian-optimized CatBoost algorithm, which dynamically adjusts model hyperparameters through Bayesian optimization. Compared to traditional machine learning algorithms (such as SVM and ordinary decision trees), it can more efficiently handle nonlinear correlations of multi-source sensor data. Simultaneously, CatBoost's advantages in category feature processing enhance model training efficiency and generalization ability, significantly reducing false alarm and false negative rates under complex operating conditions, and directly outputting three main results: "whether there is a leak, leak location, and dangerous disturbance identification."

[0017] This invention effectively identifies external threats such as third-party construction and pipeline collisions, achieving dual protection of "leakage early warning + external risk early warning," and proactively avoiding leakage causes. The dual-fiber entanglement laying method, combined with the long-distance continuous monitoring advantages of distributed sensing technology, eliminates the need for numerous relay devices, making it suitable for monitoring scenarios of long-distance buried pipelines. Simultaneously, the distributed architecture ensures no monitoring blind spots, providing more comprehensive coverage compared to point-based sensor arrays. The system as a whole possesses strong anti-interference capabilities and good real-time performance, and can operate stably in complex electromagnetic environments and harsh geological conditions, perfectly matching the actual operating environment of buried pipelines.

[0018] The preferred embodiments of the invention will be described in more detail below with reference to the accompanying drawings, so as to facilitate an understanding of the features and advantages of the invention. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly described below. The drawings are merely illustrative of some embodiments of the present invention and are not intended to limit the scope of the present invention to all embodiments.

[0020] Figure 1 A flowchart illustrating a method for monitoring leaks in buried gas pipelines under external disturbances according to an embodiment of the present invention; Figure 2This is a schematic diagram of the CatBoost algorithm according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the multi-model integration process according to an embodiment of the present invention; Figure 4 The flowchart for Bayesian optimization of the CatBoost algorithm hyperparameters is shown in the embodiment of the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. The same reference numerals in the drawings represent the same components. It should be noted that the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0022] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms “first,” “second,” and similar terms used in this patent application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, “an” or “a” and similar terms do not necessarily indicate a quantity limitation. Terms such as “comprising” or “including” mean that the element or object preceding the word encompasses the element or object listed following the word and its equivalents, without excluding other elements or objects. Terms such as “connected” or “linked” are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as “upper,” “lower,” “left,” and “right” are used only to indicate relative positional relationships; these relative positional relationships may change accordingly when the absolute position of the described object changes.

[0023] According to a first aspect of the present invention, a method for monitoring leaks in buried gas pipelines under external disturbance is provided, such as... Figure 1 , Figure 3 As shown, it includes the following steps: S10: Collect multi-source data from the gas pipeline, including temperature signal, disturbance sound wave signal and leakage sound signal, and preprocess the multi-source data to obtain input features; Specifically, the fiber optic monitoring system is first laid: a distributed temperature fiber and a distributed acoustic sensing fiber are wound and laid on the pipeline, with the fibers spaced at regular intervals and parallel to each other, and bonded together with epoxy adhesive throughout. The distributed temperature fiber is connected to a distributed fiber optic temperature demodulator, and the distributed acoustic sensing fiber is connected to the DAS (Distributed Acoustic Analyzer). Then, an oscilloscope and data acquisition card software are turned on. The oscilloscope can be used as a real-time signal display screen, allowing for real-time observation of signal fluctuations and facilitating the observation of leakage signals. The data acquisition card software acquires data, and the stored data is processed in Matlab. Finally, after each set of data for a given operating condition is acquired, the data is recorded, and the equipment power is promptly turned off.

[0024] S20: Construct a dataset based on the input features, and divide the dataset into a training set and a test set; S30: Establish three CatBoost regression prediction models, namely the first model, the second model, and the third model, and train the CatBoost regression prediction models using the training set data; among them, the first model is used to determine whether the pipeline is leaking, the second model is used to identify external disturbances, and the third model is used to identify the pipeline leak point / external disturbance location. S40: The CatBoost regression prediction model is subjected to Bayesian optimization iteration to obtain the optimal hyperparameter combination. Then, the prediction effect of the CatBoost regression prediction model is tested based on the test set data. The prediction results of the CatBoost regression prediction model are visualized and its performance is evaluated. In other words, using the Bayesian-optimized CatBoost algorithm, with dual-fiber and distributed acoustic monitoring signals as input, three results are output: 1) whether a pipeline leak has occurred, 2) the location of the leak point, and 3) whether there are dangerous disturbances around the pipeline.

[0025] This monitoring method employs a combined architecture of "dual distributed optical fibers + multi-model integration," specifically addressing the shortcomings of traditional single-fiber or single-sensor technologies. The two optical fibers are directly intertwined and laid, sensing the thermal effect of leakage through temperature signals. Distributed acoustic sensors capture the vibration of leaking airflow and the acoustic signals from external disturbances such as construction and collisions, as well as the acoustic signals of pipeline leaks. These three sensors complement each other from both thermal and acoustic perspectives. Compared to existing single-fiber monitoring systems, which suffer from insufficient sensitivity to minute leaks and weak anti-interference capabilities, this multi-sensor combination can simultaneously acquire heterogeneous data from multiple sources, significantly improving the comprehensiveness of signal identification in complex buried environments. It avoids misjudgment based on a single parameter and accurately distinguishes between "leakage signals" and "environmental interference signals."

[0026] The monitoring method employs normalization in the data preprocessing stage to effectively eliminate dimensional differences between different sensing devices and the influence of environmental noise, providing a high-quality data foundation for subsequent algorithm analysis. The method introduces the Bayesian-optimized CatBoost algorithm, which dynamically adjusts model hyperparameters through Bayesian optimization. Compared to traditional machine learning algorithms (such as SVM and ordinary decision trees), it can more efficiently handle nonlinear correlations of multi-source sensor data. Furthermore, CatBoost's advantages in category feature processing enhance model training efficiency and generalization ability, significantly reducing false alarm and false negative rates under complex operating conditions, and directly outputting three main results: "whether there is a leak, leak location, and dangerous disturbance identification."

[0027] This monitoring method effectively identifies external threats such as third-party construction and pipeline collisions, achieving dual protection of "leakage early warning + external risk early warning" and proactively avoiding leakage causes. The dual-fiber entanglement laying method, combined with the long-distance continuous monitoring advantages of distributed sensing technology, eliminates the need for numerous relay devices, making it suitable for monitoring long-distance buried pipelines. Simultaneously, the distributed architecture ensures no monitoring blind spots, providing more comprehensive coverage compared to point-based sensor arrays. The system as a whole possesses strong anti-interference capabilities and good real-time performance, enabling stable operation in complex electromagnetic environments and harsh geological conditions, perfectly suited to the actual operating environment of buried pipelines.

[0028] In one example of the present invention, step S10 involves preprocessing the multi-source data, including the following steps: S11: To mitigate the impact of different dimensional features on the model training process, this method applies Min-Max normalization to all input variables, the mathematical definition of which is as follows: In the formula, The value is the normalized value, which is generally between 0 and 1; This represents the maximum value of the feature (variable) in the original dataset; It is the minimum value of this feature (variable) in the original dataset; During the inference phase, in order to restore the model output to actual unit units, an inverse normalization transformation is performed on the normalized predicted values: S12: To improve the model's ability to fit higher-order variable relationships, this method further introduces a second-order polynomial feature expansion and cross-term construction mechanism. Let the input vector be: Its expanded expression is: In the formula, x represents the original input feature vector. For the i-th feature component, For feature dimension, These are second-order interaction terms between different features, used to model the interaction relationships between features.

[0029] In one example of the present invention, in step S30, the detailed steps of monitoring the pipeline by three CatBoost regression prediction models are as follows: firstly, the first model and the second model monitor pipeline leakage and external disturbances respectively. After the above monitoring, the third model monitors abnormal locations, wherein the abnormal location monitoring includes dangerous external disturbance locations and dangerous external leakage points.

[0030] In one example of the present invention, in step S40, as Figure 2 As shown, the training process of the CatBoost model is as follows: S41: Initialize a weak learner, usually a decision tree, denoted as... The prediction result is the initial predicted value. There is an error between the initial predicted value and the actual value. S42: In regression tasks, calculate the residuals for each sample, i.e., the true values. Compared with the current model predictions The difference ,in, Indicates the number of iterations; in classification tasks, it calculates the negative gradient of the loss function with respect to the current model predictions. S43: Using the calculated residual (for regression tasks) or negative gradient (for classification tasks) as the new target value, construct a new decision tree using a symmetric tree structure. At the same time, regularization methods are used to limit the depth of the decision tree and control the number of leaf nodes. S44: Update the current model based on the newly trained decision tree, using the following update formula: ,in, It is the learning rate (also known as the step size), which controls the degree to which each tree contributes to model updates; S45: Repeat steps S42-S44 to continuously train new decision trees and update the model until the preset number of iterations is reached and the loss function converges to a certain extent.

[0031] In one example of the present invention, prior to step S41, the method further includes: preprocessing the raw data using an improved GTBS method, specifically including: Using weighting coefficients and prior distribution term Smoothing is performed to reduce the impact of data at specific frequencies on the overall distribution, thereby improving the GTBS method and addressing the discrete feature problem of GBDT. The expression for the improved GTBS method is as follows: In the formula, Let be the smoothed feature value of the k-th sample along the i-th feature dimension. For the first The sample at the th The values ​​taken on each feature dimension For the first The target variable corresponding to each sample For smoothing coefficients, These are prior values.

[0032] In one example of the present invention, in step S41, obtaining the learner includes: Assumption and These are the strong and weak learners obtained in the previous iteration. Expectation function If it is a loss function, then the objective function in each iteration is... : application The negative gradient is used to fit an approximate value of the loss in each iteration, then in the equation... for In summary, the strong learner for this iteration is obtained. : In the formula: This is the step size for each update round.

[0033] In one example of the present invention, step S43 further includes: The model's base predictor uses a symmetric tree and applies the same splitting criterion to the tree. S41 binarizes the information, thus outputting the predicted data fed into the model and storing it. In vectors; definition Given the depth of the tree, store the leaf node values ​​in the form of In the vector; then numbered Convert the indices of the leaf nodes of the tree into binary vectors. As shown in the following formula: Among them, in specific samples The characteristics of binary obtained from for ;No. The number of binary features of the tree is .

[0034] In one example of the present invention, in step S40, as Figure 4 As shown, the optimal hyperparameter combination is obtained by Bayesian optimization iteration on the CatBoost regression prediction model, including: To optimize the key hyperparameters of CatBoost, this method introduces a Bayesian optimization approach, the basic idea of ​​which is to optimize the parameter space. Build a proxy function To approximate the objective function (such as the RMSE after cross-validation): S401: Defines the search space for CatBoost hyperparameters; S402: Establish a Gaussian process as a surrogate model to approximate the mapping relationship between the objective function and hyperparameters of the CatBoost model; S403: The root mean square error of the CatBoost model under K-fold cross-validation is used as the objective function, and K sets of hyperparameter combinations are randomly selected as the initial evaluation points; S404: Perform L rounds of iterative optimization on the hyperparameters of CatBoost. In each round of iterative optimization, based on the data of all currently evaluated points, update the surrogate model and use the expected improvement acquisition function to calculate the expected improvement acquisition function EI value of all points in the search space. Select the hyperparameter combination with the largest EI value as the point to be evaluated in the next round, and train CatBoost using the hyperparameter combination with the largest EI value. At the same time, calculate the root mean square error of the hyperparameter combination with the largest EI value under five-fold cross-validation. S405: Complete a total of K+L objective function evaluations, including K initial evaluation points and L rounds of iterative optimization, and select the hyperparameter combination with the smallest root mean square error as the optimal hyperparameter combination.

[0035] In one example of this invention, the expression for the acquisition function EI is expected to be improved as follows: In the formula, { In order to improve, For hyperparameter combination, The mean of the predictions from the surrogate model. This is the best performance value observed so far. This represents the current optimal combination of hyperparameters. For adjustable exploration parameters, For intermediate parameters, The standard deviation of the surrogate model. The cumulative distribution function of the standard normal distribution. It is the probability density function of the standard normal distribution.

[0036] A monitoring system for buried gas pipeline leakage under external disturbance according to a second aspect of the present invention includes: The data acquisition module is configured to acquire multi-source data from gas pipelines, including temperature signals, disturbance sound wave signals, and leakage sound signals. The module also preprocesses the multi-source data to obtain input features. Specifically, the fiber optic monitoring system is first laid: a distributed temperature fiber and a distributed acoustic sensing fiber are wound and laid on the pipeline, with the fibers spaced at regular intervals and parallel to each other, and bonded together with epoxy adhesive throughout. The distributed temperature fiber is connected to a distributed fiber optic temperature demodulator, and the distributed acoustic sensing fiber is connected to the DAS (Distributed Acoustic Analyzer). Then, an oscilloscope and data acquisition card software are turned on. The oscilloscope can be used as a real-time signal display screen, allowing for real-time observation of signal fluctuations and facilitating the observation of leakage signals. The data acquisition card software acquires data, and the stored data is processed in Matlab. Finally, after each set of data for a given operating condition is acquired, the data is recorded, and the equipment power is promptly turned off.

[0037] The data partitioning module is configured to construct a dataset based on input features and divide the dataset into a training set and a test set. The model training module is configured to build three CatBoost regression prediction models, namely the first model, the second model, and the third model, and to train the CatBoost regression prediction models using training set data. The first model is used to determine whether a pipeline leak has occurred, the second model is used to identify external disturbances, and the third model is used to identify the pipeline leak point / location of external disturbances. The model prediction module is configured to perform Bayesian optimization iteration on the CatBoost regression prediction model to obtain the optimal hyperparameter combination, then verify the prediction effect of the CatBoost regression prediction model based on the test set data, and visualize and evaluate the prediction results of the CatBoost regression prediction model.

[0038] This monitoring system employs a combined architecture of "dual distributed optical fibers + multi-model integration," specifically addressing the shortcomings of traditional single-fiber or single-sensor technologies. The two optical fibers are directly intertwined and laid, sensing the thermal effect of leakage through temperature signals. Distributed acoustic sensors capture the vibration of leaking airflow and the acoustic signals from external disturbances such as construction and collisions, as well as the acoustic signals of pipeline leaks. These three sensors complement each other from both thermal and acoustic perspectives. Compared to existing single-fiber monitoring systems, which suffer from insufficient sensitivity to minute leaks and weak anti-interference capabilities, this multi-sensor combination can simultaneously acquire heterogeneous data from multiple sources, significantly improving the comprehensiveness of signal identification in complex buried environments. It avoids misjudgment based on a single parameter and accurately distinguishes between "leakage signals" and "environmental interference signals."

[0039] The monitoring system employs normalization in its data preprocessing stage to effectively eliminate dimensional differences between different sensors and the impact of environmental noise, providing a high-quality data foundation for subsequent algorithm analysis. The system introduces the Bayesian-optimized CatBoost algorithm, which dynamically adjusts model hyperparameters through Bayesian optimization. Compared to traditional machine learning algorithms (such as SVM and ordinary decision trees), it can more efficiently handle nonlinear correlations of multi-source sensor data. Furthermore, CatBoost's superior category feature processing enhances model training efficiency and generalization ability, significantly reducing false alarm and false negative rates under complex operating conditions, and directly outputting three main results: "whether there is a leak, leak location, and dangerous disturbance identification."

[0040] This monitoring system effectively identifies external threats such as third-party construction and pipeline collisions, achieving dual protection of "leakage early warning + external risk early warning" to proactively avoid potential leaks. The dual-fiber optic cable winding installation combined with the long-distance continuous monitoring advantages of distributed sensing technology eliminates the need for numerous relay devices, making it suitable for monitoring long-distance buried pipelines. Simultaneously, the distributed architecture ensures no monitoring blind spots, providing more comprehensive coverage compared to point-based sensor arrays. The system as a whole possesses strong anti-interference capabilities and excellent real-time performance, enabling stable operation in complex electromagnetic environments and harsh geological conditions, perfectly suited to the actual operating environment of buried pipelines.

[0041] In one example of the present invention, the data partitioning module includes: The normalization unit, configured to apply Min-Max normalization to all input variables, is mathematically defined as follows to mitigate the impact of features with different dimensions on the model training process: In the formula, The value is the normalized value, which is generally between 0 and 1; This represents the maximum value of the feature (variable) in the original dataset; It is the minimum value of this feature (variable) in the original dataset; During the inference phase, in order to restore the model output to actual unit units, an inverse normalization transformation is performed on the normalized predicted values: The data expansion unit, configured to further introduce a second-order polynomial feature expansion and cross-term construction mechanism, aims to improve the model's ability to fit higher-order variable relationships. The input vector is set as follows: Its expanded expression is: In the formula, x represents the original input feature vector. For the i-th feature component, For feature dimension, These are second-order interaction terms between different features, used to model the interaction relationships between features.

[0042] In one example of the present invention, the model prediction module includes: The initial prediction unit is configured to initialize a weak learner, typically a decision tree, denoted as . The prediction result is the initial predicted value. There is an error between the initial predicted value and the actual value. The loss function unit is configured to calculate the residual, i.e., the true value, for each sample in a regression task. Compared with the current model predictions The difference ,in, Indicates the number of iterations; in classification tasks, it calculates the negative gradient of the loss function with respect to the current model predictions. Decision tree units use the calculated residuals (for regression tasks) or negative gradients (for classification tasks) as new target values, and construct a new decision tree using a symmetric tree structure. At the same time, regularization techniques are used to limit the depth of the decision tree and control the number of leaf nodes. The model update unit is configured to update the current model based on the newly trained decision tree, and the update formula is as follows: ,in, It is the learning rate (also known as the step size), which controls the degree to which each tree contributes to model updates; The iterative unit is configured to continuously train new decision trees and update the model until a preset number of iterations is reached and the loss function converges to a certain extent.

[0043] In one example of the present invention, the model prediction module further includes: Improve the GTBS cell by configuring it to use weighting factors. and prior distribution term Smoothing is performed to reduce the impact of data at specific frequencies on the overall distribution, thereby improving the GTBS method and addressing the discrete feature problem of GBDT. The expression for the improved GTBS method is as follows: In the formula, Let be the smoothed feature value of the k-th sample along the i-th feature dimension. For the first The sample at the th The values ​​taken on each feature dimension For the first The target variable corresponding to each sample For smoothing coefficients, These are prior values.

[0044] It should be noted that the monitoring system for buried gas pipeline leakage under external disturbance of the present invention can also perform any of the processing described in the previously described method for monitoring buried gas pipeline leakage under external disturbance, the specific details of which will not be repeated here.

[0045] The foregoing description, with reference to preferred embodiments, details an exemplary implementation of the method and system for monitoring leaks in buried gas pipelines under external disturbances proposed in this invention. However, those skilled in the art will understand that various modifications and alterations can be made to the above specific embodiments without departing from the concept of this invention, and various combinations can be made to the various technical features and structures proposed in this invention without exceeding the protection scope of this invention, which is determined by the appended claims.

Claims

1. A method for monitoring leaks in buried gas pipelines under external disturbance, characterized in that, Includes the following steps: S10: Collect multi-source data from the gas pipeline, including temperature signal, disturbance sound wave signal and leakage sound signal, and preprocess the multi-source data to obtain input features; S20: Construct a dataset based on the input features, and divide the dataset into a training set and a test set; S30: Establish three CatBoost regression prediction models, namely the first model, the second model, and the third model, and train the CatBoost regression prediction models using the training set data; among them, the first model is used to determine whether the pipeline is leaking, the second model is used to identify external disturbances, and the third model is used to identify the pipeline leak point / external disturbance location. S40: Perform Bayesian optimization iteration on the CatBoost regression prediction model to obtain the optimal hyperparameter combination, then verify the prediction effect of the CatBoost regression prediction model based on the test set data, and visualize and evaluate the prediction results of the CatBoost regression prediction model.

2. The method for monitoring leaks in buried gas pipelines under external disturbance as described in claim 1, characterized in that, In step S10, the multi-source data is preprocessed, including the following steps: S11: Apply Min-Max normalization to all input variables, its mathematical definition is as follows: In the formula, The value is the normalized value, which is generally between 0 and 1; This is the maximum value of this feature in the original dataset; This is the minimum value of this feature in the original dataset; During the inference phase, in order to restore the model output to actual unit units, an inverse normalization transformation is performed on the normalized predicted values: S12: Further introduces a second-order polynomial feature expansion and cross-term construction mechanism, assuming the input vector is: Its expanded expression is: In the formula, x represents the original input feature vector. For the i-th feature component, For feature dimension, This represents the second-order interaction term between different features.

3. The method for monitoring leaks in buried gas pipelines under external disturbance as described in claim 1, characterized in that, In step S40, the training process of the CatBoost regression prediction model is as follows: S41: Initialize a weak learner, usually a decision tree, denoted as... The prediction result is the initial predicted value. There is an error between the initial predicted value and the actual value. S42: In regression tasks, calculate the residuals for each sample, i.e., the true values. Compared with the current model predictions The difference ,in, Indicates the number of iterations; in classification tasks, it calculates the negative gradient of the loss function with respect to the current model prediction. Its expression is: S43: Use the calculated residual or negative gradient as the new target value and construct a new decision tree using a symmetric tree structure. ; S44: Update the current model based on the newly trained decision tree, using the following update formula: ,in, It is the learning rate, used to control the degree to which each tree contributes to model updates; S45: Repeat steps S42-S44 to continuously train new decision trees and update the model until the preset number of iterations is reached and the loss function converges to a certain extent.

4. The method for monitoring leaks in buried gas pipelines under external disturbance as described in claim 1, characterized in that, Before step S41, the method further includes: preprocessing the raw data using an improved GTBS method, specifically including: Using weighting coefficients and prior distribution term Smoothing is performed to obtain an improved GTBS method and handle the discrete feature problem of GBDT. The expression of the improved GTBS method is as follows: In the formula, Let be the smoothed feature value of the k-th sample along the i-th feature dimension. For the first The sample at the th The values ​​taken on each feature dimension For the first The target variable corresponding to each sample For smoothing coefficients, These are prior values.

5. The method for monitoring leakage of buried gas pipelines under external disturbance as described in claim 1, characterized in that, In step S40, the optimal hyperparameter combination is obtained by performing Bayesian optimization iteration on the CatBoost regression prediction model, including: S401: Defines the search space for CatBoost hyperparameters; S402: Establish a Gaussian process as a surrogate model to approximate the mapping relationship between the objective function and hyperparameters of the CatBoost model; S403: The root mean square error of the CatBoost model under K-fold cross-validation is used as the objective function, and K sets of hyperparameter combinations are randomly selected as the initial evaluation points. S404: Perform L rounds of iterative optimization on the hyperparameters of CatBoost. In each round of iterative optimization, based on the data of all currently evaluated points, update the surrogate model and use the expected improvement acquisition function to calculate the expected improvement acquisition function EI value of all points in the search space. Select the hyperparameter combination with the largest EI value as the point to be evaluated in the next round, and train CatBoost using the hyperparameter combination with the largest EI value. At the same time, calculate the root mean square error of the hyperparameter combination with the largest EI value under five-fold cross-validation. S405: Complete a total of K+L objective function evaluations, including K initial evaluation points and L rounds of iterative optimization, and select the hyperparameter combination with the smallest root mean square error as the optimal hyperparameter combination.

6. The method for monitoring leaks in buried gas pipelines under external disturbance as described in claim 5, characterized in that, The desired improved expression for the acquisition function EI is: In the formula, { In order to improve, For hyperparameter combination, The mean of the predictions from the surrogate model. This is the best performance value observed so far. This represents the current optimal combination of hyperparameters. For adjustable exploration parameters, For intermediate parameters, The standard deviation of the surrogate model. The cumulative distribution function of the standard normal distribution. It is the probability density function of the standard normal distribution.

7. A monitoring system for leaks in buried gas pipelines under external disturbance, characterized in that, include: The data acquisition module is configured to acquire multi-source data from gas pipelines, including temperature signals, disturbance sound wave signals, and leakage sound signals. The module also preprocesses the multi-source data to obtain input features. The data partitioning module is configured to construct a dataset based on input features and divide the dataset into a training set and a test set. The model training module is configured to build three CatBoost regression prediction models, namely the first model, the second model, and the third model, and to train the CatBoost regression prediction models using training set data. The first model is used to determine whether a pipeline leak has occurred, the second model is used to identify external disturbances, and the third model is used to identify the pipeline leak point / location of external disturbances. The model prediction module is configured to perform Bayesian optimization iteration on the CatBoost regression prediction model to obtain the optimal hyperparameter combination, then verify the prediction effect of the CatBoost regression prediction model based on the test set data, and visualize and evaluate the prediction results of the CatBoost regression prediction model.

8. The monitoring system for buried gas pipeline leakage under external disturbance as described in claim 7, characterized in that, The data partitioning module includes: The normalization unit, configured to apply Min-Max normalization to all input variables, is mathematically defined as follows: In the formula, The value is the normalized value, which is generally between 0 and 1; This is the maximum value of this feature in the original dataset; This is the minimum value of this feature in the original dataset; During the inference phase, in order to restore the model output to actual unit units, an inverse normalization transformation is performed on the normalized predicted values: The data expansion unit is configured to further introduce a second-order polynomial feature expansion and cross-term construction mechanism, assuming the input vector is: Its expanded expression is: In the formula, x represents the original input feature vector. For the i-th feature component, For feature dimension, This represents the second-order interaction term between different features.

9. The monitoring system for buried gas pipeline leakage under external disturbance as described in claim 7, characterized in that, The model prediction module includes: The initial prediction unit is configured to initialize a weak learner, denoted as . The prediction result is the initial predicted value. There is an error between the initial predicted value and the actual value. The loss function unit is configured to calculate the residual, i.e., the true value, for each sample in a regression task. Compared with the current model predictions The difference ,in, Indicates the number of iterations; in classification tasks, it calculates the negative gradient of the loss function with respect to the current model predictions. Decision tree units use the calculated residuals or negative gradients as new target values ​​and construct a new decision tree using a symmetric tree structure. ; The model update unit is configured to update the current model based on the newly trained decision tree, and the update formula is as follows: ,in, It is the learning rate, used to control the degree to which each tree contributes to model updates; The iterative unit is configured to continuously train new decision trees and update the model until a preset number of iterations is reached and the loss function converges to a certain extent.

10. The monitoring system for buried gas pipeline leakage under external disturbance as described in claim 7, characterized in that, The model prediction module also includes: Improve the GTBS cell by configuring it to use weighting coefficients. and prior distribution term Smoothing is performed to obtain an improved GTBS method and handle the discrete feature problem of GBDT. The expression of the improved GTBS method is as follows: In the formula, Let be the smoothed feature value of the k-th sample along the i-th feature dimension. For the first The sample at the th The values ​​taken on each feature dimension For the first The target variable corresponding to each sample For smoothing coefficients, These are prior values.

Citation Information

Patent Citations

  • Gas pipeline leakage detection method, system and device and storage medium

    CN114352947A

  • Gas pipeline leakage sound wave signal processing selection method and system

    CN116246659A

  • Method for monitoring leakage of low-pressure gas pipeline through distributed optical fibers

    CN118423622A