A pipe network leakage prediction method based on dynamic weight

CN122550143APending Publication Date: 2026-08-11TIANJIN TANGGU SINO FRENCH WATER SUPPLY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]本申请提供了一种基于动态权重的管网漏损预测方法,以至少解决相关技术中传统漏损预测方法依赖单一数据来源或专家经验、缺乏基于预测模型特征重要性的动态反馈与权重自适应修正机制,导致模型预测精度和泛化能力不足的问题:

Benefits of technology

[0034] This invention achieves a dynamic balance between subjective experience and objective data by integrating multi-source heterogeneous data, combining the subjective weights of the analytic hierarchy process (AHP) with the feature importance of the random forest model, and introducing a constrained optimization model for iterative correction. This significantly improves the accuracy and generalization ability of leakage prediction. At the same time, it has adaptive weight optimization and convergence control mechanisms, effectively overcoming the bias caused by traditional methods relying on a single data source or expert experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550143A_ABST
    Figure CN122550143A_ABST
Patent Text Reader

Abstract

This invention discloses a pipeline leakage prediction method based on dynamic weights, belonging to the field of pipeline leakage prediction technology. It addresses the problem of insufficient prediction accuracy caused by traditional methods relying on a single data source or expert experience and lacking dynamic weight correction. The method includes: acquiring multi-source data from the water supply network and processing it into multiple risk factors; calculating the initial weights of each risk factor using the analytic hierarchy process (AHP); constructing a leakage prediction model based on random forest, inputting the comprehensive risk index and the original risk factors into the model, and outputting the leakage probability of the pipeline segment; calculating the importance of each risk factor based on the random forest, integrating them into the initial weights to obtain corrected weights, solving for the optimal weights using a constrained optimization model, and iteratively updating until convergence. This invention integrates subjective weights and the importance of objective features, achieving adaptive weight optimization and significantly improving the accuracy and generalization ability of leakage prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pipeline leakage prediction technology, and in particular to a pipeline leakage prediction method based on dynamic weights. Background Technology

[0002] Water supply network leakage is a significant issue of water resource waste and operational safety, and accurate leakage risk prediction is crucial for network maintenance and management. Existing leakage prediction methods often rely on single types of statistical data or expert experience, such as statistical models based on historical leakage records or using only the analytic hierarchy process (AHP) to determine risk factor weights. These methods suffer from high subjectivity and difficulty in comprehensively integrating multi-source heterogeneous data. Furthermore, traditional methods lack dynamic feedback and adaptive weight correction mechanisms for the importance of prediction model features, limiting the model's prediction accuracy and generalization ability. Therefore, there is an urgent need for a water supply network leakage prediction method that can comprehensively utilize multi-source data, combine subjective and objective weights, and improve prediction performance through iterative optimization. Summary of the Invention

[0003] This application provides a pipeline leakage prediction method based on dynamic weights, which at least addresses the problems of traditional leakage prediction methods in related technologies that rely on a single data source or expert experience and lack dynamic feedback and weight adaptive correction mechanisms based on the importance of prediction model features, resulting in insufficient model prediction accuracy and generalization ability.

[0004] This application provides a method for predicting pipeline leakage based on dynamic weights, including the following steps:

[0005] S1: Obtain multi-source data of the water supply network, including network structure attributes, operating status, external environment, historical leakage and operation and maintenance management data, and process the multi-source data into risk factors, including pipe length, pipe age, pipe material, pipe diameter, burial depth, average pressure, road grade, soil corrosivity, number of repairs and historical leakage distance;

[0006] S2: The initial weights of each risk factor are calculated using the analytic hierarchy process (AHP).

[0007] S3: Construct a leakage prediction model based on random forest, using the comprehensive risk index of each risk factor and the original risk factor category feature value as the model input feature, and outputting the leakage probability of the pipe segment.

[0008] S4: Calculate the importance of each risk factor based on the random forest model, integrate the importance into the initial weights to obtain the corrected weights, solve the optimal weights using the constrained optimization model, and then input the optimal weights into the random forest model to repeat the model.

[0009] Furthermore, S3 also includes: combining the comprehensive weights of each risk factor calculated by the analytic hierarchy process with the risk factor type feature values ​​to construct a multidimensional input feature vector, wherein the comprehensive weights are expressed as:

[0010] ;

[0011] in, Let be the weight of the i-th risk factor.

[0012] Furthermore, historical leakage points are imported into the GIS map as point features, and a buffer analysis is established based on the location of the leakage points to obtain the distance of each pipe segment from the nearest historical leakage point, which is used as the historical leakage distance of that pipe segment.

[0013] Furthermore, in S4, the importance of each risk factor output by the random forest model is normalized and then fused with the initial weights of the analytic hierarchy process (AHP). The fusion formula is as follows:

[0014] ;

[0015] ;

[0016] in, The initial weights calculated for the analytic hierarchy process; The importance of normalized features in random forests; This is the balance coefficient, and its value ranges from 0 to 1. The importance of risk factor i; These are the weighted values ​​of the risk factors after fusion.

[0017] Furthermore, the objective function of the constrained optimization model is:

[0018] ;

[0019] The constraints are: , ;

[0020] in, This is the initial weight vector for the analytic hierarchy process. This represents the feature importance vector of the random forest. This is the adjustment coefficient.

[0021] Furthermore, in S2, the three-layer structure model constructed by the analytic hierarchy process includes:

[0022] The target layer is the comprehensive risk assessment of pipeline leakage.

[0023] The criteria layer, i.e., the various risk factors;

[0024] The indicator layer refers to the internal classification of risk factors.

[0025] Furthermore, the convergence condition in S4 is:

[0026] ;

[0027] in, Let be the weight vector obtained in the t-th iteration. This is the preset convergence threshold.

[0028] Furthermore, the comprehensive risk index is calculated as follows:

[0029] ;

[0030] in, Let be the initial weight of the i-th risk factor. Let be the feature value of the i-th risk factor category.

[0031] Furthermore, the random forest model can be replaced by a gradient boosting tree model or an XGBoost model to output the leakage probability of the pipeline segment.

[0032] Furthermore, in S1, the data on operating status, external environment, and operation and maintenance management are matched to the attribute table through the spatial connection tool of GIS, and historical leakage points are imported into the GIS map as point elements to form a set of risk factors containing the distance of historical leakage.

[0033] The application employs the above technical solution and has at least the following beneficial effects:

[0034] This invention achieves a dynamic balance between subjective experience and objective data by integrating multi-source heterogeneous data, combining the subjective weights of the analytic hierarchy process (AHP) with the feature importance of the random forest model, and introducing a constrained optimization model for iterative correction. This significantly improves the accuracy and generalization ability of leakage prediction. At the same time, it has adaptive weight optimization and convergence control mechanisms, effectively overcoming the bias caused by traditional methods relying on a single data source or expert experience.

[0035] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1This is an overall flowchart of the water supply network leakage prediction method provided in the embodiments of the present invention;

[0038] Figure 2 This is an AHP hierarchical structure diagram provided in an embodiment of the present invention. Detailed Implementation

[0039] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0040] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] Example 1

[0042] This embodiment provides a pipeline leakage prediction method based on dynamic weights, the overall process of which is as follows: Figure 1 As shown, the steps are performed sequentially as follows.

[0043] S1: Multi-source data acquisition and risk factor processing for water supply networks;

[0044] Acquire multi-source data of the target water supply network, including network structural attributes, operational status, external environment, historical leakage, and operation and maintenance management data. Specifically:

[0045] Pipeline network structure attributes: Extract the geometric attributes of the pipe segments from the GIS system, including pipe length, pipe age, pipe material, pipe diameter, and burial depth.

[0046] Operating status: Obtain the average pressure monitoring data of the pipeline section for the past year from the SCADA system.

[0047] External environment: Road grade and soil corrosivity grade at the location of the collection pipe section.

[0048] Historical Leakage and Operation and Maintenance Management: Obtain the location of historical leakage points, pipe burst records, and maintenance records.

[0049] The aforementioned operational status, external environment, and maintenance management data are matched to the attribute tables of each pipe segment using GIS spatial connectivity tools. Specifically, historical leakage points are imported into the GIS map as point features. A buffer analysis is established with each leakage point as the center, and the spatial distance from each pipe segment to its nearest historical leakage point is calculated and defined as the historical leakage distance. If multiple leakage points exist around a pipe segment, the nearest distance value is used. Finally, a set of risk factors is generated for each pipe segment, including: pipe length, pipe age, pipe material, pipe diameter, burial depth, average pressure, road grade, soil corrosivity, number of maintenance visits, and historical leakage distance.

[0050] S2: Calculate the initial weights of each risk factor based on the analytic hierarchy process (AHP);

[0051] Because the aforementioned risk factors have different dimensions and large differences in values, directly inputting them into a machine learning model would lead to numerical dominance bias. Therefore, the analytic hierarchy process (AHP) is used for dimensionless processing and initial weights are determined. A three-layer hierarchical model is constructed, including:

[0052] Target layer: Comprehensive risk assessment of pipeline leakage; Criteria layer: The above ten risk factors; Indicator layer: Internal categories of each risk factor.

[0053] Based on expert experience and historical data, the relative importance of each risk factor is compared pairwise to construct a judgment matrix. The eigenvalue method is then used to calculate the largest eigenvalue of the judgment matrix. The initial weight vectors of each risk factor are obtained by normalizing the corresponding eigenvectors. Calculate the consistency ratio CR. If CR < 0.1, the judgment matrix is ​​considered to meet the consistency requirements; otherwise, the judgment matrix is ​​readjusted.

[0054] Based on this, calculate the comprehensive risk index R for any pipe segment sample:

[0055] ;

[0056] in, Let be the category feature value of the i-th risk factor in this pipe segment.

[0057] S3: Construct a leakage prediction model based on random forest;

[0058] The comprehensive risk index R calculated by S2 is combined with the original 10 risk factor category feature values ​​to form a new model input feature vector:

[0059] ;

[0060] Target variable Y: Y=0 indicates that the pipe section is normal, and Y=1 indicates that the pipe section has leaked.

[0061] Building a training dataset , where m is the total number of training samples.

[0062] The Bootstrap method is used to randomly sample the training set with replacement to produce k subsets of the training set. In this embodiment, k=100. Each subset of the training set is used to train a decision tree. All decision trees are integrated to form a random forest model:

[0063] ;

[0064] For any pipe segment sample to be predicted, after inputting it into a random forest, each decision tree outputs the leakage probability. That is, the proportion of lost samples in the leaf nodes of the pipe section, and the final loss probability is:

[0065] .

[0066] S4: Dynamic weight adjustment and iterative optimization;

[0067] 4.1 Calculation of the importance of risk factors;

[0068] Based on a trained random forest model, the importance of each risk factor is calculated using the Gini index descent during the partitioning of each decision tree node. For risk factor i, its importance is... Defined as:

[0069] ;

[0070] in, Let $\frac{i}{k}$ be the total decrease in the Gini index resulting from splitting the $k$-th tree using risk factor $i$.

[0071] 4.2, Dynamic weight adjustment;

[0072] The initial weights of AHP are fused with the feature importance of Random Forest. First, the importance is normalized:

[0073] ;

[0074] Then through the balance coefficient To merge:

[0075] ;

[0076] in, The value range is 0-1. In this embodiment, the optimal value is determined by five-fold cross-validation. This minimizes the prediction error of the model on the validation set.

[0077] 4.3 Solving for the optimal weights using a constrained optimization model;

[0078] To improve the stability and rationality of the weights, a constrained optimization model is introduced, with the objective function being:

[0079] ;

[0080] The constraints are:

[0081] , ;

[0082] in, This is the adjustment coefficient, and in this embodiment, it is taken as... This is used to balance the deviations from AHP and RF weights. The optimization problem is a convex quadratic programming problem, which can be solved using the Lagrange multiplier method or numerical optimization algorithms. Numerical optimization algorithms include sequential least squares to obtain the optimal weight vector. .

[0083] 4.4 Iterative convergence mechanism;

[0084] The result As new risk factor weights, they are reintroduced into S3 to construct the random forest model. S3 and S4 are then executed repeatedly. The weight vector obtained in the t-th iteration is defined as... The convergence condition is:

[0085] ;

[0086] in, To preset the convergence threshold, this embodiment takes... When the above conditions are met, the iteration stops, and the current random forest model becomes the final prediction model.

[0087] Model application;

[0088] For any pipe segment to be predicted in the target water supply network, its risk factors and historical leakage distance are extracted according to S1. The comprehensive risk index is calculated according to the optimal weights that finally converge, and together with the original risk factors, it forms the input feature vector. This vector is then input into the final random forest model, which outputs the leakage probability of the pipe segment. When the probability exceeds a preset risk threshold, for example, exceeding 0.6, the segment is identified as a high-risk leakage segment, triggering an early warning.

[0089] Example 2

[0090] This embodiment is essentially the same as Embodiment 1, except that the Random Forest model in S3 is replaced by the XGBoost model. During training, the same input feature vector and target variable are used, and training is performed through the gradient boosting mechanism of XGBoost. For feature importance calculation, the weight or gain metric built into XGBoost is used instead of the Gini index decrease. The remaining dynamic weight adjustment, constraint optimization, and iterative convergence mechanisms remain unchanged. Practice shows that this alternative scheme can also achieve good leakage prediction accuracy and adaptive weight optimization capability.

[0091] Example 3

[0092] This embodiment is essentially the same as Embodiment 1, except that the analytic hierarchy process (AHP) in S2 is replaced by the entropy weight method. Specifically, based on the sample data of each risk factor, the information entropy of each risk factor is calculated, thereby determining its initial weight. The entropy weight method determines the weight entirely based on the data dispersion, avoiding expert subjectivity. In subsequent steps, the initial weights from the entropy weight method are fused and iteratively optimized with the importance of random forest features, which also achieves the purpose of this invention.

[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting pipeline network leakage based on dynamic weights, characterized in that, Includes the following steps: S1: Obtain multi-source data of the water supply network, including network structure attributes, operating status, external environment, historical leakage and operation and maintenance management data, and process the multi-source data into risk factors, including pipe length, pipe age, pipe material, pipe diameter, burial depth, average pressure, road grade, soil corrosivity, number of repairs and historical leakage distance; S2: The initial weights of each risk factor are calculated using the analytic hierarchy process (AHP). S3: Construct a leakage prediction model based on random forest, using the comprehensive risk index of each risk factor and the original risk factor category feature value as the model input feature, and outputting the leakage probability of the pipe segment. S4: Calculate the importance of each risk factor based on the random forest model, integrate the importance into the initial weights to obtain the corrected weights, solve the optimal weights using the constrained optimization model, and then input the optimal weights into the random forest model to repeat the model.

2. The pipeline leakage prediction method based on dynamic weights according to claim 1, characterized in that, S3 further includes: combining the comprehensive weights of each risk factor calculated by the analytic hierarchy process (AHP) with the risk factor type feature values ​​to construct a multidimensional input feature vector, wherein the comprehensive weights of each risk factor are expressed as: ; in, Let be the weight of the i-th risk factor.

3. The pipeline leakage prediction method based on dynamic weights according to claim 2, characterized in that, The historical leakage distance is obtained as follows: Import historical leakage points as point features into the GIS map, establish a buffer analysis based on the location of the leakage points, and obtain the distance of each pipe segment from the nearest historical leakage point, which is used as the historical leakage distance of that pipe segment.

4. The pipeline leakage prediction method based on dynamic weights according to claim 3, characterized in that, In step S4, the importance of each risk factor output by the random forest model is normalized and then fused with the initial weights of the analytic hierarchy process (AHP). The fusion formula is as follows: ; ; in, The initial weights calculated for the analytic hierarchy process; The importance of normalized features in random forests; This is the balance coefficient, and its value ranges from 0 to 1. The importance of risk factor i.

5. The pipeline leakage prediction method based on dynamic weights according to claim 4, characterized in that, The objective function of the constrained optimization model is: ; The constraints are: , ; in, This is the initial weight vector for the analytic hierarchy process. This represents the feature importance vector of the random forest. This is the adjustment coefficient.

6. The pipeline leakage prediction method based on dynamic weights according to claim 1, characterized in that, In S2, the three-layer structure model constructed by the analytic hierarchy process includes: The target layer is the comprehensive risk assessment of pipeline leakage. The criteria layer, i.e., the various risk factors; The indicator layer refers to the internal classification of risk factors.

7. The pipeline leakage prediction method based on dynamic weights according to claim 1, characterized in that, The convergence condition in S4 is: ; in, Let be the weight vector obtained in the t-th iteration. This is the preset convergence threshold.

8. The pipeline leakage prediction method based on dynamic weights according to claim 1, characterized in that, The calculation method for the comprehensive risk index is as follows: ; in, Let be the initial weight of the i-th risk factor. Let be the feature value of the i-th risk factor category.

9. The pipeline leakage prediction method based on dynamic weights according to claim 1, characterized in that: The random forest model can be replaced by a gradient boosting tree model or an XGBoost model to output the leakage probability of the pipeline segment.

10. The pipeline leakage prediction method based on dynamic weights according to claim 1, characterized in that: In step S1, the data on operating status, external environment, and operation and maintenance management are matched to the attribute table through the spatial connection tool of GIS, and historical leakage points are imported into the GIS map as point elements to form a set of risk factors containing the distance of historical leakage.