Municipal pipe network overflow troubleshooting and risk assessment method
Through comprehensive data processing and machine learning technology, a municipal pipeline overflow prediction model and an ecological environment impact assessment system are built, which solves the problems of scarcity and inaccurate assessment in the existing technology, and realizes accurate prediction and risk assessment of pipeline overflow.
Patent Information
- Application Number
- CN202510211097.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
The municipal pipeline overflow investigation and risk assessment methods have problems such as scarce data, weak generalization capabilities of model and inaccurate ecological environment impact assessment.
By comprehensively collecting and processing municipal pipeline network and its surrounding environment data, using machine learning and data analysis technology, an overflow prediction model based on generative adversarial network (GAN) and deep neural networks is constructed, and the risk assessment is carried out by combining hierarchical analysis method and entropy weight method.
It realizes accurate prediction and risk assessment of pipeline overflow, improves the generalization ability of the model and the accuracy of ecological and environmental impact assessment, and provides a scientific basis for municipal pipeline management.
Smart Images

Figure CN120069556A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of municipal engineering, and specifically to a method for detecting and risk assessing municipal pipe network overflows. Background Art
[0002] As a key component of the urban water cycle system, the operation status of the municipal pipe network is directly related to the efficiency of urban drainage and the quality of the environment. However, in actual operation, the pipe network often overflows due to various reasons, which not only causes urban waterlogging but also may have a serious impact on the groundwater, soil and vegetation ecological environment. Therefore, it is particularly important to carry out the work of detecting and risk assessing municipal pipe network overflows.
[0003] Traditional methods for detecting and risk assessing municipal pipe network overflows have certain limitations. For example, limited monitoring data leads to inaccurate model predictions, and the risk assessment indicators are single, without fully considering the impact on the ecological environment.
[0004] In summary, the work of detecting and risk assessing municipal pipe network overflows faces multiple challenges such as scarce data, weak model generalization ability, and inaccurate ecological environment impact assessment. Therefore, it is particularly important to develop a method for detecting and risk assessing municipal pipe network overflows. Summary of the Invention
[0005] The purpose of the present invention is to make up for the deficiencies of the prior art, and provide a method for detecting and risk assessing municipal pipe network overflows, which can accurately predict pipe network overflows and risk assess through comprehensively collecting and processing data of the municipal pipe network and its surrounding environment, and using machine learning and data analysis technologies.
[0006] In order to solve the above technical problems, the present invention provides the following technical solution: A method for detecting and risk assessing municipal pipe network overflows, and the specific steps of the method are as follows:
[0007] S1. Data collection and preprocessing
[0008] Collect the water level data H of the municipal pipe network at different times t , the flow rate data Q t , and the water quality data W t, as well as the corresponding timestamp t, geographical location coordinates (x, y) monitoring data. Meanwhile, collect the topographic and geomorphic data T, soil type data S, vegetation cover data V of the surrounding area, and meteorological data, including daily weather conditions, including sunny days, rainy days, rainfall amount, rainfall duration information, record the weather conditions corresponding to each monitoring data, and clarify whether the monitoring data is for sunny days or rainy days, where t = 1, 2, …, n, and n is the total number of monitoring data. In addition, collect information on the distribution of the surrounding water systems, water flow directions, water system types, and the connection conditions with the municipal pipe networks, clean the collected pipe network monitoring data, remove outliers and duplicate values, and use the moving average method to process the water level data H t for denoising, and the formula is where m is the size of the moving window. For the flow rate data Q t , use wavelet transform for denoising. Let the wavelet function be ψ(t) and the decomposition level be j, then the denoised flow rate data is where α j,k is the wavelet coefficient. Normalize the water quality data W t , and the formula is min(W) and max(W) are the minimum and maximum values of the water quality data respectively;
[0009] Associate the meteorological data with the original pipe network monitoring data, add fields such as weather conditions, rainfall amount, and rainfall duration in the data table, corresponding one by one with the water level, flow rate, and water quality data, mark and classify and store the data under different weather conditions, sort out and code the collected municipal pipe network distribution data and inspection well distribution data, convert the pipe network alignment information into digital path data, classify and code different pipe diameters and pipe materials. For the inspection well distribution data, associate the position coordinates with the geographical location coordinates in the pipe network monitoring data, and check whether there are missing values or outliers in the pipe network distribution data and inspection well distribution data. If so, perform corresponding processing;
[0010] Perform classification and coding processing on the environment-related data so that it can be integrated with the pipe network monitoring data. For the processing of the topographic and geomorphic data T, use the digital elevation model DEM for analysis, convert the collected topographic survey data into DEM data, and through interpolation and grid processing of the DEM data, obtain the topographic data of regular grids, and use this grid data to calculate the slope S lope and aspect A z , and the formulas are respectively:
[0011] where z is the topographic elevation, and x and y are the plane coordinates;
[0012] S2. Data augmentation based on the generative adversarial network (GAN)
[0013] Construct a generative adversarial network (GAN) model, which consists of a generator G and a discriminator D. The input of the generator G is a random noise vector z with a dimension of d z , which is mapped through a multi-layer neural network into virtual data with the same dimension as the real pipeline network monitoring data Let the weight matrix of the l-th layer of the generator G be W G , and the bias vector be The activation function is Then the output of the generator has the following calculation process: where L is the number of layers of the generator;
[0014] The input of the discriminator D is the real pipeline network monitoring data X or the virtual data generated by the generator The output is a scalar value representing the probability that the input data is real data. Let the weight matrix of the m-th layer of the discriminator D be The bias vector is The activation function is Then the calculation process of the output y of the discriminator is: or ... where M is the number of layers of the discriminator. Define the loss function of the generative adversarial network. For the generator G, its loss function L G is L G =-E z [log(D(G(z)))] indicates that the generator hopes that the generated virtual data can deceive the discriminator as much as possible. For the discriminator D, its loss function L D is L D =-E X [log(D(X))]-E z [log(1 - D(G(z)))] indicates that the discriminator hopes to accurately distinguish real data and virtual data. By alternately training the generator and the discriminator and continuously adjusting their weights, the generator can generate high-quality virtual pipeline operation data. During the training process, the random gradient descent method (SGD) is used to update the weights, and the learning rate η is dynamically adjusted according to the training situation. The adjustment strategy is that every k iterations, if the decrease in the loss function is less than the threshold ∈, then the learning rate is multiplied by the decay factor γ;
[0015] S3. Training of the overflow prediction model
[0016] Build an overflow prediction model based on a deep neural network. The model consists of an input layer, multiple hidden layers, and an output layer. Meteorological data is added as new features to the input layer of the overflow prediction model. The meteorological weather includes weather conditions, rainfall, and rainfall duration. The input layer receives the augmented pipeline network monitoring data, including preprocessed water level, flow rate, water quality data, and meteorological data. Let the input vector be X in , whose dimension is d in . The rectified linear unit is used as the activation function for the hidden layer. Let the weight matrix of the i-th hidden layer be and the bias vector be . Then the output h i of the i-th hidden layer is where h 0 = X in . The linear activation function is used for the output layer, and the predicted overflow probability P is output. Let the weight matrix of the output layer be W o , and the bias vector be b o . Then P = W o h n + b o , where n is the number of hidden layers;
[0017] The model is trained separately for the data subsets of sunny days and rainy days, and different model parameter settings are tried. When evaluating the model performance, the accuracy, recall rate, and F1 value metrics of the data on sunny days and rainy days in the validation set are calculated respectively, and the model structure parameters with the optimal comprehensive performance under different weather conditions are selected. Define the loss function of the overflow prediction model, and use the cross-entropy loss function L p . The formula is: where y i is the true overflow label, and P i is the predicted overflow probability. The model is trained using the adaptive moment estimation optimizer to adjust the weights of the model to minimize the loss function;
[0018] S4. Quantification of ecological and environmental impact factors
[0019] For the risk of groundwater pollution, according to the distance d gw between the pipeline leakage location and the groundwater level, the concentration C of the leaked substance, and the soil permeability coefficient k, a groundwater pollution risk index R gw is constructed. The formula is: where δ is a very small positive number. Considering the influence of the surrounding water system, if the pipeline leakage point is close to the water system, the weight of the leaked substance concentration C can be increased, and C 调整 = C × θ 水系, θ 水系> 1, and the specific value can be determined according to factors such as the distance between the water system and the pipe network and the importance of the water system. Regarding the impact on soil structure and fertility, consider the impact of the leaked substance on the soil porosity φ, pH value, and nutrient content N, and construct a soil impact index R s , and the formula is: R s = ω 1 (φ - φ 0 ) + ω 2 (pH - pH 0 ) + ω 3 (N - N 0 ), where ω 1 , ω 2 , ω 3 are weights, φ 0 , pH 0 , N 0 are the initial porosity, pH value, and nutrient content of the soil respectively. Regarding the inhibitory effect on the growth of surrounding vegetation, according to the polluted area A v of the vegetation, the vegetation type S v and the vegetation growth cycle T v , construct a vegetation impact index R v , and the formula is: where A t otal is the total area of the surrounding vegetation, λ 1 is the coefficient related to the vegetation type, and λ 2 is the coefficient related to the vegetation growth cycle;
[0020] Regarding the damage to the water ecosystem, consider the changes in the biodiversity index B, dissolved oxygen content DO, and chemical oxygen demand COD in the water, and construct a water ecological impact index R w , and the formula is: R w = ω 4 (B - B 0 ) + ω 5 (DO - DO 0 ) + ω 6 (COD - COD 0 ), where ω 4 , ω 5 , ω 6 are weights, B 0 , DO 0 , COD 0 are the initial biodiversity index, dissolved oxygen content, and chemical oxygen demand of the water respectively. At the same time, increase the monitoring and analysis of surface water quality indicators, and consider the impact of the flow rate and velocity of the water system on the water ecosystem;
[0021] S5. Risk assessment considering ecological environment impact
[0022] Determine the weights ω of each ecological environment impact factor gw 、ω s 、ω v 、ω w 。These weights are determined by using the analytic hierarchy process combined with the entropy weight method. First, construct a judgment matrix. Experts score the relative importance of different ecological environment impact factors to obtain the judgment matrix A. Then, calculate the eigenvector and the maximum eigenvalue of the judgment matrix, and normalize the eigenvector to obtain the subjective weight W sub 。At the same time, use the entropy weight method to calculate the objective weight W obj of each factor. The final weight ω is ω = αW sub +(1 - α)W obj ,where α is the balance coefficient. Construct a risk assessment model that comprehensively considers environmental ecological factors, fuse the quantified ecological environment impact index with traditional risk assessment indicators, and calculate the risk degree R of pipe network overflow. The formula is: where is the weight of traditional risk assessment indicators, and T i is the value of traditional risk assessment indicators. Determine the corresponding risk level according to the calculated risk degree R
[0023] Furthermore, in the data collection and preprocessing step, for the processing of topographic data T, use a digital elevation model DEM for analysis. Convert the collected topographic survey data into DEM data, and through interpolation and grid processing of the DEM data, obtain regular grid topographic data. Use this grid data to calculate the slope S lope and aspect A z of the terrain. The formulas are respectively where z is the terrain elevation, and x and y are planar coordinates. These topographic parameters are used for subsequent pipe network overflow analysis and risk assessment to consider the influence of terrain on the water flow direction and drainage capacity. Through the detailed processing and analysis of topographic data, the topographic conditions around the pipe network can be understood more accurately, providing richer and more accurate basic data for subsequent steps
[0024] Even further, in the data augmentation step based on the generative adversarial network GAN, the random noise vector z of the generator G is generated in a Gaussian distribution randomly. Let the mean of z be μ and the standard deviation be σ, then each element z i of z follows a Gaussian distribution N(μ, σ 2) Generally, μ = 0 and σ = 1. The randomly generated noise vectors can provide rich variations for the generator, making the generated virtual pipe network operation data diverse. Meanwhile, during the training process of the generator, to avoid the problems of gradient vanishing or gradient explosion, batch normalization technology is adopted. After the output of each layer of the neural network in the generator, the data is batch-normalized. The formula is as follows: where x i is the input data, μ B and are the mean and variance of a batch of data respectively, ∈ is a small constant to prevent the denominator from being 0, and γ and β are learnable parameters. The batch normalization technology can accelerate the training process of the generator, improve the stability and convergence speed of the model.
[0025] Furthermore, in the training step of the overflow prediction model, the number of hidden layers n of the deep neural network is determined by the method of cross-validation. The expanded dataset is divided into k non-overlapping subsets. Each time, one of the subsets is selected as the validation set, and the remaining k - 1 subsets are used as the training set. For different numbers of hidden layers n, the model is trained on the training set and the performance of the model is evaluated on the validation set. The accuracy, recall rate, and F1 value metrics are used for evaluation. The number of hidden layers n with the optimal performance on the validation set is selected as the final model structure parameter. Determining the number of hidden layers by the method of cross-validation can avoid overfitting or underfitting of the model and improve the generalization ability of the model. Meanwhile, during the model training process, to prevent overfitting, dropout technology is adopted. After the output of the neurons in each hidden layer, the output of the neurons is randomly set to 0 with a certain probability p. In this way, during the training process, each neuron has a certain probability not to participate in the calculation, thereby reducing the co-adaptation between neurons and improving the generalization ability of the model.
[0026] Furthermore, in the step of quantifying the ecological environment impact factors, for the soil impact index R s when calculating, the determination process of the weights ω 1 、ω 2 、ω 3 is as follows: First, a judgment matrix A s is constructed by the analytic hierarchy process. Experts in the field of soil science are invited to score the relative importance of soil porosity, pH value, and nutrient content on soil structure and fertility to obtain this matrix. Calculate the maximum eigenvalue λ s of the judgment matrix A max and the corresponding eigenvector W s . For the eigenvector W s . For the eigenvector W s , perform normalization processing to obtain the subjective weight Then use the entropy weight method to calculate the objective weight The specific steps are as follows: Calculate the information entropy E of each factor i , and the formula is: where p ij is the proportion of the i-th factor in the j-th sample, and calculate the entropy weight of each factor The formula is
[0027] On this basis, finally determine the weights ω 1 , ω 2 , ω 3 in the following way: where α is the balance coefficient, and its value range is between (0, 1). Through multiple experiments, compare the fitting degree between the risk assessment results and the actual situation under different α values to determine its specific value. Usually, α = 0.5 is taken.
[0029] Furthermore, in the risk assessment step considering the impact of the ecological environment, the determination process of the weights of traditional risk assessment indicators is as follows: First, classify the traditional risk assessment indicators, which can be divided into pipeline network own attribute indicators, meteorological related indicators, and surrounding geographical environment indicators. For each type of indicator, use the Delphi method to determine the weights respectively. The specific operation is as follows: Invite experts from multiple fields such as municipal engineering, meteorology, and geography to form an expert group, and distribute questionnaires to the experts. The questionnaires require the experts to score the relative importance of each factor in the same type of indicator according to their experience and professional knowledge. The scoring range is set to 1 - 9, and the larger the number, the more important the factor is. After collecting the scoring results of the experts, conduct statistical analysis on the scores of each factor, and calculate the average value and the standard deviation σ i , and the formulas are respectively: where S ij represents the score of the j-th expert for the i-th factor, and N is the number of experts;
[0030] Eliminate the outliers whose scores deviate too much from the average value, that is, eliminate the scores that satisfy , recalculate the average value of the remaining scores as the final score of this factor, sort the factors in the same type of indicator according to the final scores, and then determine their relative weights in this type of indicator. For example, for the pipeline network own attribute indicator type, if the final scores of the pipe diameter D, pipe material type M, and pipe age Y are S D , S M , S Y , then their relative weights in this type of indicator are respectively Finally, the analytic hierarchy process is used to determine the weight relationship between different categories of indicators, construct the judgment matrix of different categories of indicators, invite experts to score the relative importance of different categories of indicators, and obtain the judgment matrix A t , calculate the judgment matrix A t 's maximum eigenvalue and eigenvector, and normalize the eigenvector to obtain the weights ω c1 、ω c2 、ω c3 , and comprehensively obtain the weights of all traditional risk assessment indicators Through this way of determining weights by combining multiple methods, the opinions of experts in different fields and the mutual relationships between various indicators are fully considered, making the weight allocation of traditional risk assessment indicators more reasonable and accurate, and improving the reliability and scientific nature of the entire risk assessment model.
[0031] Furthermore, when calculating the risk degree R of pipe network overflow in the risk assessment model considering environmental and ecological factors, the fusion method of the ecological environment impact index and traditional risk assessment indicators is as follows: before fusing the quantified ecological environment impact index with traditional risk assessment indicators, first normalize all the data participating in the fusion to map the data uniformly to the interval [0, 1]. For the ecological environment impact index R gw 、R s 、R v 、R w , the following formulas are used for normalization respectively:
[0032]
[0033] For the traditional risk assessment indicator T i , the same normalization formula is also used: Then, according to the formula: Calculate the risk degree R. Through this way of fusing after normalization, it can effectively avoid the problem that a certain category of indicators dominates in the risk assessment due to the too large difference in the data magnitude of different indicators, enabling each indicator to play a role based on the same scale in the risk assessment, ensuring that the calculation result of the risk degree can more accurately reflect the actual overflow risk situation, and enhancing the stability and reliability of the risk assessment model.
[0034] Furthermore, during the process of determining the corresponding risk levels, the clustering analysis method is used to classify the calculated risk degree R. Based on historical data and practical experience, a certain number of representative sample points are selected. These sample points cover different pipe network operation conditions, ecological environment conditions, and known overflow risk situations, and the corresponding risk degree R values form a sample data set {R 1 ,R 2,…,R m}, select a clustering algorithm, preset the number of clustering categories k, and the value of k can be adjusted according to the actual situation and experience. Generally, the value range is 3 - 5, corresponding to different levels such as low risk, medium risk, and high risk. In the initial stage of the K - means clustering algorithm, randomly select k points as the initial clustering centers C = {C 1 , C 2 ,…, C k}. For each sample point R j in the sample dataset, calculate its distance to each clustering center C i (i = 1, 2,…, k) using the Euclidean distance formula where d is the dimension of the data. In risk assessment, the dimension of the data here is 1, that is, only considering the dimension of the risk level R. According to the nearest distance principle, assign each sample point to the cluster corresponding to the nearest clustering center. After all sample points are assigned, recalculate the center position of each cluster. The new clustering center C′ i is the mean of all sample points in the cluster, that is where n i is the number of sample points in the i - th cluster. Continuously repeat the above process of sample point assignment and clustering center update until the change amount of the clustering center is less than a preset threshold ∈. At this time, the clustering process ends. According to the final clustering result, divide the risk level R range corresponding to different clusters into different risk levels. For example, if k = 3, then the risk level R range can be divided into R ∈ [0, R low as the low - risk level, R ∈ (R l ow, R medium as the medium - risk level, R ∈ (R medium , 1] as the high - risk level, where R low and R medium are the division thresholds determined according to the clustering result. Through this method of determining risk levels based on cluster analysis, it is possible to make full use of the information of historical data and actual situations, objectively and reasonably divide risk levels, provide more targeted basis for the management and decision - making of municipal pipe networks, and help take more effective measures to deal with different levels of overflow risks.
[0035] Compared with the prior art, the method for detecting and risk - assessing overflows in municipal pipe networks has the following beneficial effects:
[0036] 1. This method comprehensively collects and processes data on water levels, flows, and water quality in municipal pipe networks, as well as data related to topography, soil types, and vegetation cover in the surrounding environment. This method can more comprehensively reflect the actual operating conditions of the pipe network and the surrounding environmental conditions. At the same time, data augmentation is carried out using a generative adversarial network (GAN), effectively expanding the dataset and improving the generalization ability of the model. In addition, a deep neural network is combined to construct an overflow prediction model, and cross-validation and dropout techniques are used to optimize the model structure, further improving the accuracy of the prediction. The quantification and consideration of ecological environment impact factors also make the risk assessment more comprehensive and can better reflect the impact of pipe network overflows on the surrounding environment.
[0037] 2. This method quantifies the ecological environment impact index and integrates it with traditional risk assessment indicators to calculate the risk level of pipe network overflows, and determines the corresponding risk levels based on the results of cluster analysis. This provides an intuitive and scientific basis for the management of municipal pipe networks. Management departments can take different management and maintenance measures for pipe network areas with different risk levels according to the risk assessment results, effectively preventing and controlling the occurrence of pipe network overflow events. At the same time, this method also provides a useful reference for the planning and design of municipal pipe networks, helping to reduce the risk of pipe network overflows from the source.
[0038] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on an examination of the following text, or can be learned from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0040] Figure 1 It is a flow operation diagram for the investigation and risk assessment method of municipal pipe network overflows. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention objective, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail the specific embodiments, structures, features, and their effects according to the present invention.
[0042] Embodiment 1
[0043] This embodiment describes that there are some old municipal pipe networks in a small town, and overflow problems often occur. The local relevant department decides to adopt this inspection and risk assessment method.
[0044] Collect the water level data H of the pipe network in this area at different times t , flow rate data Q t , water quality data W t and the corresponding timestamps, geographical location coordinate monitoring data. At the same time, collect the surrounding topographic and geomorphic data T, soil type data S, and vegetation coverage data V. For example, obtain the pipe network monitoring data every hour within a week, a total of n = 168 pieces. Denoise the water level data using the moving average method, with the moving window size m = 5, and calculate according to the formula Perform calculations. Denoise the flow rate data using wavelet transform, with the decomposition level j = 3. The denoised flow rate data is Normalize the water quality data according to the formula Analyze the topographic and geomorphic data using the digital elevation model DEM to obtain the slope and aspect
[0045] Build a GAN model. The random noise vector z of the generator G follows the Gaussian distribution N(0,1). The generator has 3 layers, and the calculation process of each layer is as The discriminator D also has 3 layers. By alternately training the generator and the discriminator, update the weights using the stochastic gradient descent method. The initial learning rate is 0.001. Every 100 iterations, if the decrease in the loss function is less than the threshold 0.001, then the learning rate is multiplied by the decay factor 0.9.
[0046] Build a deep neural network overflow prediction model. The input layer receives the preprocessed data. The number of hidden layers is determined to be 3 layers through 5-fold cross-validation. The hidden layers use the rectified linear unit activation function, such as h i = max(0, W i h h i-1 + b i h ), and the output layer uses the linear activation function to output the overflow probability P, P = W o h n + b o , and use the adaptive moment estimation optimizer to train the model, using the cross-entropy loss function
[0047] According to the distance d between the pipe network leakage location and the groundwater level gw , the concentration C of the leaked substance, and the soil permeability coefficient k, calculate the groundwater pollution risk index Let d gw= 5m, C = 0.5 mg / L, k = 0.01 m / d, δ = 0.001. Considering the influence of the leaked substance on the soil porosity φ, pH value, and nutrient content N, a soil impact index R is constructed. s = ω 1 (φ - φ 0 ) + ω 2 (pH - pH 0 ) + ω 3 (N - N 0 ), where the weights are determined by the analytic hierarchy process and the entropy weight method.
[0048] The analytic hierarchy process combined with the entropy weight method is used to determine the weights of ecological environment impact factors. The traditional risk assessment indicators are divided into three categories: the attributes of the pipeline network itself, meteorology, and the surrounding geographical environment. The Delphi method is used to determine the weights of each category of indicators. After normalizing the ecological environment impact index and the traditional risk assessment indicators, the risk degree is calculated by fusion. The clustering analysis method is used to classify R to determine the risk level.
[0049] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed as above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or variations equivalent to the equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent variations, and decorations made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A municipal pipe network overflow investigation and risk assessment method, characterized in that: The specific steps of this method are: S1. Data collection and preprocessing Collect water level data of municipal pipe network at different time periods t , Traffic data Q t , water quality data t , and the corresponding timestamp t, geographic location coordinates (x, y) monitoring data, and collect the topographic data T, soil type data S, vegetation coverage data V, and meteorological data of the surrounding area, including daily weather conditions, including sunny days, rainy days, rainfall, and rainfall duration information. Record the weather conditions corresponding to each monitoring data, and clearly indicate whether it is sunny or rainy day monitoring data, where t = 1, 2, ..., n, n is the total number of monitoring data. In addition, collect the distribution of surrounding water systems, water flow direction, water system type, and connectivity with the municipal pipe network. Clean the collected pipe network monitoring data, remove outliers and duplicate values, and use the sliding average method to calculate the water level data H t For denoising, the formula is Where m is the size of the sliding window, for the flow data Q t , wavelet transform is used for denoising, assuming that the wavelet function is ψ(t) and the number of decomposition layers is j, then the denoised traffic data for where α j,k is the wavelet coefficient, for water quality data W t After normalization, the formula is and max(W) are the minimum and maximum values of the water quality data, respectively; Associate meteorological data with existing pipe network monitoring data, add weather conditions, rainfall, and rainfall duration fields in the data table, correspond one-to-one with water level, flow, and water quality data, mark and classify data under different weather conditions, organize and encode the collected municipal pipe network distribution data and inspection well distribution data, convert pipe network direction information into digital path data, classify and encode different pipe diameters and pipe materials, and for inspection well distribution data, associate the location coordinates with the geographical location coordinates in the pipe network monitoring data, check whether there are missing values or abnormal values in the pipe network distribution data and inspection well distribution data, and if so, perform corresponding processing; The environmental data are classified and coded so that they can be integrated with the pipe network monitoring data. For the processing of topographic data T, the digital elevation model DEM is used for analysis. The collected topographic measurement data is converted into DEM data. The DEM data is interpolated and gridded to obtain regular grid terrain data. The grid data is used to calculate the slope S of the terrain. lope and slope A z , the formulas are: Where z is the terrain elevation, and x and y are plane coordinates; S2. Data enhancement based on generative adversarial network (GAN) Construct a generative adversarial network (GAN) model, which consists of a generator G and a discriminator D. The input of the generator G is a random noise vector z with a dimension of d. z , and maps it into virtual data of the same dimension as the real pipe network monitoring data through a multi-layer neural network. Let the weight matrix of the lth layer of generator G be W G , the bias vector is The activation function is The output of the generator is The calculation process is: Where L is the number of layers of the generator; The input of the discriminator D is the real pipe network monitoring data X or the virtual data generated by the generator The output is a scalar value, indicating the probability that the input data is real data. Suppose the weight matrix of the mth layer of the discriminator D is The bias vector is The activation function is Then the calculation process of the discriminator's output y is: or Where M is the number of layers of the discriminator, and the loss function of the generative adversarial network is defined. For the generator G, its loss function L G For L G =-E z [log(D(G(z)))] indicates that the generator hopes that the generated virtual data can deceive the discriminator as much as possible. For the discriminator D, its loss function L D For L D =-E X [log(D(X))]-E z [log(1-D(G(z)))] indicates that the discriminator hopes to accurately distinguish between real data and virtual data. By alternately training the generator and the discriminator and continuously adjusting their weights, the generator can generate high-quality virtual pipe network operation data. During the training process, the stochastic gradient descent method (SGD) is used to update the weights. The learning rate η is dynamically adjusted according to the training situation. The adjustment strategy is that after every k iterations, if the decrease in the loss function is less than the threshold ∈, the learning rate is multiplied by the attenuation factor γ; S3. Overflow prediction model training An overflow prediction model based on a deep neural network is constructed. The model consists of an input layer, multiple hidden layers, and an output layer. Meteorological data is added as a new feature to the input layer of the overflow prediction model. Meteorological weather includes weather conditions, rainfall, and rainfall duration. The input layer receives the expanded pipe network monitoring data, including pre-processed water level, flow, water quality data, and meteorological data. Let the input vector be X in , whose dimension is d in , the hidden layer uses the rectified linear unit as the activation function, and the weight matrix of the i-th hidden layer is The bias vector is Then the output h of the i-th hidden layer is i for where h0 = X in , the output layer uses a linear activation function to output the predicted overflow probability P, and the weight matrix of the output layer is set to W o , the bias vector is b o , then P=W o h n +b o , where n is the number of hidden layers; The model was trained on the data subsets of sunny days and rainy days respectively, and different model parameter settings were tried. When evaluating the model performance, the accuracy, recall rate, and F1 value indicators of the sunny and rainy day data on the validation set were calculated respectively. The model structure parameters with the best comprehensive performance under different weather conditions were selected, and the loss function of the overflow prediction model was defined. The cross entropy loss function L was used. p , the formula is: where y i is the real overflow label, P i To predict the overflow probability, the adaptive moment estimation optimizer is used to train the model and adjust the model weights to minimize the loss function. S4. Quantification of factors affecting the ecological environment For groundwater pollution risk, the distance d between the leakage location and the groundwater level is gw , the concentration of the leaked substance C and the permeability coefficient k of the soil, and the groundwater pollution risk index R is constructed. gw , the formula is: Where δ is a very small positive number. Considering the influence of the surrounding water system, if the pipe network leakage point is close to the water system, the weight of the leakage substance concentration C can be increased. Set C 调整 =C×θ 水系 ,θ 水系 >1. The specific value can be determined according to the distance between the water system and the pipe network and the importance of the water system. For the impact on soil structure and fertility, the influence of the leaked material on soil porosity φ, pH and nutrient content N is considered to construct the soil impact index R. s , the formula is: R s =ω1(φ-φ0)+ω2(pH-pH0)+ω3(N-N0), where ω1, ω2, ω3 are weights, φ0, pH0, N0 are the initial porosity, pH value and nutrient content of the soil, respectively. The inhibitory effect on the growth of surrounding vegetation is calculated based on the polluted area A of the vegetation. v , vegetation type S v and vegetation growth cycle T v , construct vegetation impact index R v , the formula is: Among them A t otal is the total area of surrounding vegetation, λ1 is the coefficient related to vegetation type, and λ2 is the coefficient related to vegetation growth cycle; For the destruction of water ecosystems, the water ecological impact index R is constructed by considering the changes in the biodiversity index B, dissolved oxygen content DO and chemical oxygen demand COD in the water. w , the formula is: R w =ω4(B-B0)+ω5(DO-DO0)+ω6(COD-COD0), where ω4, ω5, ω6 are weights, B0, DO0, COD0 are the initial biodiversity index, dissolved oxygen content and chemical oxygen demand of the water body, respectively. At the same time, the monitoring and analysis of surface water quality indicators are increased, and the impact of flow and velocity factors of the water system on the water ecology is considered; S5. Risk assessment considering ecological and environmental impacts Determine the weight of each ecological environment influencing factorω gw ,ω s ,ω v ,ω w , these weights are determined by combining the hierarchical analysis method with the entropy weight method. First, a judgment matrix is constructed. Experts score the relative importance of different ecological environment influencing factors to obtain the judgment matrix A. Then, the eigenvector and maximum eigenvalue of the judgment matrix are calculated, and the eigenvector is normalized to obtain the subjective weight W sub , and use the entropy weight method to calculate the objective weight W of each factor obj , the final weight ω is ω=αW sub +(1-α)W obj , where α is the balance coefficient, a risk assessment model that comprehensively considers environmental and ecological factors is constructed, the quantified ecological and environmental impact index is integrated with the traditional risk assessment index, and the risk level R of pipe network overflow is calculated. The formula is: in is the weight of traditional risk assessment indicators, T i is the value of the traditional risk assessment indicator. According to the calculated risk level R, the corresponding risk level is determined.
2. The municipal pipe network overflow investigation and risk assessment method according to claim 1 is characterized in that: In the data collection and pre-processing step, the terrain data T is processed by using a digital elevation model DEM for analysis, the collected terrain measurement data is converted into DEM data, and the terrain data of a regular grid is obtained by interpolating and gridding the DEM data. The slope S of the terrain is calculated using the grid data. lope and slope A z The formulas are Where z is the terrain elevation, and x and y are plane coordinates.
3. The municipal pipe network overflow investigation and risk assessment method according to claim 1 is characterized in that: In the data enhancement step based on the generative adversarial network GAN, the random noise vector z of the generator G is generated by random generation using a Gaussian distribution. Assuming the mean of z is μ and the standard deviation is σ, then each element z of z i Obey Gaussian distribution N(μ,σ 2 ), and in the training process of the generator, in order to avoid the problem of gradient disappearance or gradient explosion, the batch normalization technology is used. After each layer of the neural network of the generator is output, the data is batch normalized. The formula is: where x i is the input data, μ B and are the mean and variance of a batch of data respectively, ∈ is a small constant to prevent the denominator from being 0, γ and β are learnable parameters, and batch normalization technology can speed up the training process of the generator.
4. The municipal pipe network overflow investigation and risk assessment method according to claim 1 is characterized in that: In the overflow prediction model training step, the number of hidden layers n of the deep neural network is determined by a cross-validation method, and the expanded data set is divided into k non-overlapping subsets, one of which is selected as a validation set each time, and the remaining k-1 subsets are used as training sets. For different numbers of hidden layers n, the models are trained on the training sets, and the performance of the models is evaluated on the validation set. The accuracy, recall rate, and F1 value indicators are used for evaluation, and the number of hidden layers n with the best performance on the validation set is selected as the final model structure parameter. At the same time, in the model training process, in order to prevent overfitting, the dropout technology is used. After the neurons in each hidden layer are dropped out, the output of the neurons is randomly set to 0 with a certain probability p.
5. The municipal pipe network overflow investigation and risk assessment method according to claim 1 is characterized in that: In the step of quantifying the ecological environment impact factors, the soil impact index R s The process of determining the weights ω1, ω2, and ω3 during calculation is as follows: First, the judgment matrix A is constructed by the hierarchical analysis method. s , calculate the judgment matrix A s The maximum eigenvalue λ max and the corresponding eigenvector W s , for the eigenvector W s , for the eigenvector W s Normalize to get subjective weight Then the objective weight is calculated using the entropy weight method The specific steps are: Calculate the information entropy E of each factor i , the formula is: in p ij Calculate the entropy weight W of each factor as the proportion of the i-th factor in the j-th sample i e , the formula is On this basis, the final method to determine the weights ω1, ω2, and ω3 is: Where α is the balance coefficient.
6. The municipal pipe network overflow investigation and risk assessment method according to claim 1 is characterized in that: In the risk assessment step considering the impact of ecological environment, the weight of traditional risk assessment indicators The determination process is as follows: First, the traditional risk assessment indicators are classified into pipeline network attribute indicators, meteorological related indicators and surrounding geographical environment indicators. For each type of indicator, the Delphi method is used to determine the weight. The specific operation is as follows: Invite experts in municipal engineering, meteorology, geography and other fields to form an expert group, distribute questionnaires to the experts, collect the experts' scoring results, conduct statistical analysis on the scores of each factor, and calculate the average score of each factor. and standard deviation σ i , the formulas are: Where S ij represents the score of the jth expert on the i-th factor, and N is the number of experts; Remove outliers whose scores deviate too much from the average value, that is, remove The average of the remaining scores is recalculated as the final score of the factor. Finally, the weight relationship between different categories of indicators is determined by the hierarchical analysis method, and the judgment matrix of different categories of indicators is constructed. Experts are invited to score the relative importance of different categories of indicators to obtain the judgment matrix A. t , calculate the judgment matrix A t The maximum eigenvalue and eigenvector of , the eigenvector is normalized to obtain the weights ω of different categories of indicators c1 ,ω c2 ,ω c3 , and the weights of all traditional risk assessment indicators are obtained comprehensively 7. The municipal pipe network overflow investigation and risk assessment method according to claim 1 is characterized in that: When the risk assessment model that comprehensively considers environmental and ecological factors calculates the risk level R of pipe network overflow, the fusion method of the ecological environmental impact index and the traditional risk assessment index is as follows: before the quantified ecological environmental impact index is fused with the traditional risk assessment index, all the data involved in the fusion are normalized and uniformly mapped to the [0, 1] interval. gw , R s , R v , R w , respectively, using the following formulas for normalization: For the traditional risk assessment indicator T i , also using the normalization formula: Then according to the formula: Calculate the risk level R.
8. The municipal pipe network overflow investigation and risk assessment method according to claim 1 is characterized in that: In the process of determining the corresponding risk level, the calculated risk level R is classified by cluster analysis method. Based on historical data and actual experience, a certain number of representative sample points are selected. These sample points cover different pipe network operation conditions, ecological environment conditions and known overflow risk conditions. The corresponding risk level R values constitute the sample data set {R1, R2, …, R m }, select a clustering algorithm, pre-set the number of clustering categories k, and in the initial stage of the K-means clustering algorithm, randomly select k points as the initial cluster centers C = {C1, C2, ..., C k }, for each sample point R in the sample data set j , calculate it to each cluster center C i (i=1,2,…,k) using the Euclidean distance formula Where d is the dimension of the data. According to the principle of closest distance, each sample point is assigned to the cluster corresponding to the cluster center closest to it. After all sample points are assigned, the center position of each cluster is recalculated, and the new cluster center C′ i is the mean of all sample points in the cluster, that is, where n i is the number of sample points in the ith cluster. The above process of sample point allocation and cluster center update is repeated until the change in the cluster center is less than a preset threshold ∈. At this time, the clustering process ends. According to the final clustering result, the risk degree R range corresponding to different clusters is divided into different risk levels.
Citation Information
Cited By
PM2.5 pollution prediction method and equipment based on adaptive sample expansion
CN120354219A