A water supply pipeline leakage signal data enhancement method, system, device and medium
By combining Smote, LSTM, and 1D_CGAN, the problem of difficult data acquisition for water supply pipeline leakage signals was solved, a diverse dataset was constructed, and high-precision leakage detection was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies for detecting leaks in water supply pipelines face difficulties in acquiring leak signal data, resulting in insufficient dataset diversity. This leads to overfitting during model training, making it difficult to achieve high-precision automated detection.
The synthetic minority oversampling technique (Smote) and long short-term memory network (LSTM) are used to expand and filter the imbalanced data sample set. The leak signal dataset is generated by combining conditional generative adversarial network (1D_CGAN) to construct a diverse dataset.
It improves the quantity and quality of leakage signal data, reduces noise interference, enhances the diversity of datasets, and improves the accuracy and stability of leakage detection, providing a data foundation for automated leakage detection.
Smart Images

Figure CN120386977B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer data processing, and in particular to a method, system, device, and medium for enhancing water supply pipeline leakage signal data. Background Technology
[0002] Water resources are a fundamental and strategic resource for economic and social development; human survival and development are inseparable from water. China faces challenges such as low per capita water resources and uneven spatial and temporal distribution. In response to the ever-increasing demand for water and the severe water security situation in China, many scholars have proposed targeted solutions and strategies. However, a large portion of the population still faces the predicament of insufficient drinking water. To solve the water shortage problem, in addition to finding new freshwater resources, it is also necessary to improve the utilization rate of existing freshwater resources.
[0003] Water supply pipelines are characterized by long distances, extended travel times, and complex routes. Over time, these pipelines corrode, break, and even rupture, exacerbating water waste and pollution, and impacting both residential and industrial water use. When pipeline damage is severe, leak detection personnel rely on experience to extract signal characteristics and locate and repair the leak. However, when minor damage such as corrosion or breakage occurs, the leak signal is almost indistinguishable from a non-leak signal, making it difficult to handle manually. If such situations are not addressed promptly, minor damage can escalate into serious pipe ruptures, widespread flooding, water waste, and negative environmental impacts.
[0004] Different types of water supply pipes have different consequences when damaged. Damage to ordinary water supply pipes results in the loss of freshwater resources, while contact with sewage discharge during transport can pollute the surrounding environment. If not treated promptly, the foul odor of sewage can permeate the air and cause heavy metal pollution in water resources. In the area of water supply pipe leak detection, detection equipment has become increasingly advanced over time, gradually shifting from passive to active detection. Initially, leak detection relied mainly on human judgment, but now a series of leak detectors integrating signal processing technology and intelligent computer modules have emerged. These instruments have built-in filtering functions, thereby improving the accuracy and speed of location. The core of leak detection has now become a combination of hardware and software, with research focusing on leak detection and location methods. Based on the active / passive relationship, detection methods can be divided into active and passive detection methods.
[0005] Active detection methods often only detect leaks that have persisted for a long time or major leaks, and eliminating the leak point relies solely on manual screening, resulting in low efficiency and high resource consumption. Currently, active detection methods combine approaches from various fields with actual leak conditions, continuously innovating and improving. Examples include: direct detection methods such as acoustic detection, cable monitoring, acoustic leak detection, fiber optic sensing leak detection, and smart ball methods, as well as mass flow balance methods. Based on the type of signal detected in water supply pipelines, detection methods can be broadly categorized into those based on negative pressure wave signals and vibration signals. When a water supply pipeline leaks, due to the internal and external pressure difference, a large amount of liquid inside the pipeline sprays out from the leak location. The pressure at the leak point drops rapidly, and the density decreases. Liquid around the leak point experiences a pressure reduction after flowing out, and adjacent liquids flow towards the leak area from both upstream and downstream directions, repeating this process. The resulting negative pressure wave propagates upstream and downstream. Detection methods based on negative pressure wave signals only require sensors at both ends of the pipeline to detect this negative pressure wave, enabling a rapid response. The vibration-based detection principle uses methods such as cross-correlation, adaptive time delay estimation, and online vibration recorders to process the leakage signal.
[0006] With the rise of machine learning and continuous research by scholars, neural networks have made significant progress in image processing, speech recognition, and text analysis. Considering the superior performance of neural networks, many scholars are now combining them with technologies such as the Internet of Things (IoT), image processing, and speech recognition, gradually applying them to water leak detection. This involves leveraging the communication capacity of the IoT to achieve wireless detection of water leak signals. Due to the complex and variable nature of underground environments, their lack of visibility, the large number of detection points required, and the significant time commitment involved, traditional sensor networks' quantity and storage capacity are insufficient to meet practical needs. The IoT, with its advantages of low cost, low power consumption, short latency, large network capacity, and flexible networking, is gradually being applied to water leak detection, using technologies such as ZigBee and NB-IoT. To realize a wireless water leak monitoring and early warning system, Ma Jun designed a water supply pipeline leak triggering network scheme and simulated and verified the superior performance of the leak triggering network under two ZigBee routing protocols on OPBNET software. Based on the three-layer structure of the IoT, with different roles for the perception layer, network layer, and application layer, the combination of fiber optic distributed sensing technology and IoT technology achieves a distributed and high-precision wireless water leak detection system. The literature "Huo Daowei. Design of Main Waterway Monitoring System under Internet of Things Environment [D]. Anhui University of Science and Technology, 2020" proposes a regional water leakage monitoring system based on ZigBee IoT, which is more flexible, reliable and accurate in detection. With the development of other IoT technologies, the integration of sensor technology, NB-IoT technology, cloud platform technology and positioning algorithms has become a new development approach for water leakage monitoring systems. Water leakage detection based on IoT has the advantages of wide coverage, flexibility and high accuracy, but it generally requires hardware and network construction tailored to actual conditions and has a limited scope of application.
[0007] In recent years, technologies such as machine vision, image processing, and deep learning have begun to be applied to water leakage detection. In many complex and harsh environments, these technologies have made some breakthroughs, addressing issues such as the easy corrosion and damage of sensors in humid conditions and their instability during the detection process. The literature "Tian Youliang, Fan Tingli, Tang Chao. Rapid detection technology for surface water leakage in subway tunnels [J]. Surveying and Mapping Bulletin, 2022, (09): 29-33" proposes a lossless, convenient and fast automatic alarm system for water leakage in mine pump rooms. This system uses machine vision knowledge combined with deep learning algorithms to detect pipe leaks. In addition to the damage to sensors affecting water leakage detection, insufficient lighting for camera devices to capture images is also a major problem in underground environments. In response to the difficulty of lighting tunnels, Deng Banghong constructed a dataset by converting point cloud data to grayscale images, binarizing images, and annotating real water leaks. He then used a convolutional neural network based on masks and regions for automatic water leakage detection, which greatly improved the accuracy of detecting leaking pipes in dimly lit underground environments. Faced with the complex and ever-changing nature of underground pipelines, and the inability of stationary monitoring equipment such as cameras to fully detect leaks, Liu Zhiqiang and his team placed a miniature robot equipped with a video acquisition system and LED light sources into the water supply pipeline. The robot moved along the inside of the pipeline. Monitoring personnel located the leak point by identifying the low-frequency alarms emitted by the robot and the movement of the LED light sources. However, image processing is highly dependent on the completeness of the acquired data; this method is ineffective for data that cannot be acquired.
[0008] Acoustic leak detection, introduced from abroad in the late 1980s, differs from other leak detection methods in that it is cumbersome, experience-dependent, and costly. Over the past 20 years, it has been widely adopted by water supply companies in China. Regarding how to detect the sound signals of leaking pipes, some scholars use pressure accelerometers to collect sound signals, amplify and filter them, and estimate the power spectrum to determine if a leak has occurred. For the collected acoustic signals, the time series is visualized as input data, and a convolutional neural network is used for testing, resulting in high accuracy. Compared to image processing, acoustic leak detection does not require a large number of cameras and is suitable for certain narrow terrains. By utilizing the negative pressure wave and vibration signals of leaking water, acoustic leak detection becomes more cost-effective and has a wider range of applications.
[0009] The methods described above all rely on the premise of collecting sufficient leakage signals through physical devices such as cameras, detectors, and sensors to detect leaks. However, in cases where signals cannot be collected, insufficient training samples are available, or other levels of leakage signals in the pipe are not captured, these methods often suffer from overfitting due to insufficient or low-quality training samples, and may even fail to detect leaks in water supply pipes. Therefore, there is an urgent need for a method to address the difficulty in collecting leakage signal data from water supply pipes, and to construct a diverse dataset through augmented data collection.
[0010] Neural networks are characterized by high accuracy, low latency, and fast response in image processing and speech recognition. However, when training data is insufficient, the model is prone to overfitting. Although methods such as principal component analysis jitter, Dropout, and regularization have been proven to prevent model overfitting, these methods have little effect on improving the model's generalization ability when overfitting is caused by dataset issues. Traditional image data augmentation methods, such as rotation, cropping, scaling, and adding noise, can expand the amount of data, but they have certain shortcomings, such as limiting the diversity of the dataset and affecting the stability of model training. The literature "Liu Zhiqiang, Sun Yujing. Application of acoustic detection technology in regional water leakage detection [J]. Water Supply Technology, 2009, 3(01): 47-49." In order to solve the problem of uneven illumination in the image inside the water supply pipeline, combined with the actual situation of the water supply pipeline, and through comparison and verification, the histogram equalization algorithm is more suitable for image enhancement of water supply pipelines than the fast median filtering algorithm. However, image-based data augmentation for expanding datasets has limitations in terms of quantity, restricting not only the diversity of the datasets themselves but also affecting the stability of model training due to the added noise. As classification tasks increasingly demand diverse datasets, adversarial networks are gradually being applied to image, speech, and language processing. With the development of convolutional neural networks, combining convolutional neural networks and adversarial networks has become a new trend.
[0011] The paper "Tang Xianlun, Du Yiming, Liu Yuwei, Li Jiaxin, Ma Yiwei. Image recognition method based on conditional deep convolutional generative adversarial network [J]. Acta Automatica Sinica, 2018, 44(05): 855-864" utilizes the powerful generative capabilities of generative adversarial networks (GANs), which have shone brightly in the field of image processing. It proposes an intrusion sample enhancement method based on a combination of Conditional Generative Adversarial Networks (CGANs) and Convolutional Neural Networks (CNNs), and proposes an attention-based cropping data augmentation method. This method uses attention weights to guide image cropping, improving the network's ability to learn key features. Combining CGANs with deep convolutional generative adversarial networks, a conditional deep convolutional adversarial network (DCN) is proposed. This method not only accelerates the convergence speed and reduces the number of iterations, but also effectively improves the image classification and recognition rate. Furthermore, there are also improvements to Generative Adversarial Networks (GANs) to enhance their performance. The paper “Yang Yu. Research on Image Data Augmentation Method Based on Generative Adversarial Networks [D]. Information Engineering University of Strategic Support Force, 2022” proposes an improved StarGAN, which expands the dataset at the semantic level, solves the overfitting problem of small sample databases, and improves the model’s recognition rate and generalization ability.
[0012] However, existing methods for enhancing water leakage signal data are limited to time-series processing. While techniques like flipping, translating, and cropping followed by translation exist, similar to image processing, these methods only expand the dataset to a limited extent within the specific pipe and restrict its diversity. They cannot be used to detect other degrees of leakage signals in the same type of water supply pipe. Given the superiority of adversarial networks (ANNs) in image enhancement, combining AANs with one-dimensional convolutional neural networks to increase the quantity and diversity of water leakage signal data samples is a future development trend. Summary of the Invention
[0013] The purpose of this invention is to provide a method, system, device, and medium for enhancing water supply pipeline leakage signals, which can enhance water supply pipeline leakage signals, construct diverse datasets, and provide a data foundation for realizing automated leakage detection.
[0014] To achieve the above objectives, the present invention provides the following solution:
[0015] A method for enhancing water supply pipeline leakage signal data, the method comprising:
[0016] Acquire leakage signal samples and non-leakage signal samples from the water supply pipeline;
[0017] An unbalanced data sample set is constructed based on the leakage signal samples and the non-leakage signal samples;
[0018] The imbalanced data sample set is augmented using a synthetic minority oversampling technique to obtain an augmented data sample set;
[0019] The expanded data sample set was filtered using a long short-term memory network model to obtain a noise-reduced data sample set;
[0020] Construct a balanced data sample set based on the noise reduction data sample set;
[0021] A conditional generative adversarial network model is used to generate a leak signal dataset based on the balanced data sample set.
[0022] Optionally, the imbalanced data sample set is augmented using a synthetic minority oversampling technique to obtain an augmented data sample set, specifically including:
[0023] The leak signal samples and non-leak signal samples in the imbalanced data sample set are augmented by an equal-scale using a synthetic minority oversampling technique to obtain an augmented data sample set; wherein:
[0024] For any leak signal sample, determine several nearest neighbor samples of the leak signal sample from all leak signal samples, and synthesize several augmented data samples of the leak signal sample based on the distance between each nearest neighbor sample and the leak signal sample.
[0025] For any non-leaking signal sample, determine several nearest neighbor samples of the non-leaking signal sample from all non-leaking signal samples, and synthesize several augmented data samples of the non-leaking signal sample based on the distance between each nearest neighbor sample and the non-leaking signal sample.
[0026] Optionally, a long short-term memory network model is used to filter the expanded data sample set to obtain a denoised data sample set, specifically including:
[0027] The expanded data samples in the expanded data sample set are respectively input into the long short-term memory network model for prediction to obtain the prediction results;
[0028] The augmented data samples whose prediction results match the category labels are identified as denoised data samples, thus obtaining the denoised data sample set.
[0029] Optionally, constructing a balanced data sample set based on the noise-reduced data sample set specifically includes:
[0030] From the noise reduction data sample set, select the same number of noise reduction data samples with the category label "leaking water" and the category label "not leaking water" to construct a balanced data sample set.
[0031] Optionally, a conditional generative adversarial network model is used to generate a leakage signal dataset based on the balanced data sample set, specifically including:
[0032] A conditional generative adversarial network (GAN) model is trained based on the balanced data sample set. The conditional GAN model includes a generator and an adversary. The generator generates data based on random noise and a set category. The adversary distinguishes between the generated data and the sample data in the balanced data sample set to obtain the loss rate of the generated data and the loss rate of the sample data. When the loss rate of the generated data and the loss rate of the sample data reach Nash equilibrium, the conditional GAN model converges.
[0033] Once the conditional generative adversarial network model converges, the generator is used to generate a water leakage signal dataset.
[0034] Optionally, the water supply pipeline leakage signal data enhancement method further includes:
[0035] A one-dimensional convolutional neural network model is trained using the leak signal dataset and the non-leak signal dataset composed of the non-leak signal samples.
[0036] Optionally, the water supply pipeline leakage signal data enhancement method further includes:
[0037] Detect new leaks in the water supply pipeline;
[0038] The newly added water leakage signal is input into the one-dimensional convolutional neural network model for identification, and the identification result is obtained;
[0039] If the identification result is inconsistent with the category label, the newly added water leakage signal is input into the conditional generative adversarial network model to generate data and then added to the water leakage signal dataset.
[0040] If the identification result matches the category label, return to the steps of obtaining the new water leakage signal and the corresponding category label of the water supply pipeline.
[0041] A water supply pipeline leakage signal data enhancement system, the water supply pipeline leakage signal data enhancement system comprising:
[0042] The sample acquisition module is used to acquire leakage signal samples and non-leakage signal samples of the water supply pipeline.
[0043] The first sample set construction module is used to construct an unbalanced data sample set based on the leakage signal sample and the non-leakage signal sample.
[0044] The sample augmentation module is used to augment the imbalanced data sample set using synthetic minority oversampling technology to obtain an augmented data sample set.
[0045] The sample screening module is used to screen the expanded data sample set using a long short-term memory network model to obtain a noise-reduced data sample set;
[0046] The second sample set construction module is used to construct a balanced data sample set based on the noise reduction data sample set.
[0047] The sample augmentation module is used to generate a leak signal dataset based on the balanced data sample set using a conditional generative adversarial network model.
[0048] An electronic device includes a memory and a processor, the memory storing a computer program, and the processor running the computer program to cause the electronic device to perform the above-described method for enhancing water supply pipeline leakage signal data.
[0049] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for enhancing water supply pipeline leakage signal data.
[0050] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0051] This invention provides a data enhancement method for water supply pipeline leakage signals. It proposes a data balancing method that integrates Synthetic Minority Over-sampling Technique (Smote) and Long Short-Term Memory (LSTM) networks, as well as a data enhancement method based on One-Dimensional Conditional Adversarial Networks (1D_CGAN). Addressing the noise interference problem in actual water supply pipeline leakage signals, this invention creates an imbalanced data sample set to reduce noise interference. Smote is used to expand the imbalanced data sample set, and LSTM is used to filter the expanded data sample set to obtain a denoised data sample set. By fusing the results of Smote and LSTM, mixed sampling is achieved, and noise data is removed, thereby improving both the quantity and quality of leakage signal samples. Furthermore, this invention uses 1D_CGAN to generate a water leakage signal dataset, overcoming the problems of unstable training and difficulty in convergence when performing data augmentation with GAN. It also combines the characteristics of water leakage signals to enhance water leakage signals in water supply pipelines, constructing diverse datasets and providing a data foundation for achieving automated water leakage detection. Attached Figure Description
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 A flowchart of the water supply pipeline leakage signal data enhancement method provided by the present invention;
[0054] Figure 2 The overall flowchart of the water supply pipeline leakage signal data enhancement method provided by the present invention;
[0055] Figure 3 The schematic diagram of the Smote algorithm provided by this invention;
[0056] Figure 4 A diagram illustrating the data generation process for the 1D_CGAN model provided by this invention. Detailed Implementation
[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] Compared to manual feature extraction, neural networks can extract deeper features and learn feature distributions. Applying these learned features to classification problems results in higher detection and classification accuracy. When the dataset used to train a neural network model is large enough and of sufficient quality to meet task requirements, the model can often learn the features of the dataset well, achieving good classification results in various classification, detection, and recognition domains. However, due to the complex and variable realities of water supply pipelines, the difficulty in collecting signals, the wide variety of signals, and the inability to establish a unified standard, a serious problem of missing data exists in water supply pipeline datasets.
[0059] The purpose of this invention is to provide a method, system, device, and medium for enhancing water supply pipeline leakage signals, which can enhance water supply pipeline leakage signals, construct diverse datasets, and provide a data foundation for realizing automated leakage detection.
[0060] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0061] This invention provides a method for enhancing water supply pipeline leakage signal data, such as... Figure 1 and Figure 2 As shown, the method includes:
[0062] Step S1: Obtain leakage signal samples and non-leakage signal samples from the water supply pipeline.
[0063] The piezoelectric sensor probe can only collect a limited number of leakage and non-leakage signals at a few frequencies. First, based on the periodicity of the continuous signal, the collected signal is converted into multiple rows of 1024 sample data. Second, the 1024 leakage sample data are labeled as 1, while the 1024 non-leakage sample data are labeled as 0.
[0064] Step S2: Construct an unbalanced data sample set based on the leakage signal sample and the non-leakage signal sample.
[0065] To ensure the constructed dataset better reflects real-world needs, and considering the limited variety of water supply pipe leakage signals, it's impossible to collect all leakage signals of varying degrees for a single pipe type in each acquisition. Therefore, leakage signal samples are considered the minority class, while non-leakage signal samples are considered the majority class, creating an imbalanced dataset. For example, in an experiment with 10 rows of 1024 leakage samples labeled 1 and 40 rows of 1024 non-leakage samples labeled 0, the leakage and non-leakage samples are combined to create an imbalanced dataset.
[0066] Step S3: The imbalanced data sample set is expanded using synthetic minority oversampling technique to obtain an expanded data sample set.
[0067] Current research on imbalanced data mainly focuses on data preprocessing, feature analysis, and classification algorithms to ensure high accuracy for both the majority and minority classes. Achieving balanced data primarily involves data preprocessing techniques such as oversampling, undersampling, and mixed sampling. Undersampling improves minority class accuracy by reducing the number of majority class samples; oversampling improves minority class performance by increasing the number of minority class samples; and mixed sampling improves accuracy by increasing the number of minority class samples while decreasing the number of majority class samples. Synthetic Minority Over-sampling Technique (Smote) is a classic and widely applicable oversampling method.
[0068] Smote generates new, unique minority class samples by interpolating randomly selected nearest-neighbor samples of the same class. Its principle is as follows: Figure 3As shown, specifically: First, a sample Xi is selected from the minority class samples. Second, N samples Xzi are randomly selected from the K nearest neighbors of Xi according to the sampling ratio N. Finally, new samples are randomly synthesized between Xzi and Xi in turn, with the synthesis formula: Xn = Xi + β × (Xzi - Xi), where β is a random number. By generating synthetic samples, the Smote algorithm can increase the number of minority class samples, thereby balancing the importance of different class samples when training the model and improving the accuracy and stability of the model.
[0069] In this embodiment, a synthetic minority oversampling technique is used to augment the leak signal samples and non-leak signal samples in the imbalanced data sample set by an equal multiple, resulting in an augmented data sample set. Specifically: for any leak signal sample, several nearest neighbor samples are determined from all leak signal samples, and several augmented data samples of the leak signal sample are synthesized based on the distance between each nearest neighbor sample and the leak signal sample; for any non-leak signal sample, several nearest neighbor samples are determined from all non-leak signal samples, and several augmented data samples of the non-leak signal sample are synthesized based on the distance between each nearest neighbor sample and the non-leak signal sample.
[0070] Step S4: Use a long short-term memory network model to filter the expanded data sample set to obtain a noise-reduced data sample set.
[0071] For imbalanced data samples, the Smote algorithm only addresses the issue of insufficient sample size for different levels of leakage during sampling. However, it doesn't increase the data for other levels of pipe leakage, failing to satisfy the diversity of various leakage levels. Therefore, it's necessary to construct balanced data samples (e.g., 50 rows of various leakage signals and 50 rows of no-leakage signals) based on the limited sample size of the generated leakage signals, ensuring that the number of leakage signal samples is equal to the number of no-leakage signal samples.
[0072] Long Short-Term Memory (LSTM) networks are the most effective sequence models for processing time series data in practical applications. First, we compare the changes in accuracy and loss rate under different network structures to determine the network layers, and then determine the network parameters. We select automatic hyperparameter tuning techniques, exhaustively exploring the number of data samples passed to the program for training in a single run, the training batch size, and the optimizer within the parameter range. Finally, we output the optimal parameters to obtain the LSTM model, as shown in Table 1.
[0073] Table 1 Optimal Parameters for the LSTM Model
[0074]
[0075] As shown in Table 1, the LSTM network structure includes a TimeDistributed layer, an LSTM layer, a Dropout layer, a Dense layer, and an Activation layer, with a total of 43,846 parameters, all of which are trainable. The TimeDistributed layer provides the model with one-to-many and many-to-many capabilities, increasing the model's dimensionality. A series of layers, including a Conv1D layer (filters=32, kernel_size=4, activation='relu'), a MaxPooling1D layer (pool_size=2), and a Flatten layer, are added through the TimeDistributed layer. These layers process the input data at various time steps. Stacking LSTM layers allows for greater model complexity, the Dropout layer mitigates overfitting, and finally, a fully connected layer and activation function are added to achieve binary classification, followed by comparison to select the optimal parameters. The TimeDistributed layer is applied to a model with an input of shape (None, None, 5, 32). The length of the None parameter dimension is not fixed and can be determined based on the batch size of the input; the None parameter in other layers has the same representation.
[0076] In this embodiment, the expanded data samples in the expanded data sample set are respectively input into the long short-term memory network model for prediction to obtain prediction results; the expanded data samples whose prediction results are consistent with the category labels are determined as denoised data samples to obtain a denoised data sample set.
[0077] The training and prediction process of the LSTM model includes: using a pre-constructed balanced dataset and dividing it in a 7:3 ratio for training and testing. The initial LSTM model is trained to obtain a trained LSTM model. Then, the test data is loaded into the trained LSTM model for prediction. Predictions that match the classification labels are retained as meeting the criteria, while those that do not are deleted. This process of selecting balanced data samples is a noise reduction process, thus combining Smote and LSTM to complete the noise reduction processing of the water leakage signal data.
[0078] Step S5: Construct a balanced data sample set based on the noise reduction data sample set.
[0079] In this embodiment, an equal number of noise-reduced data samples with the category label "leaking water" and noise-reduced data samples with the category label "not leaking water" are selected from the noise-reduced data sample set to construct a balanced data sample set.
[0080] Step S6: Use a conditional generative adversarial network model to generate a water leakage signal dataset based on the balanced data sample set.
[0081] Based on GAN (Generative Adversarial Network) and DCGAN (Deep Convolutional Generative Adversarial Network), a 1D_CGAN (Conditional Generative Adversarial Network) data augmentation method for water supply pipelines is proposed. Addressing the issues of unstable training and difficulty in convergence when using GAN for data augmentation, and considering the characteristics of leakage signals, the convolutional layers in DCGAN are changed to one-dimensional convolutions, and a conditional guidance model is added to generate only leakage signals. Through extensive comparative experiments, the network structure and parameters are continuously modified to determine a network model suitable for the characteristics of water supply pipeline signals. The specific structures of the generator and adversary are shown in Tables 2 and 3 after extensive experiments. The generator consists of a 17-layer structure. The first layer is a fully connected layer. Layers 2 to 16 are stacked structures of convolutional layers, batch normalization layers, and ReLU activation function layers, stacked in 5 groups. The number of convolutional kernels in the convolutional layers are 64, 32, 64, 128, and 512 respectively, with a convolutional window of 3 for each layer. The last layer is a fully connected layer, using Tanh activation function. The adversarial mechanism consists of a 12-layer structure. Layers 1 to 6 are stacked structures of convolutional layers (32 convolutional kernels, 3 convolutional windows), normalization layers (BatchNormalization), and activation functions (LeakyReLU: 0.2), with a total of 2 stacks. Layer 7 is a one-dimensional max pooling (MaxPooling1D) layer. Layers 8 to 10 are convolutional layers (64 convolutional kernels, 3 convolutional windows), normalization layers (BatchNormalization), and activation functions (LeakyReLU: 0.2), respectively. Layer 11 is a one-dimensional global average pooling (GlobalAveragePooling1D) layer. The last layer is a dropout layer with a dropout rate of 0.2.
[0082] Table 2. Structure of the generator for 1D_CGAN
[0083]
[0084]
[0085] Table 3. Structure of the Adversarial Unit for 1D_CGAN
[0086] number of floors Countermeasure structure 1 Convolutional layer: 32 convolutional kernels, convolutional window size 3 2 BatchNormalization 3 LeakyReLU: 0.2 4 Convolutional layer: 32 convolutional kernels, convolutional window size 3 5 BatchNormalization 6 LeakyReLU: 0.2 7 MaxPooling1D 8 Convolutional layer: 64 convolutional kernels, convolutional window size 3 9 BatchNormalization 10 LeakyReLU: 0.2 11 GlobalAveragePooling1D 12 Dropout: 0.2
[0087] After denoising the sample data by fusing Smote and LSTM, the data is then trained using a generative adversarial network (GAN) to generate sample data diversity. This method does not directly copy or average the sample data. The GAN actually learns an approximate distribution of the training data; it is not simply a reproduction of the real data, but rather possesses data interpolation and extrapolation capabilities, thus achieving data augmentation. 1D represents one-dimensional convolution. Since the water leakage signal is a continuous signal in the time domain, it is represented by one-dimensional convolution. For example... Figure 4 As shown, the 1D_CGAN data generation process consists of a generator and an adversarial unit (i.e., a discriminator).
[0088] In this embodiment, a conditional generative adversarial network (GAN) model is trained based on the balanced data sample set. The conditional GAN model includes a generator and an adversary. The generator generates data based on random noise and a set category to obtain generated data. The adversary distinguishes between the generated data and the sample data in the balanced data sample set to obtain the loss rate of the generated data and the loss rate of the sample data. When the loss rate of the generated data and the loss rate of the sample data reach Nash equilibrium, the conditional GAN model converges. After the conditional GAN model converges, the generator is used to generate a water leakage signal dataset.
[0089] 1) The principle of generators and adversaries generating diverse data:
[0090] Generator: The generator is responsible for producing fake data. It takes a random noise vector as input and gradually transforms it into fake data that resembles real data through multiple neural network layers.
[0091] Discriminator: The discriminator is responsible for distinguishing between real and fake data. It receives input from the generator and real data, and through transformation operations of multiple neural network layers, gradually transforms the input into a binary classification result, i.e., real data or fake data.
[0092] The principle behind adversarial training for generating diverse data is that a generator and an adversarial processor engage in a game of cat and mouse through alternating training. The generator's goal is to deceive the adversarial processor as much as possible, making it unable to distinguish between real and fake data; the adversarial processor's goal is to distinguish between real and fake data as much as possible, so that the generator can get closer to real data. Through this adversarial training method, the generator continuously learns how to generate more realistic fake data, while the adversarial processor continuously learns how to accurately distinguish between real and fake data. Ultimately, the generator can generate high-quality fake data that is similar to real data, while the adversarial processor can accurately distinguish between real and fake data.
[0093] Generative Adversarial Networks (GANs) can generate new data with similar characteristics to the original data; these new data can be considered, to some extent, "variants" of the original data. GANs achieve this through a game between two neural networks (a generator and an adversarial network). During training, the generator attempts to generate data similar to the original data, while the adversarial network tries to distinguish between the generated and original data. As training progresses, the generator gradually learns to generate more realistic data, and diversity may emerge in the generated data. For example, it might generate leak signals of different degrees. This diversity does not mean that GAN has changed the distribution of the original data, but rather that the generator has learned to generate multiple possible data. These different data may have different characteristics in some aspects, but still belong to a data distribution similar to the original data. For example, it might generate leak signals of different degrees instead of environmental noise. Therefore, diversity is not necessarily a bad thing; on the contrary, it may help increase the richness and diversity of the generated data. The initial motivation for GANs was to solve the problem of generating highly complex and diverse data, and diversity is precisely a key characteristic of GANs in achieving this goal. Therefore, the motivation behind GANs can be considered reasonable, and the diversity generated by GANs is not a problem; rather, it can be seen as proof that GANs have successfully achieved their original purpose.
[0094] 2) Training process of 1D_CGAN conditional generative adversarial network:
[0095] The generator takes random Gaussian noise as input and combines it with a specified generated data category y to output generated data. The generator uses a fully connected (Dense) layer as input, specifying the size (input_shape) of the input noise. Upsampling layers copy the input data and change its dimensions so that the generator's output can be the input size of the adversarial generator. A batch normalization (BN) layer is added after each upsampling layer to normalize and prevent gradient vanishing or exploding, thus preventing overfitting. The Rectified Linear Activation Function (ReLU) is chosen to further prevent gradient vanishing and speed up training, as the input random noise does not contain negative numbers. Over-embedding layers convert the original one-hot encoding of the labels into a fixed-size dense vector, and then multiply the transformed label vector by the random noise vector as the generator's input. The adversarial generator, after local parameter adjustments to reduce its loss rate, consists of convolutional layers, BN layers, linear activation layers, and pooling layers. The adversarial mechanism selects either softmax or sigmoid as the activation function only on the Dense layer, because softmax and sigmoid are equivalent in binary classification problems. Furthermore, a max-pooling layer is added after two convolutional layers, and an average-pooling layer is added after the last convolutional layer.
[0096] The generator aims to produce samples of a specified category and deceive the adversary as much as possible, meaning the adversary classifies the generated data as real data. The adversary's goal is to distinguish between real and generated input data. They continuously engage in a game of cat and mouse. Eventually, the adversary becomes unable to detect the input data as fake, indicating that the generated data conforms to the characteristic distribution of real data. In other words, when the adversary's judgment score for both real and fake samples is 0.5, the GAN network has reached Nash equilibrium. The generator's loss function is represented by cross-entropy, a loss function used to measure the difference between the predicted output and the true label. The adversary's loss function is obtained by a weighted sum of the loss values for real samples and generated samples.
[0097] Furthermore, the method also includes:
[0098] Step S7: Train a one-dimensional convolutional neural network model using the leak signal dataset and the non-leak signal dataset composed of the non-leak signal samples.
[0099] Through extensive comparative experiments, the number of convolutional layers, activation functions, pooling layer sizes (average in the early stages, maximum in the later stages), batch size, and optimizer were repeatedly adjusted to obtain the network structure. The model's hyperparameters, such as the number of network layers, number of network nodes, number of iterations, and learning rate, were determined based on water supply pipeline signals. Ultimately, a 1D_CNN (one-dimensional convolutional neural network) model was selected, including 6 convolutional layers, alternating between Tanh and ReLU activation functions, a pooling layer size of 7 (maximum pooling and average pooling), a batch size of 20, and the adam optimizer. The 1D_CNN model was trained on diverse data samples of leak signals generated by 1D_CGAN and collected non-leak signal samples. The trained 1D_CNN model was used for subsequent validation of the diversity and accuracy of the constructed dataset.
[0100] Step S8: Obtain new water leakage signals from the water supply pipeline; input the new water leakage signals into the one-dimensional convolutional neural network model for identification to obtain identification results; if the identification results are inconsistent with the category labels, input the new water leakage signals into the conditional generative adversarial network model for data generation and add them to the water leakage signal dataset; if the identification results are consistent with the category labels, return to the step of obtaining new water leakage signals and corresponding category labels from the water supply pipeline.
[0101] The diversity validation of the dataset was performed by fusing the predicted and actual values of leakage signals of different levels and at the same sampling frequency and other sampling frequencies using a one-dimensional convolutional neural network. First, after data denoising and incorporating the characteristics of water supply pipeline leakage signals using the aforementioned model, the model parameters were tuned based on the loss rate and accuracy during the training of the one-dimensional convolutional conditional adversarial network model, ultimately leading to model convergence. Comparing the actual and predicted values of the one-dimensional convolutional neural network trained on the data-augmented dataset with those trained on the unaugmented dataset for other levels of leakage signals, the former showed greater data diversity. Second, diversity validation was conducted using multiple methods at the same sampling frequency, comparing data augmentation using different offset sampling algorithms and adversarial network-augmented data. Comparing the data augmented by the adversarial network with the original data revealed that the augmented data was richer and more diverse. Finally, diversity validation was performed at other sampling frequencies. Taking 3000Hz as an example, experiments were conducted starting at 1000Hz, and subsequent experiments were conducted at 1000-1500Hz, revealing some unmeasurable data. As the sampling frequency increases, more information is contained, but the situation above 3000Hz is similar to that at 3000Hz. This demonstrates the diversity of data enhanced by adversarial networks. By using adversarial networks to enhance the collected data, the enhanced data exhibits characteristics across different frequency ranges. This is because as the sampling frequency increases, the number of points collected per unit time increases, resulting in a greater variety of leakage signal characteristics. According to the technical parameters of the sensor probes, each type of probe has a limited frequency range. When a sensor with a large frequency range is selected, its sensitivity decreases. Therefore, when balancing frequency and sensitivity, it is advisable to select a sensor with a large sampling frequency range for initial data collection, provided conditions permit. Thus, this verifies that 1D_CGAN correctly generates diverse data.
[0102] In summary, this invention addresses the problem of insufficient data diversity in water supply pipeline leakage signal data, which is difficult to collect and thus fails to meet the quantity and quality requirements for model training samples, leading to overfitting and hindering automated leakage detection. It proposes a data augmentation method for water supply pipeline leakage signals based on conditional generative adversarial networks (GANs). This method aims to expand the diversity of leakage signal data, imbuing it with more features to enhance the features of uncollected leakage signals and improve leakage detection efficiency. This invention has the following advantages:
[0103] (1) To address the noise interference problem in actual water supply pipeline leakage signals, an imbalanced sample dataset is created to reduce noise interference. To improve the training and testing accuracy of the model, a data balancing method that integrates Smote and LSTM is proposed. The results of LSTM and Smote are fused to complete mixed sampling and remove noise data, thereby improving the quantity and quality of leakage signal samples.
[0104] (2) A 1D_CGAN data augmentation method for water supply pipelines is proposed based on GAN and DCGAN. Addressing the issues of unstable training and difficulty in convergence when using GAN for data augmentation, and considering the characteristics of leakage signals, the convolutional layers in DCGAN are changed to one-dimensional convolutions, and a conditional guidance model is added to generate only leakage signals. Through extensive comparative experiments, the network structure and parameters are continuously modified to determine a network model suitable for the characteristics of water supply pipeline signals. Comparison with datasets augmented by simple offsets verifies the diversity of data augmented by 1D_CGAN.
[0105] (3) To verify the accuracy and diversity of the created dataset, a 1D_CNN model matching the characteristics of water leakage signals was used as the training model. The network structure and parameters were adjusted to improve the model's accuracy and resistance to noise interference. The enhanced dataset was used for training. After the model automatically extracted and learned features, newly collected data was used as the test set. The model's recognition results on the test set were compared with the real labels. If the model's recognition results did not match the labels on the test set, it was considered that the enhanced dataset did not contain the features of that test set. The water leakage signal was then used as an input sample, and the generated data was added to the water leakage signal dataset using the adversarial network data augmentation method. This process was repeated continuously to identify and correct any omissions, ultimately resulting in a complete dataset of water supply pipeline leakage signals with varying degrees of leakage.
[0106] To implement the above methods and achieve the corresponding functions and technical effects, a water supply pipeline leakage signal data enhancement system is provided below. This system includes: a sample acquisition module for acquiring leakage signal samples and non-leakage signal samples from the water supply pipeline; a first sample set construction module for constructing an imbalanced data sample set based on the leakage signal samples and the non-leakage signal samples; a sample expansion module for expanding the imbalanced data sample set using synthetic minority oversampling technology to obtain an expanded data sample set; a sample filtering module for filtering the expanded data sample set using a long short-term memory network model to obtain a denoised data sample set; a second sample set construction module for constructing a balanced data sample set based on the denoised data sample set; and a sample enhancement module for generating a leakage signal dataset using a conditional generative adversarial network model based on the balanced data sample set.
[0107] The present invention also provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor runs the computer program to enable the electronic device to perform the above-described method for enhancing water supply pipeline leakage signal data. The electronic device may be a server.
[0108] In addition, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for enhancing water supply pipeline leakage signal data.
[0109] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0110] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for enhancing water supply pipeline leakage signal data, characterized in that, include: Acquire leakage signal samples and non-leakage signal samples from the water supply pipeline; An unbalanced data sample set is constructed based on the leakage signal samples and the non-leakage signal samples; The imbalanced data sample set is augmented using a synthetic minority oversampling technique to obtain an augmented data sample set; The expanded data sample set was filtered using a long short-term memory network model to obtain a noise-reduced data sample set; Construct a balanced data sample set based on the noise reduction data sample set; A conditional generative adversarial network (GAN) model is used to generate a water leakage signal dataset based on the balanced data sample set; the convolutional layers of the GAN model are all one-dimensional convolutions; the GAN model includes a conditional guidance model. The generator of the conditional generative adversarial network model includes a 17-layer structure. The first layer is a fully connected layer. Layers 2 to 16 are stacked structures of convolutional layers, normalization layers, and activation function layers, with a total of 5 stacks. The activation function of the activation function layer is the rectified linear function ReLU. The number of convolutional kernels in the convolutional layers are 64, 32, 64, 128, and 512, respectively, and the convolutional window is 3. The last layer is a fully connected layer, which uses Tanh as the activation function. The adversarial mechanism of the conditional generative adversarial network model comprises a 12-layer structure. Layers 1 to 6 are stacked structures of convolutional layers, normalization layers, and activation function layers, with a total of 2 stacks. The number of convolutional kernels in each convolutional layer is 32, and the number of convolutional windows is 3. Layer 7 is a one-dimensional max pooling layer. Layers 8 to 10 are convolutional layers, normalization layers, and activation function layers, respectively. The number of convolutional kernels in each convolutional layer is 64, and the number of convolutional windows is 3. Layer 11 is a one-dimensional global average pooling layer. The last layer is a Dropout layer with a dropout rate of 0.
2.
2. The method for enhancing water supply pipeline leakage signal data according to claim 1, characterized in that, The imbalanced data sample set is augmented using a synthetic minority oversampling technique to obtain an augmented data sample set, specifically including: The leak signal samples and non-leak signal samples in the imbalanced data sample set are augmented by an equal-scale using synthetic minority oversampling technology to obtain an augmented data sample set; wherein: For any leak signal sample, determine several nearest neighbor samples of the leak signal sample from all leak signal samples, and synthesize several augmented data samples of the leak signal sample based on the distance between each nearest neighbor sample and the leak signal sample. For any non-leaking signal sample, determine several nearest neighbor samples of the non-leaking signal sample from all non-leaking signal samples, and synthesize several augmented data samples of the non-leaking signal sample based on the distance between each nearest neighbor sample and the non-leaking signal sample.
3. The method for enhancing water supply pipeline leakage signal data according to claim 1, characterized in that, The augmented data sample set is filtered using a Long Short-Term Memory (LSTM) network model to obtain a denoised data sample set, which specifically includes: The expanded data samples in the expanded data sample set are respectively input into the long short-term memory network model for prediction to obtain the prediction results; The augmented data samples whose prediction results match the category labels are identified as denoised data samples, thus obtaining the denoised data sample set.
4. The method for enhancing water supply pipeline leakage signal data according to claim 1, characterized in that, Constructing a balanced data sample set based on the aforementioned noise reduction data sample set specifically includes: From the noise reduction data sample set, select the same number of noise reduction data samples with the category label "leaking water" and the category label "not leaking water" to construct a balanced data sample set.
5. The method for enhancing water supply pipeline leakage signal data according to claim 1, characterized in that, A conditional generative adversarial network model is used to generate a leakage signal dataset based on the balanced data sample set, specifically including: A conditional generative adversarial network (GAN) model is trained based on the balanced data sample set. The conditional GAN model includes a generator and an adversary. The generator generates data based on random noise and a set category. The adversary distinguishes between the generated data and the sample data in the balanced data sample set to obtain the loss rate of the generated data and the loss rate of the sample data. When the loss rate of the generated data and the loss rate of the sample data reach Nash equilibrium, the conditional GAN model converges. Once the conditional generative adversarial network model converges, the generator is used to generate a water leakage signal dataset.
6. The method for enhancing water supply pipeline leakage signal data according to claim 1, characterized in that, Also includes: A one-dimensional convolutional neural network model is trained using the leak signal dataset and the non-leak signal dataset composed of the non-leak signal samples.
7. The method for enhancing water supply pipeline leakage signal data according to claim 6, characterized in that, Also includes: Detect new leaks in the water supply pipeline; The newly added water leakage signal is input into the one-dimensional convolutional neural network model for identification, and the identification result is obtained; If the identification result is inconsistent with the category label, the newly added water leakage signal is input into the conditional generative adversarial network model to generate data and then added to the water leakage signal dataset. If the identification result matches the category label, return to the steps of obtaining the new water leakage signal and the corresponding category label of the water supply pipeline.
8. A water supply pipeline leakage signal data enhancement system, characterized in that, include: The sample acquisition module is used to acquire leakage signal samples and non-leakage signal samples of the water supply pipeline. The first sample set construction module is used to construct an unbalanced data sample set based on the leakage signal sample and the non-leakage signal sample. The sample augmentation module is used to augment the imbalanced data sample set using synthetic minority oversampling technology to obtain an augmented data sample set. The sample screening module is used to screen the expanded data sample set using a long short-term memory network model to obtain a noise-reduced data sample set; The second sample set construction module is used to construct a balanced data sample set based on the noise reduction data sample set. The sample augmentation module is used to generate a water leakage signal dataset based on the balanced data sample set using a conditional generative adversarial network (GAN) model; the convolutional layers of the GAN model are all one-dimensional convolutions; the GAN model includes a conditional guidance model. The generator of the conditional generative adversarial network model includes a 17-layer structure. The first layer is a fully connected layer. Layers 2 to 16 are stacked structures of convolutional layers, normalization layers, and activation function layers, with a total of 5 stacks. The activation function of the activation function layer is the rectified linear function ReLU. The number of convolutional kernels in the convolutional layers are 64, 32, 64, 128, and 512, respectively, and the convolutional window is 3. The last layer is a fully connected layer, which uses Tanh as the activation function. The adversarial mechanism of the conditional generative adversarial network model comprises a 12-layer structure. Layers 1 to 6 are stacked structures of convolutional layers, normalization layers, and activation function layers, with a total of 2 stacks. The number of convolutional kernels in each convolutional layer is 32, and the number of convolutional windows is 3. Layer 7 is a one-dimensional max pooling layer. Layers 8 to 10 are convolutional layers, normalization layers, and activation function layers, respectively. The number of convolutional kernels in each convolutional layer is 64, and the number of convolutional windows is 3. Layer 11 is a one-dimensional global average pooling layer. The last layer is a Dropout layer with a dropout rate of 0.
2.
9. An electronic device, characterized in that, The device includes a memory and a processor, the memory being used to store a computer program, and the processor running the computer program to cause the electronic device to perform the water supply pipeline leakage signal data enhancement method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the water supply pipeline leakage signal data enhancement method as described in any one of claims 1 to 7.