Water supply pipeline water leakage signal data enhancement method, system, equipment and medium
Through the combination of Smote, LSTM and 1D_CGAN, the imbalance of the water supply pipeline leakage signal data set is solved, and a diverse data set is generated, which improves the detection accuracy and stability of the model, and provides data support for automated water leakage detection.
Patent Information
- Application Number
- CN202311538486.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-11-17
AI Technical Summary
The existing leaky signal data sets have difficulty in data acquisition in water supply pipeline detection, resulting in insufficient model training samples and easy overfitting. The traditional data enhancement method limits the diversity of the data set and the stability of the model training.
The unbalanced data sample set is expanded and screened by synthetic minority oversampling technology (Smote) and long and short-term memory network (LSTM), and combined with the conditional generation adversarial network (1D_CGAN) to generate leakage signal data sets to build a diverse data set.
The number and quality of leaky signal samples are improved, the problem of GAN training instability is overcome, and a diverse data set is constructed, which provides a data basis for automated leak detection and enhances the generalization ability of the model.
Smart Images

Figure CN120386977A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer data processing, and in particular, to a method, system, device and medium for enhancing water leakage signal data of a water supply pipeline. Background Art
[0002] Water resources are the basic and strategic resources for economic and social development, and human survival and development are always inseparable from water. The per capita water resources in China are relatively scarce, and the temporal and spatial distribution of water resources is uneven. In the face of the growing water demand and the severe situation of water security in China, many scholars have put forward targeted solutions and strategies, but many people still face the dilemma of lacking domestic water. To solve the water shortage problem, in addition to searching for new fresh water resources, attention should also be paid to improving the utilization rate of existing fresh water resources.
[0003] The transportation distance of the water supply pipeline is long, the time is long, and the pipeline route is complex. As the pipeline is used over time, corrosion, damage or even fracture will occur, which will aggravate the waste and pollution of water resources and affect people's normal water use and industrial water use. When the pipeline is severely damaged, the leakage detection personnel extract signal features based on past experience and locate and repair the leakage point according to the features. When there are minor damages such as corrosion and damage in the water supply pipeline, the leakage signal is almost the same as the non-leakage signal, and it is difficult to handle solely by manpower. If this situation cannot be handled in time, the minor damage will expand into a serious water pipe rupture and large-scale water accumulation on the ground, resulting in waste of water resources and negative impacts on the surrounding environment.
[0004] The types of water supply pipelines are different, and the consequences of damage are also different. The damage of ordinary water supply pipes will result in the loss of fresh water resources. If they come into contact with sewage discharge during the transportation process, it will cause pollution to the surrounding environment. If not handled immediately, the stench of sewage will not only penetrate into the air but also cause problems such as heavy metal pollution to water resources. In terms of water supply pipeline leakage detection, as time goes by, the detection equipment has become more and more advanced, gradually changing from passive detection to active detection. Initially, leakage detection mainly relied on manual judgment, but now a series of leak detectors integrating signal processing technology and intelligent computer modules have emerged. These instruments have built-in filtering functions, thus improving the positioning accuracy and speed. Now the core of leakage detection has become the combination of hardware and software, and the research focus is concentrated on leakage detection and positioning methods. According to the active and passive relationship, the detection methods can be divided into active detection methods and passive detection methods.
[0005] Active detection methods often can only detect water leakage that has occurred for a long time or major water leakage problems, and the leakage points can only be excluded through manual screening, resulting in low work efficiency and consuming a lot of manpower and material resources. At present, active detection methods combine methods in various fields with the actual water leakage situation and are constantly innovated and improved. For example: direct detection methods such as acoustic listening detection method, cable monitoring method, acoustic leak detection method, fiber optic sensing leak detection method, and intelligent ball method, as well as mass flow balance method, etc. According to the type of detection signals of water supply pipelines, they can be roughly divided into detection methods based on negative pressure wave signals and vibration signals. When a water supply pipeline leaks, due to the internal and external pressure difference, a large amount of liquid in the pipeline sprays out from the leakage position, the pressure at the leakage point drops rapidly, the density decreases, the pressure of the liquid around the leakage point decreases after flowing out, and the liquid adjacent to this layer flows from the upstream and downstream directions to the leakage area and repeats this process, generating a negative pressure wave that propagates along the upstream and downstream. The detection method based on the negative pressure wave signal only needs to set sensors at both ends of the pipeline to detect this negative pressure wave and can respond quickly. The detection principle based on vibration is to use methods such as cross-correlation method, adaptive time delay estimation method, and online monitoring vibration recorder to process the water leakage signal.
[0006] With the rise of machine learning, through the continuous research of scholars, neural networks have made significant progress in image processing, speech recognition, text analysis, etc. Considering the excellent performance of neural networks, many scholars now combine neural networks with technologies such as the Internet of Things, image processing, and speech recognition, and gradually apply them to leak detection. With the communication capacity of the Internet of Things, wireless detection of leak signals is realized. Due to the complex and changeable, invisible underground environment, the large number of detection points required, and the long time spent, the number and storage capacity of traditional sensor networks are difficult to meet the actual needs. The Internet of Things, with advantages such as low cost, low power consumption, short delay, large network capacity, and flexible networking, has gradually been applied to leak detection, such as ZigBee, NB-IoT, etc. In order to implement a wireless leak monitoring and warning system, Ma Jun designed a water supply pipeline leak-triggered networking scheme and verified the excellent performance of leak-triggered networking under two ZigBee routing protocols through simulation on the OPBNET software. Based on the three-layer structure of the Internet of Things, the roles of the perception layer, network layer, and application layer are different. By combining fiber optic distributed sensing technology and Internet of Things technology, a distributed and high-precision wireless leak detection system is realized. The literature "Huo Daowei. Design of the Tap Water Main Road Guardianship System in the Internet of Things Environment [D]. Anhui University of Science and Technology, 2020." proposed a regional leak monitoring system based on ZigBee Internet of Things, which is more flexible, reliable, and has higher accuracy in detection. With the development of other Internet of Things technologies, the integration of sensor technology, NB-IoT technology, cloud platform technology, and positioning algorithms has become a new development idea for leak monitoring systems. Leak detection based on the Internet of Things has advantages such as wide coverage, flexibility, and high accuracy, but most of them need to build the hardware and networking of leak detection according to the actual situation and have a single scope of application.
[0007] In recent years, related technologies such as machine vision, image processing, and deep learning have begun to be applied to leak detection. In many complex and harsh environments, these technologies have made some breakthroughs in addressing issues such as the easy corrosion and damage of sensors in humid conditions and their instability during the detection process. The literature "Tian Youliang, Fan Tingli, Tang Chao. Rapid Detection Technology for Surface Seepage and Leakage in Subway Tunnels [J]. Bulletin of Surveying and Mapping, 2022, (09): 29-33." proposed a lossless, convenient, and fast automatic alarm system for water leakage in mine pump rooms. This system uses machine vision knowledge and combines deep learning algorithms to detect pipeline leaks. In addition to sensor damage affecting leak detection in the underground environment, insufficient lighting for camera devices to take pictures is also a major problem. For the difficult lighting situation in tunnels, Deng Banghong constructed a dataset through point cloud data grayscale image conversion, image binarization, and real leakage annotation, and used a convolutional neural network based on masks and regions for automatic seepage and leakage detection, greatly improving the accuracy of detecting leaky pipes in dim underground environments. In the face of the complex and variable underground pipelines and the problem that static monitoring devices such as cameras cannot monitor completely, Liu Zhiqiang et al. put a micro-robot equipped with a video acquisition system and an LED light source into the water supply pipeline, and the robot moved along the inside of the pipeline. Monitoring personnel located the leak points by identifying the low-frequency alarms emitted by the robot and the movement of the LED light source. However, image processing is extremely dependent on the completeness of the collected data, and this method is powerless for data that cannot be collected.
[0008] Since the end of the 1980s, acoustic leak detection methods have been introduced from abroad. Different from other leak detection methods that are cumbersome to operate, highly empirical, and costly, they have been widely promoted in domestic water supply enterprises in the past 20 years. Regarding how to detect the sound signals of leaky pipes, some scholars collect sound signals through pressure acceleration sensors, amplify and filter the signals, form the power spectrum estimation of the signals, and determine whether there is a leak. For the collected acoustic signals, after visualizing the time series as input data, a convolutional neural network is used to test it, and the accuracy is relatively high. Compared with image processing, acoustic leak detection methods do not require a large number of camera devices and are also more suitable for some narrow terrains. Acoustic leak detection methods utilize the characteristics of the negative pressure wave signals and vibration signals of leaks, making the leak detection cost lower and the applicable range wider.
[0009] The above methods all take the premise that sufficient leak signals can be collected through physical devices such as camera devices, detectors, and sensors to detect whether there is a leak. In cases where data cannot be collected, insufficient training samples are collected, or leak signals of other degrees of the pipeline are not collected, the above methods often suffer from overfitting due to insufficient and low-quality model training samples, and even cannot detect whether there is a leak in the water supply pipeline. Therefore, there is an urgent need for a method to solve the difficulty of collecting water leakage signal data for water supply pipelines and construct a diverse dataset by enhancing the data.
[0010] Neural networks have the characteristics of high accuracy, low latency, and fast response in the fields of image processing and speech recognition. However, when the training data is insufficient, the model is prone to overfitting. Although methods such as principal component analysis jitter, Dropout, and regularization have been proven to have a preventive effect on preventing model overfitting, for the overfitting caused by dataset problems, these methods have a small improvement on the generalization ability of the model. Traditional image data augmentation methods, such as rotation, cropping, scaling, and adding noise, although they can expand the number of datasets, have certain deficiencies, such as restricting the diversity of the datasets and affecting the stability of model training. The literature "Liu Zhiqiang, Sun Yujing. Application of Acoustic Wave Detection Technology in Regional Leak Detection [J]. Water Supply Technology, 2009, 3(01): 47-49." To solve the problem of uneven illumination of the internal image of the water supply pipeline, combined with the actual situation of the water supply pipeline, and through comparison, it is verified that the histogram equalization algorithm for image enhancement is more suitable for the image enhancement of the water supply pipeline than the fast median filtering algorithm. However, there are certain deficiencies in expanding the number of datasets based on image data augmentation, which not only restricts the diversity of the datasets themselves, but also adding noise will affect the stability of model training. Due to the increasing demand for dataset diversity in classification tasks, adversarial networks have gradually been applied to image, speech, and language processing. With the development of convolutional neural networks, combining convolutional neural networks and adversarial networks has become a new trend.
[0011] The literature "Tang Xianlun, Du Yiming, Liu Yuwei, Li Jiaxin, Ma Yiwei. Image Recognition Method Based on Conditional Deep Convolutional Generative Adversarial Network [J]. Acta Automatica Sinica, 2018, 44(05): 855-864." utilizes the powerful generation ability of the generative adversarial network that shines in the field of image processing, and proposes an intrusion sample enhancement method based on the combination of conditional generative adversarial network (Conditional Adversarial Nets, CGAN) and (Convolutional Neural Network, CNN). Moreover, an attention cropping data enhancement method is proposed to crop images guided by attention weights, improving the network's ability to learn key features. By combining the conditional generative adversarial network with the deep convolutional generative adversarial network, a conditional deep convolutional adversarial network is proposed. This method not only speeds up the convergence rate and reduces the number of iterations, but also effectively improves the image classification recognition rate. In addition, there are also improvements to the generative adversarial network (Generative Adversarial Network, GAN) to improve its performance. The literature "Yang Yu. Research on Image Data Enhancement Method Based on Generative Adversarial Network [D]. Information Engineering University of Strategic Support Force, 2022." proposes an improved StarGAN, which expands the dataset at the semantic level, solves the overfitting problem of small sample databases, and improves the recognition rate and generalization ability of the model.
[0012] However, the existing leakage signal data enhancement only stays at the research on the processing of time series. Although there are methods such as flipping, translation, and translation after cropping the data, similar to image processing, using these methods to expand the dataset is only an expansion to a certain extent in this pipeline and limits the diversity of the dataset itself, and cannot be used to detect leakage signals at other levels of this kind of water supply pipeline. Given the superiority of the adversarial network in image enhancement, combining the adversarial network with a one-dimensional convolutional neural network to enhance the number and diversity of leakage signal data samples is the future development trend. Summary of the Invention
[0013] The purpose of the present invention is to provide a method, system, device and medium for enhancing leakage signals of water supply pipelines, which can enhance the leakage signals of water supply pipelines, construct a diverse dataset, and provide a data basis for realizing automatic leakage detection.
[0014] To achieve the above object, the present invention provides the following solutions:
[0015] A method for enhancing leakage signals of water supply pipelines, the method for enhancing leakage signals of water supply pipelines includes:
[0016] Obtain leakage signal samples and non-leakage signal samples of the water supply pipeline;
[0017] Construct an imbalanced data sample set based on the water leakage signal samples and the non-water leakage signal samples;
[0018] Use the Synthetic Minority Over-sampling Technique (SMOTE) to augment the imbalanced data sample set to obtain an augmented data sample set;
[0019] Use a Long Short-Term Memory (LSTM) network model to screen the augmented data sample set to obtain a noise-reduced data sample set;
[0020] Construct a balanced data sample set based on the noise-reduced data sample set;
[0021] Use a Conditional Generative Adversarial Network (CGAN) model to generate a water leakage signal data set based on the balanced data sample set.
[0022] Optionally, using the Synthetic Minority Over-sampling Technique (SMOTE) to augment the imbalanced data sample set to obtain an augmented data sample set specifically includes:
[0023] Use the Synthetic Minority Over-sampling Technique (SMOTE) to equally magnify the water leakage signal samples and non-water leakage signal samples in the imbalanced data sample set to obtain an augmented data sample set; where:
[0024] For any water leakage signal sample, determine several neighboring samples of this water leakage signal sample from all water leakage signal samples, and synthesize several augmented data samples of this water leakage signal sample according to the distance between each neighboring sample and this water leakage signal sample respectively;
[0025] For any non-water leakage signal sample, determine several neighboring samples of this non-water leakage signal sample from all non-water leakage signal samples, and synthesize several augmented data samples of this non-water leakage signal sample according to the distance between each neighboring sample and this non-water leakage signal sample respectively.
[0026] Optionally, using a Long Short-Term Memory (LSTM) network model to screen the augmented data sample set to obtain a noise-reduced data sample set specifically includes:
[0027] Input the augmented data samples in the augmented data sample set into the Long Short-Term Memory (LSTM) network model for prediction to obtain prediction results;
[0028] Determine the augmented data samples with prediction results consistent with the class labels as noise-reduced data samples to obtain a noise-reduced data sample set.
[0029] Optionally, constructing a balanced data sample set based on the noise-reduced data sample set specifically includes:
[0030] Select the same number of noise-reduced data samples with the class label of water leakage and noise-reduced data samples with the class label of non-water leakage from the noise-reduced data sample set to construct a balanced data sample set.
[0031] Optionally, a conditional generative adversarial network model is used to generate a water leakage signal data set according to the balanced data sample set, specifically including:
[0032] Training a conditional generative adversarial network model according to the balanced data sample set; the conditional generative adversarial network model includes: a generator and a discriminator; the generator is used to generate data according to random noise and a set category to obtain generated data; the discriminator is used to discriminate between the generated data and the sample data in the balanced data sample set to obtain the loss rate of the generated data and the loss rate of the sample data; when the loss rate of the generated data and the loss rate of the sample data reach Nash equilibrium, the conditional generative adversarial network model converges;
[0033] After the conditional generative adversarial network model converges, the generator is used to generate a water leakage signal data set.
[0034] Optionally, the water supply pipeline water leakage signal data enhancement method further includes:
[0035] Training a one-dimensional convolutional neural network model using the water leakage signal data set and a non-water leakage signal data set composed of the non-water leakage signal samples.
[0036] Optionally, the water supply pipeline water leakage signal data enhancement method further includes:
[0037] Obtaining a new water leakage signal of the water supply pipeline;
[0038] Inputting the new water leakage signal into the one-dimensional convolutional neural network model for identification to obtain an identification result;
[0039] If the identification result is inconsistent with the category label, input the new water leakage signal into the conditional generative adversarial network model for data generation and then add it to the water leakage signal data set;
[0040] If the identification result is consistent with the category label, return to the steps of obtaining the new water leakage signal of the water supply pipeline and the corresponding category label.
[0041] A water supply pipeline water leakage signal data enhancement system, the water supply pipeline water leakage signal data enhancement system includes:
[0042] A sample acquisition module for acquiring water leakage signal samples and non-water leakage signal samples of the water supply pipeline;
[0043] A first sample set construction module for constructing an unbalanced data sample set according to the water leakage signal samples and the non-water leakage signal samples;
[0044] A sample augmentation module for augmenting the unbalanced data sample set using the Synthetic Minority Over-sampling Technique (SMOTE) to obtain an augmented data sample set;
[0045] A sample screening module for screening the augmented data sample set using a Long Short-Term Memory (LSTM) network model to obtain a noise-reduced data sample set;
[0046] A second sample set construction module for constructing a balanced data sample set based on the noise-reduced data sample set;
[0047] A sample enhancement module for generating a water leakage signal data set using a conditional generative adversarial network (CGAN) model based on the balanced data sample set.
[0048] An electronic device comprising a memory and a processor, the memory for storing a computer program, the processor running the computer program to cause the electronic device to execute the above-described water supply pipeline leakage signal data enhancement method.
[0049] A computer-readable storage medium storing a computer program, the computer program, when executed by a processor, implementing the above-described water supply pipeline leakage signal data enhancement method.
[0050] According to the specific embodiments provided by the present invention, the following technical effects are disclosed:
[0051] The water supply pipeline leakage signal data enhancement method provided by the present invention proposes a data balancing method that combines the Synthetic Minority Over-sampling Technique (SMOTE) and the Long Short-Term Memory (LSTM), and a water supply pipeline data enhancement method based on a one-dimensional conditional generative adversarial network (1D-CGAN). Aiming at the noise interference problem existing in the actual water supply pipeline leakage signal, an unbalanced data sample set is made to reduce noise interference. The unbalanced data sample set is augmented using SMOTE to obtain an augmented data sample set, and the augmented data sample set is screened using LSTM to obtain a noise-reduced data sample set. By fusing the results of SMOTE and LSTM, hybrid sampling is completed and noise data is removed, which can improve the quantity and quality of leakage signal samples. In addition, the present invention uses 1D-CGAN to generate a water leakage signal data set, overcomes the problems of unstable training and difficult convergence when using GAN for data enhancement, and combines the characteristics of the water leakage signal, can enhance the water supply pipeline leakage signal, construct a diverse data set, and provide a data basis for realizing automatic leakage detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0053] Figure 1 It is a flowchart of the method for enhancing the water leakage signal data of the water supply pipeline provided by the present invention;
[0054] Figure 2 It is the overall flowchart of the method for enhancing the water leakage signal data of the water supply pipeline provided by the present invention;
[0055] Figure 3 It is the schematic diagram of the Smote algorithm provided by the present invention;
[0056] Figure 4 It is the process diagram of generating data by the 1D_CGAN model provided by the present invention. Specific Embodiments
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0058] Compared with manually extracting features, neural networks can extract deeper features, learn the feature distribution, and use the learned features for classification problems, with higher detection accuracy and better classification accuracy. When the dataset for training the neural network model has a sufficient quantity and quality that can meet the task requirements, the model can often well learn the features of the dataset, thus achieving good classification effects in various fields of classification, detection, and recognition. However, due to the complex and changeable actual situation of the water supply pipeline, signals are difficult to collect, there are various types, and it is impossible to form a unified standard, resulting in a serious lack problem in the water supply pipeline dataset.
[0059] The purpose of the present invention is to provide a method, system, device, and medium for enhancing the water leakage signal data of the water supply pipeline, which can enhance the water leakage signal of the water supply pipeline, construct a diverse dataset, and provide a data basis for realizing automatic water leakage detection.
[0060] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will further elaborate on the present invention in conjunction with the accompanying drawings and specific embodiments.
[0061] The present invention provides a method for enhancing the leakage signal data of a water supply pipeline, as Figure 1 and Figure 2 shown. The method includes:
[0062] Step S1: Obtain the leakage signal samples and non-leakage signal samples of the water supply pipeline.
[0063] The piezoelectric sensor probe can only collect leakage signals and non-leakage signals of a limited number of frequencies. First, according to the periodicity of the continuous signal, the collected signals are converted into sample data of multiple rows of 1024. Secondly, the multiple rows of 1024 leakage sample data are labeled as 1, while the multiple rows of 1024 non-leakage sample data are labeled as 0.
[0064] Step S2: Construct an imbalanced data sample set according to the leakage signal samples and the non-leakage signal samples.
[0065] In order to make the constructed data set more in line with the actual requirements, since the types of water supply pipeline leakage signal data are very few, it is impossible to collect all the leakage signals of various degrees of a certain pipeline every time. Therefore, the water supply pipeline leakage signal samples are regarded as minority class samples, while the non-leakage signal samples are regarded as majority class samples, and an imbalanced data sample set is made. For example: during the experiment, there are 10 rows of 1024 leakage sample data labeled as 1 and 40 rows of 1024 non-leakage sample data labeled as 0, then the leakage sample data and the non-leakage sample data are jointly made into an imbalanced data sample.
[0066] Step S3: Use the Synthetic Minority Over-sampling Technique (Smote) to expand the imbalanced data sample set to obtain an expanded data sample set.
[0067] Currently, the research on the data imbalance problem mainly focuses on the data preprocessing level, the feature level, and the classification algorithm level to ensure that the classifier has a high classification accuracy for both the majority class and the minority class data. Achieving balanced data mainly focuses on data preprocessing, such as oversampling, undersampling, and hybrid sampling. The undersampling method improves the classification accuracy of the minority class by reducing the number of majority class samples; the oversampling method improves the classification performance of the minority class by increasing the number of minority class samples; the hybrid sampling method improves the classification accuracy by increasing the number of minority class samples in the sample set and at the same time reducing the number of majority class samples. The Synthetic Minority Over-sampling Technique (Smote) is a classic oversampling method and also one with the best applicability.
[0068] Smote generates new non-repeated minority class samples by randomly selecting samples of the same class for interpolation, and its principle is as Figure 3As shown in the figure, specifically: First, select a sample Xi from the minority class samples. Second, according to the sampling magnification N, randomly select N samples Xzi from the K nearest neighbors of Xi. Finally, randomly synthesize new samples between Xzi and Xi in turn, and the synthesis formula is: Xn = Xi + β × (Xzi - Xi), where β is a random number. By generating synthetic samples, the Smote algorithm can increase the number of minority class samples, thereby balancing the importance of different class samples during model training and improving the accuracy and stability of the model.
[0069] In this embodiment, the Synthetic Minority Over-sampling Technique (SMOTE) is used to equally magnify the leakage signal samples and non-leakage signal samples in the unbalanced data sample set to obtain an augmented data sample set. Among them: For any leakage signal sample, determine several nearest neighbor samples of this leakage signal sample from all leakage signal samples, and respectively synthesize several augmented data samples of this leakage signal sample according to the distance between each nearest neighbor sample and this leakage signal sample; For any non-leakage signal sample, determine several nearest neighbor samples of this non-leakage signal sample from all non-leakage signal samples, and respectively synthesize several augmented data samples of this non-leakage signal sample according to the distance between each nearest neighbor sample and this non-leakage signal sample.
[0070] Step S4: Use a Long Short-Term Memory (LSTM) network model to screen the augmented data sample set to obtain a noise-reduced data sample set.
[0071] For unbalanced data samples, using the Smote algorithm only solves the problem of the small amount of data samples of the leakage degree during sampling, but the data of other degrees of the pipeline leakage signal will not increase, that is, it cannot meet the diversity of various degrees of pipeline leakage. Therefore, it is necessary to construct balanced data samples (for example: 50 rows of various leakage signals and 50 rows of non-leakage signals) based on the leakage signals with a small amount of generated data samples, so that the number of samples of the leakage signal is equal to the number of samples of the non-leakage signal.
[0072] The Long Short-Term Memory (LSTM) network is the most effective sequence model for processing time series data in practical applications. First, compare the changes in the accuracy rate and loss rate of the model under different network structures to determine the network layer, and then determine the network parameters. Select the automatic parameter tuning technology to exhaustively search the number of data samples passed to the program for training at one time, the training batch size, and the optimizer within the parameter range during training, and finally output the best parameters to obtain the LSTM model, as shown in Table 1.
[0073] Table 1 Best Parameter Table of the LSTM Model
[0074]
[0075] As shown in Table 1, the network structure of LSTM includes a TimeDistributed layer, an LSTM layer, a Dropout layer, a Dense layer, and an Activation layer. The total number of parameters is 43,846, and all are trainable parameters. The TimeDistributed layer gives the model an ability of one-to-many and many-to-many, increasing the dimension of the model. A series of layers including a Conv1D layer (filters = 32, kernel_size = 4, activation ='relu'), a MaxPooling1D layer (pool_size = 2), and a Flatten layer are added through the TimeDistributed layer. These layers are used to process each time step of the input data. Stacking the LSTM layer allows for greater model complexity, the Dropout layer can slow down overfitting, and finally a fully connected layer and an activation function are added to achieve binary classification, and then better parameters are selected by comparison. The Time_distributed layer is applied to a model with an input of shape (None, None, 5, 32). Among them, the length of the None dimension of the parameter is not fixed and can be determined according to the batch size of the input. The same representation is used for the None parameters of other layers.
[0076] In this embodiment, the augmented data samples in the augmented data sample set are respectively input into the long short-term memory network model for prediction to obtain prediction results; the augmented data samples with prediction results consistent with the class labels are determined as noise reduction data samples to obtain a noise reduction data sample set.
[0077] The training and prediction process of the LSTM model includes: using a pre-constructed balanced data set and dividing it into 7:3, which are used for training and testing respectively. By training the initial LSTM model, a trained LSTM model is obtained. Then the test data is loaded into the trained LSTM model for prediction. The prediction results that are consistent with the classification labels are retained and considered as sample data that meet the conditions, and the prediction results that are inconsistent with the classification labels are deleted and considered as sample data that do not meet the conditions. The above process of selecting balanced data samples is a process of eliminating noise, so as to fuse Smote and LSTM to complete the noise reduction processing of the leakage signal data.
[0078] Step S5: Construct a balanced data sample set according to the noise reduction data sample set.
[0079] In this embodiment, the same number of noise reduction data samples with the class label of leakage and noise reduction data samples with the class label of non-leakage are selected from the noise reduction data sample set to construct a balanced data sample set.
[0080] Step S6: Use a conditional generative adversarial network model to generate a leakage signal data set according to the balanced data sample set.
[0081] Based on GAN (Generative Adversarial Network) and DCGAN (Deep Convolutional Generative Adversarial Network), a data augmentation method for water supply pipeline using 1D_CGAN (Conditional Generative Adversarial Network) is proposed. Aiming at the problems of unstable training and difficult convergence when using GAN for data augmentation and combining with the characteristics of leakage signals, the convolutional layer in DCGAN is changed to a one-dimensional convolution, and conditional guidance is added to the model to only generate leakage signals. Through a large number of comparative experiments, continuously changing the network structure and parameters, a network model suitable for the signal characteristics of water supply pipelines is determined. Finally, through a large number of experiments, the specific structures of the generator and discriminator are shown in Tables 2 and 3. Among them, the generator consists of 17 layers. The first layer is a fully connected layer, and the second to the sixteenth layers are a stacked structure of convolutional layer, normalization layer (BatchNormalization), and activation function (ReLU) layer, stacked in 5 groups. The number of convolutional kernels in the convolutional layer is 64, 32, 64, 128, and 512 in sequence, and the convolutional window is 3 for all. The last layer is a fully connected layer, and the activation function used is Tanh. The discriminator consists of 12 layers. The first to the sixth layers are a stacked structure of convolutional layer (32 convolutional kernels, convolutional window is 3), normalization layer (BatchNormalization), and activation function (LeakyReLU: 0.2) layer, stacked in 2 groups. The seventh layer is a one-dimensional max pooling (MaxPooling1D) layer, and the eighth to the tenth layers are convolutional layer (64 convolutional kernels, convolutional window is 3), normalization layer (BatchNormalization), and activation function (LeakyReLU: 0.2) layer in sequence. The eleventh layer is a one-dimensional global average pooling (GlobalAveragePooling1D) layer, and the last layer is a Dropout layer with a dropout rate of 0.2.
[0082] Table 2 Structure table of the generator of 1D_CGAN
[0083]
[0084]
[0085] Table 3 Structure table of the discriminator of 1D_CGAN
[0086] Number of layers Discriminator structure 1 Convolutional layer: 32 convolutional kernels, convolution window is 3 2 BatchNormalization 3 LeakyReLU: 0.2 4 Convolutional layer: 32 convolutional kernels, convolution window is 3 5 BatchNormalization 6 LeakyReLU: 0.2 7 MaxPooling1D 8 Convolutional layer: 64 convolutional kernels, convolution window is 3 9 BatchNormalization 10 LeakyReLU: 0.2 11 GlobalAveragePooling1D 12 Dropout: 0.2
[0087] After denoising the sample data by integrating Smote and LSTM, the diversity of the sample data is generated through the adversarial generation network training method, which does not directly copy or average the sample data. The generative adversarial network actually learns an approximate distribution of the training data, rather than simply reproducing the real data, but has a certain role in data interpolation and extrapolation, so the purpose of data augmentation can be achieved. 1D represents one-dimensional convolution. Since the leakage signal is a time-domain continuous signal, it is represented by one-dimensional convolution. As Figure 4 shown, the 1D_CGAN data generation process consists of a generator and an adversary (i.e., a discriminator).
[0088] In this embodiment, a conditional generative adversarial network model is trained according to the balanced data sample set training conditions; the conditional generative adversarial network model includes: a generator and an adversary; the generator is used to generate data according to a random noise and a set category to obtain generated data; the adversary is used to discriminate between the generated data and the sample data in the balanced data sample set to obtain the loss rate of the generated data and the loss rate of the sample data; when the loss rate of the generated data and the loss rate of the sample data reach the Nash equilibrium, the conditional generative adversarial network model converges; after the conditional generative adversarial network model converges, the generator is used to generate a leakage signal data set.
[0089] 1) The principle of the generator and the adversary generating diverse data:
[0090] Generator: The generator is responsible for generating fake data. It receives a random noise vector as input and gradually transforms the noise vector into fake data similar to the real data through the transformation operations of multiple neural network layers.
[0091] Adversary: The adversary is responsible for distinguishing between real data and fake data. It receives the input from the generator and the real data and gradually transforms the input into a binary classification result, that is, real data or fake data.
[0092] Principle of adversarial training for generating diverse data: The generator and the discriminator play against each other through alternating training. The goal of the generator is to deceive the discriminator as much as possible so that the discriminator cannot distinguish between fake data and real data; the goal of the discriminator is to distinguish between real data and fake data as accurately as possible in order to make the generator closer to real data. Through this adversarial training method, the generator continuously learns how to generate more realistic fake data, and the discriminator also continuously learns how to accurately distinguish between real data and fake data. Eventually, the generator can generate high-quality fake data similar to real data, and the discriminator can also accurately distinguish between real data and fake data.
[0093] Generative adversarial networks can generate new data with similar features to the original data, and these new data can be regarded as "variants" of the original data to some extent. GAN achieves this goal through the game between two neural networks (the generator and the discriminator). During the training process, the generator attempts to generate data similar to the original data, while the discriminator tries to distinguish between the generated data and the original data. As the training progresses, the generator gradually learns to generate more realistic data, and diversity may appear in the generated data. For example, leakage signals of other degrees are generated. This diversity does not mean that GAN has changed the distribution of the original data, but rather indicates that the generator has learned to generate data with multiple possibilities. These different data may have different characteristics in some aspects, but still belong to the data distribution similar to the original data. For example, what is generated is leakage signals of other degrees instead of environmental noise. Therefore, diversity is not necessarily wrong, but may instead help to increase the richness and diversity of the generated data. The initial motivation of GAN was to solve the problem of generating data with high complexity and diversity, and diversity is precisely an important feature for GAN to achieve this goal. Therefore, it can be considered that the motivation of the GAN method is reasonable, and the diversity generated by GAN is not a problem. On the contrary, it can be regarded as proof that GAN has successfully achieved its original intention.
[0094] 2) Training process of 1D_CGAN conditional generative adversarial network:
[0095] The generator uses random Gaussian noise as input, combines the specified generated data category y, and outputs the generated data. The generator takes a fully connected layer (Dense) as input and specifies the size of the input noise (input_shape); the upsampling layer duplicates the input data and changes the data dimension so that the output of the generator can be the input size of the discriminator; a batch normalization (BatchNormalization, BN) layer is added after each upsampling layer for normalization to prevent gradient vanishing or explosion and to prevent overfitting; the activation function selects the rectified linear unit (ReLU) to further prevent gradient vanishing and speed up the training because the input random noise does not contain negative numbers; the embedding layer converts the original one-hot encoding of the label into a dense vector of a fixed size, and then multiplies the converted label vector by the random noise vector as the input of the generator. After the discriminator makes local parameter adjustments to make the loss rate on the discriminator smaller, it consists of a convolutional layer, a BN layer, a linear activation layer, and a pooling layer. The discriminator only selects softmax or Sigmoid as the activation function on the Dense layer because softmax and Sigmoid are equivalent in the case of binary classification problems. And a max pooling layer is added after 2 convolutional layers, and an average pooling layer is added to the last convolutional layer.
[0096] The purpose of the generator is to generate samples of the specified category and deceive the discriminator as much as possible, that is, the discriminator classifies the generated data as real data; while the purpose of the discriminator is to distinguish as much as possible whether the input data is real data or generated data. The two continuously play against each other. Finally, when the discriminator cannot distinguish the falsity of the input data, it can be considered that the generated data at this time conforms to the characteristic distribution of real data. That is to say, when the discriminator's discrimination results for real and fake samples are both 0.5, then the training of the GAN network reaches the Nash equilibrium at this time. Among them, the loss function of the generator is represented by cross-entropy, which is a loss function used to measure the difference between the predicted output and the true label. The loss function of the discriminator can be obtained by weighted summation of the loss values of real samples and generated samples.
[0097] Furthermore, the method further includes:
[0098] Step S7: Train a one-dimensional convolutional neural network model using the leak signal data set and the non-leak signal data set composed of the non-leak signal samples.
[0099] Through a large number of comparative experiments, the number of convolutional layers, activation functions, pooling layer sizes (average pooling in the early stage and max pooling in the later stage), batch_size, and optimizers are repeatedly adjusted to obtain the network structure. The hyperparameters of the model, such as the number of network layers, the number of network nodes, the number of iterations, and the learning rate, are determined according to the water supply pipeline signals. Finally, a 1D_CNN (one-dimensional convolutional neural network) model is determined, including 6 convolutional layers, alternately using Tanh and ReLU activation functions, selecting a pooling layer size of 7 (max pooling and average pooling), a batch_size of 20, and an optimizer of adam. The 1D_CNN model is trained with the diverse data samples of the leakage signals generated by the 1D_CGAN and the non-leakage signal samples collected. The trained 1D_CNN model is used for subsequent verification of the diversity and accuracy of the constructed dataset.
[0100] Step S8: Obtain the newly added leakage signals of the water supply pipeline; input the newly added leakage signals into the one-dimensional convolutional neural network model for recognition to obtain a recognition result; if the recognition result is inconsistent with the class label, input the newly added leakage signals into the conditional generative adversarial network model for data generation and then add them to the leakage signal dataset; if the recognition result is consistent with the class label, return to the step of obtaining the newly added leakage signals of the water supply pipeline and the corresponding class labels.
[0101] The diversity verification of the dataset is carried out by fusing the predicted values and true values of the one-dimensional convolutional neural network for leakage signals at the same sampling frequency, other sampling frequencies, and other degrees of leakage. First, after the above model completes data denoising and combines the characteristics of the water supply pipeline leakage signals, the model is tuned according to the loss rate and accuracy during the training process of the one-dimensional convolutional conditional adversarial network model, and finally the model converges. Comparing the true values and predicted values of the one-dimensional convolutional neural network trained with the data-augmented dataset and the model trained with the non-data-augmented dataset for other leakage degree signals, the former data has diversity. Secondly, the diversity verification of multiple methods at the same sampling frequency is carried out by comparing the data augmentation of different offset sampling algorithms and the data augmented by the adversarial network. Comparing the data augmented by the adversarial network with the original data, it is found that the augmented data is more rich and diverse. Finally, the diversity verification is carried out at other sampling frequencies. Taking 3000Hz as an example, experiments are carried out starting from 1000Hz, and comparisons are made in the range of 1000 - 1500Hz in subsequent experiments. There are some data that cannot be measured. As the sampling frequency increases, more information is included, but the situation above 3000Hz is similar to that of 3000Hz. This shows that the data augmented by the adversarial network has diversity. The data collected is augmented in diversity by the adversarial network, and the augmented data has the characteristics of other frequency ranges. The reason is that as the sampling frequency increases, the number of points collected per unit time increases, and the characteristics of the leakage signals included increase. According to the technical parameters of the sensor probe, for each type of probe, the operating frequency range is limited. When choosing a sensor with a large operating frequency range, the sensitivity of the corresponding sensor decreases. Weighing the parameters of frequency and sensitivity, and when conditions permit, a sensor with a large collection frequency range should be selected as much as possible to collect the initial data. Therefore, it is verified that it is correct for 1D_CGAN to generate diverse data.
[0102] In summary, aiming at the problem that it is difficult to collect the water supply pipeline leakage signal data, resulting in insufficient diversity of the dataset, unable to meet the quantity and quality requirements of the model training samples, thus prone to overfitting and unable to achieve automatic leakage detection, the present invention proposes a method for enhancing water supply pipeline leakage signal data based on a conditional generative adversarial network, aiming to expand the diversity of the leakage signal data, make it have more features, so as to increase the features of the uncollected leakage signals, and improve the efficiency of leakage detection. The present invention has the following advantages:
[0103] (1) Aiming at the problem of noise interference existing in the actual water supply pipeline leakage signals, an imbalanced sample dataset is made to reduce noise interference. In order to improve the training and test accuracy of the model, a data balancing method integrating Smote and LSTM is proposed. The results of LSTM and Smote are fused to complete hybrid sampling and remove noise data, improving the quantity and quality of the leakage signal samples.
[0104] (2) Based on GAN and DCGAN, a data augmentation method for water supply pipeline data, namely 1D_CGAN, is proposed. Aiming at the problems of unstable training and difficult convergence when using GAN for data augmentation and combining with the characteristics of leakage signals, the convolutional layer in DCGAN is changed to one-dimensional convolution, and conditional guidance is added to the model to only generate leakage signals. Through a large number of comparative experiments, continuously changing the network structure and parameters, a network model suitable for the signal characteristics of water supply pipelines is determined. Compared with the dataset augmented by simple offset, the diversity of the data augmented by 1D_CGAN is verified.
[0105] (3) To verify the accuracy and diversity of the made dataset, a 1D_CNN that conforms to the characteristics of leakage signals is used as the training model, and the network structure and parameters are adjusted to improve the model's accuracy and anti-noise interference. After training with the augmented dataset, the model automatically extracts features and learns. Then, the newly collected data is used as the test set, and the recognition results of the model in the test set are compared and analyzed with the true labels. If the recognition results of the model are inconsistent with the labels of the test set, it is considered that the augmented dataset does not contain the features of this test set. The leakage signal is used as an input sample according to the method of augmenting data by the adversarial network, and the generated data is added to the leakage signal dataset. The above process is continuously repeated to check for omissions and make up for deficiencies, and finally a complete water supply pipeline leakage signal dataset with various degrees is obtained.
[0106] To implement the above method to achieve the corresponding functions and technical effects, a water supply pipeline leakage signal data augmentation system is provided below. The system includes: a sample acquisition module for acquiring leakage signal samples and non-leakage signal samples of the water supply pipeline. A first sample set construction module for constructing an imbalanced data sample set according to the leakage signal samples and the non-leakage signal samples. A sample augmentation module for augmenting the imbalanced data sample set using the synthetic minority over-sampling technique to obtain an augmented data sample set. A sample screening module for screening the augmented data sample set using a long short-term memory network model to obtain a noise-reduced data sample set. A second sample set construction module for constructing a balanced data sample set according to the noise-reduced data sample set. A sample enhancement module for generating a leakage signal dataset according to the balanced data sample set using a conditional generative adversarial network model.
[0107] The present invention also provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor is used to run the computer program so that the electronic device executes the above-mentioned water supply pipeline leakage signal data augmentation method. The electronic device can be a server.
[0108] In addition, the present invention further provides a computer-readable storage medium, which stores a computer program that, when executed by a processor, implements the above-described method for enhancing water leakage signal data of a water supply pipeline. <> <>
[0109] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section. <> <>
[0110] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for enhancing water leakage signal data of a water supply pipeline, characterized in that, Including: Obtaining leakage signal samples and non-leakage signal samples of a water supply pipeline; Constructing an imbalanced data sample set according to the leakage signal samples and the non-leakage signal samples; Using the Synthetic Minority Over-sampling Technique (SMOTE) to expand the imbalanced data sample set to obtain an expanded data sample set; Using a Long Short-Term Memory (LSTM) network model to screen the expanded data sample set to obtain a noise-reduced data sample set; Constructing a balanced data sample set according to the noise-reduced data sample set; Using a Conditional Generative Adversarial Network (CGAN) model to generate a leakage signal data set according to the balanced data sample set.
2. The method for enhancing water leakage signal data of a water supply pipeline according to claim 1, wherein, Using the Synthetic Minority Over-sampling Technique (SMOTE) to expand the imbalanced data sample set to obtain an expanded data sample set, specifically including: Using the Synthetic Minority Over-sampling Technique (SMOTE) to equally expand the leakage signal samples and non-leakage signal samples in the imbalanced data sample set to obtain an expanded data sample set; where: For any leakage signal sample, determine several neighboring samples of the leakage signal sample from all leakage signal samples, and synthesize several expanded data samples of the leakage signal sample according to the distance between each neighboring sample and the leakage signal sample respectively; For any non-leakage signal sample, determine several neighboring samples of the non-leakage signal sample from all non-leakage signal samples, and synthesize several expanded data samples of the non-leakage signal sample according to the distance between each neighboring sample and the non-leakage signal sample respectively.
3. The method for enhancing the water leakage signal data of the water supply pipeline according to claim 1, wherein, Using a Long Short-Term Memory (LSTM) network model to screen the expanded data sample set to obtain a noise-reduced data sample set, specifically including: Inputting the expanded data samples in the expanded data sample set into the Long Short-Term Memory (LSTM) network model respectively for prediction to obtain prediction results; Determining the expanded data samples with prediction results consistent with the class labels as noise-reduced data samples to obtain a noise-reduced data sample set.
4. The method for enhancing the water leakage signal data of a water supply pipeline according to claim 1, wherein, Constructing a balanced data sample set according to the noise-reduced data sample set, specifically including: Selecting the same number of noise-reduced data samples with the class label of leakage and noise-reduced data samples with the class label of non-leakage from the noise-reduced data sample set to construct a balanced data sample set.
5. The water supply pipeline leakage signal data enhancement method according to claim 1, characterized in that Using a Conditional Generative Adversarial Network (CGAN) model to generate a leakage signal data set according to the balanced data sample set, specifically including: Training a Conditional Generative Adversarial Network (CGAN) model according to the balanced data sample set; the Conditional Generative Adversarial Network (CGAN) model includes: a generator and a discriminator; the generator is used to generate data according to random noise and a set class to obtain generated data; the discriminator is used to discriminate between the generated data and the sample data in the balanced data sample set to obtain the loss rate of the generated data and the loss rate of the sample data; when the loss rate of the generated data and the loss rate of the sample data reach Nash equilibrium, the Conditional Generative Adversarial Network (CGAN) model converges; After the Conditional Generative Adversarial Network (CGAN) model converges, use the generator to generate a leakage signal data set.
6. The method for enhancing water leakage signal data of a water supply pipeline according to claim 1, wherein Also including: Training a one-dimensional convolutional neural network model using the leakage signal data set and a non-leakage signal data set composed of the non-leakage signal samples.
7. The method for enhancing the water leakage signal data of the water supply pipeline according to claim 6, wherein, Also including: Obtaining new leakage signals of the water supply pipeline. Input the newly added leakage signal into the one-dimensional convolutional neural network model for identification to obtain an identification result; If the identification result is inconsistent with the class label, input the newly added leakage signal into the conditional generative adversarial network model for data generation and then add it to the leakage signal data set; If the identification result is consistent with the class label, return to the step of obtaining the newly added leakage signal of the water supply pipeline and the corresponding class label.
8. A water supply pipeline leakage signal data enhancement system, characterized in that, It includes: A sample acquisition module for acquiring leakage signal samples and non-leakage signal samples of the water supply pipeline; A first sample set construction module for constructing an imbalanced data sample set according to the leakage signal samples and the non-leakage signal samples; A sample augmentation module for augmenting the imbalanced data sample set using the synthetic minority over-sampling technique to obtain an augmented data sample set; A sample screening module for screening the augmented data sample set using a long short-term memory network model to obtain a noise-reduced data sample set; A second sample set construction module for constructing a balanced data sample set according to the noise-reduced data sample set; A sample enhancement module for generating a leakage signal data set using a conditional generative adversarial network model according to the balanced data sample set.
9. An electronic device, characterized in that, It includes a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the water supply pipeline leakage signal data enhancement method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by a processor, it implements the water supply pipeline leakage signal data enhancement method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Radar interference multi-domain feature adversarial learning and detection identification method
CN114429156A
Offshore wind turbine generator gearbox fault diagnosis method based on vibration signals under noise background
CN116340859A
Synthetic data augmentation for ECG using deep learning
US20230225660A1