A method, system, medium, and device for improving leak detection performance.
By employing a semi-supervised approach of pre-training, knowledge transfer, and fine-tuning, a leak detection and identification model is constructed using unlabeled monitoring data. This approach solves the problem of high cost of labeled data and achieves a significant improvement in leak identification performance.
Patent Information
- Application Number
- CN202411652573.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-11-19
AI Technical Summary
In existing technologies, deep learning-based leak identification methods require a large amount of labeled data, which leads to high cost and difficulty in labeling data, making it difficult to effectively build accurate leak identification models.
A semi-supervised approach of pre-training-knowledge transfer-fine-tuning is adopted to build a leak detection and identification model using unlabeled monitoring data. The model parameters are optimized through feature extraction and backpropagation, and fine-tuned by combining labeled data to improve the leak identification performance.
It reduces model development costs, improves leak identification performance, and enables effective leak identification without requiring a large amount of labeled data.
Smart Images

Figure CN119622451B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water supply network leak detection technology, and in particular to a method, system, medium and equipment for improving leak detection performance. Background Technology
[0002] Water supply networks are an indispensable urban infrastructure, bearing the core responsibility of delivering clean and safe water resources. However, with the increasing prevalence of problems such as pipe aging, corrosion, external damage, and poor management, leakage in water supply networks has gradually become a major obstacle to the efficient utilization and sustainable management of urban water resources. To effectively address this challenge, leak identification has become a crucial step in leak mitigation. Early detection and location of leaks in water supply networks through advanced technologies can not only significantly reduce leakage rates but also provide strong support for the construction of a water-saving society, thereby better ensuring the long-term sustainable use of water resources.
[0003] Utilizing machine learning techniques to automatically detect leakage signals in monitoring data is a promising and efficient method for leak identification. In recent years, deep learning-based leak identification methods have received widespread attention. However, these methods are typically supervised, requiring large amounts of labeled data (i.e., data that explicitly indicates the presence of leaks) for model training. Obtaining labeled data necessitates on-site verification by professionals and often requires data mining and backfilling, leading to high costs in building accurate and practical supervised leak identification models. Summary of the Invention
[0004] To address the aforementioned problems, the present invention aims to provide a method, system, medium, and device for improving leakage identification performance, effectively solving the problems of high cost and difficulty in labeling data.
[0005] To achieve the above objectives, in a first aspect, the technical solution adopted by the present invention is as follows: a method for improving leak identification performance, comprising: dividing unlabeled monitoring data into homogeneous data pairs and non-homogeneous data pairs after data processing, and inputting both types of data pairs into a pre-trained model for feature extraction and pre-training; constructing a leak detection and identification model based on the features extracted from the pre-trained model, and transferring the ability to distinguish data differences learned from unlabeled data to the leak detection and identification model; fine-tuning the leak detection and identification model using labeled data, and continuously optimizing the parameters of the leak detection and identification model through backpropagation to improve leak identification performance.
[0006] Furthermore, the unlabeled monitoring data is divided into homologous data pairs and heterologous data pairs after data processing. This includes using data augmentation methods to generate two different versions for each unlabeled monitoring data time series, thereby forming homologous data pairs and heterologous data pairs.
[0007] Furthermore, both types of data pairs are input into the pre-trained model for feature extraction, including: inputting the generated homologous data pairs and non-homologous data pairs into the feature extractor in the pre-trained model to extract low-dimensional features of the time series, and then converting the low-dimensional features into feature vectors through the mapping layer of the pre-trained model to calculate the contrastive loss.
[0008] Furthermore, both data pairs are input into the pre-trained model for pre-training, including:
[0009] The performance of the pre-training process is evaluated using a contrastive loss function, and the parameters of the pre-trained model are continuously optimized through backpropagation.
[0010] The objective of the contrastive loss function is to maximize the similarity between the two data points in each pair of data from the same source, while minimizing the similarity between the two data points in each pair of data from different sources, so that the feature extractor has the ability to distinguish between the differences in the data.
[0011] Furthermore, the similarity is quantified using the cosine similarity index.
[0012] Furthermore, a leak detection and identification model is constructed based on features extracted from the pre-trained model, and the ability to distinguish data differences learned from unlabeled data is transferred to the leak detection and identification model, including:
[0013] Fine-tuning techniques are used to transfer the ability of the feature extractor to distinguish data differences learned from unlabeled data to the leak detection and identification model. In the training process of the leak detection and identification model, the pre-trained parameters are used as the initial parameters of the feature extractor, while the parameters of the classification layer are randomly initialized.
[0014] Furthermore, the leak detection and identification model is fine-tuned using labeled data, including: training the leak detection and identification model with labeled data, evaluating the fine-tuning process using the cross-entropy loss function, and continuously optimizing the parameters of the leak detection and identification model through backpropagation.
[0015] Secondly, the technical solution adopted by the present invention is: a system for improving leakage identification performance, comprising:
[0016] The pre-training module divides the unlabeled monitoring data into homogeneous data pairs and heterogeneous data pairs after data processing, and inputs both types of data pairs into the pre-training model for feature extraction and pre-training.
[0017] The knowledge transfer module constructs a leak detection and identification model based on features extracted from the pre-trained model, and transfers the ability to distinguish data differences learned from unlabeled data to the leak detection and identification model.
[0018] The model fine-tuning module uses labeled data to train the leak detection and identification model, and continuously optimizes the parameters of the leak detection and identification model through backpropagation to improve leak identification performance.
[0019] Thirdly, the technical solution adopted by the present invention is: a computer-readable storage medium for storing one or more programs, wherein the one or more programs include instructions, which, when executed by a computing device, cause the computing device to perform any of the methods described above.
[0020] Fourthly, the technical solution adopted by the present invention is: a computing device comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods described above.
[0021] The present invention has the following advantages due to the adoption of the above technical solutions:
[0022] This invention breaks through the traditional development approach of supervised leak detection models, constructing a three-stage semi-supervised leak detection model development method: pre-training, knowledge transfer, and fine-tuning. In the pre-training stage, unlabeled data is used to train a feature extractor capable of distinguishing data differences. In the knowledge transfer stage, a leak detection and identification model is built based on the feature extractor from the pre-trained model, transferring the ability to distinguish data differences learned from unlabeled data to the leak detection and identification model. In the fine-tuning stage, labeled data is used to enable the leak detection and identification model to be quickly applied to leak detection tasks in real-world pipeline networks. This effectively solves the problem of poor performance of leak detection models under conditions of high cost and difficulty in labeling data. Attached Figure Description
[0023] Figure 1 This is a flowchart of a method for improving leakage identification performance in an embodiment of the present invention. Detailed Implementation
[0024] To address the high cost and difficulty of acquiring labeled data in traditional supervised models, this invention proposes a method, system, medium, and equipment for improving leak detection performance. This method has low data requirements while still achieving competitive detection performance. In water supply networks, the scale of raw monitoring data (i.e., unlabeled data) that has not been verified by professionals is enormous and can be obtained after installing monitoring equipment. This invention utilizes this unlabeled data to build models, thus significantly reducing model development costs.
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.
[0026] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0027] In one embodiment of the present invention, a method, system, medium, and device for improving leakage identification performance are provided. In this embodiment, as... Figure 1 As shown, the method includes the following steps:
[0028] 1) Pre-training stage: After data processing, the unlabeled detection data is divided into homologous data pairs and non-homologous data pairs. Both types of data pairs are then input into the pre-training model for feature extraction and pre-training.
[0029] 2) Knowledge transfer stage: Based on the features extracted from the pre-trained model, a leak detection and identification model is constructed, and the ability to distinguish data differences learned from unlabeled data is transferred to the leak detection and identification model;
[0030] 3) Model fine-tuning stage: The leak detection and identification model is fine-tuned using labeled data, and the parameters of the leak detection and identification model are continuously optimized through backpropagation to improve the leak identification performance.
[0031] In step 1) above, the unlabeled monitoring data is divided into homologous data pairs and heterologous data pairs after data processing. Specifically, a data amplification method is used to generate two different versions for each unlabeled monitoring data time series, thus forming homologous data pairs and heterologous data pairs. Homologous data pairs consist of two amplified versions of the same time series, while heterologous data pairs consist of two amplified versions of different time series.
[0032] In this embodiment, the data augmentation method can be, for example, adding white noise to the time series or randomly adjusting the data arrangement order.
[0033] In step 1) above, both data pairs are input into the pre-trained model for feature extraction, specifically as follows:
[0034] Both the generated homologous and non-homologous data pairs are input into the feature extractor in the pre-trained model to extract low-dimensional features of the time series. The low-dimensional features are then transformed into feature vectors through the mapping layer of the pre-trained model to calculate the contrastive loss.
[0035] In this embodiment, both data pairs are input into the pre-trained model for pre-training. Specifically, the pre-training process is evaluated using a contrastive loss function, and the parameters of the pre-trained model are continuously optimized through backpropagation.
[0036] The objective of the contrastive loss function is to maximize the similarity between the two data points in each pair of data from the same source, while minimizing the similarity between the two data points in each pair of data from different sources, so that the feature extractor has the ability to distinguish between the differences in the data.
[0037] In this embodiment, similarity is quantified using the cosine similarity index.
[0038] In step 2) above, the pre-trained feature extractor is directly connected to the classification layer to build a leak detection and identification model.
[0039] In step 2) above, a leak detection and identification model is constructed based on the features extracted from the pre-trained model. The ability to distinguish data differences learned from unlabeled data is transferred to the leak detection and identification model. Specifically, fine-tuning techniques are used to transfer the ability of the feature extractor to distinguish data differences learned from unlabeled data to the leak detection and identification model. During the training of the leak detection and identification model, the pre-trained parameters serve as the initial parameters of the feature extractor, while the parameters of the classification layer are randomly initialized. This is significantly different from the random initialization of all parameters in traditional supervised learning methods.
[0040] In step 3) above, the leak detection and identification model is fine-tuned using labeled data. Specifically, the leak detection and identification model is trained using labeled data, and the fine-tuning process is evaluated using the cross-entropy loss function.
[0041] In the above embodiments, unlabeled and labeled data: the data sources are the same, and the monitoring equipment is of the same type, such as flow meters in water supply networks or noise monitors; however, the latter has been verified by professionals to determine whether it represents a leak and can be used for training supervised models.
[0042] In the above embodiments, the feature extractors in the pre-trained model and the leak detection and identification model can be neural networks of any structure, such as convolutional neural networks, recurrent neural networks, etc. The model structure can be adjusted according to the input data. If it is a time series, a recurrent neural network is preferred; if it is a time-spectrum graph, a convolutional neural network is preferred.
[0043] In the above embodiments, the mapping of the pre-trained model and the classification layer in the leakage detection and identification model are composed of a simple fully connected neural network, for example, it may contain two fully connected layers.
[0044] In one embodiment of the present invention, a system for improving leakage detection performance is provided, comprising:
[0045] The pre-training module divides the unlabeled monitoring data into homogeneous data pairs and heterogeneous data pairs after data processing, and inputs both types of data pairs into the pre-training model for feature extraction and pre-training.
[0046] The knowledge transfer module constructs a leak detection and identification model based on features extracted from the pre-trained model, and transfers the ability to distinguish data differences learned from unlabeled data to the leak detection and identification model.
[0047] The model fine-tuning module uses labeled data to fine-tune the leak detection and identification model, and continuously optimizes the parameters of the leak detection and identification model through backpropagation to improve leak identification performance.
[0048] In the above embodiments, the unlabeled monitoring data is divided into homologous data pairs and heterologous data pairs after data processing. Specifically, a data amplification method is used to generate two different versions for each unlabeled monitoring data time series, thereby forming homologous data pairs and heterologous data pairs.
[0049] In the above embodiment, both data pairs are input into the pre-trained model for feature extraction. Specifically, the generated homologous data pairs and non-homologous data pairs are input into the feature extractor in the pre-trained model to extract low-dimensional features of the time series. Then, the low-dimensional features are transformed into feature vectors through the mapping layer of the pre-trained model to calculate the contrastive loss.
[0050] In this embodiment, both data pairs are input into the pre-trained model for pre-training. Specifically, the pre-training process is evaluated using a contrastive loss function, and the parameters of the pre-trained model are continuously optimized through backpropagation. The goal of the contrastive loss function is to maximize the similarity between the two data points in each homologous data pair and minimize the similarity between the two data points in each non-homologous data pair, so that the feature extractor has the ability to distinguish data differences.
[0051] The similarity is quantified using the cosine similarity index.
[0052] In the above embodiments, a leak detection and identification model is constructed based on features extracted from the pre-trained model, and the ability to distinguish data differences learned from unlabeled data is transferred to the leak detection and identification model, including:
[0053] The ability of the feature extractor to distinguish data differences learned from unlabeled data is transferred to the leak detection and identification model using fine-tuning techniques;
[0054] In the training process of the leak detection and identification model, the pre-trained parameters are used as the initial parameters of the feature extractor, while the parameters of the classification layer are randomly initialized.
[0055] In the above embodiments, fine-tuning the leak detection and identification model using labeled data includes: training the leak detection and identification model using labeled data and evaluating the knowledge transfer process using the cross-entropy loss function.
[0056] The system provided in this embodiment is used to execute the above-described method embodiments. For specific processes and details, please refer to the above embodiments, which will not be repeated here.
[0057] In one embodiment of the present invention, a computing device is provided. This computing device can be a terminal and may include a processor, a communication interface, memory, a display screen, and an input device. The processor, communication interface, and memory communicate with each other via a communication bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. When the computer programs are executed by the processor, they implement the methods described in the above embodiments. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals. Wireless communication can be achieved through Wi-Fi, a management network, NFC (Near Field Communication), or other technologies. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input device can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad mounted on the casing of the computing device, or an external keyboard, touchpad, or mouse. The processor can call logical instructions stored in the memory.
[0058] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0059] In one embodiment of the present invention, a computer program product is provided, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, and when the program instructions are executed by a computer, the computer is able to perform the methods provided in the above-described method embodiments.
[0060] In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided, which stores server instructions that cause a computer to perform the methods provided in the above embodiments.
[0061] The computer-readable storage medium provided in the above embodiments has a similar implementation principle and technical effect to the above method embodiments, and will not be described again here.
[0062] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0063] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0064] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for improving leakage identification performance, characterized in that, include: After data processing, the unlabeled monitoring data is divided into homologous data pairs and heterologous data pairs. Both types of data pairs are then input into the pre-trained model for feature extraction and pre-training. A leak detection and identification model is constructed based on features extracted from a pre-trained model, and the ability to distinguish data differences learned from unlabeled data is transferred to the leak detection and identification model. The leak detection and identification model is fine-tuned using labeled data, and the parameters of the leak detection and identification model are continuously optimized through backpropagation to improve the leak identification performance. After data processing, the unlabeled monitoring data were divided into homologous data pairs and heterologous data pairs. Specifically, a data augmentation method was used to generate two different versions for each unlabeled monitoring data time series, thus forming homologous data pairs and heterologous data pairs. Homologous data pairs consist of two augmented versions of the same time series, while heterologous data pairs consist of two augmented versions of different time series. A leak detection and identification model is constructed based on features extracted from a pre-trained model, and the ability to distinguish data differences learned from unlabeled data is transferred to the leak detection and identification model, including: The ability of the feature extractor to distinguish data differences learned from unlabeled data is transferred to the leak detection and identification model using fine-tuning techniques; In the training process of the leak detection and identification model, the pre-trained parameters are used as the initial parameters of the feature extractor, while the parameters of the classification layer are randomly initialized. Fine-tuning the leak detection and identification model using labeled data includes: training the leak detection and identification model using labeled data, evaluating the fine-tuning process using the cross-entropy loss function, and continuously optimizing the parameters of the leak detection and identification model through backpropagation.
2. The method for improving leakage identification performance as described in claim 1, characterized in that, Both data pairs are input into the pre-trained model for feature extraction, including: Both the generated homologous and non-homologous data pairs are input into the feature extractor in the pre-trained model to extract low-dimensional features of the time series. The low-dimensional features are then transformed into feature vectors through the mapping layer of the pre-trained model to calculate the contrastive loss.
3. The method for improving leakage identification performance as described in claim 2, characterized in that, Both data pairs are input into the pre-trained model for pre-training, including: The performance of the pre-training process is evaluated using a contrastive loss function, and the parameters of the pre-trained model are continuously optimized through backpropagation. The objective of the contrastive loss function is to maximize the similarity between the two data points in each pair of data from the same source, while minimizing the similarity between the two data points in each pair of data from different sources, so that the feature extractor has the ability to distinguish between the differences in the data.
4. The method for improving leakage identification performance as described in claim 3, characterized in that, Similarity is quantified using the cosine similarity index.
5. A system for improving leak detection performance, characterized in that, include: The pre-training module divides the unlabeled monitoring data into homogeneous data pairs and heterogeneous data pairs after data processing, and inputs both types of data pairs into the pre-training model for feature extraction and pre-training. The knowledge transfer module constructs a leak detection and identification model based on features extracted from the pre-trained model, and transfers the ability to distinguish data differences learned from unlabeled data to the leak detection and identification model. The model fine-tuning module uses labeled data to train the leak detection and identification model, and continuously optimizes the parameters of the leak detection and identification model through backpropagation to improve the leak identification performance. After data processing, the unlabeled monitoring data were divided into homologous data pairs and heterologous data pairs. Specifically, a data augmentation method was used to generate two different versions for each unlabeled monitoring data time series, thus forming homologous data pairs and heterologous data pairs. Homologous data pairs consist of two augmented versions of the same time series, while heterologous data pairs consist of two augmented versions of different time series. A leak detection and identification model is constructed based on features extracted from a pre-trained model, and the ability to distinguish data differences learned from unlabeled data is transferred to the leak detection and identification model, including: The ability of the feature extractor to distinguish data differences learned from unlabeled data is transferred to the leak detection and identification model using fine-tuning techniques; In the training process of the leak detection and identification model, the pre-trained parameters are used as the initial parameters of the feature extractor, while the parameters of the classification layer are randomly initialized. Fine-tuning the leak detection and identification model using labeled data includes: training the leak detection and identification model using labeled data, evaluating the fine-tuning process using the cross-entropy loss function, and continuously optimizing the parameters of the leak detection and identification model through backpropagation.
6. A computer-readable storage medium for storing one or more programs, characterized in that, The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods described in claims 1 to 4.
7. A computing device, characterized in that, include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods described in claims 1 to 4.
Citation Information
Patent Citations
Pipeline leakage detection method and device, electronic equipment and storage medium
CN117332324A
Method and system for identifying unbalanced encrypted traffic based on comparative learning pre-training
CN118523943A
Cited By
Lightweight water supply pipe network leakage identification method and system based on time-frequency cross-domain feature alignment, medium and equipment
CN121528242A