Water supply network sound leakage monitoring method and device based on multi-modal data fusion and multi-task learning

Through multimodal data fusion and multitask learning methods, combined with acoustic time series, time-frequency characteristics and spectrum diagrams, the problem of difficult-to-capture leak detection, evaluation and positioning tasks in the prior art is solved, and a high-precision and robust sound leakage monitoring of water supply pipeline networks is achieved.

CN120145219APending Publication Date: 2025-06-13INNOVATION CENTER OF YANGTZE RIVER DELTA ZHEJIANG UNIVERSITY +2

Patent Information

Application Number
CN202510093661.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing sound leakage monitoring technology of water supply networks is difficult to effectively capture the inherent correlation between leak detection, evaluation and positioning tasks, resulting in insufficient leakage monitoring accuracy and robustness of the model.

Method used

Using a method based on multimodal data fusion and multitask learning, a multi-level task prediction network is constructed through the residual convolutional neural network framework, combining acoustic time series, time-frequency characteristics and spectrograms to perform joint learning of leakage detection, evaluation and positioning.

Benefits of technology

The accuracy and robustness of the model in leak monitoring are improved, especially in tasks with scarce data, and the pipeline leakage detection, evaluation and positioning are achieved with high accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145219A_ABST
    Figure CN120145219A_ABST
Patent Text Reader

Abstract

The invention discloses a water supply network sound leakage monitoring method based on multi-modal data fusion and multi-task learning, and the method comprises the steps: inputting an original sound signal, and carrying out the labeling of a label; converting the original sound signal to obtain a corresponding logarithmic Mel spectrogram; performing feature extraction on the original sound signal to construct a time-frequency feature set; forming a data set by the original sound signal, the logarithmic Mel spectrogram, the label and the time-frequency feature set; constructing a multistage task prediction network based on a residual convolutional neural network framework; performing multi-task training on the multi-level task prediction network by using the data set to obtain a water supply network sound leakage monitoring model; and inputting the original sound signals of the collection points into the water supply network sound leakage monitoring model to obtain a prediction result. The invention further provides a water supply network sound leakage monitoring device. According to the method provided by the invention, the leakage detection aspect is promoted to more comprehensive sound leakage monitoring application through a data-driven acoustic detection technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of water supply pipeline leakage detection, and particularly relates to a method and device for acoustic leakage monitoring of water supply pipe networks based on multi-modal data fusion and multi-task learning. Background Art

[0002] The integrity and efficiency of water supply pipe networks are crucial for sustainable urban development, especially in the context of increasingly scarce water resources and continuously growing demand. Pipe leakage not only causes huge economic losses but also may pose environmental risks and threaten public health. With the acceleration of urbanization and the aging of infrastructure, the demand for efficient and comprehensive leakage monitoring technologies has become increasingly urgent. The problem of acoustic leakage monitoring of water supply pipe networks involves multiple subtasks, including leakage detection, leakage assessment, and leakage location. For each subtask, an independent model needs to be established separately, and most existing studies also adopt this method. However, separately modeling these three subtasks not only causes waste in computing and human resources, but more importantly, these subtasks share the same dataset and there are strong implicit correlations between them. Independent models often fail to effectively capture and utilize this inherent correlation. Therefore, how to systematically model these three tasks in a unified paradigm has become the key to improving the accuracy of acoustic leakage monitoring, which is also an issue not explored by existing studies.

[0003] The academic literature "Leakage Detection in Water Distribution Systems Based on Time–Frequency Convolutional Neural Network [J]" discloses that the use of short-time Fourier transform and time-frequency convolutional neural network improves the accuracy of leakage detection.

[0004] The academic literature "Feature selection of acoustic signals for leak detection in water pipelines [J]" proposes a feature selection method combining maximum discriminability and minimum redundancy, and uses the SHAP technique to identify four key leakage features. However, most current detection methods rely on single-modal acoustic data and are often limited by environmental noise and insufficient feature extraction, affecting the on-site application performance of the model. In recent years, some researchers have begun to pay attention to methods of multi-modal data fusion to obtain more comprehensive leakage features.

[0005] Patent document CN118959911A discloses a method and system for identifying the working condition sound signals of a water supply pipeline. The method includes: acquiring the working condition sound signals on the inner wall of the water supply pipeline; inputting the working condition sound signals into a leakage identification model: inputting the working condition sound signals into a first processing branch line to obtain a first result feature related to the features of the working condition sound signals in the time domain; inputting the working condition sound signals into a second processing branch line to obtain a second result feature related to the features of the working condition sound signals in the frequency domain; inputting the working condition sound signals into a third processing branch line to obtain a third result feature related to the features of the working condition sound signals in the mode; splicing the three result features, and the output result is that the water supply pipeline leaks or the water supply pipeline does not leak. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and device for monitoring sound leakage in a water supply network based on multi-modal data fusion and multi-task learning. This method advances from leakage detection to a more comprehensive sound leakage monitoring application through data-driven acoustic detection technology.

[0007] To achieve the first object of the present invention, the following technical solutions are provided: A method for monitoring sound leakage in a water supply network based on multi-modal data fusion and multi-task learning, including the following steps: Input the original sound signal, where the original sound signal includes normal noise signals, periodic interference signals, and pipeline leakage signals, and label the original sound signal according to whether there is leakage, the leakage position information opposite to the original sound signal acquisition point, and the leakage level. Convert the original sound signal to obtain the corresponding logarithmic Mel spectrogram. Extract the time domain features and frequency domain features in the original sound signal to construct a time-frequency feature set corresponding to the original signal. Form a data set from the original sound signal, logarithmic Mel spectrogram, label, and time-frequency feature set. Construct a multi-level task prediction network based on the residual convolutional neural network framework. The multi-level task prediction network includes a feature extraction module, a multi-modal feature fusion module, and a multi-task prediction module. The feature extraction module includes a first feature extractor, a second feature extractor, and a third feature extractor. The first feature extractor is used to extract the sound signal sequence features in the input original sound signal. The second feature extractor is used to extract the time domain features and frequency domain features in the input original sound signal. The third feature extractor is used to extract the spectral features of the logarithmic Mel spectrogram corresponding to the input original sound signal. The multi-modal feature fusion module splices the extracted sound signal sequence, time domain features, frequency domain features, and spectral features to obtain a corresponding fusion feature sequence. The multi-task prediction module performs task prediction based on the fused feature sequence to output a prediction result including whether there is leakage, the leakage level, and the leakage location information; Use the data set to perform multi-task training on the multi-level task prediction network to obtain a water supply network acoustic leakage monitoring model for predicting multi-tasks; Input the original acoustic signal at the collection point into the water supply network acoustic leakage monitoring model to obtain a prediction result.

[0008] The present invention utilizes the multi-modal advantages of acoustic data, combines acoustic time series, manually extracted time-frequency features, and spectrograms to provide more comprehensive acoustic feature information; jointly learns leakage detection, evaluation, and localization tasks, captures shared features between different tasks, thereby improving the accuracy and robustness of the model in each task, especially in tasks with scarce data.

[0009] Specifically, the periodic interference signal includes the dripping sound after rain or electromagnetic pulse interference.

[0010] Specifically, the pipeline leakage signal includes metal pipeline leakage signals and plastic pipeline leakage signals.

[0011] Specifically, the process of converting the original acoustic signal into a logarithmic Mel spectrogram is as follows: Perform pre-emphasis on the original acoustic signal, frame and window the pre-emphasized original acoustic signal; Among them, the purpose of windowing is to reduce the boundary effect of each frame of signal to reduce spectral leakage.

[0012] Perform a fast Fourier transform on each frame of signal and input it into the Mel filter bank to output a Mel spectrum; Use the Mel spectrum for logarithmic transformation to obtain a logarithmic Mel spectrogram; The main purpose of the above process is to convert the original audio signal into a spectral representation suitable for human hearing while being able to retain the key information in the original pipeline vibration signal.

[0013] Specifically, during the multi-task training process, a multi-loss function is used to adjust the parameters in the multi-level task prediction network, and the weights of the tasks are dynamically adjusted during the adjustment process according to the loss reduction speed.

[0014] Specifically, the multi-loss function includes the first multi-class cross-entropy, the second multi-class cross-entropy, and the root mean square error; The expression of the first multi-class cross-entropy is as follows:

[0015] The expression of the second multi-class cross-entropy is as follows: ; ; Among them, is the true label value in the multi-class cross-entropy; is the predicted probability value in the multi-class cross-entropy; is the number of samples; is the number of classes; The expression of the root mean square error is as follows: ; Among them, represents the true label value of the root mean square error, represents the predicted probability value of the root mean square error.

[0016] Specifically, the process of dynamically adjusting the weights of tasks according to the loss decrease rate is as follows: When the multi-loss function values of task i in the t-th epoch and the (t - 1)-th epoch in all tasks are respectively and , then the expression of the loss change rate of task i is as follows:

[0017] Among them, is the loss change rate of task ; is the loss of task in the -th epoch, is the loss of task in the -th epoch; According to the loss change rate of task i , update the weight of task i in all tasks, and its expression is as follows: ; Among them, is the weight of task in the -th epoch; is a constant, taking 1e - 7; Normalize the weights of all tasks so that the sum of the weights is 1.

[0018] To achieve the second object of the present invention, the following technical solution is provided: A sound leakage monitoring device for a water supply network, which is used to implement the steps of the above-mentioned sound leakage monitoring method for a water supply network based on multi-modal data fusion and multi-task learning.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: It has high accuracy and robustness for acoustic leakage monitoring of water distribution networks, can simultaneously achieve pipeline leakage detection, evaluation and localization, and aims to advance the data-driven acoustic detection technology from leakage detection to a more comprehensive acoustic leakage monitoring application. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a flowchart of the water supply network acoustic leakage monitoring method based on multi-modal data fusion and multi-task learning provided in this embodiment; Figure 2 It is a schematic diagram of the original acoustic signal provided in this embodiment; Figure 3 It is a logarithmic Mel spectrogram of the original acoustic signal provided in this embodiment; Figure 4 It is a schematic diagram of the residual block structure of the multi-level task prediction network provided in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0022] As Figure 1 shown, a water supply network acoustic leakage monitoring method based on multi-modal data fusion and multi-task learning provided in this embodiment includes the following steps: Input the original acoustic signal, which includes normal noise signals, periodic interference signals, and pipeline leakage signals, and label the original acoustic signal according to whether there is leakage, the leakage position information opposite to the original acoustic signal acquisition point, and the leakage level; Convert the original acoustic signal to obtain the corresponding logarithmic Mel spectrogram; Extract the time-domain features and frequency-domain features in the original acoustic signal to construct a time-frequency feature set corresponding to the original signal.

[0023] Construct a multi - level task prediction network based on the residual convolutional neural network framework. The multi - level task prediction network includes a feature extraction module, a multi - modal feature fusion module, and a multi - task prediction module; The feature extraction module includes a first feature extractor, a second feature extractor, and a third feature extractor. The first feature extractor is used to extract the acoustic signal sequence features in the input original acoustic signal. The second feature extractor is used to extract the time - domain features and frequency - domain features in the input original acoustic signal. The third feature extractor is used to extract the spectral features of the logarithmic Mel spectrogram corresponding to the input original acoustic signal; The multi - modal feature fusion module splices the obtained acoustic signal sequence, time - domain features, frequency - domain features, and spectral features to obtain a corresponding fused feature sequence; The multi - task prediction module performs task prediction according to the fused feature sequence to output a prediction result including information on whether there is a leakage, the leakage level, and the leakage location.

[0024] Use a data set to perform multi - task training on the multi - level task prediction network to obtain a water supply network acoustic leakage monitoring model for predicting multi - tasks.

[0025] Input the original acoustic signal at the collection point into the water supply network acoustic leakage monitoring model to obtain a prediction result.

[0026] More specifically, in this embodiment, for the initial acoustic data at the leakage location, perform data processing on it to obtain data features in multiple modalities.

[0027] Modality 1: For the original acoustic signal, since the original acoustic signal can reveal information such as the waveform, amplitude, amplitude change, periodicity, etc. In the technical solution of this embodiment, four types of original acoustic signals are selected as shown in Figure 2 Normal noise, periodic interference (dripping sound after rain or electromagnetic pulse interference), pipeline leakage signals (divided into metal pipeline leakage signals and plastic pipeline leakage signals).

[0028] Modality 2: In acoustic leakage monitoring, basic statistical parameters from the time - domain and frequency - domain can be used as features for pipeline state identification and prediction. In the technical method of this embodiment, 8 time - domain features and 6 time - frequency features are selected, and their feature numbers are T1 - T8 and F1 - F6: Mean value (T1), ; Variance (T2), ; Skewness (T3), ; Maximum value (T4), ; Root mean square (T5), ; Kurtosis (T6), ; Energy (T7), ; Zero-crossing rate (T8), ; Mean frequency (F1), ; Peak frequency (F2), ; Frequency centroid (F3), ; Root mean square frequency (F4), ; Frequency standard deviation (F5), ; Frequency energy (F6), .

[0029] Mode 3: In this embodiment, the original acoustic signal is converted into a logarithmic Mel spectrogram. First, pre-emphasis is performed on the pipeline vibration signal, and then it is framed and windowed. The purpose of windowing is to reduce the boundary effect of each frame of the signal and reduce spectral leakage. Then, the fast Fourier transform (FFT) is performed on each frame of the signal to obtain the spectrum. After that, the Mel filter bank is applied to convert the spectrum into a Mel spectrum, and a logarithmic transformation is performed to obtain the logarithmic Mel spectrogram as shown in Figure 3 , where Figure 3 (a) in is the logarithmic Mel spectrogram of background noise, Figure 3 (b) in is the logarithmic Mel spectrogram of periodic interference, Figure 3 (c) in is the logarithmic Mel spectrogram of metal pipeline leakage, Figure 4 (d) in is the logarithmic Mel spectrogram of plastic pipeline leakage.

[0030] The main purpose of this process is to convert the original audio signal into a spectral representation suitable for human hearing while being able to retain the key information in the original pipeline vibration signal.

[0031] In the multi-level task prediction network mentioned in this embodiment, a 1D residual convolutional neural network (1D R-CNN) with weight sharing is mainly used as the backbone to extract preliminary features from the three-modal data, and then the Concatenate layer is applied for fusion. 1D R-CNN is a network that stacks one-dimensional convolutional layers through residual blocks (Residual Block). The structure of the residual block can prevent information loss in the multi-layer network and help the network train deeper. Another reason for choosing 1D R-CNN is its flexibility, which can well extract features from various modal data.

[0032] In the convolutional layer, the input is convolved by a learnable convolutional kernel, and the resulting value passes through the ReLU activation function to obtain the input of the next layer. The convolution operation and ReLU are described by equations (1) and (2).

[0033] (1) (2) Among them, is the th convolution vector of the layer, is the number of input feature vectors, is the th input feature vector of the layer, is the correlation operation, is the th layer th bias vector. is the th th output after the convolution operation of the layer, is the value of the activation function

[0034] To efficiently extract more features and improve the network training speed, a pooling layer is used to downsample the features learned by the convolutional layer, while preventing network overfitting to a certain extent. The pooling operation is described by equation (3): (3) Among them, is the activation value of the th neuron in the th feature vector in the layer;

[0035] (4) The main idea of the residual structure is to use a shortcut connection to combine multiple hidden layers into a whole, that is, a residual module, as Figure 4 shown. When the residual module input passes through the main path, its other branch bypasses the implicit layer through the shortcut connection and directly sums with the output of the main path to form the residual module output as shown in equation (4).

[0036] After the data of the three modalities are extracted by 1D R-CNN features, the Concatenate layer is used to splice the features and fuse the information from different modalities.

[0037] Three tasks are learned in parallel, and the output results affect each other. The differences and connections between tasks are taken into account simultaneously. This multi-task paradigm allows each task to learn highly targeted features on the basis of sharing information. At the same time, it can also weight each task through the total loss function to ensure that MDFML achieves good overall performance while balancing each task. In this work, Task 1 is leakage detection, and its output includes four types. Therefore, the multi-class cross-entropy is used as the loss function.

[0038] (5) Among them, is the true label value; is the predicted probability value; is the number of samples; is the number of classes.

[0039] In addition, since some labels in Task 2 and Task 3 are missing, as shown in Table 3. When calculating the total loss, it is necessary to assign values (such as -1) to the missing labels in Task 2 and Task 3, and then perform mask processing, as shown in Equation (6). Task 2 is leakage assessment, and its output is three leakage levels. Therefore, the multi-class cross-entropy is also used as the loss function, as shown in Equation (7).

[0040] (6) (7) Task 3 is leakage location, and its output is the leakage distance. Therefore, the root mean square error is also used as the loss function.

[0041] (8) To achieve the joint optimization training of the three tasks, combined with Equations (5), (7) and (8), the total loss function of MDFML is constructed, as shown in Equation (9).

[0042] (9) Among them, is the total loss function; are the loss function weights of Task 1, Task 2 and Task 3 respectively.

[0043] Considering that the data volumes and loss functions of the three tasks are quite different, the optimization difficulties are also different. Therefore, Dynamic Weight Averaging (DWA) is introduced to adaptively optimize the loss weights of each task. DWA dynamically adjusts the weights of tasks according to the loss reduction speed of each task during the training process, so that tasks with slower learning speeds receive more attention. It mainly consists of the following three steps: (1)When all tasks iThe multi-loss function values at the t-th epoch and the (t - 1)-th epoch are respectively and , then the expression of the loss change rate of task i is as shown in Equation (10).

[0044] (10) where is the loss change rate of task ; is the loss of task at the -th epoch, is the loss of task at the -th epoch.

[0045] (2) Weight update. According to the loss change rate of task i , update the weight of task i among all tasks, and the calculation formula is as shown in Equation (11).

[0046] (11) where is the weight of task at the -th epoch; is a constant, taking 1e-7.

[0047] (3) Weight normalization. Normalize the weights of all tasks so that the sum of the weights is 1, as shown in Equations (12) and (13).

[0048] (12) (13) where is the total number of tasks; are the weights of tasks 1, 2, and 3 at the -th epoch respectively.

[0049] In summary, MDFML adopts the stochastic gradient descent method and finds its best convergence by minimizing the total loss function of Equation (8) in the optimizer Adam. During training, MDFML combines dynamic weighted averaging to optimize the weights of each task, and at the same time updates , and , optimizes the parameters of each convolutional layer and fully connected layer and shares them with other tasks.

[0050] In addition, accuracy, precision, F1-score, sensitivity, and specificity are used as classification evaluation metrics for Tasks 1 and 2. In Table 1, mean squared error (MSE), root mean squared error (RMSE), coefficient of determination (R 2 ), mean absolute error (MAE), and mean absolute percentage error (MAPE) are used to evaluate the performance of Task 3. Among them, is the label, is the predicted value, is the average value of the predicted values, and its evaluation metrics are shown in Table 1.

[0051] Table 1

[0052] To better illustrate the technical effects of the method provided in this embodiment, a dataset is constructed, which includes 2,295 metal pipeline leakage signals, 441 plastic pipeline leakage signals, 2,349 normal noise signals, and 2,322 periodic interference signals. In addition to having leakage status and noise labels, some of the collected signal samples also have leakage flow rate and leakage distance labels, as shown in Table 2, where the leakage distance is defined as the distance between the leakage point and the sensor deployment location. This distance is identified using a correlator and then verified through on-site excavation, and the corresponding leakage flow rate is measured.

[0053] Table 2

[0054] To illustrate the superiority of the model provided in this embodiment (hereinafter collectively referred to as MDFML) in leakage monitoring, the performance of independent models for each task is compared and analyzed.

[0055] Table 3

[0056] As shown in Table 3, for Task 1, the accuracy, precision, F1-score, and sensitivity of the independent model are only 96.49%, 96.57%, 96.65%, and 96.90%, respectively. While the accuracy, precision, F1-score, and sensitivity of MDFML reach 99.49%, 99.40%, 99.49%, and 99.57%, respectively. While the robustness is increased, these four evaluation metrics have all increased by more than 2 percentage points. The data volume of Task 1 is relatively large compared to Tasks 2 and 3, which helps MDFML extract richer and more robust feature representations. These features are not only applicable to Task 1 itself, but can also be used as prior knowledge for Tasks 2 and 3 through shared learning representations, providing important input information for Tasks 2 and 3, thereby improving their recognition ability and prediction accuracy.

[0057] For Task 2, the accuracy of the independent model is only 96.13%, while that of MDFML reaches 99.53%, which is more than 3% higher than that of the independent model. In addition, this improvement is not only reflected in accuracy, but also shows similar advantages in precision, F1-score, and sensitivity. By training multiple related tasks simultaneously, MDFML can better capture the correlations between different tasks (such as time-frequency features), utilize shared feature information to improve leakage monitoring performance, and thus enhance the ability to handle complex problems.

[0058] In Task 3, MDFML also demonstrates its excellent performance. The mean squared error (MSE) and mean absolute error (MAE) are 0.84 and 0.66, significantly lower than 5.40 and 1.68 of the independent model, fully meeting the requirements of on-site engineering detection (MAE less than 1m). In addition, the R² of MDFML reaches 99.53%, while that of the independent model is 97.12%, indicating that MDFML can better fit the leakage distance data.

[0059] In the ablation experiment of multi-modal data, the contributions of three modalities, namely Modality 1 (acoustic signal time series), Modality 2 (handcrafted time-frequency features), and Modality 3 (log Mel spectrogram), are evaluated one by one to reveal the impact of each modality on the performance of MDFML and the advantages of multi-modal learning.

[0060] As shown in Table 4, MDFML with the complete combination of the three modalities of data outperforms the other three cases in Tasks 1, 2, and 3. When Modality 1 is removed, the performance of MDFML in Tasks 1 and 2 both decreases slightly, with the accuracy, precision, F1-score, and sensitivity all decreasing by about 0.7% on average. When Modality 2 is removed, compared with the removal of Modality 1, the average decline of each evaluation index in Tasks 1 and 2 is slightly larger, about 1.0%. Generally speaking, the original acoustic signal sequence contains a large amount of time-domain noise and redundant information, so it is reasonable that its contribution to the performance improvement of MDFML is small. The performance of MDFML without Modality 2 is only slightly lower than that without Modality 1, which may be due to the number of time-domain and frequency-domain features of Modality 2. Because too many handcrafted features will cause feature information redundancy, while fewer features may not be able to fully describe the characteristics of the leakage acoustic signal. When Modality 3 is removed, the accuracy, precision, F1-score, and sensitivity of Task 1 all decrease by more than 1.5%, indicating that Modality 3 makes a particularly significant contribution to the performance improvement of MDFML compared with Modality 1 and 2. The same trend is also shown in Task 2 when Modality 3 is removed. Modality 3 converts the original audio into a graph-structured data, which can enhance the frequency-domain feature expression of the signal and is more conducive to the deep learning model to capture the local features of its spectrum.

[0061] In Task 3, the mean absolute error (MAE) of MDFML is 0.53 m. After removing Mode 1, it increases to 0.69 m respectively, showing that the influence of Mode 1 on the leakage localization performance is limited. However, after removing Modes 2 and 3, the values of MAE increase to 0.97 m and 1.09 m respectively, proving the importance of Modes 2 and 3 in acoustic leakage localization.

[0062] Table 4

[0063] In acoustic leakage monitoring, although different tasks are highly correlated, they also have different focuses during the training process. During the training of MDFML, it is found that the loss values of Tasks 1 and 2 are at the same order of magnitude (classification tasks), while the loss of Task 3 (regression task) is much larger than that of Tasks 1 and 2. If the loss weights of the three tasks are the same, it may cause Task 3 to become the main task for the model to optimize. However, the data volume of Task 3 is the smallest and its ability to resist noise is also the weakest, which may lead to insufficient training of Tasks 1 and 2. Therefore, it is necessary to optimize and explore the loss weights of each task to ensure that MDFML can achieve the optimal leakage monitoring performance.

[0064] Therefore, three methods are selected: (1) Proportion method: Based on the above prior knowledge, fix the loss weight coefficients α and β in Tasks 1 and 2, and conduct a sensitivity analysis of the loss weight γ only based on the initial numerical scale of the loss function of Task 3. (2) Normalization: Normalize the loss functions of the three tasks. (3) Dynamic weighted average. The results of the three methods are shown in Table 5 below.

[0065] Table 5

[0066] For the proportion method, with the change of the loss weight γ in Task 3, the recognition results of each task have changed significantly. As γ changes from 0.001 to 1 in turn, the evaluation indicators of Tasks 1 and 2 in MDFML show a trend of first rising and then falling, indicating that the leakage detection and evaluation performance of MDFML first increases and then decreases. The MSE, RMSE, MAE, and MAPE of Task 3 show a trend of first falling and then rising, and R 2 shows a trend of rising and then falling, indicating that the leakage localization prediction performance of MDFML also first increases and then decreases.

[0067] When conducting the regularization method, it is necessary to regularize the loss functions of tasks 1, 2, and 3 simultaneously. In MDFML, the evaluation indicators of tasks 1 and 2 are very close to the performance when γ is 0.1 in the proportional method, and even exceed it in terms of the accuracy of task 1. However, the MSE and MAE of task 3 are only 1.74m and 0.93m respectively, which are significantly lower than 0.84m and 0.66m when γ is 0.1.

[0068] As can be seen from Table 5, compared with the above two methods, the dynamic weighted average method has obvious advantages in all three tasks. The indicators in tasks 1 and 2 all reach about 99.5%, and the differences between the evaluation indicators are small. The average MAE in task 3 reaches 0.53m, which is significantly lower than 0.66m and 0.93m of the proportional method and normalization. To sum up, to ensure that all tasks of MDFML can exhibit excellent performance, the dynamic weighted average is more suitable than the proportional method and normalization.

[0069] This embodiment also provides a sound leakage monitoring device for a water supply network, which is used to implement the steps of the sound leakage monitoring method for a water supply network based on multi-modal data fusion and multi-task learning provided in the above embodiment.

[0070] The present invention has high accuracy and robustness for acoustic leakage monitoring of a water distribution network, and can simultaneously achieve pipeline leakage detection, evaluation, and localization. It aims to advance the data-driven acoustic detection technology from leakage detection to a more comprehensive acoustic leakage monitoring application. The invention can make full use of the multi-modal advantages of acoustic data, combine acoustic time series, manually extracted time-frequency features, and spectrograms to provide more comprehensive acoustic feature information. At the same time, the research jointly learns the leakage detection, evaluation, and localization tasks to capture the shared features between different tasks, thereby improving the accuracy and robustness of the model in each task, especially in tasks with scarce data. The MDFML method not only provides a new technical path for leakage monitoring, but also provides stronger technical support for urban water supply management and sustainable development.

Claims

1. A method for monitoring acoustic leakage in a water supply network based on multimodal data fusion and multi-task learning, characterized in that: The following steps are involved: Input an original sound signal, which includes a normal noise signal, a periodic interference signal, and a pipeline leakage signal, and label the original sound signal with whether there is leakage, leakage position information relative to the original sound signal collection point, and leakage level; Convert the original sound signal to obtain the corresponding logarithmic Mel spectrum; Extracting time domain features and frequency domain features from the original sound signal to construct a time-frequency feature set corresponding to the original signal; The original sound signal, logarithmic Mel spectrogram, label and time-frequency feature set are combined into a data set; Building a multi-level task prediction network based on a residual convolutional neural network framework, wherein the multi-level task prediction network includes a feature extraction module, a multimodal feature fusion module, and a multi-task prediction module; The feature extraction module includes a first feature extractor, a second feature extractor and a third feature extractor, wherein the first feature extractor is used to extract the acoustic signal sequence features in the input original acoustic signal, the second feature extractor is used to extract the time domain features and the frequency domain features in the input original acoustic signal, and the third feature extractor is used to extract the spectrum features of the logarithmic Mel spectrum corresponding to the input original acoustic signal; The multimodal feature fusion module splices the extracted acoustic signal sequence, time domain features, frequency domain features and spectrum features to obtain a corresponding fusion feature sequence; The multi-task prediction module performs task prediction based on the fused feature sequence to output a prediction result including information on leakage, leakage level and leakage location; The multi-task prediction network is trained using the dataset to obtain a water supply network acoustic leakage monitoring model for predicting multiple tasks. The original acoustic signal of the collection point is input into the water supply network acoustic leakage monitoring model to obtain the prediction result.

2. The water supply network acoustic leakage monitoring method based on multimodal data fusion and multi-task learning according to requirement 1 is characterized in that: The periodic interference signal includes the sound of dripping water after rain or electromagnetic pulse interference.

3. The method for monitoring acoustic leakage of a water supply network based on multimodal data fusion and multi-task learning according to claim 1 is characterized in that: The pipeline leakage signal includes a metal pipeline leakage signal and a plastic pipeline leakage signal.

4. The method for monitoring acoustic leakage of a water supply network based on multimodal data fusion and multi-task learning according to claim 1 is characterized in that: The process of converting the original sound signal into a logarithmic Mel spectrum is as follows: Pre-emphasize the original sound signal, divide the pre-emphasized original sound signal into frames and add windows; Perform fast Fourier transform on each frame signal and input it into the Mel filter bank to output the Mel spectrum; The Mel spectrum is logarithmically transformed to obtain a logarithmic Mel spectrum graph.

5. The method for monitoring acoustic leakage of a water supply network based on multimodal data fusion and multi-task learning according to claim 1, characterized in that: During the multi-task training process, multiple loss functions are used to adjust the parameters in the multi-level task prediction network, and the weights of the tasks are dynamically adjusted according to the loss reduction rate during the adjustment process.

6. The method for monitoring acoustic leakage of a water supply network based on multimodal data fusion and multi-task learning according to claim 5 is characterized in that: The multiple loss functions include a first multi-classification cross entropy, a second multi-classification cross entropy, and a root mean square error; The expression of the first multi-classification cross entropy is as follows: The expression of the second multi-classification cross entropy is as follows: ; ; in, is the true label value in multi-classification cross entropy; is the predicted probability value in multi-classification cross entropy; is the number of samples; is the number of categories; The expression of the root mean square error is as follows: ; in, represents the true label value of the root mean square error, The predicted probability value represents the root mean square error.

7. The method for monitoring acoustic leakage of a water supply network based on multimodal data fusion and multi-task learning according to claim 5, characterized in that: The process of dynamically adjusting the weight of the task by the rate of loss decrease is as follows: When all tasks i The multiple loss function values ​​at the tth epoch and t-1th epoch are and , then the task i The expression of the loss change rate is as follows: in, For the task The rate of change of loss; For the task In the The loss of epochs, For the task In the The loss of epochs; According to the task i The loss change rate, update task i The weights in all tasks are expressed as follows: ; in, For the task In the The weight of each epoch; is a constant, take 1e-7; Normalize the weights of all tasks so that the weights are 1.

8. A water supply network acoustic leakage monitoring device, characterized in that: Steps for implementing the method for monitoring acoustic leakage of a water supply network based on multimodal data fusion and multi-task learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and system for identifying working condition sound signals of water supply pipeline

    CN118959911A

Cited By

  • Intelligent water supply management system and method based on wind, solar and water multi-source data fusion

    CN120580092B

  • Multi-task language processing model training method and device and related equipment

    CN120598061A

  • A multi-task language processing model training method and device and related equipment

    CN120598061B

  • Hydrogen-doped natural gas pipeline leakage detection method and system based on multi-task Mamba-CNN

    CN120932683A

  • Agricultural operation environment state parameter detection method

    CN121033486A