Machine learning water supply pipe network false alarm identification method based on domain knowledge
Through a machine learning method based on domain knowledge, combined with TNFS and GES algorithms, the acoustic signal data in the leak detection system is processed and enhanced, and the problems of false alarms and data imbalance are solved, achieving more accurate leakage identification and reducing false alarm rates.
Patent Information
- Application Number
- CN202510289311.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
Existing leak detection systems are susceptible to background environmental noise, resulting in false positives, and classical methods are difficult to deal with data scarcity and imbalance.
Using a machine learning method based on domain knowledge, acoustic wave signals are collected through acoustic sensors, fast Fourier transform and minimum-maximum normalization are performed, transformer noise characteristics are extracted in combination with the TNFS algorithm, and data augmentation is used by the GES algorithm to generate a balanced data set for training machine learning models.
Effectively identifying leaked data and false alarm data, reducing the false alarm rate and solving the problem of high false alarm rate of leakage detection system caused by electrical noise.
Smart Images

Figure CN120217232A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of pipeline leakage detection and classification, and particularly relates to a method for identifying false alarms in a water supply pipe network based on domain knowledge and machine learning. Background Art
[0002] Leakage detection has always been a key aspect of water loss management in water distribution networks (WDNs). However, the continuous occurrence of leaks in current infrastructure poses a problem, and financial constraints make it challenging to replace aging facilities, further exacerbating this problem. This continuous leakage results in the waste of precious water resources.
[0003] In recent years, research has been actively exploring the application of artificial intelligence (AI) technology in Internet of Things-based acoustic recorders to detect leaks in WDNs. The main problem encountered by researchers in related work on detecting WDNs is not the detection of leak points, but the handling of false alarms, because a significant drawback of the acoustic recorder system it relies on is that it is vulnerable to background environmental noise, and the noise level around water distribution pipes may cause false alarm problems. This is a problem that previous research has not successfully solved. This leads to a decline in the performance of leakage detection systems and an increase in human resource consumption.
[0004] In the prior art, using bidirectional long short-term memory (BiLSTM) and deep autoencoder (AE) to reduce false alarms caused is a common and effective method. However, when collecting actual data through acoustic recorders in WDNs, a common problem is the scarcity and imbalance of data: normal sounds and leakage sounds, leakage sounds and false alarms. The first and second items represent the main and secondary categories respectively. And these classical methods are difficult to handle the above problems. While linear prediction extracts features from the original signal and the noise-removed signal, singular spectrum analysis method, short-time Fourier transform identification, and other feature representation methods do improve the accuracy of leakage detection, but they do not provide a clear solution to the false alarm problem. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a method for identifying false alarms in a water supply pipe network based on domain knowledge and machine learning, including: an offline stage and an online stage;
[0006] The offline stage includes:
[0007] S1: Use a sound sensor to collect acoustic wave signals generated by pipeline leakage or other reasons, and then perform a fast Fourier transform (FFT) on the obtained acoustic wave signals to convert them into frequency domain data;
[0008] The acoustic wave signals include: indoor leakage signals, outdoor leakage signals, interference signals, electrical / mechanical noise interference signals, living interference noise signals, and normal signals; among them, the transformer noise in the electrical / mechanical noise interference signals is relatively close to the electrical noise frequency, and the system cannot identify the transformer noise, which is an important factor leading to false alarms of the system;
[0009] S2: Perform min-max normalization on the electrical / mechanical noise interference signals in the frequency domain data to obtain the normalized frequency domain data;
[0010] S3: Extract the transformer noise in the electrical / mechanical noise interference signals based on the normalized frequency domain data through the TNFS algorithm, and form a data set with other acoustic wave signals;
[0011] S4: Use the GES algorithm for oversampling to achieve data augmentation, complete the balance of the data set, and obtain a new data set;
[0012] S5: Use the new data set as the training data of the machine learning model to classify the leakage sound and interference noise and obtain preliminary results;
[0013] The online stage includes:
[0014] S6: Embed the algorithm trained in the offline stage in the acoustic recorder, and the acoustic recorder monitors and collects pipeline leakage signals in real time and extracts abnormal signals;
[0015] S7: Process the extracted abnormal signals through the min-max normalization and TNFS mentioned above;
[0016] S8: Use the trained machine learning algorithm to judge leakage / false alarms, and upload the results to the server for the operator to handle actually.
[0017] Advantages of the present invention:
[0018] Through the leakage feature selection method based on domain knowledge drive, the present invention introduces the transformer noise characteristics as domain knowledge for frequency selection, and proposes the synthetic minority oversampling technique (SMOTE) of the generative adversarial network (GAN) to perform data augmentation to solve the problems of data scarcity and sample imbalance. The introduction of the synthetic minority oversampling technique can solve the problem of mode collapse of the generative adversarial network, and can accurately identify leakage data and false alarm data, thus effectively solving the problem of high false alarm rate of the leakage detection system caused by electrical noise. Description of the drawings
[0019] Figure 1 It is a schematic diagram of the overall framework of the present invention;
[0020] Figure 2 Schematic diagram of the data enhancement method for GES of the present invention. Specific embodiments
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work shall fall within the protection scope of the present invention.
[0022] As Figure 1 shown, Figure 1 shows a machine learning false leakage alarm prevention framework for water supply pipe networks driven by domain knowledge. The proposed framework is divided into two stages: online and offline. Figure 1 On the left is the offline stage and on the right is the online stage;
[0023] 1) Offline stage process:
[0024] 1: For the data collected on site, Hz is used as the feature, and the type and problem are used as the training labels of the network model.
[0025] 2: Min-max normalization and TNFS process. Among them, min-max normalization is a data scaling technique that uses the minimum and maximum values of the features to linearly transform each feature in the dataset into a predetermined range, usually [0,1].
[0026] 3: Data enhancement. After the above preprocessing, the collected data is divided into leakage sounds or electrical noises according to different types and problems. Subsequently, data enhancement techniques are used to oversample the electrical noises in the dataset. To effectively enhance the data, a method called GES is proposed.
[0027] 4: Network training classification. The balanced dataset is used as the training data for the machine learning model to classify leakage sounds and electrical noises and obtain preliminary results.
[0028] 2) Online stage process
[0029] After the offline stage is completed, the trained noise classifier is applied in an acoustic logger
[0030] 1: Data detection and processing. An acoustic sensor, FFT, and leakage detection algorithm constitute an acoustic logger. The acoustic sensor monitors the water supply pipe network in real time, collects the data suspected of leakage, and performs fast Fourier transform (FFT) to obtain the frequency domain signal, which is convenient for the leakage detection algorithm to classify.
[0031] 2: Abnormal data classification. The data classified as abnormal by the leakage detection algorithm will be further classified as leakage and noise by the embedded noise classifier again.
[0032] 3: False alarm prediction. The data with leakage or noise labels is then transmitted to the server. This information can significantly help the system operator to identify false alarms. The final leakage diagnosis is determined by the system operator using the data on the server. To distinguish between leakage and noise events, the following decision rules are used:
[0033]
[0034] Among them, N represents the number of acoustic recorders near the potential leakage location, leak represents determined as leakage, and noise represents determined as noise (no leakage). dn is the label of the data received from the nth acoustic recorder, where 1 represents leakage and 0 represents noise. In WDN, false leakage alarms are more harmful than missed leakage events. Although missed leakage events only result in increased water loss before the next inspection cycle, false leakage alarms will lead to unnecessary on-site inspection and excavation work costs, causing greater economic losses, thus this rule is formulated.
[0035] A method for identifying false alarms in water supply pipe networks based on domain knowledge machine learning, including: an offline stage and an online stage;
[0036] The offline stage includes:
[0037] S1: Use acoustic sensors to collect the acoustic wave signals generated by pipeline leakage or other reasons, and then perform fast Fourier transform (FFT) on the obtained acoustic wave signals to convert them into frequency domain data;
[0038] The acoustic wave signals include: indoor leakage signals, outdoor leakage signals, interference signals, electrical / mechanical noise interference signals, living interference noise signals, and normal signals; among them, the transformer noise in the electrical / mechanical noise interference signals is relatively close to the electrical noise frequency, and the system cannot identify the transformer noise, which is an important factor leading to false alarms in the system;
[0039] S2: Perform min-max normalization on the electrical / mechanical noise interference signals in the frequency domain data to obtain the normalized frequency domain data;
[0040] S3: Extract the transformer noise in the electrical / mechanical noise interference signals based on the normalized frequency domain data through the TNFS algorithm, and form a data set with other acoustic wave signals;
[0041] S4: Use the GES algorithm for oversampling to achieve data enhancement, complete the balance of the data set, and obtain a new data set;
[0042] S5: Use the new data set as the training data for the machine learning model to classify the leakage sound and interference noise and obtain preliminary results;
[0043] The online phase includes:
[0044] S6: Embed the algorithm trained in the offline phase in the acoustic recorder, and the acoustic recorder monitors and collects pipeline leakage signals in real time and extracts abnormal signals;
[0045] S7: Process the extracted abnormal signals through the min-max normalization and TNFS mentioned above;
[0046] S8: Use the trained machine learning algorithm to judge leakage / false alarms, and upload the results to the server for the operator to perform actual processing.
[0047] The min-max normalization processing includes:
[0048]
[0049] where Xnorm represents the normalized frequency-domain data, X represents the input frequency-domain data, X min represents the minimum value in the frequency-domain data, X max represents the maximum value in the frequency-domain data.
[0050] In the acoustic system, false alarms caused by electrical noise are similar to transformer noise. The present invention proposes a TNFS algorithm, which utilizes this similarity relationship to effectively extract features from electrical noise. The present invention utilizes the transformer noise data measured in existing research. In order to improve the performance of the proposed noise classifier and ensure its applicability to low-specification IoT devices such as acoustic recorders, effective feature representation of high-dimensional frequency data is required. And from the information provided by the domain knowledge obtained through background noise analysis for feature representation, it can be seen that more than 70% of the electrical noise comes from transformer vibration. Based on this premise, an effective feature representation technique called TNFS (Transformer Noise-Based Frequency Selection) is proposed.
[0051] Based on the normalized frequency-domain data, extract the transformer noise in the electrical / mechanical noise interference signal through the TNFS algorithm, including:
[0052] By calculating the cosine similarity and Euclidean distance between the transformer noise and the electrical noise, the similarity in the frequency domain is judged. The smaller the Euclidean distance, the more similar the two frequencies are. The closer the cosine similarity is to 1, the more similar the two frequencies are. The more similar the frequencies between the transformer noise and the electrical noise are, the more difficult it is for the acoustic recorder to distinguish them. Then the TNFS algorithm is used for discrimination;
[0053] The Euclidean threshold and the cosine similarity threshold are respectively set to screen suitable electrical noise. The electrical noise with the cosine similarity between the transformer noise and the electrical noise greater than the set cosine similarity threshold and the Euclidean distance between the transformer noise and the electrical noise less than the Euclidean threshold is screened out. Here, the Euclidean threshold should be between 0.95 and 1.11, and the cosine similarity threshold should be between 0.56 and 0.68;
[0054] Through the TNFS algorithm, using the preset amplitude threshold, the electrical noise at high frequencies is used to indirectly represent the transformer noise.
[0055] Transformer noise TN (Transformer - noise) is an important factor leading to misjudgment in leakage detection. TN mainly comes from electromagnetic vibration, involving vibrating iron cores, magnetic deformation, windings, and structural components. The main reason is magnetostriction, which causes shape changes due to the influence of the magnetic field. Voltage and current respectively cause vibrations in the transformer iron core and windings. These induced vibrations will ultimately propagate to the transformer tank, resulting in noise interference. However, since there is no complete physical model of transformer noise, it makes the quantitative modeling of transformer noise a challenging task.
[0056] However, research has found that electrical noise EN (electrical noise) has similar frequency characteristics to TN in certain specific environments. Using the Euclidean distance formula and the cosine similarity formula, a connection can be established between EN and TN, making it possible to indirectly obtain TN by observing EN.
[0057] Formula (2) gives the Euclidean distance between the frequency domain signals of EN and TN after FFT. The smaller the Euclidean distance, the more similar the two frequencies are. Formula (3) is the cosine similarity formula between two signals. The closer the cosine similarity is to 1, the more similar the two frequencies are.
[0058]
[0059] Among them, C s (f,F) represents the cosine similarity between the transformer noise and the electrical noise, d(f,F) represents the Euclidean distance between the transformer noise and the electrical noise, N represents the number of characteristics of the transformer noise or the electrical noise, f represents the electrical noise, and F represents the transformer noise.
[0060] By using the TNFS algorithm and a preset amplitude threshold, the electrical noise at high frequencies is used to indirectly represent the transformer noise, including:
[0061] Input: Transformer noise data F. Output: Selected feature set O.
[0062] Step 1: Count the number of data in the input data F and denote it as n_F. Define a threshold epsilon between 0 and 1 to judge the amplitude level;
[0063] Step 2: Initialize the selected feature set O and the intermediate set D as empty sets respectively;
[0064] Step 3: In the outer loop, i traverses each noise data sample F_i from 1 to n_F;
[0065] Step 4: Normalize each sample F_i to the range of 0 to 1;
[0066] Step 5: In the inner loop, judge for the specific index value j of the normalized sample. If F_{ij} is greater than the threshold epsilon, add the index value j to the set D_i and add each sample's D_i set to D; j takes values of 0, 10, 20, …, 20000; where F_{ij} represents the j-th feature value of the i-th sample;
[0067] Step 6: Remove duplicate elements from the set D, merge all D_i sets, store the result in the set O, and output the finally selected feature set O to obtain the transformer noise.
[0068] For frequencies from 0 to 2000 Hz, the frequencies exceeding the amplitude level threshold (denoted as ∈) are stored in Di. The parameter ∈, as a hyperparameter, helps to select the main frequency and prevent over-selection of features. This process is iteratively executed for all data. After completion, the selected feature set O is obtained by eliminating the duplicates of the selected frequencies. In the invention, a value of ∈ = 0.1 is used, thus selecting 24 frequencies as features.
[0069] Due to the particularity of the pipeline leakage detection field itself, the obtained dataset often has problems such as sample scarcity or data imbalance, which may cause the network to overfit, thereby reducing the classification performance. This is one of the classic problems that need to be solved. To solve the above problems, a data augmentation method called GES is introduced as Figure 2 shown, which can well solve a thorny problem in GAN - mode collapse. Mode collapse means that when the generation network in GAN is over-optimized, it will cause the generated samples to have limited types, making it difficult for the adversarial network (discriminator) to escape the trap and thus resulting in mode collapse.
[0070] GES Workflow:
[0071] 1. Data Classification: The collected data is divided into majority classes and minority classes.
[0072] 2. Minority Class Data Augmentation: This process is initially carried out through GAN. In GAN, the generator is trained to generate data similar to real data from random noise vectors. On the other hand, the discriminator is trained to correctly classify data as generated data or original data. The generator and the discriminator participate in a competitive process and are trained using their respective loss functions. Among them, the random noise vector is randomly selected from the majority class and obtained by adding random noise; the data similar to real data is the minority class data; this can be interpreted as a model for solving the following min-max problem:
[0073]
[0074] where G represents the generator network, D represents the discriminator network, and V represents the value function used to measure the performance of the discriminator and the generator. denotes the expectation with respect to the sample x from the real data distribution px(x), denotes the expectation with respect to the noise z from the noise distribution pz(z), x represents the real data sample, and z represents the noise vector. px(x) represents the distribution of the real data, while pz(z) represents the distribution vector of the noise. The updates of the discriminator and the generator respectively use stochastic gradient ascent (5) and stochastic gradient descent (6).
[0075] logD w (x)+log(1-D w (G γ (z)))(5)
[0076] log(1-D w (G γ (z)))(6)
[0077] where Dw represents the discriminator with parameter w, and Gγ(z) represents the generator with parameter γ.
[0078] 3. Data Combination: After the initial augmentation using GAN, the original data and the generated data are combined through the process outlined in Algorithm 2. The generated data G is divided into n segments, each segment is called s, and all segments have equal size. The Wasserstein distance is used to quantitatively evaluate mode collapse, which is defined by the following equation:
[0079]
[0080] In this equation, W(p1, p2) represents the Wasserstein distance between two probability distributions p1 and p2. The goal is to find the infimum (the largest lower bound) of the expected cost c(x, y) over all possible joint distributions γ that align p1 and p2. When the two distributions are similar, the value of the Wasserstein distance is low, and it can be used to quantify overfitting and mode collapse. When C contains a large portion of the generated data, the distributions of C and s tend to become more similar, making it more likely for the value of W(C, s) to decrease.
[0081] Therefore, an algorithm that treats it as a pattern has been developed. As shown in Algorithm 2, if W(C, s) decreases to less than β times of σ [the initial W(O, s)], then it collapses. If no mode collapse is observed between the segments and the original data, they are combined into C.
[0082] After combining the data generated by GAN with the original data, SMOTE is applied to balance the majority-class data. This completes the GES data augmentation process, and by combining the augmented data with the original majority class, balanced data is obtained.
[0083] Oversampling is performed using the GES algorithm to achieve data augmentation, balance the dataset, and obtain a new dataset, including:
[0084] The data in the dataset is divided into the majority class and the minority class. Among them, the minority class represents the transformer noise that causes false alarms in the system, and the other data is used as the majority class;
[0085] By training the generator and discriminator in GAN, during the training process, the generator and discriminator participate in a competitive process to respectively generate data similar to real data from a random noise vector and correctly classify the data as generated data or original data. Through the generator and discriminator after training, the preliminary data augmentation of the minority class is completed;
[0086] After the discriminator and generator are trained, they use stochastic gradient ascent and stochastic gradient descent respectively
[0087] After the preliminary augmentation using GAN, the data in the dataset and the data generated by the preliminary data augmentation are combined through a mode collapse detection algorithm, and the Wasserstein distance is used to quantitatively evaluate the mode collapse;
[0088] After combining the data generated by GAN with the original data, SMOTE is used to balance the majority-class data, completing the data augmentation process of the GES algorithm. By combining the augmented data with the original majority class, balanced data is obtained, and a new dataset is obtained.
[0089] The mode collapse detection algorithm includes:
[0090] Input: original data O and generated data Q, Output: merged data C;
[0091] Step 1: Use the transformer noise as the original data O, and divide the generated data Q after preliminary data augmentation into n segments, where n > 2, and each segment is called s, and the sizes of all segments are equal;
[0092] Step 2: Set the threshold hyperparameter beta, and determine the maximum threshold sigma for mode collapse detection;
[0093] Step 3: Initialize the merged data set C as an empty set, and store the generated data Q divided into n segments in the set S;
[0094] Step 4: Loop through each segment, and randomly pop a segment s from the set S;
[0095] Step 5: When segment i = 1, calculate the initial threshold sigma (the Wasserstein distance between the original data O and segment s multiplied by beta), and initialize C as the original data O;
[0096] Step 6: If the current threshold sigma is greater than the metric W(C, s) between the merged data C and segment s, then skip this segment;
[0097] Step 7: If the skip condition is not met, then merge segment s into C, and then return the finally merged data C.
[0098] If n is set to a smaller value, resulting in larger segments, the probability of having similar patterns within a single segment will increase. Therefore, the number of combined segments increases. On the contrary, setting n to a larger value will result in smaller segments, and the probability of finding similar patterns in each segment will decrease. Therefore, the number of combined segments decreases. Therefore, in most cases, when n is set to a value that is neither too small nor too large, an appropriate number of data can be effectively combined.
[0099] In addition, the threshold σ for mode collapse detection is affected by the Wasserstein distance between the initial combined segments and the original data set. Therefore, the value of σ is determined relative to n.
[0100] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for identifying false alarms in water supply networks based on machine learning based on domain knowledge, characterized in that: include: Offline phase and online phase; The offline stage includes: S1: Use an acoustic sensor to collect the acoustic wave signal generated by the pipeline due to leakage or other reasons, and then perform fast Fourier transform FFT on the acquired acoustic wave signal to convert it into frequency domain data; The acoustic wave signals include: indoor leakage signals, outdoor leakage signals, interference signals, electrical / mechanical noise interference signals, life interference noise signals, and normal signals; among them, the transformer noise in the electrical / mechanical noise interference signal is close to the frequency of the electrical noise, and the system cannot identify the transformer noise, which is an important factor leading to the system's false alarm; S2: performing minimum-maximum normalization processing on the electrical / mechanical noise interference signal under the frequency domain data to obtain normalized frequency domain data; S3: Based on the normalized frequency domain data, the transformer noise in the electrical / mechanical noise interference signal is extracted by the TNFS algorithm, and the data set is formed with other sound wave signals; S4: Use the GES algorithm to perform oversampling to achieve data enhancement, complete the balance of the data set, and obtain a new data set; S5: Use the new dataset as training data for machine learning models to classify leakage sounds and interference noises and obtain preliminary results; The online stage includes: S6: The algorithm trained in the offline stage is embedded in the acoustic recorder, which monitors and collects pipeline leakage signals in real time and extracts abnormal signals; S7: The extracted abnormal signals are processed by the minimum-maximum normalization and TNFS mentioned above; S8: Use the trained machine learning algorithm to make leak / false alarm judgments, and upload the results to the server for actual processing by the operator.
2. According to the method of claim 1, the false alarm identification method of water supply network based on machine learning of domain knowledge is characterized in that: The minimum-maximum normalization process comprises: Among them, Xnorm represents the normalized frequency domain data, X represents the input frequency domain data, and X min Represents the minimum value in the frequency domain data, X max Represents the maximum value in the frequency domain data.
3. According to the method of claim 1, the false alarm identification method of water supply network based on machine learning of domain knowledge is characterized in that: Based on the normalized frequency domain data, the transformer noise in the electrical / mechanical noise interference signal is extracted by the TNFS algorithm, including: By calculating the cosine similarity and Euclidean distance between the transformer noise and the electrical noise, the similarity in the frequency domain is determined. The smaller the Euclidean distance, the more similar the two frequencies are. The closer the cosine similarity is to 1, the more similar the two frequencies are. The more similar the frequencies between the transformer noise and the electrical noise are, the more difficult it is for the acoustic recorder to distinguish them. The TNFS algorithm is then used for discrimination. The Euclidean threshold and the cosine similarity threshold are respectively set to screen suitable electrical noises, and the electrical noises whose cosine similarity between the transformer noise and the electrical noise is greater than the set cosine similarity threshold and whose Euclidean distance between the transformer noise and the electrical noise is less than the Euclidean threshold are screened out; The TNFS algorithm uses a preset amplitude threshold to filter out electrical noise to indirectly represent the transformer noise.
4. The method for identifying false alarms in water supply networks based on machine learning based on domain knowledge according to claim 3 is characterized in that: Calculate the cosine similarity between transformer noise and electrical noise, including: Among them, C s (f, F) represents the cosine similarity between transformer noise and electrical noise, N represents the number of features of transformer noise or electrical noise, f represents electrical noise, and F represents transformer noise.
5. The method for identifying false alarms in water supply networks based on machine learning based on domain knowledge according to claim 3 is characterized in that: Calculate the Euclidean distance between transformer noise and electrical noise, including: Wherein, d(f,F) represents the Euclidean distance between transformer noise and electrical noise, N represents the characteristic number of transformer noise or electrical noise, f represents electrical noise, and F represents transformer noise.
6. The method for identifying false alarms in water supply networks based on machine learning based on domain knowledge according to claim 3 is characterized in that: The TNFS algorithm uses a preset amplitude threshold to indirectly represent the transformer noise by using high-frequency electrical noise, including: Step 1: Input transformer noise data F, count the number of data in the input data F and record it as n_F, and define a threshold epsilon between 0 and 1 to judge the amplitude level; Step 2: Initialize the selected feature set O and the intermediate set D as empty sets respectively; Step 3: The outer loop from i 1 to n_F traverses each noise data sample F_i; Step 4: Normalize each sample F_i to the range of 0 to 1; Step 5: The inner loop judges the specific index value j of the normalized sample. If F_{ij} is greater than the threshold epsilon, the index value j is added to the D_i set and the D_i set of each sample is added to D. The value of j is 0, 10, 20, ..., 20000. F_{ij} represents the jth eigenvalue of the i-th sample. Step 6: Remove duplicate elements from set D and merge all D_i sets. Store the results in set O and output the final selected feature set O to obtain transformer noise.
7. The method for identifying false alarms in water supply networks based on machine learning based on domain knowledge according to claim 1 is characterized in that: Use the GES algorithm to perform oversampling to achieve data enhancement, balance the data set, and obtain a new data set, including: The data in the dataset is divided into majority class and minority class, where the minority class represents the transformer noise that causes the system to falsely report, and the other data is the majority class; By training the generator and discriminator in GAN, the generator and discriminator participate in the competition process during the training process to respectively generate data similar to real data from random noise vectors and correctly classify data as generated data or original data. The generator and discriminator after training are used to complete the preliminary data enhancement of minority classes. The discriminator and generator are trained using stochastic gradient ascent and stochastic gradient descent respectively. After the initial enhancement using GAN, the data in the dataset and the data generated by the initial data enhancement are combined through the mode collapse detection algorithm, and the Wasserstein distance is used to quantitatively evaluate the mode collapse; After combining the data generated by GAN with the original data, SMOTE is applied to balance the majority class data to complete the GES algorithm data enhancement process. The enhanced data is combined with the original majority class to obtain balanced data and a new data set.
8. The method for identifying false alarms in water supply networks based on machine learning based on domain knowledge according to claim 7 is characterized in that: By training the generator and discriminator in GAN, the generator and discriminator participate in the competition process during the training process to respectively generate data similar to real data from random noise vectors and correctly classify data as generated data or original data, including: Among them, G represents the generator, D represents the discriminator, and V represents the value function used to measure the performance of the discriminator and the generator. It means to find the expectation of sample x from the real data distribution px(x), represents the expectation of the noise z from the noise distribution pz(z), p x (x) represents the distribution of real data, p z (z) represents the distribution of noise data, x represents the real data sample, and z represents the noise vector.
9. The method for identifying false alarms in water supply networks based on machine learning based on domain knowledge according to claim 7, characterized in that: The mode collapse detection algorithm comprises: Input: original data O and generated data Q, output: merged data C; Step 1: Take the transformer noise as the original data O, and divide the generated data Q with preliminary data enhancement into n segments, n>2, each segment is called s, and the size of all segments is equal; Step 2: Set the threshold hyperparameter beta and determine the maximum threshold sigma for mode collapse detection; Step 3: Initialize the merged data set C as an empty set, and store the generated data Q after being divided into n segments in the set S; Step 4: Loop through each fragment and randomly pop a fragment s from the set S; Step 5: When segment i=1, the initial threshold sigma is calculated by multiplying the Wasserstein distance between the original data O and the segment s by beta, and C is initialized to the original data O; Step 6: If the current threshold sigma is greater than the metric W(C,s) of the merged data C and the segment s, skip the segment; Step 7: If the current threshold sigma is greater than the metric W(C,s) of the merged data C and fragment s, merge fragment s into C and return the final merged data C.
10. The method for identifying false alarms in water supply networks based on machine learning based on domain knowledge according to claim 1, characterized in that: Use trained machine learning algorithms to detect leaks and false alarms, including: Where N represents the number of acoustic recorders near the potential leak location, leak represents a leak, noise represents noise, and dn represents the label of the data received from the nth acoustic recorder.