Generative adversarial network dual-data representation-based natural gas pipeline audio anomaly detection method

By adopting a dual data representation method based on a generative adversarial network in the audio anomaly detection of natural gas pipelines, abnormal detection is decoupled and potential representation consistency losses are introduced, and the problems of excessive generalization and semantic feature in the prior art are solved, and the stability and generalization of abnormal detection are improved.

CN119993201AActive Publication Date: 2025-05-13CHANGZHOU UNIV

Patent Information

Application Number
CN202510059614.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The existing natural gas pipeline audio anomaly detection method based on generative adversarial networks has problems related to excessive generalization and inability to focus on normal data definition in actual industrial scenarios, resulting in poor abnormal detection effect.

Method used

Using a dual data representation method based on a generative adversarial network, anomaly detection model is constructed through a generator and a discriminator, anomaly detection decouples are carried out, and an improved potential representation consistency loss is introduced to prevent information leakage between the potential spaces.

Benefits of technology

It effectively solves the problem of over-generalization, promotes the model to eliminate information-free features from anomaly detection tasks in a type of learning environment, and improves the stability and generalization of anomaly detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993201A_ABST
    Figure CN119993201A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio anomaly detection, in particular to a natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation, which comprises the following steps: S1, collecting normal audio and abnormal audio of a natural gas pipeline under different working conditions, preprocessing the collected normal audio, and taking the abnormal audio as a test set; s2, constructing an anomaly detection model based on the generative adversarial network by using a generator and a discriminator; s3, performing anomaly detection decoupling, separating semantic information related to pipeline detection from global spectrum information, and eliminating interference of irrelevant information; s4, introducing improved potential representation consistency loss to prevent information leakage between potential spaces; and S5, using the improved anomaly detection model based on the generative adversarial network to detect the natural gas pipeline audio data set, and outputting a natural gas pipeline audio anomaly detection result. According to the method, the efficiency and accuracy of audio anomaly detection of the natural gas pipeline can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio anomaly detection, and in particular to a natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation. Background Art

[0002] As the artery of energy transportation, the safe and stable operation of oil and gas pipelines is crucial. Therefore, real-time anomaly detection of pipelines is required. Among the many anomaly detection technologies, the acoustic method is usually the preferred method for anomaly detection of natural gas pipelines due to its advantages of fast response time, high positioning accuracy and low cost.

[0003] Lu Jingyi et al. proposed a pipeline detection model based on integrated 1D-CNN-VAPSOSVM from the perspective of adaptive learning, and adaptively extracted features by studying one-dimensional convolutions of different depths. Yao Lizhong et al. proposed a pipeline anomaly detection model that integrates acoustic feature processing technology and feature reconstruction. A feature encoder (FAE) was introduced into the one-dimensional convolutional network to extract effective fault features while reconstructing local spatial features. Most of the existing pipeline audio anomaly detection methods combined with deep learning assume that there are sufficient abnormal samples, simple environmental noise, and a single working state; however, in actual industrial scenarios, due to the stable operation of the pipeline and the huge amount of monitoring data, abnormal data is scarce and the cost of data annotation is high. These factors will affect the effectiveness of acoustic detection. Therefore, natural gas pipeline anomaly detection based on a class of normal data has become a research hotspot.

[0004] At present, anomaly detection based on one-class data is mainly based on reconstruction-oriented models, relying on encoders or generative adversarial networks with outstanding reconstruction capabilities. Reconstruction-based methods have been shown to perform well in a wide range of anomaly detection tasks and also provide competitive results in natural gas pipeline audio anomaly detection. Gu Xiaohua et al. proposed an unsupervised adversarial domain adaptation method based on adversarial domain adaptation and one-class SVM. Through an adversarial learning strategy, the source domain and target domain data are aligned in an unsupervised way to detect abnormal sounds in industrial scenes under complex working conditions. Li Jiliang et al. proposed a generative adversarial network with multi-attention enhanced discriminator, integrating the attention mechanism of multiple dimensions into the discriminator to highlight the abnormal feature areas in the mechanical working audio spectrogram, thereby enhancing its ability to distinguish test samples from reconstructed samples. However, the existing learning objectives for training reconstruction models usually focus on low-level pixel comparisons rather than semantic comparisons intrinsically related to the training data, which leads to the potential representation encoding more low-level data features shared between normal and abnormal data samples rather than more discriminative high-level semantic features.

[0005] Furthermore, when trained on data with rich visual features and complex appearance, the likelihood of high-fidelity reconstruction of anomaly data increases, making reconstruction-based models less effective at detecting anomalies in such situations. This problem is further exacerbated by the high generalization capabilities of modern generative models, as high-quality reconstruction of anomaly samples can be expected under looser assumptions.

[0006] Therefore, the key challenge of anomaly detection methods based on generative adversarial networks is to learn latent representations that can encode important semantics and are insensitive to low-level visual features that are commonly shared by normal and abnormal data. These semantics are crucial for successfully describing normal data, helping to better distinguish normal image samples from abnormal samples and solving the over-generalization problem common in reconstruction-based anomaly detection models. Summary of the invention

[0007] The technical problem to be solved by the present invention is: in order to solve the problems existing in the prior art in the above-mentioned background technology, a natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation is provided.

[0008] The technical solution adopted by the present invention to solve the technical problem is: a natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation, comprising the following steps:

[0009] S1. Collect normal and abnormal audios of natural gas pipelines under different working conditions, pre-process the collected audios to meet the needs of model training, and use the abnormal audios as test sets;

[0010] S2. Use the generator and discriminator to build an anomaly detection model based on the generative adversarial network;

[0011] S3, perform anomaly detection decoupling, separate the semantic information related to pipeline detection from the global spectrum information, and eliminate the interference of irrelevant information;

[0012] S4, introduces an improved latent representation consistency loss to prevent information leakage between latent spaces;

[0013] S5. Use the improved anomaly detection model based on generative adversarial network to detect the natural gas pipeline audio dataset and output the natural gas pipeline audio anomaly detection results.

[0014] Furthermore, the sampling rate of the natural gas pipeline audio collected in step S1 is 16000, and the collected audio is preprocessed, specifically:

[0015] S11, dividing the collected audio into overlapping frames of 20 ms, applying a Hamming window function to each frame, applying a 2048-point short-time Fourier transform to each frame, overlapping each frame by 512 points, and obtaining an overlapping representation of each frame;

[0016] S12, mapping the short-time Fourier transform spectrum to the Mel scale, and using 128 Mel filters to obtain a Mel spectrum graph;

[0017] S13. Perform logarithmic transformation on the Mel spectrum in decibels to obtain a logarithmic Mel spectrum graph.

[0018] Furthermore, the anomaly detection model based on the generative adversarial network in step S2 consists of a generator, a discriminator, and an automatic encoder E m and a classifier C, the generator is a dual data representation generator, which consists of an encoder E s 、Encoder E r and decoder D, the encoder E s and encoder E r The architecture is the same and can learn different parameters independently;

[0019] Among them, the encoder E s and encoder E r They are composed of convolutional layers with a stride of 2, each of which is followed by a batch normalization layer and a LeakyReLU activation function with a negative slope of 0.2. The potential representation z is obtained through the dual data representation generator. s and the potential representation z r ; The decoder D performs decoding by transposed convolution with a stride of 2, each transposed convolution is followed by a batch normalization layer, a ReLU activation function, and the last convolution layer uses a tanh activation function to constrain the boundaries; the discriminator has the same s and encoder E r Same architecture, with the last convolutional layer followed by a sigmoid activation function; classifier C is a multilayer perceptron with 1 hidden layer and 30 hidden units. The size of the input layer is determined by the dimensionality of the potential representation, and the size of the output layer is determined by the number of classes in the training data; autoencoder E m With encoder E s The same architecture, network parameters and encoder E s same.

[0020] Furthermore, the classifier operation steps of the anomaly detection model based on the generative adversarial network are as follows:

[0021] S21. Train the classifier C based on the potential representation z s Predict the correct class label and use the latent representation z rThe training process controls the information content in the dual latent representation and learns meaningful data features for anomaly detection in the absence of abnormal data examples.

[0022] S22, use the gradient reversal layer to transform the potential representation z r Transform to ensure that the potential representation z r There is no useful semantic information for classification.

[0023] Furthermore, the abnormality detection decoupling in step S3 is specifically performed as follows:

[0024] S31, encoder E s Capture spectral semantic information, which is the key to characterizing normal training data; encoder E r Encode low-level residual features to ensure that the latent representation z s and the potential representation z r Ability to encode mutually independent information;

[0025] S32, using the decoupling process based on classifier C, the semantic information related to normal data is included in the potential representation z s and encode irrelevant residual information into the latent representation z r middle.

[0026] Furthermore, the improved latent representation consistency loss in step S4 is when the latent representation z r When the encoder E changes, s and encoder E m Extracting inconsistent semantically relevant information will cause the encoder E s Punishment, specifically:

[0027] S41, using cosine similarity to measure encoder E s and encoder E m The latent representation z s and the potential representation z′ s , focusing on the directions of the two to capture the local linear correlation, their cosine similarity CosSim(z s ,z′ s )The formula is expressed as follows:

[0028]

[0029] Among them, z s ·z′ s Yes s and z′ s The dot product of ||z s ||、||z′ s || are respectively ||zs || and ||z′ s ||'s modulus length;

[0030] S42. Improved latent representation consistency loss introduces an objective function based on mutual information improvement: First calculate z s and z′ s The joint density function p(z s ,z′ s ), the marginal probability density functions are p(z s ) and p(z′ s ), mutual information I(z s ; z′ s )The formula is expressed as follows:

[0031]

[0032] Then calculate z separately s and z′ s The entropy H(z s ) and H(z′ s ), according to the minimum entropy normalization formula, the mutual information is normalized to between [0,1];

[0033] S43, combining the cosine similarity of step S41 and the objective function based on mutual information improvement of step S42, an improved potential representation consistency loss is obtained, and its formula is expressed as follows:

[0034] L con = -α·NMI(z s ,z′ s )-(1-α)·CosSim(z s ,z′ s ).

[0035] Furthermore, the specific steps of step S5 are: uploading the audio data collected by the microphone collection array installed near the natural gas pipeline to the deep learning processor; the deep learning processor adopts the method of step S1 to pre-process the natural gas pipeline audio data, and converts the audio into a logarithmic Mel spectrum graph; the pre-processed data is input into the model for detection, and if an abnormality occurs in the natural gas pipeline, an alarm signal is issued. During this period, the microphone collection array will continue to collect signals to determine whether the abnormality has been eliminated. If it has been eliminated, the alarm will be stopped.

[0036] Beneficial effects of the present invention:

[0037] The gas pipeline audio anomaly detection model based on the generative adversarial network dual data representation of the present invention selectively encodes without the need for a pre-trained classification model, thereby solving the problems of over-generalization and failure to focus on semantic features related to the definition of normal data associated with the reconstruction of the anomaly detection model.

[0038] Construct an anomaly detection model based on generative adversarial networks to decouple anomaly detection, promote the model to exclude uninformative features from anomaly detection tasks in a class of learning environments, and separate uninformative data features and semantic features related to normal training data through two latent spaces; this anomaly detection decoupling strategy is also applicable to a class of classification problems in other fields;

[0039] An improved latent representation consistency loss is introduced to prevent information leakage between latent spaces, ensuring that the effective latent representation does not change even if the uninformative latent representation changes, improving the stability and generalization of anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0041] Figure 1 It is a step diagram of the natural gas pipeline audio anomaly detection method based on the generative adversarial network dual data representation of the present invention.

[0042] Figure 2 It is a flow chart of a natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation of the present invention.

[0043] Figure 3 This is a network architecture diagram of the natural gas pipeline audio anomaly detection method based on the generative adversarial network dual data representation of the present invention.

[0044] Figure 4 It is a specific flow chart of the abnormality detection decoupling method in the method of the present invention. DETAILED DESCRIPTION

[0045] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0046] like Figure 1 and Figure 2 As shown, a natural gas pipeline audio anomaly detection method based on a generative adversarial network is specifically as follows:

[0047] 1. Collecting audio signals from natural gas pipelines

[0048] Normal and abnormal audio signals of natural gas pipelines under different working conditions are collected. Normal audio is used as training data, and abnormal audio is used as test data. The normal audio signal is denoted as K = {K 1 ,K 2 ,K 3 ,…,K i …,K n}: K irepresents the i-th category of audio data, each category contains t audio data, and n represents the total number of categories collected; each recording is a single-channel audio with a length of 10s, including pipeline working sounds and environmental sounds.

[0049] 2. Preprocess the collected audio signals

[0050] (1) Framing: Divide the audio into overlapping frames of 20 ms and apply the Hamming window function to each frame; (2) Apply a 2048-point short-time Fourier transform to each frame, with 512 points overlapping each frame to obtain an overlapping representation of each frame; map the short-time Fourier transform spectrum to the Mel scale and use 128 Mel filters to obtain a Mel spectrum graph; (3) Perform a logarithmic transformation on the Mel spectrum in decibels to obtain a logarithmic Mel spectrum graph; select j types of normal spectra from the preprocessed logarithmic Mel spectrum graph as training data, and select the remaining normal spectra and abnormal spectra as mixed test data, where j < n.

[0051] 3. Use the generator and discriminator to build an anomaly detection network model. Its network architecture is as follows Figure 3 shown.

[0052] The anomaly detection model based on the generative adversarial network consists of a generator, a discriminator, and an autoencoder E m and a classifier C. The generator is a dual data representation generator, which consists of an encoder E s 、Encoder E r and decoder D, encoder E s and encoder E r The architecture of the two networks is the same and they can learn different parameters independently.

[0053] Among them, the encoder E s and encoder E r They are composed of convolutional layers with a stride of 2, each of which is followed by a batch normalization layer and a LeakyReLU activation function with a negative slope of 0.2. The potential representation z is obtained through the dual data representation generator. s and the potential representation z r The decoder D performs decoding by transposed convolutions with a stride of 2, each followed by a batch normalization layer, a ReLU activation function, and the last convolution layer uses a tanh activation function to constrain the boundaries. The classifier C is a multi-layer perceptron with 1 hidden layer and 30 hidden units. The size of the input layer is determined by the dimensionality of the potential representation, and the size of the output layer is determined by the number of classes in the training data; the autoencoder E m With encoder E s The same architecture and network parameters are the same as the encoder Es. The discriminator has the same s and encoder E rSame architecture, the last convolutional layer is followed by a Sigmoid activation function.

[0054] 4. Model training:

[0055] (a) Generator training: The generator part consists of the encoder E s and encoder E r , decoder D e , classifier C, encoder E s The input spectrum x∈R W×H×C Mapped to a semantically relevant latent representation z s , encoder E r Map the input spectrum x to the dual residual representation z r , the formula is as follows:

[0056] z s =E s (x) (1),

[0057] z r =E r (x) (2),

[0058] The complete latent representation of the input spectrum x is computed as the concatenation of the two partial representations, i.e. and passed to the decoder D e Reconstruction is performed, and the formula is expressed as follows:

[0059]

[0060] in, To reconstruct the spectrum; To ensure that all image content in the input spectrum x is captured by the concatenated representation z, the standard L is used when learning the parameters of the autoencoder 1 The reconstruction loss is expressed as follows:

[0061]

[0062] (b) Discriminator training: Discriminator D s It consists of an encoder that encodes the input through 5 convolutions, batch normalization, and LeakyReLU activation function, and finally uses the Softmax activation function to distinguish the input spectrum from the reconstructed spectrum; the potential representation z s and z r The expressive power of depends largely on the fidelity of image reconstruction; to improve the fidelity and ensure that the reconstructed samples follow the distribution of normal training data, an adversarial loss L is used during the training of the model. adv The encoder weights are updated based on the following feature matching objectives to reduce training instability and avoid GAN overtraining, as shown in the following formula:

[0063]

[0064] Where f(·) is D s The activation of the last convolutional layer; in addition, the discriminator distinguishes between real and fake spectra by optimizing the standard binary cross entropy loss, which is expressed as follows:

[0065]

[0066] Anomaly detection decoupling methods such as Figure 4 As shown, the latent representation z is forced to s and z r Encode mutually exclusive information and maximize z using cross entropy loss s The classification performance is expressed as follows:

[0067]

[0068] Where y∈R n is the one-hot encoded ground truth label of the input spectrum x, is the classifier prediction; the classifier C learns to accurately represent the semantically relevant latent representation z s Classification; during the back propagation process, the gradient of the classification loss passes through the corresponding (semantically related) encoder E s Propagation, encoder E s Update to produce a more semantically relevant and accurate latent representation z s ; Potential representation z s Provide more information to the classifier.

[0069] L s The minimization process will promote the potential representation z s Provide more information about normal data classification, so that z r To achieve the opposite effect, use the gradient reversal layer R to represent the dual residual z r Transform; the gradient reversal layer R acts as the identity function in the forward propagation of the model, that is, R(z r )=z r , but the gradients of subsequent layers are reversed during back-propagation, which is expressed as follows:

[0070]

[0071] Among them, λ R is a hyperparameter, I∈R d×d is the identity matrix, d is the potential representation z s And the dual residual represents z r During training, the encoder E rProduce a latent representation that is uninformative to the classifier C, and achieve z through the following goal r Minimize the relevant information content in , the formula is as follows:

[0072]

[0073] in,

[0074] Formula (10) and Formula (12) provide a first level of decoupling, but they cannot completely prevent the potential representation z s And the dual residual represents z r information leakage between them; introducing an improved latent representation consistency loss, if the encoder E s and encoder E m If the semantically relevant latent representations extracted from s and encoder E m To punish.

[0075] First, randomly shuffle the latent representation z of a given training batch r to ensure that each potential representation z s With the randomly selected z' r Concatenate; pass the concatenated potential representation to the decoder D e , decoder D e Generate hybrid reconstructed image Make Next, the reconstructed image is input to the encoder E m , extracting latent representation should be the same as the z originally used to generate the hybrid reconstructed image s Equal; using improved latent representation consistency loss penalty z s and The difference between them is expressed as follows:

[0076]

[0077] Among them, α is the weight parameter;

[0078] Cosine similarity is used to measure the encoder E s and the autoencoder E m The latent representation z s and It focuses on the direction of the two rather than their size, so it can capture the similarity between the two well without being affected by the length, capturing the local linear correlation, and their cosine similarity The formula is as follows:

[0079]

[0080] in, Yes s and The dot product of ||z s ||、 They are ||z s || and The mold length.

[0081] Calculate z s and The joint density function of The marginal probability density functions are p(z s )and Mutual Information The formula is as follows:

[0082]

[0083] Normalization is used to map the mutual information values ​​to a fixed range to enhance the comparability and scale consistency of different indicators. s and The entropy H(z s )and The formula is as follows:

[0084]

[0085] Among them, P(Z=z s ) is the potential representation to get the potential representation z s The mutual information is normalized to [0,1], and the formula is expressed as follows:

[0086]

[0087] When training the generator, the network weights of the discriminator are fixed, and the discriminator is made to identify the reconstructed spectrum as the real spectrum as much as possible to achieve the purpose of deceiving the discriminator. First, the loss terms of formula (7), formula (8), formula (10), formula (12) and formula (13) are calculated, and then the loss terms of the combined objective function L are calculated. G Update the weight of the generator, the formula is as follows:

[0088] L G =λ 1 L rec +λ 2 L adv +λ 3 L s +λ 4 L con (15)

[0089] Among them, λ 1 , 2 ,3 , 4 is the weight parameter.

[0090] When training the discriminator, the network weights of the generator are fixed, and formulas (7) and (9) are calculated according to the combined objective function L D Update, the formula is as follows:

[0091] L D =λ 1 L rec +λ 6 L bce (16)

[0092] Among them, λ 1 , 6 are weight parameters; during training, the generator and discriminator are updated alternately for a fixed number of rounds.

[0093] The experiment is implemented based on Python 3.7, PyTorch 1.5 and CUDA 10.2, and the hardware devices used are Intel Core i7-8700K CPU and NVIDIA GeForce RTX 2080Ti GPU. The Adam optimizer is used to minimize formula (18) and formula (19). The learning rate of the Adam optimizer is lr = 0.0002, the momentum β1 = 0.5, and the momentum β2 = 0.999. In the experiment, λ is taken 1 =λ 5 =50,λ 2 =λ 3 =λ 4 =λ 6 =1; the weight of the gradient reversal layer R is initialized to λ R =0, and then gradually updated as the training progresses; the training data of the simulated natural gas pipeline audio dataset is loaded into the improved generative adversarial anomaly detection network for training, and the model with the best effect is saved.

[0094] 5. Construct anomaly score indicator: input the detection sample into the generator to calculate the potential representation z s =E s (x), which is then fed into the classifier; the activation value of the classifier output layer is then Normalize it to satisfy the probability form, that is, The anomaly score formula is as follows:

[0095] s=1-p i (x) (17),

[0096] The normalization process generates anomaly scores s∈[0,1], where 0 represents ideal normal data; calculating anomaly scores only requires the encoder Es and classifier C, which significantly shortens the inference time.

[0097] 6. Deploy and train the optimal model for natural gas pipeline anomaly detection: The microphone acquisition array installed near the natural gas pipeline collects audio data and uploads it to the deep learning processor. The processor first uses the above method to preprocess the natural gas pipeline audio data and converts the audio into a logarithmic Mel spectrum graph; the preprocessed data is input into the model for detection. If an abnormality occurs in the natural gas pipeline and the obtained anomaly score exceeds the threshold, an alarm signal is issued. During this period, the microphone acquisition array will continue to collect signals to determine whether the anomaly has been eliminated. If it has been eliminated, the alarm will stop.

[0098] In order to evaluate the effectiveness of the above method for natural gas pipeline audio anomaly detection, the experiment uses the area under the ROC curve (AUC), partial area under the ROC curve (pAUC), accuracy, and F1 score to measure the effectiveness of the above method; the verification results are as follows:

[0099] Table 1 Experiments on simulated natural gas pipeline audio dataset

[0100]

[0101] Table 2 DCASE Challenge 2023 Task 2 dataset experiments

[0102]

[0103] In Table 1, AnoGAN is an anomaly detection algorithm based on generative adversarial networks (GAN); EfficientGAN is a method designed to improve the training efficiency and stability of generative adversarial networks (GAN); AEGAN (Auto-Encoder Generative Adversarial Networks) is a model that combines autoencoders and generative adversarial networks (GAN); MeSkipGANomaly is an improved model based on Skip-GANomaly for anomaly detection tasks. Skip-GANomaly itself is a model that combines autoencoders and generative adversarial networks (GAN), which introduces skip connections to better capture the normal data distribution at multiple scales. MeSkipGANomaly is further optimized and improved on this basis to improve the performance and efficiency of the model.

[0104] In Table 2, AE-GAN-AD is an unsupervised method for machine audio anomaly detection, combining the advantages of autoencoders and generative adversarial networks (GANs); the core of this method is to reconstruct the input spectrogram through a generator (also an autoencoder), and to assist the generator with a discriminator to improve performance in the training and detection stages. A new method for anomaly detection proposed by Fujimura et al. combines ResNeXt neural network, improved loss function SCAdaCos, and Gaussian mixture model (GMM) for anomaly judgment. A method based on subclustering AdaCos proposed by Kevin Wilkinghoff for abnormal sound detection under domain offset conditions.

[0105] It can be seen from Table 1 and Table 2 that the overall anomaly detection effect of the model in this embodiment is better than the existing method.

[0106] Based on the above ideal embodiments of the present invention, the relevant staff can make various changes and modifications without departing from the technical concept of the present invention through the above description. The technical scope of the present invention is not limited to the contents of the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. A natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation, characterized in that: The steps include: S1. Collect normal and abnormal audios of natural gas pipelines under different working conditions, pre-process the collected audios to meet the needs of model training, and use the abnormal audios as test sets; S2. Use the generator and discriminator to build an anomaly detection model based on the generative adversarial network; S3, perform anomaly detection decoupling, separate the semantic information related to pipeline detection from the global spectrum information, and eliminate the interference of irrelevant information; S4, introduces an improved latent representation consistency loss to prevent information leakage between latent spaces; S5. Use the improved anomaly detection model based on generative adversarial network to detect the natural gas pipeline audio dataset and output the natural gas pipeline audio anomaly detection results.

2. The natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation according to claim 1 is characterized by: The sampling rate of the natural gas pipeline audio in step S1 is 16000, and the collected audio is preprocessed as follows: S11, dividing the collected audio into overlapping frames of 20 ms, applying a Hamming window function to each frame, applying a 2048-point short-time Fourier transform to each frame, overlapping each frame by 512 points, and obtaining an overlapping representation of each frame; S12, mapping the short-time Fourier transform spectrum to the Mel scale, and using 128 Mel filters to obtain a Mel spectrum graph; S13. Perform logarithmic transformation on the Mel spectrum in decibels to obtain a logarithmic Mel spectrum graph.

3. The natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation according to claim 1 is characterized by: The anomaly detection model based on the generative adversarial network in step S2 consists of a generator, a discriminator, and an automatic encoder E m and a classifier C, the generator is a dual data representation generator, which consists of an encoder E s 、Encoder E r and decoder D, the encoder E s and encoder E r The architecture is the same and can learn different parameters independently; Among them, the encoder E s and encoder E r They are composed of convolutional layers with a stride of 2, each of which is followed by a batch normalization layer and a LeakyReLU activation function with a negative slope of 0.

2. The potential representation z is obtained through the dual data representation generator. s and the potential representation z r ; The decoder D performs decoding by transposed convolution with a stride of 2, each transposed convolution is followed by a batch normalization layer, a ReLU activation function, and the last convolution layer uses a tanh activation function to constrain the boundaries; the discriminator has the same s and encoder E r Same architecture, with the last convolutional layer followed by a sigmoid activation function; classifier C is a multilayer perceptron with 1 hidden layer and 30 hidden units. The size of the input layer is determined by the dimensionality of the potential representation, and the size of the output layer is determined by the number of classes in the training data; autoencoder E m With encoder E s The same architecture, network parameters and encoder E s same.

4. The natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation according to claim 3 is characterized by: The specific steps of running the classifier of the anomaly detection model based on the generative adversarial network are as follows: S21, train the classifier C according to the potential representation z s Predict the correct class label and use the latent representation z r The training process controls the information content in the dual latent representation and learns meaningful data features for anomaly detection in the absence of abnormal data examples. S22, use the gradient reversal layer to transform the potential representation z r Transform to ensure that the potential representation z r There is no useful semantic information for classification.

5. The natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation according to claim 3 is characterized by: The abnormality detection decoupling in step S3 is specifically performed as follows: S31, encoder E s Capture spectral semantic information, which is the key to characterizing normal training data; encoder E r Encode low-level residual features to ensure that the latent representation z s and the potential representation z r Ability to encode mutually independent information; S32, using the decoupling process based on classifier C, the semantic information related to normal data is included in the potential representation z s and encode irrelevant residual information into the latent representation z r middle.

6. The natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation according to claim 3 is characterized by: The improved latent representation consistency loss in step S4 is when the latent representation z r When the encoder E changes, s and encoder E m Extracting inconsistent semantically relevant information will cause the encoder E s Punishment, specifically: S41, using cosine similarity to measure encoder E s and encoder E m The latent representation z s and the potential representation z' s , focusing on the directions of the two to capture the local linear correlation, their cosine similarity CosSim(z s ,z' s )The formula is expressed as follows: Among them, z s ·z' s Yes s and z' s The dot product of ||z s ||、||z' s || are respectively ||z s || and ||z' s ||'s modulus length; S42. Improved latent representation consistency loss introduces an objective function based on mutual information improvement: First calculate z s and z' s The joint density function p(z s ,z' s ), the marginal probability density functions are p(z s ) and p(z' s ), mutual information I(z s ;z' s )The formula is expressed as follows: Then calculate z separately s and z' s The entropy H(z s ) and H(z' s ), according to the minimum entropy normalization formula, the mutual information is normalized to between [0,1]; S43, combining the cosine similarity of step S41 and the objective function based on mutual information improvement of step S42, an improved potential representation consistency loss is obtained, and its formula is expressed as follows: L con =-α NMI(z s ,With' s )-(1-α)·CosSim(z s ,With' s )。 7. The natural gas pipeline audio anomaly detection method based on generative adversarial network dual data representation according to claim 2 is characterized by: The specific steps of step S5 are: uploading the audio data collected by the microphone collection array installed near the natural gas pipeline to the deep learning processor; the deep learning processor uses the method of step S1 to pre-process the natural gas pipeline audio data and convert the audio into a logarithmic Mel spectrum graph; The preprocessed data is input into the model for detection. If an abnormality occurs in the natural gas pipeline, an alarm signal is issued. During this period, the microphone acquisition array will continue to collect signals to determine whether the abnormality has been eliminated. If it has been eliminated, the alarm will stop.

Citation Information

Patent Citations

  • Audio abnormality detection method based on confrontation network generation

    CN109461458A

  • Railway signal infrastructure data management method and system

    CN119046409A

  • Equipment anomaly detection method based on audio analysis

    CN119207463A

  • Defect-detecting device and defect-detecting method for an audio device

    US20220130411A1

Cited By

  • Escalator abnormal sound detection method and system based on domain invariant feature transfer and clustering

    CN122511299A