Rotating machinery air domain acoustic diagnosis method based on generative adversarial denoising

By using a generative adversarial noise reduction network to compress and decode the local time-frequency features of acoustic signals from rotating machinery, and combining a discriminator and an information fusion classification network, the diagnostic challenge of rotating machinery under complex background noise is solved, and efficient fault identification of rotating machinery is achieved.

CN121237122BActive Publication Date: 2026-03-31INST OF ACOUSTICS CHINA ACAD OF TESTING TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In real industrial scenarios, the radiated sound source signals of rotating machinery are masked or altered by complex background noise, resulting in the inaccurate transmission of equipment status information, increasing the difficulty of diagnosis or causing misdiagnosis.

Method used

A generative adversarial noise reduction method for acoustic diagnosis of rotating machinery in the air domain is adopted. The method uses an encoder to perform local time-frequency sensing mechanism segmentation and encoding feature compression, a decoder to perform hierarchical information interaction time-frequency decoding, a discriminator to perform local noise reduction, and finally an information fusion classification network to perform acoustic diagnosis.

Benefits of technology

It effectively filters out non-static background noise, accurately restores the acoustic signals of rotating machinery, improves diagnostic accuracy, and enables efficient identification of rotating machinery faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237122B_ABST
    Figure CN121237122B_ABST
Patent Text Reader

Abstract

The application provides a rotating machinery air domain acoustic diagnosis method based on generative adversarial denoising, relates to the technical field of mechanical fault diagnosis, and the method is characterized in that: an encoder is used to segment and encode feature compression of a collected two-dimensional time-frequency acoustic sample of rotating machinery, so that deep semantic compression encoding features are obtained; a decoder is used to perform hierarchical information interaction time-frequency decoding on the deep semantic compression encoding features, so that time-frequency information is restored and reconstructed; a discriminator is used to perform local denoising performance discrimination on the reconstructed time-frequency information, so that a denoised multi-channel two-dimensional time-frequency signal is obtained; an information fusion classification network is used to process the denoised multi-channel two-dimensional time-frequency signal, so that fusion features are obtained; a classification prediction network is used to perform classification prediction on the fusion features, so that a rotating machinery air domain acoustic diagnosis result is obtained. The application solves the problem of how to use an air domain acoustic diagnosis method to effectively identify the abnormal state of rotating machinery under the condition of non-static noise in a real industrial scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of mechanical fault diagnosis technology, and in particular to an acoustic diagnostic method for the air domain of rotating machinery based on generative adversarial noise reduction. Background Technology

[0002] The self-excited radiated noise and vibration signals of rotating machinery actuators are mutually causal, containing rich equipment status information. They possess clear physical meaning, are intuitive, and are easy to identify and use for decision-making. Thanks to the rapid development of signal processing technology, the knowledge transfer and quantitative expression of airborne acoustic signals have become feasible. This has highlighted the advantages of airborne acoustic diagnostic methods in predictive maintenance and health management of rotating machinery, including non-contact measurement, no impact on equipment operation, simple measuring instruments, and easy signal acquisition. These methods can address the limitations of contact-based sensing methods for abnormal performance during the complex dynamic service process of rotating machinery actuators, making them a highly promising non-contact, non-invasive fault diagnosis method in recent years, and attracting significant attention in the field of rotating machinery health management.

[0003] The core idea of ​​acoustic diagnostic technology for rotating machinery is to use the self-excited noise generated during the motion of the rotating machinery's actuators as the carrier of equipment status information, using air as the propagation medium, and employing air coupling to acquire the acoustic energy information radiated by the rotating machinery. Fault mode identification and anomaly warning are then achieved by analyzing the changes in radiated acoustic energy acquired during normal and abnormal operation. However, sound waves are highly susceptible to external environmental factors during propagation in air. In the case of rotating machinery, in actual industrial environments, the radiated sound source signal is often surrounded by complex background noise (knocking sounds, collision sounds, falling sounds, wear sounds of mechanical parts, worker voices, and their mixed noises) permeating the production workshop or workplace. These noise signals appear randomly, producing time-varying energy intensity and frequency distributions, mixed with different signal-to-noise ratios in each frame of the rotating machinery's radiated sound source signal, exhibiting non-static additivity. Under the interference of non-static background noise, the radiated acoustic energy of the rotating machinery is easily masked or altered, causing the equipment status information contained in the sound source signal to be unable to be accurately and effectively transmitted, thus increasing the difficulty of diagnosis or leading to misdiagnosis. Summary of the Invention

[0004] To address the aforementioned shortcomings in existing technologies, the present invention provides a method for air-domain acoustic diagnosis of rotating machinery based on generative adversarial noise reduction, which solves the problem of how to effectively identify abnormal states of rotating machinery using air-domain acoustic diagnosis methods under non-static noise conditions in real industrial scenarios.

[0005] To achieve the aforementioned objectives, the technical solution adopted by this invention is: a method for acoustic diagnosis of rotating machinery in the air domain based on generative adversarial noise reduction, comprising:

[0006] S1: Using the local time-frequency sensing mechanism of the encoder, the collected two-dimensional time-frequency acoustic samples of rotating machinery are segmented and encoded to compress features, resulting in deep semantic compressed encoding features.

[0007] S2: Using the feature aggregation and restoration module of the decoder, the deep semantic compression coding features are subjected to hierarchical information interaction time-frequency decoding, and the time-frequency spatial details of the sparse periodic acoustic signal are restored step by step from local to global, and the time-frequency information is restored and reconstructed.

[0008] S3: Using a discriminator, the reconstructed time-frequency information is subjected to equal-scale local noise reduction performance discrimination, and non-static noise components mixed in the reconstructed time-frequency information are tracked and suppressed to perform air domain acoustic noise reduction, thereby obtaining a noise-reduced multi-channel two-dimensional time-frequency signal; wherein, the encoder, the decoder and the discriminator belong to a generative adversarial noise reduction network, which is obtained through training;

[0009] S4: Using an information fusion classification network, the denoised multi-channel two-dimensional time-frequency signal is processed, and multi-channel time-frequency information is fused using depthwise separable convolution and grouped convolution to obtain fused features;

[0010] S5: Using a classification prediction network, perform acoustic diagnosis on the fused features to obtain acoustic diagnosis results for the air domain of rotating machinery, thus completing the acoustic diagnosis of rotating machinery faults.

[0011] The beneficial effects of the present invention are as follows: The present invention provides an acoustic diagnostic method for the air domain of rotating machinery based on generative adversarial denoising. By using a discriminator, the reconstructed time-frequency information is subjected to local denoising performance discrimination at the same scale, and non-static noise components mixed in the reconstructed time-frequency information are tracked and suppressed to perform air domain acoustic denoising, and a denoised multi-channel two-dimensional time-frequency signal is obtained. The denoised multi-channel two-dimensional time-frequency signal is processed by an information fusion classification network, and multi-channel time-frequency information is fused by depthwise separable convolution and group convolution to obtain fused features. The fused features are then subjected to acoustic diagnosis by a classification prediction network to obtain the acoustic diagnostic results for the air domain of rotating machinery. (1) By using a local time-frequency sensing mechanism, block-based correlation compression coding is carried out for medium-area block regions of multi-scale time-frequency features to improve the ability to perceive time-frequency details of sparse sound signals and obtain high-quality compressed coding features. (2) Using the feature aggregation and restoration module, the information transmission methods of skip connections and dense connections are integrated, and a time-frequency decoding method based on hierarchical information interaction is proposed. Skip connections are used to help fuse shallow detail features of encoding downsampling with deep semantic features of corresponding decoding upsampling, to restore time-frequency spatial resolution. Based on dense connections, a multi-level feature reuse combination is formed inside the decoder to enhance the diversity and richness of time-frequency features and improve the local detail restoration capability. A multi-level feature aggregation and restoration mechanism is formed under the upsampling framework to solve the fine-grained restoration problem of sparse signals. (3) Using a discriminator, a region discriminator is designed to traverse the restored signal features in the medium-scale local time-frequency region to evaluate the noise reduction performance and the quality of sparse periodic information restoration, and guide the generator (1, 2 above) to continuously optimize in reverse, so as to achieve the purpose of effectively estimating and filtering the time-varying statistical characteristics of non-static noise. Finally, the efficient filtering of non-static background noise in real industrial scenarios and the refined restoration output of acoustic signals of rotating machinery are realized, improving the accuracy of acoustic diagnosis.

[0012] Further, S1 includes:

[0013] The two-dimensional time-frequency acoustic samples of the rotating machinery were converted using a Mel filter bank to obtain the two-dimensional Mel time spectrum.

[0014] By using patch partitioning, the two-dimensional Mel-time spectrum is divided according to the frequency domain axis and the time domain axis to obtain a local time-frequency region matrix;

[0015] Based on the local time-frequency region matrix, corresponding information is extracted according to the similarity score, and the perceptual coding matrix in the time domain and frequency domain is obtained by segmentation coding feature compression.

[0016] Using a feedforward neural network, the perceptual coding matrices in the time domain and frequency domain are concatenated and reconstructed based on the original input dimension to obtain coded two-dimensional features;

[0017] The patch aggregation module is used to aggregate the encoded two-dimensional features to obtain deep semantic compression encoded features; wherein, the Mel filter bank, the patch partition, the local time-frequency region matrix, the feedforward neural network and the patch aggregation module belong to the encoder.

[0018] In this way, the function of the multi-level segmentation time-frequency coding module based on the feature pyramid downsampling framework can be realized, and the multi-level perception of local time-frequency information can be completed. (1) After the feature map is patched and partitioned, there is no need to divide the local window. Self-attention operation is carried out with the global block region as the object, which more directly associates local and global information. (2) The multi-level downsampling process of patch aggregation is executed at the pixel level of the feature map, i.e., the time-frequency unit, and is not executed at the patch level, i.e., the block region. It is also necessary to re-patch and partition the block region according to the downsampling compression coding feature scale, so as to promote the information fusion and interaction of adjacent time-frequency units. (3) Considering the time-frequency sparse structure characteristics implied by the two-dimensional feature map, a time-frequency block multi-head self-attention mechanism is constructed to enhance the ability to perceive details of time and frequency domain information and improve the feature compression coding quality in the multi-level downsampling architecture.

[0019] Furthermore, the expressions for the time-domain and frequency-domain perceptual coding matrices are as follows:

[0020] ;

[0021] ;

[0022] ;

[0023] ;

[0024] in, This represents the "query" matrix of the time-domain and frequency-domain multi-head self-attention matrices in the linear projection space. Represents the time and frequency multi-head self-attention matrix. express The corresponding linear projection space, Represents the set of real numbers. This represents the formed time-frequency block region. This indicates the embedded vector dimension in a linear projection. The "key" matrix represents the multi-head self-attention matrix in the time and frequency domains projected into the linear projection space. express The corresponding linear projection space, The matrix representing the "values" of the time-domain and frequency-domain multi-head self-attention matrices in the linear projection space. express The corresponding linear projection space, Represents the self-attention function. express function, To represent the transpose of a matrix, Indicates the relative encoding position.

[0025] Furthermore, the expression for the encoded two-dimensional feature is:

[0026] ;

[0027] ;

[0028] ;

[0029] ;

[0030] in, This represents the equivalent of a fully connected output with the initial dimension. This represents a fully connected layer equivalent to the initial dimension. This represents the time and frequency of the multi-head self-attention integrated output matrix. express, This represents the non-linear expression of the activation function, i.e., the output of a fully connected circuit. This represents a fully connected layer with a four-fold higher-dimensional mapping. This represents the time-frequency fusion sensing coding matrix. Indicates a combination stacking transformation. Represents the time-domain aware coding matrix. Represents the frequency domain sensing coding matrix. This represents the encoded two-dimensional output features. This means that the splicing and reconstruction transformation is to restore the original two-dimensional matrix dimension by splicing and combining the block regions according to the time and frequency axes.

[0031] Further, S2 includes:

[0032] The patch extension component of the feature aggregation and restoration module is used to extend the time-frequency resolution of the deep semantic compressed coding features to obtain upsampled features;

[0033] By using a skip connection method, the upsampled features and time-frequency compressed coding features are concatenated to obtain fused representation information;

[0034] The fused representation information is aggregated and restored using a hierarchical information interaction time-frequency decoding network to obtain initial reconstructed time-frequency information. The initial reconstructed time-frequency information is stacked with the fused representation information along the channel dimension through dense connections, and then subjected to hierarchical information interaction time-frequency decoding via a convolutional neural network. The time-frequency spatial details of the sparse periodic acoustic signal are restored step by step from local to global to obtain the reconstructed time-frequency information. The patch extension component, the skip connection method, and the hierarchical information interaction time-frequency decoding network belong to the feature aggregation and restoration module.

[0035] Furthermore, the loss function expression for training the generative adversarial denoising network is:

[0036] ;

[0037] ;

[0038] ;

[0039] ;

[0040] in, This represents the discriminant loss function. express, Indicates the discriminator, This represents the clean Mel-time spectrum unaffected by non-static noise. Indicates the first real data label, express, This indicates that the generator, i.e., the encoding-decoding mechanism, restores and recovers the output. This represents the mixed Mel-time spectrum affected by non-static noise. Indicates the second real data label, This represents the generation loss function. Indicates the total number of samples. Indicates the first A clean frequency, Indicates the first One restored sample, This represents the compressed encoding features of a clean sample with a dimension of 5×4. These represent the compressed coding features of noisy samples of the same dimension. This represents the magnitude of the eigenvector.

[0041] Furthermore, the overall objective function expression of the generative adversarial denoising network is:

[0042] ;

[0043] ;

[0044] in, express, express, This represents the discriminant loss function. express, This represents the mean squared error loss. This represents the generation loss function. This represents the cosine similarity loss. Attached Figure Description

[0045] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0046] Figure 1 This is an exemplary flowchart of a rotating machinery air domain acoustic diagnostic method based on generative adversarial noise reduction, as shown in some embodiments of this specification.

[0047] Figure 2 This is an exemplary schematic diagram illustrating the ablation visualization analysis of time-frequency noise reduction performance according to some embodiments of this specification;

[0048] Figure 3 This is an exemplary schematic diagram illustrating a visual comparison analysis of time-frequency noise reduction performance based on some embodiments of this specification. Detailed Implementation

[0049] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0050] Example

[0051] Figure 1 This is an exemplary flowchart illustrating a method for acoustic diagnostics of rotating machinery in the air domain based on generative adversarial noise reduction, according to some embodiments of this specification. Figure 1 As shown, the process includes the following steps. In some embodiments, the process may be executed by a processor.

[0052] S1: Using the local time-frequency sensing mechanism of the encoder, the collected two-dimensional time-frequency acoustic samples of rotating machinery are segmented and encoded to compress features, resulting in deep semantic compressed encoding features.

[0053] The local time-frequency sensing mechanism is a mechanism that performs local time-frequency sensing depth processing on two-dimensional time-frequency acoustic samples of rotating machinery. For example, the local time-frequency sensing mechanism may include a feature pyramid downsampling multi-level segmentation coding module.

[0054] In some embodiments, the processor can utilize a local time-frequency sensing mechanism to construct a multi-level segmentation time-frequency coding module based on a feature pyramid downsampling framework, which is used to perform segmentation coding feature compression on the collected two-dimensional time-frequency acoustic samples of rotating machinery to obtain deep semantic compression coding features.

[0055] Two-dimensional time-frequency acoustic samples of rotating machinery are sound samples generated during the operation of rotating machinery. For example, two-dimensional time-frequency acoustic samples of rotating machinery can include radiated sound and non-static background noise. Specifically, two-dimensional time-frequency acoustic samples of rotating machinery can include datasets covering 12 combination modes of different operating conditions and health states of rotating machinery, with a balanced number of categories. Each dataset contains 18,000 four-channel training samples and an equivalent number of 3,600 four-channel validation and test samples.

[0056] In some embodiments, the processor can symmetrically install four identical microphones around the rotating machinery using a spherical array frame to construct a microphone array. The processor can then use a data acquisition unit to collect the time-domain acoustic signals radiated by the rotating machinery in real time at a sampling rate of 16kHz. The collected time-domain acoustic signals are divided into data samples in seconds, i.e., each second of time-domain signal is a two-dimensional time-frequency acoustic sample of the rotating machinery.

[0057] Deep semantic compression coding features are features resulting from multi-level deep semantic encoding. For example, deep semantic compression coding features can include first-level deep semantic compression coding features, second-level deep semantic compression coding features, and third-level deep semantic compression coding features.

[0058] In some embodiments, the processor can split each deep compressed time-frequency unit in the matrix into a finer-grained region of adjacent 2×2 (a total of 4 sub-blocks) through spatial rearrangement operations, and then use a channel compression mechanism to balance computational load and information preservation, thereby constructing a new 10×8×1 dimension matrix with doubled resolution and restored channel number, achieving efficient and adaptive upsampling, forming a higher resolution feature map, and obtaining compressed coded features.

[0059] In some embodiments, the processor can utilize a Mel filter bank to convert the acquired two-dimensional time-frequency acoustic samples of rotating machinery to obtain a two-dimensional Mel time-frequency spectrum; utilize patch partitioning to evenly divide the two-dimensional Mel time-frequency spectrum according to the frequency domain axis and the time domain axis to obtain a local time-frequency region matrix; based on the local time-frequency region matrix, extract corresponding information according to similarity scores, and obtain a perceptual coding matrix in the time domain and frequency domain by segmentation coding feature compression; utilize a feedforward neural network to concatenate and reconstruct the perceptual coding matrix in the time domain and frequency domain based on the original input dimension to obtain coded two-dimensional features; utilize a patch aggregation module to aggregate the coded two-dimensional features to obtain downsampled deep semantic compression coding features; wherein, the Mel filter bank, the patch partitioning, the local time-frequency region matrix, the feedforward neural network, and the patch aggregation module belong to the encoder.

[0060] The two-dimensional Mel-time spectrum is the characteristic form of the two-dimensional Mel-time spectrum of a two-dimensional time-frequency acoustic sample of rotating machinery.

[0061] In some embodiments, the processor can perform frame-by-frame windowing of the signal for each data sample using a 64ms Hamming window with a 50% overlap rate, and apply a Mel filter bank consisting of 40 filters to each frame of the signal to form a 40-dimensional vector output. Thus, each 1-second two-dimensional time-frequency acoustic sample of rotating machinery is converted into a 40×32 two-dimensional Mel time-frequency spectrum, where 40 represents the number of output features of the Mel filter bank and 32 represents the frame number. Subsequently, the Mel time-frequency spectrum is normalized, its amplitude is scaled and converted to the [-1,1] interval to obtain the two-dimensional Mel time-frequency spectrum.

[0062] Patch partitions are partitions used to divide regions according to time and frequency axes.

[0063] The local time-frequency region matrix is ​​a matrix that reflects the information of a local region in the two-dimensional Mel-time spectrum.

[0064] In some embodiments, the processor can use patch partitioning to evenly divide the 40×32 two-dimensional time spectrum along the frequency domain axis and the time domain axis to form 4×4 time-frequency region blocks (Patch) with equal information scale, thereby realizing information deconstruction from "global" to "local"; based on the time-frequency region blocks, 16 10×8-dimensional local time-frequency region matrices are constructed.

[0065] The perceptual coding matrix in the time and frequency domains is the encoded local time-frequency region matrix.

[0066] In some embodiments, the processor can construct an 8×16×10 temporal multi-head self-attention matrix and a 10×16×8 frequency multi-head self-attention matrix at both time and frequency scales. The first dimension of the matrix represents the temporal or frequency domain resolution, i.e., the meaning of "multi-head," the second dimension represents the divided local time-frequency region blocks, and the third dimension represents the amount of frequency or temporal information carried by each region block in the temporal or frequency multi-head self-attention matrix, forming a temporal and frequency domain awareness mechanism. The temporal and frequency multi-head self-attention matrices are projected onto three different vector spaces through a fully connected layer to obtain a "query" matrix, a "key" matrix, and a "value" matrix of equal dimensions, denoted by Q, K, and V, respectively. Considering that multi-head self-attention requires sufficient dimensions to segment different attention heads, and that higher-dimensional spaces can encode more complex temporal and frequency-domain local patterns, enhancing feature representation to improve detail perception, an embedding dimension design is used during the linear projection process of the fully connected layer to expand the original feature dimension to a three-fold higher-dimensional feature space, resulting in the final feature embedding vector used for time-frequency perception computation. Thus, the Q, K, and V matrix dimensions for temporal multi-head self-attention are 8×16×30, while the matrix dimension for frequency multi-head self-attention is 10×16×24. Subsequently, in temporal and frequency multi-head self-attention, the response relationship between each block position of the "query" matrix and all block positions of the "key" matrix is ​​calculated through the dot product of the "query" matrix and the transpose of the "key" matrix, forming the self-attention weights. The dot product result is scaled by the square root of the embedding vector dimension to control the variance of matrix elements, maintain gradient stability, and avoid gradient vanishing or exploding problems. Simultaneously, a block position bias representation is introduced to form the relative position encoding of the global block region. Finally, the weight coefficients are normalized using the SoftMax activation function to obtain similarity scores that describe the correlation between the positions of each block in the "query" matrix and the information in the "key" matrix. The similarity matrix and the "value" matrix are then subjected to a dot product operation. The corresponding information is extracted according to the similarity measurement relationship between the block regions, i.e., the similarity scores, to obtain the perceptual coding matrices in the time domain and frequency domain.

[0067] In some embodiments, the expressions for the time-domain and frequency-domain perceptual coding matrices can be:

[0068] ;

[0069] ;

[0070] ;

[0071] ;

[0072] in, This represents the "query" matrix of the time-domain and frequency-domain multi-head self-attention matrices in the linear projection space. Represents the time and frequency multi-head self-attention matrix. express The corresponding linear projection space, Represents the set of real numbers. This represents the formed time-frequency block region. This indicates the embedded vector dimension in a linear projection. The "key" matrix represents the multi-head self-attention matrix in the time and frequency domains projected into the linear projection space. express The corresponding linear projection space, The matrix representing the "values" of the time-domain and frequency-domain multi-head self-attention matrices in the linear projection space. express The corresponding linear projection space, Represents the self-attention function. express function, To represent the transpose of a matrix, Indicates the relative encoding position.

[0073] Encoding two-dimensional features is a perceptual encoded feature with the same dimension as the original input.

[0074] In some embodiments, the processor can integrate the outputs of all heads in the time and frequency multi-head self-attention to form a 16×240 dimension time-domain and frequency-domain perceptual matrix. The input fully connected layer restores the linear projection-expanded feature embedding dimension to the initial dimension of 16×80. Subsequently, using a feedforward neural network, the embedding vector dimension is mapped to a four-fold higher-dimensional space through a two-layer fully connected structure, and then the dimension is restored to form the time-domain and frequency-domain perceptual coding representation of the original acoustic signal. The time-domain and frequency-domain representation matrices are then stacked along the embedding vector dimension, and the time-frequency coding information is fused using a fully connected layer to form a time-frequency fused perceptual coding output. The time-frequency block embedding information is then concatenated and reconstructed according to the original input dimension to restore a 40×32 two-dimensional matrix, thus obtaining the coded two-dimensional feature.

[0075] In some embodiments, the expression for encoding two-dimensional features can be:

[0076] ;

[0077] ;

[0078] ;

[0079] ;

[0080] in, This represents the equivalent of a fully connected output with the initial dimension. This represents a fully connected layer equivalent to the initial dimension. This represents the time and frequency of the multi-head self-attention integrated output matrix. express, This represents the non-linear expression of the activation function, i.e., the output of a fully connected circuit. This represents a fully connected layer with a four-fold higher-dimensional mapping. This represents the time-frequency fusion sensing coding matrix. Indicates a combination stacking transformation. Represents the time-domain aware coding matrix. Represents the frequency domain sensing coding matrix. This represents the encoded two-dimensional output features. This means that the splicing and reconstruction transformation is to restore the original two-dimensional matrix dimension by splicing and combining the block regions according to the time and frequency axes.

[0081] The patch aggregation module is a module that aggregates features along the time axis and frequency axis.

[0082] In some embodiments, the processor can use a patch aggregation module to process the encoded two-dimensional features, extract the encoded features at intervals of 2 time-frequency units along the time axis (row direction) and frequency axis (column direction), and splice and combine them into a multi-channel feature with a dimension of 20×16×4. The multi-channel information is aggregated using a convolutional neural network with a receptive field of 2×2 and a stride of 1×1, and the output is a downsampled compressed feature with a dimension of 20×16, forming a deep expression of the time-frequency features, and obtaining the first-level deep semantic compressed encoded features.

[0083] In some embodiments, the processor can use 20×16 downsampled compressed features as the second-level perceptual input, and after the above-mentioned patch partitioning, time-frequency block multi-head self-attention mechanism construction, time-frequency fusion perceptual coding and patch merging processes, output downsampled compressed features with a dimension of 10×8, complete the second-level perceptual compressed coding of local time-frequency information, and obtain the second-level deep semantic compressed coding features.

[0084] In some embodiments, the processor can use a 4×4 time-frequency region partitioning method for the patch partition, with the time and frequency multi-head self-attention matrices being 4×16×5 and 5×16×4 respectively. Using 10×8 downsampled features as the third-level perceptual input, and using a 2×2 patch partition, the processor can obtain 5×4 deep perceptual compressed coding features by following the same processing procedure as in the above embodiments, thus obtaining the third-level deep semantic compressed coding features.

[0085] In this way, the function of the multi-level segmentation time-frequency coding module based on the feature pyramid downsampling framework can be realized, and the multi-level perception of local time-frequency information can be completed. (1) After the feature map is patched and partitioned, there is no need to divide the local window. Self-attention operation is carried out with the global block region as the object, which more directly associates local and global information. (2) The multi-level downsampling process of patch aggregation is executed at the pixel level of the feature map, i.e., the time-frequency unit, and is not executed at the patch level, i.e., the block region. It is also necessary to re-patch and partition the block region according to the downsampling compression coding feature scale, so as to promote the information fusion and interaction of adjacent time-frequency units. (3) Considering the time-frequency sparse structure characteristics implied by the two-dimensional feature map, a time-frequency block multi-head self-attention mechanism is constructed to enhance the ability to perceive details of time and frequency domain information and improve the feature compression coding quality in the multi-level downsampling architecture.

[0086] S2: Using the feature aggregation and restoration module of the decoder, the deep semantic compression coding features are subjected to hierarchical information interaction time-frequency decoding, and the time-frequency spatial details of the sparse periodic sound signal are restored step by step from local to global, thus restoring and reconstructing the time-frequency information.

[0087] A multi-level feature aggregation and restoration mechanism is used to decode and restore features at multiple levels. For example, a multi-level feature aggregation and restoration mechanism may include patch extension components, skip connection methods, and hierarchical information interaction time-frequency decoding networks; the feature aggregation and restoration module can be divided into a first feature aggregation and restoration module, a second feature aggregation and restoration module, and a third feature aggregation and restoration module for deep semantic compression coding features at different levels.

[0088] Reconstructing time-frequency information is the initial decoded representation of time-frequency spatial detail information.

[0089] In some embodiments, the processor can utilize the patch extension component of the feature aggregation and restoration module to extend the time-frequency resolution of the deep semantic compressed coding features to obtain upsampled features; use a skip connection method to concatenate the upsampled features and the time-frequency compressed coding features to obtain fused representation information; use a hierarchical information interaction time-frequency decoding network to aggregate and restore the fused representation information to obtain initial reconstructed time-frequency information; the initial reconstructed time-frequency information is stacked with the fused representation information along the channel dimension through a dense connection method, and through a convolutional neural network, hierarchical information interaction time-frequency decoding is performed to restore the time-frequency spatial details of the sparse periodic acoustic signal from local to global level to obtain reconstructed time-frequency information.

[0090] The patch expanding module in the feature aggregation and restoration module is used to expand the number of channels for features.

[0091] Upsampling features are deep semantic compression coding features resulting from channel expansion.

[0092] In some embodiments, the processor can apply a linear layer, i.e. a fully connected layer, to the input deep semantic compression coding features based on the feature aggregation and restoration module using patch expansion, expanding the number of channels to four times the original dimension, forming a 5×4×4 dimension matrix. Then, through a channel compression mechanism, the computational load and information retention are balanced, and a new 10×8×1 dimension matrix is ​​constructed in a way that doubles the resolution and restores the number of channels, thus obtaining the upsampled features.

[0093] Fusion representation information is the concatenation of upsampled features and time-frequency compressed coding features. For example, fusion representation information can compensate for the loss of time-frequency details in compressed coding features during downsampling.

[0094] In some embodiments, the processor can use skip connections to concatenate the third-level compressed coding features of local time-frequency information with the upsampled features after patch expanding in the first-level feature aggregation and restoration module along the channel dimension to obtain fused representation information.

[0095] Hierarchical information interaction time-frequency decoding networks are used to aggregate and restore fused representation information. For example, hierarchical information interaction time-frequency decoding networks may include patch partitioning, multi-head self-attention mechanisms, common modules of feedforward neural networks, and private modules for hierarchical information interaction fusion.

[0096] In some embodiments, the processor can aggregate and restore the fused representation information based on a hierarchical information interaction time-frequency decoding network. Using patch partitioning, the 10×8×2 input fused representation features are divided into 2×2 time-frequency region blocks of equal information scale. Four 5×4×2-dimensional local time-frequency regions are established and expanded along the time-frequency dimension, then embedded and integrated into a 4×20×2 matrix, which is then input into a multi-head self-attention mechanism. Subsequently, a multi-head self-attention matrix with four heads is constructed in the multi-head self-attention mechanism, transforming the input matrix dimension to 4×4×10. This matrix is ​​then projected to three different vector spaces using a fully connected layer with linear mapping of the embedding dimension, forming a "query" matrix, a "key" matrix, and a "value" matrix. Standard multi-head self-attention operations are performed, and after head output integration and linear mapping of the fully connected layer to restore the embedding dimension, a 4×40 output matrix is ​​obtained and fed into a feedforward neural network. In the feedforward neural network, a two-layer fully connected structure is used to first map the embedding vector dimension to a four-fold higher-dimensional space and then compress it to half of the original input dimension to aggregate feature information. Based on this, the time-frequency segmented embedded information is spliced ​​and reconstructed according to the upsampled time-frequency feature structure to form a 10×8×1 dimension matrix output. Finally, dense connections are used to splice and reconstruct the features along the channel dimension by the upsampled fused representation feature information and the feedforward neural network output. The channel stacked information is then aggregated through a convolutional neural network with a receptive field of 3×3 and a stride of 1×1 to achieve the interactive fusion of hierarchical information, outputting a 10×8 dimension decoded feature, thus obtaining the primary aggregated restored feature.

[0097] In some embodiments, the processor can, within an upsampling framework, use 10×8 primary aggregated and restored features as input to the second-level feature aggregation and restoration module. Following the same patch expanding, patch partitioning, multi-head self-attention mechanism construction, feedforward neural network mapping, and hierarchical information interaction fusion processes, it outputs secondary aggregated and restored features with a dimension of 20×16, completing the secondary decoding representation of time-frequency spatial detail information. The patch partitioning uses a 4×4 time-frequency region division method, and the number of heads in the multi-head self-attention mechanism is maintained at 4, resulting in a multi-head self-attention matrix with a dimension of 4×16×10. Subsequently, the 20×16 secondary aggregated and restored features are used as input to the third-level feature aggregation and restoration module. Again, a 4×4 patch partitioning method is used, and the multi-head self-attention mechanism is set with 8 heads. The above process is repeated without retaining the hierarchical information interaction fusion module, ultimately restoring a 40×32 Mel-time spectrum with the same dimension as the original input time-frequency signal. This achieves a step-by-step restoration of the time-frequency spatial detail information of the sparse periodic acoustic signal from local to global, obtaining the reconstructed time-frequency information.

[0098] S3: Using a discriminator, the reconstructed time-frequency information is subjected to equal-scale local noise reduction performance discrimination, and non-static noise components mixed in the reconstructed time-frequency information are tracked and suppressed to perform air domain acoustic noise reduction, thereby obtaining a noise-reduced multi-channel two-dimensional time-frequency signal; wherein, the encoder, the decoder and the discriminator belong to a generative adversarial noise reduction network, which is obtained through training.

[0099] The discriminator is used to identify and estimate non-static noise components, guiding the generator to continuously optimize in reverse and generate near-clean acoustic samples. For example, the discriminator may include a 4-layer convolutional module; where each convolutional module includes a convolutional layer, a batch normalization layer (BN layer), and an activation layer.

[0100] The denoised multi-channel two-dimensional time-frequency signal is an anti-generated multi-channel two-dimensional time-frequency signal without non-static noise components.

[0101] In some embodiments, the processor can input time-frequency reconstruction information into the discriminator, which is then processed by two convolutional modules with a receptive field of 4×4, a stride of 2×2, and a Leak ReLU activation function, with the activation function remaining unchanged, a single convolutional module with a receptive field of 2×2 and a stride of 2×2, and an output convolutional module with a receptive field of 1×1 and a stride of 1×1, to obtain an output discriminant matrix with a dimension of 5×4×1, thereby obtaining the discrimination results of the noise reduction performance of local regions at various scales.

[0102] In some embodiments, the processor can pass the target sound source signal, which is not affected by non-static background noise, through a discriminator in the same manner to obtain the true label result. The discriminator estimates the result, and the true label is input into a least-squares generative adversarial network (LEAN) computation framework to measure the error between the generated denoised samples and the true samples. The measurement result guides the generator (encoder-decoder mechanism) to continuously optimize in reverse, completing the training of the generative adversarial denoising network. Specifically, the processor can take 18,000 four-channel training data points containing non-static background noise collected from real industrial scenarios as input, expand them along the channels to form 72,000 single-channel data points, perform frame-segmentation and windowing preprocessing, and Mel-time-frequency conversion to construct a two-dimensional time-frequency signal. This signal is then fed into the generator of the generative adversarial denoising network. Through a local time-frequency sensing mechanism and a multi-level feature aggregation and restoration mechanism, the restored Mel-time spectrum (i.e., the reconstructed time-frequency information) is output. Simultaneously, 18,000 identical four-channel training data collected in a semi-anechoic environment are expanded along the channels to form 72,000 clean label samples. These, along with the restored Mel-time spectrum, are input into the discriminator of the generative adversarial denoising network. Under the optimization framework composed of local least squares generative adversarial, mean square error, and cosine similarity, the generator is guided to carry out supervised learning, continuously improving the quality of sparse periodic signal restoration, i.e., the generation effect, until convergence is achieved, i.e., the difference between the denoised and restored samples and the clean label samples is minimal under the evaluation metrics of mean square error and peak signal-to-noise ratio, thus completing the training of the generative adversarial denoising network.

[0103] In some embodiments, the loss function expression for training a generative adversarial denoising network can be:

[0104] ;

[0105] ;

[0106] ;

[0107] ;

[0108] in, This represents the discriminant loss function. express, Indicates the discriminator, This represents the clean Mel-time spectrum unaffected by non-static noise. Indicates the first real data label, express, This indicates that the generator, i.e., the encoding-decoding mechanism, restores and recovers the output. This represents the mixed Mel-time spectrum affected by non-static noise. Indicates the second real data label, This represents the generation loss function. Indicates the total number of samples. Indicates the first A clean frequency, Indicates the first One restored sample, This represents the compressed encoding features of a clean sample with a dimension of 5×4. These represent the compressed coding features of noisy samples of the same dimension. This represents the magnitude of the eigenvector.

[0109] In some embodiments, to make the objective function have a smoother optimization path and reduce oscillations during the optimization process, the processor uses a label smoothing strategy to set the parameter to 0.9. This represents the second real data label, and is set to 0 according to the LSGAN standard.

[0110] In some embodiments, the overall objective function expression of a generative adversarial denoising network can be:

[0111] ;

[0112] ;

[0113] in, express, express, This represents the discriminant loss function. express, This represents the mean squared error loss. This represents the generation loss function. This represents the cosine similarity loss.

[0114] S4: Using an information fusion classification network, the denoised multi-channel two-dimensional time-frequency signal is processed, and multi-channel time-frequency information is fused using depthwise separable convolution and grouped convolution to obtain fused features.

[0115] Information fusion classification network is a network used to perform feature fusion on denoised multi-channel two-dimensional time-frequency signals.

[0116] Fusion features are the features of multi-channel two-dimensional time-frequency signals after noise reduction and information fusion.

[0117] In some embodiments, the processor can fuse the noise-reduced multi-channel two-dimensional time-frequency signal input information into a classification network, and use a depthwise separable convolution with a receptive field of 3×3, a stride of 1×1, and a channel expansion of 8, and a grouped convolution with a receptive field of 1×1, a stride of 1×1, a group number of 4, and a channel expansion of 16 to sequentially apply to the input 40×32×4 three-dimensional matrix to obtain a 40×32×64 fused feature.

[0118] S5: Using a classification prediction network, perform acoustic diagnosis on the fused features to obtain acoustic diagnosis results for the air domain of rotating machinery, thus completing the acoustic diagnosis of rotating machinery faults.

[0119] Classification prediction networks are used for acoustic diagnostics in the aerodynamic domain of rotating machinery. For example, a classification prediction network may include a head convolutional module, a multi-scale convolutional module, a tail convolutional module, and an output module; each module includes convolutional layers, batch normalization (BN) layers, and activation layers.

[0120] In some embodiments, the processor can use 18,000 four-channel training data samples collected in a semi-anechoic environment as input to the classification prediction network, and train the classification prediction network under a supervised learning framework with known fault category labels. Performance verification and parameter fine-tuning of the noise reduction and diagnosis parts are performed using 14,400 single-channel validation samples from a real industrial scenario dataset and 3,600 four-channel validation samples from a semi-anechoic environment dataset, respectively, to form a defined parameter configuration that meets the verification requirements, thus completing the training of the classification prediction network.

[0121] The air-domain acoustic diagnostic results of rotating machinery are the identification results of the fault type under a certain operating condition of the rotating machinery.

[0122] In some embodiments, the processor can input fused features into a classification prediction network, and further aggregate and fuse the features using a head convolutional module with a receptive field of 7×7, a stride of 1×1, and a ReLU activation function to obtain dimension-invariant feature outputs. A multi-scale convolutional module with a three-level mapping architecture is applied to the output features of the head convolutional module. First, the time-frequency space is compressed using a first-level mapping with a receptive field of 5×5, a stride of 2×2, a ReLU activation function, and 2×2 max pooling, achieving 10×8×64-dimensional feature extraction. Then, a second-level mapping with a receptive field of 3×3 and the same stride, pooling, and activation function configuration is used to expand the channel dimension while compressing the time-frequency space, forming 10×8×128-dimensional output features. Finally, in the third-level mapping, a linear convolutional module with a receptive field of 3×3, a stride of 1×1, and 2×2 max pooling is continuously applied to the receptive field of 3×3. A standard convolutional module with a stride of 1×1 and a ReLU activation function is used to extract a feature matrix with a dimension of 10×8×128. The features obtained from the three sets of parallel mappings are stacked along the channel dimension and fed into a tail convolutional module with a receptive field of 3×3, a stride of 1×1, and a ReLU activation function to achieve multi-scale feature fusion, outputting a feature matrix with a dimension of 10×8×384. Then, the output module uses depthwise separable convolution with a receptive field of 3×3 and a stride of 1×1 and global average pooling to compress the time-frequency space to a 1×1 unit size while keeping the channel dimension unchanged, and integrates the feature dimension to 384. This is then input into a fully connected layer, activated by Softmax, and output to obtain a 12-dimensional discriminant matrix, representing the predicted probability of 12 typical fault modes of rotating machinery under different working conditions. The category corresponding to the highest probability is selected as the acoustic diagnosis result of the air domain of the rotating machinery.

[0123] In some embodiments, the processor can take 14,400 single-channel test sample data from a real industrial scenario dataset as input, perform frame segmentation and windowing, and then send the data to a generator with fixed parameters after Mel-time-frequency conversion to recover and restore the noise-free Mel-time spectrum. Subsequently, the 14,400 single-channel denoised Mel-time spectra are integrated into 3,600 four-channel Mel-time spectra according to the corresponding acquisition channels, and input into a trained classification and prediction network to obtain the acoustic diagnosis results of the rotating machinery in the air domain, realizing the acoustic diagnosis of rotating machinery under non-static background noise interference in a real industrial scenario; the performance of the trained adversarial denoising network in the rotating machinery fault identification task is shown in Table 1.

[0124] Table 1 Comparison of Ablation Test Performance

[0125]

[0126] In some embodiments, Figure 2This is a visualization analysis chart of the time-frequency noise reduction performance ablation. A comparison of the experimental data in Table 1 shows that the method of this invention significantly outperforms other ablation schemes with functional module variant combinations in noise reduction diagnostic tasks. Analysis Figure 2 The visualization results of the noise reduction and restoration show that the traditional Transformer method can completely preserve the time-frequency structure characteristics of the signal and restore periodic information, but it lacks good spatial detail restoration capabilities. The restored signal after noise suppression exhibits a smooth transition, with significant loss of time-frequency spatial detail information. While the Swin Transformer method can restore local detail information, it is mostly concentrated in the high-frequency and low-frequency regions, lacking sufficient attention to the mid-frequency information containing the operating state of the rotating machinery. Similarly, for the other three schemes, the lack of corresponding configurations for time-frequency block sensing encoding, hierarchical information interaction time-frequency decoding, and local generative adversarial noise reduction models affects the performance of the local time-frequency multi-level sensing mechanism and multi-level feature aggregation restoration mechanism, leading to different types of defects in the restoration of time-frequency spatial detail information. The sparse time-frequency information of the rotating machinery radiation signal is closely related to its operating state. In the two-stage cascaded architecture of noise reduction and diagnosis, the defects in the noise reduction and restoration of time-frequency detail information will directly affect the diagnostic performance.

[0127] In some embodiments, Figure 3 This is a visual comparison and analysis chart of the time-frequency noise reduction performance based on filtering noise reduction methods. Figure 3 This can be achieved by comparing this method with traditional filtering-based noise reduction techniques under the same evaluation metrics. The experimental data in Table 2 shows that the performance of this invention's method in noise reduction diagnostic tasks is significantly superior to that of traditional methods. Analysis Figure 3 The visualization results of noise reduction and restoration show that filtering and noise reduction rely on noise-aware training, which cannot efficiently estimate and track global non-static noise. This leads to distortion or warping of the restored signal, resulting in the loss of signal periodic information and affecting diagnostic performance. Furthermore, due to the sparse structural characteristics of the radiated acoustic signals from rotating machinery, the weighted averaging signal recovery method within a local time window results in a restored signal with only a clear envelope outline but blurred time-frequency details. This makes the rotating machinery state characteristics carried by the acoustic signal singular, making it difficult to provide a basis for accurate prediction of fault conditions and resulting in low diagnostic accuracy.

[0128] Related Figure 2 , Figure 3Further analysis of Tables 1 and 2 leads to the conclusion that the generative adversarial denoising technique, using clean labeled samples for fine-grained reconstruction, provides rich time-frequency details for the diagnostic model, ensuring its diagnostic performance. Compared to traditional filtering denoising schemes, it exhibits significant technical advantages. This further demonstrates the innovative contribution and potential value of the present invention; Table 2 is a comparison table of the diagnostic performance of various methods.

[0129] Table 2 Performance Comparison with Other Existing Methods

[0130]

[0131] In some embodiments of this specification, a method for acoustic diagnosis of rotating machinery in the air domain based on generative adversarial noise reduction is proposed. (1) By utilizing a local time-frequency sensing mechanism, block-based correlation compression coding is carried out for medium-area block regions of multi-scale time-frequency features to improve the ability to perceive time-frequency details of sparse acoustic signals and obtain high-quality compressed coding features. (2) By utilizing a multi-level feature aggregation and restoration mechanism, and integrating the information transmission methods of skip connections and dense connections, a time-frequency decoding method based on hierarchical information interaction is proposed. Skip connections are used to help fuse the shallow detail features of the coding downsampling with the corresponding deep semantic features of the decoding upsampling, to restore the time-frequency spatial resolution. Based on dense connections, a multi-level feature reuse combination is formed inside the decoder to enhance the diversity and richness of time-frequency features, improve the ability to restore local details, and solve the problem of fine-grained restoration of sparse signals. (3) Using a discriminator, design a regional discriminator to traverse the local time-frequency region of the restored signal features at a medium scale to evaluate the noise suppression and sparse restoration performance, guide the reverse continuous optimization, so as to achieve the purpose of effectively estimating and tracking the time-varying statistical characteristics of non-static noise, and finally realize the efficient suppression of non-static background noise in real industrial scenarios and the refined restoration output of acoustic signals of rotating machinery, thereby improving the accuracy of acoustic diagnosis.

Claims

1. A method for air domain acoustic diagnosis of rotating machinery based on generative adversarial denoising, characterized in that, The method comprises the following steps: S1: using a local time-frequency perception mechanism of an encoder to perform segmented coding feature compression on a collected two-dimensional time-frequency acoustic sample of a rotating machine to obtain deep semantic compressed coding features; S1 comprises the following steps: using a mel filter bank to convert the collected two-dimensional time-frequency acoustic sample of the rotating machine to obtain a two-dimensional mel time-frequency spectrum; using patch partitioning to equally divide the two-dimensional mel time-frequency spectrum along the frequency domain axis and the time domain axis to obtain a local time-frequency region matrix; based on the local time-frequency region matrix, extracting corresponding information according to a similarity score, and performing segmented coding feature compression to obtain a time-frequency perceptual coding matrix; using a feedforward neural network to perform concatenation reconstruction on the time-frequency perceptual coding matrix based on an original input dimension to obtain encoded two-dimensional features; using a patch aggregation module to aggregate the encoded two-dimensional features to obtain deep semantic compressed coding features; wherein the mel filter bank, the patch partitioning, the local time-frequency region matrix, the feedforward neural network and the patch aggregation module belong to the encoder; S2: using a feature aggregation restoration module of a decoder to perform hierarchical information interaction time-frequency decoding on the deep semantic compressed coding features to gradually restore time-frequency spatial detail information of a sparse periodic sound signal from local to global, and restore reconstructed time-frequency information; S3: using a discriminator to perform equal-scale local noise reduction performance discrimination on the reconstructed time-frequency information, track and suppress non-static noise components mixed in the reconstructed time-frequency information, perform air domain acoustic noise reduction, and obtain a multi-channel two-dimensional time-frequency signal after noise reduction; wherein the encoder, the decoder and the discriminator belong to a generative adversarial noise reduction network, which is obtained by training; S4: using an information fusion classification network to process the multi-channel two-dimensional time-frequency signal after noise reduction, and using a depth separable convolution and a grouped convolution to fuse multi-channel time-frequency information to obtain fusion features; S5: using a classification prediction network to perform acoustic diagnosis on the fusion features to obtain a rotating machine air domain acoustic diagnosis result, and complete acoustic diagnosis of the rotating machine fault.

2. The method of claim 1, wherein, The expression of the time-frequency perceptual coding matrix is: ; ; ; ; wherein, denotes the "query" matrix of the time- and frequency-head self-attention matrix in the linear projection space, denotes the time- and frequency-head self-attention matrix, denotes the corresponding linear projection space, denotes a set of real numbers, denotes the formed time- frequency tile region, denotes the embedding vector dimension in the linear projection, denotes the "query" matrix of the time- and frequency-head self-attention matrix in the linear projection space, denotes the corresponding linear projection space, denotes the "value" matrix of the time- and frequency-head self-attention matrix in the linear projection space, denotes the corresponding linear projection space, denotes the self-attention function, denotes the function, denotes the transpose of a matrix, denotes the relative encoding position.

3. The method of claim 1, wherein, The expression of the encoded two-dimensional features is: ; ; ; ; wherein, represents a fully connected output equivalent to the initial dimension, represents a fully connected layer equivalent to the initial dimension, represents a time and frequency multi-head self-attention all-heads integrated output matrix, represents, represents an activation function, i.e., a nonlinear representation of the fully connected output result, represents a four times high-dimensional mapping fully connected layer, represents a time-frequency fusion perception encoding matrix, represents a combination stacking transformation, represents a time-domain perception encoding matrix, represents a frequency-domain perception encoding matrix, represents an encoded two-dimensional output feature, represents a splicing reconstruction transformation, i.e., a splicing combination patch area recovery to an original two-dimensional matrix dimension along the time and frequency axes.

4. The method of claim 1, wherein, S2 comprises the following steps: using a patch expansion component of the feature aggregation restoration module to perform time-frequency resolution expansion on the deep semantic compressed coding features to obtain up-sampling features; using a skip connection method to concatenate the up-sampling features and the time-frequency compressed coding features to obtain fusion representation information; using a hierarchical information interaction time-frequency decoding network to aggregate and restore the fusion representation information to obtain initial reconstructed time-frequency information; the initial reconstructed time-frequency information is stacked along the channel dimension with the fusion representation information through a dense connection method, and the hierarchical information interaction time-frequency decoding is performed on the initial reconstructed time-frequency information through a convolutional neural network to gradually restore time-frequency spatial detail information of a sparse periodic sound signal from local to global, and the reconstructed time-frequency information is obtained; the patch expansion component, the skip connection method and the hierarchical information interaction time-frequency decoding network belong to the feature aggregation restoration module.

5. The method of claim 1, wherein, The loss function expression of the generative adversarial denoising network training is: ; ; ; ; wherein, represents a discriminative loss function, represents, represents a discriminator, represents a clean mel-spectrogram without non-stationary noise interference, represents a first real data label, represents, represents a generator, i.e. an encoding-decoding mechanism, restores an output, represents a mixed mel-spectrogram with non-stationary noise interference, represents a second real data label, represents a generative loss function, represents a total number of samples, represents a th clean time-frequency, represents a th restored sample, represents a clean sample compressed encoding feature with dimension 5x4, represents a noisy sample compressed encoding feature with the same dimension, represents a feature vector module length.

6. The method of claim 1, wherein, The overall objective function expression of the generative adversarial denoising network is: ; ; wherein, denotes an optimization operation that minimizes the objective function, denotes an objective function of the discriminator, denotes a discriminative loss function, denotes a generative loss function, denotes a mean squared error loss, denotes a generative loss function, denotes a cosine similarity loss.

Citation Information

Patent Citations

  • Speech enhancement method and device based on multi-scale feature learning, equipment and medium

    CN120220712A

  • Water turbine fault classification diagnosis method based on multi-modal fusion and meta learning

    CN120356482A