An abnormal sound sample generation method, a part abnormal sound fault diagnosis method, device and equipment
By superimposing target noise onto the spectrogram of real abnormal noise audio data and performing feature encoding and adaptive calibration, an abnormal noise sample set that is highly similar to the real abnormal noise audio data is generated, which solves the problem of insufficient real fault samples and improves the accuracy and stability of abnormal noise fault diagnosis of components.
Patent Information
- Application Number
- CN202511053293.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing methods for diagnosing abnormal noises in components suffer from insufficient real-world fault samples, resulting in inadequate model generalization ability and difficulty in maintaining stable and accurate detection performance under diverse real-world operating conditions.
By superimposing target noise onto the spectrogram of real abnormal noise audio data, a noisy abnormal noise spectrogram is generated. Then, a joint calibration feature map is generated layer by layer through feature encoding, channel dimension adaptive calibration, and spatial dimension adaptive calibration. Combined with upsampling operation, an abnormal noise sample set that is highly similar to the real abnormal noise audio data is generated.
The fault diagnosis model for abnormal noises in components has been improved in terms of detection accuracy and generalization ability under complex working conditions, thereby enhancing the accuracy and stability of diagnosis.
Smart Images

Figure CN120561824B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of abnormal sound fault diagnosis, in particular to an abnormal sound sample generation method, a part abnormal sound fault diagnosis method, device and equipment. BACKGROUND
[0002] With the rapid development of modern manufacturing industry, mechanical part abnormal sound fault detection is crucial to ensure product quality and enhance user satisfaction.
[0003] The existing part abnormal sound fault diagnosis method is based on deep learning technology, which extracts time-frequency domain features of audio signals through convolutional neural network, recurrent neural network and other models to improve the classification accuracy of abnormal sound signals. However, in the industrial scene, the occurrence of part abnormal sound has high randomness and occasionality, which leads to difficulty in collecting real fault samples and lack of data, so that the existing model has insufficient generalization ability and is difficult to maintain stable and accurate detection performance in diversified actual working conditions. SUMMARY
[0004] The present application provides an abnormal sound sample generation method, a part abnormal sound fault diagnosis method, device and equipment to solve the problem of low accuracy of abnormal sound fault diagnosis caused by insufficient real fault samples in the prior art.
[0005] In order to solve the above problems, the present application discloses an abnormal sound sample generation method, which comprises:
[0006] Superimpose target noise on the real abnormal sound audio data spectrum graph to generate a noisy abnormal sound spectrum graph;
[0007] Perform N times of feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration on the noisy abnormal sound spectrum graph in turn, and generate N joint calibration feature maps layer by layer, each joint calibration feature map at least including feature information of fusion of target noise suppression and real abnormal sound audio data enhancement;
[0008] After feature encoding of the Nth joint calibration feature map, perform N times of upsampling operation, and each upsampling operation combines the corresponding joint calibration feature map for feature fusion to generate an abnormal sound sample set highly close to the real abnormal sound audio data.
[0009] Further, the N joint calibration feature maps are generated by a feature extraction module of a sample expansion model, and the feature extraction module at least includes N feature extraction layers;
[0010] Through the nth feature extraction layer, perform N times of feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration on the noisy abnormal sound spectrum graph in turn, and generate N joint calibration feature maps layer by layer, including:
[0011] The (n-1)th joint calibration feature map is feature-encoded to generate a preliminary feature map, where 2≤n≤N;
[0012] Based on the dependencies between channels in the preliminary feature map, weight parameters for each channel feature in the preliminary feature map are generated, and after weighted calculation, a channel dimension calibration feature map is generated.
[0013] For each channel of the channel dimension calibration feature map, the width aggregation feature and height aggregation feature of the channel are extracted along the row dimension and column dimension respectively. The row weight and column weight of each channel are generated through the position attention mechanism. After weighted operation, the nth joint calibration feature map is generated.
[0014] Furthermore, each feature extraction layer includes at least one channel dimension adaptive calibration layer;
[0015] The weight parameters for each channel feature in the generated preliminary feature map include:
[0016] For the preliminary feature map, the global feature value of each feature channel is calculated through the channel dimension adaptive calibration layer;
[0017] Based on the global feature value of each feature channel, the local interaction between each channel and its k neighboring channels is captured to generate the weight parameters of each channel feature in the preliminary feature map.
[0018] Furthermore, each feature extraction layer also includes at least one spatial dimension adaptive calibration layer;
[0019] The generation of row weights and column weights for each channel includes:
[0020] For each channel of the channel dimension calibration feature map, the local feature values of the channel dimension calibration feature map at each position are calculated through the spatial dimension adaptive calibration layer;
[0021] Based on the local feature values of the channel dimension calibration feature map at various locations, determine the width aggregation feature and height aggregation feature of each channel:
[0022] Along the channel dimension, the width and height aggregated features of the channel dimension calibration feature map are concatenated in each channel to obtain the spatial statistical feature matrix of each channel, and the position attention coordinates of each channel are generated. The position attention coordinates of each channel represent the attention intensity of the channel to different positions.
[0023] Based on the row attention coordinates and column attention coordinates of the position attention coordinates of each channel, the row weight and column weight of that channel are generated respectively, thus obtaining the row weight and column weight of each channel.
[0024] To solve the above problems, the application further discloses a part abnormal sound fault diagnosis method, the method comprising:
[0025] inputting the audio graph to be diagnosed into an abnormal sound fault diagnosis model to obtain an initial feature graph;
[0026] generating each channel weight according to the dependency relationship between each channel of the initial feature graph through the abnormal sound fault diagnosis model, wherein the each channel weight represents the correlation between the channel and the abnormal sound feature;
[0027] weighting and adjusting the audio graph to be diagnosed according to each channel weight through the abnormal sound fault diagnosis model to obtain a weighted feature graph,
[0028] performing fault diagnosis based on the weighted feature graph through the abnormal sound fault diagnosis model, and outputting a fault diagnosis result;
[0029] wherein the abnormal sound fault diagnosis model is trained based on an abnormal sound sample set, and each abnormal sound sample in the abnormal sound sample set is generated according to the abnormal sound sample generation method.
[0030] Further, the method further comprises:
[0031] collecting part normal audio data to generate a normal abnormal sound sample set;
[0032] annotating a label for each abnormal sound sample in the abnormal sound sample set and each normal sample in the normal abnormal sound sample set according to a true result to obtain a normal training sample set and the abnormal sound training sample set;
[0033] inputting each training sample in the normal training sample set and the abnormal sound training sample set into a pre-trained abnormal sound fault diagnosis model to perform model training to obtain a diagnosis result of the each training sample;
[0034] updating model parameters of the pre-trained abnormal sound fault diagnosis model according to the difference between the diagnosis result of the each training sample and the true result of the training sample to obtain a trained abnormal sound fault diagnosis model.
[0035] Further, the abnormal sound fault diagnosis model at least comprises a channel dimension adaptive calibration layer;
[0036] generating each channel weight according to the dependency relationship between each channel of the initial feature graph comprises:
[0037] calculating a global feature value of each feature channel through the channel dimension adaptive calibration layer for the initial feature graph;
[0038] Based on the global feature value of each feature channel, the local interaction of each channel with its m adjacent channels is captured through the channel dimension adaptive calibration layer to generate the channel weight of the initial feature map.
[0039] To solve the above problems, the application further discloses an abnormal sound sample generation device, comprising:
[0040] The superposition module is configured to superimpose target noise on the real abnormal sound audio data spectrum graph to generate a noisy abnormal sound spectrum graph.
[0041] The feature extraction module is configured to sequentially perform feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration on the noisy abnormal sound spectrum graph N times through the feature extraction module of the sample expansion model, and generate N joint calibration feature maps layer by layer, each joint calibration feature map comprising at least feature information fused by suppressing the target noise and enhancing the real abnormal sound audio data.
[0042] The feature reconstruction module is configured to perform N times of upsampling operation on the Nth joint calibration feature map after feature encoding, and perform feature fusion on the corresponding joint calibration feature map each time to generate an abnormal sound sample set highly close to the real abnormal sound audio data.
[0043] To solve the above problems, the application further discloses a component abnormal sound fault diagnosis device, comprising:
[0044] The to-be-diagnosed audio graph is input into the abnormal sound fault diagnosis model to obtain an initial feature map.
[0045] The weight generation module is configured to generate channel weights according to the dependency relationship between channels of the initial feature map through the abnormal sound fault diagnosis model, wherein the channel weights represent the correlation between the channel and abnormal sound features.
[0046] The weighted adjustment module is configured to perform weighted adjustment on the to-be-diagnosed audio graph according to each channel weight through the abnormal sound fault diagnosis model to obtain a weighted feature map.
[0047] The fault diagnosis module is configured to perform fault diagnosis based on the weighted feature map through the abnormal sound fault diagnosis model and output a fault diagnosis result.
[0048] The abnormal sound fault diagnosis model is trained based on an abnormal sound sample set, and each abnormal sound sample in the abnormal sound sample set is generated according to the abnormal sound sample generation method.
[0049] To solve the above problems, the application further discloses an electronic device, comprising:
[0050] A processor;
[0051] A memory for storing instructions executable by the processor;
[0052] The processor is configured to execute the instructions to implement any one of the abnormal sound sample generation method or any one of the component abnormal sound fault diagnosis method.
[0053] Compared with the prior art, the present application has the following advantages:
[0054] Based on real abnormal sound audio data, real abnormal sound audio data spectrum is generated through spectrum conversion technology, target noise is generated according to preset parameters, and real abnormal sound audio data spectrum is fused through frequency domain superposition or time domain aliasing mode, and noise interference under different noise interference is generated through gradual noise addition diffusion process. The noise-containing abnormal sound spectrum is sequentially executed N times of feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration, and the joint calibration feature map of fusion noise suppression and abnormal sound enhancement is generated layer by layer. After feature encoding of the Nth joint calibration feature map, the corresponding joint calibration feature map generated at each level is combined, and N times of up-sampling operation is performed, and finally the abnormal sound sample set highly close to the real abnormal sound audio data is generated. The problem that the traditional component abnormal sound fault diagnosis method has weak generalization ability due to insufficient real abnormal sound fault samples, and low detection accuracy in diversified actual working conditions is solved. The sample diversity is expanded through parameterized noise generation and gradual noise addition, the noise is accurately suppressed and the abnormal sound key features are strengthened through double-dimension adaptive calibration, and finally the high-quality sample set highly close to the real abnormal sound audio data is generated, the detection precision and generalization ability of the model for weak abnormal sound under complex working conditions are improved, and the accuracy and stability of the component abnormal sound fault diagnosis are enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0055] Figure 1 The method flowchart for generating abnormal sound samples provided by an embodiment of the present application is shown;
[0056] Figure 2 The example method flowchart for generating abnormal sound samples provided by an embodiment of the present application is shown;
[0057] Figure 3 The method flowchart for generating N joint calibration feature maps layer by layer provided by an embodiment of the present application is shown;
[0058] Figure 4 The example structure diagram of the sample expansion model provided by an embodiment of the present application is shown;
[0059] Figure 5 The example structure diagram of the feature extraction model provided by an embodiment of the present application is shown;
[0060] Figure 6 A flow chart of a method for diagnosing abnormal sound of a component is shown according to an embodiment of the present application;
[0061] Figure 7 A flow chart of a method for training an abnormal sound fault diagnosis model is shown according to an embodiment of the present application;
[0062] Figure 8 A structural schematic diagram of an abnormal sound sample generation device is shown according to an embodiment of the present application;
[0063] Figure 9 A structural schematic diagram of an abnormal sound fault diagnosis device for a component is shown according to an embodiment of the present application;
[0064] Figure 10 A structural schematic diagram of an electronic device is shown according to an embodiment of the present application. DETAILED DESCRIPTION
[0065] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0066] With the rapid development of modern manufacturing industry, mechanical components as the core components of industrial equipment, their running stability directly affects the overall performance and production efficiency of the equipment. Abnormal sound of components not only reduces the reliability of equipment operation, but also may indicate potential failure of mechanical system. If not detected in time, it may have a serious impact on the service life of the equipment, product quality and industrial production safety. Therefore, accurately identifying the abnormal sound of mechanical components is of great significance to improve the reliability of industrial equipment.
[0067] Traditional abnormal sound fault diagnosis methods mostly rely on artificial auscultation or simple sensor detection methods. Artificial auscultation highly depends on the experience and subjective judgment of the detector, and is low in efficiency and poor in consistency. The determination results of different detectors often have obvious differences. The detection method of simple sensor is limited by the accuracy of sensor and the simplicity of algorithm, and often uses threshold judgment, simple frequency spectrum analysis or time domain statistical characteristics, which is difficult to effectively capture the weak abnormal sound characteristics under complex working conditions, and cannot meet the needs of modern industrial manufacturing for high precision and high reliability detection.
[0068] With the rapid development of deep learning technology, convolutional neural networks, recurrent neural networks and other models have shown strong feature extraction and pattern recognition capabilities in the field of audio signal processing. By learning the complex features of abnormal sound signals, the subjectivity and limitations of traditional artificial auscultation are avoided, providing a more efficient and accurate solution for component abnormal sound diagnosis.
[0069] However, in actual industrial scenarios, the occurrence of part abnormal sound is highly random and sporadic, resulting in great difficulty in collecting real abnormal sound samples and a serious lack of data. Deep learning models are sensitive to data volume, and a small amount of samples can easily cause the model to overfit the noise or specific working condition characteristics in the training data, rather than the general abnormal sound pattern, resulting in decreased model generalization ability and difficulty in maintaining stable and accurate detection performance under diversified actual working conditions.
[0070] Taking the automobile industry as an example, the vehicle-mounted light machine curtain, as the main part of the in-vehicle entertainment system, the abnormal sound problem in its running process not only seriously affects the passenger's riding experience, but also may reflect the potential failure of internal parts, and has an impact on the overall quality and brand image of the vehicle. Therefore, accurately detecting the abnormal sound of the vehicle-mounted curtain is of great significance for improving the quality of automobile products and enhancing user satisfaction. Since the occurrence of abnormal sound of vehicle-mounted light machine curtain in actual production scenarios is highly random, the number of real abnormal sound samples collected is extremely limited. Under the condition that a deep learning model needs a large amount of training data, a small amount of samples can easily cause overfitting, resulting in insufficient model generalization ability and difficulty in maintaining stable and accurate detection performance in diversified actual working conditions.
[0071] Based on the above technical problems, the embodiments of the present application provide an abnormal sound sample production method, a part abnormal sound fault diagnosis method, device and electronic equipment. Referring to Figure 1 The embodiment of the present application provides a method flow chart for generating abnormal sound samples. The embodiment of the present application proposes an abnormal sound sample generation method, which includes:
[0072] S10, superimposing target noise on the real abnormal sound audio data spectrum graph to generate a noisy abnormal sound spectrum graph.
[0073] Specifically, the real abnormal sound audio data spectrum graph refers to a two-dimensional spectrum graph generated by mapping a one-dimensional real abnormal sound audio signal to time-frequency domain through spectrum conversion technology, which can be implemented by Fourier transform. The target noise refers to a noise spectrum graph generated according to a preset parameter, which is fused with the real abnormal sound audio data spectrum graph through frequency domain superposition or time domain aliasing to generate a noisy abnormal sound spectrum graph.
[0074] Exemplarily, in the process of generating the noisy abnormal sound spectrogram, first, the real abnormal sound audio signal is collected, and after the pre-processing of the real abnormal sound audio signal, the real abnormal sound audio signal is converted into a real abnormal sound audio data spectrogram through Fourier transform, realizing the mapping of the audio signal from the time domain to the frequency domain; then, according to the noise spectrogram generated by the preset parameter, the noise spectrogram is gradually added to the real abnormal sound audio data spectrogram to form a gradual noise adding diffusion process until the noisy abnormal sound spectrogram becomes approximately Gaussian distribution. The preset parameter can be Gaussian noise, environmental noise or specific frequency band interference noise. According to the type of the real abnormal sound audio signal, the preset parameter, the noise distribution parameter and the noise time-varying characteristic are adjusted, the gradual noise adding diffusion process is implemented on the same real abnormal sound audio data spectrogram, and finally the noisy abnormal sound spectrogram covering different working condition noise interference is generated, so that a single real abnormal sound audio data spectrogram can derive a plurality of noisy abnormal sound spectrograms with different noise characteristics, thereby improving the number and diversity of generated abnormal sound samples, and providing more abundant and representative data support for the part abnormal sound fault diagnosis model.
[0075] S20, sequentially performing N times of feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration on the noisy abnormal sound spectrogram, and generating N joint calibration feature maps layer by layer, each joint calibration feature map at least including: feature information fused by target noise suppression and real abnormal sound audio enhancement.
[0076] Specifically, the feature encoding refers to feature extraction and spatial scale compression of the noisy abnormal sound spectrogram through the convolution layer and the down-sampling operation. The convolution layer is used to capture the local time-frequency domain mode of the noisy abnormal sound spectrogram and each joint calibration feature map, and the down-sampling gradually reduces the spatial resolution of the noisy abnormal sound spectrogram and each joint calibration feature map, so as to make the subsequent convolution layer aggregate information in a larger spatial range, expand the receptive field range layer by layer, and extract multi-scale information from fine-grained local features to coarse-grained global features; in the feature encoding stage, the mixed features of the real abnormal sound audio data spectrogram and the target noise are reserved to provide basic feature representation for each channel dimension adaptive calibration and spatial dimension adaptive calibration.
[0077] To solve the problems of local convolution operation in equal processing of each channel feature, not distinguishing key semantics and redundant information, lacking dynamic evaluation of cross-layer feature fusion, insufficient efficiency of multi-scale information integration, and difficulty in capturing long-distance spatial correlation, the embodiment of the present application proposes channel dimension adaptive calibration and spatial dimension adaptive calibration. Channel dimension adaptive calibration refers to generating channel-level statistical features through global average pooling, then learning the importance weight of each channel through a fully connected layer or a one-dimensional convolution, enhancing the channel related to the noise, and suppressing the channel related to the noise. Through weight re-calibration of the channel dimension, the high discriminative feature channel is strengthened, realizing channel suppression of the target noise and channel enhancement of the real noise audio data.
[0078] Spatial dimension adaptive calibration refers to performing average pooling operation along the channel dimension after channel calibration, merging and then learning the weight distribution of the spatial position through the convolution layer, focusing on the local spatial area where the noise occurs, and suppressing the noise interference of the irrelevant spatial area. Through feature re-calibration of the spatial dimension, the key area is focused to capture long-distance dependence, realizing spatial-level suppression of the target noise and spatial-level enhancement of the real noise feature.
[0079] Joint calibration feature map refers to the feature map generated after feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration. The joint calibration feature map generated at each level not only retains the multi-scale time-frequency domain features, but also fuses the double information of suppression of the target noise and enhancement of the real noise audio data through double calibration of the channel and spatial dimensions. With the deepening of the level, the joint calibration feature map gradually focuses on more global and abstract noise patterns, providing multi-scale and noise-robust feature representation for subsequent execution of the upsampling operation to generate a noise sample set highly close to the real noise audio data.
[0080] In the embodiment, feature encoding retains the time-frequency domain information from fine granularity to coarse granularity, channel calibration suppresses noise dominant channels and enhances noise related channels, spatial calibration focuses on noise occurrence areas and suppresses irrelevant spatial noise, accurately suppresses the target noise and directionally enhances the real noise feature, and the finally generated joint calibration feature map fuses the double information of noise suppression and noise enhancement, and gradually focuses on more abstract noise patterns with the deepening of the level, providing a multi-scale and high representation feature basis for subsequent generation of a noise sample set highly close to the real noise audio data, effectively improving the detection ability and generalization performance of the part noise fault diagnosis model for weak noise under complex working conditions.
[0081] S30, after feature encoding of the Nth joint calibration feature map, performing N times of upsampling operation, and each time of upsampling operation combines the corresponding joint calibration feature map to perform feature fusion, generating a noise sample set highly close to the real noise audio data.
[0082] Specifically, the upsampling operation, as the core step of the decoder, gradually restores the feature map obtained after feature encoding of the Nth joint calibration feature map to the heterophonic sample data highly close to the real heterophonic audio data through layer-by-layer denoising. After each upsampling operation, the feature map generated by the current upsampling is fused with the joint calibration feature map generated at the corresponding level. Specifically, the jump connection can be used to combine the low-level detail features of the joint calibration feature map with the high-level semantic features generated in the upsampling process, avoid the loss of part of the feature information due to downsampling, and improve the fineness of the heterophonic sample data. By gradually reversing the adding process of the target noise, the noisy heterophonic spectrogram is denoised, and the heterophonic sample highly close to the real heterophonic audio data is reconstructed. Through layer-by-layer denoising and feature enhancement, the interference of the target noise is effectively suppressed, and the key features of the real heterophonic audio data are preserved.
[0083] For example, N times of feature encoding is implemented by using an encoder, multi-scale abstract expression of features is realized through hierarchical combination of convolution and pooling operations, upsampling operation is combined with deconvolution operation to gradually restore the feature space resolution of the real heterophonic audio data spectrogram, and a heterophonic sample set highly close to the real heterophonic audio data is generated. The jump connection directly maps each joint calibration feature map to each upsampling operation, forming a cross-layer information transmission channel, which not only preserves the low-level features such as edges and textures captured by the shallow network, but also combines the high-level semantic information of upsampling, preserves multi-scale information, improves the image clarity of the heterophonic sample set, and improves the completeness and accuracy of feature expression.
[0084] For example, Figure 2 An example method flowchart for generating a heterophonic sample provided by an embodiment of the present application is shown. Referring to Figure 2 Taking a vehicle-mounted light machine curtain as an example, the original data refers to the real heterophonic audio data spectrogram, the Gaussian noise refers to the target noise, and the generated data refers to the heterophonic sample set. After collecting the real heterophonic audio data of the vehicle-mounted light machine curtain, the one-dimensional real heterophonic audio data is first converted into a two-dimensional real heterophonic audio data spectrogram to obtain the original data, and then the Gaussian noise is added to the original data to generate a noisy heterophonic spectrogram close to the standard Gaussian distribution through step-by-step noise addition. By predicting and eliminating noise, the real data distribution is restored from the noisy heterophonic spectrogram through step-by-step denoising of the noisy heterophonic spectrogram, and the generated data highly close to the real heterophonic audio data spectrogram is generated. The scale and diversity of the generated data are effectively increased, the generated data is used to train the part heterophonic fault diagnosis model, which not only makes up for the limitation of data deficiency on the performance of the model, but also ensures that the model can accurately identify various vehicle-mounted curtain heterophones in the actual production environment, thereby providing solid technical support for improving the automobile quality control level and user experience.
[0085] The embodiment of the application generates a real abnormal sound audio data spectrum graph through a spectrum conversion technology based on real abnormal sound audio data, generates a target noise according to a preset parameter, and fuses the real abnormal sound audio data spectrum graph through a frequency domain superposition or time domain aliasing manner. A noisy abnormal sound spectrum graph under different noise interference is generated through a gradual noise adding diffusion process. The noisy abnormal sound spectrum graph is sequentially executed N times of feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration, and a joint calibration feature map fused with noise suppression and abnormal sound enhancement is generated layer by layer. After feature encoding of the Nth joint calibration feature map, combined with the joint calibration feature map generated at the corresponding level, N times of upsampling operation is performed, and finally an abnormal sound sample set highly close to the real abnormal sound audio data is generated. The problem that the traditional component abnormal sound fault diagnosis method has weak generalization ability due to insufficient real abnormal sound fault samples, and low detection accuracy in diversified actual working conditions is solved. The sample diversity is expanded through parameterized noise generation and gradual noise adding, the noise is accurately suppressed and the abnormal sound key features are strengthened through double-dimensional adaptive calibration, and finally a high-quality sample set highly close to the real abnormal sound audio data is generated, so as to improve the detection precision and generalization ability of the model to weak abnormal sound under complex working conditions, and enhance the accuracy and stability of the component abnormal sound fault diagnosis.
[0086] Figure 3 A method flowchart for generating N joint calibration feature maps layer by layer is shown, Figure 4 An exemplary structure diagram of the sample expansion model is shown, Figure 5 An exemplary structure diagram of the feature extraction model is shown. Referring to Figures 3 to 5 The N joint calibration feature maps are generated by the feature extraction module of the sample expansion model, and the feature extraction module at least includes N feature extraction layers.
[0087] Referring to Figure 4The sample expansion model is used to generate a noisy heterophonic spectrogram, and a heterophonic sample set highly close to the real heterophonic audio data is generated by performing N times of feature encoding, channel dimension adaptive calibration, spatial dimension adaptive calibration and upsampling operation. Specifically, the sample expansion model comprises a feature extraction module and a sample generation module. The feature extraction module comprises at least N feature extraction layers. Each feature extraction layer comprises at least a convolution layer and a double-dimension adaptive calibration layer connected in sequence. The double-dimension adaptive calibration layer comprises at least a channel dimension adaptive calibration layer and a spatial dimension adaptive calibration layer, which are used to perform N times of channel dimension adaptive calibration and spatial dimension adaptive calibration, respectively. Each feature extraction layer performs feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration once. The first feature extraction layer is used to sequentially perform feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration on the noisy heterophonic spectrogram to generate a first joint calibration feature map. The joint calibration feature map generated by each feature extraction layer is used as the input of the next feature extraction layer.
[0088] S20, sequentially performing N times of feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration on the noisy heterophonic spectrogram to generate N joint calibration feature maps layer by layer, comprising:
[0089] S21, performing feature encoding on the (n-1)th joint calibration feature map by the nth feature extraction layer to generate a preliminary feature map, wherein 2≤n≤N.
[0090] Referring to Figure 4 In the process of performing feature encoding on the (n-1)th joint calibration feature map to generate a preliminary feature map, local time-frequency domain pattern capture is performed by the convolution layer of the nth feature extraction layer. The convolution kernel slides on the (n-1)th joint calibration feature map to extract the time-frequency joint features of the local region and generate a preliminary feature map reflecting the initial feature mapping.
[0091] S22, generating a weight parameter of each channel feature in the preliminary feature map according to the dependency relationship between the channels in the preliminary feature map, and generating a channel dimension calibration feature map through a weighting operation.
[0092] Specifically, each feature extraction layer comprises at least one channel dimension adaptive calibration layer.
[0093] The generation of the weight parameter of each channel feature in the preliminary feature map comprises:
[0094] S22-1, for the preliminary feature map, calculating a global feature value of each feature channel by the channel dimension adaptive calibration layer.
[0095] S22-2, according to the global feature value of each feature channel, capturing the local interaction of each channel with its k adjacent channels, generating the weight parameter of each channel feature in the preliminary feature map.
[0096] For example, referring to Figure 5 , the channel attention refers to the channel dimension adaptive calibration layer, the nth channel dimension adaptive calibration layer is connected after the nth feature extraction layer, the preliminary feature map is taken as the input of the channel dimension adaptive calibration layer, the global feature value of each feature channel is generated by aggregating the spatial dimension information through the global average pooling, so as to construct the dependency relationship between the feature channels, the local interaction of each channel with k adjacent channels is captured through one-dimensional convolution, and the channel weight is generated through the sigmoid activation function, the adaptive weighting of each channel is realized through the weighted fusion with the preliminary feature map, and the channel dimension calibration feature map which strengthens the frequency or semantic features related to the target task is generated. Specifically, the specific value range of k is set according to actual requirements, which is not limited here.
[0097] S23, for each channel of the channel dimension calibration feature map, the width aggregation feature and the height aggregation feature of the channel are extracted along the row dimension and the column dimension respectively, and the row weight and the column weight of each channel are generated through the position attention mechanism, and the nth joint calibration feature map is generated after the weighted operation.
[0098] For example, referring to Figure 5 , each feature extraction layer further includes at least one spatial dimension adaptive calibration layer. Wherein, Figure 5 The spatial attention in the spatial dimension adaptive calibration layer refers to the spatial dimension adaptive calibration layer, and each spatial dimension adaptive calibration layer is connected between the channel dimension adaptive calibration layer of the current feature extraction layer and the convolution layer of the next feature extraction layer.
[0099] Specifically, the row weight and the column weight of each channel are generated, including:
[0100] S23-1, for each channel of the channel dimension calibration feature map, the local feature value of the channel dimension calibration feature map at each position is calculated through the spatial dimension adaptive calibration layer.
[0101] For example, in the process of generating the nth joint calibration feature map, the spatial dimension adaptive calibration layer first calculates the local feature value of each position in each channel of the channel dimension calibration feature map. The local feature value of each position is the statistical value of the pixels at the position.
[0102] S23-2, according to the local feature value of the channel dimension calibration feature map at each position, the width aggregation feature and the height aggregation feature of each channel are determined.
[0103] Specifically, for each feature channel, the width aggregated feature refers to a feature representation obtained by aggregating the channel-dimensionally calibrated feature map along the width direction, and the height aggregated feature refers to a feature representation obtained by aggregating the channel-dimensionally calibrated feature map along the height direction. Through an average pooling operation, the width aggregated feature and the height aggregated feature of each channel can be obtained.
[0104] For example, the width aggregated feature and the height aggregated feature of each channel are calculated according to the following formula:
[0105]
[0106]
[0107] wherein, and are the width and the height of the input feature map respectively, i refers to the row index of the channel-dimensionally calibrated feature map, and j refers to the column index of the channel-dimensionally calibrated feature map, is a local feature value of the channel-dimensionally calibrated feature map at the position of i and j, refers to the width aggregated feature, which is a vector with a height dimension of H and a width dimension of 1, and reflects the global feature distribution of the channel in the horizontal direction, refers to the height aggregated feature, which is a vector with a height dimension of 1 and a width dimension of W, and reflects the global feature distribution of the channel in the vertical direction.
[0108] S23-3, along the channel dimension, the width aggregated feature and the height aggregated feature of the channel-dimensionally calibrated feature map in each channel are spliced to obtain a spatial statistical feature matrix of each channel, and a position attention coordinate of each channel is generated, wherein the position attention coordinate of each channel represents the attention intensity of the channel to different positions.
[0109] For example, the position attention coordinate of each channel is calculated according to the following formula:
[0110]
[0111] wherein, refers to the position attention coordinate of each channel, represents a 1*1 convolution, refers to the width aggregated feature, refers to the height aggregated feature, represents a splicing operation.
[0112] S23-4, according to the row attention coordinate and the column attention coordinate of the position attention coordinate of each channel, a row weight and a column weight of the channel are respectively generated, to obtain the row weight and the column weight of each channel.
[0113] Specifically, according to the row attention coordinates and the column attention coordinates of the position attention coordinates of each channel, the row weight and the column weight of each channel can be obtained through a batch normalization operation, a ReLu function and a sigmoid function.
[0114] For example, the row weight and the column weight of each channel are obtained according to the following formula:
[0115]
[0116]
[0117] The nth joint calibration feature map is generated according to the following formula:
[0118]
[0119] wherein, represents the row weight, represents the column weight, and E represents the channel dimension calibration feature map.
[0120] The embodiments of the present application capture local cross-channel interaction by comprehensively considering each feature channel and its k nearest neighbor feature channels, solve the technical problems of feature loss caused by capturing nonlinear cross-channel interaction through dimension reduction operation in the prior art, and interaction efficiency reduction caused by capturing the dependency relationship between all channels. Not only improves the efficiency of cross-channel information interaction, but also inputs the channel dimension calibration feature map into the spatial dimension adaptive calibration layer, extracts key position information from each channel of each channel dimension calibration feature map. Realize the complementation of information with the channel dimension calibration feature map, retain multi-scale time-frequency information through feature coding, combine channel and spatial double dimension adaptive calibration, accurately suppress target noise, solve the problem of insufficient spatial correlation capture in traditional methods, directional enhancement of real abnormal feature, generate feature map with noise robustness and abnormal feature representation, provide key support for subsequent generation of abnormal sample set close to real abnormal audio data, effectively improve the detection precision and generalization ability of the part abnormal fault diagnosis model to weak abnormality under complex working conditions.
[0121] Figure 6 A flow chart of a part abnormal fault diagnosis method provided by an embodiment of the present application is shown. Referring to Figure 6 The embodiments of the present application also provide a part abnormal fault diagnosis method, comprising:
[0122] S100, input the audio graph to be diagnosed into an abnormal fault diagnosis model to obtain an initial feature map.
[0123] The abnormal sound fault diagnosis model includes an input layer, a hidden layer and an output layer, and is used for part abnormal sound fault diagnosis and outputs an abnormal sound fault diagnosis result. The input layer is used for receiving a to-be-diagnosed audio graph, the hidden layer includes three convolution layers, three pooling layers and two full connection layers, the convolution layers are respectively composed of 16 groups, 6 groups and 1 group of convolution kernels, are activated by a ReLU function, a channel dimension adaptive calibration layer is connected after the last convolution layer, is used for extracting important parts in each feature information of an initial feature map, and is used for improving the recognition and classification ability of the model. The first full connection layer includes 64 neurons, a random inactivation layer with a dropout rate of 0.1 is connected between the first full connection layer and the second full connection layer to inhibit overfitting, the second full connection layer includes 4 neurons, and finally a softmax activation function is used to output a classification probability.
[0124] In S200, according to the dependency relationship between each channel of the initial feature map, each channel weight is generated by an abnormal sound fault diagnosis model, and the each channel weight represents the correlation between the channel and the abnormal sound feature.
[0125] Specifically, according to the dependency relationship between each channel of the initial feature map, each channel weight is generated, including:
[0126] In S210, for the initial feature map, a global feature value of each feature channel is calculated by a channel dimension adaptive calibration layer.
[0127] In S220, based on the global feature value of each feature channel, a local interaction between each channel and its m adjacent channels is captured by the channel dimension adaptive calibration layer, and each channel weight of the initial feature map is generated.
[0128] For example, in the process of generating each channel weight, first, a global feature value of each feature channel of the initial feature map is calculated. The global feature value can be a global feature statistical value of each feature channel, and can be a pixel average value. By capturing the local interaction between each channel and its m adjacent channels, each channel weight of the initial feature map is generated, which solves the feature loss caused by the dimension reduction operation in the prior art to capture the nonlinear cross-channel interaction, and the interaction efficiency caused by capturing the dependency relationship between all channels.
[0129] In S300, according to each channel weight, a to-be-diagnosed audio graph is weighted and adjusted by an abnormal sound fault diagnosis model, and a weighted feature map is obtained.
[0130] In S400, based on the weighted feature map, fault diagnosis is performed by an abnormal sound fault diagnosis model, and a fault diagnosis result is output.
[0131] Specifically, based on each channel weight, the weighted feature map obtained after the abnormal sound fault diagnosis model is used to determine the fault diagnosis result after the diagnosis audio graph is adjusted by weighting. The fault diagnosis result includes two specific cases of normal and abnormal, and the abnormal fault diagnosis result at least includes a specific fault type.
[0132] For example, the abnormal sound fault diagnosis model performs a part abnormal sound fault diagnosis process as follows:
[0133] The input layer receives the audio graph to be diagnosed, extracts multi-scale time-frequency features through the 3 convolutional layers and pooling layers of the hidden layer, and generates an initial feature map. Then, the channel dimension adaptive calibration layer after the last convolutional layer is processed, the global feature values of each channel of the initial feature map are calculated, the local interaction of each feature channel with its m adjacent channels is captured, and the channel weight representing the correlation between the channel and the abnormal sound is generated. The weighting feature map is obtained by weighting and adjusting the to-be-diagnosed audio graph, which strengthens the key features of abnormal sound and suppresses redundant information. Finally, through the softmax activation, the classification probability of normal or abnormal fault type is output through the full connection layer containing 64 neurons, the random inactivation layer with a dropout rate of 0.1, and the full connection layer with 4 neurons.
[0134] The abnormal sound fault diagnosis model is trained based on the abnormal sound sample set. Each abnormal sound sample in the abnormal sound sample set is generated according to any one of the abnormal sound sample generation methods. The abnormal sound fault diagnosis model trained using sufficient abnormal sound samples effectively solves the technical problem of insufficient accuracy of abnormal sound fault diagnosis caused by data scarcity in the prior art, and achieves the technical effect of ensuring the accuracy of abnormal sound fault diagnosis under complex working conditions.
[0135] Figure 7 A method for training an abnormal sound fault diagnosis model is shown in the flowchart. Referring to Figure 7 , the training process of the abnormal sound fault diagnosis model includes:
[0136] A1, collecting part normal audio data, and generating a normal abnormal sound sample set.
[0137] A2, according to the true result, each abnormal sound sample in the abnormal sound sample set and each normal sample in the normal abnormal sound sample set is labeled, and a normal training sample set and an abnormal training sample set are obtained.
[0138] A3, each training sample in the normal training sample set and the abnormal training sample set is input into the pre-trained abnormal sound fault diagnosis model for model training, and the diagnosis result of each training sample is obtained.
[0139] A4, according to the difference between the diagnosis result of each training sample and the true result of the training sample, the model parameters of the pre-trained abnormal sound fault diagnosis model are updated, and the trained abnormal sound fault diagnosis model is obtained.
[0140] In the training process of the abnormal sound fault diagnosis model, normal audio data of parts is collected to generate a normal abnormal sound sample set, and an abnormal sound sample set is generated through an abnormal sound sample generation method. Each abnormal sound sample in the abnormal sound sample set and each normal sample in the normal abnormal sound sample set are labeled, that is, each abnormal sound sample is labeled with a label reflecting the abnormal sound category, and each normal audio data is labeled with a label reflecting the normal result. Through model training, the difference between the real result and the diagnosis result is compared, the model parameters are updated, and the trained abnormal sound fault diagnosis model is obtained.
[0141] The abnormal sound sample set generated by the abnormal sound sample generation method trains the abnormal sound fault diagnosis model, which can solve the technical problem of insufficient model generalization ability caused by insufficient training samples. Training the abnormal sound fault diagnosis model with a large number of training samples effectively improves the accuracy of abnormal sound fault diagnosis under complex working conditions.
[0142] Figure 8 The structure of the abnormal sound sample generation device provided by an embodiment of the application is shown. Referring to Figure 8 , the abnormal sound sample generation device comprises:
[0143] The superposition module 10 is configured to superimpose target noise on the real abnormal sound audio spectrum graph to generate a noisy abnormal sound spectrum graph.
[0144] The feature extraction module 20 is configured to sequentially perform N times of feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration on the noisy abnormal sound spectrum graph to generate N joint calibration feature maps layer by layer, and each joint calibration feature map at least includes feature information that fuses suppression of target noise and enhancement of real abnormal sound audio data.
[0145] The feature reconstruction module 30 is configured to perform N times of upsampling operation on the Nth joint calibration feature map after feature encoding, and each upsampling operation is combined with the corresponding joint calibration feature map for feature fusion to generate an abnormal sound sample set highly close to the real abnormal sound audio data.
[0146] In some embodiments, the N joint calibration feature maps are generated by a feature extraction module of a sample expansion model, and the feature extraction module at least includes N feature extraction layers.
[0147] The feature extraction module 20 further comprises:
[0148] The feature encoding unit is configured to perform feature encoding on the (n-1)th joint calibration feature map through the nth feature extraction layer to generate a preliminary feature map, wherein 2≤n≤N.
[0149] The weight parameter generation unit is configured to generate a weight parameter of each channel feature in the preliminary feature map according to a dependency relationship between the channels in the preliminary feature map, and generate a channel-dimension calibration feature map through a weighting operation.
[0150] The weight parameter generation unit further includes:
[0151] The calculation subunit is configured to calculate a global feature value of each feature channel through the channel-dimension adaptive calibration layer for the preliminary feature map.
[0152] The capture subunit is configured to capture a local interaction of each channel with its k adjacent channels according to the global feature value of each feature channel, and generate the weight parameter of each channel feature in the preliminary feature map.
[0153] The joint calibration feature map generation unit is configured to extract a width aggregation feature and a height aggregation feature of each channel of the channel-dimension calibration feature map along a row dimension and a column dimension respectively, generate a row weight and a column weight of each channel through a position attention mechanism, and generate an nth joint calibration feature map through a weighting operation.
[0154] Each feature extraction layer further includes at least one spatial-dimension adaptive calibration layer. The joint calibration feature map generation unit further includes:
[0155] The local feature value generation subunit is configured to calculate a local feature value of the channel-dimension calibration feature map at each position through the spatial-dimension adaptive calibration layer for each channel of the channel-dimension calibration feature map.
[0156] The determination subunit is configured to determine the width aggregation feature and the height aggregation feature of each channel according to the local feature value of the channel-dimension calibration feature map at each position.
[0157] The splicing subunit is configured to splice the width aggregation feature and the height aggregation feature of each channel of the channel-dimension calibration feature map along a channel dimension, obtain a spatial statistical feature matrix of each channel, and generate a position attention coordinate of each channel, where the position attention coordinate of each channel represents an attention intensity of the channel to different positions.
[0158] The generation subunit is configured to generate the row weight and the column weight of each channel according to a row attention coordinate and a column attention coordinate of the position attention coordinate of each channel.
[0159] Figure 9 A structure schematic diagram of a part abnormal sound fault diagnosis device is shown, and the part abnormal sound fault diagnosis device includes:
[0160] The input module 100 is configured to input an audio graph to be diagnosed into the abnormal sound fault diagnosis model to obtain an initial feature graph.
[0161] The channel weight generation module 200 is configured to generate, by the abnormal sound fault diagnosis model, a channel weight of each channel according to a dependency relationship between the channels of the initial feature graph, where the channel weight represents a correlation between the channel and an abnormal sound feature.
[0162] The adjustment unit 300 is configured to perform, by the abnormal sound fault diagnosis model, weighted adjustment on the audio graph to be diagnosed according to each channel weight to obtain a weighted feature graph.
[0163] The fault diagnosis unit 400 is configured to perform, by the abnormal sound fault diagnosis model, fault diagnosis based on the weighted feature graph to output a fault diagnosis result.
[0164] In some embodiments, the channel weight generation module 200 comprises:
[0165] The calculation unit is configured to calculate, for the initial feature graph, a global feature value of each feature channel by a channel-dimension adaptive calibration layer.
[0166] The capture unit is configured to capture, based on the global feature value of each feature channel, a local interaction between each channel and its m adjacent channels by the channel-dimension adaptive calibration layer to generate the channel weight of the initial feature graph.
[0167] In some embodiments, the abnormal sound fault diagnosis device for the component further comprises:
[0168] The collection module is configured to collect normal audio data of the component to generate a normal abnormal sound sample set.
[0169] The labeling module is configured to label a label for each abnormal sound sample in the abnormal sound sample set and each normal sample in the normal abnormal sound sample set according to the true result to obtain a normal training sample set and an abnormal sound training sample set.
[0170] The training module is configured to input each training sample in the normal training sample set and the abnormal sound training sample set into the pre-trained abnormal sound fault diagnosis model to perform model training to obtain a diagnosis result of each training sample.
[0171] The updating module is configured to update a model parameter of the pre-trained abnormal sound fault diagnosis model according to a difference between the diagnosis result of each training sample and a true result of the training sample to obtain a trained abnormal sound fault diagnosis model.
[0172] Figure 10 An electronic device structure diagram provided by an embodiment of the present application is shown. Referring to Figure 10 The present application further provides an electronic device, comprising:
[0173] a processor.
[0174] a memory for storing processor-executable instructions.
[0175] The processor is configured to execute the instructions to implement any one of the abnormal sound sample generation methods or any one of the component abnormal sound fault diagnosis methods.
[0176] In the embodiment, the computer device includes a processor, a memory and a network interface connected by a system bus.
[0177] The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is configured to store data samples. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement any one of the abnormal sound sample generation methods or any one of the component abnormal sound fault diagnosis methods.
[0178] Those skilled in the art can understand that, Figure 10 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0179] The embodiments of the present application also provide a computer readable storage medium, when the instructions in the computer readable storage medium are executed by the processor of the terminal, the terminal can execute any one of the abnormal sound sample generation methods or any one of the component abnormal sound fault diagnosis methods.
[0180] The computer readable storage medium described above can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic storage, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0181] Optionally, a readable storage medium is coupled to the processor, such that the processor is enabled to read information from, and write information to, the readable storage medium. Of course, the readable storage medium could be a part of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the readable storage medium could also be located in a device.
[0182] The embodiments of the present application further provide a computer program product, which comprises a computer program. The computer program is executed by a processor to implement any one of the abnormal sound sample generation method or any one of the part abnormal sound fault diagnosis method.
[0183] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, a device or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0184] The present application is described with reference to the flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device implemented in the flowcharts and / or block diagrams. Figure 1 The device that implements the function specified in one or more flows and / or blocks. Figure 1 The device that implements the function specified in one or more flows and / or blocks.
[0185] These computer program instructions can also be stored in a computer readable storage medium that can direct the computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable storage medium produce a manufactured product including instruction devices, which implement the flowcharts and / or block diagrams. Figure 1 The device that implements the function specified in one or more flows and / or blocks. Figure 1 The device that implements the function specified in one or more flows and / or blocks.
[0186] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 one or more flowcharts and / or blocks
[0187] Although preferred embodiments of the application have been described herein, changes and modifications can be suggested to one skilled in the art and are intended to be encompassed within the scope of the appended claims. It is the intent, therefore, to be limited only as indicated by the scope of the claims appended hereto.
[0188] It will be apparent to those skilled in the art that various modifications and variations can be made to the present application without departing from the spirit or scope of the application. Thus, it is intended that the present application cover the modifications and variations of this application provided they come within the scope of the appended claims and their equivalents.
[0189] The above provides a noise sample generation method, a part noise fault diagnosis method, device and equipment, which are described in detail. The principles and implementation manners of the present application are described by using specific examples. The above examples are only used to help understand the method and its core idea of the present application. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application scope will be changed. In summary, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method for generating an abnormal sound sample, characterized by, The method comprises: superimposing the target noise on the real abnormal sound audio spectrum graph to generate a noisy abnormal sound spectrum graph; performing feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration on the noisy abnormal sound spectrum graph in turn N times to generate N joint calibration feature maps layer by layer, each joint calibration feature map at least comprising feature information of fusion of the target noise suppression and the real abnormal sound audio enhancement; after feature encoding on the Nth joint calibration feature map, performing N times of upsampling operation, each upsampling operation combining corresponding joint calibration feature map for feature fusion to generate an abnormal sound sample set highly close to the real abnormal sound audio data; wherein the channel dimension adaptive calibration refers to generating channel level statistical features through global average pooling, then learning importance weights of each channel through a fully connected layer or a one-dimensional convolution to enhance channels strongly related to abnormal sound and suppress channels strongly related to noise; the spatial dimension adaptive calibration refers to performing average pooling operation along the channel dimension after channel calibration, then learning weight distribution of spatial position through a convolution layer to focus on local spatial area where abnormal sound occurs and suppress noise interference in irrelevant spatial area.
2. The method of claim 1, wherein, The N joint calibration feature maps are generated by a feature extraction module of a sample expansion model, and the feature extraction module at least comprises N feature extraction layers; performing feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration on the noisy abnormal sound spectrum graph in turn N times to generate N joint calibration feature maps layer by layer, comprising: generating a preliminary feature map by performing feature encoding on the (n-1)th joint calibration feature map through the nth feature extraction layer, wherein 2≤n≤N; generating weight parameters of each channel feature in the preliminary feature map according to the dependency relationship between each channel in the preliminary feature map, and generating a channel dimension calibration feature map through weighted operation; for each channel of the channel dimension calibration feature map, width aggregation features and height aggregation features of the channel are extracted along the row dimension and the column dimension respectively, and row weights and column weights of each channel are generated through a position attention mechanism, and the nth joint calibration feature map is generated through weighted operation.
3. The method of claim 2, wherein, Each feature extraction layer at least comprises a channel dimension adaptive calibration layer; the generation of the weight parameters of each channel feature in the preliminary feature map comprises: for the preliminary feature map, the global feature value of each feature channel is calculated through the channel dimension adaptive calibration layer; according to the global feature value of each feature channel, the local interaction of each channel with its k adjacent channels is captured to generate the weight parameters of each channel feature in the preliminary feature map.
4. The method of claim 2, wherein, Each feature extraction layer at least further comprises a spatial dimension adaptive calibration layer; the generation of the row weights and the column weights of each channel comprises: for each channel of the channel dimension calibration feature map, the local feature value of the channel dimension calibration feature map at each position is calculated through the spatial dimension adaptive calibration layer; according to the local feature value of the channel dimension calibration feature map at each position, the width aggregation features and the height aggregation features of each channel are determined: Along the channel dimension, the channel dimension calibration feature map is spliced with the width aggregation features and the height aggregation features of each channel to obtain a spatial statistical feature matrix of each channel, and a position attention coordinate of each channel is generated, the position attention coordinate of each channel representing an attention intensity of the channel to different positions; According to the row attention coordinate and the column attention coordinate of the position attention coordinate of each channel, a row weight and a column weight of the channel are respectively generated, to obtain the row weight and the column weight of each channel.
5. A method for diagnosing abnormal noise of a component, characterized by, The method comprises: inputting the audio image to be diagnosed into an abnormal sound fault diagnosis model to obtain an initial feature map; According to the dependency relationship between each channel of the initial feature map, the abnormal sound fault diagnosis model generates a channel weight, which represents the correlation between the channel and the abnormal sound feature; According to each channel weight, the abnormal sound fault diagnosis model performs weighted adjustment on the audio image to be diagnosed to obtain a weighted feature map; According to the weighted feature map, the abnormal sound fault diagnosis model performs fault diagnosis to output a fault diagnosis result; The abnormal sound fault diagnosis model is trained based on an abnormal sound sample set, and each abnormal sound sample in the abnormal sound sample set is generated according to the abnormal sound sample generation method in any one of claims 1 to 4.
6. The method of claim 5, wherein, The method further comprises: collecting normal audio data of parts to generate a normal abnormal sound sample set; According to the true result, each abnormal sound sample in the abnormal sound sample set and each normal sample in the normal abnormal sound sample set are labeled to obtain a normal training sample set and an abnormal training sample set; Each training sample in the normal training sample set and the abnormal training sample set is input into a pre-trained abnormal sound fault diagnosis model for model training to obtain a diagnosis result of each training sample; According to the difference between the diagnosis result of each training sample and the true result of the training sample, the model parameters of the pre-trained abnormal sound fault diagnosis model are updated to obtain a trained abnormal sound fault diagnosis model.
7. The method of claim 5, wherein, The abnormal sound fault diagnosis model at least comprises a channel dimension adaptive calibration layer; According to the dependency relationship between each channel of the initial feature map, the abnormal sound fault diagnosis model generates a channel weight, which comprises: For the initial feature map, the abnormal sound fault diagnosis model calculates a global feature value of each feature channel through the channel dimension adaptive calibration layer; Based on the global feature value of each feature channel, the abnormal sound fault diagnosis model captures the local interaction between each channel and its m adjacent channels through the channel dimension adaptive calibration layer to generate a channel weight of the initial feature map.
8. An abnormal sound sample generation device characterized by comprising: It comprises: The superposition module is used for superimposing target noise on the real abnormal sound audio data spectrum graph to generate a noisy abnormal sound spectrum graph; The feature extraction module is used for sequentially performing N times of feature encoding, channel dimension adaptive calibration and spatial dimension adaptive calibration on the noisy abnormal sound spectrum graph through the feature extraction module of the sample expansion model to generate N joint calibration feature maps layer by layer, each joint calibration feature map at least comprising feature information fused by suppressing the target noise and enhancing the real abnormal sound audio data; The feature reconstruction module is configured to perform N times of upsampling operations on the Nth joint calibration feature map after feature coding by the feature reconstruction module of the sample expansion model, and each upsampling operation is combined with the corresponding joint calibration feature map to perform feature fusion, thereby generating an abnormal sound sample set highly close to the real abnormal sound data. The channel dimension adaptive calibration is achieved by generating channel-level statistical features through global average pooling, and then learning the importance weights of each channel through a fully connected layer or a one-dimensional convolution, thereby enhancing the channels strongly related to abnormal sound and suppressing the channels strongly related to noise. The spatial dimension adaptive calibration is achieved by performing an average pooling operation along the channel dimension after the channel calibration, and then learning the weight distribution of the spatial position through a convolution layer, thereby focusing on the local spatial area where the abnormal sound occurs and suppressing the noise interference in the irrelevant spatial area.
9. A device for diagnosing abnormal noises from components, characterized in that, The method comprises: inputting the audio graph to be diagnosed into the abnormal sound fault diagnosis model to obtain an initial feature map; a weight generation module configured to generate channel weights according to the dependency relationship between each channel of the initial feature map by the abnormal sound fault diagnosis model, wherein the channel weights represent the correlation between the channel and the abnormal sound feature; a weighting adjustment module configured to perform weighting adjustment on the audio graph to be diagnosed according to each channel weight by the abnormal sound fault diagnosis model to obtain a weighted feature map; a fault diagnosis module configured to perform fault diagnosis based on the weighted feature map by the abnormal sound fault diagnosis model, and output a fault diagnosis result; wherein the abnormal sound fault diagnosis model is trained based on an abnormal sound sample set, and each abnormal sound sample in the abnormal sound sample set is generated according to the abnormal sound sample generation method of any one of claims 1 to 4.
10. An electronic device, comprising: The method comprises: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the abnormal sound sample generation method of any one of claims 1 to 4, or the part abnormal sound fault diagnosis method of any one of claims 5 to 7.
Citation Information
Patent Citations
Small sample rolling bearing fault diagnosis method based on time-frequency enhancement and generative learning
CN118626929A
Complete fault vibration signal generation method for elevator guide wheel bearing
CN118839587A