Machine abnormal sound detection method, system and device
By fusing delta features in logarithmic mel spectrograms to capture short-term dynamic evolution and transient changes of audio signals in machine equipment, the problem of difficulty in capturing transient signals in edge computing scenarios in the prior art is solved, and efficient and accurate detection of machine abnormal sounds is achieved.
Patent Information
- Application Number
- CN202510363001.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-10
AI Technical Summary
In edge computing application scenarios where resources are constrained or have strict requirements on real-time, it is difficult to effectively capture bursty and transient signal characteristics in machines and equipment, resulting in complex detection models and weak detection capabilities.
By fusing delta features on the basis of the logarithmic mel spectral diagram, a multi-channel logarithmic mel spectral representation is formed, short-term dynamic evolution and transient changes in the audio signal are captured, and input it into the sound detection algorithm framework for abnormal sound detection.
It realizes improving the accuracy of machine abnormal sound detection under limited resource conditions, simplifies the model architecture, enables the model to run efficiently, and improves the stability and security of the system.
Smart Images

Figure CN120126488A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of sound detection algorithm framework architecture, and in particular to a method, system and device for detecting abnormal sound of a machine. Background Art
[0002] At present, although the sound anomaly detection algorithm based on deep learning has made significant progress, it still faces many difficulties when resources are constrained or there are strict requirements for real-time performance, especially in the widely used edge computing application scenarios, such as sensor networks, embedded devices and IoT terminals, whose hardware resource configuration is often extremely limited, and generally exhibits characteristics such as weak computing power, small storage space and controlled energy supply. Therefore, when building an anomaly detection model that adapts to this environment, it is necessary not only to improve the detection accuracy, but also to simplify the model architecture and reduce resource consumption. This is the key issue that needs to be solved in the current research on abnormal sound detection.
[0003] In the existing abnormal sound detection research, most methods use logarithmic Mel-spectrogram as a preprocessing method before model input. Although this representation method is good at capturing time-frequency domain features, it usually calculates the spectrum based on a fixed time window and has certain limitations in revealing the characteristics of sudden and transient signals. In practical applications, sudden abnormal events appear in the form of transient signals. Therefore, capturing transient signals is crucial for abnormal sound detection. Summary of the invention
[0004] The embodiments of the present application provide a method, system and device for detecting abnormal sounds of machines. In order to deeply explore the subtle features in the audio signal, delta features are integrated on the basis of the logarithmic Mel spectrum diagram to form a multi-channel logarithmic Mel spectrum representation, which aims to not only capture conventional time-frequency domain information, but also keenly capture the short-term dynamic evolution and transient changes in the audio signal. It solves the technical problems in the prior art of detecting abnormal sounds of machine equipment, such as certain limitations in the detection of sudden and transient signal characteristics, high complexity of the sound detection model, and weak detection capabilities. It achieves the technical effect of improving the detection accuracy and simplifying the model architecture when constructing a model for detecting abnormal sounds of machines, so that the model can run efficiently under limited resource conditions, and can more effectively improve the stability and security of the entire system.
[0005] In a first aspect, an embodiment of the present application provides a method for detecting abnormal machine sounds, including: obtaining an audio signal generated when the machine is operating; performing a preprocessing operation on the audio signal to obtain a logarithmic Mel spectrogram; performing a difference calculation on the logarithmic Mel spectrogram to obtain delta features, and splicing the logarithmic Mel spectrogram and the delta features to obtain a spliced logarithmic Mel spectrogram, where the delta features include first-order delta features and second-order delta features; constructing a sound detection algorithm framework, and inputting the spliced logarithmic Mel spectrogram into the sound detection algorithm framework to obtain an abnormal sound detection result.
[0006] In combination with the first aspect, in a possible implementation manner, the performing a preprocessing operation on the audio signal to obtain a logarithmic Mel spectrogram includes: performing a discrete Fourier transform on the audio signal to convert the time-domain signal to the frequency domain to obtain a discrete Fourier transform result; extracting an amplitude spectrum from the discrete Fourier transform result, and performing Mel filter bank filtering on the amplitude spectrum to obtain a filtering result; performing a normalization process according to the filtering result and the amplitude of the filtering result to obtain a logarithmic Mel spectrogram.
[0007] In combination with the first aspect, in a second possible implementation manner, the performing a difference calculation on the logarithmic Mel spectrogram to obtain delta features includes: performing a first-order difference calculation on the logarithmic Mel spectrogram to obtain first-order delta features; performing a second-order difference calculation based on the first-order delta features to obtain second-order delta features.
[0008] In combination with the first aspect, in a third possible implementation manner, the sound detection algorithm framework includes a generator network and a discriminator network.
[0009] In combination with the third possible implementation manner of the first aspect, in a fourth possible implementation manner, the inputting the spliced logarithmic Mel spectrogram into the sound detection algorithm framework to obtain an abnormal sound detection result includes: inputting the spliced logarithmic Mel spectrogram into the generator network in the sound detection algorithm framework for reconstruction to obtain a reconstruction result; performing abnormal discrimination on the reconstruction result based on the discriminator network to obtain an abnormal sound detection result.
[0010] In a second aspect, an embodiment of the present application provides a machine abnormal sound detection system, including: an audio acquisition module, configured to acquire an audio signal generated when the machine is operating; a data processing module, configured to preprocess the audio signal to obtain a logarithmic mel spectrogram; a differential calculation module, configured to perform differential calculation on the logarithmic mel spectrogram to obtain delta features, and splice the logarithmic mel spectrogram and the delta features to obtain a spliced logarithmic mel spectrogram; a sound detection algorithm framework, including a generator network and a discriminator network, wherein the generator network is configured to reconstruct the spliced logarithmic mel spectrogram to obtain a reconstruction result, and the discriminator network is configured to perform abnormal discrimination on the reconstruction result to obtain an abnormal sound detection result; a user interaction module, configured to visualize the abnormal sound detection result and respond to an operation request of a client.
[0011] In combination with the second aspect, in a possible implementation manner, the generator network includes a plurality of rBlock convolutional blocks; the rBlock convolutional block includes a plurality of convolutional layers, a normalization layer, and a Leaky ReLU activation function, and the rBlock convolutional block is configured to extract features of the spliced logarithmic mel spectrogram and reconstruct the spliced logarithmic mel spectrogram to obtain a reconstruction result.
[0012] In combination with the second aspect, in a second possible implementation manner, the discriminator network includes a plurality of rBlock convolutional blocks, a linear linear downsampling module, and a sigmoid activation function; the linear linear downsampling module is configured to gradually reduce the resolution of the reconstruction result and extract features of the reconstruction result; the discriminator network performs abnormal discrimination based on the features of the reconstruction result to obtain an abnormal sound detection result, and optimizes and trains the process of the generator network reconstructing the spliced logarithmic mel spectrogram.
[0013] In a third aspect, an embodiment of the present application provides a machine abnormal sound detection device, including: a receiving unit, configured to receive an audio signal generated by the machine in a working mode; a preprocessing unit, configured to preprocess the audio signal to obtain a logarithmic mel spectrogram; a feature extraction unit, configured to obtain delta features according to the logarithmic mel spectrogram; a splicing unit, configured to splice the logarithmic mel spectrogram and the delta features to obtain a spliced logarithmic mel spectrogram; an abnormal sound detection unit, configured to obtain an abnormal sound detection result according to the spliced logarithmic mel spectrogram.
[0014] In a fourth aspect, an embodiment of the present application provides a device, the device includes: a processor; a memory for storing processor-executable instructions; when the processor executes the executable instructions, the method described in the first aspect or any possible implementation manner of the first aspect is implemented.
[0015] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: In the embodiments of the present application, the audio signal generated during the operation of the machine is processed to obtain a logarithmic mel spectrogram; the delta feature is obtained based on the logarithmic mel spectrogram, and the logarithmic mel spectrogram and the delta feature are spliced; the result is input into the sound detection algorithm framework constructed in the embodiments of the present application to obtain the result of the detection model; the feature extraction is performed on the result of the detection model to obtain the feature representation; the logarithmic mel spectrogram and the feature representation are reconstructed, and finally an abnormal sound detection report is generated. The present application effectively solves the technical problem that it is difficult to capture transient and short-term audio signals in the traditional machine sound detection process, and further realizes the technical effects of reducing the amount of calculation, conveniently deploying mobile machine terminals, and keenly capturing the short-term dynamic evolution in the audio signal during the process of capturing and diagnosing machine fault sounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is a flowchart of a method for detecting abnormal machine sounds provided by the embodiments of the present application; Figure 2 It is a system block diagram of a system for detecting abnormal machine sounds provided by the embodiments of the present application; Figure 3 It is a block diagram of a sound detection algorithm framework provided by the embodiments of the present application; Figure 4 It is an example diagram of the convolutional block network structure provided by the embodiments of the present application (the left figure is an example diagram of the traditional convolutional block network structure, and the right figure is an example diagram of the convolutional block network structure provided by the present application); Figure 5 It is an example diagram of the network structure of the generator network provided by the embodiments of the present application; Figure 6 It is an example diagram of the network structure of the discriminator network provided by the embodiments of the present application; Figure 7 It is an example diagram of the network structure of the lightweight CMT channel attention provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] Figure 1 This is a flowchart of a method for detecting abnormal machine sounds provided by an embodiment of the present application. As Figure 1 shown, the present application collects, preprocesses, judges, and generates an abnormal sound detection report for the audio signal during the operation of factory equipment based on a sound detection algorithm framework. Specifically: S1: Collect the sounds generated by factory equipment based on a mobile terminal with an audio collection function to obtain the audio signal generated during machine operation.
[0020] It should be noted that this audio signal is an analog signal, and when in use, it needs to be converted into a digital signal for subsequent calculations.
[0021] S2: Perform preprocessing operations on the audio signal to obtain a log Mel spectrogram. Specifically: Perform a discrete Fourier transform on the audio signal. At this time, the purpose of performing the discrete Fourier transform is to convert the time-domain signal in the audio signal to the frequency domain to obtain the discrete Fourier transform result. Among them, this result is a complex number array, representing the amplitude and phase of each frame of the audio signal at different frequency components.
[0022] Extract the amplitude spectrum from this discrete Fourier transform and perform Mel filter bank filtering on this amplitude spectrum to obtain a filtering result. This filtering result can reflect the energy distribution of the audio signal on the Mel scale.
[0023] Perform normalization processing according to the filtering result and the amplitude spectrum of the filtering result to obtain a log Mel spectrogram.
[0024] Those skilled in the art should be aware that the Mel spectrogram is a representation method that mimics the sound frequency perception characteristics of the human ear and is usually extracted from the original audio signal through steps such as discrete Fourier transform, Mel filter bank, and amplitude normalization, expressed as: (1) In formula (1), represents the value of the Mel spectrogram at time index t and frequency index f, STFT(x[n]) is the short-time discrete Fourier transform result of the input audio signal x[n], and Mel(f) is the Mel scale mapping function.
[0025] The mel spectrogram mainly focuses on the spectral statistical characteristics over a long time period. For short-term and sudden abnormal signals, due to its window length and overlapping processing, these transient information may not be completely retained. However, transient information is a characteristic feature of machine sudden anomalies. Therefore, in this embodiment, by fusing delta features with the mel spectrogram, time-frequency domain information and auxiliary information are provided for the device abnormal sound detection algorithm framework.
[0026] In the process of detecting abnormal sounds of machine equipment provided by the embodiments of this application, the discrete Fourier transform is used to process audio signals because the discrete Fourier transform can clearly show the frequency components and amplitude distribution of signals when processing periodic signals, and can be efficiently calculated through the fast discrete Fourier transform algorithm. Its lower computational complexity is more suitable for extracting the frequency domain features of machine signals.
[0027] S3: Obtain delta features based on the logarithmic mel spectrogram, splice the logarithmic mel spectrogram and the delta features to obtain the spliced logarithmic mel spectrogram. Specifically: Perform a first-order difference calculation on the logarithmic mel spectrogram to obtain the first-order delta features.
[0028] Perform a second-order difference calculation on the first-order delta features to obtain the second-order delta features.
[0029] It should be noted that the first-order delta features and the second-order delta features are collectively referred to as delta features.
[0030] Those skilled in the art should be aware that delta features generally refer to the features obtained after performing a difference operation on the features in the original audio signal (such as MFCC, that is, Mel Frequency Cepstral Coefficients). The MFCC feature is an important acoustic feature extracted from speech signals, which reflects the spectral envelope information of speech. In order to capture the time-varying characteristics of speech signals, a difference operation is usually performed on the MFCC features to generate first-order difference (delta) features and second-order difference (delta-delta) features.
[0031] Exemplarily, the first-order difference features can reflect the rate of change of data over time, capturing the subtle changes of sound signals on the time axis, which plays a certain role in detecting mutation points, acceleration or deceleration processes in signals. In many anomaly detection scenarios, abnormal events are often accompanied by rapid changes in data, such as increased vibration or sudden temperature rise before equipment failure.
[0032] The second-order difference features further reveal the change in the rate of signal change, highlighting the speed and acceleration of the spectral change of sound signals, which is crucial for detecting sharp and instantaneous transitions, such as: impact sounds and explosion sounds generated by mechanical failures.
[0033] Specifically, in order to further capture the time dynamic characteristics of the audio signal, the first-order difference feature and the second-order difference feature of the Mel spectrogram can be calculated respectively. The first-order difference feature is shown in formula (2).
[0034] (2) Based on formula (2), the second-order difference feature is calculated as follows: (3) Combining and simplifying formula (2) and formula (3), we get: (4) In formula (4), MelS(n) represents the Mel spectrum value of the nth frame. represents the first-order difference result from the nth frame to the n + 1th frame.
[0035] S4: Construct a sound detection algorithm framework, and input the spliced logarithmic Mel spectrogram into the sound detection algorithm framework to obtain the abnormal sound detection result. Specifically: Input the spliced logarithmic Mel spectrogram into the generator network in the sound detection algorithm framework for reconstruction to obtain the reconstruction result.
[0036] Based on the discriminator network, perform abnormal discrimination on the reconstruction result to obtain the abnormal sound detection result.
[0037] In the embodiment of the present application, the audio signal generated during the operation of the machine is processed to obtain a logarithmic Mel spectrogram; the delta feature is obtained based on the logarithmic Mel spectrogram, and the logarithmic Mel spectrogram and the delta feature are spliced; the result is input into the sound detection algorithm framework constructed in the embodiment of the present application to obtain the detection model result; the feature representation is obtained by extracting features from the detection model result; the logarithmic Mel spectrogram and the feature representation are reconstructed, and finally an abnormal sound detection report is generated. The present application effectively solves the technical problem that it is difficult to capture transient and short-term audio signals in the traditional machine sound detection process, and further realizes the technical effects of reducing the calculation amount, conveniently deploying mobile machine terminals, and keenly capturing the short-term dynamic evolution in the audio signal during the process of capturing and diagnosing machine fault sounds.
[0038] Figure 2 This is the system block diagram of a machine abnormal sound detection system provided by the embodiment of the present application. As Figure 2 shown, the present application proposes a machine abnormal sound detection system specifically for the problem of abnormal sound detection of factory machine equipment, which includes: An audio acquisition module for acquiring the audio signal generated during the operation of the machine.
[0039] A data processing module for preprocessing an audio signal to obtain a log Mel spectrogram. The data processing module converts an audio signal with an analog signal as the signal source into a digital signal suitable for a sound detection algorithm framework. It should be noted that the preprocessing process is a commonly used technical means in the art, so it will not be described in detail here.
[0040] A differential calculation module for performing differential calculation on the log Mel spectrogram to obtain delta features, and splicing the log Mel spectrogram and the delta features to obtain a spliced log Mel spectrogram.
[0041] A sound detection algorithm framework for generating samples of an audio signal and determining whether an abnormal sound appears. The sound detection algorithm framework includes a generator network and a discriminator network. Among them, the generator network is used to reconstruct the log Mel spectrogram to obtain a reconstruction result, and the discriminator network is used to perform abnormal discrimination on the reconstruction result to obtain an abnormal sound detection result. The discriminator network is also used to optimize the generator network based on the discrimination result.
[0042] A user interaction module for visualizing the abnormal sound detection result and responding to the operation request of the client.
[0043] Exemplarily, in the embodiment of the present application, a high-resolution abnormal sound detection algorithm framework Light GAN is proposed, as Figure 3 shown.
[0044] Figure 3 It is a block diagram of the sound detection algorithm framework provided by the embodiment of the present application. The algorithm framework includes a generator network and a discriminator network.
[0045] The generator network includes one or more rBlock convolutional blocks. The rBlock convolutional block contains one or more convolutional layers, a normalization layer, and a Leaky ReLU activation function. The rBlock convolutional block is used to extract the features of the spliced log Mel spectrogram and reconstruct the spliced log Mel spectrogram to obtain a reconstruction result.
[0046] The discriminator network includes one or more rBlock convolutional blocks, a linear linear downsampling module, and a sigmoid activation function.
[0047] The linear linear downsampling module is used to gradually reduce the resolution of the reconstruction result and extract the features of the reconstruction result.
[0048] Based on the features of the reconstruction result, the discriminator network performs abnormal discrimination to obtain an abnormal sound detection result, and optimizes the training process of the generator network for reconstructing the spliced log Mel spectrogram.
[0049] It should be noted that the generator network also includes a lightweight encoder and a lightweight decoder, and the discriminator network also includes a lightweight encoder and a CMT-Att (attention in three aspects of Channel, Mel spectrogram, and Time) module.
[0050] The steps to establish the generator network are as follows: First, design a lightweight convolutional block to replace the original convolutional block, and pass multi-scale feature information through skip connections in the generator, aiming to extract rich time-frequency domain information and perform sample reconstruction.
[0051] Second, streamline the CMT-Att attention mechanism network structure to further lightweight the model.
[0052] Finally, introduce delta features to capture the short-term trends and transient changes of audio signals and enhance the expressive ability of the model. The specific operation is as follows: calculate the first-order difference features and second-order difference features of the log Mel spectrogram respectively, and splice them with the original log Mel spectrogram along the channel dimension to reconstruct a three-channel log Mel spectrogram. The processes of calculating the first-order difference features and second-order difference features of the log Mel spectrogram are shown in formulas (2) and (3).
[0053] It should be noted that in the embodiments of the present application, an improved rBlock depthwise separable convolutional block is used to replace the ordinary convolution in the generator, aiming to greatly reduce the computational complexity and parameter scale of the convolutional neural network on the premise of maintaining high model performance. The specific operation is as follows.
[0054] Depthwise separable convolution decomposes the standard convolution into two parts: depthwise convolution and pointwise convolution. Depthwise convolution focuses on extracting spatial features within the same channel; pointwise convolution focuses on cross-channel feature integration and transformation. Compared with the traditional convolution operation process, depthwise separable convolution optimizes the resource requirements in the calculation process of the convolutional neural network to a certain extent, and can encode the input features by reducing the effective parameters.
[0055] The lightweight convolutional module adopted in the embodiments of the present application is designed as an rBlock convolutional block, and its structural schematic is as Figure 4 shown.
[0056] Figure 4This is an example diagram of the convolutional block network structure provided by the embodiments of the present application. The left figure is an example diagram of the overall structure of a traditional convolutional neural network model, and the right figure is an example diagram of the overall framework structure of the convolutional neural network model proposed by the embodiments of the present application. Among them, the core architecture design of the overall framework of the traditional convolutional neural network model is mainly the MobileNet V1 architecture. During the application of the MobileNet V1 architecture, a large number of zero values frequently appear in the output feature layers. These zero values hinder the effective propagation and backpropagation of gradients. The root cause of this problem is that the ReLU activation function may produce a saturation effect, resulting in a large number of zero values in the output feature layers.
[0057] As Figure 4 shown in the right figure of [], for this situation that occurs in the traditional convolutional neural network model, the embodiments of the present application propose a method of using the Leaky ReLU activation function to replace the traditional ReLU function, and at the same time introduce a linear layer, in order to minimize the adverse phenomenon that the ReLU activation function may produce a saturation effect, and ensure that the gradients can fully flow and update the weights during the training process of the convolutional neural network model.
[0058] The functional service of the generator network based on rBlock provided by the embodiments of the present application is for the reconstruction task of samples, and its structure is as Figure 5 shown. Specifically, the generator starts with an initial ordinary convolutional layer, and the convolutional layer is responsible for mapping the input 3-channel spatial data to a 64-channel feature space of a higher dimension. Subsequently, a fully convolutional encoder-decoder architecture based on depthwise separable convolution and bilinear upsampling / downsampling operations is used to generate a feature tensor with a dimension of H×W×64. All convolutional blocks in the entire network are combined with the Leaky ReLU activation function to enhance the non-linear expression ability. In addition, based on the design concept of U-net, the embodiments of the present application introduce a skip connection mechanism between different levels of the encoder and the decoder, so that the feature maps captured by the encoder at different resolutions can be directly fed to the corresponding levels of the decoder. This mechanism is crucial during the upsampling restoration process of the decoder, and it can ensure that low-level but detail-rich feature information is effectively transmitted to the decoder, thereby enriching the detail expressiveness and fidelity of the reconstructed samples.
[0059] As Figure 6 shown, this is an example diagram of the network structure of the discriminator network provided by the embodiments of the present application. The main function of the discriminator network is to accurately distinguish between the original samples and the reconstructed feature maps, and perform targeted lightweight processing on the convolutional blocks used inside the discriminator. The embodiments of the present application replace the ordinary convolutional blocks in the prior art based on rBlock.
[0060] To further ensure the overall lightweight characteristics of the model, the embodiments of the present application also implement corresponding lightweight improvements for the CMT-Att module. As Figure 7 shown, it is the specific detailed design of the lightweight improvement, where GAP (Global Average Pooling) represents global average pooling. GMP (Goroutine, Machine, and Processor) is a model for dispatch. The specific implementation of the lightweight improvement is as follows: Perform global max pooling and global average pooling operations on the feature map.
[0061] Perform concatenation on the feature map and perform non-linear operations.
[0062] Simplify the linear layer in the convolutional neural network, and use the convolutional layer to replace the original linear layer, thereby further reducing the number of model parameters.
[0063] The embodiments of the present application also provide a machine abnormal sound detection device. It should be noted that this machine abnormal sound detection device is applied based on the machine abnormal sound detection method. It includes: A receiving unit for receiving the audio signal generated by the machine in the working mode.
[0064] A preprocessing unit for preprocessing the audio signal to obtain a log Mel spectrogram.
[0065] A feature extraction unit for obtaining delta features according to the log Mel spectrogram.
[0066] A splicing unit for performing splicing on the log Mel spectrogram and delta features and obtaining the spliced log Mel spectrogram.
[0067] An abnormal sound detection unit for generating an abnormal sound detection report according to the spliced log Mel spectrogram.
[0068] The embodiments of the present application provide a lightweight abnormal sound detection algorithm framework Light GAN for the problem of abnormal sound detection of factory machine equipment. Based on the generative adversarial network, this structure integrates a lightweight convolutional block rBlock constructed by depthwise separable convolution and combines the idea of U-Net skip connections. On the one hand, it retains the high-level abstract features refined by depthwise separable convolution, and on the other hand, it can fully utilize the low-level detailed features to generate more accurate and detailed segmentation results during the upsampling process.
[0069] Convolutional neural networks are widely used in deep learning models due to their excellent performance in extracting image features. However, the significant increase in their computational complexity and number of parameters inevitably limits their use in resource-constrained environments, especially in small mobile devices and embedded devices. Traditional convolution operations involve a large number of weight parameters and calculation operations when processing each pixel of an image, which places high demands on memory usage and computing efficiency.
[0070] In the embodiments of the present application, the core technical means used can be summarized as follows: 1) A lightweight sound detection algorithm framework LightGAN is proposed for the problem of abnormal sound detection of factory machinery and equipment.
[0071] 2) In order to deeply explore the subtle features in the audio signal, delta features are fused on the basis of the logarithmic Mel spectrum to form a multi-channel logarithmic Mel spectrum representation, which aims to not only capture the conventional time-frequency domain information, but also keenly capture the short-term dynamic evolution and transient changes in the audio signal.
[0072] 3) In order to minimize the computational cost and ensure the model performance is easy to deploy on mobile machine terminals, a lightweight convolution module is used to replace the traditional convolution layer, thereby improving the detection performance while ensuring the efficient operation of the model.
[0073] Although the present application provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative labor. The order of steps listed in this embodiment is only one way of executing the order of many steps and does not represent the only execution order. When the actual device or client product is executed, it can be executed in sequence or in parallel according to the method shown in this embodiment or the accompanying drawings (for example, in a parallel processor or multi-threaded processing environment).
[0074] Some modules in the apparatus described in the present application can be described in the general context of computer executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0075] The devices or modules described in the above application embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, when describing the above devices, they are divided into various modules according to functions and described separately. When implementing the application embodiments, the functions of each module can be implemented in the same or multiple software and / or hardware. Of course, the module implementing a certain function can also be implemented by combining multiple sub-modules or sub-units.
[0076] The methods, devices or modules described in this application can be implemented in the form of computer-readable program codes. The controller can be implemented in any appropriate manner. For example, the controller can take the form of, for example, a microprocessor or a processor, and a computer-readable medium storing computer-readable program codes (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0077] The embodiments of the present application also provide a device, which includes: a processor; a memory for storing executable instructions of the processor; when the processor executes the executable instructions, the method described in the embodiments of the present application is implemented.
[0078] In addition, in each embodiment of the present invention, the functional modules can be integrated into one processing module, or each module can exist independently, or two or more modules can be integrated into one module.
[0079] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product or can also be reflected in the implementation process of data migration. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0080] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. All or part of this application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0081] The above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit this application; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.
Claims
1. A method for detecting abnormal sound of a machine, characterized in that: include: Acquire the audio signal generated when the machine is running; Perform preprocessing operations on the audio signal to obtain a logarithmic Mel-spectrogram; Perform differential calculation on the logarithmic Mel spectrum graph to obtain delta features, and concatenate the logarithmic Mel spectrum graph and the delta features to obtain a concatenated logarithmic Mel spectrum graph, wherein the delta features include first-order delta features and second-order delta features; Construct a sound detection algorithm framework, and input the concatenated logarithmic Mel-spectrogram into the sound detection algorithm framework to obtain abnormal sound detection results.
2. The method according to claim 1, characterized in that The preprocessing operation is performed on the audio signal to obtain a logarithmic Mel-frequency spectrum diagram, including: Performing a discrete Fourier transform on the audio signal to convert the time domain signal into the frequency domain to obtain a discrete Fourier transform result; Extract the amplitude spectrum from the discrete Fourier transform result, and perform Mel filter bank filtering on the amplitude spectrum to obtain the filtering result; Normalization processing is performed according to the filtering result and the amplitude of the filtering result to obtain a logarithmic Mel spectrum graph.
3. The method according to claim 1, characterized in that The delta feature is obtained by performing differential calculation on the logarithmic Mel spectrum graph, including: Perform first-order difference calculation on the logarithmic Mel spectrum to obtain the first-order delta feature; A second-order difference calculation is performed based on the first-order delta feature to obtain the second-order delta feature.
4. The method according to claim 1, characterized in that: The sound detection algorithm framework includes a generator network and a discriminator network.
5. The method according to claim 4, characterized in that The spliced logarithmic Mel-spectrogram is input into the sound detection algorithm framework to obtain abnormal sound detection results, including: The concatenated logarithmic Mel-spectrogram is input into the generator network in the sound detection algorithm framework for reconstruction to obtain a reconstruction result; The reconstruction result is judged as abnormal based on the discriminator network to obtain the abnormal sound detection result.
6. A machine abnormal sound detection system, characterized in that: include: An audio acquisition module, used to acquire the audio signal generated when the machine is running; A data processing module is used to pre-process the audio signal to obtain a logarithmic Mel spectrum graph; A differential calculation module is used to perform differential calculation on the logarithmic Mel spectrum graph to obtain delta features, and to concatenate the logarithmic Mel spectrum graph and the delta features to obtain a concatenated logarithmic Mel spectrum graph; The sound detection algorithm framework includes a generator network and a discriminator network, wherein the generator network is used to reconstruct the concatenated logarithmic Mel-spectrogram to obtain a reconstruction result, and the discriminator network is used to perform abnormal discrimination on the reconstruction result to obtain an abnormal sound detection result; The user interaction module is used to visualize the abnormal sound detection results and respond to the operation requests of the user end.
7. The machine abnormal sound detection system according to claim 6, characterized in that: The generator network includes several rBlock convolution blocks; The rBlock convolution block contains several convolution layers, normalization layers and Leaky ReLU activation functions. The rBlock convolution block is used to extract the features of the concatenated log-Mel spectrum graph and reconstruct the concatenated log-Mel spectrum graph to obtain the reconstruction result.
8. The machine abnormal sound detection system according to claim 6, characterized in that: The discriminator network includes several rBlock convolution blocks, linear downsampling modules and sigmoid activation functions; The linear downsampling module is used to gradually reduce the resolution of the reconstruction result and extract the features of the reconstruction result; The discriminator network performs abnormal discrimination based on the features of the reconstruction results to obtain abnormal sound detection results, and optimizes the training process of the generator network's reconstructed and spliced logarithmic Mel-spectrogram.
9. A device for detecting abnormal sound of a machine, characterized in that: include: A receiving unit, used for receiving an audio signal generated by the machine in a working mode; A preprocessing unit, used for preprocessing the audio signal to obtain a logarithmic Mel spectrum graph; A feature extraction unit, used to obtain delta features according to the logarithmic Mel spectrum graph; A concatenation unit, used for concatenating the logarithmic Mel spectrum graph and the delta feature to obtain a concatenated logarithmic Mel spectrum graph; The abnormal sound detection unit is used to obtain an abnormal sound detection result according to the spliced logarithmic Mel-spectrogram.
10. A device, characterized in that: include: processor; a memory for storing processor-executable instructions; When the processor executes the executable instructions, the method according to any one of claims 1 to 5 is implemented.