Microphone anomaly detection method based on LSTM-FCN network
Real-time abnormality detection of microphones is performed through the LSTM-FCN network and the minimum mean square error noise reduction algorithm, which solves the problems of inaccurate and unreal-time detection in the prior art, and realizes efficient microphone status monitoring and abnormal response, improving the stability and security of the Internet of Things system.
Patent Information
- Application Number
- CN202510617504.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-05
AI Technical Summary
The existing microphone anomaly detection methods are not accurate and robust enough when processing complex or variable environmental data, lack real-time and adaptability, and are difficult to respond to rapid changes in device status in a timely manner, affecting the stability and security of IoT systems.
The LSTM-FCN network is used to conduct in-depth analysis of the hardware characteristics and real-time data flow of the microphone, combined with the minimum mean square error noise reduction algorithm and classifier training, to monitor and identify the abnormal state of the microphone in real time, and improve signal quality and detection accuracy through preprocessing and spectrum analysis.
Real-time and adaptive microphone abnormality detection is realized, which improves the accuracy and efficiency of detection, enhances the stability and security of IoT devices, and provides user-friendly operation experience and system integration.
Smart Images

Figure CN120434579A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of abnormality monitoring for Internet of Things (IoT) devices, focusing on real-time monitoring of microphone status through audio signal processing technology. By employing advanced methods such as LSTM-FCN networks to conduct in-depth analysis and comparison of collected audio data, the invention can effectively determine whether microphone operating status is abnormal. This method not only improves monitoring accuracy but also plays a significant role in timely diagnosis and response to potential device issues, thus playing a key role in ensuring the stability and reliability of IoT systems. Background Art
[0002] With the rapid development and widespread adoption of the Internet of Things (IoT) technology, a large number of smart devices are being deployed in various environments to collect and transmit critical data. These devices, which contain microphones, play a vital role in fields such as speech recognition, intelligent surveillance, and industrial automation. However, due to various factors such as environmental changes, device aging, and external interference, microphones may experience performance degradation or failure, resulting in data anomalies. This not only affects system accuracy and reliability but can also lead to safety risks and financial losses.
[0003] In IoT systems, an abnormal microphone state typically refers to a situation where the sound signal it outputs differs significantly from the signal in normal operation. This discrepancy may be caused by internal device failures (such as damage to the microphone's electronic components or deformation of the acoustic structure), external environmental interference (such as strong electromagnetic interference, extreme temperature fluctuations, high humidity), or improper device installation (such as proximity to strong noise sources or obstruction by objects). These abnormal conditions can lead to misjudgments in speech recognition systems, false alarms or missed alarms in intelligent monitoring systems, and data collection errors in industrial automation systems, thus impacting the normal operation of the entire IoT system.
[0004] Traditional microphone anomaly detection methods typically rely on fixed thresholds or simple statistical analysis. These methods can be inaccurate or inflexible when dealing with complex or changing environmental data. Furthermore, these methods often lack real-time and adaptability, making it difficult to respond to rapid changes in device status. Therefore, developing a technology that can accurately and adaptively detect and diagnose microphone anomalies in real time is crucial for improving the performance and reliability of IoT devices.
[0005] This invention aims to address these issues by combining advanced audio signal processing techniques with deep learning algorithms, particularly methods such as LSTM-FCN networks, to conduct in-depth analysis of microphone hardware characteristics and real-time data streams, enabling real-time monitoring of device status and anomaly detection. This approach effectively identifies abnormal microphone conditions and takes appropriate action, improving the stability and security of IoT systems. Summary of the Invention
[0006] This paper addresses the problem of detecting abnormal states of sound sensors (i.e., microphones) in Internet of Things (IoT) devices. It proposes a microphone anomaly detection method based on an LSTM-FCN network. This method analyzes the microphone data stream in depth to detect and determine whether the device is in an abnormal state in real time.
[0007] During the implementation process, the hardware is first initialized. Next, the real-time microphone data stream is collected through a specific hardware device. The collected sound signal undergoes preprocessing, including pre-emphasis, framing, and windowing, to enhance the high-frequency portion of the signal and reduce spectral leakage. The audio signal's spectrum is then obtained using a Fast Fourier Transform (FFT) and the Mel spectrum is calculated. This Mel spectrum is then input into an LSTM-FCN network for training, resulting in a model for anomaly detection. The minimum mean square error (MMSE) denoising algorithm is used to reduce background noise and improve signal quality and intelligibility. During the registration phase, audio signals are collected multiple times in different environments and locations, and a classifier is used to train the model to eliminate the influence of acoustic echoes. Finally, the trained LSTM-FCN network model is used to detect anomalies in the sound signal. The model classifies the input sound signal and determines whether it is an anomaly, thereby identifying various unknown anomalies.
[0008] After confirming that there is a suspected abnormal microphone, the system will promptly feedback the abnormal situation to the user until the user handles the abnormality, thereby confirming its normal working status.
[0009] The invention not only improves the accuracy and real-time performance of microphone anomaly detection, but also enhances the ability to respond to abnormal conditions, which is of great significance for ensuring the stable operation of IoT devices and data security.
[0010] The idea behind the anomaly detection method of the present invention is as follows:
[0011] As a crucial component of IoT devices, the proper functioning of microphones is crucial to the overall system performance. Microphone hardware characteristics, such as frequency response, signal-to-noise ratio, and sensitivity, are key indicators of their operational status. These characteristics can change during microphone operation due to factors such as device aging, environmental changes, and external interference, potentially leading to data anomalies.
[0012] The present invention collects microphone data streams in real time and analyzes these data using algorithms such as LSTM-FCN networks to identify abnormal patterns that are inconsistent with normal working patterns.
[0013] Once an abnormal point is detected, the present invention will further determine the cause of the abnormality and promptly feedback relevant information to the user.
[0014] The core advantage of the present invention is that it can monitor the microphone status in real time and adaptively, and improve the accuracy and efficiency of anomaly detection through strategies such as frequency domain analysis and interference detection.
[0015] The technical solutions adopted by the present invention to achieve the above-mentioned objectives are:
[0016] Step 1. After the system is started, the data acquisition module connected to the microphone is used to obtain the microphone data stream in real time;
[0017] Step 2. Preprocess the collected data stream, including pre-emphasis, framing, windowing, etc., to enhance the high-frequency part of the signal and reduce spectral leakage;
[0018] Step 3. Perform a fast Fourier transform (FFT) on each frame of the preprocessed signal to convert the time domain signal into a frequency domain signal, obtain the spectrum of each frame, and calculate the Mel spectrum. The frequency response Hm(k) of the mth Mel filter is in the form of a triangle wave: Among them, f m is the center frequency of the mth filter (Mel scale conversion formula: f Mel =2595log 10 (1+f Hz / 700).
[0019] Step 4. Input the Mel spectrum into the LSTM-FCN network for training to obtain a model for anomaly detection. The LSTM-FCN network structure is as follows: LSTM layer: The LSTM layer consists of multiple LSTM units, each of which contains an input gate, a forget gate, and an output gate. The calculation formula for the LSTM unit is as follows: i t =σ(W xi xt +W hi h t-1 +b i ) f t =σ(W xf x t +W hi h t-1 +b f ) o t =σ(W xo x t +W hi h t-1 +b o ) c t =f t c t-1 +i t tanh(W xc x t +W hc h t-1 +b c ) h t =o t tanh(c t ) Among them, x t is the input, h t is the output, c t is the cell state, i t 、f t and o t are the activation values of the input gate, forget gate, and output gate, respectively. W and b are weight and bias parameters, and σ is the sigmoid activation function. FCN layer: The FCN layer consists of multiple convolutional layers and pooling layers to extract local features. The calculation formula of the convolutional layer is as follows: Among them, x (l-1) is the input feature map, y (l) is the output feature map, ω (l) and b (l) is the convolution kernel weight and bias parameter, σ is the ReLU activation function, M l-1 and N l-1 are the width and height of the input feature map, M l and N l are the width and height of the convolution kernel respectively; The output layer of the LSTM-FCN network uses the softmax function for classification, and the calculation formula is as follows: Among them, z k is the output of the kth neuron in the output layer, K is the total number of categories;
[0020] Step 5. Use the minimum mean square error (MMSE) noise reduction algorithm to reduce the noise of the sound signal, remove the environmental background noise, and improve the quality and intelligibility of the signal. The MMSE filter coefficient is calculated as: in, is the power spectrum of the speech signal, is the noise power spectrum (estimated by the silent segment);
[0021] Step 6. During the registration phase, multiple audio signals are collected in different environments and locations, and a classifier is used to train the model to eliminate the influence of acoustic echoes.
[0022] Step 7. Use the trained LSTM-FCN network model to detect anomalies in the sound signal. The model classifies the input sound signal to determine whether it is an abnormal signal, thereby identifying various unknown abnormal situations.
[0023] Step 8. Generate an anomaly report based on the anomaly detection results, indicating the time, location and possible cause of the anomaly, and providing decision support, such as adjusting the monitoring strategy, repairing or replacing the microphone, etc.
[0024] In the above technical solution, we further adopt advanced data mining technology and sound signal processing algorithms to improve the accuracy and adaptability of anomaly detection.
[0025] Furthermore, the system can be configured to promptly provide abnormal information to the user when an abnormality is detected, and guide the user to perform precise positioning and inspection.
[0026] The beneficial effects of the present invention are:
[0027] Real-time monitoring: By collecting microphone data in real time, the present invention can promptly and accurately identify abnormal microphone conditions. In addition, the system can adaptively adjust thresholds to adapt to data changes and the characteristics of different microphones, thereby improving the accuracy and efficiency of anomaly detection.
[0028] Versatility and Scalability: This invention is highly versatile and applicable to a wide range of microphone types because it does not rely on the detailed electromagnetic radiation characteristics of a specific microphone. As IoT technology continues to evolve, new microphone types and application scenarios continue to emerge. This invention can adapt to these changes by updating its algorithms and parameters, maintaining its leading position in the field of microphone anomaly detection.
[0029] User-friendly operation experience: The present invention provides an intuitive user interface, allowing users to observe the microphone status through simple operations. When the system detects an anomaly, it will issue a clear alarm prompt to guide the user to take necessary measures, greatly simplifying the user's operation process.
[0030] System Integrability: The anomaly detection system can be easily integrated into existing IoT devices and monitoring platforms, providing an additional layer of security. By seamlessly integrating with existing systems, the system enhances the reliability and stability of the entire IoT ecosystem. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is the overall architecture diagram of the microphone anomaly detection system;
[0032] Figure 2 It is the core flow chart of audio signal processing;
[0033] Figure 3 This is a schematic diagram of the internal structure of the LSTM-FCN network;
[0034] Figure 4 It is the ROC curve display diagram of model performance verification;
[0035] Figure 5 This is a diagram of the user interface when the system detects an anomaly; DETAILED DESCRIPTION
[0036] In order to describe the present invention in more detail, it is described below with reference to the accompanying drawings.
[0037] Reference Figure 1 The microphone anomaly detection system of the present invention primarily consists of the following modules: an audio signal acquisition module, an audio signal processing module, a device fingerprint acquisition module, a device anomaly identification module, and a user interface and report generation module. These modules work together to achieve real-time monitoring of microphones and effective diagnosis of anomalies.
[0038] Sound signal collection stage:
[0039] Step 1: The sound signal acquisition module is started and running. The user turns on the data acquisition device, and the module obtains the microphone data stream in real time. The working principle of the microphone is as follows Figure 2 As shown in Figure 1, it captures sound information in the environment by converting sound waves into electrical signals. During the acquisition process, the system records the timestamp of the sound signal for subsequent analysis and anomaly detection.
[0040] Sound signal processing stage:
[0041] Step 2: Preprocess the received original sound signal. First, pre-emphasize the signal to enhance the high-frequency portion of the signal by boosting the high-frequency components, thereby improving the overall signal quality. Next, the signal undergoes framing and windowing, dividing the continuous signal into multiple small segments (frames). A windowing function is applied to each frame to reduce spectral leakage. The audio signal's spectrum is then obtained through a fast Fourier transform (FFT), and the Mel spectrum is further calculated. The Mel spectrum uses the Mel scale for frequency conversion, which better aligns with the human ear's auditory characteristics and prepares for subsequent feature extraction.
[0042] Step 3: Input the Mel spectrum into the LSTM-FCN network for training to obtain a model for anomaly detection. The LSTM-FCN network structure is as follows: Figure 3 As shown in the figure, the time series features are extracted through the LSTM layer, the local features are extracted through the FCN layer, and finally the output layer is used for classification to determine whether the signal is an abnormal signal.
[0043] Device fingerprint acquisition stage:
[0044] Step 4: Use the Minimum Mean Square Error (MMSE) denoising algorithm to remove background noise and improve signal quality and intelligibility. The MMSE denoising algorithm effectively removes background noise from the original audio signal by modeling and estimating it.
[0045] Step 5: During the registration phase, multiple audio signals are collected from different environments and locations, and a classifier is used to train the model to eliminate the influence of acoustic echoes. Through the multiple samples collected, the classifier can learn the sound characteristics of different environmental conditions, effectively eliminating echo interference in practical applications.
[0046] Equipment anomaly identification stage:
[0047] Step 6: Use the trained LSTM-FCN network model to detect anomalies in the sound signal. The model classifies the input sound signal and determines whether it is an abnormal signal, thereby identifying various unknown abnormal situations.
[0048] Step 7: Generate an anomaly report based on the anomaly detection results, indicating the time, location, and possible causes of the anomaly. This report also provides decision support, such as adjusting monitoring strategies or repairing or replacing microphones. The anomaly report should include the nature of the anomaly, the time and location of the anomaly, and its potential impact, providing a basis for users to take appropriate maintenance or adjustment measures.
[0049] User interface and report generation phase:
[0050] Step 8: Display the processed audio signal and the feature extraction results of the LSTM-FCN network, provide a user interaction interface, generate and display anomaly reports, and provide decision support based on anomalies. Figure 5 This is a feedback interface for when the microphone is disturbed. Users can use it to intuitively understand the abnormal state of the microphone and take appropriate measures based on the system's suggestions. The user interface is simple and clear, making it easy for users to quickly understand and operate.
[0051] The microphone anomaly detection method of the present invention offers advantages such as real-time monitoring, adaptive diagnosis, and user-friendly operation. By collecting and analyzing microphone data in real time, the system can promptly and accurately identify abnormal conditions. Furthermore, the constructed anomaly detection algorithm is adaptable to different types of microphones and has high versatility. Users can perform anomaly detection with simple operations, and the system will also provide clear alarm prompts to help users deal with abnormal situations. This efficient and versatile method greatly improves the monitoring capabilities and anomaly response efficiency of IoT devices, providing a new solution for ensuring data security and stable device operation.
[0052] The above is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, use the technical content disclosed above to make many possible changes and modifications to the technical solution of the present invention, or modify it into an equivalent embodiment with equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention should fall within the scope of protection of the technical solution of the present invention.
Claims
1. Patent Name: A microphone anomaly detection method based on LSTM-FCN network, characterized by: The following steps are involved: Step 1: The sound signal is collected through a microphone and preprocessed, including pre-emphasis, framing, and windowing, to enhance the high-frequency part of the signal and reduce spectrum leakage. Step 2: Perform fast Fourier transform (FFT) on the preprocessed signal to convert the time domain signal into a frequency domain signal, obtain the spectrum of each frame, and calculate the Mel spectrum; Step 3: Input the Mel spectrum into the LSTM-FCN network for training to obtain a model for anomaly detection. The LSTM-FCN network combines the characteristics of the long short-term memory network (LSTM) and the fully convolutional network (FCN). It can effectively extract the temporal and spatial features of the sound signal, thereby more accurately identifying abnormal microphone conditions. Step 4: Use the minimum mean square error (MMSE) noise reduction algorithm to reduce the noise of the sound signal, remove the background noise, and improve the quality and intelligibility of the signal; Step 5: During the registration phase, audio signals are collected multiple times in different environments and locations, and a classifier is used to train the model to eliminate the influence of acoustic echoes. Step 6: Use the trained LSTM-FCN network model to detect anomalies in the sound signal. The model classifies the input sound signal to determine whether it is an abnormal signal, thereby identifying various unknown abnormal situations. Step 7: Generate an anomaly report based on the anomaly detection results, indicating the time, location and possible cause of the anomaly, and providing decision support, such as adjusting the monitoring strategy, repairing or replacing the microphone, etc.
2. The microphone anomaly detection method based on the LSTM-FCN network according to claim 1 is characterized in that In step 1, the purpose of pre-emphasis is to enhance the high-frequency portion, flatten the signal spectrum, eliminate the effects of the vocal cords and lips during the occurrence process, compensate for the high-frequency portion of the speech signal suppressed by the pronunciation system, and highlight the high-frequency resonance peak; framing is to group a certain number of sampling points into an observation unit, called a frame, with an overlapping area between two adjacent frames, usually half the frame length; windowing is to add a Hamming window to make the framed signal continuous and make each frame exhibit the characteristics of a periodic function.
3. The microphone anomaly detection method based on LSTM-FCN network according to claim 1 is characterized in that In step 2, the fast Fourier transform converts the time domain signal into a frequency domain signal to obtain the spectrum of each frame, and the spectrum of the speech signal is squared modulo the energy distribution of the sound signal on the spectrum, that is, the power spectrum of the sound signal; the Mel spectrum is calculated by passing the energy spectrum obtained by the Fourier transform through a set of Mel-scale triangular filters to obtain the Mel spectrum.
4. The microphone anomaly detection method based on the LSTM-FCN network according to claim 1 is characterized in that In step 3, the LSTM-FCN network can automatically learn the feature representation of the sound signal without the need for manual feature extraction, thereby improving the efficiency and accuracy of anomaly detection.
5. The microphone anomaly detection method based on LSTM-FCN network according to claim 1 is characterized in that, In step 4, the minimum mean square error noise reduction algorithm models and estimates the environmental background noise, thereby effectively removing it from the original audio signal, thereby improving the quality and clarity of the speech signal.
6. The microphone anomaly detection method based on LSTM-FCN network according to claim 1 is characterized in that: In step 5, audio signals in different environments and positions are collected multiple times during the registration phase, and a classifier is used for model training, which can effectively offset the echo effect and make the system more robust and reliable.
7. The microphone anomaly detection method based on LSTM-FCN network according to claim 1 is characterized in that: In step 6, the LSTM-FCN network model can efficiently detect anomalies in sound signals, has high accuracy and recall rate, and can detect abnormalities in the microphone in a timely manner.
8. The microphone anomaly detection method based on LSTM-FCN network according to claim 1 is characterized in that: In step 7, the abnormality report should include the time, location, possible cause and impact of the abnormality, and provide decision support information, such as recommended maintenance measures, adjustment of monitoring strategies or suggestions for replacing microphones.
9. An abnormal state detection system for IoT microphones, characterized in that: include: A sound signal acquisition module, used to collect sound signals from a microphone; The sound signal processing module is used to pre-process the collected sound signals, perform fast Fourier transform, and calculate Mel spectrum; LSTM-FCN network training module, used to input the Mel spectrum into the LSTM-FCN network for training to obtain a model for anomaly detection; A noise reduction module is used to perform noise reduction processing on the sound signal using a minimum mean square error (MMSE) noise reduction algorithm; The echo cancellation module is used to collect audio signals in different environments and locations multiple times during the registration phase and use a classifier to train a model to eliminate the impact of acoustic echoes. Anomaly detection module, used to detect anomalies in sound signals using the trained LSTM-FCN network model; The status assessment and decision support module is used to generate anomaly reports based on anomaly detection results and provide corresponding decision support and suggestions; The user interface and report generation module is used to display the processed audio signal and anomaly detection results, provide a user interaction interface, generate and display anomaly reports, and provide decision support based on anomaly situations; The present invention realizes real-time monitoring of microphones and effective diagnosis of abnormal states through the collaborative work of the above-mentioned method and system, improves the accuracy and real-time performance of abnormality detection, and provides strong guarantees for the stable operation of the microphones and data security.
Citation Information
Cited By
Intelligent predictive maintenance system for audio equipment fault
CN120996781A
An intelligent predictive maintenance system for audio equipment failure
CN120996781B