Audio device fault detection method and apparatus, electronic device, and readable storage medium
By acquiring reference and test signals from audio devices, performing frame-by-frame processing and wavelet transform denoising, and generating two-dimensional time-frequency characteristic curves, the problems of high cost and low efficiency in traditional detection are solved, enabling efficient audio device fault detection without the need for specialized equipment and engineers.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-04-14
AI Technical Summary
Traditional audio equipment fault detection requires specialized equipment and engineers, which is costly and inefficient, and Fourier transform has insufficient resolution at low and high frequencies.
By acquiring reference and test signals from audio devices, frame segmentation and wavelet transform denoising are performed to generate two-dimensional time-frequency characteristic curves, and device faults are determined based on these curves.
It can detect faults without the need for specialized equipment and engineers, reduce costs, improve efficiency, and process multiple audio signals simultaneously.
Smart Images

Figure CN121331165B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of detection technology, and in particular to a method, apparatus, electronic device, and readable storage medium for detecting audio equipment faults. Background Technology
[0002] Currently, we widely use various audio devices in our lives, such as speakers, amplifiers, audio processors, and mixing consoles. These devices are used in various settings, including conference rooms, cinemas, shopping malls, schools, hospitals, squares, exhibition halls, and parks. During use, these audio devices are prone to malfunctions, such as silence, sound distortion, excessive or insufficient loudness, noise, high-frequency damage, and broken cones. In important locations, such as conference rooms and exhibition halls, maintenance personnel need to proactively identify faulty audio equipment for timely replacement or repair to avoid disruptions and serious losses.
[0003] Traditional audio equipment fault diagnosis requires purchasing specialized testing equipment. On one hand, this equipment is expensive, leading to high testing costs. On the other hand, it necessitates assigning professional audio engineers to carry the equipment to each audio device on-site for testing, which is time-consuming, labor-intensive, and inefficient. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a method, apparatus, electronic device, and readable storage medium for audio equipment fault detection. This eliminates the need to purchase specialized testing equipment; fault detection can be performed using existing equipment, thus reducing testing costs. Furthermore, by automating the analysis of audio signals, minimal human intervention is required, and multiple audio signals can be processed simultaneously, thereby improving testing efficiency.
[0005] In a first aspect, embodiments of this application provide an audio device fault detection method, including:
[0006] Acquire a reference audio signal and a test audio signal of the audio device to be tested; the reference audio signal is the audio signal played by the audio device during the playback of a reference audio signal when the audio device is not malfunctioning; the test audio signal is the audio signal played by the audio device during the playback of the reference audio signal during the fault detection phase of the audio device.
[0007] For each audio signal in the reference audio signal and the test audio signal, the audio signal is divided into frames to obtain multiple audio signal frames of the audio signal;
[0008] For each audio signal frame of the audio signal, wavelet transform is used to remove noise from the audio signal frame, resulting in a denoised audio signal frame.
[0009] Calculate the amplitude of each denoised audio signal frame at its center frequency, and generate a two-dimensional time-frequency characteristic curve corresponding to the audio signal based on the amplitude of each denoised audio signal frame at its center frequency.
[0010] Based on the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal, it is determined whether the audio device has malfunctioned.
[0011] In conjunction with the first aspect, this application provides a first possible implementation of the first aspect, wherein the reference audio signal is a sweep frequency signal with exponentially increased phase continuity; the starting frequency of the reference audio signal is low frequency, and the ending frequency is high frequency.
[0012] In conjunction with the first aspect, this application provides a second possible implementation of the first aspect, wherein the audio signal is a discrete audio signal, and the signal sequence of the discrete audio signal includes the amplitude of the audio signal at each sampling point; the step of performing frame segmentation processing on each audio signal in the reference audio signal and the test audio signal to obtain multiple audio signal frames of the audio signal includes:
[0013] For each audio signal in the reference audio signal and the test audio signal, the audio signal is divided into frames according to a preset frame length and frame shift to obtain multiple audio signal frames of the audio signal; wherein, the frame length is used to control the number of sampling points in each audio signal frame; the frame shift is used to control the number of sampling points shifted backward relative to the starting point of the previous frame in two adjacent audio signal frames; the frame length is greater than the frame shift and less than the total duration of the reference audio signal.
[0014] In conjunction with the first possible implementation of the first aspect, this application provides a third possible implementation of the first aspect, wherein the audio signal is a discrete audio signal, and the signal sequence of the discrete audio signal includes the amplitude of the audio signal at each sampling point; the step of removing noise from each audio signal frame of the audio signal using wavelet transform to obtain a denoised audio signal frame includes:
[0015] For each audio signal frame, a discrete wavelet transform is performed on the audio signal frame based on the number of wavelet decomposition levels set in the discrete wavelet transform, to obtain the detail coefficients of the audio signal frame at each frequency level; wherein, the detail coefficients represent the energy intensity of the audio signal at that frequency level; each frequency level corresponds to a frequency band, and the frequency band corresponding to each frequency level of the audio signal frame is located within the frequency band of the audio signal frame; in the same audio signal frame, the higher the number of frequency levels, the closer the frequency band of that frequency level is to the low-frequency end of the frequency band of the audio signal frame;
[0016] Based on the center frequency of the audio signal frame and the center frequencies of each frequency layer corresponding to the audio signal frame, the frequency layer closest to the center frequency of the audio signal frame is determined from each frequency layer corresponding to the audio signal frame and is used as the target frequency layer of the audio signal frame.
[0017] The detail coefficients of the target frequency layer of the audio signal frame are retained, and the detail coefficients of other frequency layers of the audio signal frame are set to zero to obtain the denoised detail coefficient sequence corresponding to the audio signal frame.
[0018] Wavelet reconstruction is performed on the denoised detail coefficient sequence corresponding to the audio signal frame to obtain the reconstructed audio signal frame, which is then used as the denoised audio signal frame.
[0019] In conjunction with the first aspect, this application provides a fourth possible implementation of the first aspect, wherein calculating the amplitude of each denoised audio signal frame at its center frequency includes:
[0020] For each denoised audio signal frame, calculate the continuous wavelet transform coefficients of the denoised audio signal frame, and determine the amplitude of the denoised audio signal frame at its center frequency based on the continuous wavelet transform coefficients.
[0021] In conjunction with the fourth possible implementation of the first aspect, this application provides a fifth possible implementation of the first aspect, wherein calculating the continuous wavelet transform coefficients of the denoised audio signal frame and determining the amplitude of the denoised audio signal frame at its center frequency based on the continuous wavelet transform coefficients includes:
[0022] Calculate the two-dimensional coefficient matrix of the denoised audio signal frame using the following continuous wavelet transform function:
[0023]
[0024] Among them, W k(a,b) represents the two-dimensional coefficient matrix of the k-th audio signal frame after denoising; a is the scaling factor, representing the vertical coordinate of the continuous wavelet transform coefficients contained in the two-dimensional coefficient matrix; b is the translation factor, representing the horizontal coordinate of the continuous wavelet transform coefficients contained in the two-dimensional coefficient matrix. Indicates the mother wavelet; Indicates complex conjugation; This represents the k-th frame of the denoised audio signal; n represents the n-th sampling point of the k-th frame of the denoised audio signal; and N is the total number of sampling points in the k-th frame of the denoised audio signal. The sampling period;
[0025] The scale factor corresponding to the center frequency of the denoised k-th frame of the audio signal is calculated using the following formula. :
[0026]
[0027] in, This represents the center frequency of the continuous wavelet transform function; This represents the center frequency of the k-th audio signal frame;
[0028] From the two-dimensional coefficient matrix corresponding to the k-th frame of the denoised audio signal, determine the ordinate as the scale factor. The absolute value of the largest coefficient among the multiple continuous wavelet transform coefficients is taken as the amplitude of the k-th frame of the denoised audio signal at its center frequency.
[0029] In conjunction with the first aspect, this application provides a sixth possible implementation of the first aspect, wherein the horizontal axis of the two-dimensional time-frequency characteristic curve corresponding to the audio signal is the center point time of the denoised audio signal frame, and the vertical axis is the amplitude of each denoised audio signal frame at its center frequency; there is a one-to-one correspondence between the center point time of each denoised audio signal frame and the number of frames; the step of determining whether the audio device has malfunctioned based on the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal includes:
[0030] For each frame number of the denoised audio signal, the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the test audio signal is calculated, and the ratio between the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the reference audio signal is obtained. The amplitude ratio corresponding to that frame number of audio signal frames is then logarithmically transformed to obtain the logarithmic amplitude difference corresponding to that frame number of audio signal frames; the logarithmic amplitude difference is used to represent decibels.
[0031] A frequency response difference curve is constructed based on the logarithmic amplitude difference corresponding to each frame number; wherein, the horizontal axis of the frequency response difference curve is the center point time of the denoised audio signal frame, and the vertical axis is decibels;
[0032] The degree of deviation of the frequency response difference curve from the horizontal axis is used to determine whether the audio device has malfunctioned.
[0033] In conjunction with the sixth possible implementation of the first aspect, this application provides a seventh possible implementation of the first aspect, wherein, for each frame number of the denoised audio signal frame, the ratio between the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve corresponding to the test audio signal and the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal is calculated to obtain the amplitude ratio corresponding to that frame number of audio signal frames, and the logarithmic transformation is performed on the amplitude ratio to obtain the logarithmic amplitude difference corresponding to that frame number of audio signal frames, includes:
[0034] The logarithmic magnitude difference corresponding to the audio signal frames of this frame number is calculated using the following formula:
[0035]
[0036] in, This represents the amplitude of the k-th frame of the audio signal after denoising in the two-dimensional time-frequency characteristic curve corresponding to the test audio signal; This represents the amplitude of the k-th frame of the audio signal after denoising in the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal; This represents the logarithmic magnitude difference corresponding to the k-th frame of the audio signal after denoising.
[0037] In conjunction with the sixth possible implementation of the first aspect, this application provides an eighth possible implementation of the first aspect, wherein determining whether the audio device has malfunctioned based on the degree of deviation of the frequency response difference curve from the horizontal axis includes:
[0038] Using the constructed smooth convolution kernel, the frequency response difference curve is smoothed and convolved to obtain the smoothed time-decibel curve; the horizontal axis of the time-decibel curve is the center point time of the denoised audio signal frame, and the vertical axis is decibels;
[0039] Based on the number of frames of each audio signal frame, the center frequency of each audio signal frame is determined. Based on the center frequency of each audio signal frame, the time-decibel curve is converted into a frequency-decibel curve with the center frequency on the horizontal axis and decibels on the vertical axis.
[0040] The degree of deviation of the frequency-decibel curve from the horizontal axis is used to determine whether the audio device is malfunctioning; wherein, the degree of deviation is positively correlated with the probability of malfunction.
[0041] Secondly, embodiments of this application also provide an audio device fault detection apparatus, comprising:
[0042] The acquisition module is used to acquire a reference audio signal and a test audio signal of the audio device to be tested; the reference audio signal is the audio signal played by the audio device during the playback of the reference audio signal when the audio device is not malfunctioning; the test audio signal is the audio signal played by the audio device during the playback of the reference audio signal during the fault detection phase of the audio device.
[0043] The framing module is used to perform framing processing on each audio signal in the reference audio signal and the test audio signal to obtain multiple audio signal frames of the audio signal.
[0044] The noise reduction module is used to remove noise from each audio signal frame of the audio signal using wavelet transform, so as to obtain a noise-reduced audio signal frame.
[0045] The calculation module is used to calculate the amplitude of each denoised audio signal frame at its center frequency, and generate a two-dimensional time-frequency characteristic curve corresponding to the audio signal based on the amplitude of each denoised audio signal frame at its center frequency.
[0046] The judgment module is used to determine whether the audio device has malfunctioned based on the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal.
[0047] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps in any of the possible implementations of the first aspect described above are performed.
[0048] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps in any of the possible implementations of the first aspect described above.
[0049] This application provides an audio device fault detection method, apparatus, electronic device, and readable storage medium. The method acquires a reference audio signal and a test audio signal of the audio device to be detected. The reference audio signal is an audio signal played by the audio device during the playback of a reference audio signal when the audio device is not malfunctioning. The test audio signal is an audio signal played by the audio device during the playback of the reference audio signal during the audio device fault detection phase. The reference audio signal is segmented into frames to obtain multiple audio signal frames. Wavelet transform is used to remove noise from each audio signal frame, resulting in denoised audio signal frames corresponding to the reference audio signal. The amplitude of each denoised audio signal frame at its center frequency is calculated, and a two-dimensional time-frequency characteristic curve corresponding to the reference audio signal is generated based on the amplitude of each denoised audio signal frame at its center frequency. Furthermore, the test audio signal is segmented into frames to obtain multiple audio signal frames corresponding to the test audio signal. Wavelet transform is then used to remove noise from each audio signal frame, resulting in denoised audio signal frames. The amplitude of each denoised audio signal frame at its center frequency is calculated, and a two-dimensional time-frequency characteristic curve is generated based on this amplitude. Finally, based on the two-dimensional time-frequency characteristic curves of both the reference audio signal and the test audio signal, it is determined whether the audio device has malfunctioned.
[0050] In this embodiment, the above-described detection method can run on any existing device (such as a server), eliminating the need to purchase specialized testing equipment. Fault detection can be performed using existing equipment, which helps reduce testing costs. Furthermore, by automating the analysis of audio signals, minimal human intervention is required, and multiple audio signals can be processed simultaneously (i.e., multiple audio signals can be analyzed concurrently), which improves detection efficiency and automation.
[0051] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0052] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1A flowchart of an audio device fault detection method provided in an embodiment of this application is shown;
[0054] Figure 2 The illustration shows a sample diagram of a reference audio signal and a test audio signal provided in an embodiment of this application;
[0055] Figure 3 This illustration shows a schematic diagram of performing discrete wavelet transform on the k-th audio signal frame based on the number of wavelet decomposition layers, according to an embodiment of this application.
[0056] Figure 4 This paper shows an example diagram of a two-dimensional time-frequency characteristic curve of an audio signal provided in an embodiment of this application;
[0057] Figure 5 An example diagram of a frequency-decibels curve provided in an embodiment of this application is shown;
[0058] Figure 6 This illustration shows a structural schematic diagram of an audio equipment fault detection device provided in an embodiment of this application;
[0059] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0060] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0061] Traditional audio equipment fault diagnosis requires purchasing specialized testing equipment. On one hand, this equipment is expensive, leading to high testing costs. On the other hand, it necessitates assigning professional audio engineers to carry the equipment to each audio device on-site for testing, which is time-consuming, labor-intensive, and inefficient.
[0062] Furthermore, traditional detection methods can also be based on Fourier transform, but there are some shortcomings, mainly in the insufficient frequency resolution at low frequencies and the insufficient time resolution at high frequencies.
[0063] Based on this, embodiments of this application provide an audio device fault detection method, apparatus, electronic device, and readable storage medium, which are described below through embodiments.
[0064] To facilitate understanding of this embodiment, a detailed description of an audio device fault detection method disclosed in this application embodiment will be provided first. This audio device fault detection method can be applied to existing detection equipment (such as servers, mobile devices with data processing capabilities, or handheld devices, etc.), such as... Figure 1 As shown, the audio device fault detection method includes the following steps S101-S105:
[0065] S101: Acquire the reference audio signal and the test audio signal of the audio device to be tested; the reference audio signal is the audio signal played by the audio device during the playback of the reference audio signal when the audio device is not malfunctioning; the test audio signal is the audio signal played by the audio device during the playback of the reference audio signal during the fault detection phase of the audio device.
[0066] S102: For each audio signal in the reference audio signal and the test audio signal, perform frame segmentation processing on the audio signal to obtain multiple audio signal frames of the audio signal.
[0067] S103: For each audio signal frame of the audio signal, use wavelet transform to remove noise from the audio signal frame to obtain a denoised audio signal frame.
[0068] S104: Calculate the amplitude of each denoised audio signal frame at its center frequency, and generate a two-dimensional time-frequency characteristic curve corresponding to the audio signal based on the amplitude of each denoised audio signal frame at its center frequency.
[0069] S105: Based on the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal, determine whether the audio device has malfunctioned.
[0070] In step S101, the audio device refers to an audio playback device, which is physically composed of one or more devices and has the ability to amplify audio signals and play sound. Audio devices include, but are not limited to: power amplifiers, mixers, speakers, audio cables, etc.
[0071] An audio device malfunction refers to a problem or abnormality in a physical component of the audio device, resulting in abnormal sound playback, such as noise, distortion, lack of sound at a certain frequency, or excessively loud or soft volume. Conversely, an audio device functioning correctly means that all physical components within the audio device are functioning normally and without malfunction.
[0072] The audio device to be tested refers to the audio device that needs to be fault-tested. Generally, each audio device can be tested periodically; the audio device currently requiring testing is the one to be tested.
[0073] This embodiment is divided into two stages: the first stage is the preparation stage (the stage where the equipment has not malfunctioned), and the second stage is the fault detection stage. The preparation stage can be the stage when the audio equipment has just been purchased or is about to be put into use. The audio equipment in this stage is generally new and unused. New audio equipment is mostly fault-free (i.e., normal). To further ensure that the audio equipment in the preparation stage is fault-free, it can be further tested manually or through other testing methods to ensure that the audio equipment in the preparation stage is fault-free. The fault detection stage generally refers to the stage where the audio equipment is periodically tested for faults after it has been in use for a period of time.
[0074] In one possible implementation, when performing step S101 to acquire the reference audio signal and test audio signal of the audio device to be detected, the following steps S1001-S1002 can be specifically performed:
[0075] S1001: For the audio device to be tested, during the preparation phase (i.e., the phase in which no fault has occurred), the audio device that has not experienced a fault is controlled to play a reference audio signal. During the process of the audio device that has not experienced a fault playing the reference audio signal during the preparation phase, the audio signal played by the audio device is collected and used as a reference audio signal.
[0076] S1002: During the fault detection phase, the audio device is controlled to play a reference audio signal, so that during the process of the audio device playing the reference audio signal during the fault detection phase, the audio signal played by the audio device is collected and used as a test audio signal.
[0077] In this embodiment, the detection device can be a mobile device or handheld device with data processing capabilities. The detection device is typically positioned near the audio device to be detected to collect the audio signal played by that device.
[0078] In another possible implementation, when performing step S101 to acquire the reference audio signal and test audio signal of the audio device to be detected, the following steps may also be performed:
[0079] The reference audio signal and test audio signal of the audio device to be tested can be obtained from the database, or the reference audio signal and test audio signal of the audio device to be tested can be obtained from the audio acquisition device.
[0080] In this embodiment, a reference audio signal and a test audio signal of the audio device to be tested can be acquired using an audio acquisition device. Specifically, the audio acquisition device can acquire the reference audio signal and the test audio signal of the audio device to be tested through steps S1001-S1002. After acquiring the reference audio signal and the test audio signal, the audio acquisition device can transmit the acquired reference audio signal and the test audio signal to the testing device; alternatively, it can store the acquired reference audio signal and the test audio signal in a database so that the testing device can retrieve them from the database.
[0081] In this embodiment, since the reference and test audio signals of the audio device are acquired by the audio acquisition device, and the detection device is only used for data analysis, the detection device does not need to be placed near the audio device to be detected. However, the audio acquisition device needs to be placed near the audio device to be detected during the acquisition of the reference and test audio signals.
[0082] like Figure 2 As shown, sample diagrams of the reference audio signal and the test audio signal are displayed, where the red curve represents the reference audio signal and the green curve represents the test audio signal.
[0083] The reference audio signal is a standard audio signal. In order to better test the audio playback performance of audio devices at various frequencies (or frequency bands), or to detect whether audio devices have malfunctioned at various frequencies (or frequency bands), in one possible implementation, the reference audio signal is a frequency sweep signal with exponentially increasing phase; the starting frequency of the reference audio signal is low frequency, and the ending frequency is high frequency.
[0084] A swept frequency signal, also called a linear frequency modulated signal, refers to a signal whose frequency is not fixed, but rather changes smoothly from a starting frequency to a ending frequency over a period of time. In this embodiment, the starting frequency is low frequency, and the ending frequency is high frequency. The division between high and low frequencies can be done in a conventional way. Generally, low frequency refers to around 20Hz-250Hz; mid frequency refers to 250Hz-4000Hz, which is the most sensitive frequency range for the human ear, containing the core timbre of human voices and most musical instruments; high frequency refers to 4000Hz-20,000Hz. Exponential boost describes the pattern of frequency change. It specifies how the instantaneous frequency changes over time. Phase continuity means that the phase function of the signal is a continuous, smooth function without sudden jumps or angles.
[0085] Using this reference audio signal, the response characteristics at various frequencies from low to high can be tested, thus providing a more comprehensive understanding of the frequency response characteristics of audio devices.
[0086] In this embodiment, the formula for the reference audio signal is as follows:
[0087]
[0088] Where x(t) is the amplitude of the swept frequency signal (i.e., the reference audio signal) at time t; amplitude refers to the intensity or magnitude of the audio signal at any given moment, and for the human ear, its most direct perception is the loudness of the sound. Amplitude and loudness are positively correlated.
[0089] f1 is the starting frequency (low frequency end), f2 is the ending frequency (high frequency end), T is the total sweep duration (i.e., the total duration of the reference audio signal); ln is the natural logarithm; t is the time variable, ranging from [0, T].
[0090] In this embodiment, exponential increase is selected, which can provide better resolution at low frequencies. This aligns with the human ear's sensitivity to low frequencies and insensitivity to high frequencies, allowing detection of even minor changes at low frequencies.
[0091] In this embodiment, whether it is a detection device or an audio acquisition device, when acquiring the reference audio signal and test audio signal of the audio device to be detected, the sampling rate can be set to 44100 Hz or other sampling rates according to the required resolution. After recording, it is saved as an audio file, such as a WAV format file or other format file, for later analysis. Therefore, the acquired reference audio signal and test audio signal are both discrete audio signals, denoted by x[n], where 0 ≤ n ≤ T × f S f S Where is the sampling rate, and T is the total duration of recording the audio signals (reference audio signal and test audio signal).
[0092] In step S102, since the recorded audio signals (including reference audio signals and test audio signals) may be relatively long, and given the high resolution of the signal, using wavelet transform for such long signals could result in a large amount of data processing and require a significant amount of memory. To accelerate processing efficiency, this embodiment first performs frame-by-frame processing on the reference audio signals and test audio signals.
[0093] In one possible implementation, the audio signal (including a reference audio signal and a test audio signal) is a discrete audio signal, and the signal sequence of the discrete audio signal contains the amplitude of the audio signal at each sampling point. When performing step S102, the following steps can be specifically followed:
[0094] For each audio signal in the reference audio signal and the test audio signal, the audio signal is divided into frames according to the preset frame length and frame shift to obtain multiple audio signal frames. The frame length is used to control the number of sampling points in each audio signal frame. The frame shift is used to control the number of sampling points shifted backward from the starting point of the previous frame in two adjacent audio signal frames. The frame length is greater than the frame shift and less than the total duration of the reference audio signal.
[0095] In this embodiment, frame length refers to the number of milliseconds (ms) of audio signal contained in each frame, or the number of sampling points. Frame shift refers to how many milliseconds (ms) or sampling points the subsequent frame is offset from the starting point of the previous frame.
[0096] The frame length and frame shift are set to fixed values. By making the frame length greater than the frame shift, the next frame overlaps with the previous frame after the frame shift, making the signal smoother. For example, the frame length is 30 milliseconds (the length of each segment), and the frame shift is 10 milliseconds (the distance moved forward each time).
[0097] For example, the first frame starts at 0ms and cuts to 30ms. The second frame doesn't start at 30ms, but moves forward 10ms, so it starts at 10ms and cuts to 40ms. The third frame moves forward another 10ms, starting at 20ms and cuts to 50ms.
[0098] In this embodiment, regardless of whether the audio signal is the reference audio signal or the test audio signal, let the sampled audio signal be x[n], and divide it into frames according to the frame length L and the frame shift R. The formula for the k-th frame is as follows:
[0099]
[0100] Where x[n] is the audio signal (reference audio signal or test audio signal); x k [n] represents the amplitude at the nth sampling point in the kth frame of the audio signal; L represents the length of each frame of the audio signal; R represents the frame shift; L represents the total number of frames of an audio signal; K represents the total number of frames of an audio signal. Specifically, the total number of frames of the reference audio signal is the same as the total number of frames of the test audio signal.
[0101]
[0102] Where N is the total number of sampling points for the audio signal.
[0103] In step S103, considering that environmental noise is usually included during the recording of the reference audio signal and the test audio signal, and since the reference audio signal and the test audio signal are recorded at different times, the noise they contain is usually different. Therefore, it is necessary to first perform noise reduction processing on each audio signal frame corresponding to the reference audio signal and the test audio signal.
[0104] Since the reference audio signal in this embodiment is a non-stationary signal, Fourier analysis is not very effective for denoising. Therefore, wavelet transform is used for denoising in this embodiment, which is more effective.
[0105] In this embodiment, the wavelet transform is specifically the discrete wavelet transform (DWT), such as standard multiscale decomposition, holed wavelet, lifting wavelet, WPD, FWT, etc. The wavelet basis can be Haar, dbN, symN, coifN, Morlet, etc. As a preferred choice, dbN is more suitable for the audio signal in this embodiment. The Haar wavelet basis is suitable for signals with abrupt changes. This embodiment uses a swept frequency signal, which is a non-abrupt signal.
[0106] Wavelet transform is a time-scale (time-frequency) analysis method for signals. It features multi-resolution analysis and the ability to characterize local signal features in both the time and frequency domains. It's a time-frequency localization analysis method with a fixed window size but a variable shape, allowing both time and frequency windows to be modified. Specifically, it has lower time resolution and higher frequency resolution in the low-frequency range, and higher time resolution and lower frequency resolution in the high-frequency range. It is well-suited for analyzing non-stationary signals and extracting local signal features; therefore, wavelet transform is often referred to as a microscope for signal analysis and processing.
[0107] This method, based on wavelet transform, has the following advantages:
[0108] High detection sensitivity: Wavelet transform can capture minute waveform changes, making it suitable for detecting slight distortions or noise;
[0109] Strong anti-interference capability: It has strong robustness to environmental noise and nonlinear distortion;
[0110] Low cost: It does not rely on expensive professional equipment; it can be deployed with ordinary microphones and processing systems.
[0111] Automation and remote support: It can be integrated into the operation and maintenance system to automatically complete scheduled tests and generate reports;
[0112] Adaptable to various devices and scenarios: Supports audio equipment, speakers, amplifiers, audio processors, etc., suitable for conference rooms, exhibition halls, cinemas, and other venues.
[0113] In one possible implementation, when performing step S103, the following steps S1031-S1034 can be specifically performed:
[0114] S1031: For each audio signal frame of the audio signal, based on the number of wavelet decomposition layers set in the discrete wavelet transform, perform a discrete wavelet transform on the audio signal frame to obtain the detail coefficients of the wavelet decomposition of the audio signal frame at each frequency layer; wherein, the detail coefficients represent the energy intensity of the audio signal at that frequency layer; each frequency layer corresponds to a frequency band, and the frequency band corresponding to each frequency layer of the audio signal frame is located within the frequency band of the audio signal frame; in the same audio signal frame, the higher the number of frequency layers, the closer the frequency band of that frequency layer is to the low-frequency end of the frequency band of the audio signal frame.
[0115] S1032: Based on the center frequency of the audio signal frame and the center frequencies of each frequency layer corresponding to the audio signal frame, determine the frequency layer that is closest to the center frequency of the audio signal frame from each frequency layer corresponding to the audio signal frame, and use it as the target frequency layer of the audio signal frame.
[0116] S1033: Retain the detail coefficients of the target frequency layer of the audio signal frame, and set the detail coefficients of other frequency layers of the audio signal frame to zero, to obtain the denoised detail coefficient sequence corresponding to the audio signal frame.
[0117] S1034: Perform wavelet reconstruction on the denoised detail coefficient sequence corresponding to the audio signal frame to obtain the reconstructed audio signal frame, and use the reconstructed audio signal frame as the denoised audio signal frame.
[0118] In step S1031, for the k-th audio signal frame, based on the wavelet decomposition level J set in the discrete wavelet transform, a discrete wavelet transform is performed on the k-th audio signal frame to obtain the detail coefficients of the wavelet decomposition of the k-th audio signal frame at each frequency level j. The sequence of detail coefficients for each frequency level wavelet decomposition corresponding to the k-th audio signal frame is as follows:
[0119]
[0120] in, Let be the detail coefficients of the k-th frame of the audio signal in the first frequency layer wavelet decomposition. The detail coefficients of the k-th frame of the audio signal are obtained from the wavelet decomposition at the second frequency level. denoted as the detail coefficients of the k-th frame of the audio signal in the wavelet decomposition at the J-frequency level.
[0121] like Figure 3The diagram illustrates the process of performing a discrete wavelet transform on the k-th audio signal frame based on the wavelet decomposition level J when the wavelet decomposition level J=3, to obtain the detail coefficients of the k-th audio signal frame at each frequency level j. For ease of demonstration, Figure 3 In this context, D1(k) represents the above. Let D2(k) represent the above Let D3(k) represent .
[0122] In this embodiment, each frequency layer (the j-th frequency layer) does not correspond to just a single precise frequency, but rather to a frequency band (frequency range). For example, the third frequency layer might correspond to the 800Hz-1000Hz frequency band. Therefore, D3(k) (i.e. This describes the situation of audio signals in the frequency range of 800Hz-1000Hz.
[0123] Each frequency layer of an audio signal frame corresponds to a frequency band within the frequency band of that audio signal frame. For example, taking the k-th audio signal frame as an example, assuming the frequency band of the k-th audio signal frame is 20Hz-2000Hz, then the frequency band corresponding to the first frequency layer can be 1800Hz-2000Hz; the frequency band corresponding to the second frequency layer can be 1200Hz-1500Hz; and the frequency band corresponding to the third frequency layer can be 800Hz-1000Hz.
[0124] Within the same audio signal frame, the higher the number of frequency layers, the closer the frequency band of that layer is to the low-frequency end of the frequency band of that audio signal frame. Specifically, for example... Figure 3 As shown, within the same audio signal frame, each frequency layer j corresponds to a center frequency and a bandwidth. The smaller the frequency layer j, the higher the corresponding center frequency and the wider the bandwidth. The larger j is, the lower the corresponding center frequency. Therefore, This detail coefficient comprehensively depicts the activity of the audio signal in the j-th frequency band.
[0125] The detail factor of each frequency level represents the energy intensity of the audio signal at that frequency level. The absolute value of directly reflects the degree of fluctuation and intensity of the audio signal within the frequency band corresponding to the k-th audio signal frame. The larger the absolute value of , the more intense the fluctuation of the audio signal, the higher the intensity, and the stronger the energy within the frequency band corresponding to the k-th frame of the audio signal. The smaller the absolute value, the weaker the fluctuation, the lower the intensity, and the weaker the energy of the audio signal within the frequency band corresponding to the k-th frame of the audio signal.
[0126] In step S1032, each audio signal frame corresponds to a center frequency. For example, if the frequency band of the kth audio signal frame is 20Hz-2000Hz, then its center frequency is 1010Hz.
[0127] Each frequency layer of this audio signal frame corresponds to a center frequency. For example, the first frequency layer has a frequency band of 1800Hz-2000Hz, and its corresponding center frequency is 1900Hz. The second frequency layer has a frequency band of 1200Hz-1500Hz, and its corresponding center frequency is 1350Hz; the third frequency layer has a frequency band of 800Hz-1000Hz, and its corresponding center frequency is 900Hz.
[0128] Among them, the center frequency of 900Hz corresponding to the third frequency layer is closest to the center frequency of 1010Hz corresponding to the audio signal frame. Therefore, the third frequency layer is taken as the target frequency layer of the audio signal frame.
[0129] In step S1033, "other frequency layers" refers to all frequency layers corresponding to the audio signal frame, excluding the target frequency layer. In this embodiment, considering that signals far from the center frequency of each audio signal frame are usually noise signals, the noise signals in the audio signal frame are removed by setting the detail coefficients of the other frequency layers of the audio signal frame to zero, thus obtaining the denoised detail coefficient sequence corresponding to the audio signal frame.
[0130] For example, suppose the target frequency layer is the 1st. Therefore, the denoised detail coefficient sequence corresponding to the audio signal frame is as follows:
[0131]
[0132] In step S1034, the denoised detail coefficient sequence corresponding to the audio signal frame is reconstructed using the inverse discrete wavelet transform (IDWT) to obtain the reconstructed audio signal frame:
[0133]
[0134] in, The reconstructed audio signal frame (i.e., the denoised audio signal frame). The inverse discrete wavelet transform (IDWT), corresponding to the previously selected discrete wavelet transform (DWT) algorithm, is an inverse algorithm.
[0135] In this embodiment, each audio signal frame corresponds to a denoised audio signal frame.
[0136] In step S104, the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal is used to reflect the amplitude change over time in the response of the audio device to the reference audio signal when no fault occurs. The two-dimensional time-frequency characteristic curve corresponding to the test audio signal is used to reflect the amplitude change over time in the response of the audio device to the reference audio signal during the fault detection phase.
[0137] In one possible implementation, when performing step S104 to calculate the amplitude of each denoised audio signal frame at its center frequency, the following steps can be specifically performed:
[0138] S1041: For each denoised audio signal frame, calculate the continuous wavelet transform coefficients of the denoised audio signal frame, and determine the amplitude of the denoised audio signal frame at its center frequency based on the continuous wavelet transform coefficients.
[0139] In this embodiment, the two-dimensional coefficient matrix of the denoised audio signal frame is calculated using the following continuous wavelet transform function:
[0140]
[0141] Among them, W k (a,b) represents the two-dimensional coefficient matrix of the k-th audio signal frame after denoising; a is the scaling factor, representing the vertical coordinate of the continuous wavelet transform coefficients (i.e., CWT coefficients) contained in the two-dimensional coefficient matrix; b is the translation factor, representing the horizontal coordinate of the continuous wavelet transform coefficients contained in the two-dimensional coefficient matrix. Indicates the mother wavelet; Indicates complex conjugation; This represents the k-th frame of the denoised audio signal; n represents the n-th sampling point of the k-th frame of the denoised audio signal; and N is the total number of sampling points in the k-th frame of the denoised audio signal. The sampling period.
[0142] The scale factor corresponding to the center frequency of the denoised k-th frame of the audio signal is calculated using the following formula. :
[0143]
[0144] in, This represents the center frequency of the continuous wavelet transform function (e.g., the center frequency of the Morlet wavelet is approximately 0.8125). The center frequency (known) of the k-th frame of the audio signal is represented.
[0145] In this embodiment, the scale factor corresponding to the center frequency of the k-th audio signal frame is calculated. This is a scaling factor in the two-dimensional coefficient matrix of the k-th audio signal frame. The scaling factors correspond to different audio signal frames. They may be different.
[0146] Next, from the two-dimensional coefficient matrix corresponding to the k-th frame of the denoised audio signal, the vertical axis is determined to be the scale factor. The absolute value of the largest coefficient among the multiple continuous wavelet transform coefficients is taken as the amplitude of the k-th frame of the denoised audio signal at its center frequency.
[0147] Specifically, the x-axis corresponding to the largest coefficient among multiple continuous wavelet transform coefficients is used as the scaling factor. Corresponding translation factor The amplitude A of the denoised k-th frame audio signal at its center frequency is extracted using the following formula. k :
[0148]
[0149] in, It is the largest coefficient among the continuous wavelet transform coefficients.
[0150] After calculating the center frequency amplitude of each frame of the audio signal, the characteristics of the entire audio signal are summarized into a two-dimensional time-frequency characteristic curve, and each characteristic point is defined as: , where t k A represents the time of the center point of the k-th frame, which is also equivalent to representing the frame number of the k-th frame. k This represents the amplitude (characteristic intensity) of the center frequency corresponding to the k-th frame.
[0151] The two-dimensional time-frequency characteristic curve corresponding to this audio signal is defined as follows:
[0152]
[0153] This two-dimensional time-frequency characteristic curve reflects the amplitude change over time in response to a reference sweep signal when the audio device is operating. It has the following characteristics: time-domain resolution is controlled by frame length and step size; frequency-domain information is focused near the center frequency; interference noise has been removed; and it can be considered as the "dynamic frequency response characteristic" of the audio device. Figure 4 The image shows a sample two-dimensional time-frequency characteristic curve of an audio signal. The horizontal axis of the two-dimensional time-frequency characteristic curve corresponding to the audio signal is the center point time (or frame number) of the denoised audio signal frame. Specifically, each denoised audio signal frame corresponds to its own frame number and center point time. The center point time increases with the increase of the frame number. The smaller the frame number, the smaller the center point time, and the larger the frame number, the larger the center point time. The vertical axis is the amplitude of each denoised audio signal frame at its center frequency.
[0154] In this embodiment, by performing the above steps S102-S104 on the reference audio signal and the test audio signal respectively, the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal can be obtained.
[0155] In step S105, after obtaining the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal, the audio device is judged to have malfunctioned by analyzing these two two-dimensional time-frequency characteristic curves.
[0156] In one possible implementation, the horizontal axis of the two-dimensional time-frequency characteristic curve corresponding to the audio signal is the center point time (or frame number) of the denoised audio signal frame, and the vertical axis is the amplitude of each denoised audio signal frame at its center frequency; there is a one-to-one correspondence between the center point time and the frame number of each denoised audio signal frame; when executing step S105, it can be specifically executed according to the following steps S1051-S1053:
[0157] S1051: For each frame number of the denoised audio signal frame, calculate the ratio between the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the test audio signal and the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the reference audio signal, and obtain the amplitude ratio of the audio signal frame corresponding to that frame number. Perform a logarithmic transformation on the amplitude ratio to obtain the logarithmic amplitude difference of the audio signal frame corresponding to that frame number; the logarithmic amplitude difference is used to represent decibels.
[0158] S1052: Construct a frequency response difference curve based on the logarithmic amplitude difference corresponding to each frame number; where the horizontal axis of the frequency response difference curve is the center point time of the denoised audio signal frame, and the vertical axis is decibels.
[0159] S1053: Determine whether the audio device is malfunctioning based on the degree of deviation between the frequency response difference curve and the horizontal axis.
[0160] In step S1051, the two-dimensional time-frequency characteristic curve corresponding to the test audio signal is as follows:
[0161]
[0162] The two-dimensional time-frequency characteristic curve corresponding to the reference audio signal is as follows:
[0163]
[0164] In this embodiment, the logarithmic magnitude difference corresponding to the audio signal frames of this frame number is calculated using the following formula:
[0165]
[0166] in, This represents the amplitude of the k-th frame of the audio signal after denoising in the two-dimensional time-frequency characteristic curve corresponding to the test audio signal. This represents the amplitude of the k-th frame of the denoised audio signal in the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal. This represents the logarithmic magnitude difference (in dB) corresponding to the k-th frame of the audio signal after denoising.
[0167] In this embodiment, the logarithmic amplitude difference is used to represent decibels. It can be seen from the above formula that when... Greater than When, after taking the logarithm If it is a positive number, it indicates that the amplitude of the test audio signal is greater than the amplitude of the reference audio signal. When Less than When, after taking the logarithm If it is a negative number, it means that the amplitude of the test audio signal is less than the amplitude of the reference audio signal.
[0168] In step S1052, a frequency response difference curve is constructed based on the logarithmic magnitude difference corresponding to each frame number; wherein, the horizontal axis of the frequency response difference curve is the center point time (or frame number) of the denoised audio signal frame, and the vertical axis is decibels (i.e., logarithmic magnitude).
[0169] In step S1053, under ideal conditions (i.e., the audio device is free of any faults and in an ideal, noise-free external environment), the frequency response difference curve should be coaxial with the horizontal axis (i.e., it should be on the same straight line as the horizontal axis, and the curve value should be 0 at all points).
[0170] If the value of the frequency response difference curve (vertical axis) is positive, it indicates that the audio signal played by the audio device during the fault detection phase is stronger (stronger than the audio signal played during the non-fault phase). Conversely, if the value of the frequency response difference curve (vertical axis) is negative, it indicates that the audio signal played by the audio device during the fault detection phase is weaker (weaker than the audio signal played during the non-fault phase).
[0171] Both excessively strong and excessively weak audio signals indicate an abnormality (malfunction) in the audio equipment. Therefore, when determining whether an audio device is malfunctioning based on the degree of deviation of the frequency response difference curve from the horizontal axis, the specific steps could be as follows:
[0172] If the deviation is greater than the preset level, it indicates that the audio device has malfunctioned; if the deviation is less than or equal to the preset level, it indicates that the audio device has not malfunctioned. Both positive and negative deviations can be represented by the absolute value of the distance between the frequency response difference curve and the horizontal axis.
[0173] In this embodiment, the degree of deviation is positively correlated with the probability (i.e., likelihood) of a malfunction. The greater the deviation, the more likely the audio device is to malfunction, and the higher the probability of a malfunction; the smaller the deviation, the more likely the audio device is to function normally, and the lower the probability of a malfunction.
[0174] In one possible implementation, in the frequency response difference curve, the horizontal axis represents not only the center point time (or frame number) of the audio signal frame, but also the frequency band (frequency range) of the audio signal frame. If the deviation of the frequency response difference curve is too large within a relatively long frequency band, it is considered that the audio device has malfunctioned. Conversely, if the deviation of the frequency response difference curve is too large within a very short frequency band, it may indicate interference. Therefore, to avoid misjudgment, in this embodiment, the frequency response difference curve is first smoothed to prevent abrupt changes that could lead to misjudgment. Specifically, when executing step S1053, the following steps can be followed:
[0175] S10531: Using the constructed smooth convolution kernel, the frequency response difference curve is smoothly convolved to obtain the smoothed time-decibel curve; the horizontal axis of the time-decibel curve is the center point time of the denoised audio signal frame, and the vertical axis is decibel.
[0176] S10532: Determine the center frequency of each audio signal frame based on the number of frames of each audio signal frame, and convert the time-decibel curve into a frequency-decibel curve with the center frequency on the horizontal axis and decibels on the vertical axis based on the center frequency of each audio signal frame.
[0177] S10533: Determine whether the audio device is malfunctioning based on the degree of deviation of the frequency-decibel curve from the horizontal axis; wherein, the degree of deviation is positively correlated with the probability of malfunction.
[0178] In step S10531, a smooth convolution kernel h is first constructed:
[0179]
[0180] Where M represents the length of the smooth convolution kernel.
[0181] The frequency response difference curve D is smoothed by a smooth convolution kernel h:
[0182]
[0183] in, This represents the time-decibel curve after smoothing the convolution; This represents the logarithmic magnitude difference corresponding to the (k+i)th frame of the audio signal; This represents the convolution operation.
[0184] For the k-th point in the time-decibel curve (i.e. the k-th audio signal frame), if it is higher or lower than the threshold, such as higher than +6dB or lower than -6dB, it can be considered that the audio device has a fault at the center frequency corresponding to the k-th frame.
[0185] In step S10532, the center point time t of the audio signal frame corresponding to the horizontal axis is converted into the center frequency f of the audio signal frame, and then mapped onto the frequency axis. The conversion formula is as follows:
[0186]
[0187] Where f1 represents the starting frequency of the reference audio signal (e.g., 20Hz); f2 represents the ending frequency of the reference audio signal (e.g., 20kHz); T represents the total sweep time, i.e., the total duration of the reference audio signal; t k This represents the center point time of the k-th audio signal frame.
[0188] In step S10533, as Figure 5 As shown, the degree of deviation of the frequency-decibel curve from the horizontal axis is positively correlated with the probability (likelihood) of a malfunction. The greater the deviation, the more likely the audio device is to malfunction; the smaller the deviation, the more likely the audio device is to function normally.
[0189] Based on the same technical concept, embodiments of this application also provide an audio device fault detection device, such as... Figure 6 As shown, the device includes:
[0190] The acquisition module 601 is used to acquire a reference audio signal and a test audio signal of the audio device to be tested; the reference audio signal is the audio signal played by the audio device during the playback of the reference audio signal when the audio device is not malfunctioning; the test audio signal is the audio signal played by the audio device during the playback of the reference audio signal during the fault detection phase of the audio device.
[0191] The framing module 602 is used to perform framing processing on each audio signal in the reference audio signal and the test audio signal to obtain multiple audio signal frames of the audio signal.
[0192] The noise reduction module 603 is used to remove noise from each audio signal frame of the audio signal using wavelet transform to obtain a noise-reduced audio signal frame.
[0193] The calculation module 604 is used to calculate the amplitude of each denoised audio signal frame at its center frequency, and generate a two-dimensional time-frequency characteristic curve corresponding to the audio signal based on the amplitude of each denoised audio signal frame at its center frequency.
[0194] The judgment module 605 is used to determine whether the audio device has malfunctioned based on the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal.
[0195] Optionally, the reference audio signal is a frequency sweep signal with an exponentially increasing phase continuity; the starting frequency of the reference audio signal is low frequency, and the ending frequency is high frequency.
[0196] Optionally, the audio signal is a discrete audio signal, and the signal sequence of the discrete audio signal includes the amplitude of the audio signal at each sampling point;
[0197] When the framing module 602 performs framing processing on each audio signal in the reference audio signal and the test audio signal to obtain multiple audio signal frames, it is specifically used for:
[0198] For each audio signal in the reference audio signal and the test audio signal, the audio signal is divided into frames according to a preset frame length and frame shift to obtain multiple audio signal frames of the audio signal; wherein, the frame length is used to control the number of sampling points in each audio signal frame; the frame shift is used to control the number of sampling points shifted backward relative to the starting point of the previous frame in two adjacent audio signal frames; the frame length is greater than the frame shift and less than the total duration of the reference audio signal.
[0199] Optionally, the audio signal is a discrete audio signal, and the signal sequence of the discrete audio signal includes the amplitude of the audio signal at each sampling point; the denoising module 603, when used to remove noise from each audio signal frame of the audio signal using wavelet transform to obtain a denoised audio signal frame, is specifically used for:
[0200] For each audio signal frame, a discrete wavelet transform is performed on the audio signal frame based on the number of wavelet decomposition levels set in the discrete wavelet transform, to obtain the detail coefficients of the audio signal frame at each frequency level; wherein, the detail coefficients represent the energy intensity of the audio signal at that frequency level; each frequency level corresponds to a frequency band, and the frequency band corresponding to each frequency level of the audio signal frame is located within the frequency band of the audio signal frame; in the same audio signal frame, the higher the number of frequency levels, the closer the frequency band of that frequency level is to the low-frequency end of the frequency band of the audio signal frame;
[0201] Based on the center frequency of the audio signal frame and the center frequencies of each frequency layer corresponding to the audio signal frame, the frequency layer closest to the center frequency of the audio signal frame is determined from each frequency layer corresponding to the audio signal frame and is used as the target frequency layer of the audio signal frame.
[0202] The detail coefficients of the target frequency layer of the audio signal frame are retained, and the detail coefficients of other frequency layers of the audio signal frame are set to zero to obtain the denoised detail coefficient sequence corresponding to the audio signal frame.
[0203] Wavelet reconstruction is performed on the denoised detail coefficient sequence corresponding to the audio signal frame to obtain the reconstructed audio signal frame, which is then used as the denoised audio signal frame.
[0204] Optionally, when calculating the amplitude of each denoised audio signal frame at its center frequency, the calculation module 604 is specifically used for:
[0205] For each denoised audio signal frame, calculate the continuous wavelet transform coefficients of the denoised audio signal frame, and determine the amplitude of the denoised audio signal frame at its center frequency based on the continuous wavelet transform coefficients.
[0206] Optionally, when the calculation module 604 calculates the continuous wavelet transform coefficients of the denoised audio signal frame and determines the amplitude of the denoised audio signal frame at its center frequency based on the continuous wavelet transform coefficients, it specifically performs the following:
[0207] Calculate the two-dimensional coefficient matrix of the denoised audio signal frame using the following continuous wavelet transform function:
[0208]
[0209] Among them, W k (a,b) represents the two-dimensional coefficient matrix of the k-th audio signal frame after denoising; a is the scaling factor, representing the vertical coordinate of the continuous wavelet transform coefficients contained in the two-dimensional coefficient matrix; b is the translation factor, representing the horizontal coordinate of the continuous wavelet transform coefficients contained in the two-dimensional coefficient matrix. Indicates the mother wavelet; Indicates complex conjugation; This represents the k-th frame of the denoised audio signal; n represents the n-th sampling point of the k-th frame of the denoised audio signal; and N is the total number of sampling points in the k-th frame of the denoised audio signal. The sampling period;
[0210] The scale factor corresponding to the center frequency of the denoised k-th frame of the audio signal is calculated using the following formula. :
[0211]
[0212] in, This represents the center frequency of the continuous wavelet transform function; This represents the center frequency of the k-th audio signal frame;
[0213] From the two-dimensional coefficient matrix corresponding to the k-th frame of the denoised audio signal, determine the ordinate as the scale factor. The absolute value of the largest coefficient among the multiple continuous wavelet transform coefficients is taken as the amplitude of the k-th frame of the denoised audio signal at its center frequency.
[0214] Optionally, the horizontal axis of the two-dimensional time-frequency characteristic curve corresponding to the audio signal is the center point time of the denoised audio signal frame, and the vertical axis is the amplitude of each denoised audio signal frame at its center frequency; there is a one-to-one correspondence between the center point time of each denoised audio signal frame and the frame number; when the judgment module 605 is used to determine whether the audio device has malfunctioned based on the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal, it is specifically used for:
[0215] For each frame number of the denoised audio signal, the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the test audio signal is calculated, and the ratio between the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the reference audio signal is obtained. The amplitude ratio corresponding to that frame number of audio signal frames is then logarithmically transformed to obtain the logarithmic amplitude difference corresponding to that frame number of audio signal frames; the logarithmic amplitude difference is used to represent decibels.
[0216] A frequency response difference curve is constructed based on the logarithmic amplitude difference corresponding to each frame number; wherein, the horizontal axis of the frequency response difference curve is the center point time of the denoised audio signal frame, and the vertical axis is decibels;
[0217] The degree of deviation of the frequency response difference curve from the horizontal axis is used to determine whether the audio device has malfunctioned.
[0218] Optionally, the judgment module 605, when calculating the amplitude of the audio signal frame corresponding to the number of frames in the two-dimensional time-frequency characteristic curve corresponding to the test audio signal for each frame number of the denoised audio signal frame, and the amplitude of the audio signal frame corresponding to the number of frames in the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal for each frame number, to obtain the amplitude ratio corresponding to the number of audio signal frames, and performing a logarithmic transformation on the amplitude ratio to obtain the logarithmic amplitude difference corresponding to the number of audio signal frames, is specifically used for:
[0219] The logarithmic magnitude difference corresponding to the audio signal frames of this frame number is calculated using the following formula:
[0220]
[0221] in, This represents the amplitude of the k-th frame of the audio signal after denoising in the two-dimensional time-frequency characteristic curve corresponding to the test audio signal; This represents the amplitude of the k-th frame of the audio signal after denoising in the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal; This represents the logarithmic magnitude difference corresponding to the k-th frame of the audio signal after denoising.
[0222] Optionally, when the judgment module 605 is used to determine whether the audio device has malfunctioned based on the degree of deviation of the frequency response difference curve from the horizontal axis, it is specifically used for:
[0223] Using the constructed smooth convolution kernel, the frequency response difference curve is smoothed and convolved to obtain the smoothed time-decibel curve; the horizontal axis of the time-decibel curve is the center point time of the denoised audio signal frame, and the vertical axis is decibels;
[0224] Based on the number of frames of each audio signal frame, the center frequency of each audio signal frame is determined. Based on the center frequency of each audio signal frame, the time-decibel curve is converted into a frequency-decibel curve with the center frequency on the horizontal axis and decibels on the vertical axis.
[0225] The degree of deviation of the frequency-decibel curve from the horizontal axis is used to determine whether the audio device is malfunctioning; wherein, the degree of deviation is positively correlated with the probability of malfunction.
[0226] Figure 7 A schematic diagram of an electronic device provided in this application embodiment includes: a processor 701, a memory 702, and a bus 703. The memory 702 stores machine-readable instructions executable by the processor 701. When the electronic device runs the above-described information processing method, the processor 701 and the memory 702 communicate through the bus 703. The processor 701 executes the machine-readable instructions to perform the steps of the method described in Embodiment 1.
[0227] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps described in Embodiment 1.
[0228] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, electronic devices, and computer-readable storage media described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0229] In the several embodiments provided in this application, it should be understood that the disclosed methods, apparatuses, electronic devices, and computer-readable storage media can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or modules may be electrical, mechanical, or other forms.
[0230] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0231] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0232] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0233] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the claims.
Claims
1. A method for detecting faults in audio equipment, characterized in that, include: Acquire the reference audio signal and test audio signal of the audio device to be tested; The reference audio signal is the audio signal played by the audio device during the playback of the reference audio signal when the audio device is not malfunctioning; the test audio signal is the audio signal played by the audio device during the playback of the reference audio signal during the fault detection phase of the audio device. For each audio signal in the reference audio signal and the test audio signal, the audio signal is divided into frames to obtain multiple audio signal frames of the audio signal; For each audio signal frame of the audio signal, wavelet transform is used to remove noise from the audio signal frame, resulting in a denoised audio signal frame. Calculate the amplitude of each denoised audio signal frame at its center frequency, and generate a two-dimensional time-frequency characteristic curve corresponding to the audio signal based on the amplitude of each denoised audio signal frame at its center frequency. Based on the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal, it is determined whether the audio device has malfunctioned. The horizontal axis of the two-dimensional time-frequency characteristic curve corresponding to the audio signal is the center point time of the denoised audio signal frame, and the vertical axis is the amplitude of each denoised audio signal frame at its center frequency; there is a one-to-one correspondence between the center point time of each denoised audio signal frame and the frame number; the determination of whether the audio device has malfunctioned based on the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal includes: For each frame number of the denoised audio signal, the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the test audio signal is calculated, and the ratio between the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the reference audio signal is obtained. The amplitude ratio corresponding to that frame number of audio signal frames is then logarithmically transformed to obtain the logarithmic amplitude difference corresponding to that frame number of audio signal frames; the logarithmic amplitude difference is used to represent decibels. A frequency response difference curve is constructed based on the logarithmic amplitude difference corresponding to each frame number; wherein, the horizontal axis of the frequency response difference curve is the center point time of the denoised audio signal frame, and the vertical axis is decibels; The degree of deviation of the frequency response difference curve from the horizontal axis is used to determine whether the audio device has malfunctioned.
2. The method according to claim 1, characterized in that, The reference audio signal is a frequency sweep signal with a continuously increasing phase by an exponential rate; the starting frequency of the reference audio signal is low frequency, and the ending frequency is high frequency.
3. The method according to claim 1, characterized in that, The audio signal is a discrete audio signal, and the signal sequence of the discrete audio signal contains the amplitude of the audio signal at each sampling point; for each audio signal in the reference audio signal and the test audio signal, the audio signal is divided into frames to obtain multiple audio signal frames of the audio signal, including: For each audio signal in the reference audio signal and the test audio signal, the audio signal is divided into frames according to a preset frame length and frame shift to obtain multiple audio signal frames of the audio signal; wherein, the frame length is used to control the number of sampling points in each audio signal frame; the frame shift is used to control the number of sampling points shifted backward relative to the starting point of the previous frame in two adjacent audio signal frames; the frame length is greater than the frame shift and less than the total duration of the reference audio signal.
4. The method according to claim 2, characterized in that, The audio signal is a discrete audio signal, and the signal sequence of the discrete audio signal contains the amplitude of the audio signal at each sampling point; the step of removing noise from each audio signal frame using wavelet transform to obtain a denoised audio signal frame includes: For each audio signal frame, a discrete wavelet transform is performed on the audio signal frame based on the number of wavelet decomposition levels set in the discrete wavelet transform, to obtain the detail coefficients of the audio signal frame at each frequency level; wherein, the detail coefficients represent the energy intensity of the audio signal at that frequency level; each frequency level corresponds to a frequency band, and the frequency band corresponding to each frequency level of the audio signal frame is located within the frequency band of the audio signal frame; in the same audio signal frame, the higher the number of frequency levels, the closer the frequency band of that frequency level is to the low-frequency end of the frequency band of the audio signal frame; Based on the center frequency of the audio signal frame and the center frequencies of each frequency layer corresponding to the audio signal frame, the frequency layer closest to the center frequency of the audio signal frame is determined from each frequency layer corresponding to the audio signal frame and is used as the target frequency layer of the audio signal frame. The detail coefficients of the target frequency layer of the audio signal frame are retained, and the detail coefficients of other frequency layers of the audio signal frame are set to zero to obtain the denoised detail coefficient sequence corresponding to the audio signal frame. Wavelet reconstruction is performed on the denoised detail coefficient sequence corresponding to the audio signal frame to obtain the reconstructed audio signal frame, which is then used as the denoised audio signal frame.
5. The method according to claim 1, characterized in that, The calculation of the amplitude of each denoised audio signal frame at its center frequency includes: For each denoised audio signal frame, calculate the continuous wavelet transform coefficients of the denoised audio signal frame, and determine the amplitude of the denoised audio signal frame at its center frequency based on the continuous wavelet transform coefficients.
6. The method according to claim 5, characterized in that, The calculation of the continuous wavelet transform coefficients of the denoised audio signal frame, and the determination of the amplitude of the denoised audio signal frame at its center frequency based on the continuous wavelet transform coefficients, includes: Calculate the two-dimensional coefficient matrix of the denoised audio signal frame using the following continuous wavelet transform function: Among them, W k (a,b) represents the two-dimensional coefficient matrix of the k-th audio signal frame after denoising; a is the scaling factor, representing the vertical coordinate of the continuous wavelet transform coefficients contained in the two-dimensional coefficient matrix; b is the translation factor, representing the horizontal coordinate of the continuous wavelet transform coefficients contained in the two-dimensional coefficient matrix. Indicates the mother wavelet; Indicates complex conjugation; This represents the k-th frame of the denoised audio signal; n represents the n-th sampling point of the k-th frame of the denoised audio signal; and N is the total number of sampling points in the k-th frame of the denoised audio signal. The sampling period; The scale factor corresponding to the center frequency of the denoised k-th frame of the audio signal is calculated using the following formula. : in, This represents the center frequency of the continuous wavelet transform function; This represents the center frequency of the k-th audio signal frame; From the two-dimensional coefficient matrix corresponding to the k-th frame of the denoised audio signal, determine the ordinate as the scale factor. The absolute value of the largest coefficient among the multiple continuous wavelet transform coefficients is taken as the amplitude of the k-th frame of the denoised audio signal at its center frequency.
7. The method according to claim 1, characterized in that, For each frame number of the denoised audio signal frame, the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the test audio signal is calculated, and the ratio between this amplitude and the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the reference audio signal is obtained. This amplitude ratio is then logarithmically transformed to obtain the logarithmic amplitude difference of the audio signal frames corresponding to that frame number, including: The logarithmic magnitude difference corresponding to the audio signal frames of this frame number is calculated using the following formula: in, This represents the amplitude of the k-th frame of the audio signal after denoising in the two-dimensional time-frequency characteristic curve corresponding to the test audio signal; This represents the amplitude of the k-th frame of the audio signal after denoising in the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal; This represents the logarithmic magnitude difference corresponding to the k-th frame of the audio signal after denoising.
8. The method according to claim 1, characterized in that, The step of determining whether the audio device has malfunctioned based on the degree of deviation of the frequency response difference curve from the horizontal axis includes: Using the constructed smooth convolution kernel, the frequency response difference curve is smoothed and convolved to obtain the smoothed time-decibel curve; the horizontal axis of the time-decibel curve is the center point time of the denoised audio signal frame, and the vertical axis is decibels; Based on the number of frames of each audio signal frame, the center frequency of each audio signal frame is determined. Based on the center frequency of each audio signal frame, the time-decibel curve is converted into a frequency-decibel curve with the center frequency on the horizontal axis and decibels on the vertical axis. The degree of deviation of the frequency-decibel curve from the horizontal axis is used to determine whether the audio device is malfunctioning; wherein, the degree of deviation is positively correlated with the probability of malfunction.
9. An audio equipment fault detection device, characterized in that, include: The acquisition module is used to acquire the reference audio signal and the test audio signal of the audio device to be tested; The reference audio signal is the audio signal played by the audio device during the playback of the reference audio signal when the audio device is not malfunctioning; the test audio signal is the audio signal played by the audio device during the playback of the reference audio signal during the fault detection phase of the audio device. The framing module is used to perform framing processing on each audio signal in the reference audio signal and the test audio signal to obtain multiple audio signal frames of the audio signal. The noise reduction module is used to remove noise from each audio signal frame of the audio signal using wavelet transform, so as to obtain a noise-reduced audio signal frame. The calculation module is used to calculate the amplitude of each denoised audio signal frame at its center frequency, and generate a two-dimensional time-frequency characteristic curve corresponding to the audio signal based on the amplitude of each denoised audio signal frame at its center frequency. The judgment module is used to determine whether the audio device has malfunctioned based on the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal. The horizontal axis of the two-dimensional time-frequency characteristic curve corresponding to the audio signal represents the center point time of the denoised audio signal frame, and the vertical axis represents the amplitude of each denoised audio signal frame at its center frequency; there is a one-to-one correspondence between the center point time of each denoised audio signal frame and the frame number; the judgment module, when used to determine whether the audio device has malfunctioned based on the two-dimensional time-frequency characteristic curve corresponding to the reference audio signal and the two-dimensional time-frequency characteristic curve corresponding to the test audio signal, is specifically used for: For each frame number of the denoised audio signal, the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the test audio signal is calculated, and the ratio between the amplitude of the audio signal frame corresponding to that frame number in the two-dimensional time-frequency characteristic curve of the reference audio signal is obtained. The amplitude ratio corresponding to that frame number of audio signal frames is then logarithmically transformed to obtain the logarithmic amplitude difference corresponding to that frame number of audio signal frames; the logarithmic amplitude difference is used to represent decibels. A frequency response difference curve is constructed based on the logarithmic amplitude difference corresponding to each frame number; wherein, the horizontal axis of the frequency response difference curve is the center point time of the denoised audio signal frame, and the vertical axis is decibels; The degree of deviation of the frequency response difference curve from the horizontal axis is used to determine whether the audio device has malfunctioned.
10. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the memory via the bus, and the machine-readable instructions, when executed by the processor, perform the steps of the method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Equipment fault monitoring method and system and storage medium
CN114136600A
Audio equipment detection method and device, electronic equipment and readable storage medium
CN118764810A