Methods, devices and storage media for detecting plosive sounds
By performing spectral energy normalization and short-time Fourier transform on the audio signal, combined with confidence intervals and detection thresholds, the problem of difficult identification of plosive sounds in speaker electroacoustic performance testing was solved, and accurate detection of plosive sounds was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-08
- Publication Date
- 2026-03-13
AI Technical Summary
Existing loudspeaker electroacoustic performance testing systems cannot effectively detect the brief popping noise (POP sound) generated during the sound production process, which affects the testing results.
By collecting the spectral energy of the audio signal, normalizing it and performing a short-time Fourier transform, determining the quantization value, and detecting plosive sounds based on the confidence interval and detection threshold, and removing broad-spectrum noise to accurately identify plosive sounds.
It enables accurate detection of plosive sounds in audio signals, including faint and masked plosive sounds, improving the accuracy and reliability of detection.
Smart Images

Figure CN119601038B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of speaker unit audio performance testing, and in particular to a method, apparatus and storage medium for detecting plosive sounds. Background Technology
[0002] Among related technologies, there are many schemes based on SoundCheck for testing the electroacoustic performance of loudspeakers. SoundCheck itself is a precise and powerful software-based electroacoustic and audio electronic measurement system capable of quickly measuring parameters such as frequency response, phase, sensitivity, distortion, directivity, impedance, and noise. However, these technologies cannot measure the brief popping noise emitted by the loudspeaker during sound production, thus affecting the actual effectiveness of popping noise detection. Summary of the Invention
[0003] To overcome the problems existing in related technologies, this disclosure provides a method, apparatus and storage medium for detecting plosive sounds.
[0004] According to a first aspect of the present disclosure, a method for detecting plosive sounds is provided, comprising:
[0005] The audio signal emitted by the audio device is acquired, and multiple spectral energies of the audio signal in the time domain are determined; the multiple spectral energies are normalized to obtain multiple quantization values, and there is a one-to-one correspondence between the multiple quantization values and the multiple spectral energies; the audio signal is subjected to a short-time Fourier transform to obtain the sum of multiple spectral energies in the time series; based on the sum of the multiple spectral energies and quantization values in the time series, plosive sounds are detected in the audio signal emitted by the audio device.
[0006] In one embodiment, the normalization processing of the plurality of spectral energies to obtain a plurality of quantization values, and the normalization processing of the time-frequency information of the audio signal to obtain a plurality of quantization values, include:
[0007] Based on the time-frequency information of the audio signal, a first confidence interval is determined according to a first confidence level, and the spectral data energy located in the first confidence interval among the multiple spectral data energies is determined; the mean of the spectral energy of the spectral data located in the first confidence interval is determined; a parameter value for normalizing the mean value is determined, and the spectral energy of the multiple spectral data is normalized based on the parameter value to obtain multiple quantized values corresponding to the spectral energy of the multiple spectral data, wherein the multiple quantized values have a one-to-one correspondence with the multiple spectral data.
[0008] In one embodiment, detecting plosive sounds in the audio signal emitted by the audio device based on the plurality of quantization values includes:
[0009] Based on the multiple quantization values, a broad-spectrum noise quantization value corresponding to the broad-spectrum noise value is determined, wherein the broad-spectrum noise is the steady-state noise emitted by the audio signal for testing; the difference between the multiple quantization values and the broad-spectrum noise quantization value is determined; the spectral energy and the represented audio of the quantization value whose difference is greater than the plosive sound detection threshold are determined as plosive sounds.
[0010] In one embodiment, determining the broad-spectrum noise quantization value corresponding to the broad-spectrum noise value based on the plurality of quantization values includes:
[0011] A second confidence interval is determined according to a second confidence level, and the quantized value that falls within the second confidence interval among the plurality of quantized values is determined; the maximum quantized value among the mean of the quantized values that fall within the second confidence interval is taken as the quantized value of the spectral noise corresponding to the spectral noise value.
[0012] In one implementation, the spectral data determines multiple spectral energies of the audio signal in the time domain, including:
[0013] The audio signal is transformed in the time-frequency domain to obtain the spectral energy of the audio signal in the time domain; the time domain is divided into multiple time intervals according to a preset time interval, and the spectral energy in each of the multiple time intervals is summed to obtain the sum of the spectral energy corresponding to the multiple time intervals; the sum of the multiple spectral energies is taken as the multiple spectral energies of the audio signal in the time domain.
[0014] According to a second aspect of the present disclosure, a plosive sound detection device is provided, comprising:
[0015] The acquisition unit is used to acquire the audio signal emitted by the audio device and determine multiple spectral energies of the audio signal in the time domain; the processing unit is used to normalize the multiple spectral energies to obtain multiple quantization values, wherein there is a one-to-one correspondence between the multiple quantization values and the multiple spectral energies; the detection unit is used to detect plosive sounds present in the audio signal emitted by the audio device based on the multiple quantization values.
[0016] In one embodiment, the processing unit normalizes the plurality of spectral energies to obtain a plurality of quantized values in the following manner:
[0017] A first confidence interval is determined according to a first confidence level, and the spectral energy located in the first confidence interval among the plurality of spectral energies is determined; the mean of the spectral energy located in the first confidence interval is determined; a parameter value for normalizing the mean is determined, and the plurality of spectral energies are normalized based on the parameter value to obtain a plurality of quantized values corresponding to the plurality of spectral energies.
[0018] In one embodiment, the detection unit detects plosive sounds in the audio signal emitted by the audio device based on the plurality of quantization values in the following manner:
[0019] Based on the multiple quantization values, a broad-spectrum noise quantization value is determined, whereby the broad-spectrum noise is the steady-state noise emitted by the audio signal for testing; the difference between the multiple quantization values and the broad-spectrum noise quantization value is determined; and the audio represented by the spectral energy of the quantization value whose difference is greater than the plosive sound detection threshold is identified as a plosive sound.
[0020] In one embodiment, the detection unit determines the broad-spectrum noise quantization value based on the plurality of quantization values in the following manner:
[0021] A second confidence interval is determined according to a second confidence level, and the quantized value that falls within the second confidence interval among the plurality of quantized values is determined; the largest quantized value among the quantized values that falls within the second confidence interval is taken as the broad-spectrum noise quantized value.
[0022] In one embodiment, the acquisition unit determines multiple spectral energies of the audio signal in the time domain in the following manner:
[0023] The audio signal is transformed in the time-frequency domain to obtain the spectral energy of the audio signal in the time domain; the time domain is divided into multiple time intervals according to a preset time interval, and the spectral energy in each of the multiple time intervals is summed to obtain the sum of the spectral energy corresponding to the multiple time intervals; the sum of the multiple spectral energies is taken as the multiple spectral energies of the audio signal in the time domain.
[0024] According to a third aspect of the present disclosure, an audio control device is provided, comprising:
[0025] A processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the function control method described in the first aspect or any embodiment of the first aspect.
[0026] According to a fourth aspect of the present disclosure, a storage medium is provided, the storage medium storing instructions that, when executed by a processor of a terminal, enable the terminal to perform the method described in the first aspect or any one of the embodiments of the first aspect.
[0027] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: the test equipment collects the audio signal emitted by the audio device, determines the spectral energy of the collected audio signal in the time domain, performs normalization processing, and obtains a quantization value to measure the plosive sounds in the audio signal. This enables testing of any audio signal, and detection of slight plosive sounds or even masked plosive sounds.
[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0029] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0030] Figure 1 This is a schematic diagram illustrating a plosive sound testing environment according to an exemplary embodiment.
[0031] Figure 2 This is a schematic diagram illustrating a time spectrum containing a pop sound according to an exemplary embodiment.
[0032] Figure 3 This is a flowchart illustrating a method for detecting plosive sounds according to an exemplary embodiment.
[0033] Figure 4 This is a flowchart illustrating a spectral energy normalization method according to an exemplary embodiment.
[0034] Figure 5 This is a flowchart illustrating a method for detecting plosive sounds according to an exemplary embodiment.
[0035] Figure 6 This is a flowchart illustrating a method for determining a broad spectrum noise quantization value according to an exemplary embodiment.
[0036] Figure 7 This is a flowchart illustrating a method for obtaining spectral energy according to an exemplary embodiment.
[0037] Figure 8 This is a schematic diagram illustrating the correspondence between time and spectrum information according to an exemplary embodiment.
[0038] Figure 9 This is a flowchart illustrating a plosive sound detection method according to an exemplary embodiment.
[0039] Figure 10 This is a schematic diagram illustrating the difference between time and multiple quantization values and a broad spectrum noise quantization value according to an exemplary embodiment.
[0040] Figure 11 This is a block diagram illustrating a plosive sound detection device according to an exemplary embodiment.
[0041] Figure 12 This is a block diagram illustrating an apparatus for detecting plosive sounds according to an exemplary embodiment. Detailed Implementation
[0042] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure.
[0043] In the accompanying drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, of the embodiments of this disclosure. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure. The embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.
[0044] The plosive sound detection method provided in this disclosure can be applied to scenarios involving testing the audio performance of a speaker unit. By introducing audio detection that includes confidence interval calculation, it is possible to accurately determine whether a plosive sound exists in a sound-generating audio device. It should be understood that a plosive sound can also be represented as a pop sound. This disclosure solves the problems of stringent pop sound detection conditions, the inability to calculate pop sounds generated during steady-state signal transitions, the requirement that pop sounds be very distinct, and the inability to detect pop sounds with lower energy or those caused by other factors such as distortion, especially pop sounds that are masked by the original signal in the time domain.
[0045] Figure 1 This is a schematic diagram illustrating a plosive sound testing environment according to an exemplary embodiment.
[0046] In this embodiment, the audio control method described above is illustrated using the following example: the detection device 100 is a computer, the signal conversion device 101 is a sound card, the audio output device 102 is a true wireless stereo (TWS) earphone, and the sound receiving device 103 is a microphone. First, the TWS earphone continuously plays a specific signal using the Advanced Audio Distribution Profile (A2DP) protocol. Then, the signal is recorded by an artificial ear and transmitted back to the computer via the sound card. The computer locally calculates and analyzes the recorded signal to produce pop sounds.
[0047] Figure 2 This is a schematic diagram illustrating a time spectrum containing a pop sound according to an exemplary embodiment.
[0048] In related technologies, only the energy difference between adjacent frames of an audio signal is calculated, which can only calculate the pop sound generated during the transition of a steady-state signal, for example... Figure 2 The identifiable pop sound signals marked in the text cannot distinguish pop sounds with low energy or those caused by other reasons such as distortion, especially pop sounds that are masked in the time domain by the original signal, for example... Figure 2 The difficult-to-identify pop sound signals are marked in the image.
[0049] In view of this, the present disclosure provides a method for detecting plosive sounds. In this method, the acquired audio data is quantized through operations such as normalization and removal of broad-spectrum noise, and then compared with a detection threshold. Audio frequencies exceeding the detection threshold are identified as plosive sounds. This enables the detection of even small or masked pop sounds.
[0050] Figure 3 This is a flowchart illustrating a plosive sound detection method according to an exemplary embodiment, such as... Figure 3 As shown, the plosive sound control method includes the following steps.
[0051] In step S11, the audio signal emitted by the audio device is acquired, and the multiple spectral energies of the audio signal in the time domain are determined.
[0052] In this embodiment of the disclosure, the audio signal includes at least a test fixed-frequency signal emitted by the audio device and a plosive sound signal. Before determining the multiple spectral energies of the audio signal in the time domain, the audio signal can also be filtered to remove ambient noise that may interfere with detection.
[0053] In this embodiment of the disclosure, the audio device needs to play a steady-state signal during the test, such as steady-state noise like white noise or a fixed-frequency signal, to avoid test errors caused by sudden changes in the test audio signal.
[0054] In step S12, the multiple spectral energies are normalized to obtain multiple quantization values, and there is a one-to-one correspondence between the multiple quantization values and the multiple spectral energies.
[0055] In this embodiment of the disclosure, the spectral energy of the audio data includes both the time-domain and frequency-domain characteristics of the audio data. For example, the time-spectrum graph of the audio data is obtained through a short-time Fourier transform. It should be understood that the method for obtaining the spectral energy can be other than the short-time Fourier transform, and this disclosure does not limit it.
[0056] In this embodiment of the disclosure, normalization of spectral energy data can standardize the data from each test, unifying data obtained from different batches and under different conditions. Normalization maps the data to a certain range, giving it a uniform numerical range for easier analysis and comparison.
[0057] In step S13, based on multiple quantization values in the time series, popping sounds are detected in the audio signal emitted by the audio device.
[0058] In this embodiment of the disclosure, plosive sounds contained in the audio signal are obtained by detecting the normalized quantization value. The quantization value of the normalized plosive sound is significantly different from the quantization value that does not contain plosive sounds.
[0059] In this embodiment of the disclosure, the normalization of the spectral energy further includes normalization processing using the mean of the confidence interval.
[0060] Figure 4 This is a flowchart illustrating a spectral energy normalization method according to an exemplary embodiment, such as... Figure 4 As shown, the spectral energy normalization method includes the following steps.
[0061] In step S21, a first confidence interval is determined according to a first confidence level, and the spectrum data energy located in the first confidence interval among multiple spectrum data energies is determined.
[0062] In this embodiment of the disclosure, the first confidence interval represents a range of spectral energy, and the first confidence level is used to represent the probability that a randomly selected spectral energy falls within the range represented by the corresponding first confidence interval. Therefore, if the first confidence interval has a first confidence level of n, then out of a total of m spectral energies, there are n×m spectral energies that fall within the first confidence interval. For example, assuming there are 10 spectral energies and the first confidence level corresponding to the first confidence interval is 0.9, then 9 of the 10 time-frequency information energies exist within the first confidence interval.
[0063] In step S22, the mean value of the spectral energy located in the first confidence interval is determined.
[0064] Standardization of spectral energy data is achieved by determining the mean of the spectral energy within the first confidence interval.
[0065] In step S23, the parameter value for mean normalization is determined, and the multiple spectral energies are normalized based on the parameter value to obtain multiple quantization values corresponding to the multiple spectral energies.
[0066] In this embodiment of the disclosure, if the mean of the spectral data in the first confidence interval is multiplied by a parameter value equal to 1, then the parameter value is the reciprocal of the mean of the spectral data in the first confidence interval. Normalization of the spectral data is achieved by multiplying all spectral data by this parameter value. Normalization standardizes the spectral data, making it comparable to other normalized audio data.
[0067] In this embodiment, the spectral energy is normalized. However, based on the acquired multiple quantization values, it is also necessary to remove broad-spectrum noise and compare it with a detection threshold to detect plosive sounds in the audio signal.
[0068] Figure 5 This is a flowchart illustrating a plosive sound detection method according to an exemplary embodiment, such as... Figure 5 As shown, the method for detecting plosive sounds includes the following steps.
[0069] In step S31, a quantization value for the broad spectrum noise is determined based on multiple quantization values. The broad spectrum noise is the steady-state noise emitted by the audio signal for testing.
[0070] In this embodiment of the disclosure, the broad-spectrum noise quantization value is used to represent the quantization value of the spectral energy contained in the audio signal emitted by the audio device during the test. By determining the broad-spectrum noise quantization value, the dynamic adjustment process of the broad-spectrum noise after normalization can be reflected.
[0071] In step S32, the differences between multiple quantization values and the quantization values of the broad spectrum noise are determined.
[0072] In this embodiment of the disclosure, the normalized data is subtracted from the quantized value of the broad-spectrum noise to obtain the quantized value of the spectral energy after removing most of the broad-spectrum noise. By determining the differences between multiple quantized values and the quantized value of the broad-spectrum noise, the quantized value data of the spectral energy mainly containing plosive sounds can be obtained.
[0073] In step S33, the audio with spectral energy characterization corresponding to the quantization value whose difference is greater than the plosive sound detection threshold is identified as a plosive sound.
[0074] In this embodiment of the disclosure, the difference is compared with a plosive sound detection threshold. If the difference is greater than the threshold, it is determined to be a plosive sound; if it is less than the threshold, it is determined to be a non-plosive sound. It should be understood that the plosive sound detection threshold can be adjusted by the user based on subjective perception. By comparing the difference with the plosive sound detection threshold, accurate detection of plosive sounds can be achieved.
[0075] In this embodiment of the disclosure, plosive sounds are detected by representing the spectral energy of the audio signal whose difference is greater than the quantization value corresponding to the plosive sound detection threshold. However, to achieve plosive sound detection, it is necessary to specifically confirm the quantization value of the broad-spectrum noise.
[0076] Figure 6 This is a flowchart illustrating a method for determining a broad spectrum noise quantization value according to an exemplary embodiment, such as... Figure 6 As shown, the method for determining the quantization value of broad-spectrum noise includes the following steps.
[0077] In step S41, a second confidence interval is determined according to the second confidence level, and the quantized value that is located in the second confidence interval among multiple quantized values is determined.
[0078] In this embodiment of the disclosure, the second confidence interval is used to represent the range of a quantized value, and the second confidence level is used to represent the probability that a randomly selected quantized value falls within the range represented by the second confidence interval. The quantized value located within the second confidence interval can be determined using the second confidence level and the second confidence interval.
[0079] In step S42, the maximum quantization value among the quantization values located in the second confidence interval is taken as the quantization value of the broad spectrum noise.
[0080] In this embodiment of the disclosure, the broad-spectrum noise quantization value is used to represent the quantization value of the audio energy emitted by the steady-state noise of the test in the audio signal. By using the broad-spectrum noise quantization value, the magnitude of the broad-spectrum noise contained in the quantization value at each time point is represented, allowing for dynamic adjustment of the detection threshold.
[0081] In this embodiment of the disclosure, by setting a second confidence interval and quantizing the broad spectrum noise, the spectral energy can be determined by the time domain information and the frequency domain information corresponding to each time unit.
[0082] Figure 7 This is a flowchart illustrating a method for obtaining spectral energy according to an exemplary embodiment, such as... Figure 7 As shown, the method for obtaining spectral energy includes the following steps.
[0083] In step S51, the audio signal is transformed in the time-frequency domain to obtain the spectral energy of the audio signal in the time domain.
[0084] In this embodiment of the disclosure, time-frequency domain transformation can be used to obtain the time-domain and frequency-domain characteristics of an audio signal, thereby acquiring the frequency-domain energy information of the audio signal in the time domain. For example, performing time-frequency domain transformation on an audio signal using STFT utilizes the sampling rate and window width. The sampling rate defines the number of samples extracted from a continuous signal per unit time to form a discrete signal, while the window width controls the accuracy of the Fourier transform. The correspondence between time and spectral information can be obtained through the time-domain information, sampling rate, and window width of the audio signal.
[0085] In step S52, the time domain is divided into multiple time intervals according to a preset time interval, and the spectral energy in each of the multiple time intervals is summed to obtain the sum of the spectral energy of the corresponding multiple time intervals.
[0086] In this embodiment of the disclosure, the preset time interval can be determined by a time-frequency domain transformation method. For example, if an STFT is used to perform a time-frequency domain transformation on the audio signal, the time interval is determined by the sampling rate and the window width in the STFT. The spectral energy within each interval is obtained by performing a Fourier transform on the time-domain information within that interval.
[0087] In step S53, the sum of multiple spectral energies is used as the multiple spectral energies of the audio signal in the time domain.
[0088] In this embodiment of the disclosure, since the spectral energy contained in the plosive sound is much greater than that contained in the steady-state noise, summing the spectral energy values contained at each time point can make the spectral energy value of the plosive sound more prominent.
[0089] Figure 8 This is a schematic diagram illustrating the correspondence between time and spectral energy according to an exemplary embodiment.
[0090] The correspondence between time and spectrum information in this exemplary embodiment is as follows: Figure 8 As shown, assuming this graph was recorded using a 48000 Hz sampling rate, the time-frequency domain graph is obtained through a short-time Fourier transform (STFT) with a window width of 256. The unit of time in the obtained time-frequency domain graph is the quotient of 48000 and 256. The vertical axis of the time-frequency domain graph represents frequency, and the brightness represents the energy contained in the frequency. By summing the energy of all frequencies corresponding to each time interval, the correspondence between time and spectral energy can be obtained.
[0091] Figure 9 This is a flowchart illustrating a plosive sound detection method according to an exemplary embodiment.
[0092] The plosive sound detection process of this exemplary embodiment is as follows: Figure 9As shown, the audio device is assumed to be a TWS earphone. It should be understood that the mean spectral energy of the first confidence interval can also be called the broad-spectrum baseline, and the quantized value of the broad-spectrum noise can also be called the basis of the broad-spectrum noise. First, a steady-state signal needs to be played during the test. Considering that TWS earphones can transmit any signal in A2DP mode, the earphones can play steady-state noise such as white noise or a fixed-frequency signal during recording. The signal is recorded at 48000 samples. The signal is acquired by a microphone and sound card. The test device processes the signal through STFT to obtain time-frequency domain information. The window width can be 256. Since the pop sound does not end instantly, the window width does not need to be too high. From the sampling rate and window width, the time unit can be calculated as 256 / 48000 seconds, approximately 5 milliseconds. Then, frequency band data filtering is performed, extracting spectral data within the range above 1000 Hz to reduce the additional influence of ambient noise. The spectral energy is summed over time to obtain the spectral energy sum as it changes over time. This sum represents the total spectral energy within the current 5 milliseconds. Since the energy contained in pop sounds is much greater than that of steady-state noise, summing the spectral energy better reflects the difference between pop sounds and steady-state noise; therefore, the spectral energy sum can clearly indicate the presence of pop sounds. For example, if the spectral energy detected in one 5-millisecond period is 10dB, 20dB, and 15dB, while the spectral energy detected in another 5-millisecond period is 30dB, 15dB, and 40dB, then the sum of the spectral energy clearly shows the presence of pop sounds in the latter. The spectral data is then normalized. A confidence interval for the curve is calculated using a 0.9 confidence level, and the average value of the data within the confidence interval is used as a broad-spectrum baseline. All data are then normalized to 1 based on this broad-spectrum baseline to adjust their magnitude. To remove broad-spectrum noise, a normalized curve confidence interval is calculated using a 0.8 confidence level. The maximum value within this confidence interval is used as the basis for the broad-spectrum noise. Subtracting the basis of the broad-spectrum noise from the normalized data yields the correspondence between time and the difference between the broad-spectrum reference and the basis of the broad-spectrum noise. This process allows for the dynamic removal of broad-spectrum noise from audio. A threshold is also defined; any difference exceeding this threshold is considered a pop sound.
[0093] Figure 10 This is a schematic diagram illustrating the difference between time and multiple quantization values and a broad spectrum noise quantization value according to an exemplary embodiment.
[0094] The difference between the time and quantization value and the broad-spectrum noise quantization value in this example is as follows: Figure 10 As shown in the figure, the horizontal axis represents time, the vertical axis represents the difference between the quantized value and the quantized value of the broad-spectrum noise, delta represents the relationship curve between time and the difference between the quantized value and the quantized value of the broad-spectrum noise, and limit represents the plosive sound detection threshold. When delta exceeds limit, it means that a pop sound occurred at the time corresponding to the delta exceeding limit.
[0095] In this embodiment of the disclosure, by normalizing the spectral energy and removing broad-spectrum noise through the correspondence between time and spectral energy, the accurate judgment of existing plosive sounds is achieved, avoiding the omission of plosive sounds emitted by audio devices.
[0096] Based on the same concept, this disclosure also provides a plosive sound detection device.
[0097] It is understood that the blast sound detection device provided in this disclosure includes hardware structures and / or software modules corresponding to each function in order to achieve the above-mentioned functions. In conjunction with the units and algorithm steps of the various examples disclosed in this disclosure, this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of this disclosure.
[0098] Figure 11 This is a block diagram illustrating a plosive sound detection device 200 according to an exemplary embodiment. (Refer to...) Figure 11 The device includes a data acquisition unit 201, a processing unit 202, and a detection unit 203.
[0099] The acquisition unit 201 is used to acquire the audio signal emitted by the audio device and determine the multiple spectral energies of the audio signal in the time domain.
[0100] The processing unit 202 is used to normalize multiple spectral energies to obtain multiple quantization values, and there is a one-to-one correspondence between the multiple quantization values and the multiple spectral energies.
[0101] The detection unit 203 is used to detect popping sounds in the audio signal emitted by the audio device based on multiple quantization values.
[0102] In one embodiment, the processing unit 202 normalizes multiple spectral energies in the following manner to obtain multiple quantization values: determining a first confidence interval according to a first confidence level, and determining the spectral energies located in the first confidence interval among the multiple spectral energies; determining the mean of the spectral energies located in the first confidence interval; determining a parameter value for normalizing the mean, and normalizing the multiple spectral energies based on the parameter value to obtain multiple quantization values corresponding to the multiple spectral energies.
[0103] In one embodiment, the detection unit 203 detects plosive sounds in the audio signal emitted by the audio device based on multiple quantization values in the following manner: determining a broad-spectrum noise quantization value based on multiple quantization values, where broad-spectrum noise is the steady-state noise emitted by the audio signal for testing; determining the difference between the multiple quantization values and the broad-spectrum noise quantization value; and identifying the audio with spectral energy characterization corresponding to the quantization value whose difference is greater than the plosive sound detection threshold as a plosive sound.
[0104] In one embodiment, the detection unit 203 determines a broad-spectrum noise quantization value based on multiple quantization values in the following manner: a second confidence interval is determined according to a second confidence level; the quantization value located in the second confidence interval among the multiple quantization values is determined; and the largest quantization value located in the second confidence interval is taken as the broad-spectrum noise quantization value.
[0105] In one embodiment, the acquisition unit 201 determines multiple spectral energies of the audio signal in the time domain in the following manner: based on the time domain signal, sampling rate, and window width of the audio signal, it obtains the correspondence between time and spectrum information, wherein the time in the correspondence is determined by the sampling rate and window width; based on the correspondence between time and spectrum information, it determines the energy contained in all frequencies corresponding to each time moment; and based on the energy contained in all frequencies corresponding to each time moment, it obtains the sum of the energies corresponding to each time moment.
[0106] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0107] Figure 12 This is a block diagram illustrating an apparatus 300 for detecting plosive sounds according to an exemplary embodiment. For example, apparatus 300 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0108] Reference Figure 12 The device 300 may include one or more of the following components: processing component 302, memory 304, power component 306, multimedia component 308, audio component 310, input / output (I / O) interface 312, sensor component 314, and communication component 316.
[0109] Processing component 302 typically controls the overall operation of device 300, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 302 may include one or more processors 320 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 302 may include one or more modules to facilitate interaction between processing component 302 and other components. For example, processing component 302 may include a multimedia module to facilitate interaction between multimedia component 308 and processing component 302.
[0110] Memory 304 is configured to store various types of data to support the operation of device 300. Examples of such data include instructions for any application or method operating on device 300, contact data, phonebook data, messages, pictures, videos, etc. Memory 304 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0111] The power supply component 306 provides power to the various components of the device 300. The power supply component 306 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 300.
[0112] Multimedia component 308 includes a screen that provides an output interface between the device 300 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 308 includes a front-facing camera and / or a rear-facing camera. When the device 300 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0113] Audio component 310 is configured to output and / or input audio signals. For example, audio component 310 includes a microphone (MIC) configured to receive external audio signals when device 300 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 304 or transmitted via communication component 316. In some embodiments, audio component 310 also includes a speaker for outputting audio signals.
[0114] I / O interface 312 provides an interface between processing component 302 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0115] Sensor assembly 314 includes one or more sensors for providing status assessments of various aspects of device 300. For example, sensor assembly 314 may detect the on / off state of device 300, the relative positioning of components such as the display and keypad of device 300, changes in the position of device 300 or a component of device 300, the presence or absence of user contact with device 300, the orientation or acceleration / deceleration of device 300, and temperature changes of device 300. Sensor assembly 314 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 314 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 314 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0116] Communication component 316 is configured to facilitate wired or wireless communication between device 300 and other devices. Device 300 can access wireless networks based on communication standards, such as WiFi, 3G, or a combination thereof. In one exemplary embodiment, communication component 316 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 316 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0117] In an exemplary embodiment, the apparatus 300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0118] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 304 including instructions, which can be executed by a processor 320 of the device 300 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0119] It is understood that in this disclosure, "multiple" refers to two or more, and other quantifiers are similar. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. The singular forms "a," "the," and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0120] It is further understood that the terms "first," "second," etc., are used to describe various types of information, but this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another, and do not indicate a specific order or degree of importance. In fact, the expressions "first," "second," etc., are completely interchangeable. For example, without departing from the scope of this disclosure, first information can also be referred to as second information, and similarly, second information can also be referred to as first information.
[0121] It is further understood that the terms “center,” “longitudinal,” “lateral,” “lower,” “left,” “right,” “vertical,” “horizontal,” “top,” “bottom,” “inner,” and “outer,” etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this embodiment and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation.
[0122] It can be further understood that, unless otherwise specified, "connection" includes both direct connections where no other components exist between the two parties and indirect connections where other components exist between them.
[0123] It is further understood that although operations are described in a specific order in the accompanying drawings in the embodiments of this disclosure, this should not be construed as requiring these operations to be performed in the specific order or serial order shown, or requiring all of the shown operations to be performed to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous.
[0124] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein.
[0125] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for detecting plosive sounds, characterized in that, include: Acquire audio signals emitted by an audio device and determine multiple spectral energies of the audio signal in the time domain; The multiple spectral energies are normalized to obtain multiple quantized values, and there is a one-to-one correspondence between the multiple quantized values and the multiple spectral energies. Based on the multiple quantization values, the plosive sounds present in the audio signal emitted by the audio device are detected. The normalization process for the plurality of spectral energies, resulting in a plurality of quantized values, includes: A first confidence interval is determined according to a first confidence level, and the spectral energy located within the first confidence interval among the plurality of spectral energies is determined; Determine the mean of the spectral energy located in the first confidence interval; A parameter value for normalizing the mean is determined, and the plurality of spectral energies are normalized based on the parameter value to obtain a plurality of quantized values corresponding to the plurality of spectral energies.
2. The method according to claim 1, characterized in that, The step of detecting popping sounds in the audio signal emitted by the audio device based on the plurality of quantization values includes: Based on the multiple quantization values, a broad-spectrum noise quantization value is determined, wherein the broad-spectrum noise is the steady-state noise emitted by the audio signal for testing. Determine the difference between the plurality of quantization values and the broad-spectrum noise quantization value; Audio frequencies whose spectral energy characterization corresponds to quantization values with a difference greater than the plosive detection threshold are identified as plosives.
3. The method according to claim 2, characterized in that, The step of determining the quantization value of the broad-spectrum noise based on the plurality of quantization values includes: A second confidence interval is determined according to a second confidence level, and the quantized value that falls within the second confidence interval among the plurality of quantized values is determined; The largest quantized value among the quantized values located in the second confidence interval is taken as the quantized value of the broad-spectrum noise.
4. The method according to claim 1, characterized in that, Determining the multiple spectral energies of the audio signal in the time domain includes: Perform time-frequency domain transformation on the audio signal to obtain the spectral energy of the audio signal in the time domain; The time domain is divided into multiple time intervals according to a preset time interval, and the spectral energy in each of the multiple time intervals is summed to obtain the sum of the spectral energy corresponding to the multiple time intervals; The sum of the multiple spectral energies is taken as the multiple spectral energies of the audio signal in the time domain.
5. A device for detecting plosive sounds, characterized in that, include: The acquisition unit is used to acquire audio signals emitted by the audio device and determine multiple spectral energies of the audio signal in the time domain; The processing unit is used to normalize the plurality of spectral energies to obtain a plurality of quantized values, wherein there is a one-to-one correspondence between the plurality of quantized values and the plurality of spectral energies. The detection unit is used to detect plosive sounds present in the audio signal emitted by the audio device based on the plurality of quantization values. The processing unit normalizes the multiple spectral energies to obtain multiple quantized values in the following manner: A first confidence interval is determined according to a first confidence level, and the spectral energy located within the first confidence interval among the plurality of spectral energies is determined; Determine the mean of the spectral energy located in the first confidence interval; A parameter value for normalizing the mean is determined, and the plurality of spectral energies are normalized based on the parameter value to obtain a plurality of quantized values corresponding to the plurality of spectral energies.
6. The apparatus according to claim 5, characterized in that, The detection unit detects plosive sounds in the audio signal emitted by the audio device based on the multiple quantization values in the following manner: Based on the multiple quantization values, a broad-spectrum noise quantization value is determined, wherein the broad-spectrum noise is the steady-state noise emitted by the audio signal for testing. Determine the difference between the plurality of quantization values and the broad-spectrum noise quantization value; Audio frequencies whose spectral energy characterization corresponds to quantization values with a difference greater than the plosive detection threshold are identified as plosives.
7. An audio control device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the method described in any one of claims 1 to 4.
8. A storage medium, characterized in that, The storage medium stores instructions that, when executed by the terminal's processor, enable the terminal to perform the method described in any one of claims 1 to 4.
Citation Information
Patent Citations
Sonic boom detection method and device
CN104143341A
Methods and apparatus to fingerprint an audio signal via normalization
CN113614828A