Microphone Biting Detection Method, Audio Recording Method and Computer Device
By extracting the consistency detection of the vocal fundamental frequency information and the peak point pair of interest in the sliding window frame in the audio frame, the existing problem of low accuracy of spray detection is solved, and higher precision spray detection and audio recording quality assurance is achieved.
Patent Information
- Application Number
- CN202210192723.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-02-28
AI Technical Summary
The existing method of spraying spray detects spraying by setting a loudness threshold, resulting in low detection accuracy and inability to adapt to the differences between different devices.
By obtaining the human voice audio signal of the audio frame to be detected, extracting the basic frequency information of the human voice, traversing the peak points of interest in the continuous window frame using a preset sliding window, and performing spray-spray detection based on its consistency with the human voice signal point to be detected, improving detection accuracy.
Improve the accuracy of spray-spray detection, reduce false detection, and ensure the quality of audio recording.
Smart Images

Figure CN114566169B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech recognition, and in particular to a method for detecting microphone popping, an audio recording method, an apparatus, a computer device, a storage medium, and a computer program product. Background Art
[0002] With the rise of KTV software, users can sing anytime and anywhere through mobile devices such as mobile phones. When singing, users will use the built-in microphone of the mobile phone or the earpiece attached to ordinary headphones for recording. During the singing recording process, if the microphone is placed too close, it will cause popping, resulting in the occurrence of microphone popping phenomenon, and further leading to the damage of the recording sound quality. Therefore, it is necessary to detect and suppress microphone popping behavior during song recording. Currently, the common way to deal with microphone popping behavior is to detect whether the sound is microphone popping by setting a loudness threshold. However, when detecting microphone popping by the method of loudness threshold, the threshold standard is not universal for different devices. Using the same threshold for detection is prone to false detection.
[0003] Therefore, the current method for detecting microphone popping has the defect of low detection accuracy. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method for detecting microphone popping, an audio recording method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product that can improve the detection accuracy.
[0005] In a first aspect, the present application provides a method for detecting microphone popping, the method comprising:
[0006] Obtaining an audio frame to be detected, and extracting the human voice audio signal in the audio frame to be detected;
[0007] Determining a set of human voice signal points to be detected in the human voice audio signal according to the amplitude change between adjacent human voice signal points in the human voice audio signal;
[0008] Traversing the human voice audio signal through a preset sliding window, and obtaining a pair of peak points of interest in consecutive window frames obtained by traversing the preset sliding window;
[0009] Obtaining a microphone popping detection result of the audio frame to be detected according to the consistency between the pair of peak points of interest in consecutive window frames and the human voice signal points to be detected in the set of human voice signal points to be detected.
[0010] In one of the embodiments, the extracting the human voice audio signal in the audio frame to be detected includes:
[0011] Determining the fundamental frequency information of the human voice in the audio frame to be detected;
[0012] Extract the human voice audio signal in the audio frame to be detected according to the fundamental frequency information of the human voice.
[0013] In one embodiment, the determining the fundamental frequency information of the human voice in the audio frame to be detected includes:
[0014] Calculate a plurality of average amplitude differences corresponding to the audio frame to be detected in different amplitude difference calculation periods, and determine the relative threshold of the audio frame to be detected according to the relative magnitudes of the plurality of average amplitude differences;
[0015] Perform audio period division on the audio frame to be detected according to the relative threshold, and obtain the fundamental frequency information of the human voice in the audio frame to be detected according to the audio period division result.
[0016] In one embodiment, the determining the relative threshold of the audio frame to be detected according to the relative magnitudes of the plurality of average amplitude differences includes:
[0017] Obtain the maximum average amplitude difference among the plurality of average amplitude differences;
[0018] Take the product of the maximum average amplitude difference and a set coefficient as the relative threshold.
[0019] In one embodiment, the performing audio period division on the audio frame to be detected according to the relative threshold, and obtaining the fundamental frequency information of the human voice in the audio frame to be detected according to the audio period division result includes:
[0020] Obtain a plurality of target audio points in the audio frame to be detected with amplitudes less than the relative threshold, and obtain the minimum period presented by the plurality of target audio points;
[0021] Obtain the reciprocal of the minimum period to obtain the fundamental frequency information of the human voice.
[0022] In one embodiment, the extracting the human voice audio signal in the audio frame to be detected according to the fundamental frequency information of the human voice includes:
[0023] Perform low-pass filtering on the audio frame to be detected according to the fundamental frequency information of the human voice to obtain the human voice audio signal in the audio frame to be detected.
[0024] In one embodiment, the determining the set of human voice signal points to be detected in the human voice audio signal according to the amplitude change between adjacent human voice signal points in the human voice audio signal includes:
[0025] Perform differential operation on the human voice audio signal to obtain the amplitude change rate between adjacent human voice signal points in the audio frame to be detected;
[0026] Determine multiple human voice signal points corresponding to the amplitude change rate satisfying a preset amplitude change rate threshold as the set of to-be-detected human voice signal points in the human voice audio signal.
[0027] In one embodiment, for each of the continuous window frame interested peak point pairs, its window frame peak is greater than a preset window frame peak threshold.
[0028] In one embodiment, obtaining the pop noise detection result of the to-be-detected audio frame according to the consistency between the continuous window frame interested peak point pairs and the to-be-detected human voice signal points in the set of to-be-detected human voice signal points includes:
[0029] If all of the continuous window frame interested peak point pairs are in the set of to-be-detected human voice sample points, determine that the pop noise detection result is that pop noise occurs.
[0030] In a second aspect, the present application provides an audio recording method, and the method includes:
[0031] Obtain the real-time input recorded audio;
[0032] Detect whether pop noise occurs in the recorded audio according to the pop noise detection method as described above;
[0033] If so, display a pop noise reminder message on the recording interface of the recorded audio.
[0034] In a third aspect, the present application provides a pop noise detection device, and the device includes:
[0035] An acquisition module, configured to acquire a to-be-detected audio frame and extract the human voice audio signal in the to-be-detected audio frame;
[0036] A human voice detection module, configured to determine a set of to-be-detected human voice signal points in the human voice audio signal according to the amplitude change between adjacent human voice signal points in the human voice audio signal;
[0037] A peak detection module, configured to traverse the human voice audio signal through a preset sliding window and obtain continuous window frame interested peak point pairs obtained by traversing the preset sliding window;
[0038] A pop noise detection module, configured to obtain the pop noise detection result of the to-be-detected audio frame according to the consistency between the continuous window frame interested peak point pairs and the to-be-detected human voice signal points in the set of to-be-detected human voice signal points.
[0039] In a fourth aspect, the present application provides an audio recording device, and the device includes:
[0040] A recording module, configured to acquire the real-time input recorded audio;
[0041] A checking module, configured to check whether the recorded audio has a popping sound according to the popping sound detection method as described above;
[0042] A reminding module, configured to, if so, display a popping sound reminder message on the recording interface of the recorded audio.
[0043] In a fifth aspect, the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0044] In a sixth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0045] In a seventh aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0046] For the above popping sound detection method, audio recording method, device, computer device, storage medium, and computer program product, by determining the relative threshold of the audio frame to be detected based on the relative magnitudes of multiple average amplitude differences of the audio frame to be detected under different amplitude difference calculation periods, and performing audio period division on the audio frame to be detected according to the relative threshold to obtain the fundamental frequency information of the human voice, extracting the human voice audio signal according to the fundamental frequency information of the human voice, and determining the set of human voice signal points to be detected in the human voice audio signal according to the amplitude change between adjacent human voice signal points in the human voice audio signal, and traversing the human voice audio signal through a preset sliding window to obtain the continuous window frame interested peak point pairs. According to the consistency of the continuous window frame interested peak point pairs with the human voice signal points to be detected in the set of human voice signal points to be detected, the popping sound detection result of the audio frame to be detected is obtained. Compared with the traditional method of detecting whether a user has a popping sound according to a loudness threshold, this solution detects whether a popping sound occurs in the audio frame by checking the consistency between multiple human voice signal points with relatively small amplitude changes and the continuous window frame interested peak point pairs, thereby improving the accuracy of popping sound detection. Description of the Drawings
[0047] Figure 1 It is an application environment diagram of the popping sound detection method in an embodiment;
[0048] Figure 2 It is a flowchart of the popping sound detection method in an embodiment;
[0049] Figure 3 It is a flowchart of the audio recording method in an embodiment;
[0050] Figure 4 It is a schematic diagram of the popping sound reminder interface in an embodiment;
[0051] Figure 5 It is a structural block diagram of a microphone blasting detection device in an embodiment;
[0052] Figure 6 It is a structural block diagram of an audio recording device in an embodiment;
[0053] Figure 7 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0054] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0055] The microphone blasting detection method provided by the embodiments of the present application can be applied to, for example Figure 1 the application environment shown. The terminal 102 can obtain the audio frame to be detected input by the user, and based on the amplitude change between adjacent human voice audio signal points in the audio frame to be detected, obtain the set of human voice signal points to be detected therein, and obtain the pair of peak points of interest of the continuous window frames obtained by traversing through a preset sliding window. Determine the microphone blasting detection result of the audio frame to be detected through the consistency between the pair of peak points of interest of the continuous window frames and the human voice signal points to be detected. In addition, in some embodiments, a server 104 is further included. Among them, the terminal 102 communicates with the server 104 through a network. The terminal 102 can upload the microphone blasting detection result to the server 104 for storage. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed on the cloud or other network servers. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0056] In one embodiment, as Figure 2 shown, a microphone blasting detection method is provided. Taking the terminal in Figure 1 as an example for description, the method includes the following steps:
[0057] Step S202, obtain the audio frame to be detected, and extract the human voice audio signal in the audio frame to be detected.
[0058] Among them, the audio frame to be detected can be an audio frame obtained from the audio input by the user. The user can input the audio to be detected in the audio recording interface, so that the terminal 102 can obtain multiple audio frames to be detected from the audio to be detected. Each audio frame to be detected can contain multiple sample points, and there is a corresponding audio amplitude change in each audio frame to be detected. The terminal 102 can perform pop detection on the audio input by the user in units of the audio frames to be detected. The terminal 102 can calculate multiple average amplitude differences corresponding to the audio frames to be detected under different amplitude difference calculation periods. Among them, since the above-mentioned audio frames to be detected include multiple sample points, the terminal 102 can calculate multiple short-time average amplitude differences according to different point intervals. For example, the terminal 102 can set the time series of the audio signal of the audio to be detected input by the user as x(n), and the terminal 102 can perform windowing and framing on the audio signal to obtain multiple frames of audio signals. For the audio signal x k (m) of the k-th audio frame to be detected, where the subscript k represents the k-th frame, and m represents the m-th sample point in this frame. The terminal 102 can make the length of each audio frame to be detected be N, and the short-time average amplitude difference d k (i) of this frame is calculated as follows: where i represents the amplitude difference calculation period, 1 < i < N, that is, the terminal 102 can obtain multiple average amplitude differences based on different amplitude difference calculation periods.
[0059] After the terminal 102 obtains the above multiple average amplitude differences, it can determine the relative threshold of the audio frame to be detected according to the relative magnitudes of the multiple average amplitude differences. The relative threshold can be used to perform cycle division on the audio frame to be detected. For example, other waveforms that cannot be used to determine the cycle in the audio frame to be detected are filtered out through the relative threshold, and the cycle is determined based on the filtered waveform diagram. The terminal 102 can calculate the above relative threshold through a set formula.
[0060] Specifically, in one embodiment, extracting the human voice audio signal in the audio frame to be detected includes: determining the fundamental frequency information of the human voice in the audio frame to be detected; extracting the human voice audio signal in the audio frame to be detected according to the fundamental frequency information of the human voice. The specific process of determining the fundamental frequency information of the human voice in the audio frame to be detected can include: calculating multiple average amplitude differences corresponding to the audio frame to be detected under different amplitude difference calculation periods, and determining the relative threshold of the audio frame to be detected according to the relative magnitudes of the multiple average amplitude differences; performing audio cycle division on the audio frame to be detected according to the relative threshold, and obtaining the fundamental frequency information of the human voice in the audio frame to be detected according to the audio cycle division result.
[0061] Among them, the relative threshold can be a threshold determined according to the relative magnitudes of multiple average amplitude differences. The terminal 102 can perform audio cycle division on the audio frame to be detected based on the relative threshold. For example, the terminal 102 filters out the waveforms with amplitudes greater than or equal to the relative threshold based on the relative threshold, and divides the audio cycle based on the remaining waveforms. The terminal 102 can also obtain the fundamental frequency information of the human voice in the audio frame to be detected based on the divided audio cycle division result. The terminal 102 can obtain the above-mentioned fundamental frequency information of the human voice by performing a specific operation on the above audio cycle. Among them, the fundamental frequency information of the human voice can be information that determines the pitch of the human voice.
[0062] Step S204: Determine the set of human voice signal points to be detected in the human voice audio signal according to the amplitude change between adjacent human voice signal points in the human voice audio signal.
[0063] Among them, the fundamental frequency information of the human voice can be a type of fundamental frequency information. The fundamental frequency is the lowest oscillation frequency of a free oscillation system, the lowest frequency in a complex wave, and in the human voice signal, the fundamental frequency determines the pitch. Then the terminal 102 can extract the human voice audio signal in the audio to be detected based on the fundamental frequency information of the human voice. For example, the human voice audio signal can be obtained by filtering based on the fundamental frequency information of the human voice. Among them, the human voice audio signal includes multiple human voice signal points. After the terminal 102 obtains the human voice audio signal, it can determine the set of human voice signal points to be detected in the human voice audio signal according to the amplitude change between adjacent human voice signal points in the human voice audio signal. For example, the terminal 102 can use adjacent human voice signal points with smaller amplitude changes as the human voice signal points to be detected. There can be multiple human voice signal points to be detected, and the terminal 102 can form a set of human voice signal points to be detected from multiple human voice signal points to be detected.
[0064] Step S206: Traverse the human voice audio signal through a preset sliding window, and obtain the pairs of peak points of interest in the continuous window frames obtained by traversing the preset sliding window.
[0065] Among them, the preset sliding window can be a rectangular window, and the human voice audio signal can be a waveform signal. The terminal 102 can traverse the human voice audio signal through the preset sliding window. For example, the terminal 102 can set the window length of the rectangular window to O, where O can be less than the length of the audio frame to be detected above. The terminal 102 can obtain the corresponding continuous window frame interested peak points based on the sliding of the above rectangular window in the human voice audio signal. There can be multiple such continuous window frame interested peak points, which are obtained based on the peak points in the current preset sliding window when the preset sliding window traverses different positions in the human voice signal. Among them, the above continuous window frame interested peak point pairs can be two adjacent window frame interested peak point pairs formed by the window frame interested peak point in the previous window and the window frame interested peak point obtained in the window after sliding after one sliding, that is, the continuous window frame interested peak point pairs are continuous peak point pairs. Specifically, the terminal 102 sets the current rectangular window length to O, starts from the waveform starting point of the above human voice audio signal, and superimposes a set number of samples in the preset sliding window. For example, 75% of the samples are superimposed, that is, the terminal 102 can slide the window every O / 4 samples. After each sliding, the terminal 102 can calculate the frame peak in the current sliding window, denoted as p k , then the above continuous window frame interested peak point pairs can be composed of two consecutive p k .
[0066] In addition, since the peak points in the human voice audio signal are not necessarily the peaks caused by plosives, after the terminal 102 obtains the peak points in the preset sliding window, it is also necessary to determine whether the peak point is greater than a certain threshold. For example, in one embodiment, the window frame peaks of the above continuous window frame interested peak point pairs are all greater than the preset window frame peak threshold. In this embodiment, the window frame peak is the peak point obtained from the preset sliding window. The terminal 102 can set the preset window frame peak threshold to aT. When the terminal 102 detects that each window frame peak p k in the above continuous window frame interested peak point pair is greater than the absolute threshold aT, the terminal 102 can determine that the continuous window frame peak point is the peak caused by plosives. Then the terminal 102 can use the continuous window frame peak point pair as the continuous window frame interested peak point pair and continue to detect plosives based on this continuous window frame interested peak point. If the terminal 102 detects that there is a p k less than the absolute threshold aT among each window frame peak p k in the above continuous window frame interested peak point pair, the terminal 102 can determine that the continuous window frame peak point is not the peak caused by plosives. Then the terminal 102 can discard the continuous window frame peak point pair and recalculate the next window frame peak.
[0067] Step S208: Obtain the plosive microphone detection result of the to-be-detected audio frame based on the consistency between the peak point pairs of the consecutive window frames of interest and the to-be-detected voice signal points in the to-be-detected voice signal point set.
[0068] Among them, plosive microphone means that when singing or recording, the mouth is too close to the microphone, causing the microphone to be blown by the air in the mouth and making a puffing sound, which affects the sound effect. For example, pronunciations with / b / , / p / , / g / , / k / , / t / , / d / will cause a strong airflow to directly impact the diaphragm of the microphone, resulting in plosive microphone. Plosive microphone is generally caused by excessive low-frequency energy in the proximity effect. The proximity effect is that the sound pressure generated by a sound source at a certain point is inversely proportional to the distance from this point to the sound source. The closer the microphone is to the sound source, the greater the change in sound pressure, resulting in the abnormal increase in the low-frequency of the picked-up sound. However, if it is too close to the microphone, it will cause the microphone to generate excessive low-frequency harmonic resonance, resulting in distortion of the low-frequency. The peak point pairs of the consecutive window frames of interest can be the peak points of the window frames in two consecutive sliding windows. The to-be-detected voice signal point set can be a set composed of multiple to-be-detected voice signal points, which includes voice signal points with little amplitude change between adjacent voice signal points. For example, the voice signal points where the line formed by adjacent voice signal points is parallel to the abscissa. The terminal 102 can obtain the plosive microphone detection result of the to-be-detected audio frame based on the consistency between the peak point pairs of the consecutive window frames of interest and the to-be-detected voice signal points in the to-be-detected voice signal point set.
[0069] For example, the terminal 102 can determine the plosive microphone detection result by comparing whether the peak point pairs of the consecutive window frames of interest exist in the to-be-detected voice sample set. In one embodiment, obtaining the plosive microphone detection result of the to-be-detected audio frame based on the consistency between the peak point pairs of the consecutive window frames of interest and the to-be-detected voice signal points in the to-be-detected voice signal point set includes: if the peak point pairs of the consecutive window frames of interest all exist in the to-be-detected voice sample set, then determine that the plosive microphone detection result is that plosive microphone occurs. In this embodiment, the terminal 102 can detect whether the peak point pairs of the consecutive window frames of interest exist in the to-be-detected voice sample set. If the terminal 102 detects that the peak point pairs of the consecutive window frames of interest all exist in the to-be-detected voice sample set, then the terminal 102 can determine that the plosive microphone detection result is that plosive microphone occurs. That is, the terminal 102 can detect whether the peak point pairs of the consecutive window frames of interest coincide with the voice signal points in the to-be-detected voice sample set. Specifically, the voice signal points in the above to-be-detected voice signal point set can be denoted as t k , during one sliding of the preset sliding window, the terminal 102 can use the p k in the current sliding window and compare it with t k to determine whether p k coincides with t kWhether they coincide. If they coincide, the terminal 102 can slide the sliding window by O / 4 samples and repeat the calculation of t k , and p of the current sliding window k , and continue to determine p in the current sliding window k and t k Whether they coincide. If it is determined that they coincide twice in a row, the terminal 102 can determine that the current microphone blasting situation has occurred. That is, the terminal 102 can determine whether the points in the waveform parallel to the coordinate horizontal axis in the human voice audio signal coincide with the points where the frame peak occurs due to microphone blasting. If they coincide, it is determined that the user has a microphone blasting situation
[0070] In the above microphone blasting detection method, by determining the relative size of multiple average amplitude differences of the audio frame to be detected under different amplitude difference calculation periods, the relative threshold of the audio frame to be detected is determined, and the audio frame to be detected is divided into audio periods according to the relative threshold to obtain the fundamental frequency information of the human voice. The human voice audio signal is extracted according to the fundamental frequency information of the human voice, and the set of human voice signal points to be detected in the human voice audio signal is determined according to the amplitude change between adjacent human voice signal points in the human voice audio signal. The sliding window is traversed through the human voice audio signal to obtain the pair of peak points of interest of the continuous window frames obtained by traversing. According to the consistency of the pair of peak points of interest of the continuous window frames with the human voice signal points to be detected in the set of human voice signal points to be detected, the microphone blasting detection result of the audio frame to be detected is obtained. Compared with the traditional method of detecting whether the user has a microphone blasting situation according to the loudness threshold, this solution detects whether there is a microphone blasting in the audio frame by the consistency of multiple human voice signal points with small amplitude changes and the pair of peak points of interest of the continuous window frames, thereby improving the accuracy of detecting microphone blasting
[0071] In one embodiment, determining the relative threshold of the audio frame to be detected according to the relative size of multiple average amplitude differences includes: obtaining the maximum average amplitude difference among multiple average amplitude differences; using the product of the maximum average amplitude difference and the set coefficient as the relative threshold
[0072] In this embodiment, the average amplitude difference can be a value calculated by the terminal 102 under different amplitude difference calculation periods. According to the above formula for calculating the average amplitude difference, since i changes, there will be multiple average amplitude differences obtained. The terminal 102 can obtain the maximum average amplitude difference among multiple average amplitude differences, and the terminal 102 can use the product of the maximum average amplitude difference and the set coefficient as the relative threshold. Among them, the above audio frame to be detected can be an audio frame segmented from the complete audio to be detected. If the audio frame to be detected is the kth frame in the audio to be detected, the formula for calculating the relative threshold of the kth frame can be as follows: rT = 0.618 * max(d k ). Among them, d kThe maximum average amplitude difference, 0.618 is the set coefficient, and rT is the relative threshold.
[0073] Through this embodiment, the terminal 102 can obtain the relative threshold of the audio frame to be detected based on the maximum average amplitude difference among multiple average amplitude differences, so that the terminal 102 can perform pop detection on the audio frame to be detected based on the relative threshold, improving the accuracy of pop detection.
[0074] In one embodiment, dividing the audio period of the audio frame to be detected according to the relative threshold and obtaining the fundamental frequency information of the human voice in the audio frame to be detected according to the audio period division result includes: obtaining multiple target audio points with amplitudes less than the relative threshold in the audio frame to be detected, and obtaining the minimum period presented by the multiple target audio points; obtaining the reciprocal of the minimum period to obtain the fundamental frequency information of the human voice.
[0075] In this embodiment, the terminal 102 can divide the audio period of the audio frame to be detected based on the relative threshold. The terminal 102 can use the relative threshold to filter out other waveforms in the waveform of the audio frame to be detected that cannot be used to determine the period. Specifically, the terminal 102 can obtain multiple target audio points with amplitudes less than the relative threshold in the audio frame to be detected, and obtain the minimum period presented by the multiple target audio points. The terminal 102 can also obtain the reciprocal of the minimum period to obtain the above-mentioned fundamental frequency information of the human voice. Among them, the fundamental frequency information of the human voice can be denoted as f0. The fundamental frequency is generally distributed near the low frequency of 100 - 200 hz. Its time-domain period T0 = 1 / f0. Therefore, the terminal 102 can linearly transform the fundamental frequency extraction into the extraction of the period T0. That is, after filtering out the audio points greater than or equal to the relative threshold, the terminal 102 can determine the minimum period T0 from the remaining multiple target audio points as the current frame period, and obtain the fundamental frequency f0 through the inverse transformation formula such as f0 = 1 / T0.
[0076] Through this embodiment, the terminal 102 can divide the minimum period in the audio frame to be detected through the relative threshold and obtain the fundamental frequency information of the human voice based on the minimum period, so that the terminal 102 can perform pop detection based on the fundamental frequency information of the human voice, improving the accuracy of pop detection.
[0077] In one embodiment, extracting the human voice audio signal in the audio frame to be detected according to the fundamental frequency information of the human voice includes: performing low-pass filtering on the audio frame to be detected according to the fundamental frequency information of the human voice to obtain the human voice audio signal in the audio frame to be detected.
[0078] In this embodiment, after the terminal 102 obtains the fundamental frequency information of the human voice, it can extract the human voice audio signal from the audio frame to be detected. That is, the terminal 102 can perform low-pass filtering on the audio frame to be detected based on the fundamental frequency information of the human voice, and use the waveform less than the fundamental frequency information of the human voice as the human voice audio signal in the audio frame to be detected. Among them, the above-mentioned plosive microphone signal belongs to the clipping distortion caused by excessive low-frequency energy, and the clipping distortion generally occurs when the power amplifier works in the oversaturated state. In digital signal processing, clipping occurs when the audio signal is limited by the selected representation range, and this part of the original signal is lost and cannot be restored. Therefore, the terminal 102 can perform low-pass filtering through the above-mentioned fundamental frequency information f0 of the human voice to obtain the signal below f0, and its calculation formula can be as follows: y k (m)=filter lp (x k (m)), where y k (m) can be the human voice audio signal obtained after low-pass filtering, x k (m) can be the audio signal in the audio frame to be detected, and filter lp is a low-pass filtering function, where lp (Low-pass) indicates that this filtering function is a low-pass filter.
[0079] Through this embodiment, the terminal 102 can perform low-pass filtering on the audio frame to be detected based on the fundamental frequency information of the human voice to obtain the human voice audio signal, so that the terminal 102 can perform plosive microphone detection based on the human voice audio signal, improving the accuracy of plosive microphone detection.
[0080] In one embodiment, according to the amplitude change between adjacent human voice signal points in the human voice audio signal, determining the set of human voice signal points to be detected in the human voice audio signal includes: performing a difference operation on the human voice audio signal to obtain the amplitude change rate between adjacent human voice signal points in the audio frame to be detected; determining multiple human voice signal points corresponding to the amplitude change rate satisfying the preset amplitude change rate threshold as the set of human voice signal points to be detected in the human voice audio signal.
[0081] In this embodiment, the human voice audio signal can be the audio signal representing the human voice in the audio frame to be detected. The human voice audio signal includes multiple human voice signal points, and the terminal 102 can select multiple human voice signal points from them to form a set of human voice signal points to be detected. For example, the terminal 102 can perform a difference operation on the human voice audio signal obtained by the above low-pass filtering to obtain the amplitude change rate between adjacent human voice signal points in the audio frame to be detected. Among them, the above difference operation can be a first-order difference operation performed in the time domain, and its calculation formula is as follows: Dif k (i)=|y k (i)-y k (i - 1)|. Where Dif k(i) can be the rate of change of the amplitudes of two adjacent human voice signal points, that is, the slope, y k (i) is the i-th human voice audio signal point. The terminal 102 calculates multiple rates of change of amplitudes, and determines the multiple human voice signal points corresponding to which the above rates of change of amplitudes satisfy a preset rate-of-change-of-amplitude threshold as the to-be-detected human voice signal points in the human voice audio signal. Thus, the terminal 102 can form a to-be-detected human voice signal point set based on the multiple to-be-detected human voice signal points. Among them, the above to-be-detected audio frame can be one of the frames in the to-be-detected audio, for example, the k-th frame. The terminal 102 can find, based on multiple slopes Dif k (i), the sample points where the slopes of adjacent samples tend to 0. For example, by taking Dif k (i) less than a preset slope threshold for two adjacent human voice signal points as the to-be-detected human voice signal points, denoted as t k .
[0082] Through this embodiment, the terminal 102 can obtain a to-be-detected human voice signal point set where clipping distortion occurs based on differential operation and slope judgment. Thus, the terminal 102 can perform pop noise detection based on the to-be-detected human voice signal points, improving the accuracy of pop noise detection.
[0083] In one embodiment, as Figure 3 shown, an audio recording method is provided. Taking the example that this method is applied to the terminal in Figure 1 , the method includes the following steps:
[0084] Step S302, obtain the real-time input recorded audio.
[0085] Among them, the recorded audio can be the audio signal that the terminal 102 receives from the user in real time. For example, the user can input audio in real time through the audio recording software in the terminal 102, and then the terminal 102 can obtain the recorded audio input by the user in real time. Specifically, taking the terminal 102 as a mobile phone as an example, the user can sing through the mobile phone microphone or earphone receiver, and the mobile phone can obtain the recorded audio input by the user in real time.
[0086] Step S304, detect whether the recorded audio has pop noise according to the above pop noise detection method.
[0087] Among them, after the terminal 102 obtains the recorded audio input by the user, it can perform pop noise detection on the recorded audio. For example, the terminal 102 can detect whether there is pop noise in the recorded audio input by the user through the above pop noise detection method. Taking the terminal 102 as a mobile phone as an example, the mobile phone can detect whether there is pop noise in the singing audio input by the user.
[0088] Step S306, if so, display a pop noise reminder message on the recording interface of the recorded audio.
[0089] Among them, the terminal 102 can detect whether plosives occur in the real-time recorded audio input by the user. If the terminal 102 determines that plosives occur in the recorded audio, the terminal 102 can display a plosive reminder message on the recording interface of the recorded audio. For example, as Figure 4 shown Figure 4 is a schematic diagram of the plosive reminder interface in an embodiment. When the terminal 102 detects plosives, it can display the information that plosives occur on the audio recording interface and remind the user to stay away from the microphone, so as to achieve the effect of reducing the impact of plosives after detecting plosives.
[0090] In the above audio recording method, by using the microphone of the terminal 102 to provide real-time and accurate plosive detection during audio recording, the terminal 102 can more accurately detect the human voice part based on the fundamental frequency information of the human voice in the above plosive detection method, and use the relative threshold after differential operation for real-time tracking detection. Compared with the traditional method of using a pop filter, this solution can better ensure the permeability of audio pickup during the entire recording and singing process, and also get rid of the limitation that it is impossible to sing anytime and anywhere due to relying on such physical anti-plosive devices, simplifies the entire recording process, improves the frequency of users using audio recording software, and moreover, the terminal 102 corrects the user's incorrect audio recording operation in a timely manner by reminding of plosives, solves the plosive problem that mainly affects the recorded audio, and provides audio quality assurance for audio recording in terms of sound quality improvement and post-audio production.
[0091] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0092] Based on the same inventive concept, the embodiments of the present application also provide a plosive detection device for implementing the plosive detection method involved above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in one or more embodiments of the plosive detection device provided below can refer to the limitations on the plosive detection method in the above text, and will not be repeated here.
[0093] In one embodiment, as Figure 5As shown in the figure, a pop microphone detection device is provided, including: an acquisition module 500, a human voice detection module 502, a peak detection module 504, and a pop microphone detection module 506, where:
[0094] The acquisition module 500 is configured to acquire an audio frame to be detected and extract the human voice audio signal in the audio frame to be detected.
[0095] The human voice detection module 502 is configured to determine a set of human voice signal points to be detected in the human voice audio signal according to the amplitude change between adjacent human voice signal points in the human voice audio signal.
[0096] The peak detection module 504 is configured to traverse the human voice audio signal through a preset sliding window and obtain a pair of peak points of interest of consecutive window frames obtained by traversing the preset sliding window.
[0097] The pop microphone detection module 506 is configured to obtain the pop microphone detection result of the audio frame to be detected according to the consistency between the pair of peak points of interest of consecutive window frames and the human voice signal points to be detected in the set of human voice signal points to be detected.
[0098] In one embodiment, the above acquisition module 500 is specifically configured to determine the fundamental frequency information of the human voice in the audio frame to be detected; extract the human voice audio signal in the audio frame to be detected according to the fundamental frequency information of the human voice.
[0099] In one embodiment, the above acquisition module 500 is specifically configured to calculate multiple average amplitude differences corresponding to the audio frame to be detected under different amplitude difference calculation periods, and determine the relative threshold of the audio frame to be detected according to the relative magnitudes of the multiple average amplitude differences; divide the audio frame to be detected according to the relative threshold to obtain the fundamental frequency information of the human voice in the audio frame to be detected according to the audio period division result.
[0100] In one embodiment, the above acquisition module 500 is specifically configured to obtain the maximum average amplitude difference among the multiple average amplitude differences; use the product of the maximum average amplitude difference and a set coefficient as the relative threshold.
[0101] In one embodiment, the above acquisition module 500 is specifically configured to obtain multiple target audio points in the audio frame to be detected with amplitudes less than the relative threshold, and obtain the minimum period presented by the multiple target audio points; obtain the reciprocal of the minimum period to obtain the fundamental frequency information of the human voice.
[0102] In one embodiment, the above human voice detection module 504 is specifically configured to perform low-pass filtering on the audio frame to be detected according to the fundamental frequency information of the human voice to obtain the human voice audio signal in the audio frame to be detected.
[0103] In one embodiment, the above-mentioned voice detection module 504 is specifically configured to perform differential operation on the voice audio signal to obtain the amplitude change rate between adjacent voice signal points in the audio frame to be detected; and determine multiple voice signal points corresponding to the amplitude change rate satisfying the preset amplitude change rate threshold as the set of voice signal points to be detected in the voice audio signal.
[0104] In one embodiment, for each of the continuous window frame interested peak points, its window frame peak is greater than the preset window frame peak threshold.
[0105] In one embodiment, the above-mentioned pop microphone detection module 508 is specifically configured to determine that the pop microphone detection result is a pop microphone occurrence if the continuous window frame interested peak point pairs are all in the set of voice sample points to be detected.
[0106] In one embodiment, as Figure 6 shown, an audio recording device is provided, including: a recording module 600, an inspection module 602, and a reminder module 604, where:
[0107] The recording module 600 is configured to obtain the real-time input recorded audio.
[0108] The inspection module 602 is configured to detect whether a pop microphone occurs in the recorded audio according to the pop microphone detection method as described above.
[0109] The reminder module 604 is configured to, if so, display a pop microphone reminder message on the recording interface of the recorded audio.
[0110] Each module in the above-mentioned pop microphone detection device and audio recording device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0111] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a method for detecting microphone squelch. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0112] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0113] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the above-mentioned method for detecting microphone squelch.
[0114] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the above-mentioned method for detecting microphone squelch.
[0115] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the above-mentioned method for detecting microphone squelch.
[0116] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties.
[0117] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0118] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0119] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for detecting microphone popping, characterized in that, The method includes: Obtaining a to-be-detected audio frame and extracting the human voice audio signal in the to-be-detected audio frame, including: determining the fundamental frequency information of the human voice in the to-be-detected audio frame, based on the fundamental frequency information of the human voice, performing low-pass filtering on the to-be-detected audio frame, and taking the waveform smaller than the fundamental frequency information of the human voice as the human voice audio signal in the to-be-detected audio frame; Determining a set of to-be-detected human voice signal points in the human voice audio signal according to the amplitude change between adjacent human voice signal points in the human voice audio signal; Traversing the human voice audio signal through a preset sliding window and obtaining pairs of peak points of interest of consecutive window frames obtained by traversing the preset sliding window; Obtaining the pop microphone detection result of the to-be-detected audio frame according to the consistency between the pair of peak points of interest of the consecutive window frames and the to-be-detected human voice signal points in the set of to-be-detected human voice signal points.
2. The method according to claim 1, characterized in that, The determining the fundamental frequency information of the human voice in the to-be-detected audio frame includes: Calculating a plurality of average amplitude differences corresponding to the to-be-detected audio frame under different amplitude difference calculation periods, and determining the relative threshold of the to-be-detected audio frame according to the relative magnitudes of the plurality of average amplitude differences; Performing audio period division on the to-be-detected audio frame according to the relative threshold, and obtaining the fundamental frequency information of the human voice in the to-be-detected audio frame according to the audio period division result.
3. The method according to claim 2, wherein The determining the relative threshold of the to-be-detected audio frame according to the relative magnitudes of the plurality of average amplitude differences includes: Obtaining the maximum average amplitude difference among the plurality of average amplitude differences; Taking the product of the maximum average amplitude difference and a set coefficient as the relative threshold.
4. The method according to claim 2, wherein The performing audio period division on the to-be-detected audio frame according to the relative threshold and obtaining the fundamental frequency information of the human voice in the to-be-detected audio frame according to the audio period division result includes: Obtaining a plurality of target audio points in the to-be-detected audio frame whose amplitudes are smaller than the relative threshold, and obtaining the minimum period presented by the plurality of target audio points; Obtaining the reciprocal of the minimum period to obtain the fundamental frequency information of the human voice.
5. The method according to claim 1, wherein The determining the set of to-be-detected human voice signal points in the human voice audio signal according to the amplitude change between adjacent human voice signal points in the human voice audio signal includes: Performing a difference operation on the human voice audio signal to obtain the amplitude change rate between adjacent human voice signal points in the to-be-detected audio frame; Determining a plurality of human voice signal points corresponding to the amplitude change rate satisfying a preset amplitude change rate threshold as the set of to-be-detected human voice signal points in the human voice audio signal.
6. The method according to claim 1, characterized in that, Each window frame peak in the pair of peak points of interest of the consecutive window frames is greater than a preset window frame peak threshold.
7. The method according to any one of claims 1 to 6, characterized in that, The obtaining the pop microphone detection result of the to-be-detected audio frame according to the consistency between the pair of peak points of interest of the consecutive window frames and the to-be-detected human voice signal points in the set of to-be-detected human voice signal points includes: If the pair of peak points of interest of the consecutive window frames are all in the set of to-be-detected human voice signal points, determining that the pop microphone detection result is that pop microphone occurs.
8. An audio recording method, characterized in that, The method includes: Obtaining a real-time input recorded audio; Detecting whether pop microphone occurs in the recorded audio according to the pop microphone detection method according to any one of claims 1 to 7; If so, a pop - up microphone reminder message is displayed on the recording interface of the recorded audio.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Speech signal processing method and device
CN104409081A
Noise suppression method and device, medium and electronic equipment
CN113571078A