Voice Break Detection Method, Device, Equipment and Storage Medium

By dividing the audio signal into multiple time domain signal frames and determining the cutoff frequency value based on the upper spectrum limit value, the problem of degradation of sound break detection accuracy caused by fixed threshold value in the prior art is solved, and higher detection accuracy and versatility are achieved.

CN113393862BActive Publication Date: 2025-06-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011224319.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-05
Publication Date
2025-06-20
Estimated Expiration
2040-11-05

AI Technical Summary

Technical Problem

In the sound break detection, the prior art cannot be applied to situations where the time domain signal fluctuates greatly, resulting in a decrease in detection accuracy.

Method used

By dividing the audio signal into N time domain signal frames, and obtaining the upper spectrum limit value of each frame, the cutoff frequency value is determined, thereby distinguishing between normal audio signals and sound-breaking audio signals.

Benefits of technology

It improves the accuracy of sound break detection, and dynamically adjusts the cutoff frequency value, which is suitable for multiple time domain signal frames, enhancing the universality of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113393862B_ABST
    Figure CN113393862B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method, device, equipment, and storage medium for detecting broken sound, which relates to the technical field of audio detection. The method includes: dividing the audio signal to be detected into N time-domain signal frames; obtaining the spectral upper limit values respectively corresponding to the N time-domain signal frames, where the spectral upper limit value refers to the frequency value corresponding to the maximum power value in the frequency-domain power spectrum of the time-domain signal frame; determining a cut-off frequency value based on the spectral upper limit values respectively corresponding to the N time-domain signal frames; wherein the cut-off frequency value is used to distinguish between a normal audio signal and a broken sound audio signal from the perspective of signal power; for the i-th time-domain signal frame among the N time-domain signal frames, if the magnitude relationship between the spectral upper limit value corresponding to the i-th time-domain signal frame and the cut-off frequency value satisfies a condition, it is determined that the i-th time-domain signal frame belongs to a broken sound signal frame. The technical solution provided by the embodiment of the present application can improve the accuracy of broken sound detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of audio detection, and particularly to a method, device, equipment and storage medium for detecting audio breakage. Background Art

[0002] When processing an audio signal, there may sometimes be a broken sound part in the audio signal, so it is necessary to detect the broken sound part of the audio.

[0003] In the related art, the time-domain signal of the audio signal is segmented into multiple segments, and each segment is sampled and statistically analyzed respectively to obtain a statistical chart of each segment. If the statistical chart corresponding to a certain segment does not conform to the normal distribution, there is an abnormal peak, and the peak value of the abnormal peak is greater than the threshold value, then it is determined that the segment is the broken sound part of the audio signal.

[0004] In the above related art, when the fluctuation of the time-domain signal is large, if a fixed threshold value is used to detect the broken sound of all segments, it will cause the threshold value not to be applicable to all segments, thus reducing the accuracy of broken sound detection. Summary of the Invention

[0005] Embodiments of the present application provide a method, device, equipment and storage medium for detecting audio breakage, which can improve the accuracy of broken sound detection. The technical solution is as follows:

[0006] According to one aspect of the embodiments of the present application, a method for detecting audio breakage is provided. The method includes:

[0007] Dividing the audio signal to be detected into N time-domain signal frames, where N is a positive integer;

[0008] Obtaining the spectral upper limit values respectively corresponding to the N time-domain signal frames, where the spectral upper limit value refers to the frequency value corresponding to the maximum power value in the frequency-domain power spectrum of the time-domain signal frame;

[0009] Based on the spectral upper limit values respectively corresponding to the N time-domain signal frames, determining a cut-off frequency value; wherein, the cut-off frequency value is used to distinguish a normal audio signal and a broken sound audio signal from the perspective of signal power;

[0010] For the i-th time-domain signal frame among the N time-domain signal frames, if the magnitude relationship between the spectral upper limit value corresponding to the i-th time-domain signal frame and the cut-off frequency value satisfies a condition, it is determined that the i-th time-domain signal frame belongs to a broken sound signal frame, where i is a positive integer less than or equal to N.

[0011] According to one aspect of the embodiments of the present application, a device for detecting audio breakage is provided. The device includes:

[0012] A time domain division module, configured to divide an audio signal to be detected into N time domain signal frames, where N is a positive integer;

[0013] An upper limit value acquisition module, configured to acquire the upper limit values of the spectra corresponding to the N time domain signal frames respectively, where the upper limit value of the spectrum refers to the frequency value corresponding to the maximum power value in the frequency domain power spectrum of the time domain signal frame;

[0014] A frequency determination module, configured to determine a cut-off frequency value based on the upper limit values of the spectra corresponding to the N time domain signal frames respectively; wherein, the cut-off frequency value is used to distinguish a normal audio signal and a crackling audio signal from the perspective of signal power;

[0015] A crackling determination module, configured to, for the i-th time domain signal frame among the N time domain signal frames, if the magnitude relationship between the upper limit value of the spectrum corresponding to the i-th time domain signal frame and the cut-off frequency value satisfies a condition, determine that the i-th time domain signal frame belongs to a crackling signal frame, where i is a positive integer less than or equal to N.

[0016] According to one aspect of the embodiments of the present application, there is provided a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above crackling detection method.

[0017] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above crackling detection method.

[0018] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above crackling detection method.

[0019] The technical solution provided by the embodiments of the present application may include the following beneficial effects:

[0020] By dividing an audio signal into N time-domain signal frames, obtaining a determined cut-off frequency based on the spectral upper limit values respectively corresponding to the N time-domain signal frames, and determining a distorted sound signal frame based on the numerical relationship between each spectral upper limit value and the cut-off frequency, this cut-off frequency can be used for distorted sound detection of multiple time-domain signal frames, with strong versatility, thereby improving the accuracy of distorted sound detection.

[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 is a schematic diagram of a distorted sound signal provided by an embodiment of this application;

[0024] Figure 2 is a flowchart of a distorted sound detection method provided by an embodiment of this application;

[0025] Figure 3 is a statistical histogram of spectral upper limit values provided by an embodiment of this application;

[0026] Figure 4 is a flowchart of a distorted sound detection method provided by another embodiment of this application;

[0027] Figure 5 is a block diagram of a distorted sound detection device provided by an embodiment of this application;

[0028] Figure 6 is a block diagram of a distorted sound detection device provided by another embodiment of this application;

[0029] Figure 7 is a block diagram of a computer device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description involves the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this application. On the contrary, they are only examples of methods consistent with some aspects of this application as detailed in the appended claims.

[0031] Clipping is a problem where the input audio signal exceeds the maximum range of the digital representation of the current audio signal, resulting in a significant impairment of the sound quality. As Figure 1 shown in the spectrogram, the clipped signal 11 is a signal that clearly exceeds the normal frequency spectrum range. The cause of clipping can be that during the process of microphone audio acquisition, the volume of the audio signal exceeds the upper limit of the ADC (Analogue-to-Digital Conversion) of the recording device, causing the output digital signal value to reach the upper limit of the ADC. Another cause of clipping can be that during the audio signal processing, the audio signal is amplified, and the amplification result exceeds the maximum digital representation of the sound bit depth. For example, for a 16-bit audio digital representation, the maximum value is +32767 / -32768, which locks the final output value at the maximum value (i.e., +32767 or -32768). Subjectively, clipping is manifested as a hoarse and noisy sound.

[0032] Clipping detection is to detect whether there is a clipping phenomenon in the audio signal and the specific location of the clipping, so as to adjust the recording parameters or audio signal processing parameters in a timely manner to avoid clipping. In addition, it is also used for the preliminary judgment of the clipping repair algorithm. The clipping repair algorithm is to repair the detected clipping location based on the clipping detection result, reducing the obvious distortion problem brought by clipping to the audio signal.

[0033] Optionally, the clipping detection method is used in the audio processing of the target application program. The target application program can be any application program with audio processing functions, such as social application programs, music application programs, video application programs, news application programs, game application programs, etc. The embodiments of the present application do not limit this. If the target application program is a social application program, by using the clipping detection method provided by the embodiments of the present application, when there is a voice or video call between different accounts of the target application program, it can detect in real time whether there is clipping in the voice during the call and determine the location of the clipped part, so as to determine the cause of the clipping. If it is determined that the cause of the clipping is the device setting, such as too high volume, the user will be prompted to turn down the volume or automatically reduce the call volume; if it is determined that the cause of the clipping is network transmission, the user will be prompted to switch the network or move to a place with better network signal. If the target application program is a music application program, by using the clipping detection method provided by the embodiments of the present application, it can detect in real time whether there is clipping in the played music and determine the location of the clipped part, so as to determine the cause of the clipping. If the cause of the clipping is determined, if it is determined that the cause of the clipping is a poor sound source, the user will be prompted to replace the sound source or automatically switch to a sound source with better sound quality.

[0034] In the method provided by the embodiments of the present application, the execution subject of each step may be a computer device, which refers to an electronic device with data calculation, processing, and storage capabilities. The computer device may be a terminal such as a PC (Personal Computer), a tablet computer, a smart phone, a wearable device, a smart robot, etc.; or it may be a server. Among them, the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Next, the technical solution of the present application will be introduced and illustrated through several embodiments.

[0035] Next, the technical solution of the present application will be introduced and illustrated through several embodiments.

[0036] Please refer to Figure 2 , which shows a flowchart of a method for detecting broken sound provided by an embodiment of the present application. In this embodiment, it is exemplified that the method is applied to the computer device introduced above. The method may include the following steps (201-204):

[0037] Step 201, divide the audio signal to be detected into N time-domain signal frames, where N is a positive integer.

[0038] When the audio signal is a time-domain signal, divide the audio signal to be detected according to a fixed duration to obtain N time-domain signal frames. Optionally, the specific duration of the fixed duration is 5 milliseconds, 10 milliseconds, 15 milliseconds, 20 milliseconds, etc. The specific duration of the fixed duration is set by relevant technical personnel according to the actual situation, and the embodiments of the present application do not limit this. In some embodiments, each time-domain signal frame is subjected to windowing processing to reduce spectral energy leakage. Among them, the window function used for windowing processing may be a Hanning window, a Hamming window, etc. Among them, the function expression of the Hanning window is:

[0039]

[0040] where N is the number of time-domain signal frames, n is the independent variable of the function, n belongs to [0, N-1], and win(n) is the dependent variable of the function. The Hanning window is a special case of the raised cosine window. The Hanning window can be regarded as the sum of the spectra of 3 rectangular time windows, that is, the sum of 3 sinc(t) type functions, and the two terms in the parentheses are shifted to the left and right by π / T (T represents the period) relative to the first spectral window, so that the side lobes cancel each other out and reduce spectral energy leakage.

[0041] Step 202, obtain the spectral upper limit values corresponding to the N time-domain signal frames respectively.

[0042] The upper limit value of the spectrum may refer to the frequency value corresponding to the maximum power value in the frequency-domain power spectrum of a time-domain signal frame. The frequency-domain power spectrum is used to indicate the power values corresponding to each frequency in the time-domain signal frame.

[0043] In some embodiments, the above step 202 further includes the following sub-steps:

[0044] 1. For the i-th time-domain signal frame, perform a time-domain to frequency-domain conversion process on the i-th time-domain signal frame to obtain the i-th frame spectrum;

[0045] 2. Determine the power corresponding to each frequency value in the i-th frame spectrum to obtain the frequency-domain power spectrum corresponding to the i-th time-domain signal frame;

[0046] 3. Determine the frequency value corresponding to the maximum power in the i-th frequency-domain power spectrum as the upper limit value of the spectrum corresponding to the i-th time-domain signal frame.

[0047] In some embodiments, for the i-th time-domain signal frame among N time-domain signal frames, calculate the power values corresponding to each frequency corresponding to the i-th time-domain signal frame, so as to obtain the i-th frequency-domain power spectrum; determine the frequency corresponding to the maximum power value in the i-th frequency-domain power spectrum as the upper limit value of the spectrum corresponding to the i-th time-domain signal frame, where i is a positive integer less than or equal to N. In this implementation manner, directly determine the frequency value corresponding to the maximum power in the i-th frequency-domain power spectrum as the upper limit value of the spectrum corresponding to the i-th time-domain signal frame, which is simple and convenient, reduces the execution steps of the solution, and thus saves processing resources.

[0048] In some embodiments, the calculation formula for the power value corresponding to each frequency refers to the following formula 1:

[0049] Formula 1:

[0050] Where p(i, k) represents the power value of the k-th frequency in the i-th frequency-domain power spectrum, N is the number of frequency-domain power spectra, and x(n) represents the sample value of the i-th time-domain signal frame. Formula 1 means that the power value of the k-th frequency in the i-th frequency-domain power spectrum is the sum of the powers corresponding to the k-th frequency of each sample point of the i-th time-domain signal frame.

[0051] In some embodiments, the frequencies involved in the embodiments of the present application are frequency points, and the frequency points are used to represent the corresponding frequency ranges. For example, frequency point 1, frequency point 2, and frequency point 3 respectively correspond to the frequency ranges of 0 to 200 Hz, 200 to 400 Hz, and 400 to 600 Hz.

[0052] Step 203, determine the cut-off frequency value based on the upper limit values of the spectra corresponding to N time-domain signal frames respectively.

[0053] Among them, the cut-off frequency value is used to distinguish normal audio signals and distorted audio signals from the perspective of signal power. By obtaining the upper frequency limit values corresponding to N time-domain signal frames respectively, that is, obtaining N upper frequency limit values, the N upper frequency limit values can be the same or different from each other. By statistically analyzing the N upper frequency limit values, the cut-off frequency can be determined.

[0054] In some embodiments, step 203 further includes the following sub-steps:

[0055] 1. Count the number of each different upper frequency limit value;

[0056] 2. Determine the upper frequency limit value with the largest number as the cut-off frequency value.

[0057] Accumulatively count the same upper frequency limit values to obtain the number of each different upper frequency limit value; determine the upper frequency limit value with the largest number as the cut-off frequency value.

[0058] Optionally, counting the number of each different upper frequency limit value includes: using the histogram statistical method to count the number of each different upper frequency limit value.

[0059] In some embodiments, when the proportion of the number of upper frequency limit values in the total statistical number is greater than a%, the upper frequency limit value is determined as the cut-off frequency, where a is a positive number less than or equal to 100. Among them, a can be 50, 60, 70, 75, etc. The specific value of a is set by relevant technical personnel according to the actual situation, and the embodiments of the present application do not limit this. Optionally, as Figure 3 shown in the histogram 30, the number of the upper frequency limit value 31 is 2, the number of the upper frequency limit value 32 is 425, and the number of the upper frequency limit value 33 is 12; assuming that the value of a is 70, then 425 / (2 + 425 + 23)>70%, that is, the upper frequency limit value 32 is determined as the cut-off frequency value.

[0060] Optionally, counting the number of each different upper frequency limit value further includes: using the table statistical method to count the number of each different upper frequency limit value and directly mark it in the table, so that the number of each different upper frequency limit value can be obtained intuitively and conveniently, and it is convenient to calculate the proportion of the number of each upper frequency limit value in the total number.

[0061] In other embodiments, counting the number of each different upper frequency limit value can also be using the line chart statistical method, pie chart statistical method, etc., and the embodiments of the present application do not limit this.

[0062] In this implementation, by limiting that the number of cut-off frequency values must reach a% of the total number of upper limit values of the spectrum, the problem of poor generality of the cut-off frequency values caused by the overly scattered statistical results of the histogram is avoided, thereby improving the accuracy of crosstalk detection.

[0063] Step 204: For the i-th time-domain signal frame among N time-domain signal frames, if the magnitude relationship between the upper limit value of the spectrum corresponding to the i-th time-domain signal frame and the cut-off frequency value satisfies the condition, it is determined that the i-th time-domain signal frame belongs to a crosstalk signal frame, where i is a positive integer less than or equal to N.

[0064] Compare the upper limit value of the spectrum corresponding to the i-th time-domain signal frame with the cut-off frequency value. If the upper limit value of the spectrum corresponding to the i-th time-domain signal frame is greater than the cut-off frequency value, it can be determined that the i-th time-domain signal frame belongs to a crosstalk signal frame.

[0065] In some embodiments, the above step 204 further includes the following sub-steps: If the upper limit value of the spectrum corresponding to the i-th time-domain signal frame is greater than a preset threshold, it is determined that the i-th time-domain signal frame belongs to a crosstalk signal frame; where the preset threshold is determined based on the cut-off frequency value and is a value greater than the cut-off frequency value. Optionally, in some embodiments, the ratio between the preset threshold and the cut-off frequency value is greater than 1 and less than 2, such as 1.6, 1.7, etc.; in other embodiments, the ratio between the preset threshold and the cut-off frequency value is greater than 1 and less than 1.5, such as 1.1, 1.2, etc., so as to ensure that the cut-off frequency is close to the frequency upper limit of the normal audio signal, thereby reducing the probability of missed detection of crosstalk signal frames. In the embodiments of the present application, the specific value of the ratio between the preset threshold and the cut-off frequency value is set by those skilled in the art according to the actual situation, and the embodiments of the present application do not make specific limitations thereon. In other embodiments, the ratio between the preset threshold and the cut-off frequency value is 1, that is, the preset threshold is equal to the cut-off frequency value, so as to directly compare the cut-off frequency value with the upper limit value of the spectrum corresponding to each time-domain signal frame, without calculating the preset threshold, reducing the calculation amount, and thus saving processing resources.

[0066] In some embodiments, taking the ratio between the preset threshold and the cut-off frequency value needing to be greater than 1.2 as an example, the judgment code corresponding to this step 204 is:

[0067]

[0068]

[0069] Among them, y(i) represents the upper limit value of the spectrum of the i-th time-domain signal frame, T represents the cut-off frequency, flag(i)=1 indicates that the i-th time-domain signal frame is a crosstalk signal frame, and flag(i)=0 indicates that the i-th time-domain signal frame is a normal signal frame.

[0070] In summary, in the technical solution provided by the embodiments of the present application, by dividing the audio signal into N time-domain signal frames, a determined cut-off frequency is obtained based on the upper spectral limit values respectively corresponding to the N time-domain signal frames, and the break signal frames are determined based on the numerical relationship between each upper spectral limit value and the cut-off frequency. This cut-off frequency can be used for break detection of multiple time-domain signal frames, with strong versatility, thereby improving the accuracy of break detection.

[0071] In some embodiments, the above step 202 further includes the following sub-steps:

[0072] 1. For the i-th time-domain signal frame, perform a time-domain to frequency-domain conversion process on the i-th time-domain signal frame to obtain the i-th frame spectrum.

[0073] Among them, the i-th time-domain signal frame is transformed into the i-th frame spectrum through Fourier transform, and the i-th frame spectrum is the frequency-domain signal corresponding to the i-th time-domain signal frame.

[0074] 2. Determine the power corresponding to each frequency value in the i-th frame spectrum to obtain the frequency-domain power spectrum corresponding to the i-th time-domain signal frame.

[0075] The calculation of the power corresponding to each frequency value can refer to the above formula (1), which will not be elaborated here.

[0076] 3. Determine the frequency value corresponding to the maximum power greater than the power threshold value in the i-th frequency-domain power spectrum as the upper spectral limit value corresponding to the i-th time-domain signal frame.

[0077] Optionally, determine the maximum power in the i-th frequency-domain power spectrum, and judge whether the maximum power is greater than the power threshold value. If the maximum power is greater than the power threshold value, then determine the frequency corresponding to the maximum power as the upper spectral limit value corresponding to the i-th time-domain signal frame; if the maximum power is less than or equal to the power threshold value, it is considered that the i-th time-domain signal frame has no upper spectral limit value.

[0078] In some embodiments, the judgment code corresponding to this step 202 can be:

[0079]

[0080]

[0081] Among them, p(i, k) is the maximum power in the i-th frequency-domain power spectrum, thres_p is the power threshold value, y(i) represents the upper spectral limit value corresponding to the i-th time-domain signal frame, and k is the frequency value corresponding to the maximum power in the i-th frequency-domain power spectrum.

[0082] In this implementation, if the maximum power in the i-th frequency-domain power spectrum is still less than or equal to the power threshold value, it indicates that the powers in the i-th frequency-domain power spectrum are relatively small. Then, it is considered that the i-th time-domain signal frame cannot be a cracked voice signal frame, and the upper frequency limit of the i-th frequency-domain power spectrum is not statistically calculated. Thus, before statistically calculating the upper frequency limit values corresponding to each time-domain signal frame, the time-domain signal frames are screened, reducing the amount of interfering data in the statistical data, improving the generality of the cut-off frequency, and further improving the accuracy of cracked voice detection.

[0083] Next, Figure 4 The cracked voice detection method provided in the embodiments of the present application is summarized as follows (Steps 401 to 404):

[0084] Step 401: Perform frame windowing processing and time-frequency transformation on the time-domain signal to obtain multiple frame spectra.

[0085] The audio signal to be detected is divided according to a fixed duration to obtain multiple time-domain signal frames; then, the multiple time-domain signal frames are subjected to Fourier transform to obtain multiple frame spectra.

[0086] Step 402: Perform upper frequency limit detection on multiple frame spectra respectively to obtain multiple upper frequency limit values.

[0087] Among them, the upper frequency limit value may refer to the frequency value corresponding to the maximum power value in the frequency-domain power spectrum of the time-domain signal frame.

[0088] Step 403: Perform histogram statistics on multiple upper frequency limit values to determine the cut-off frequency value.

[0089] Among them, the cut-off frequency value is used to distinguish normal audio signals and cracked voice audio signals from the perspective of signal power. Statistically analyze multiple upper frequency limit values to determine the cut-off frequency.

[0090] Step 404: Determine the cracked voice signal frame based on multiple upper frequency limit values and the cut-off frequency value.

[0091] Compare the upper frequency limit value corresponding to the i-th time-domain signal frame with the cut-off frequency value. If the upper frequency limit value corresponding to the i-th time-domain signal frame is greater than the cut-off frequency value, it can be determined that the i-th time-domain signal frame belongs to the cracked voice signal frame.

[0092] The following is an embodiment of the device of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the present application.

[0093] Please refer to Figure 5, which shows a block diagram of a crackling detection device provided by an embodiment of the present application. The device has the function of implementing the method example of the above crackling detection, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the computer device introduced above or can be set on the computer device. The device 500 may include: a time domain division module 510, an upper limit value acquisition module 520, a frequency determination module 530, and a crackling determination module 540.

[0094] The time domain division module 510 is configured to divide the audio signal to be detected into N time domain signal frames, where N is a positive integer.

[0095] The upper limit value acquisition module 520 is configured to acquire the spectral upper limit values corresponding to the N time domain signal frames respectively, where the spectral upper limit value refers to the frequency value corresponding to the maximum power value in the frequency domain power spectrum of the time domain signal frame.

[0096] The frequency determination module 530 is configured to determine a cut-off frequency value based on the spectral upper limit values corresponding to the N time domain signal frames respectively; wherein, the cut-off frequency value is used to distinguish a normal audio signal and a crackling audio signal from the perspective of signal power.

[0097] For the i-th time domain signal frame among the N time domain signal frames, the crackling determination module 540 is configured to determine that the i-th time domain signal frame belongs to a crackling signal frame if the magnitude relationship between the spectral upper limit value corresponding to the i-th time domain signal frame and the cut-off frequency value satisfies a condition, where i is a positive integer less than or equal to N.

[0098] In summary, in the technical solution provided by the embodiment of the present application, by dividing the audio signal into N time domain signal frames, obtaining a determined cut-off frequency based on the spectral upper limit values corresponding to the N time domain signal frames respectively, and determining the crackling signal frames based on the numerical relationship between each spectral upper limit value and the cut-off frequency, the cut-off frequency can be used to perform crackling detection on multiple time domain signal frames, with strong versatility, thereby improving the accuracy of crackling detection.

[0099] In some embodiments, as Figure 6 shown, the frequency determination module 530 includes: a quantity statistics sub-module 531 and a frequency determination sub-module 532.

[0100] The quantity statistics sub-module 531 is configured to count the quantities of each different spectral upper limit value.

[0101] The frequency determination sub-module 532 is configured to determine the spectral upper limit value with the largest quantity as the cut-off frequency value.

[0102] In some embodiments, as Figure 6As shown, the quantity statistics sub-module 531 is used to: adopt the histogram statistics method to count the quantities of the respective different spectral upper limit values.

[0103] In some embodiments, the crackling determination module 540 is used to:

[0104] If the spectral upper limit value corresponding to the i-th time-domain signal frame is greater than a preset threshold, it is determined that the i-th time-domain signal frame belongs to the crackling signal frame;

[0105] Wherein, the preset threshold is a value greater than the cut-off frequency value and is determined based on the cut-off frequency value.

[0106] In some embodiments, the ratio between the preset threshold and the cut-off frequency value is greater than 1 and less than 2.

[0107] In some embodiments, the upper limit value acquisition module 520 is used to:

[0108] For the i-th time-domain signal frame, perform a time-domain to frequency-domain conversion process on the i-th time-domain signal frame to obtain the i-th frame spectrum;

[0109] Determine the power corresponding to each frequency value in the i-th frame spectrum to obtain the frequency-domain power spectrum corresponding to the i-th time-domain signal frame;

[0110] Determine the frequency value corresponding to the maximum power greater than the power threshold value in the i-th frequency-domain power spectrum as the spectral upper limit value corresponding to the i-th time-domain signal frame.

[0111] In some embodiments, the upper limit value acquisition module 520 is used to:

[0112] For the i-th time-domain signal frame, perform a time-domain to frequency-domain conversion process on the i-th time-domain signal frame to obtain the i-th frame spectrum;

[0113] Determine the power corresponding to each frequency value in the i-th frame spectrum to obtain the frequency-domain power spectrum corresponding to the i-th time-domain signal frame;

[0114] Determine the frequency value corresponding to the maximum power in the i-th frequency-domain power spectrum as the spectral upper limit value corresponding to the i-th time-domain signal frame.

[0115] It should be noted that for the device provided in the above embodiments, when implementing its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.

[0116] Please refer to Figure 7 , which shows a structural block diagram of a computer device provided in an embodiment of the present application. This computer device is used to implement the popping sound detection method provided in the above embodiments. Specifically:

[0117] The computer device 700 includes a CPU (Central Processing Unit) 701, a system memory 704 including a RAM (Random Access Memory) 702 and a ROM (Read-Only Memory) 703, and a system bus 705 connecting the system memory 704 and the central processing unit 701. The computer device 700 also includes a basic I / O (Input / Output) system 706 for facilitating the transfer of information between various components within the computer, and a mass storage device 707 for storing an operating system 713, application programs 714, and other program modules 715.

[0118] The basic input / output system 706 includes a display 708 for displaying information and input devices 709 such as a mouse and a keyboard for user input. The display 708 and the input devices 709 are both connected to the central processing unit 701 through an input / output controller 710 connected to the system bus 705. The basic input / output system 706 may also include an input / output controller 710 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 710 also provides output to a display screen, a printer, or other types of output devices.

[0119] The mass storage device 707 is connected to the central processing unit 701 through a mass storage controller (not shown) connected to the system bus 705. The mass storage device 707 and its associated computer-readable medium provide non-volatile storage for the computer device 700. That is to say, the mass storage device 707 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.

[0120] Without loss of generality, the computer-readable medium may include a computer storage medium and a communication medium. The computer storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer storage medium includes RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc), or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage medium is not limited to the above several types. The above system memory 704 and the mass storage device 707 can be collectively referred to as memory.

[0121] According to various embodiments of the present application, the computer device 700 may also run by connecting to a remote computer on the network through a network such as the Internet. That is, the computer device 700 may be connected to the network 712 through the network interface unit 711 connected to the system bus 705, or in other words, the network interface unit 711 may also be used to connect to other types of networks or remote computer systems (not shown).

[0122] In an exemplary embodiment, a computer-readable storage medium is also provided, in which at least one instruction, at least one segment of program, a code set, or an instruction set is stored, and when the at least one instruction, the at least one segment of program, the code set, or the instruction set is executed by a processor, the above-mentioned popping sound detection method is implemented.

[0123] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0124] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned popping sound detection method.

[0125] It should be understood that the "plurality" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0126] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for detecting voice break, characterized in that, The method includes: Dividing the audio signal to be detected into N time-domain signal frames, where N is a positive integer; Obtaining the spectral upper limit values respectively corresponding to the N time-domain signal frames, where the spectral upper limit value refers to the frequency value corresponding to the maximum power value in the frequency-domain power spectrum of the time-domain signal frame; Counting the number of each different spectral upper limit value; Determining the spectral upper limit value with the largest number as the cut-off frequency value; wherein, the cut-off frequency value is used to distinguish between normal audio signals and distorted audio signals from the perspective of signal power; For the i-th time-domain signal frame among the N time-domain signal frames, if the magnitude relationship between the spectral upper limit value corresponding to the i-th time-domain signal frame and the cut-off frequency value satisfies the condition, it is determined that the i-th time-domain signal frame belongs to the distorted signal frame, where i is a positive integer less than or equal to N.

2. The method according to claim 1, characterized in that, The counting the number of each different spectral upper limit value includes: Using the histogram statistics method to count the number of each different spectral upper limit value.

3. The method according to claim 1, characterized in that, The if the magnitude relationship between the spectral upper limit value corresponding to the i-th time-domain signal frame and the cut-off frequency value satisfies the condition, it is determined that the i-th time-domain signal frame belongs to the distorted signal frame includes: If the spectral upper limit value corresponding to the i-th time-domain signal frame is greater than a preset threshold, it is determined that the i-th time-domain signal frame belongs to the distorted signal frame; Wherein, the preset threshold is determined based on the cut-off frequency value and is a value greater than the cut-off frequency value.

4. The method according to claim 3, characterized in that, The ratio between the preset threshold and the cut-off frequency value is greater than 1 and less than 2.

5. The method according to any one of claims 1 to 4, characterized in that, The obtaining the spectral upper limit values respectively corresponding to the N time-domain signal frames includes: For the i-th time-domain signal frame, performing a time-domain to frequency-domain conversion process on the i-th time-domain signal frame to obtain the i-th frame spectrum; Determining the power corresponding to each frequency value in the i-th frame spectrum to obtain the frequency-domain power spectrum corresponding to the i-th time-domain signal frame; Determining the frequency value corresponding to the maximum power in the i-th frequency-domain power spectrum as the spectral upper limit value corresponding to the i-th time-domain signal frame.

6. The method according to any one of claims 1 to 4, characterized in that, The obtaining the spectral upper limit values respectively corresponding to the N time-domain signal frames includes: For the i-th time-domain signal frame, performing a time-domain to frequency-domain conversion process on the i-th time-domain signal frame to obtain the i-th frame spectrum; Determining the power corresponding to each frequency value in the i-th frame spectrum to obtain the frequency-domain power spectrum corresponding to the i-th time-domain signal frame; Determining the frequency value corresponding to the maximum power greater than the power threshold value in the i-th frequency-domain power spectrum as the spectral upper limit value corresponding to the i-th time-domain signal frame.

7. A device for detecting voice break, characterized in that, The device includes: A time-domain division module, configured to divide the audio signal to be detected into N time-domain signal frames, where N is a positive integer; An upper limit value obtaining module, configured to obtain the spectral upper limit values respectively corresponding to the N time-domain signal frames, where the spectral upper limit value refers to the frequency value corresponding to the maximum power value in the frequency-domain power spectrum of the time-domain signal frame; A frequency determination module, configured to count the number of each different spectral upper limit value; determine the spectral upper limit value with the largest number as the cut-off frequency value; wherein, the cut-off frequency value is used to distinguish a normal audio signal and a clipped audio signal from the perspective of signal power; A clipped sound determination module, for the i-th time-domain signal frame among the N time-domain signal frames, if the magnitude relationship between the spectral upper limit value corresponding to the i-th time-domain signal frame and the cut-off frequency value satisfies a condition, determine that the i-th time-domain signal frame belongs to a clipped sound signal frame, where i is a positive integer less than or equal to N.

8. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the clipped sound detection method according to any one of claims 1 to 6 above.

9. A computer-readable storage medium, characterized in that, At least one program is stored in the computer-readable storage medium, and the at least one program is loaded and executed by a processor to implement the clipped sound detection method according to any one of claims 1 to 6 above.

10. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the clipped sound detection method according to any one of claims 1 to 6 above.

Citation Information

Patent Citations

  • Signal detecting method and device

    CN106847307A

  • Audio data tone quality detection method and device

    CN111312290A