Whisper suppression methods, devices, electronic devices and readable storage media
Patent Information
- Application Number
- CN202311154363.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-07
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2043-09-07
AI Technical Summary
[0005]本申请实施例的目的是提供一种啸叫抑制方法、装置、电子设备及可读存储介质,能够解决抑制啸叫的效果较差的问题
[0012] In this embodiment, the howling signal within the frequency range of the first audio frame can be determined based on the cepstral envelope of the first audio frame. The howling signal includes howling frequency points or howling frequency bands. The howling signal is then suppressed based on the cepstral envelope. This scheme, because howling frequency points or howling frequency bands in the audio frame can be determined and suppressed based on the cepstral envelope, allows for howling detection and suppression based on the characteristics of the audio signal itself. Therefore, it is not limited by traditional signal models and can avoid suppressing spectral structures similar to howling, thereby improving the howling suppression effect.
Smart Images

Figure CN117351977B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of audio technology, specifically relating to a method, apparatus, electronic device, and readable storage medium for suppressing howling. Background Technology
[0002] In Public Address Systems (PAS) and Hands-free Communication Systems (HFCS), howling can significantly impact the quality of voice calls.
[0003] Currently, feedback can be detected and suppressed based on signal models using adaptive feedback control (AFC) or by using notch filters (NH), thereby improving voice call quality.
[0004] However, the above methods are all limited by traditional signal models, which will detect similar spectral structures as howling and suppress them, resulting in poor howling suppression. Summary of the Invention
[0005] The purpose of this application is to provide a method, apparatus, electronic device, and readable storage medium for suppressing howling, which can solve the problem of poor howling suppression effect.
[0006] In a first aspect, embodiments of this application provide a howling suppression method, the method comprising: determining howling signals within the frequency range of a first audio frame based on the cepstral envelope of a first audio frame, wherein the howling signals include howling frequency points or howling frequency bands; and suppressing howling signals based on the cepstral envelope.
[0007] Secondly, embodiments of this application provide a howling suppression device, which includes a determining module and a suppressing module; the determining module is used to determine howling signals within the frequency range of the first audio frame based on the cepstral envelope of the first audio frame, wherein the howling signals include howling frequency points or howling frequency bands; the suppressing module is used to suppress the howling signals based on the cepstral envelope.
[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0012] In this embodiment, the howling signal within the frequency range of the first audio frame can be determined based on the cepstral envelope of the first audio frame. The howling signal includes howling frequency points or howling frequency bands. The howling signal is then suppressed based on the cepstral envelope. This scheme, because howling frequency points or howling frequency bands in the audio frame can be determined and suppressed based on the cepstral envelope, allows for howling detection and suppression based on the characteristics of the audio signal itself. Therefore, it is not limited by traditional signal models and can avoid suppressing spectral structures similar to howling, thereby improving the howling suppression effect. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of a single-microphone and single-speaker sound feedback system in related technologies;
[0014] Figure 2 This is a schematic diagram of a multi-microphone and multi-speaker sound feedback system in related technologies;
[0015] Figure 3 This is one of the flowcharts of the howling suppression method provided in the embodiments of this application;
[0016] Figure 4 This is a schematic diagram of the first audio frame and its cepstral envelope in the howling suppression method provided in the embodiments of this application;
[0017] Figure 5 This is the second flowchart of the howling suppression method provided in the embodiments of this application;
[0018] Figure 6 This is the third flowchart of the howling suppression method provided in the embodiments of this application;
[0019] Figure 7 This is the fourth flowchart of the howling suppression method provided in the embodiments of this application;
[0020] Figure 8 This is the fifth flowchart of the howling suppression method provided in the embodiments of this application;
[0021] Figure 9 This is the sixth flowchart of the howling suppression method provided in the embodiments of this application;
[0022] Figure 10 This is a schematic diagram illustrating the effect of suppressing the howling frequency band in the howling suppression method provided in this application embodiment;
[0023] Figure 11 This is a schematic diagram of the howling suppression device provided in the embodiments of this application;
[0024] Figure 12 This is a schematic diagram of the electronic device provided in the embodiments of this application;
[0025] Figure 13 This is a hardware schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0026] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0027] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0028] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc., in the instruction sent. An indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.
[0029] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."
[0030] The following section will first explain some of the terms or terms used in the specification and claims of this application.
[0031] Feedback: The tail sound phenomenon of a microphone when it is at the critical point of acoustic feedback. Feedback is a common abnormal phenomenon in sound reinforcement systems. Feedback requires the following two conditions to be met simultaneously: phase condition, the acoustic signal fed back to the microphone must be in phase with the acoustic signal input from the original sound source of the microphone; amplitude condition, the acoustic feedback loop must be positive feedback, that is, the feedback gain is greater than 1.
[0032] Cepstral envelope: the envelope of the cepstral spectrum; where the cepstral spectrum is the signal obtained by performing an inverse Fourier transform on the Fourier transform spectrum of a signal after logarithmic operation. Since the Fourier spectrum is generally a complex spectrum, it is also called a complex cepstral spectrum; the envelope is a graphic composed of many interwoven elliptical curves, which looks like an envelope. In mathematics, the envelope of a family of planar straight lines (or curves) is a curve that is tangent to any one of the straight lines (or curves) in the family.
[0033] The following description, in conjunction with the accompanying drawings, details the howling suppression method, apparatus, electronic device, and readable storage medium provided in this application through specific embodiments and application scenarios.
[0034] In PAS and HFCS, sound feedback has been a prominent and long-standing problem that has plagued the industry. Sound feedback refers to the phenomenon where sound emitted by a speaker is transmitted back to the microphone through different paths.
[0035] For example, such as Figure 1 As shown, s(t) is the sound source signal, F is the transfer function of the sound feedback path, G is the transfer function of the electroacoustic forward path, and the output signal of the loudspeaker is u(t) = G*y(t); where y(t) = x(t) + s(t) represents the microphone capture signal, x(t) = F*u(t) represents the signal obtained by the loudspeaker output after passing through the sound feedback transfer function, and * represents the convolution operation.
[0036] The main cause of howling in PAS and HFCS is that certain strong frequency points in the signal, after passing through the gain of the electrical signal transmission path G and the feedback path F, their energy will become larger and larger in this positive feedback closed system, leading to system instability and eventually howling. Howling has some significant characteristics: first, it presents a periodic harsh and sharp sound in terms of audibility, and its energy becomes stronger over time cycles; second, most howling occurs in the middle or high frequency bands, which is specifically manifested as the energy of some middle and high frequency points being far greater than the energy of normal harmonics; furthermore, howling frequency points may be a single frequency point or multiple frequency points, but generally there are no continuous howling frequency points on the frequency spectrum.
[0037] The existence of howling greatly affects the quality of voice calls. To solve this problem, many methods have emerged. For example, an adaptive feedback control (AFC) signal model based on an adaptive filter or the use of NH is adopted to suppress specific howling frequencies. However, all these methods follow the approach of detecting howling first and then suppressing it, or suppressing while detecting, which makes the effect of howling suppression largely depend on the accuracy of howling detection. Once howling detection is falsely triggered, it is easy to damage normal speech; for example, many unvoiced pronunciations such as "shi", "zhi" or "qi" in Chinese have spectrum structures similar to howling. Once these pronunciations are detected as howling and trigger the howling suppression algorithm, it is easy to damage the speech.
[0038] As Figure 1 shown, most current AFC and NH related algorithms are based on this model, that is, the single feedback path of single speaker and single microphone. This model is usually also shared with Acoustic Echo Cancellation (AEC). However, in practice, howling generation is far from being limited to this single feedback structure. As Figure 2 shown, an acoustic feedback path in a game voice chat scenario is provided. Assuming that multiple people are in a team game voice chat in this scenario, and two of them are within 1 meter apart and turn on the speaker and microphone of their mobile phones at the same time, then Figure 2 the acoustic feedback path in this scenario is simulated, where G1 and G2 are the electrical signal paths of two mobile phones respectively. Under ideal conditions, F' is the propagation feedback path between the two mobile phones. Since the user will constantly operate the mobile phone or the user's own movement during the game, the position between the two mobile phones changes in real time, which also causes F' to change in real time. This makes Figure 1The single feedback structure shown cannot describe this scenario because it typically assumes that F remains constant, making it unsuitable for complex feedback paths from a fundamental modeling perspective. This description assumes an ideal situation where the acoustic feedback paths of two phones share a single F'. However, in reality, users may use different brands and models of phones, leading to structural differences. The positions of microphones and speakers, the materials used, and the conductive structures all contribute to different feedback paths for the two phones. Real-world scenarios are far more complex than this idealized model, making feedback more likely. Therefore, traditional single feedback path models cannot address this problem, resulting in poor feedback suppression and consequently impairing normal speech.
[0039] To improve the accuracy of howling detection and reduce false triggers, many howling detection metrics have been proposed. For example, the Peak-to-Threshold Power Ratio (PTPR) feature calculates the ratio of the energy of the frequency band where howling is likely to occur to a threshold; when this ratio exceeds a certain range, howling can be considered to have occurred. Other features include the Peak-to-Harmonic Power Ratio (PHPR) feature, which calculates the ratio of the energy of the howling frequency band to its harmonic energy, and Interframe Peak Magnitude Persistence (IPMP), which calculates the duration of the howling frequency band's energy across consecutive frames. Many other similar howling detection features exist, but they are all fundamentally designed based on the energy characteristics of howling.
[0040] Besides using the traditional signal models or features mentioned above, directly using neural networks (NNs) for howling suppression has also become a common method in recent years. However, NN algorithms are always limited by the howling dataset and the computing power of the terminal device's inference capabilities. Therefore, if NN algorithms are used for howling processing, the howling is usually treated as noise and fused into a noise reduction network, or noise reduction, echo cancellation, and howling suppression are solved in a large model. This places certain demands on the computing power of the terminal device, thus limiting its application scope.
[0041] To address the aforementioned problems, embodiments of this application provide a howling suppression method, apparatus, electronic device, and readable storage medium. The howling suppression method provided in this application can be applied to scenarios such as hands-free calling, game voice chat, and multi-person conferencing.
[0042] The feedback suppression method provided in this application embodiment can determine the feedback signal within the frequency range of the first audio frame based on the cepstral envelope of the first audio frame. The feedback signal includes feedback frequency points or feedback frequency bands. The feedback signal is then suppressed based on the cepstral envelope. This scheme, because it can determine and suppress feedback frequency points or feedback frequency bands in the audio frame based on the cepstral envelope, allows for feedback detection and suppression based on the characteristics of the audio signal itself. Therefore, it is not limited by traditional signal models and can avoid suppressing spectral structures similar to feedback, thereby improving the feedback suppression effect.
[0043] Furthermore, unlike the aforementioned NN algorithms, it does not require a large amount of computation on a large dataset, thus saving computing power and the power consumption of electronic devices.
[0044] It should be noted that the howling suppression method provided in this application can be executed by a howling suppression device, an electronic device, or a functional module within an electronic device. Some embodiments of this application use an electronic device executing the howling suppression method as an example to illustrate the howling suppression method provided in this application.
[0045] Figure 3 A flowchart of the howling suppression method provided in an embodiment of this application is shown. Figure 3 As shown, the howling suppression method provided in this application embodiment may include the following steps 301 and 302.
[0046] Step 301: The electronic device determines the howling signal within the frequency range of the first audio frame based on the cepstral envelope of the first audio frame.
[0047] The aforementioned howling signals include howling frequency points or howling frequency bands.
[0048] In this embodiment of the application, the above-mentioned howling frequency point is the frequency point that generates howling, that is, only an isolated frequency point in the entire spectrum generates self-excitation feedback. This kind of howling usually occurs in the mid-low frequency band below 6KHz.
[0049] In this embodiment of the application, the above-mentioned howling frequency band is the frequency band that generates howling, that is, a certain sub-band on the spectrum generates howling. This howling usually occurs in the mid-to-high frequency band above 2KHz.
[0050] Optionally, in this embodiment of the application, the number of the above-mentioned howling frequency points can be one or more.
[0051] Optionally, in this embodiment of the application, the number of the above-mentioned howling frequency bands can be one or more.
[0052] Optionally, in this embodiment of the application, the first audio frame may be an audio frame of an audio signal picked up by the electronic device at the current moment of the system.
[0053] Optionally, in this embodiment of the application, the audio signal may be a voice call audio signal.
[0054] Optionally, in this embodiment of the application, the audio signal may include multiple consecutive audio frames. Whenever the electronic device picks up an audio frame, the electronic device may determine the howling signal within the frequency range of the audio frame based on the cepstral envelope of the picked-up audio frame.
[0055] Optionally, in this embodiment of the application, after the electronic device picks up the first audio frame, it can first calculate the cepstral envelope of the first audio frame.
[0056] For example, assuming the first audio frame is s(t), the electronic device can first perform a Fourier transform on s(t) using the following formula (1) to obtain the frequency domain signal X(jω):
[0057]
[0058] Then, the electronic device calculates the cepstral C of the first audio frame based on X(jω) using the following formula (2):
[0059]
[0060] In this context, the first few points of the aforementioned cepstral C are denoted as N. The electronic device then performs a Fourier transform on these N points to obtain... Finally, the electronic device can calculate the cepstral envelope E(t, f) of the first audio frame using the following formula (3):
[0061] E(t,f)=Real(C'(t,f)). (3)
[0062] It should be noted that the cepstral envelope reflects the changing characteristics of the spectrum, and the size of N determines the fineness of the envelope's spectral characterization; the larger N is, the finer the spectral characterization. In practical applications, with an audio frame sampling rate of 16kHz and a frame length of 10ms, a value of N between 10 and 15 usually yields a relatively ideal effect, effectively characterizing the features of howling and speech.
[0063] Figure 4 A schematic diagram of the first audio frame and its cepstral envelope is shown below, as follows: Figure 4 As shown, curve 41 is the spectrum of the first audio frame, and curve 42 is the cepstral envelope of the first audio frame. It can be seen that curve 42 can accurately reflect the changing characteristics of curve 41 and effectively reduce the redundant information in curve 41.
[0064] Optionally, in the embodiments of this application, combined with Figure 3 ,like Figure 5 As shown, step 301 above can be implemented through step 301a below.
[0065] Step 301a: The electronic device determines the howling frequency point based on the frequency intensity of each peak frequency point in the first audio frame and the frequency intensity of the corresponding frequency point in the cepstral envelope of each peak frequency point.
[0066] Optionally, in this embodiment of the application, before determining the aforementioned howling frequency point, the electronic device may first define a frequency point set P, which satisfies the following condition (a):
[0067]
[0068] Where L is the window length of the Fourier transform; it can be seen that the frequency points in the above frequency point set P are all the peak frequency points in the first audio frame above.
[0069] Then, the electronic device can determine the set of frequency points Q from the set of frequency points P that satisfies the following condition (b) based on the frequency point intensity of each peak frequency point and the frequency point intensity of the corresponding frequency point in the cepstral envelope of each peak frequency point:
[0070]
[0071] Here, α and β are adjustable parameters that determine the breadth of frequency suppression for howling; the smaller α and β are, the more likely frequency point p can be considered to be suppressed. i For the above-mentioned howling frequency points, conversely, the larger α and β are, the more likely p can be considered to be... i This is not the frequency point of the howling; in practical applications, the effect is better when α = 10 and β = 20. It can be seen that each frequency point in the above frequency point set Q satisfies the following: the frequency point intensity is greater than the sum of the frequency point intensity of the previous frequency point and α, and is greater than the sum of the frequency point intensity of the corresponding frequency point in the above cepstral envelope and β of the previous frequency point;
[0072] Therefore, the electronic device can determine the frequency points in the frequency point set Q as the aforementioned howling frequency points.
[0073] In this embodiment, since the electronic device can determine the howling frequency point based on the frequency intensity of each peak frequency point of the first audio frame and the frequency intensity of the corresponding frequency point in the cepstral envelope of each peak frequency point, the howling frequency point can be determined based on the curve characteristics of the audio signal itself, thereby improving the accuracy of determining the howling frequency point.
[0074] Optionally, in the embodiments of this application, combined with Figure 3 ,like Figure 6 As shown, step 301 above can be implemented through step 301b below.
[0075] Step 301b: The electronic device determines the howling frequency band based on the frequency domain positions of the local maxima and local minima in the cepstral envelope.
[0076] Optionally, in this embodiment of the application, before determining the aforementioned howling frequency band, the electronic device may first determine the aforementioned local maximum value |E| using the following condition (c). MC |:
[0077] |E MC |=max(H max ),H max ={h i :|E(h i )|>|E(h i +1)|&&|E(h i )|>|E(h i -1)|};(c)
[0078] Then, the local minimum H is determined by the following condition (d). min :
[0079] H min ={h i :|E(h i )|<|E(h i +1)|&&|E(h i )|<|E(h i -1)|}; (d)
[0080] Next, based on the above |E MC |Corresponding frequency domain position f MC Determine the frequency domain position f MC Left and right sides, and a distance f from the frequency domain position MC The frequency domain location of the nearest local minimum is denoted as f. MC-L and f MC-R Then we have:
[0081]
[0082]
[0083] Therefore, electronic devices can f MC-L and f MC-R The frequency band between these two points was determined to be the aforementioned howling frequency band.
[0084] In this embodiment, since the electronic device can determine the above-mentioned howling frequency band based on the frequency domain positions of the local maximum and local minimum values in the cepstral envelope, the howling frequency band can be determined based on the curve characteristics of the audio signal itself, thereby improving the accuracy of determining the howling frequency band.
[0085] Step 302: The electronic device suppresses the howling signal based on the cepstral envelope and the frame type of the first audio frame.
[0086] Optionally, in this embodiment of the application, the electronic device may suppress the above-mentioned howling signal based on the above-mentioned cepstral envelope and the frame type of the above-mentioned first audio frame.
[0087] Optionally, in this embodiment of the application, the aforementioned howling signal includes the aforementioned howling frequency point. For example, in conjunction with... Figure 3 ,like Figure 7 As shown, step 302 can be implemented through steps 302a and 302b below.
[0088] Step 302a: The electronic device determines at least one target frequency point of the first audio frame based on the frame type of the first audio frame.
[0089] Among them, at least one of the target frequency points mentioned above includes the aforementioned howling frequency point.
[0090] Optionally, in the embodiments of this application, the frequency points other than the above-mentioned howling frequency points among the at least one target frequency points are those adjacent to the howling frequency point and with a higher frequency point intensity.
[0091] Optionally, in the embodiments of this application, the above-mentioned frame type can be a voiced frame, an unvoiced frame, or a non-speech frame, etc.
[0092] Optionally, in the embodiments of this application, step 302a above can be implemented by step A or step B below.
[0093] Step A: When the frame type of the first audio frame is a voiced frame, the electronic device determines at least one target frequency point within the first frequency point set.
[0094] Step B: If the frame type of the first audio frame is a non-voiced frame, the electronic device determines at least one target frequency point within the second frequency point set.
[0095] The frequency range of the first frequency set is less than or equal to the frequency range of the second frequency set.
[0096] Optionally, in the embodiments of this application, the aforementioned unvoiced frames can be unvoiced frames or non-speech frames, etc.
[0097] Optionally, in this embodiment of the application, the frequency range of the first frequency point set is less than or equal to the frequency range of the second frequency point set, that is, the number of frequency points in the first frequency point set is less than or equal to the number of frequency points in the second frequency point set.
[0098] For example, the frequency range of the first frequency set is [-a, a], and the frequency range of the second frequency set is [-b, b], where a and b are both positive integers, and a is less than or equal to b.
[0099] It should be noted that in actual implementation, the frequency range of the first frequency point set is usually smaller than that of the second frequency point set. However, for certain audio frames that meet specific spectral characteristics (such as the critical curve characteristics of voiced and unvoiced frames), the frequency range of the first frequency point set can be equal to that of the second frequency point set. That is, regardless of whether the frame type is a voiced or unvoiced frame, at least one target frequency point can be determined within the same frequency point set, thereby simplifying the computation of determining the target frequency point while ensuring the suppression effect.
[0100] In this embodiment of the application, since the electronic device can determine at least one target frequency point from the set of frequency points corresponding to the frame type of the first audio frame, it can determine different numbers of target frequency points near the howling frequency point as frequency points to be suppressed according to the frame type, thereby improving the flexibility of howling suppression.
[0101] Step 302b: The electronic device suppresses at least one target frequency point based on the frequency intensity of each target frequency point and the frequency intensity of the corresponding frequency point in the cepstral envelope of each target frequency point.
[0102] Optionally, in this embodiment of the application, the electronic device can suppress at least one of the target frequency points mentioned above using the following formula (4):
[0103]
[0104] Wherein, γ is an adjustable parameter used to control the suppression strength of at least one target frequency point. Typically, the value of γ is between (0,1). If the frame type is a voiced frame, then K=1 or K=3. If the frame type is an unvoiced frame, then K=3 or K=5.
[0105] It should be noted that in actual implementation, electronic devices can also suppress only the above-mentioned howling frequency points. In this case, it is sufficient to set m in the above formula (4) to 0.
[0106] In this embodiment of the application, since the electronic device can suppress at least one target frequency point determined by the frame type of the first audio frame based on the frequency point intensity of each target frequency point and the frequency point intensity of the corresponding frequency point in the cepstral envelope of each target frequency point, it can suppress the howling frequency point and several frequency points adjacent to the howling frequency point based on the characteristics of the audio signal itself, thereby improving the effect of suppressing the howling frequency point.
[0107] Optionally, in this embodiment of the application, the aforementioned howling signal includes the aforementioned howling frequency band. For example, in conjunction with... Figure 3 ,like Figure 8 As shown, step 302 above can be implemented through step 302c below.
[0108] Step 302c: When the first probability value is greater than or equal to the first threshold, the electronic device suppresses the howling band based on the cepstral envelope and the frame type of the first audio frame.
[0109] The first probability value is used to indicate the probability of the first audio frame experiencing frequency band howling. The first probability value is determined based on at least one of the first feature, the second feature, and the third feature. The first feature is used to indicate the ratio of the peak intensity of the frequency point in the cepstral envelope to the low-frequency energy difference. The second feature is used to indicate the bandwidth of the spectral peak of the cepstral envelope. The third feature is used to indicate the slope between the spectral peak and the spectral valley of the cepstral envelope.
[0110] It should be noted that the first, second, and third features mentioned above are all based on the cepstral envelope definition, and are not designed directly from the energy perspective of the frequency or time domain, as in related technologies such as PTPR or PHPR. This makes it easier to distinguish howling from normal speech.
[0111] Optionally, in the embodiments of this application, the first feature mentioned above can be the cepstrum envelope peak to low frequency strength ratio (CEPTLFS), which can be defined as the following formula (5):
[0112]
[0113] in, The average energy of the low-frequency peak can be obtained based on the above condition (a). If the above frame type is a voiced frame, then This can be approximated as the energy intensity of voiced sounds; if the above frame type is a non-voiced frame, then... This refers to the energy intensity in the low-frequency region.
[0114] For example, if the above frame type is a voiced frame and |E MC If the frame falls in the mid-frequency band, the larger the CEPTLFS value, the greater the probability of howling in the mid-frequency band. If the frame type is a non-voiced frame, the value of the first feature can be used as a reference for the energy difference between the mid and low frequencies, and other features need to be introduced for judgment.
[0115] Optionally, in the embodiments of this application, the second feature mentioned above can be the cepstrum envelope peak band width ratio (CEPBWR), which can be defined by the following formula (6):
[0116]
[0117] Optionally, in the embodiments of this application, if it is voice, then f MC-R -f MC-L The larger the value, the smaller the corresponding CEPBWR value; if it is a howling signal, the energy of the howling signal is usually more concentrated, mostly a narrowband signal, f MC-R -f MC-L If it is too small, the corresponding CEPBWR will be larger.
[0118] Optionally, in the embodiments of this application, the third feature mentioned above can be the cepstrum envelope peak-to-valley slope (CEPTVS), which can be defined as the following formula (7):
[0119]
[0120] Optionally, in the embodiments of this application, the above-mentioned CEPTVS describes the rate of change from the peak to the valley of the envelope spectrum. Howling usually changes rapidly from low energy to high energy in the frequency domain, while the energy distribution of speech in the frequency domain is relatively stable. Therefore, the larger the CEPTVS, the more likely it is to be howling.
[0121] For example, taking the first probability value as determined based on the first feature, the second feature, and the third feature, the electronic device can calculate the first probability value prob using the following formula (8) before suppressing the whistling frequency band:
[0122] prob=w1CEPTLFS+w2CEPBWR+w3CEPTVS; (8)
[0123] Among them, w1, w2 and w3 are all weights.
[0124] Optionally, in the embodiments of this application, the first threshold can be preset by the system or set by the user according to actual usage needs, and the embodiments of this application do not limit it.
[0125] It should be noted that the larger the first probability value, the greater the probability that the first audio frame will experience frequency band howling; conversely, the smaller the first probability value, the smaller the probability that the first audio frame will experience frequency band howling.
[0126] In this embodiment of the application, since the electronic device will only suppress the howling frequency band based on the cepstral envelope and the frame type when the first probability value used to indicate the probability of howling frequency bands occurring in the first audio frame is greater than or equal to the first threshold, it can improve and reduce unnecessary howling frequency band suppression, thereby saving the power consumption of the electronic device.
[0127] Optionally, in the embodiments of this application, combined with Figure 8 ,like Figure 9 As shown, step 302c can be implemented through steps 302c1, 302c2, or 302c3 as described below.
[0128] Step 302c1: When the first probability value is greater than or equal to the first threshold, and when the frame type is a non-voice frame, the electronic device suppresses the howling frequency band according to the frequency intensity of each frequency point in the howling frequency band.
[0129] Optionally, in this embodiment of the application, when the frame type is the non-voice frame, the electronic device can specifically suppress the howling frequency band using the following formula (9):
[0130]
[0131] Step 302c2: When the first probability value is greater than or equal to the first threshold, and when the frame type is a voiced frame, the electronic device suppresses the howling band based on the frequency intensity, peak frequency intensity, and average energy of the low-frequency peak of the cepstral envelope of each frequency point in the howling band.
[0132] Optionally, in this embodiment of the application, when the frame type is the voiced frame, the electronic device can specifically suppress the howling frequency band using the following formula (10):
[0133]
[0134] Step 302c1: When the first probability value is greater than or equal to the first threshold, and when the frame type is a clear frame, the electronic device suppresses the howling frequency band based on the frequency intensity, peak frequency intensity of each frequency point in the howling frequency band, and the intensity of the first audio frame.
[0135] Optionally, in this embodiment of the application, when the frame type is the aforementioned uncomfort frame, the electronic device can specifically suppress the aforementioned howling frequency band using the following formula (11):
[0136]
[0137] in, The intensity of the first audio frame can be determined by the intensity of the most recent voiced frame preceding the first audio frame.
[0138] In this way, electronic devices can suppress the aforementioned howling frequency bands according to different frame types, so as to effectively protect the speech when howling and speech frames overlap, suppress the howling to an intensity similar to the speech energy, and thus maximize the listening experience without damaging the speech.
[0139] In this embodiment, since the electronic device can use a suppression strategy corresponding to the frame type of the first audio frame to suppress the howling frequency band, the flexibility and accuracy of suppressing the howling frequency band can be improved.
[0140] It should be noted that if there are multiple howling frequency points or howling frequency bands, the above method can be used to suppress each howling frequency point or howling frequency band separately to ensure the quality of the picked-up speech. Moreover, the howling suppression methods provided in the embodiments of this application only adjust the amplitude of the signal, while the phase of the signal remains unchanged throughout the suppression process.
[0141] The following description, in conjunction with the accompanying drawings, illustrates the effect of howling suppression using the howling suppression method provided in the embodiments of this application.
[0142] For example, such as Figure 10 The image shows a test corpus during a feedback loop. The audio frame has a sampling rate of 16kHz, a frame length of 10ms, a Fast Fourier Transform length of 512, and an offset of 160. The first channel is the original input signal, and the second channel is the signal after feedback suppression using the feedback suppression method provided in this embodiment. Specifically, the signal in region 12 is the signal after suppressing the feedback frequency band in region 11, and the signal in region 13 is the signal after suppressing the feedback frequency point. It can be seen that the feedback suppression method provided in this embodiment can accurately suppress the feedback frequency points and bands. Furthermore, when feedback and speech overlap, it can reduce speech impairment to a limited width subband while suppressing feedback. For scenarios such as multi-person hands-free calls and game voice chat, it does not affect the user's listening experience. Moreover, based on traditional signal processing algorithms, it consumes very little computing power and is easy to deploy.
[0143] In the feedback suppression method provided in this application embodiment, since the feedback frequency point or feedback frequency band in the audio frame can be determined and suppressed based on the cepstral envelope of the audio frame, that is, feedback detection and suppression can be performed based on the characteristics of the audio signal itself, it is not limited to the traditional signal model, and can avoid suppressing spectral structures similar to feedback, thereby improving the feedback suppression effect.
[0144] The feedback suppression method provided in this application can be executed by a feedback suppression device. This application uses a feedback suppression device executing the feedback suppression method as an example to illustrate the feedback suppression device provided in this application.
[0145] like Figure 11 As shown, this application embodiment provides a howling suppression device 110, which may include a determination module 111 and a suppression module 112.
[0146] The determining module 111 can be used to determine the howling signal within the frequency range of the first audio frame based on the cepstral envelope of the first audio frame, wherein the howling signal includes a howling frequency point or a howling frequency band. The suppressing module 112 can be used to suppress the howling signal based on the cepstral envelope.
[0147] In one possible implementation, the determining module 111 can be specifically used to determine the howling frequency point based on the frequency point intensity of each peak frequency point of the first audio frame and the frequency point intensity of the corresponding frequency point in the cepstral envelope of each peak frequency point.
[0148] In one possible implementation, the aforementioned howling signal includes the aforementioned howling frequency point, and the suppression module 112 may include a determination submodule and a suppression submodule. The determination submodule may be used to determine at least one target frequency point of the first audio frame based on the frame type of the first audio frame, the at least one target frequency point including the howling frequency point. The suppression submodule may be used to suppress the at least one target frequency point based on the frequency point intensity of each of the at least one target frequency point and the frequency point intensity of the corresponding frequency point in the aforementioned cepstral envelope.
[0149] In one possible implementation, the aforementioned determining submodule can be specifically used to: determine at least one target frequency point within a first frequency point set when the frame type is a voiced frame; and determine the at least one target frequency point within a second frequency point set when the frame type is an unvoiced frame. The frequency range of the first frequency point set is less than or equal to the frequency range of the second frequency point set.
[0150] In one possible implementation, the determining module 111 can be specifically used to determine the above-mentioned howling frequency band based on the frequency domain positions of the local maximum and local minimum values in the above-mentioned cepstral envelope.
[0151] In one possible implementation, the aforementioned howling signal includes the aforementioned howling frequency band. The suppression module 112 can specifically be used to suppress the aforementioned howling frequency band based on the aforementioned cepstral envelope and the frame type of the aforementioned first audio frame when the first probability value is greater than or equal to a first threshold; wherein the first probability value indicates the probability of howling occurring in the first audio frame, and the first probability value is determined according to at least one of a first feature, a second feature, and a third feature; the first feature indicates the ratio of the peak intensity of the frequency point in the cepstral envelope to the low-frequency energy difference; the second feature indicates the bandwidth of the spectral peak of the cepstral envelope; and the third feature indicates the slope between the spectral peak and the spectral valley of the cepstral envelope.
[0152] In one possible implementation, the suppression module 112 can be specifically used to: suppress the howling frequency band based on the frequency intensity of each frequency point in the howling frequency band when the frame type of the first video frame is a non-audio frame; suppress the howling frequency band based on the frequency intensity of each frequency point in the howling frequency band, the peak value of the frequency intensity, and the average energy of the low-spectral peak of the cepstral envelope when the frame type of the first video frame is a voiced frame; and suppress the howling frequency band based on the frequency intensity of each frequency point in the howling frequency band, the peak value of the frequency intensity, and the intensity of the first audio frame when the frame type of the first video frame is a voiceless frame.
[0153] In the feedback suppression device provided in this application embodiment, since the feedback suppression device can determine and suppress feedback frequency points or feedback frequency bands in the audio frame based on the cepstral envelope of the audio frame, that is, it can perform feedback detection and suppression based on the characteristics of the audio signal itself, so it is not limited to the traditional signal model, and can avoid suppressing similar spectral structures, thereby improving the feedback suppression effect.
[0154] The howling suppression device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0155] The howling suppression device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0156] The whistling suppression device provided in this application embodiment can realize all the processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0157] like Figure 12 As shown, this application embodiment also provides an electronic device 120, including a processor 121 and a memory 122. The memory 122 stores a program or instructions that can run on the processor 121. When the program or instructions are executed by the processor 121, they implement the various steps of the above-described howling suppression method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0158] It should be noted that the electronic devices in the embodiments of this application include mobile electronic devices and non-mobile electronic devices.
[0159] Figure 13 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0160] like Figure 13As shown, the electronic device 1000 includes, but is not limited to, components such as: radio frequency unit 1001, network module 1002, audio output unit 1003, input unit 1004, sensor 1005, display unit 1006, user input unit 1007, interface unit 1008, memory 1009, and processor 1010.
[0161] Those skilled in the art will understand that the electronic device 1000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 13 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0162] The processor 1010 can be used to determine the howling signal within the frequency range of the first audio frame based on the cepstral envelope of the first audio frame, the howling signal including howling frequency points or howling frequency bands; and suppress the howling signal based on the cepstral envelope.
[0163] In one possible implementation, the processor 1010 can be specifically used to determine the howling frequency point based on the frequency point intensity of each peak frequency point of the first audio frame and the frequency point intensity of the corresponding frequency point in the cepstral envelope of each peak frequency point.
[0164] In one possible implementation, the aforementioned howling signal includes the aforementioned howling frequency point. The processor 1010 can specifically be configured to determine at least one target frequency point of the first audio frame, including the howling frequency point, based on the frame type of the first audio frame; and to suppress the at least one target frequency point based on the frequency point intensity of each of the at least one target frequency points and the frequency point intensity of the corresponding frequency point in the aforementioned cepstral envelope.
[0165] In one possible implementation, the processor 1010 is specifically configured to: determine at least one target frequency point within a first frequency point set when the frame type is a voiced frame; and determine the at least one target frequency point within a second frequency point set when the frame type is an unvoiced frame. The frequency range of the first frequency point set is less than or equal to the frequency range of the second frequency point set.
[0166] In one possible implementation, the processor 1010 can be used to determine the above-mentioned howling frequency band based on the frequency domain positions of the local maxima and local minima in the above-mentioned cepstral envelope.
[0167] In one possible implementation, the aforementioned howling signal includes the aforementioned howling frequency band. The processor 1010 can specifically be used to suppress the aforementioned howling frequency band based on the aforementioned cepstral envelope and the frame type of the aforementioned first audio frame, when a first probability value is greater than or equal to a first threshold; wherein the first probability value indicates the probability of howling occurring in the first audio frame, and the first probability value is determined based on at least one of a first feature, a second feature, and a third feature, the first feature indicating the ratio of the peak intensity of the frequency point in the cepstral envelope to the low-frequency energy difference, the second feature indicating the bandwidth of the spectral peak of the cepstral envelope, and the third feature indicating the slope between the spectral peak and the spectral valley of the cepstral envelope.
[0168] In one possible implementation, the processor 1010 can be specifically used to: suppress the howling frequency band based on the frequency intensity of each frequency point in the howling frequency band when the frame type of the first video frame is a non-audio frame; suppress the howling frequency band based on the frequency intensity of each frequency point in the howling frequency band, the peak value of the frequency intensity, and the average energy of the low-spectral peak of the cepstral envelope when the frame type of the first video frame is a voiced frame; and suppress the howling frequency band based on the frequency intensity of each frequency point in the howling frequency band, the peak value of the frequency intensity, and the intensity of the first audio frame when the frame type of the first video frame is a voiceless frame.
[0169] In the electronic device provided in this application embodiment, since the electronic device can determine and suppress the howling frequency point or howling frequency band in the audio frame based on the cepstral envelope of the audio frame, that is, it can perform howling detection and suppression based on the characteristics of the audio signal itself, it does not need to be limited by the traditional signal model, and can avoid suppressing spectral structures similar to howling, thereby improving the effect of howling suppression.
[0170] It should be understood that, in this embodiment, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0171] The memory 1009 can be used to store software programs and various data. The memory 1009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1009 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0172] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor 1010.
[0173] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described howling suppression method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0174] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0175] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described howling suppression method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0176] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0177] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described howling suppression method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0178] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0179] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0180] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for suppressing howling, characterized in that, The method includes: Based on the cepstral envelope of the first audio frame, the howling signal within the frequency range of the first audio frame is determined, and the howling signal includes howling frequency points or howling frequency bands; Based on the cepstral envelope, the howling signal is suppressed; The howling signal includes the howling frequency band; Suppressing the howling signal based on the cepstral envelope includes: If the first probability value is greater than or equal to the first threshold, the howling frequency band is suppressed based on the cepstral envelope and the frame type of the first audio frame; Wherein, the first probability value is used to indicate the probability of the first audio frame experiencing frequency band howling. The first probability value is determined based on at least one of the first feature, the second feature, and the third feature. The first feature is used to indicate the ratio of the peak intensity of the frequency point in the cepstral envelope to the difference in low-frequency energy. The second feature is used to indicate the bandwidth of the spectral peak of the cepstral envelope. The third feature is used to indicate the slope between the spectral peak and the spectral valley of the cepstral envelope.
2. The method according to claim 1, characterized in that, The determination of the howling signal within the frequency range of the first audio frame based on the cepstral envelope of the first audio frame includes: The howling frequency point is determined based on the frequency intensity of each peak frequency point in the first audio frame and the frequency intensity of the corresponding frequency point in the cepstral envelope for each peak frequency point.
3. The method according to claim 1 or 2, characterized in that, The howling signal includes the howling frequency point; The method of suppressing the howling signal based on the cepstral envelope includes: Based on the frame type of the first audio frame, at least one target frequency point of the first audio frame is determined, and the at least one target frequency point includes the howling frequency point; Based on the frequency intensity of each target frequency point and the frequency intensity of the corresponding frequency point in the cepstral envelope, the at least one target frequency point is suppressed.
4. The method according to claim 3, characterized in that, Determining at least one target frequency point of the first audio frame based on its frame type includes: In the case where the frame type is a voiced frame, the at least one target frequency point is determined within the first frequency point set; In the case where the frame type is a non-voiced frame, the at least one target frequency point is determined within the second frequency point set; Wherein, the frequency range of the first frequency point set is less than or equal to the frequency range of the second frequency point set.
5. The method according to claim 1, characterized in that, The determination of the howling signal within the frequency range of the first audio frame based on the cepstral envelope of the first audio frame includes: The howling frequency band is determined based on the frequency domain positions of the local maxima and local minima in the cepstral envelope.
6. The method according to claim 1, characterized in that, Suppressing the howling frequency band based on the cepstral envelope and the frame type of the first audio frame includes: When the frame type is a non-speech frame, the howling frequency band is suppressed according to the frequency intensity of each frequency point in the howling frequency band; When the frame type is a voiced frame, the howling frequency band is suppressed based on the frequency intensity of each frequency point in the howling frequency band, the peak value of the frequency intensity, and the average energy of the low-spectral peak of the cepstral envelope. When the frame type is a clear frame, the howling frequency band is suppressed based on the frequency intensity of each frequency point in the howling frequency band, the peak value of the frequency intensity, and the intensity of the first audio frame.
7. A whistling suppression device, characterized in that, The device includes a determining module and a suppressing module; The determining module is used to determine the howling signal within the frequency range of the first audio frame based on the cepstral envelope of the first audio frame, wherein the howling signal includes howling frequency points or howling frequency bands; The suppression module is used to suppress the howling signal based on the cepstral envelope; The howling signal includes the howling frequency band; The suppression module is specifically used to suppress the howling frequency band based on the cepstral envelope and the frame type of the first audio frame when the first probability value is greater than or equal to the first threshold. Wherein, the first probability value is used to indicate the probability of the first audio frame experiencing frequency band howling. The first probability value is determined based on at least one of the first feature, the second feature, and the third feature. The first feature is used to indicate the ratio of the peak intensity of the frequency point in the cepstral envelope to the difference in low-frequency energy. The second feature is used to indicate the bandwidth of the spectral peak of the cepstral envelope. The third feature is used to indicate the slope between the spectral peak and the spectral valley of the cepstral envelope.
8. The apparatus according to claim 7, characterized in that, The determining module is specifically used to determine the howling frequency point based on the frequency point intensity of each peak frequency point of the first audio frame and the frequency point intensity of the corresponding frequency point of each peak frequency point in the cepstral envelope.
9. The apparatus according to claim 7 or 8, characterized in that, The howling signal includes the howling frequency point, and the suppression module includes a determination submodule and a suppression submodule; The determining submodule is used to determine at least one target frequency point of the first audio frame according to the frame type of the first audio frame, wherein the at least one target frequency point includes the howling frequency point; The suppression submodule is used to suppress the at least one target frequency point based on the frequency point intensity of each target frequency point and the frequency point intensity of the corresponding frequency point of each target frequency point in the cepstral envelope.
10. The apparatus according to claim 7, characterized in that, The determining module is specifically used to determine the howling frequency band based on the frequency domain positions of the local maxima and local minima in the cepstral envelope.
11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the howling suppression method as described in any one of claims 1-6.
12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the howling suppression method as described in any one of claims 1-6.
Citation Information
Patent Citations
Audio data processing method and device, equipment and storage medium
CN115223584A
Audio signal squeal detection and suppression method and device
CN115720317A