Squeal Suppression Method, Device, Computer Equipment and Storage Medium

By performing frequency domain transformation and subband gain calculation on the audio signals of voice call devices, and combining historical gain data to suppress howling, the problem of howling in voice calls is solved and the quality of voice call is improved.

CN114333749BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011062254.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-30
Publication Date
2025-07-15
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

In voice calls, especially when multiple voice calls devices are close to each other, howling is prone to occur, resulting in a degradation of voice calls quality, and it is difficult for the prior art to effectively suppress howling.

Method used

By obtaining the current audio signal, frequency domain transformation is performed, molecular banding is divided, and the subband gain coefficient is determined based on howling detection results and speech detection results, and the current subband gain is calculated using the historical subband gain to perform howling suppression.

Benefits of technology

Accurately suppressing whistling, improving the quality of voice calls, and ensuring the clarity and intelligibility of the audio signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333749B_ABST
    Figure CN114333749B_ABST
Patent Text Reader

Abstract

The present application relates to a method, device, computer device, and storage medium for suppressing whistling. The method includes: obtaining a current audio signal corresponding to a current time period, performing a frequency-domain transformation on the current audio signal to obtain a frequency-domain audio signal; dividing the frequency-domain audio signal to obtain each sub-band, and determining a target sub-band from each sub-band; obtaining a current whistling detection result and a current speech detection result corresponding to the current audio signal, and determining a sub-band gain coefficient corresponding to the current audio signal based on the current whistling detection result and the current speech detection result; obtaining a historical sub-band gain corresponding to the audio signal in a historical time period, and calculating a current sub-band gain corresponding to the current audio signal based on the sub-band gain coefficient and the historical sub-band gain; suppressing whistling in the target sub-band based on the current sub-band gain to obtain a first target audio signal corresponding to the current time period. The accuracy of whistling suppression is improved by using this method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to a method and apparatus for suppressing howling, a computer device, and a storage medium. Background Art

[0002] With the development of Internet communication technologies, it is possible to conduct voice calls based on a network. For example, voice calls in various instant messaging applications. However, when making a voice call, especially during a voice conference, there are often two or more voice call devices in close proximity. For example, in the same room. At this time, howling is very likely to occur, which in turn affects the quality of the voice call. Currently, it is usually necessary to adjust the distance between the voice call devices to avoid howling. However, when the distance cannot be adjusted, howling will occur, resulting in a decrease in the quality of the voice call. Summary of the Invention

[0003] Based on this, it is necessary to provide a howling suppression method, apparatus, computer device, and storage medium that can suppress howling and thus improve the quality of voice calls for the above technical problems.

[0004] A howling suppression method, the method comprising:

[0005] Obtain a current audio signal corresponding to a current time period, perform a frequency-domain transformation on the current audio signal to obtain a frequency-domain audio signal;

[0006] Divide the frequency-domain audio signal to obtain respective subbands, and determine a target subband from the respective subbands;

[0007] Obtain a current howling detection result and a current voice detection result corresponding to the current audio signal, and determine a subband gain coefficient corresponding to the current audio signal based on the current howling detection result and the current voice detection result;

[0008] Obtain a historical subband gain corresponding to an audio signal in a historical time period, and calculate a current subband gain corresponding to the current audio signal based on the subband gain coefficient and the historical subband gain;

[0009] Suppress howling in the target subband based on the current subband gain to obtain a first target audio signal corresponding to the current time period.

[0010] A howling suppression apparatus, the apparatus comprising:

[0011] A signal transformation module, configured to obtain a current audio signal corresponding to a current time period, and perform a frequency-domain transformation on the current audio signal to obtain a frequency-domain audio signal;

[0012] A subband determination module, configured to divide the frequency-domain audio signal to obtain respective subbands, and determine a target subband from the respective subbands;

[0013] A coefficient determination module, configured to obtain a current howling detection result and a current speech detection result corresponding to a current audio signal, and determine a sub-band gain coefficient corresponding to the current audio signal based on the current howling detection result and the current speech detection result;

[0014] A gain determination module, configured to obtain a historical sub-band gain corresponding to an audio signal in a historical time period, and calculate a current sub-band gain corresponding to the current audio signal based on the sub-band gain coefficient and the historical sub-band gain;

[0015] A howling suppression module, configured to perform howling suppression on a target sub-band based on the current sub-band gain to obtain a first target audio signal corresponding to the current time period.

[0016] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0017] Obtain a current audio signal corresponding to the current time period, perform a frequency domain transformation on the current audio signal to obtain a frequency domain audio signal;

[0018] Divide the frequency domain audio signal to obtain each sub-band, and determine a target sub-band from each sub-band;

[0019] Obtain a current howling detection result and a current speech detection result corresponding to the current audio signal, and determine a sub-band gain coefficient corresponding to the current audio signal based on the current howling detection result and the current speech detection result;

[0020] Obtain a historical sub-band gain corresponding to an audio signal in a historical time period, and calculate a current sub-band gain corresponding to the current audio signal based on the sub-band gain coefficient and the historical sub-band gain;

[0021] Perform howling suppression on the target sub-band based on the current sub-band gain to obtain a first target audio signal corresponding to the current time period.

[0022] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0023] Obtain a current audio signal corresponding to the current time period, perform a frequency domain transformation on the current audio signal to obtain a frequency domain audio signal;

[0024] Divide the frequency domain audio signal to obtain each sub-band, and determine a target sub-band from each sub-band;

[0025] Obtain a current howling detection result and a current speech detection result corresponding to the current audio signal, and determine a sub-band gain coefficient corresponding to the current audio signal based on the current howling detection result and the current speech detection result;

[0026] Obtain the historical sub-band gain corresponding to the audio signal in the historical time period, and calculate the current sub-band gain corresponding to the current audio signal based on the sub-band gain coefficient and the historical sub-band gain;

[0027] Perform howling suppression on the target sub-band based on the current sub-band gain to obtain the first target audio signal corresponding to the current time period.

[0028] The howling suppression method, device, computer device, and storage medium provided by the embodiments of the present application obtain the current audio signal corresponding to the current time period, and then obtain the current howling detection result and the current voice detection result corresponding to the current audio signal, so as to be able to determine the sub-band gain coefficient corresponding to the current audio signal according to the current howling detection result and the current voice detection result, and calculate the current sub-band gain corresponding to the current audio signal through the sub-band gain coefficient and the historical sub-band gain, so that the obtained current sub-band gain is more accurate, and then use the current sub-band gain to perform howling suppression on the target sub-band, so as to be able to accurately suppress howling, improve the quality of the first target audio signal corresponding to the obtained current time period, and thus be able to improve the voice call quality. Description of the Drawings

[0029] Figure 1 It is an application environment diagram of the howling suppression method in an embodiment;

[0030] Figure 2 It is a flowchart of the howling suppression method in an embodiment;

[0031] Figure 2a It is a schematic diagram of the relationship between the frequency and energy of the audio signal in a specific embodiment;

[0032] Figure 3 It is a flowchart of the process of obtaining the current audio signal in an embodiment;

[0033] Figure 4 It is a flowchart of the howling detection process in an embodiment;

[0034] Figure 5 It is a flowchart of the process of obtaining the current audio signal in another embodiment;

[0035] Figure 6 It is a flowchart of the process of obtaining the current audio signal in yet another embodiment;

[0036] Figure 7 It is a flowchart of the process of obtaining the sub-band gain coefficient in an embodiment;

[0037] Figure 8 It is a flowchart of the process of obtaining the second target audio signal in an embodiment;

[0038] Figure 8a Schematic diagram of energy constraint curve in a specific embodiment;

[0039] Figure 9 Schematic diagram of the process of howling suppression method in a specific embodiment;

[0040] Figure 10 Schematic diagram of the application scenario of howling suppression method in a specific embodiment;

[0041] Figure 11 Schematic diagram of the application framework of howling suppression method in a specific embodiment;

[0042] Figure 12 Schematic diagram of the process of howling suppression method in a specific embodiment;

[0043] Figure 13 Schematic diagram of the application framework of howling suppression method in another specific embodiment;

[0044] Figure 14 Schematic diagram of the application framework of howling suppression method in yet another specific embodiment;

[0045] Figure 15 Block diagram of the structure of howling suppression device in an embodiment;

[0046] Figure 16 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0047] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0048] The howling suppression method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 106 through the network, the terminal 104 communicates with the server 106 through the network, the terminal 102 and the terminal 104 make a voice call through the server 106, and the terminal 102 and the terminal 104 are relatively close, for example, in the same room. The terminal 102 and the terminal 104 can be either a sending terminal for sending voice or a receiving terminal for receiving voice. The terminal 102 or the terminal 104 obtains the current audio signal corresponding to the current time period, performs a frequency-domain transformation on the current audio signal to obtain a frequency-domain audio signal; the terminal 102 or the terminal 104 divides the frequency-domain audio signal to obtain each sub-band, and determines a target sub-band from each sub-band; the terminal 102 or the terminal 104 obtains the current howling detection result and the current voice detection result corresponding to the current audio signal, and determines the sub-band gain coefficient corresponding to the current audio signal based on the current howling detection result and the current voice detection result; the terminal 102 or the terminal 104 obtains the historical sub-band gain corresponding to the audio signal in the historical time period, and calculates the current sub-band gain corresponding to the current audio signal based on the sub-band gain coefficient and the historical sub-band gain; the terminal 102 or the terminal 104 performs howling suppression on the target sub-band based on the current sub-band gain to obtain the first target audio signal corresponding to the current time period. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server can be implemented by an independent server or a server cluster composed of multiple servers.

[0049] In one embodiment, as Figure 2 shown, a howling suppression method is provided. Taking the terminal in Figure 1 as an example, the method includes the following steps:

[0050] Step 202, obtain the current audio signal corresponding to the current time period, and perform a frequency-domain transformation on the current audio signal to obtain a frequency-domain audio signal.

[0051] Among them, the audio signal is a carrier of the frequency and amplitude change information of sound waves with speech, music, sound effects, etc. The current audio signal refers to the audio signal for which howling suppression is required, that is, there is a howling signal in the current audio signal. Due to problems such as too short a distance between the sound source and the sound reinforcement device, energy self-excitation occurs, resulting in howling. The howling signal refers to the audio signal corresponding to the howling, and the howling is often sharp and harsh. The current audio signal can be an audio signal that needs to be subjected to howling suppression obtained by collecting an audio signal through a collection device such as a microphone and performing signal processing. This signal processing can include echo cancellation, noise suppression, howling detection, and so on. Echo cancellation means eliminating the noise generated by the air return path between a collection device such as a microphone and a playback device such as a speaker through a sound wave interference method. Noise suppression means extracting a pure original audio from the noisy audio, and the audio signal does not contain background noise. Howling detection means detecting whether there is a howling signal in the audio signal. The current audio signal can also be an audio signal that needs to be subjected to howling suppression obtained by receiving an audio signal through a network and performing processing. This signal processing can be howling detection. The current time period refers to the time period in which the current audio signal is located, that is, the time period after voice framing of the audio signal. For example, the length of the current time period can be within 10 ms to 30 ms. Frequency domain transformation means transforming the current audio signal from the time domain to the frequency domain. The time domain refers to the relationship between the audio signal and time. The time domain waveform of the audio signal can express the change of the audio signal over time. The frequency domain is a coordinate system used when describing the characteristics of the signal in terms of frequency, that is, the audio signal changes with frequency. The frequency domain diagram shows the amount of signal in each given frequency band within a frequency range. The frequency domain representation can also include information about the phase shift of each sine curve so that the frequency components can be recombined to restore the original time signal. The frequency domain audio signal refers to the audio signal obtained by transforming the current audio signal from the time domain to the frequency domain.

[0052] Feasibly, the terminal can collect voice through a collection device such as a microphone to obtain the audio signal in the current time period, and then perform howling detection on the audio signal. Among them, the howling can be detected through a machine learning model established by a neural network, or the howling can be checked through parameter criteria such as the peak / mean ratio. The howling can also be detected based on the fundamental period in the audio signal. The howling can also be detected based on the energy in the audio signal.

[0053] When there is a howling signal in the audio signal, the current audio signal corresponding to the current time period is obtained. Then, the current audio signal is subjected to frequency domain transformation through Fourier transform to obtain the frequency domain audio signal. Among them, before the terminal performs howling detection on the collected audio signal, it can also perform processing such as echo cancellation and noise suppression on the collected audio signal.

[0054] The terminal can also obtain the voice sent by other voice call terminals through the network to obtain the audio signal in the current time period, and then perform howling detection on the audio signal. When there is a howling signal in the audio signal, the current audio signal corresponding to the current time period is obtained, and then the current audio signal is subjected to frequency domain transformation through Fourier transform to obtain the frequency domain audio signal.

[0055] Step 204: Divide the frequency domain audio signal to obtain each sub-band, and determine the target sub-band from each sub-band.

[0056] Among them, the sub-band refers to the sub-frequency band obtained by dividing the frequency domain audio signal. The target sub-band refers to the sub-band that needs to perform howling suppression.

[0057] Feasibly, when the terminal divides the frequency domain audio signal, it can use a band-pass filter to divide the frequency domain audio signal to obtain each sub-band. Among them, the division of the sub-band can be performed according to the preset number of sub-bands, or according to the preset frequency band range, etc. Then calculate the energy of each sub-band, and select the target sub-band according to the energy of each sub-band. Among them, the selected target sub-band can be one, for example, the sub-band with the maximum energy is the target sub-band, or it can be multiple. For example, the selected target sub-bands can be the preset number of sub-bands selected in order from large to small according to the energy of the sub-bands.

[0058] Step 206: Obtain the current howling detection result and the current voice detection result corresponding to the current audio signal, and determine the sub-band gain coefficient corresponding to the current audio signal based on the current howling detection result and the current voice detection result.

[0059] Among them, the current howling detection result refers to the detection result obtained after performing howling detection on the current audio signal, which can include that there is a howling signal in the current audio signal and that there is no howling signal in the current audio signal. The current voice detection result refers to the detection result obtained after performing voice endpoint detection on the current audio signal. Among them, voice endpoint detection (Voice Activity Detection, VAD) refers to accurately locating the start and end of the voice in the current audio signal. The current voice detection result can include that there is a voice signal in the current audio signal and that there is no voice signal in the current audio signal. The sub-band gain coefficient is used to represent the degree of howling suppression required for the current audio signal. When the sub-band gain coefficient is smaller, it means that the degree of howling suppression required for the current audio signal is higher. When the sub-band gain coefficient is larger, it means that the degree of howling suppression required for the current audio signal is smaller.

[0060] Feasibly, the terminal can obtain the current howling detection result and the current voice detection result corresponding to the current audio signal. The current howling detection result and the current voice detection result corresponding to the current audio signal can be obtained by performing howling detection and voice endpoint detection on the current audio signal before howling suppression, and then saving the current howling detection result and the current voice detection result into the memory.

[0061] The terminal can also obtain the current howling detection result and the current voice detection result corresponding to the current audio signal from a third party, where the third party is the service provider that performs howling detection and voice endpoint detection on the current audio signal. For example, the terminal can obtain the saved current howling detection result and the current voice detection result corresponding to the current audio signal from a server.

[0062] Step 208: Obtain the historical sub-band gain corresponding to the audio signal in the historical time period, and calculate the current sub-band gain corresponding to the current audio signal based on the sub-band gain coefficient and the historical sub-band gain.

[0063] Among them, the historical time period refers to the historical time period corresponding to the current time period. The time length of the historical time period can be the same as that of the current time period, or different from that of the current time period. The historical time period can be the previous time period of the current time period, or multiple time periods before the current time period. The historical time period can have a preset interval with the current time period, or be directly connected to the current time period. For example, within the time range of 0 ms to 100 ms, the current time period can be 80 ms to 100 ms, and the historical time period can be the time period of 60 ms to 80 ms. The audio signal in the historical time period refers to the audio signal after howling suppression has been performed. The historical sub-band gain refers to the sub-band gain used for howling suppression of the audio signal in the historical time period. The current sub-band gain refers to the sub-band gain used for howling suppression of the current audio signal.

[0064] Feasibly, the terminal can obtain the historical sub-band gain corresponding to the audio signal in the historical time period from the memory, calculate the product of the sub-band gain coefficient and the historical sub-band gain, and obtain the current sub-band gain corresponding to the current audio signal. Among them, when the current time period is the starting time period, the historical sub-band gain is the pre-set initial sub-band gain value. For example, the initial sub-band gain value can be 1. An initial sub-band gain value of 1 indicates that no suppression will be performed on the current audio signal. When the sub-band gain coefficient is less than 1, it indicates that howling suppression needs to be performed on the current audio signal. When the sub-band gain coefficient is greater than 1, it indicates that the howling suppression on the current audio signal needs to be reduced.

[0065] In one embodiment, the current subband gain corresponding to the current audio signal is compared with the lower limit value of the preset subband gain. When the current subband gain corresponding to the current audio signal is less than the lower limit value of the preset subband gain, the lower limit value of the preset subband gain is used as the current subband gain corresponding to the current audio signal.

[0066] In one embodiment, the current subband gain corresponding to the current audio signal is compared with the initial subband gain value. When the current subband gain corresponding to the current audio signal is greater than the initial subband gain value, the initial subband gain value is used as the current subband gain corresponding to the current audio signal.

[0067] Step 210, perform howling suppression on the target subband based on the current subband gain to obtain a first target audio signal corresponding to the current time period.

[0068] Wherein, the first target audio signal refers to the audio signal obtained by performing howling suppression on the target subband in the current audio signal.

[0069] Feasibly, the spectrum of the target subband is gain-adjusted using the current subband gain, and then the gain-adjusted audio signal is converted from the frequency domain to the time domain using the inverse Fourier transform algorithm to obtain the first target audio signal corresponding to the current time period.

[0070] In one embodiment, if the current audio signal corresponding to the current time period is collected by a collection device such as a microphone of the terminal, the first target audio signal corresponding to the current time period can be encoded to obtain an encoded audio signal, and then the encoded audio signal is sent to other terminals for voice calls through the network interface. For example, as Figure 1 shown, the terminal 102 collects an audio signal through a microphone, performs echo cancellation and noise suppression to obtain the current audio signal corresponding to the current time period, then performs howling suppression on the current audio signal to obtain the first target audio signal corresponding to the current time period, and sends the first target audio signal corresponding to the current time period to the terminal 104 through the server 106. The terminal 104 receives the first target audio signal corresponding to the current time period, decodes it, and then plays the decoded first target audio signal.

[0071] In one embodiment, after obtaining the first target audio signal, the volume of the first target audio signal can also be adjusted. For example, the volume of the first target audio signal can be increased, and then the first target audio signal with the increased volume is encoded, and then the encoded first target audio signal is sent to other terminals for voice calls through the network interface.

[0072] In one embodiment, the current audio signal corresponding to the current time period is sent by another voice call terminal through the network interface. Then, the first target audio signal corresponding to the current time period can be directly played for voice. For example, as Figure 1 shown, the terminal 102 collects an audio signal through a microphone. After echo cancellation and noise suppression, the audio signal is encoded and sent to the terminal 104 through the server 106. The terminal 104 receives the encoded audio signal, decodes it to obtain the decoded audio signal, processes the decoded audio signal to obtain the current audio signal corresponding to the current time period, then performs howling suppression on the current audio signal to obtain the first target audio signal corresponding to the current time period, and then plays the first target audio signal.

[0073] In a feasible embodiment, as Figure 2a shown, it is a schematic diagram of the relationship between the frequency and energy of an audio signal. Among them, the abscissa in this schematic diagram represents frequency, and the ordinate represents energy. Different subbands are obtained based on frequency division. The figure shows 9 subbands. The subbands with frequencies lower than 1400HZ are low-frequency subbands, and the subbands with frequencies higher than 1400HZ are high-frequency subbands. The low-frequency subbands are the 1st to 4th subbands, and the high-frequency bands are the 5th to 9th subbands. The solid line in this figure represents the relationship curve between frequency and energy when there is only a voice signal. The dotted line represents the relationship curve between frequency and energy when there are voice signals and howling signals in the audio signal. It can be seen that the energy is significantly more when there are voice signals and howling signals in the audio signal than when there is only a voice signal. At this time, in the high-frequency subband, the 8th subband has the most energy, so the 8th subband is determined as the target subband. Howling suppression is performed on the target subband. Since howling suppression is performed on the 8th subband, the energy of the 8th subband gradually decreases until the energy of the 6th subband is the maximum subband energy. The 6th subband is determined as the target subband, and then howling suppression is performed on the 6th subband.

[0074] The above howling suppression method can determine the subband gain coefficient corresponding to the current audio signal according to the current howling detection result and the current voice detection result by obtaining the current audio signal corresponding to the current time period and then obtaining the current howling detection result and the current voice detection result corresponding to the current audio signal, and calculate the current subband gain corresponding to the current audio signal through the subband gain coefficient and the historical subband gain, so that the obtained current subband gain is more accurate. Then, the current subband gain is used to perform howling suppression on the target subband, so that howling can be accurately suppressed, the quality of the first target audio signal corresponding to the current time period is improved, and thus the voice call quality can be improved.

[0075] In one embodiment, as Figure 3 shown, step 202, obtaining the current audio signal corresponding to the current time period, includes:

[0076] Step 302: Collect the initial audio signal corresponding to the current time period, and perform echo cancellation on the initial audio signal to obtain the initial audio signal after echo cancellation.

[0077] The initial audio signal refers to the digital audio signal obtained by converting the user's voice collected by a microphone or other acquisition device.

[0078] Feasibly, when the terminal is a sending terminal for sending voice, the terminal collects the initial audio signal corresponding to the current time period, and uses an echo cancellation algorithm to perform echo cancellation on the initial audio signal to obtain the initial audio signal after echo cancellation. Among them, echo cancellation can estimate the desired signal through an adaptive algorithm, and the desired signal approximates the echo signal passing through the actual echo path, that is, the simulated echo signal. Then, the simulated echo is subtracted from the initial audio signal collected by a microphone or other acquisition device to obtain the initial audio signal after echo cancellation. The echo cancellation algorithm includes at least one of the LMS (Least Mean Square) algorithm, the RLS (Recursive Least Square) algorithm, and the APA (Affine Projection Algorithm) algorithm.

[0079] Step 304: Perform voice activity detection on the initial audio signal after echo cancellation to obtain the current voice detection result.

[0080] Feasibly, the terminal uses a voice activity detection algorithm to perform voice activity detection on the initial audio signal after echo cancellation to obtain the current voice detection result. Among them, the voice activity detection algorithm includes the double-threshold detection method, the energy-based endpoint detection algorithm, the cepstrum coefficient-based endpoint detection algorithm, the frequency band variance-based endpoint detection algorithm, the autocorrelation similarity distance-based endpoint detection algorithm, the information entropy-based endpoint detection algorithm, and so on.

[0081] Step 306: Perform noise suppression on the initial audio signal after echo cancellation based on the current voice detection result to obtain the initial audio signal after noise suppression.

[0082] Feasibly, when the current voice detection result indicates that the initial audio signal after echo cancellation does not contain a voice signal, noise estimation is performed on the initial audio signal after echo cancellation and the noise is suppressed to obtain the initial audio signal after noise suppression. Among them, a trained neural network model for noise removal can be used for noise suppression, or a filter can be used for noise suppression. When the current voice detection result indicates that the initial audio signal after echo cancellation contains a voice signal, noise suppression is performed while trying to retain the voice signal to obtain the initial audio signal after noise suppression. The voice signal refers to the signal corresponding to the user's voice.

[0083] Step 308: Perform howling detection on the initial audio signal after noise suppression to obtain the current howling detection result.

[0084] Feasibly, the terminal uses a howling detection algorithm to perform howling detection on the initial audio signal after noise suppression to obtain the current howling detection result. Among them, the howling detection algorithm can be a detection algorithm based on energy distribution, such as the peak harmonic power ratio algorithm, the peak-to-peak ratio algorithm, the frame-by-frame peak holding algorithm, etc. It can also be a detection algorithm based on a neural network, etc.

[0085] Step 310: When the current howling detection result indicates that there is a howling signal in the initial audio signal after noise suppression, use the initial audio signal after noise suppression as the current audio signal corresponding to the current time period.

[0086] Feasibly, when the terminal detects that there is a howling signal in the initial audio signal after noise suppression, it uses the initial audio signal after noise suppression as the current audio signal corresponding to the current time period, and then performs howling suppression on the current audio signal corresponding to the current time period.

[0087] In the above embodiment, by performing echo cancellation on the collected initial audio signal, performing voice endpoint detection on the initial audio signal after echo cancellation, performing noise suppression based on the current voice detection result, performing howling detection on the initial audio signal after noise suppression, and when it is detected that there is a howling signal in the initial audio signal after noise suppression, using the initial audio signal after noise suppression as the current audio signal corresponding to the current time period, it is ensured that the obtained current audio signal is the audio signal that needs to be subjected to howling suppression.

[0088] In one embodiment, Step 304, where the step of performing voice endpoint detection on the initial audio signal after echo cancellation to obtain the current voice detection result includes:

[0089] Input the initial audio signal after echo cancellation into a voice endpoint detection model for detection to obtain the current voice detection result. The voice endpoint detection model is trained using a neural network algorithm based on training audio signals and corresponding training voice detection results.

[0090] Among them, the neural network algorithm can be a BP (back propagation, feedforward neural network) neural network algorithm, an LSTM (Long Short-Term Memory) algorithm, an RNN (Recurrent Neural Network) neural network algorithm, etc. The training audio signal refers to the audio signal used when training the voice endpoint detection model, and the training voice detection result refers to the voice detection result corresponding to the training audio signal. The training voice detection result includes that the training audio signal contains a voice signal and that the training audio signal does not contain a voice signal. Among them, the loss function uses the cross-entropy loss function and is optimized using the gradient descent method, and the activation function uses the Sigmoid function.

[0091] Feasibly, the terminal uses wavelet analysis to extract audio features from the initial audio signal after echo cancellation. The audio features include the short-time zero-crossing rate, short-time energy, kurtosis of the short-time amplitude spectrum, skewness of the short-time amplitude spectrum, etc. The audio features are input into the voice endpoint detection model for detection to obtain the current voice detection result. The current voice detection result includes that the initial audio signal after echo cancellation contains a voice signal and that the initial audio signal after echo cancellation does not contain a voice signal. The voice endpoint detection model is trained using a neural network algorithm based on the training audio signal and the corresponding training voice detection result. It can be trained and saved based on the training audio signal and the corresponding training voice detection result using a neural network algorithm in the server, and the terminal obtains the voice endpoint detection model from the server for use. It can also be trained based on the training audio signal and the corresponding training voice detection result using a neural network algorithm in the terminal.

[0092] In one embodiment, step 304, that is, performing voice endpoint detection on the initial audio signal after echo cancellation to obtain the current voice detection result, includes:

[0093] Performing low-pass filtering on the initial audio signal after echo cancellation to obtain a low-frequency signal; calculating the signal energy corresponding to the low-frequency signal, calculating the energy fluctuation based on the signal energy, and determining the current voice detection result according to the energy fluctuation.

[0094] Among them, low-pass filtering is a filtering method, and the rule is that low-frequency signals can pass through normally, while high-frequency signals exceeding the set critical value are blocked and attenuated. However, the degree of blocking and attenuation will vary according to different frequencies and different filtering procedures (purposes). The signal energy refers to the short-time energy corresponding to the low-frequency signal. The energy fluctuation refers to the ratio of the signal energy between the previous frame of the low-frequency signal and the next frame of the low-frequency signal.

[0095] Feasibly, due to the different energy distributions of the speech signal and the howling signal in the audio signal, and the low-frequency energy in the howling signal is significantly weaker than that of the speech signal. Then the terminal performs low-pass filtering on the initial audio signal after echo cancellation according to a preset low-frequency value to obtain a low-frequency signal. The preset low-frequency value can be 500HZ. Then calculate the signal energy corresponding to each frame in the low-frequency signal, and the triangular filter can be used to calculate the signal energy. Then calculate the ratio between the signal energy corresponding to the previous frame and the signal energy corresponding to the next frame. When the ratio exceeds the preset energy ratio, it indicates that the initial audio signal after echo cancellation contains a speech signal. When the ratio does not exceed the preset energy ratio, it indicates that the initial audio signal after echo cancellation does not contain a speech signal, thereby obtaining the current speech detection result.

[0096] In the above embodiment, by performing low-pass filtering on the initial audio signal after echo cancellation to obtain a low-frequency signal, and then determining the current speech detection result according to the energy fluctuation of the low-frequency signal, the obtained current speech detection result can be made more accurate.

[0097] In one embodiment, step 304, that is, performing voice endpoint detection on the initial audio signal after echo cancellation to obtain the current speech detection result, includes the steps of:

[0098] Performing low-pass filtering on the initial audio signal after echo cancellation to obtain a low-frequency signal, performing pitch detection on the low-frequency signal to obtain a pitch period, and determining the current speech detection result according to the pitch period.

[0099] Among them, generally, sound is composed of a series of vibrations with different frequencies and amplitudes emitted by the sounding body. Among these vibrations, there is a vibration with the lowest frequency, and the sound emitted by it is the fundamental tone, and the rest are overtones. Pitch detection refers to the estimation of the pitch period, which is used to detect a trajectory curve that is exactly the same as or as close as possible to the vocal cord vibration frequency. The pitch period refers to the time when the vocal cords open and close once.

[0100] Feasibly, the terminal performs low-pass filtering on the initial audio signal after echo cancellation to obtain a low-frequency signal, and uses a pitch detection algorithm to perform pitch detection on the low-frequency signal to obtain a pitch period. Among them, the pitch detection algorithm can include the autocorrelation method, the average magnitude difference function method, the parallel processing method, the cepstrum method, the simplified inverse filtering method, and so on. Then determine whether the initial audio signal after echo cancellation contains a speech signal according to the pitch period, that is, if the pitch period can be detected, it indicates that the initial audio signal after echo cancellation contains a speech signal. If the pitch period cannot be detected, it indicates that the initial audio signal after echo cancellation does not contain a speech signal, thereby obtaining the current speech detection result.

[0101] In the above embodiments, the current voice detection result is obtained by detecting the pitch period, which improves the accuracy of obtaining the current voice detection result.

[0102] In one embodiment, step 308, that is, performing a howling detection on the initial audio signal after noise suppression to obtain the current howling detection result, includes the steps of:

[0103] Inputting the initial audio signal after noise suppression into a howling detection model for detection to obtain the current howling detection result, where the howling detection model is trained using a neural network algorithm based on howling training audio signals and corresponding training howling detection results.

[0104] Among them, the neural network algorithm can be a BP (back propagation, feedforward neural network) neural network algorithm, an LSTM (Long Short-Term Memory, long short-term memory artificial neural network) algorithm, an RNN (Recurrent Neural Network, recurrent neural network) neural network algorithm, and so on. The howling training audio signal refers to the audio signal used when training the howling detection model. The training howling detection result refers to the howling detection result corresponding to the howling training audio signal, including that the initial audio signal after noise suppression contains a howling signal and that the initial audio signal after noise suppression does not contain a howling signal.

[0105] Feasibly, the terminal can extract the audio features corresponding to the initial audio signal after noise suppression. The audio features include MFCC (Mel-Frequency cepstrum coefficients) dynamic features, band representative vectors, and various types of audio fingerprints. The Mel-Frequency cepstrum coefficients refer to the coefficients that make up the Mel-Frequency cepstrum. The audio fingerprint refers to the digital features in the initial audio signal after noise suppression extracted in the form of an identifier through a specific algorithm. The band representative vector is an ordered list of indices of prominent tones in a frequency band. The terminal inputs the extracted audio features into the howling detection model for detection to obtain the current howling detection result.

[0106] In the above embodiments, by using the howling detection model to perform howling detection on the initial audio signal after noise suppression, the efficiency and accuracy of howling detection are improved.

[0107] In one embodiment, as Figure 4 shown, step 308, that is, performing a howling detection on the initial audio signal after noise suppression to obtain the current howling detection result, includes:

[0108] Step 402: Extract the initial audio features corresponding to the initial audio signal after noise suppression.

[0109] The initial audio features refer to the audio features extracted from the initial audio signal after noise suppression, and the initial audio features include at least one of Mel-Frequency Cepstrum Coefficients (MFCC) dynamic features, band representative vectors, and various types of audio fingerprints.

[0110] In one embodiment, the terminal can also select corresponding audio features according to accuracy and computational complexity. When the computational resources of the terminal are limited, the band representative vectors and various types of audio fingerprints are used as the initial audio features. When higher accuracy is required, the Mel-Frequency Cepstrum Coefficients dynamic features, band representative vectors, and various types of audio fingerprints, i.e., all of them, can be used as the initial audio features.

[0111] Feasibly, the terminal extracts the initial audio features corresponding to the initial audio signal after noise suppression. For example, to extract the Mel-Frequency Cepstrum Coefficients dynamic features, the initial audio signal after noise suppression can be pre-emphasized, then framed, windowed for each frame, and subjected to a Fast Fourier Transform on the windowed result to obtain the transformed result. The logarithmic energy is calculated for the transformed result through triangular filtering, and then the Mel-Frequency Cepstrum Coefficients dynamic features are obtained after discrete cosine transform.

[0112] Step 404: Obtain the first historical audio signal corresponding to the first historical time period, and extract the first historical audio features corresponding to the first historical audio signal.

[0113] The first historical time period refers to the time period before the current time period and has the same time length as the current time period. There can be multiple such first historical time periods. For example, if the current call duration is 2500 ms and the length of the current time period is 300 ms, that is, the current time period is from 2200 ms to 2500 ms, and the preset interval is 20 ms, then the first historical time periods can be 200 ms - 500 ms, 220 ms - 520 ms, 240 ms - 540 ms,..., 1980 ms - 2280 ms, and 2000 - 2300 ms. The first historical audio signal refers to the historical audio signal corresponding to the first historical time period, which is the audio signal collected by the microphone during the first historical time period.

[0114] Feasibly, the terminal obtains the first historical audio signal corresponding to the first historical time period and extracts the first historical audio features corresponding to the first historical audio signal.

[0115] Step 406: Calculate the first similarity between the initial audio feature and the first historical audio feature, and determine the current howling detection result based on the first similarity.

[0116] The first similarity refers to the similarity between the initial audio feature and the first historical audio feature, which can be a distance similarity or a cosine similarity.

[0117] Feasibly, the terminal can use a similarity algorithm to calculate the first similarity between the initial audio feature and the first historical audio feature. When the first similarity exceeds a pre-set first similarity threshold, it indicates that there is a howling signal in the initial audio signal after noise suppression. When the first similarity does not exceed the pre-set first similarity threshold, it indicates that there is no howling signal in the initial audio signal after noise suppression, thus obtaining the current howling detection result.

[0118] In one embodiment, when there are multiple first historical time periods, multiple first historical audio signals can be obtained. Calculate the first historical audio features corresponding to each first historical audio signal respectively, and calculate the first similarity between each first historical audio feature and the initial audio feature respectively. Statistically analyze the continuous duration during which the first similarity exceeds the pre-set first similarity threshold. When the continuous duration exceeds the pre-set duration, it indicates that there is a howling signal in the initial audio signal after noise suppression. When the continuous duration does not exceed the pre-set duration, it indicates that there is no howling signal in the initial audio signal after noise suppression, thus obtaining the current howling detection result.

[0119] In the above embodiment, by calculating the first similarity between the initial audio feature and the first historical audio feature, due to the cyclic transmission in the voice sending terminal and the voice receiving terminal of the howling signal, there is a historical similarity. Then, determine the current howling detection result based on the first similarity, thereby making the obtained current howling detection result more accurate.

[0120] In one embodiment, the howling suppression method further includes the steps of:

[0121] When the current howling detection result indicates that there is a howling signal in the current audio signal, obtain the audio signal to be played and the preset audio watermark signal, add the preset audio watermark signal to the audio signal to be played and play it.

[0122] Wherein, the audio signal to be played refers to the audio signal that the terminal is about to play through the playback device while the user is speaking. This audio signal can be perceptible to the human ear (such as playing the voice of the other party) or imperceptible (such as the quiet background sound when the other party is not speaking). The preset audio watermark signal refers to the audio signal that is preset to indicate the existence of a howling signal in the audio signal sent through the network. It is an audio signal that is imperceptible to the human ear. For example, it can be a high-frequency watermark signal selected from the high-frequency band or even the ultrasonic band.

[0123] Feasibly, since howling is detected and suppressed at the sending terminal, when there are multiple receiving terminals receiving voice signals, it will cause the audio signals received by all receiving terminals of the voice signals to be the audio signals after howling suppression, affecting the audio signal quality of all receiving terminals. At this time, when the sending terminal detects that there is a howling signal in the current audio signal, it does not perform howling suppression, obtains the audio signal to be played and the preset audio watermark signal, adds the preset audio watermark signal to the audio signal to be played and plays it, and then does not perform howling suppression on the current audio signal, and directly sends the current audio signal through the network to all receiving terminals. In one embodiment, a single-frequency tone or multi-frequency tone with a preset frequency can be embedded in the high-frequency band of the audio signal to be played as the preset high-frequency watermark signal. In one embodiment, multiple preset high-frequency watermark signals can be embedded in the audio signal to be played for playback. In one embodiment, the time-domain audio watermark algorithm can also be used to add the preset audio watermark signal to the audio signal to be played. In one embodiment, the transform-domain audio watermark algorithm can also be used to add the preset audio watermark signal to the audio signal to be played. Since the receiving terminal generating howling is close to the sending terminal, the receiving terminal generating howling can receive the audio signal with the preset audio watermark signal added and the current audio signal at this time. Then the receiving terminal generating howling detects the audio signal with the preset audio watermark signal added to obtain the result that there is a howling signal in the current audio signal, and then suppresses the current audio signal to obtain the first target audio signal and plays it, avoiding reducing the quality of the audio signals received by all receiving terminals.

[0124] In one embodiment, as Figure 5 shown, the howling suppression method further includes the steps of:

[0125] 502, collect the first audio signal corresponding to the first time period, detect the first audio signal based on the audio watermark detection algorithm, and determine that the first audio signal contains the target audio watermark signal.

[0126] Among them, the first audio signal refers to the audio signal collected by a collection device such as a microphone from a nearby terminal that is sent through a playback device. There may be howling between this terminal and the nearby terminal. The audio watermark detection algorithm is used to detect the audio watermark signal added to the first audio signal. It can be the adjacent band energy ratio algorithm. The adjacent band energy ratio algorithm can be to calculate the ratio between the energies corresponding to each sub-band in the first audio signal, and extract the audio watermark signal according to the ratio. The target audio watermark signal refers to the preset audio watermark signal added by the nearby terminal in the first audio signal. The first time period refers to the time period corresponding to the first audio signal.

[0127] Feasibly, when the terminal is a receiving terminal for receiving voice, the terminal collects the first audio signal corresponding to the first time period through a collection device such as a microphone. The first audio signal is divided into sub-bands, and the energy of each sub-band is calculated. Then, the energies of adjacent sub-bands are compared to obtain the adjacent band energy ratio. When the adjacent band energy ratio exceeds the preset adjacent band energy ratio threshold, it is determined that the first audio signal contains the target high-frequency watermark signal. At this time, it indicates that the audio signal received through the network contains a howling signal. The preset adjacent band energy ratio threshold refers to the threshold of the adjacent band energy ratio set in advance, which is used to detect whether there is a preset high-frequency watermark signal. In one embodiment, the audio watermark algorithm added to the first audio signal can also be detected through a watermark extraction algorithm.

[0128] 506, receive the target network-coded audio signal corresponding to the second time period, and decode the target network-coded audio signal to obtain the target network audio signal.

[0129] Among them, the second time period refers to the time period corresponding to the target network-coded audio signal. This second time period is after the first time period. The target network-coded audio signal refers to the currently encoded audio signal received through the network. The target network audio signal refers to the currently decoded audio signal.

[0130] Feasibly, the terminal receives the target network-coded audio signal corresponding to the second time period through the network, and decodes the target network-coded audio signal to obtain the target network audio signal.

[0131] 506, based on the fact that the first audio signal contains the target audio watermark signal, use the target network audio signal as the current audio signal.

[0132] Feasibly, the terminal uses the target network audio signal as the current audio signal according to the fact that the first audio signal contains the target audio watermark signal.

[0133] In the above embodiments, when the terminal is a terminal for receiving voice, a preset audio watermark signal can be detected through the collected first audio signal. When a preset audio watermark signal is detected in the first audio signal, the target network audio signal received through the network is used as the current audio signal, and then the current audio signal is subjected to howling suppression to avoid affecting the quality of the audio signals received by all terminals. Moreover, by detecting the preset audio watermark signal to determine whether to use the target network audio signal as the current audio signal, the accuracy of the obtained current audio signal is improved.

[0134] In one embodiment, as Figure 6 shown, step 202 of obtaining the current audio signal corresponding to the current time period includes:

[0135] Step 602: Receive the current network-coded audio signal corresponding to the current time period, decode the network-coded audio signal to obtain the current network audio signal.

[0136] Wherein, the current time period refers to the time period of the current network-coded audio signal received by the terminal through the network. The current network-coded audio signal refers to the encoded audio signal received through the network.

[0137] Feasibly, when the terminal is a terminal for receiving voice, the terminal receives the current network-coded audio signal corresponding to the current time period through the network interface, and decodes the network-coded audio signal to obtain the current network audio signal.

[0138] Step 604: Perform voice activity detection on the current network audio signal to obtain a network voice detection result, and at the same time perform howling detection on the current network audio signal to obtain a network howling detection result.

[0139] Wherein, the network language detection result refers to the result obtained by performing voice activity detection on the current network audio signal, including that the current network audio signal contains a voice signal and that the current network audio signal does not contain a voice signal. The network howling detection result refers to the result obtained by performing howling detection on the current network audio signal, which may include that the current network audio signal contains a howling signal and that the current network audio signal does not contain a howling signal.

[0140] In one embodiment, a voice activity detection model is used to perform voice activity detection on the current network audio signal to obtain a network voice detection result, and a howling detection model is used to perform howling detection on the current network audio signal to obtain a network howling detection result.

[0141] In one embodiment, the current network audio signal can be low-pass filtered to obtain a low-frequency signal, the signal energy corresponding to the low-frequency signal is calculated, the energy fluctuation is calculated based on the signal energy, and the network voice detection result corresponding to the current network audio signal is determined according to the energy fluctuation.

[0142] In one embodiment, the current network audio signal can be low-pass filtered to obtain a low-frequency signal, the pitch detection is performed on the low-frequency signal to obtain a pitch period, and the network voice detection result corresponding to the current network audio signal is determined according to the pitch period.

[0143] In one embodiment, the current network audio feature corresponding to the current network audio signal can be extracted, the historical network audio feature is obtained, the similarity between the historical network audio feature and the current network audio feature is calculated, and the network howling detection result is determined based on the similarity.

[0144] Step 606: Extract the network audio feature of the current network audio signal, obtain the second historical audio signal in the second historical time period, and extract the second historical audio feature corresponding to the second historical audio signal.

[0145] Wherein, the network audio feature refers to the audio feature corresponding to the current network audio signal. The second historical time period refers to the time period corresponding to the second historical audio signal, and there can be multiple second historical time periods. The second historical audio signal refers to the historical audio signal collected by a collection device such as a microphone. The second historical audio feature refers to the audio feature corresponding to the second historical audio signal.

[0146] Feasibly, the terminal extracts the network audio feature of the current network audio signal, obtains the second historical audio signal in the second historical time period saved in the memory, and extracts the second historical audio feature corresponding to the second historical audio signal.

[0147] Step 608: Calculate the network audio similarity between the network audio feature and the second historical audio feature, and determine the network audio signal as the current audio signal corresponding to the current time period based on the network audio similarity and the network howling detection result.

[0148] Wherein, the network audio similarity refers to the similarity degree between the current network audio signal and the second historical audio signal. The higher the network audio similarity, the closer the distance between the terminal and the terminal that sends the current network audio signal.

[0149] Feasibly, the terminal calculates the network audio similarity between the network audio feature and the second historical audio feature through a similarity algorithm. When the network audio similarity exceeds a preset network audio similarity threshold and the network howling detection result indicates that there is a howling signal in the current network audio signal, the network audio signal is used as the current audio signal corresponding to the current time period. Among them, the preset network audio similarity threshold is a threshold used to determine the position of the terminal sending the current network audio signal. When the network audio similarity exceeds the preset network audio similarity threshold, it indicates that the position of the terminal is close to the terminal sending the current network audio signal and howling is likely to occur. When the network audio similarity does not exceed the preset network audio similarity threshold, it indicates that the position of the terminal is far from the terminal sending the current network audio signal and howling is not likely to occur.

[0150] In one embodiment, the terminal can obtain multiple second historical audio signals, extract the second historical audio features corresponding to each second historical audio signal, and calculate the network audio similarity between each second historical audio feature and the second historical audio feature. When the duration for which the network audio similarity exceeds the preset network audio similarity threshold exceeds a preset threshold, it indicates that the position of the terminal is close to the terminal sending the current network audio signal. When the duration for which the network audio similarity exceeds the preset network audio similarity threshold does not exceed the preset threshold, it indicates that the position of the terminal is far from the terminal sending the current network audio signal.

[0151] In the above embodiment, by calculating the network audio similarity between the network audio feature and the second historical audio feature, and determining the network audio signal as the current audio signal corresponding to the current time period based on the network audio similarity and the network howling detection result, the determined current audio signal is made more accurate.

[0152] In one embodiment, in step 204, dividing the frequency-domain audio signal to obtain each sub-band, and determining the target sub-band from each sub-band includes:

[0153] Dividing the frequency-domain audio signal according to the preset number of sub-bands to obtain each sub-band. Calculating the sub-band energy corresponding to each sub-band, and smoothing the sub-band energy of each sub-band to obtain the smoothed sub-band energy of each sub-band. Determining the target sub-band based on the smoothed sub-band energy of each sub-band.

[0154] Among them, the preset number of sub-bands is the number of sub-bands to be divided that is preset.

[0155] Feasibly, the terminal unevenly divides the frequency-domain audio signal according to the preset number of subbands to obtain each subband. The terminal then calculates the subband energy corresponding to each subband, and the subband energy can be volume or logarithmic energy. That is, in one embodiment, triangular filters can be used to calculate the subband energy corresponding to each subband. For example, the energy of each subband can be calculated by 30 triangular filters. The frequency ranges of each subband may not be equal, and there may be frequency overlap between adjacent subbands. Then, smooth processing is performed on each subband energy, that is, the energy corresponding to the subbands at the same position existing in the most recent time period is obtained, and then the average value is calculated to obtain the smoothed subband energy of the subband. For example, to smooth the subband energy of the first subband in the current audio signal, the historical subband energy of the first subband in the most recent 10 historical audio signals can be obtained, and then the average subband energy is calculated, and the average subband energy is used as the smoothed subband energy of the first subband in the current audio signal. The smoothed subband energy corresponding to each subband is calculated in turn.

[0156] Then, the smoothed subband energies of each subband are compared, and the subband with the largest subband energy is selected as the target subband, and this target subband contains the most howling energy. In one embodiment, the subband with the largest subband energy can be selected starting from a specified subband. For example, if the current audio signal is divided into 30 subbands, the subband corresponding to the largest smoothed subband energy can be selected from the 6th to the 30th subbands. In one embodiment, a preset number of subbands can be selected as the target subbands in order from largest to smallest according to the comparison result. For example, the top three subbands sorted by subband energy from largest to smallest are selected as the target subbands.

[0157] In the above embodiment, by smoothing the subband energies of each subband and selecting the target subband from each subband according to the smoothed subband energy, the selected target subband is more accurate.

[0158] In one embodiment, determining the target subband based on the smoothed subband energies of each subband includes:

[0159] Obtain the current howling detection result corresponding to the current audio signal, determine each howling subband from each subband according to the current howling detection result, and obtain the howling subband energy of each howling subband; select the target energy from the howling subband energies of each howling subband, and use the target howling subband corresponding to the target energy as the target subband.

[0160] Among them, the howling subband refers to the subband containing the howling signal. The howling subband energy refers to the energy corresponding to the howling subband. The target energy refers to the largest howling subband energy. The target howling subband refers to the howling subband corresponding to the largest howling subband energy.

[0161] Feasibly, the terminal obtains the current howling detection result corresponding to the current audio signal. When the current howling detection result indicates that there is a howling signal in the current audio signal, the sub-band corresponding to the howling signal is determined from each sub-band according to the frequency of the howling signal and the frequency of the speech signal, so as to obtain each howling sub-band. Then, the energy corresponding to each howling sub-band is determined according to the energy of each sub-band. Then, the energies of each howling sub-band are compared, and the maximum howling sub-band energy is selected as the target energy, and the target howling sub-band corresponding to the target energy is used as the target sub-band.

[0162] In one embodiment, each howling sub-band corresponding to each howling sub-band energy can be directly used as the target sub-band, that is, the sub-band gain coefficient corresponding to each howling sub-band is calculated, and the historical sub-band gain corresponding to each howling sub-band is obtained. The product of the sub-band gain coefficient and the historical sub-band gain is calculated to obtain the current sub-band gain corresponding to each howling sub-band. Based on each current sub-band gain, howling suppression is performed on each howling sub-band to obtain the first target audio signal.

[0163] In the above embodiment, each howling sub-band is determined from each sub-band through the current howling detection result, and then the target sub-band is determined from each howling sub-band, which improves the accuracy of obtaining the target sub-band.

[0164] In one embodiment, as Figure 7 shown, step 206 of determining the sub-band gain coefficient corresponding to the current audio signal based on the current howling detection result and the current speech detection result includes:

[0165] Step 702, when the current speech detection result indicates that there is no speech signal in the current audio signal and the current howling detection result indicates that there is a howling signal in the current audio signal, obtain a preset decreasing coefficient, and use the preset decreasing coefficient as the sub-band gain coefficient corresponding to the current audio signal;

[0166] Among them, the preset decreasing coefficient refers to a coefficient that is preset to make the sub-band gain decrease. It can be a value less than 1. There is a howling signal in the current audio signal

[0167] Feasibly, when the terminal detects that there is no speech signal in the current audio signal and there is a howling signal in the current audio signal, it obtains a preset decreasing coefficient and uses the preset decreasing coefficient as the sub-band gain coefficient corresponding to the current audio signal. That is, when it is detected that there is no speech signal in the current audio signal and there is a howling signal in the current audio signal, the sub-band gain needs to be gradually decreased from the initial value until there is no howling signal in the current audio signal or the sub-band gain of the current audio signal reaches the preset lower limit value. For example, 0.08.

[0168] Step 704, when the current voice detection result is that the current audio signal contains a voice signal and the current howling detection result is that the current audio signal contains a howling signal, obtain a preset first increment coefficient, and use the preset increment coefficient as the sub-band gain coefficient corresponding to the current audio signal;

[0169] Step 706, when the current howling detection result is that the current audio signal does not contain a howling signal, obtain a preset second increment coefficient, and use the preset second increment coefficient as the sub-band gain coefficient corresponding to the current audio signal, where the preset first increment coefficient is greater than the preset second increment coefficient.

[0170] Among them, the preset first increment coefficient refers to a coefficient that is pre-set to increase the sub-band gain when the current audio signal contains a voice signal and a howling signal. The preset second increment coefficient refers to a coefficient that is pre-set to increase the sub-band gain when the current audio signal does not contain a howling signal. The preset first increment coefficient is greater than the preset second increment coefficient.

[0171] Feasibly, when the terminal detects that the current audio signal contains a voice signal and a howling signal, use the preset increment coefficient as the sub-band gain coefficient corresponding to the current audio signal. At this time, in order to protect the quality of the voice signal, it is necessary to rapidly increase the sub-band gain so that the sub-band gain can be restored to the initial value. When the terminal detects that the current audio signal does not contain a howling signal, use the preset second increment coefficient as the sub-band gain coefficient corresponding to the current audio signal. At this time, restore the sub-band gain of the current audio signal to the initial value according to the preset second increment coefficient. Among them, the preset first increment coefficient is greater than the preset second increment coefficient, indicating that the speed at which the sub-band gain is restored to the initial value when the current audio signal contains a voice signal and a howling signal is greater than the restoration speed when the current audio signal does not contain a howling signal. For example, in a voice call, every 20 ms, the current audio signal is obtained, and the sub-band gain of the current audio signal is calculated. At the start of the voice call, generally there is no howling signal, so the sub-band gain remains unchanged. Then when it is detected that there is a howling signal and no voice signal, the initial value of the sub-band gain of the current audio signal is decreased according to the preset decrement coefficient, and then when it is detected that there is a howling signal and a voice signal, the sub-band gain of the current audio signal is calculated according to the preset first increment coefficient, that is, the sub-band gain of the current audio signal is rapidly increased to restore the sub-band gain to the initial value.

[0172] In the above embodiment, the sub-band gain coefficient is determined according to the current voice detection result and the howling detection result, so that the obtained sub-band gain coefficient can be more accurate, so that howling suppression is more accurate, and further improves the quality of the obtained first target audio signal.

[0173] In one embodiment, as Figure 8 shown, the howling suppression method further includes:

[0174] Step 802: Determine a target low-frequency signal and a target high-frequency signal from the current audio signal based on a preset low-frequency range.

[0175] The preset low-frequency range refers to the pre-set frequency range of human voices. For example, it is less than 1400 HZ. The target low-frequency signal refers to the audio signal within the preset low-frequency range in the current audio signal, and the target high-frequency signal refers to the audio signal exceeding the preset low-frequency range in the current audio signal.

[0176] Feasibly, the terminal divides the current audio signal according to the preset low-frequency range to obtain the target low-frequency signal and the target high-frequency signal. For example, the audio signal less than 1400 HZ in the current audio signal is used as the target low-frequency signal, and the audio signal exceeding 1400 HZ in the current audio signal is used as the target high-frequency signal.

[0177] Step 804: Calculate the low-frequency energy corresponding to the target low-frequency signal, and smooth the low-frequency energy to obtain the smoothed low-frequency energy.

[0178] The low-frequency energy refers to the energy corresponding to the target low-frequency signal.

[0179] Feasibly, the terminal directly calculates the low-frequency energy corresponding to the target low-frequency signal, or divides the target low-frequency signal to obtain sub-bands of each low-frequency signal, then calculates the energy corresponding to each sub-band of the low-frequency signal, and then calculates the sum of the energies corresponding to each sub-band of the low-frequency signal to obtain the low-frequency energy corresponding to the target low-frequency signal. Then, the low-frequency energy is smoothed to obtain the smoothed low-frequency energy. Among them, the following formula (1) can be used for smoothing processing.

[0180] E v (t) = a * E v (t - 1) + (1 - a) * E c Formula (1)

[0181] Among them, E v (t) refers to the smoothed low-frequency energy corresponding to the target low-frequency signal in the current audio signal corresponding to the current time period. E v (t - 1) refers to the historical low-frequency energy corresponding to the historical low-frequency signal in the historical audio signal corresponding to the previous historical time period. E c The low-frequency energy corresponding to the target low-frequency signal in the current audio signal corresponding to the current time period. a refers to the smoothing coefficient, which is pre-set. Among them, E c Greater than E v (t - 1), the value of a can be the same as when E c Less than E v(t - 1) has different values of a, which is used to better track the rising and falling segments of energy.

[0182] Step 806: Divide the target high-frequency signal to obtain each high-frequency sub-band, and calculate the high-frequency sub-band energy corresponding to each high-frequency sub-band.

[0183] Feasibly, the terminal can divide the target high-frequency signal to obtain each high-frequency sub-band, and use a triangular filter to calculate the high-frequency sub-band energy corresponding to each high-frequency sub-band.

[0184] Step 808: Obtain the preset energy upper limit weight corresponding to each high-frequency sub-band, and calculate the high-frequency sub-band upper limit energy corresponding to each high-frequency sub-band based on the preset energy upper limit weight corresponding to each high-frequency sub-band and the smoothed low-frequency energy.

[0185] Among them, the preset energy upper limit weight refers to the pre-set energy upper limit weight of the high-frequency sub-band. Different high-frequency sub-bands have different preset energy upper limit weights, and the energy upper limit weights of the high-frequency sub-bands can be set to decrease in order of increasing frequency. The high-frequency sub-band upper limit energy refers to the upper limit of the high-frequency sub-band energy, and the energy of the high-frequency sub-band cannot exceed this upper limit.

[0186] Feasibly, the terminal obtains the preset energy upper limit weight corresponding to each high-frequency sub-band, and calculates the product of the preset energy upper limit weight corresponding to the high-frequency sub-band and the smoothed low-frequency energy to obtain the high-frequency sub-band upper limit energy corresponding to each high-frequency sub-band. The high-frequency sub-band upper limit energy can be calculated using formula (2).

[0187] E u (k) = E v (t) * b(k) Formula (2)

[0188] Among them, k refers to the k-th high-frequency sub-band, which is a positive integer, and E u (k) is the high-frequency sub-band upper limit energy corresponding to the k-th high-frequency sub-band. E v (t) refers to the smoothed low-frequency energy corresponding to the target low-frequency signal, and b(k) refers to the preset energy upper limit weight corresponding to the k-th high-frequency sub-band. For example, the preset energy upper limit weights of each high-frequency sub-band can be (0.8, 0.7, 0.6,...) in sequence.

[0189] Step 810: Calculate the ratio of the high-frequency sub-band upper limit energy to the high-frequency sub-band energy to obtain the high-frequency sub-band upper limit gain.

[0190] Among them, the high-frequency sub-band upper limit gain refers to the upper limit gain corresponding to sub-band gain of the high-frequency sub-band, that is, the sub-band gain of the high-frequency sub-band cannot exceed the high-frequency sub-band upper limit gain.

[0191] Feasibly, the terminal calculates the ratio of the upper limit energy of each high-frequency subband to the energy of the corresponding high-frequency subband respectively, and obtains the upper limit gain of each high-frequency subband. For example, the upper limit gain of the high-frequency subband can be calculated using formula (3).

[0192]

[0193] Among them, E(k) refers to the high-frequency subband energy corresponding to the kth high-frequency subband. E u (k) refers to the upper limit energy of the kth high-frequency subband. M(k) refers to the upper limit gain of the kth high-frequency subband.

[0194] Step 812: Calculate the gains of each high-frequency subband, determine the target gain of each high-frequency subband based on the upper limit gains of each high-frequency subband and the gains of each high-frequency subband, and perform howling suppression on each high-frequency subband based on the target gains of each high-frequency subband to obtain the second target audio signal corresponding to the current time period.

[0195] Among them, the high-frequency subband gain is calculated based on the high-frequency subband gain coefficient and the historical high-frequency subband gain. The high-frequency subband gain coefficient is determined according to the current howling detection result and the current voice detection result. The historical high-frequency subband gain refers to the gain of the high-frequency subband corresponding to the historical audio signal in the historical time period. The high-frequency subband target gain refers to the gain used for howling suppression. The second target audio signal refers to the audio signal obtained after performing howling suppression on all high-frequency subbands.

[0196] Feasibly, the terminal obtains the historical high-frequency subband gains corresponding to each historical high-frequency subband, determines the high-frequency subband gain coefficients according to the current howling detection result and the current voice detection result, and calculates the product of each historical high-frequency subband gain and each high-frequency subband gain coefficient respectively to obtain the high-frequency subband gains corresponding to each high-frequency subband. Compare the upper limit gains of each high-frequency subband with the corresponding high-frequency subband gains respectively, and select the smaller gain between the upper limit gain of the high-frequency subband and the high-frequency subband gain as the high-frequency subband target gain. For example, formula (4) can be used to select the high-frequency subband target gain.

[0197] B(k) = min[G(k), M(k)] Formula (4)

[0198] Among them, B(k) refers to the high-frequency subband target gain corresponding to the kth high-frequency subband, G(k) refers to the high-frequency subband gain corresponding to the kth high-frequency subband, and M(k) refers to the upper limit gain of the kth high-frequency subband. Then the terminal performs howling suppression on each high-frequency subband using the high-frequency subband target gains, and converts the frequency-domain audio signals corresponding to the howling-suppressed high-frequency subbands into time-domain audio signals to obtain the second target audio signal corresponding to the current time period.

[0199] In a specific embodiment, as Figure 8a shown, it is a schematic diagram of the energy constraint curve. In this curve schematic diagram, the abscissa represents the frequency and the ordinate represents the energy. Different sub-bands are obtained based on the frequency division. There are 9 sub-bands shown in the figure. The sub-bands with a frequency lower than 1400 HZ are low-frequency bands, and the sub-bands with a frequency higher than 1400 HZ are high-frequency bands. The low-frequency bands are the 1st to the 4th sub-bands, and the high-frequency bands are the 5th to the 9th sub-bands. Among them, curve C is the energy curve when there is only a voice signal in the audio signal. Curve B refers to the energy constraint curve for high-frequency signals. Curve A refers to the energy curve when the audio signal contains a voice signal and a howling signal. Obviously, it can be seen that when there is a voice signal in the low-frequency band, that is, from the 1st to the 4th sub-bands, no energy constraint is performed. In the high-frequency band, that is, after the 4th sub-band, when there is a howling signal, the energy of the audio signal needs to be constrained below curve B to obtain the audio signal after howling suppression.

[0200] In the above embodiment, by using the high-frequency sub-band upper limit gain to constrain the high-frequency sub-band energy of the high-frequency sub-band, the quality of the obtained second target audio signal is guaranteed.

[0201] In a specific embodiment, as Figure 9 shown, the howling suppression method includes the following steps:

[0202] Step 902, collect the initial audio signal corresponding to the current time period through a microphone, and perform echo cancellation on the initial audio signal to obtain the initial audio signal after echo cancellation.

[0203] Step 902, input the initial audio signal after echo cancellation into a voice activity detection model for detection to obtain the current voice detection result. Based on the current voice detection result, perform noise suppression on the initial audio signal after echo cancellation to obtain the initial audio signal after noise suppression.

[0204] Step 902, extract the initial audio features corresponding to the initial audio signal after noise suppression, obtain the first historical audio signal corresponding to the first historical time period, and extract the first historical audio features corresponding to the first historical audio signal. Calculate the first similarity between the initial audio features and the first historical audio features, and determine the current howling detection result based on the first similarity.

[0205] Step 902, when the current howling detection result indicates that there is a howling signal in the initial audio signal after noise suppression, perform a frequency domain transformation on the current audio signal to obtain a frequency domain audio signal;

[0206] Step 902: Divide the frequency-domain audio signal according to the preset number of subbands to obtain each subband, calculate the subband energy corresponding to each subband, smooth the subband energy of each subband, obtain the smoothed subband energy of each subband, and determine the target subband based on the smoothed subband energy of each subband.

[0207] Step 902: When the current voice detection result is that the current audio signal does not contain a voice signal and the current howling detection result is that the current audio signal contains a howling signal, obtain a preset decreasing coefficient, and use the preset decreasing coefficient as the subband gain coefficient corresponding to the current audio signal.

[0208] Step 902: Obtain the historical subband gain corresponding to the audio signal in the historical time period, and calculate the current subband gain corresponding to the current audio signal based on the subband gain coefficient and the historical subband gain.

[0209] Step 902: Perform howling suppression on the target subband based on the current subband gain to obtain the first target audio signal corresponding to the current time period, and send the first target audio signal corresponding to the current time period to the terminal that receives the first target audio signal through the network.

[0210] The present application also provides an application scenario that applies the above howling suppression method. Feasibly, the application of the howling suppression method in this application scenario is as follows:

[0211] When a voice conference is held in the enterprise WeChat application, as Figure 10 shown, it is a specific scenario application diagram of the howling suppression method. Among them, terminals 1002 and 1004 are in the same room and conduct a voip (Voice over Internet Protocol, voice transmission based on IP) call with other terminals. At this time, the voice collected by the microphone of terminal 1002 will be sent to terminal 1004 through the network, and after being played by the speaker of terminal 1004, the microphone of terminal 1002 will collect this voice again. Therefore, an acoustic loop is formed, and this cycle repeats, producing an acoustic effect of "howling".

[0212] When performing howling suppression at this time, a schematic diagram of the architecture of a howling suppression method is provided, as Figure 11 shown, where all terminals process the audio signal collected by the microphone through upstream audio processing and then encode and send it through the network. The audio signal obtained from the network interface is processed through downstream audio processing and then played.

[0213] Feasibly, the sound collected by the terminal 1002 through the microphone will be encoded and sent to the network side after being processed by the uplink audio processing to form a network signal. The uplink audio processing includes performing echo cancellation on the audio signal, and performing voice activity detection on the audio signal after echo cancellation, that is, analyzing and identifying non-voice signals and voice signals in the audio signal. Noise printing is performed on the non-voice signal to obtain the audio signal after noise suppression. Then, howling detection is performed on the audio signal after noise suppression to obtain the howling detection result. Howling suppression is performed according to the howling detection result and the voice activity detection result to obtain the voice signal after howling suppression, and the voice signal after howling suppression is volume-controlled and then encoded and sent.

[0214] Among them, when performing howling suppression, as Figure 12 shown, it is the flowchart for performing howling suppression. The terminal 1002 performs signal analysis on the audio signal that needs to be suppressed by howling, that is, transforms the time domain to the frequency domain to obtain the audio signal after frequency domain transformation. Then, the energy of each subband is calculated according to the preset number of subbands and the subband frequency range for the audio signal after frequency domain transformation. Then, the energy of each subband is smoothed in time to obtain the smoothed energy of each subband. The largest smoothed subband energy is selected from the smoothed energy of each subband as the target subband. Based on the howling detection result and the voice detection result, the subband gain coefficient corresponding to the audio signal is determined. Specifically, howlFlag represents the howling detection result. When howlFlag is 1, it indicates that there is a howling signal in the audio signal. When howlFlag is 0, it indicates that there is no howling signal in the audio signal. When VAD is 1, it indicates that the audio signal includes a voice signal. When VAD is 0, it indicates that the audio signal does not include a voice signal. When howlFlag is 1 and VAD is 0, the preset decreasing coefficient is obtained as the subband gain coefficient. When howlFlag is 1 and VAD is 1, the preset first increasing coefficient is obtained as the subband gain coefficient. When howlFlag is 0, the preset second increasing coefficient is obtained as the subband gain coefficient. At the same time, the historical subband gain used when the previous audio signal was processed for howling is obtained, and the product of the historical subband gain and the subband gain coefficient is calculated to obtain the current subband gain. The current subband gain is used to suppress howling for the target subband to obtain the audio signal after howling suppression, and then the audio signal after howling suppression is sent from the network side.

[0215] Meanwhile, when performing howling suppression, the target low-frequency signal and the target high-frequency signal can also be determined from the current audio signal based on a preset low-frequency range; calculate the low-frequency energy corresponding to the target low-frequency signal, smooth the low-frequency energy to obtain the smoothed low-frequency energy; divide the target high-frequency signal to obtain each high-frequency subband, and calculate the high-frequency subband energy corresponding to each high-frequency subband; obtain the preset energy upper limit weight corresponding to each high-frequency subband, and calculate the high-frequency subband upper limit energy corresponding to each high-frequency subband based on the preset energy upper limit weight corresponding to each high-frequency subband and the smoothed low-frequency energy; calculate the ratio of the high-frequency subband upper limit energy to the high-frequency subband energy to obtain the high-frequency subband upper limit gain corresponding to each high-frequency subband; calculate the high-frequency subband gain corresponding to each high-frequency subband, determine the target gain of each high-frequency subband based on the high-frequency subband upper limit gain and the high-frequency subband gain corresponding to each high-frequency subband, perform howling suppression on each high-frequency subband based on the target gain of each high-frequency subband to obtain the second target audio signal corresponding to the current time period, and send the second target audio signal through the network side. When the terminal 1004 receives a network signal through the network interface, it decodes the network signal to obtain an audio signal, and then performs downlink audio processing and then audio playback. The downlink audio processing may be volume control and so on. Similarly, the uplink audio processing in the terminal 1004 can also use the same method to process the audio and then send it through the network side.

[0216] In a specific embodiment, an architecture schematic diagram of another howling suppression method is provided, as Figure 13 shown. Specifically:

[0217] As Figure 10 shown, when the terminal 1002 sends an audio signal to each terminal, since the terminal 1002 is relatively close to the terminal 1004, howling may occur. For other terminals, including the terminal 1008, the terminal 1010, and the terminal 1012, which are relatively far from the terminal 1002 and will not generate howling. At this time, howling suppression can be performed in the terminal that receives the audio signal. Specifically:

[0218] When the terminal 1004 receives the network signal sent by the terminal 1002 through the network interface, it decodes the signal to obtain an audio signal. Generally, this audio signal is the one that has undergone echo cancellation and noise suppression at the sending terminal. At this time, the terminal 1004 directly performs howling detection and voice endpoint detection on the audio signal to obtain the howling detection result and the voice endpoint detection result. Moreover, the terminal 1004 collects historical audio signals of the same time length through the microphone for local detection, which is used to detect whether the terminal 1004 and the terminal 1002 are close. Specifically: by extracting the audio features of the audio signal collected through the microphone for the same time length, and extracting the audio features of the audio signal received from the network side, and then calculating the similarity. When the similarity exceeds the pre-set similarity threshold for a period of time, it indicates that the terminal 1004 and the terminal 1002 are close, and the local detection result is that the terminal 1004 and the terminal 1002 are close, indicating that the terminal 1004 is the terminal on the audio loop causing howling. At this time, howling suppression is performed according to the local detection result, the howling detection result, and the voice endpoint detection result, that is, the process as Figure 12 is executed to suppress howling, and the audio signal after howling suppression is obtained. Then the terminal 1004 plays the audio signal after howling suppression. In one embodiment, when the howling detection result indicates that the possibility of a howling signal existing in the audio signal exceeds the pre-set local detection pause threshold, the operation of local detection is paused, and howling suppression is only performed according to the howling detection result and the voice endpoint detection result, saving terminal resources.

[0219] By performing howling suppression in the terminal receiving the audio signal, the quality of the audio signal received by other terminals receiving the audio is ensured. And by performing howling suppression according to the local detection result, the howling detection result, and the voice endpoint detection result, the accuracy of howling suppression is improved. Similarly, the method for processing the downstream audio of the terminal 1004, that is, the above process for howling processing of audio information, can also be applied to the downstream audio processing of other terminals, such as the terminal 1002.

[0220] In a specific embodiment, as Figure 14 shown, another schematic diagram of the architecture of the howling suppression method is also provided. Specifically:

[0221] The terminal 1002 collects the current audio signal through the microphone. After performing echo cancellation and noise printing on the current audio signal, howling detection is performed to obtain the current howling detection result. When the current howling detection result indicates that a howling signal exists in the current audio signal, the audio signal to be played and the preset audio watermark signal are obtained. The preset audio watermark signal is added to the audio signal to be played and played through the speaker. At the same time, the current audio signal is volume-controlled and encoded into a network signal, and sent to the terminal 1004 through the network interface.

[0222] At this time, the terminal 1004 collects the audio signal played by the speaker of the terminal 1002 through the microphone, and then performs watermark detection, that is, calculates the adjacent band energy ratio of the collected audio signal. When the adjacent band energy ratio exceeds the preset adjacent band energy ratio threshold, it is determined that the collected audio signal contains the set audio watermark signal. At this time, the terminal 1004 obtains the network signal sent by the terminal 1002, decodes it to obtain the audio signal, and suppresses the howling of the audio signal, that is, executes as Figure 12 the process shown, obtains the howling-suppressed audio signal, and plays the howling-suppressed audio signal through the speaker. By adding an audio watermark signal to the audio signal played by the transmitting terminal, since the terminal generating the howling is relatively close, the receiving terminal will collect the audio signal with the added audio watermark signal through the microphone, perform watermark detection on the collected audio signal, and then perform howling suppression, improving the efficiency and accuracy of howling suppression. Similarly, when the terminal 1004 sends an audio signal through the network side, an audio watermark signal can also be added, and then the terminal 1002 can also perform watermark detection to determine whether to perform howling suppression on the received audio signal.

[0223] It should be understood that although Figure 2 , Figures 3 - 8 and Figure 9 the steps in the flowcharts of Figure 2 , Figures 3 - 8 and Figure 9 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,

[0224] In one embodiment, as Figure 15 shown, a howling suppression device 1500 is provided. This device can be a software module or a hardware module, or a combination of both to become a part of a computer device. Specifically, this device includes: a signal transformation module 1502, a sub-band determination module 1504, a coefficient determination module 1506, a gain determination module 1508, and a howling suppression module 1510, where:

[0225] The signal transformation module 1502 is configured to obtain the current audio signal corresponding to the current time period, perform a frequency domain transformation on the current audio signal, and obtain a frequency domain audio signal;

[0226] A sub - band determination module 1504, configured to divide a frequency - domain audio signal to obtain each sub - band, and determine a target sub - band from each sub - band;

[0227] A coefficient determination module 1506, configured to obtain a current howling detection result and a current voice detection result corresponding to a current audio signal, and determine a sub - band gain coefficient corresponding to the current audio signal based on the current howling detection result and the current voice detection result;

[0228] A gain determination module 1508, configured to obtain a historical sub - band gain corresponding to an audio signal in a historical time period, and calculate a current sub - band gain corresponding to the current audio signal based on the sub - band gain coefficient and the historical sub - band gain;

[0229] A howling suppression module 1510, configured to perform howling suppression on the target sub - band based on the current sub - band gain to obtain a first target audio signal corresponding to the current time period.

[0230] In one embodiment, the signal transformation module 1502 includes:

[0231] An echo cancellation unit, configured to collect an initial audio signal corresponding to the current time period, perform echo cancellation on the initial audio signal, and obtain an initial audio signal after echo cancellation;

[0232] A voice detection unit, configured to perform voice endpoint detection on the initial audio signal after echo cancellation to obtain a current voice detection result;

[0233] A noise suppression unit, configured to perform noise suppression on the initial audio signal after echo cancellation based on the current voice detection result to obtain an initial audio signal after noise suppression;

[0234] A howling detection unit, configured to perform howling detection on the initial audio signal after noise suppression to obtain a current howling detection result;

[0235] A current audio signal determination unit, configured to use the initial audio signal after noise suppression as the current audio signal corresponding to the current time period when the current howling detection result indicates that there is a howling signal in the initial audio signal after noise suppression.

[0236] In one embodiment, the voice detection unit is further configured to input the initial audio signal after echo cancellation into a voice endpoint detection model for detection to obtain a current voice detection result, and the voice endpoint detection model is trained using a neural network algorithm based on training audio signals and corresponding training voice detection results.

[0237] In one embodiment, the voice detection unit is further configured to perform low-pass filtering on the initial audio signal after echo cancellation to obtain a low-frequency signal; calculate the signal energy corresponding to the low-frequency signal, calculate the energy fluctuation based on the signal energy, and determine the current voice detection result according to the energy fluctuation.

[0238] In one embodiment, the voice detection unit is further configured to perform low-pass filtering on the initial audio signal after echo cancellation to obtain a low-frequency signal; perform pitch detection on the low-frequency signal to obtain a pitch period, and determine the current voice detection result according to the pitch period.

[0239] In one embodiment, the howling detection unit is further configured to input the initial audio signal after noise suppression into a howling detection model for detection to obtain the current howling detection result, where the howling detection model is trained using a neural network algorithm based on howling training audio signals and corresponding training howling detection results.

[0240] In one embodiment, the howling detection unit is further configured to extract initial audio features corresponding to the initial audio signal after noise suppression; obtain the first historical audio signal corresponding to the first historical time period, and extract the first historical audio features corresponding to the first historical audio signal; calculate the first similarity between the initial audio features and the first historical audio features, and determine the current howling detection result based on the first similarity.

[0241] In one embodiment, the howling suppression method further includes:

[0242] A watermark addition module, configured to, when the current howling detection result indicates that there is a howling signal in the current audio signal, obtain the audio signal to be played and a preset audio watermark signal; add the preset audio watermark signal to the audio signal to be played and play it.

[0243] In one embodiment, the howling suppression method further includes:

[0244] A watermark detection module, configured to collect the first audio signal corresponding to the first time period, detect the first audio signal based on an audio watermark detection algorithm, and determine that the first audio signal contains a target audio watermark signal;

[0245] A signal obtaining module, configured to receive the target network-coded audio signal corresponding to the second time period, decode the target network-coded audio signal to obtain a target network audio signal;

[0246] A current audio signal determination module, configured to use the target network audio signal as the current audio signal based on the fact that the first audio signal contains a target audio watermark signal.

[0247] In one embodiment, the signal transformation module 1502 includes:

[0248] A network signal obtaining module, configured to receive a current network-coded audio signal corresponding to a current time period, decode the network-coded audio signal to obtain a current network audio signal;

[0249] A network signal detection module, configured to perform voice endpoint detection on the current network audio signal to obtain a network voice detection result, and perform howling detection on the current network audio signal to obtain a network howling detection result;

[0250] A feature extraction module, configured to extract network audio features of the current network audio signal, obtain a second historical audio signal in a second historical time period, and extract second historical audio features corresponding to the second historical audio signal;

[0251] A current audio signal obtaining module, configured to calculate a network audio similarity between the network audio features and the second historical audio features, and determine the network audio signal as the current audio signal corresponding to the current time period based on the network audio similarity and the network howling detection result.

[0252] In one embodiment, the sub-band determination module 1504 is further configured to divide the frequency-domain audio signal according to a preset number of sub-bands to obtain each sub-band; calculate sub-band energies corresponding to each sub-band, and smooth the sub-band energies of each sub-band to obtain the smoothed sub-band energies of each sub-band; determine a target sub-band based on the smoothed sub-band energies of each sub-band.

[0253] In one embodiment, the sub-band determination module 1504 is further configured to obtain a current howling detection result corresponding to the current audio signal, determine each howling sub-band from each sub-band according to the current howling detection result, and obtain howling sub-band energies of each howling sub-band; select a target energy from the howling sub-band energies of each howling sub-band, and use the target howling sub-band corresponding to the target energy as the target sub-band.

[0254] In one embodiment, the coefficient determination module 1506 is further configured to, when the current voice detection result is that the current audio signal does not contain a voice signal and the current howling detection result is that the current audio signal contains a howling signal, obtain a preset decreasing coefficient, and use the preset decreasing coefficient as the sub-band gain coefficient corresponding to the current audio signal; when the current voice detection result is that the current audio signal contains a voice signal and the current howling detection result is that the current audio signal contains a howling signal, obtain a preset first increasing coefficient, and use the preset increasing coefficient as the sub-band gain coefficient corresponding to the current audio signal; when the current howling detection result is that the current audio signal does not contain a howling signal, obtain a preset second increasing coefficient, and use the preset second increasing coefficient as the sub-band gain coefficient corresponding to the current audio signal, where the preset first increasing coefficient is greater than the preset second increasing coefficient.

[0255] In one embodiment, the howling suppression method further includes:

[0256] A signal division module, configured to determine a target low-frequency signal and a target high-frequency signal from a current audio signal based on a preset low-frequency range;

[0257] A low-frequency energy calculation module, configured to calculate the low-frequency energy corresponding to the target low-frequency signal, smooth the low-frequency energy, and obtain the smoothed low-frequency energy;

[0258] A high-frequency energy calculation module, configured to divide the target high-frequency signal to obtain each high-frequency sub-band, and calculate the high-frequency sub-band energy corresponding to each high-frequency sub-band;

[0259] An upper limit energy calculation module, configured to obtain a preset energy upper limit weight corresponding to each high-frequency sub-band, and calculate the high-frequency sub-band upper limit energy corresponding to each high-frequency sub-band based on the preset energy upper limit weight corresponding to each high-frequency sub-band and the smoothed low-frequency energy;

[0260] An upper limit gain determination module, configured to calculate the ratio of the high-frequency sub-band upper limit energy to the high-frequency sub-band energy to obtain the upper limit gain of each high-frequency sub-band;

[0261] A target audio signal obtaining module, configured to calculate the high-frequency sub-band gain corresponding to each high-frequency sub-band, determine the target gain of each high-frequency sub-band based on the upper limit gain of each high-frequency sub-band and the high-frequency sub-band gain of each high-frequency sub-band, and perform howling suppression on each high-frequency sub-band based on the target gain of each high-frequency sub-band to obtain a second target audio signal corresponding to the current time period.

[0262] For the specific limitations of the howling suppression device, reference may be made to the limitations of the howling suppression method in the above text, which will not be elaborated here. Each module in the above howling suppression device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0263] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 16As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a howling suppression method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0264] Those skilled in the art can understand that Figure 16 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0265] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0266] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0267] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0268] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0269] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0270] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several variations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A howling suppression method, characterized in that, The method includes: Obtain a current audio signal corresponding to a current time period, perform a frequency-domain transformation on the current audio signal to obtain a frequency-domain audio signal; Divide the frequency-domain audio signal to obtain each sub-band, and determine a target sub-band from the sub-bands; Obtain a current howling detection result and a current speech detection result corresponding to the current audio signal, and determine a sub-band gain coefficient corresponding to the current audio signal based on the current howling detection result and the current speech detection result. The sub-band gain coefficient includes a preset decreasing coefficient, a preset first increasing coefficient, and a preset second increasing coefficient. The preset decreasing coefficient refers to a coefficient that is preset to decrease the sub-band gain when there is no speech signal in the current audio signal and there is a howling signal in the current audio signal. The preset first increasing coefficient refers to a coefficient that is preset to increase the sub-band gain when there is a speech signal and a howling signal in the current audio signal. The preset second increasing coefficient refers to a coefficient that is preset to increase the sub-band gain when there is no howling signal in the current audio signal. The preset first increasing coefficient is greater than the preset second increasing coefficient; Obtain a historical sub-band gain corresponding to an audio signal in a historical time period, and calculate a current sub-band gain corresponding to the current audio signal based on the sub-band gain coefficient and the historical sub-band gain; Perform howling suppression on the target sub-band based on the current sub-band gain to obtain a first target audio signal corresponding to the current time period.

2. The method according to claim 1, characterized in that, The obtaining of the current audio signal corresponding to the current time period includes: Collect an initial audio signal corresponding to the current time period, perform echo cancellation on the initial audio signal to obtain an initial audio signal after echo cancellation; Perform voice activity detection on the initial audio signal after echo cancellation to obtain a current speech detection result; Perform noise suppression on the initial audio signal after echo cancellation based on the current speech detection result to obtain an initial audio signal after noise suppression; Perform howling detection on the initial audio signal after noise suppression to obtain a current howling detection result; When the current howling detection result indicates that there is a howling signal in the initial audio signal after noise suppression, use the initial audio signal after noise suppression as the current audio signal corresponding to the current time period.

3. The method according to claim 2, wherein The performing of voice activity detection on the initial audio signal after echo cancellation to obtain a current speech detection result includes: Perform low-pass filtering on the initial audio signal after echo cancellation to obtain a low-frequency signal; Calculate the signal energy corresponding to the low-frequency signal, calculate the energy fluctuation based on the signal energy, and determine the current speech detection result according to the energy fluctuation.

4. The method according to claim 2, wherein The performing of voice activity detection on the initial audio signal after echo cancellation to obtain a current speech detection result includes: Perform low-pass filtering on the initial audio signal after echo cancellation to obtain a low-frequency signal; Perform pitch detection on the low-frequency signal to obtain a pitch period, and determine the current speech detection result according to the pitch period.

5. The method according to claim 2, characterized in that Performing howling detection on the initial audio signal after the noise suppression to obtain the current howling detection result, including: Extracting the initial audio features corresponding to the initial audio signal after the noise suppression; Obtaining the first historical audio signal corresponding to the first historical time period, and extracting the first historical audio features corresponding to the first historical audio signal; Calculating a first similarity between the initial audio features and the first historical audio features, and determining the current howling detection result based on the first similarity.

6. The method according to claim 2, characterized in that, The method further includes: When the current howling detection result indicates that there is a howling signal in the current audio signal, obtaining the audio signal to be played and a preset audio watermark signal; Adding the preset audio watermark signal to the audio signal to be played and playing the result.

7. The method according to claim 6, wherein The method further includes: Collecting a first audio signal corresponding to a first time period, detecting the first audio signal based on an audio watermark detection algorithm, and determining that the first audio signal contains a target audio watermark signal; Receiving a target network-coded audio signal corresponding to a second time period, decoding the target network-coded audio signal to obtain a target network audio signal; Using the target network audio signal as the current audio signal based on the fact that the first audio signal contains a target audio watermark signal.

8. The method according to claim 1, characterized in that The obtaining of the current audio signal corresponding to the current time period includes: Receiving the current network-coded audio signal corresponding to the current time period, decoding the network-coded audio signal to obtain a current network audio signal; Performing voice activity detection on the current network audio signal to obtain a network voice detection result, and performing howling detection on the current network audio signal to obtain a network howling detection result; Extracting the network audio features of the current network audio signal, obtaining a second historical audio signal of a second historical time period, and extracting the second historical audio features corresponding to the second historical audio signal; Calculating a network audio similarity between the network audio features and the second historical audio features, and determining the network audio signal as the current audio signal corresponding to the current time period based on the network audio similarity and the network howling detection result.

9. The method according to claim 1, characterized in that The dividing of the frequency-domain audio signal to obtain each sub-band and determining a target sub-band from the each sub-band includes: Dividing the frequency-domain audio signal according to a preset number of sub-bands to obtain each sub-band; Calculating the sub-band energy corresponding to each sub-band, and smoothing the sub-band energy of each sub-band to obtain the smoothed sub-band energy of each sub-band; Determining a target sub-band based on the smoothed sub-band energy of each sub-band.

10. The method according to claim 9, wherein The determining of the target sub-band based on the smoothed sub-band energy of each sub-band includes: Obtaining the current howling detection result corresponding to the current audio signal, determining each howling sub-band from the each sub-band according to the current howling detection result, and obtaining the howling sub-band energy of each howling sub-band; Selecting a target energy from the howling sub-band energy of each howling sub-band, and using the target howling sub-band corresponding to the target energy as the target sub-band.

11. The method according to claim 1, wherein The method further includes: Determining a target low-frequency signal and a target high-frequency signal from the current audio signal based on a preset low-frequency range; Calculate the low-frequency energy corresponding to the target low-frequency signal, smooth the low-frequency energy to obtain the smoothed low-frequency energy; Divide the target high-frequency signal to obtain each high-frequency subband, and calculate the high-frequency subband energy corresponding to each high-frequency subband; Obtain the preset energy upper limit weight corresponding to each high-frequency subband, and calculate the high-frequency subband upper limit energy corresponding to each high-frequency subband based on the preset energy upper limit weight corresponding to each high-frequency subband and the smoothed low-frequency energy; Calculate the ratio of the high-frequency subband upper limit energy to the high-frequency subband energy to obtain the high-frequency subband upper limit gain for each high-frequency subband; Calculate the high-frequency subband gain corresponding to each high-frequency subband, determine the target high-frequency subband gain for each high-frequency subband based on the high-frequency subband upper limit gain and the high-frequency subband gain for each high-frequency subband, and perform howling suppression on each high-frequency subband based on the target high-frequency subband gain for each high-frequency subband to obtain the second target audio signal corresponding to the current time period.

12. A howling suppression device, characterized in that, The device includes: A signal transformation module, configured to obtain a current audio signal corresponding to a current time period, perform a frequency domain transformation on the current audio signal to obtain a frequency domain audio signal; A subband determination module, configured to divide the frequency domain audio signal to obtain each subband, and determine a target subband from each subband; A coefficient determination module, configured to obtain a current howling detection result and a current voice detection result corresponding to the current audio signal, and determine a subband gain coefficient corresponding to the current audio signal based on the current howling detection result and the current voice detection result. The subband gain coefficient includes a preset decreasing coefficient, a preset first increasing coefficient, and a preset second increasing coefficient. The preset decreasing coefficient refers to a coefficient that is preset to decrease the subband gain when the current audio signal does not contain a voice signal and there is a howling signal in the current audio signal. The preset first increasing coefficient refers to a coefficient that is preset to increase the subband gain when the current audio signal contains a voice signal and a howling signal. The preset second increasing coefficient refers to a coefficient that is preset to increase the subband gain when the current audio signal does not contain a howling signal. The preset first increasing coefficient is greater than the preset second increasing coefficient; A gain determination module, configured to obtain a historical subband gain corresponding to an audio signal in a historical time period, and calculate a current subband gain corresponding to the current audio signal based on the subband gain coefficient and the historical subband gain; A howling suppression module, configured to perform howling suppression on the target subband based on the current subband gain to obtain a first target audio signal corresponding to the current time period.

13. The device according to claim 12, characterized in that, The signal transformation module includes: An echo cancellation unit, configured to collect an initial audio signal corresponding to the current time period, perform echo cancellation on the initial audio signal to obtain an initial audio signal after echo cancellation; A voice detection unit, configured to perform voice endpoint detection on the initial audio signal after echo cancellation to obtain a current voice detection result; A noise suppression unit, configured to perform noise suppression on the initial audio signal after echo cancellation based on the current voice detection result to obtain an initial audio signal after noise suppression; A howling detection unit, configured to perform howling detection on the initial audio signal after noise suppression to obtain a current howling detection result; A current audio signal determination unit, configured to use the initial audio signal after noise suppression as the current audio signal corresponding to the current time period when the current howling detection result indicates that there is a howling signal in the initial audio signal after noise suppression.

14. The device according to claim 13, characterized in that, The voice detection unit is further configured to perform low-pass filtering on the initial audio signal after echo cancellation to obtain a low-frequency signal; calculate the signal energy corresponding to the low-frequency signal, calculate the energy fluctuation based on the signal energy, and determine the current voice detection result according to the energy fluctuation.

15. The device according to claim 13, characterized in that, The voice detection unit is further configured to perform low-pass filtering on the initial audio signal after echo cancellation to obtain a low-frequency signal; perform pitch detection on the low-frequency signal to obtain a pitch period, and determine the current voice detection result according to the pitch period.

16. The device according to claim 13, characterized in that, The voice detection unit is further configured to extract initial audio features corresponding to the initial audio signal after noise suppression; obtain a first historical audio signal corresponding to a first historical time period, and extract first historical audio features corresponding to the first historical audio signal; calculate a first similarity between the initial audio features and the first historical audio features, and determine the current howling detection result based on the first similarity.

17. The device according to claim 13, characterized in that, The apparatus further includes: A watermark addition module, configured to obtain a to-be-played audio signal and a preset audio watermark signal when the current howling detection result indicates that there is a howling signal in the current audio signal; add the preset audio watermark signal to the to-be-played audio signal and play the signal.

18. The device according to claim 17, characterized in that, The apparatus further includes: A watermark detection module, configured to collect a first audio signal corresponding to a first time period, detect the first audio signal based on an audio watermark detection algorithm, and determine that the first audio signal contains a target audio watermark signal; A signal obtaining module, configured to receive a target network-coded audio signal corresponding to a second time period, decode the target network-coded audio signal to obtain a target network audio signal; A current audio signal determination module, configured to use the target network audio signal as the current audio signal based on the fact that the first audio signal contains a target audio watermark signal.

19. The device according to claim 12, characterized in that The signal transformation module includes: A network signal obtaining module, configured to receive the current network-coded audio signal corresponding to the current time period, decode the network-coded audio signal to obtain a current network audio signal; A network signal detection module, configured to perform voice endpoint detection on the current network audio signal to obtain a network voice detection result, and perform howling detection on the current network audio signal to obtain a network howling detection result; A feature extraction module, configured to extract network audio features of the current network audio signal, and obtain a second historical audio signal of a second historical time period, and extract second historical audio features corresponding to the second historical audio signal; A current audio signal obtaining module, configured to calculate a network audio similarity between the network audio feature and the second historical audio feature, and determine the network audio signal as the current audio signal corresponding to the current time period based on the network audio similarity and the network howling detection result.

20. The device according to claim 12, characterized in that, The sub-band determination module is further configured to divide the frequency-domain audio signal according to a preset number of sub-bands to obtain each sub-band; calculate the sub-band energy corresponding to each sub-band, and smooth the sub-band energy of each sub-band to obtain the smoothed sub-band energy of each sub-band; determine a target sub-band based on the smoothed sub-band energy of each sub-band.

21. The device according to claim 20, characterized in that, The sub-band determination module is further configured to obtain a current howling detection result corresponding to the current audio signal, determine each howling sub-band from each of the sub-bands according to the current howling detection result, and obtain the sub-band energy of each howling sub-band; Select a target energy from the sub-band energies of each howling sub-band, and use the target howling sub-band corresponding to the target energy as the target sub-band.

22. The device according to claim 12, characterized in that, The apparatus further includes: A signal division module, configured to determine a target low-frequency signal and a target high-frequency signal from the current audio signal based on a preset low-frequency range; A low-frequency energy calculation module, configured to calculate the low-frequency energy corresponding to the target low-frequency signal, and smooth the low-frequency energy to obtain the smoothed low-frequency energy; A high-frequency energy calculation module, configured to divide the target high-frequency signal to obtain each high-frequency sub-band, and calculate the high-frequency sub-band energy corresponding to each high-frequency sub-band; An upper limit energy calculation module, configured to obtain a preset energy upper limit weight corresponding to each high-frequency sub-band, and calculate the high-frequency sub-band upper limit energy corresponding to each high-frequency sub-band based on the preset energy upper limit weight corresponding to each high-frequency sub-band and the smoothed low-frequency energy; An upper limit gain determination module, configured to calculate a ratio of the high-frequency sub-band upper limit energy to the high-frequency sub-band energy to obtain the upper limit gain of each high-frequency sub-band; A target audio signal obtaining module, configured to calculate the high-frequency sub-band gain corresponding to each high-frequency sub-band, determine the target gain of each high-frequency sub-band based on the upper limit gain of each high-frequency sub-band and the high-frequency sub-band gain of each high-frequency sub-band, and perform howling suppression on each high-frequency sub-band based on the target gain of each high-frequency sub-band to obtain the second target audio signal corresponding to the current time period.

23. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

24. A computer-readable storage medium stores a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.

25. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Method and device for debugging audio acquisition equipment

    CN110035374A

  • Howling detection method and equipment, storage medium and electronic equipment

    CN110148426A

  • Audio signal processing method, model training method and related device

    CN111210021A