Howling suppression method, device, chip and electronic equipment

By preprocessing the audio signal and training machine learning model processing, combined with howling residual detection, the problem of poor effect of existing howling suppression technology is solved, and efficient and accurate howling suppression effect is achieved, reducing residues and improving the quality of the audio signal.

CN114827833BActive Publication Date: 2025-08-22SPREADTRUM COMMUNICATION (SHANGHAI) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210386133.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-13
Publication Date
2025-08-22
Estimated Expiration
2042-04-13

AI Technical Summary

Technical Problem

The existing howling suppression technology is not effective in sound reinforcement systems, especially in scenarios such as multiplayer meetings and game team battles, which may damage the device. The suppression effect of machine learning solutions that rely on traditional methods does not meet the requirements.

Method used

By preprocessing the audio signal, using the trained machine learning model to suppress howling, combined with howling residual detection, it realizes rapid processing and residual detection of multi-frequency, broadband, and changing howling frequency points, and further suppression of howling residual points are used to judge the signal suppression ratio and peak near power ratio.

Benefits of technology

Improve the accuracy and efficiency of howling suppression, reduce the residue of howling, ensure the quality of audio signals, and avoid the limitations of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114827833B_ABST
    Figure CN114827833B_ABST
Patent Text Reader

Abstract

The present invention provides a howling suppression method, device, chip, and electronic device. The howling suppression method includes: preprocessing an input audio signal to obtain audio analysis data; inputting the audio analysis data into a trained machine learning model to perform machine howling suppression to obtain initial suppressed audio data; performing howling residual detection on the initial suppressed audio data to obtain a howling residual detection result; and processing the initial suppressed audio data based on the howling residual detection result to obtain target suppressed audio data. The howling suppression method of the present invention improves the howling suppression effect and accuracy, and improves suppression efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sound processing technology, and in particular to a howling suppression method, device, chip and electronic equipment. Background Art

[0002] In a sound reinforcement system, audio signals are amplified by speakers and picked up by microphones, forming a closed-loop feedback system. When the closed-loop gain of the system is greater than or equal to 1 and the phase is an integer multiple of 2π, howling occurs. With the advancement of mobile terminal technology, scenarios such as online multi-person meetings and team gaming are becoming increasingly popular. Howling can seriously affect voice quality, and excessive system gain can damage components. Therefore, howling suppression is crucial in sound reinforcement systems.

[0003] Research on howling suppression technology has a long history, primarily through disrupting the amplitude and phase conditions that produce howling. Traditional suppression methods include notch filtering, frequency shifting, and adaptive acoustic feedback suppression. The frequency shifting method directly alters the frequency components of the input signal to disrupt the conditions that produce howling, but this also damages normal speech. The adaptive acoustic feedback suppression method uses a filter to estimate the acoustic feedback path, but is affected by factors such as the sound field environment and signal coloration, placing high demands on filter length and convergence performance. Furthermore, in mobile terminal communications, it is impossible to simultaneously obtain both the reference signal and the desired signal. The notch filtering method is the most widely used and requires howling detection followed by notch filtering. The howling generated in real-world scenarios often has multi-frequency and broadband characteristics, placing high demands on detection accuracy and notch filter design, which can easily result in residual sound or speech distortion.

[0004] In recent years, machine learning technology has developed rapidly, achieving significant results in areas such as speech enhancement and speech recognition. Chip platforms and mobile terminals in the industry are already equipped with machine learning voice processing solutions. Research on howling suppression technology based on machine learning is still in its developmental stages. Existing solutions often use machine learning models to mark howling points for howling detection, and then further suppress it using traditional methods such as notch filtering. This approach replaces traditional howling detection methods with machine learning technology, improving howling detection accuracy. However, the effectiveness of the suppression still depends on the design of the subsequent suppression method, resulting in the howling suppression effect not meeting the requirements.

[0005] Therefore, it is necessary to provide a new howling suppression method, device, chip and electronic device to solve the above problems existing in the prior art. Summary of the Invention

[0006] The object of the present invention is to provide a howling suppression method, device, chip and electronic equipment, which improve the howling suppression effect and accuracy and improve the suppression efficiency.

[0007] In a first aspect, to achieve the above-mentioned objectives, the present invention provides a howling suppression method, the method comprising:

[0008] Preprocessing the input audio signal to obtain audio analysis data;

[0009] Inputting the audio analysis data into a trained machine learning model to perform machine howling suppression to obtain initial suppressed audio data;

[0010] performing a howling residual detection on the initial suppressed audio data to obtain a howling residual detection result;

[0011] The initial suppressed audio data is processed according to the howling residual detection result to obtain target suppressed audio data.

[0012] The beneficial effect of the howling suppression method described in the present invention is that: by preprocessing the input audio signal to obtain audio analysis data, and then performing machine howling processing through the trained machine learning model, it is possible to quickly process multi-frequency, broadband, and changing howling frequencies, thereby improving the howling processing efficiency, avoiding the limitations of notch filter suppression or other howling suppression methods, and being able to suppress howling in real time on the audio signal. At the same time, howling residual detection is performed on the initial suppressed audio data after processing by the machine learning model, so as to further suppress potential howling frequencies according to the howling residual detection results, effectively reducing the howling residual and improving the howling suppression effect.

[0013] Optionally, performing a howling residual detection on the initial suppressed audio data to obtain a howling residual detection result includes:

[0014] calculating a signal suppression ratio between the initial suppressed audio data and the audio analysis data;

[0015] After determining that the signal suppression ratio is lower than a first threshold, calculating a peak-to-peak power ratio of a frequency point in the initially suppressed audio data;

[0016] After determining that the peak-to-peak power ratio of the frequency point in the initial suppressed audio data is greater than a second threshold, determining that the frequency point is a howling residual point, and determining that howling residual exists in the initial suppressed audio data.

[0017] Optionally, the peak adjacent power ratio of the frequency point is the ratio of the peak power of the frequency point to the power of the mth adjacent frequency, where m is an integer.

[0018] Optionally, the absolute value of m is greater than 2 and does not exceed a preset threshold. The beneficial effect is that it facilitates the detection of broadband howling residuals.

[0019] Optionally, the processing the initial suppressed audio data to obtain target suppressed audio data according to the howling residual detection result includes:

[0020] using the frequency point determined as the howling residual point in the initial suppressed audio data as the target suppression frequency point;

[0021] A suppression coefficient of the target suppression frequency point is obtained, and the target suppression frequency point in the initial suppressed audio data is suppressed according to the suppression coefficient to obtain the target suppressed audio data.

[0022] Optionally, the suppression coefficient is a signal suppression ratio or a peak proximity power ratio.

[0023] Optionally, performing a howling residual detection on the initial suppressed audio data to obtain a howling residual detection result further includes:

[0024] After determining that the signal suppression ratio is greater than or equal to the first threshold, determining that no howling residual exists in the initial suppressed audio data;

[0025] The processing of the initial suppressed audio data to obtain target suppressed audio data according to the howling residual detection result includes:

[0026] After determining that no howling remains in the initial suppressed audio data, the initial suppressed audio data is used as the target suppressed audio data.

[0027] Optionally, the method further includes, after determining that the peak adjacent power ratio of the frequency point in the initial suppressed audio data is less than or equal to a second threshold, determining that there is no howling residual in the initial suppressed audio data, and using the initial suppressed audio data as the target suppressed audio data.

[0028] Optionally, preprocessing the input audio signal to obtain audio analysis data includes:

[0029] Performing time domain analysis or frequency domain analysis on the audio signal to obtain time domain features or frequency domain features, and using the audio signal with the time domain features or the frequency domain features as the audio analysis data.

[0030] Optionally, before performing time domain analysis or frequency domain analysis on the audio signal, the method further includes:

[0031] The audio signal is divided into sub-bands according to the frequency domain height of the audio signal. This has the beneficial effect of reducing the computational complexity of the subsequent machine learning model.

[0032] In a second aspect, the present invention further provides a howling suppression device, comprising:

[0033] A preprocessing module, used for preprocessing the input audio signal to obtain audio analysis data;

[0034] a machine learning model processing module, configured to input the audio analysis data into a trained machine learning model for machine howling suppression to obtain initial suppressed audio data;

[0035] a howling residual detection module, configured to perform a howling residual detection on the initial suppressed audio data to obtain a howling residual detection result;

[0036] The residual suppression module is configured to process the initial suppressed audio data according to the howling residual detection result to obtain target suppressed audio data.

[0037] The beneficial effects of the howling suppression device described in the present invention are: the input audio signal is preprocessed by the preprocessing module to obtain audio analysis data, and then the machine learning model processing module performs machine howling processing through the trained machine learning model, thereby realizing rapid processing of multi-frequency, broadband, and changing howling frequencies, improving the howling processing efficiency, and avoiding the limitations of notch filter suppression or other howling suppression methods. At the same time, the howling residual detection module performs howling residual detection on the initial suppressed audio data after processing by the machine learning model, so that the residual suppression module can further suppress potential howling frequencies according to the howling residual detection results, effectively reducing the howling residual and improving the howling suppression effect.

[0038] Optionally, the howling residual detection module is further used to:

[0039] calculating a signal suppression ratio between the initial suppressed audio data and the audio analysis data;

[0040] After determining that the signal suppression ratio is lower than a first threshold, calculating a peak-to-peak power ratio of a frequency point in the initially suppressed audio data;

[0041] After determining that the peak-to-peak power ratio of the frequency point in the initial suppressed audio data is greater than a second threshold, determining that the frequency point is a howling residual point, and determining that howling residual exists in the initial suppressed audio data.

[0042] Optionally, the residue suppression module is further configured to:

[0043] using the frequency point determined as the howling residual point in the initial suppressed audio data as the target suppression frequency point;

[0044] A suppression coefficient of the target suppression frequency point is obtained, and the target suppression frequency point in the initial suppressed audio data is suppressed according to the suppression coefficient to obtain the target suppressed audio data.

[0045] Optionally, the howling residual detection module is further used to:

[0046] After determining that the signal suppression ratio is greater than or equal to the first threshold, determining that no howling residual exists in the initial suppressed audio data;

[0047] The residual suppression module is further configured to:

[0048] After determining that no howling remains in the initial suppressed audio data, the initial suppressed audio data is used as the target suppressed audio data.

[0049] Optionally, the preprocessing module is further used to:

[0050] Performing time domain analysis or frequency domain analysis on the audio signal to obtain time domain features or frequency domain features, and using the audio signal with the time domain features or the frequency domain features as the audio analysis data.

[0051] In a third aspect, the present invention further discloses a chip, comprising a processor and a communication interface, wherein the processor is configured for the above-mentioned howling suppression method.

[0052] In a fourth aspect, the present invention further provides an electronic device, comprising: a processor and a memory;

[0053] The memory is used to store computer programs;

[0054] The processor is configured to execute the computer program stored in the memory, so as to enable the electronic device to perform the above-mentioned howling suppression method.

[0055] The beneficial effects of the chip and the electronic device of the present invention are similar to those of the aforementioned howling suppression method, and will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 Schematic diagram of the overall process of the howling suppression method according to an embodiment of the present invention;

[0057] Figure 2 Schematic diagram of the process of howling residual detection in the howling suppression method according to an embodiment of the present invention;

[0058] Figure 3 is a structural block diagram of the howling suppression device according to an embodiment of the present invention;

[0059] Figure 4 is a structural block diagram of the electronic device according to an embodiment of the present invention;

[0060] Figure 5 2 is a schematic diagram comparing the process of suppressing a howling speech signal by the howling suppression method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0061] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be the common meanings understood by people with ordinary skills in the field to which the present invention belongs. The words "including" and similar words used in this article mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects.

[0062] In view of the problems existing in the prior art, the embodiment of the present invention provides a method for suppressing howling. Figure 1 , the method comprises the following steps:

[0063] S101 : Preprocess an input audio signal to obtain audio analysis data.

[0064] In some embodiments, pre-processing the input audio signal to obtain audio analysis data includes:

[0065] Performing time domain analysis or frequency domain analysis on the audio signal to obtain time domain features or frequency domain features, and using the audio signal with the time domain features or the frequency domain features as the audio analysis data.

[0066] In some embodiments, before performing time domain analysis or frequency domain analysis on the audio signal, the audio signal is divided into sub-bands according to the high and low frequency domains of the audio signal.

[0067] For example, the input audio signal is first framed and then divided into sub-bands according to the frequency domain of the audio signal, so as to reduce the amount of calculation when the audio signal is subsequently processed, thereby improving the efficiency of howling suppression processing.

[0068] In yet other embodiments, since howling generally occurs in mid- and high-frequency regions of an audio signal, a coarse division is performed on the low-frequency sub-bands, while a fine division is performed on the mid- and high-frequency sub-bands, based on the frequency domain of the audio signal. Specifically, the division accuracy of the mid- and high-frequency sub-bands is three times that of the low-frequency sub-bands, to facilitate howling suppression processing in the finely divided mid- and high-frequency sub-bands.

[0069] After the audio signal is divided into frames, feature extraction and analysis are performed on the divided sub-bands, including time domain analysis and frequency domain analysis.

[0070] When performing time domain analysis, time domain features such as short-time energy and autocorrelation function are obtained from the audio signal; when performing frequency domain analysis, frequency domain features such as analysis spectrum, power spectrum, and cepstrum are obtained. The obtained time domain features or frequency domain features can be used as audio analysis data to facilitate subsequent audio signal howling suppression processing.

[0071] S102: Input the audio analysis data into the trained machine learning model to suppress machine howling to obtain initial suppressed audio data.

[0072] In this embodiment, the machine learning model is a trained neural network, and the training data of the machine learning model includes a first audio signal without howling and a second audio signal with howling, and the audio signal includes a first audio signal with multiple sampling rates and a second audio signal with multiple howling frequencies.

[0073] Specifically, the second audio signal containing howling can be collected in different sound field environments and terminal device scenarios to cover various frequency bands. By inputting the second audio signal containing howling into a machine learning model for training, the machine learning model outputs corresponding clean audio signal features without howling. Specifically, the features can be time domain or frequency domain information extracted during data preprocessing. After training, a machine learning model capable of processing howling in audio signals is obtained.

[0074] By using the trained machine learning model to perform machine howling suppression on the audio analysis data, the initial suppressed audio data is obtained, which can effectively avoid the limitations of traditional notch filters or other howling suppression methods, while improving the efficiency of howling suppression processing. Moreover, the machine learning model of the neural network can achieve real-time howling suppression of the audio signal, which can effectively improve the quality of the audio signal.

[0075] S103: Perform a howling residual detection on the initial suppressed audio data to obtain a howling residual detection result.

[0076] In some embodiments, performing a howling residual detection on the initial suppressed audio data to obtain a howling residual detection result includes:

[0077] calculating a signal suppression ratio between the initial suppressed audio data and the audio analysis data;

[0078] After determining that the signal suppression ratio is lower than a first threshold, calculating a peak-to-peak power ratio of a frequency point in the initially suppressed audio data;

[0079] After determining that the peak-to-peak power ratio of the frequency point in the initial suppressed audio data is greater than a second threshold, determining that the frequency point is a howling residual point, and determining that howling residual exists in the initial suppressed audio data.

[0080] In this embodiment, after the machine learning model performs machine howling processing to obtain initial suppressed audio data, howling residual detection is performed on the initial suppressed audio data to detect frequency points where howling is not completely suppressed, thereby reducing the probability of subsequent howling.

[0081] Exemplarily, a signal suppression ratio is calculated between the initially suppressed audio data and the audio analysis data to obtain the howling suppression of the audio analysis data by the machine learning model. In some embodiments, the signal suppression ratio can be calculated in either the time domain or the frequency domain.

[0082] Specifically, when calculating the signal suppression ratio in the time domain, the signal suppression ratio is the ratio of the time domain short-time energy between the initial suppressed audio data and the audio analysis data; and when calculating the signal suppression ratio in the frequency domain, the signal suppression ratio is the ratio of the frequency domain sub-band power between the initial suppressed audio data and the audio analysis data.

[0083] refer to Figure 2 , the calculated signal suppression ratio is recorded as g. After the signal suppression ratio g is calculated, the signal suppression ratio g is compared with the first threshold T1. After determining that the signal suppression ratio g is less than the first threshold T1, the peak proximity power ratio of the frequency point in the initial suppressed audio data is calculated, so as to facilitate subsequent further determination of whether there is residual howling based on the peak proximity power ratio PNPR of the frequency point.

[0084] After determining that the peak proximity power ratio PNPR of the frequency point in the initial suppressed audio data is greater than the second threshold T2, the frequency point is determined to be a howling residual point, and it is determined that howling residual exists in the initial suppressed audio data.

[0085] In the above process, the signal suppression ratio g and the peak-to-nearest power ratio PNPR of the frequency points of the initially suppressed audio data are calculated to facilitate determining whether there is residual howling in the initially suppressed audio data based on the first threshold T1 and the second threshold T2. The frequency points that simultaneously satisfy the conditions that the signal suppression ratio g is less than the first threshold T1 and the peak-to-nearest power ratio PNPR is greater than the second threshold T2 are regarded as residual howling points, and potential howling frequency points that are not completely suppressed in the initially suppressed audio data are determined.

[0086] In some other embodiments, the peak-to-peak power ratio PNPR of the frequency point is the peak power of the frequency point. and the mth adjacent frequency power The ratio of , m is an integer.

[0087] Exemplarily, the peak-to-peak power ratio PNPR of the frequency point satisfies the following formula:

[0088]

[0089] in, Indicates the frequency point where the peak is located, k is the frame identifier, Δf is the frequency resolution of the spectrum, and m is an integer.

[0090] In other embodiments, in the process of calculating the peak adjacent power ratio PNPR of the frequency points, taking into account the possible broadband characteristics of howling, the absolute value of m is taken to be greater than 2 and not exceed a preset threshold. That is, when taking adjacent frequency points, the most adjacent frequency points separated by one and two digits are skipped. At the same time, nearby frequency points are selected within a certain interval based on the bandwidth and spectrum resolution of the howling, so that broadband howling residuals can also be detected, thereby improving the accuracy of the detection results.

[0091] Exemplarily, the preset threshold is equal to 5, that is, in the process of calculating the peak proximity power ratio PNPR of the frequency point, the absolute value of m is greater than 2 and less than or equal to 5.

[0092] S104: Process the initial suppressed audio data according to the howling residual detection result to obtain target suppressed audio data.

[0093] In some embodiments, the processing the initial suppressed audio data to obtain target suppressed audio data according to the howling residual detection result includes:

[0094] using the frequency point determined as the howling residual point in the initial suppressed audio data as the target suppression frequency point;

[0095] A suppression coefficient of the target suppression frequency point is obtained, and the target suppression frequency point in the initial suppressed audio data is suppressed according to the suppression coefficient to obtain the target suppressed audio data.

[0096] In this embodiment, after detecting the frequency point where the howling residual point exists in the initial suppressed audio data through howling residual detection, the howling residual point is used as the target suppression frequency point, and the target suppression frequency point in the initial suppressed audio data is suppressed according to the suppression coefficient of the target suppression frequency point, thereby completing the final suppression processing process of the howling, so as to obtain the target suppressed audio data and complete the howling suppression, thereby effectively avoiding the problem of howling suppression residual of the audio signal by the machine learning model.

[0097] Exemplarily, the suppression coefficient is the signal suppression ratio g or the peak proximity power ratio PNPR, and howling suppression is performed on the target suppression frequency according to the suppression coefficient. Since there is very little howling residual in the initial suppressed audio data output after suppression by the machine learning model, only simple howling suppression is required through the suppression coefficient to suppress the remaining howling.

[0098] It should be noted that the howling suppression process for the residual howling in the target suppression frequency point may adopt a howling suppression method in the prior art, which is not particularly limited here.

[0099] In some other embodiments, after determining that the signal suppression ratio g is greater than or equal to the first threshold T1, it can be determined that the initial suppressed audio data after suppression by the machine learning model meets the requirements, that is, it is determined that there is no howling residual in the initial suppressed audio data; then, after determining that there is no howling residual in the initial suppressed audio data, the initial suppressed audio data is used as the target suppressed audio data, thereby completing the howling suppression process of the audio signal.

[0100] Exemplary, reference Figure 5 , using the howling processing method of the present invention to process a howling voice signal, Figure 5 The upper half is a schematic diagram of the frequency of the voice signal before processing, and the lower half is a schematic diagram of the frequency of the voice signal after processing. The sampling rate of the howling voice signal before processing is 16kHz, the howling in the voice is continuous, and the howling component is multi-frequency, broadband, and variable. However, after the howling processing method of the present invention is used to process the voice signal, the howling component is basically suppressed and there is almost no damage to the voice, thereby effectively improving the effect of suppressing howling in the audio signal.

[0101] In some further embodiments, after determining that the peak proximity power ratio PNPR of the frequency point in the initial suppressed audio data is less than or equal to a second threshold T2, it is determined that no howling remains in the initial suppressed audio data, and the initial suppressed audio data is used as the target suppressed audio data, thereby completing the howling suppression process.

[0102] The present invention provides a howling suppression device, referring to Figure 3 ,include:

[0103] The preprocessing module 301 is used to preprocess the input audio signal to obtain audio analysis data;

[0104] A machine learning model processing module 302 is configured to input the audio analysis data into a trained machine learning model to perform machine howling suppression to obtain initial suppressed audio data;

[0105] a howling residual detection module 303, configured to perform a howling residual detection on the initial suppressed audio data to obtain a howling residual detection result;

[0106] The residual suppression module 304 is configured to process the initial suppressed audio data according to the howling residual detection result to obtain target suppressed audio data.

[0107] In some embodiments, the howling residual detection module 303 is further configured to:

[0108] calculating a signal suppression ratio between the initial suppressed audio data and the audio analysis data;

[0109] After determining that the signal suppression ratio is lower than a first threshold, calculating a peak-to-peak power ratio of a frequency point in the initially suppressed audio data;

[0110] After determining that the peak-to-peak power ratio of the frequency point in the initial suppressed audio data is greater than a second threshold, determining that the frequency point is a howling residual point, and determining that howling residual exists in the initial suppressed audio data.

[0111] In some embodiments, the residue suppression module 304 is further configured to:

[0112] using the frequency point determined as the howling residual point in the initial suppressed audio data as the target suppression frequency point;

[0113] A suppression coefficient of the target suppression frequency point is obtained, and the target suppression frequency point in the initial suppressed audio data is suppressed according to the suppression coefficient to obtain the target suppressed audio data.

[0114] In some embodiments, the howling residual detection module 303 is further configured to:

[0115] After determining that the signal suppression ratio is greater than or equal to the first threshold, determining that no howling residual exists in the initial suppressed audio data;

[0116] The residual suppression module 304 is further configured to:

[0117] After determining that no howling remains in the initial suppressed audio data, the initial suppressed audio data is used as the target suppressed audio data.

[0118] In some embodiments, the pre-processing module 301 is further configured to:

[0119] Performing time domain analysis or frequency domain analysis on the audio signal to obtain time domain features or frequency domain features, and using the audio signal with the time domain features or the frequency domain features as the audio analysis data.

[0120] It should be noted that the structure and principle of the above-mentioned howling suppression device correspond one-to-one to the steps in the above-mentioned howling suppression method, so they are not described in detail here.

[0121] It should be noted that the division of the modules of the above apparatus is merely a division of logical functions. In actual implementation, they may be fully or partially integrated into a single physical entity or physically separated. Furthermore, these modules may be implemented entirely in software called by a processing element, or entirely in hardware. Alternatively, some modules may be implemented in software called by a processing element, while others may be implemented in hardware. For example, the selection module may be a separate processing element, or integrated into a chip of the above system. Furthermore, it may be stored in the form of program code in the memory of the above system, called by a processing element of the above system to perform the functions of the above module x. The implementation of other modules is similar. Furthermore, these modules may be fully or partially integrated or implemented independently. The processing element described herein may be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules may be performed by hardware integrated logic circuits in the processor element or by software instructions.

[0122] For example, the above modules can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, when a module is implemented by scheduling program code through a processing element, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0123] The present invention also discloses a chip, characterized in that the chip includes a processor and a communication interface, and the processor is configured to execute the above-mentioned howling suppression method.

[0124] In other embodiments of the present application, the present application discloses an electronic device, such as Figure 4As shown, the electronic device 400 may include: one or more processors 401; a memory 402; a display 403; one or more application programs (not shown); and one or more computer programs 404. The above components may be connected via one or more communication buses 405. The one or more computer programs 404 are stored in the memory 402 and configured to be executed by the one or more processors 401 to implement the above-mentioned howling suppression method.

[0125] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0126] The functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0127] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.

[0128] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

[0129] While the embodiments of the present invention have been described in detail above, it will be apparent to those skilled in the art that various modifications and variations of these embodiments are possible. However, it should be understood that such modifications and variations are within the scope and spirit of the present invention as set forth in the claims. Furthermore, the invention described herein is susceptible to other embodiments and may be practiced or implemented in a variety of ways.

Claims

1. A howling suppression method, characterized in that: The method comprises: Preprocessing the input audio signal to obtain audio analysis data includes: performing time domain analysis or frequency domain analysis on the audio signal to obtain time domain features or frequency domain features, and using the audio signal with the time domain features or the frequency domain features as the audio analysis data; Inputting the audio analysis data into a trained machine learning model to perform machine howling suppression to obtain initial suppressed audio data; Performing a howling residual detection on the initial suppressed audio data to obtain a howling residual detection result, comprising: calculating a signal suppression ratio of the initial suppressed audio data and the audio analysis data in the time domain or the frequency domain; after determining that the signal suppression ratio is lower than a first threshold, calculating a peak adjacent power ratio of a frequency point in the initial suppressed audio data; after determining that the peak adjacent power ratio of the frequency point in the initial suppressed audio data is greater than a second threshold, determining that the frequency point is a howling residual point, and determining that a howling residual exists in the initial suppressed audio data; wherein, when the signal suppression ratio is calculated in the time domain, the signal suppression ratio is a ratio of time-domain short-time energy between the initial suppressed audio data and the audio analysis data; when the signal suppression ratio is calculated in the frequency domain, the signal suppression ratio is a ratio of frequency-domain sub-band power between the initial suppressed audio data and the audio analysis data; The initial suppressed audio data is processed according to the howling residual detection result to obtain target suppressed audio data.

2. The howling suppression method according to claim 1, wherein: The peak adjacent power ratio of the frequency point is the ratio of the peak power of the frequency point to the power of the mth adjacent frequency, where m is an integer.

3. The howling suppression method according to claim 2, characterized in that: The absolute value of m is greater than 2 and does not exceed a preset threshold.

4. The howling suppression method according to claim 1, wherein: The processing of the initial suppressed audio data to obtain target suppressed audio data according to the howling residual detection result includes: using the frequency point determined as the howling residual point in the initial suppressed audio data as the target suppression frequency point; A suppression coefficient of the target suppression frequency point is obtained, and the target suppression frequency point in the initial suppressed audio data is suppressed according to the suppression coefficient to obtain the target suppressed audio data.

5. The howling suppression method according to claim 4, characterized in that: The suppression coefficient is a signal suppression ratio or a peak-to-peak power ratio.

6. The howling suppression method according to claim 1, characterized in that: Performing a howling residual detection on the initial suppressed audio data to obtain a howling residual detection result further includes: After determining that the signal suppression ratio is greater than or equal to the first threshold, determining that no howling residual exists in the initial suppressed audio data; The processing of the initial suppressed audio data to obtain target suppressed audio data according to the howling residual detection result includes: After determining that no howling remains in the initial suppressed audio data, the initial suppressed audio data is used as the target suppressed audio data.

7. The howling suppression method according to claim 1, characterized in that: The method further includes, after determining that the peak adjacent power ratio of the frequency point in the initial suppressed audio data is less than or equal to a second threshold, determining that no howling remains in the initial suppressed audio data, and using the initial suppressed audio data as the target suppressed audio data.

8. The howling suppression method according to claim 1, characterized in that: Before performing time domain analysis or frequency domain analysis on the audio signal, the method further includes: The audio signal is divided into sub-bands according to the frequency domain of the audio signal.

9. A howling suppression device, used to implement the howling suppression method according to any one of claims 1 to 8, characterized in that: The howling suppression device comprises: a preprocessing module, configured to preprocess an input audio signal to obtain audio analysis data, including: performing time domain analysis or frequency domain analysis on the audio signal to obtain time domain features or frequency domain features, and using the audio signal having the time domain features or the frequency domain features as the audio analysis data; a machine learning model processing module, configured to input the audio analysis data into a trained machine learning model for machine howling suppression to obtain initial suppressed audio data; A howling residual detection module is configured to perform a howling residual detection on the initial suppressed audio data to obtain a howling residual detection result, comprising: calculating a signal suppression ratio of the initial suppressed audio data and the audio analysis data; after determining that the signal suppression ratio is lower than a first threshold, calculating a peak adjacent power ratio of a frequency point in the initial suppressed audio data; after determining that the peak adjacent power ratio of the frequency point in the initial suppressed audio data is greater than a second threshold, determining that the frequency point is a howling residual point, and determining that a howling residual exists in the initial suppressed audio data; The residual suppression module is configured to process the initial suppressed audio data according to the howling residual detection result to obtain target suppressed audio data.

10. The howling suppression device according to claim 9, characterized in that: The residual suppression module is further configured to: using the frequency point determined as the howling residual point in the initial suppressed audio data as the target suppression frequency point; A suppression coefficient of the target suppression frequency point is obtained, and the target suppression frequency point in the initial suppressed audio data is suppressed according to the suppression coefficient to obtain the target suppressed audio data.

11. The howling suppression device according to claim 9, characterized in that: The howling residual detection module is further used for: After determining that the signal suppression ratio is greater than or equal to the first threshold, determining that no howling residual exists in the initial suppressed audio data; The residual suppression module is further configured to: After determining that no howling remains in the initial suppressed audio data, the initial suppressed audio data is used as the target suppressed audio data.

12. A chip, characterized in that: The chip includes a processor and a communication interface, and the processor is configured to execute the howling suppression method according to any one of claims 1 to 8.

13. An electronic device, characterized in that: include: processor and memory; The memory is used to store computer programs; The processor is configured to execute the computer program stored in the memory, so as to enable the electronic device to perform the howling suppression method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Howl detection method and device and howl inhibition method and device

    CN107919134A

  • Music rhythm detection method based on frequency domain and time domain and storage medium

    CN111081271A

  • Audio signal processing method, model training method and related device

    CN111210021A

  • Howling suppression method, device and equipment

    CN111583949A

  • Howling detection method and device, storage medium and computer equipment

    CN111800725A