Howling suppression method and device, storage medium and electronic device
By using a feedback detection model and a feedback suppression model to detect and suppress audio features in the acoustic loop of instant messaging, the problem of poor feedback suppression in instant messaging scenarios is solved. This achieves accurate detection and suppression of complex feedback signals and improves sound quality.
Patent Information
- Application Number
- CN202210307288.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-03-25
AI Technical Summary
Existing howling suppression techniques are ineffective in the acoustic loops of instant messaging, and cannot effectively handle nonlinear and uncertain howling characteristics, such as intermittency, multiple frequency points, and frequency shifts.
Employing a howling detection model and a howling suppression model, this approach extracts audio features and performs howling feature parameter detection and suppression to adapt to the complexity of instant messaging scenarios.
It achieves accurate detection and effective suppression of complex and non-constant howling signals in the acoustic loop of instant messaging, thus improving sound quality.
Smart Images

Figure CN114863941B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure relate to the field of audio signal processing technology, and more specifically, to howling suppression methods and apparatus, storage media, and electronic devices. Background Technology
[0002] This section is intended to provide background or context for the embodiments of this disclosure set forth in the claims, and the description herein is not acknowledged as prior art simply because it is included in this section.
[0003] The essence of howling is that the feedback system is in an unstable state. The stability of the system can be determined by its open-loop transfer function using the Nyquist stability criterion. In a typical feedback system, the input R(s) is transferred to the output C(s) via the forward transfer function G(s), and the output C(s) is transferred to the input R(s) via the feedback transfer function H(s). Therefore, the open-loop transfer function of the system can be derived as: H(s)·G(s). Thus, the stability of the system can be determined based on the Nyquist plot or Bode plot of the open-loop transfer function. In a feedback system, when the feedback signal and the input signal are in phase, and the feedback loop is positive (i.e., the corresponding open-loop gain is greater than 1), the system is in an unstable state.
[0004] In acoustic scenarios, howling is likely to occur when a closed acoustic feedback loop is formed. Summary of the Invention
[0005] In acoustic scenarios such as conference rooms, auditoriums, and karaoke rooms, microphones pick up sound, and speakers play it back. The signal played by the speakers is then picked up by the microphones, creating a loop. In these acoustic scenarios, feedback is often generated by the system itself, and the acoustic characteristics of feedback are mostly single-frequency or multi-frequency continuous feedback. The feedback in the aforementioned acoustic scenarios has relatively fixed and easily identifiable characteristics.
[0006] In the acoustic loop scenario of Real-Time Communication (RTC), factors such as differences in the built-in audio processing performance of different devices, variations in network transmission environments between devices, changes in device location, and differences in device frequency response can all introduce uncertainties to the transmission of audio signals in the acoustic loop. Furthermore, the changes and effects of these factors are not linear, making it impossible to quantitatively measure and analyze the transfer function of the acoustic loop in an RTC scenario. Simultaneously, these nonlinear factors also introduce many characteristics distinct from traditional feedback scenarios, such as the intermittency, multiple frequency points, frequency shifting, and frequency diffusion of feedback.
[0007] Current howling suppression technologies generally employ:
[0008] Option 1: Use frequency-shifting and phase-shifting, notch filtering, and adaptive filtering methods for howling suppression. Frequency-shifting and phase-shifting: This method changes the conditions that cause howling by shifting the frequency and phase, thus suppressing it. Notch filtering: First, determine the frequency of the howling, then apply a notch filter to that frequency to suppress it. Adaptive filtering: Dynamically update the filter coefficients to filter the howling signal. However, these methods are more suitable for traditional conference and hearing aid systems, where the conditions for howling are relatively fixed. They are less effective at suppressing howling in acoustic loops of real-time communication systems with many nonlinearities and uncertainties.
[0009] Option 2: First, detect the howling frequency point, then remove the signal at the howling frequency point, and finally repair the signal near the howling frequency point using a neural network. Option 2 essentially still uses a notch filtering method for howling suppression, but it provides a signal repair network compared to the traditional notch filtering method. While it improves sound quality compared to the traditional notch filtering method, like the traditional method, it is still not suitable for acoustic loop scenarios in instant messaging.
[0010] Therefore, there is a great need for an improved method and apparatus for suppressing howling, as well as storage media and electronic devices, to provide an acoustic loop that can adapt to real-time communication with many nonlinearities and uncertainties.
[0011] In this context, embodiments of the present disclosure are intended to provide a method and apparatus for suppressing howling, a storage medium, and an electronic device.
[0012] According to one aspect of this disclosure, a howling suppression method is provided, applied to a first device, the first device being used for real-time communication with a second device, the first device and the second device belonging to the same acoustic loop, the first device including a first communication module, a first audio acquisition module and a first audio playback module, the second device including a second communication module, a second audio acquisition module and a second audio playback module, the method comprising:
[0013] Extract the audio features of the audio signal to be processed, wherein the audio signal to be processed is the audio signal acquired by the first device through its first audio acquisition module, and the audio signal to be processed is the superposition of the acoustic signal emitted by the sound source and the second audio signal played by the second audio playback module;
[0014] The audio features are input into the feedback detection model, and the feedback detection model outputs the feedback feature parameters of the audio signal to be processed.
[0015] The feedback feature parameters and the audio features are input into the feedback suppression model;
[0016] Based on the output of the feedback suppression model, a feedback suppression audio signal is obtained.
[0017] According to one aspect of this disclosure, a howling suppression device is provided, applied to a first device, the first device being used for real-time communication with a second device, the first device and the second device belonging to the same acoustic loop, the first device including a first communication module, a first audio acquisition module and a first audio playback module, the second device including a second communication module, a second audio acquisition module and a second audio playback module, the device comprising:
[0018] An audio feature extraction module is used to extract audio features of an audio signal to be processed. The audio signal to be processed is an audio signal acquired by the first device through its first audio acquisition module. The audio signal to be processed is the superposition of an acoustic signal emitted by a sound source and a second audio signal played by the second audio playback module.
[0019] The feedback detection module is used to input the audio features into the feedback detection model, and the feedback detection model outputs the feedback feature parameters of the audio signal to be processed;
[0020] The feedback suppression input module is used to input the feedback feature parameters and the audio features into the feedback suppression model;
[0021] The feedback suppression output module is used to obtain the feedback suppression audio signal based on the output of the feedback suppression model.
[0022] According to one aspect of this disclosure, a storage medium is provided on which a computer program is stored, wherein the above-described howling suppression method is executed by a processor.
[0023] According to one aspect of this disclosure, an electronic device is provided, comprising:
[0024] Processor; and
[0025] Memory for storing the executable instructions of the processor;
[0026] The processor is configured to execute any of the above-described howling suppression methods by executing the executable instructions.
[0027] According to the feedback suppression method of this disclosure, the audio features of the audio signal to be processed are input into a feedback detection model for detection to obtain feedback feature parameters, and feedback suppression is performed based on the audio features and feedback feature parameters. Therefore, this disclosure is applicable to acoustic loops in instant messaging, enabling the detection of complex, non-fixed feedback feature parameters generated by uncertain feedback conditions in the acoustic loop of instant messaging based on a feedback detection model, thereby achieving effective feedback suppression based on the detected feedback feature parameters. Attached Figure Description
[0028] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which:
[0029] Figure 1 The diagram schematically illustrates the spectrum of howling signals in a common acoustic scenario in the prior art;
[0030] Figure 2 The diagram schematically illustrates the spectrum of howling signals in another common acoustic scenario in the prior art;
[0031] Figure 3 A spectrogram of a howling signal in an acoustic loop of an instant messaging scenario according to an embodiment of the present disclosure is schematically shown.
[0032] Figure 4 A flowchart of a howling suppression method according to an embodiment of the present disclosure is shown schematically;
[0033] Figure 5 A schematic diagram of the acoustic loop in an instant messaging scenario according to an embodiment of the present disclosure is shown.
[0034] Figure 6 A schematic diagram illustrating the cascaded application of a howling detection model and a howling suppression model according to embodiments of the present disclosure is shown.
[0035] Figure 7 A schematic diagram of a howling detection model according to an embodiment of the present disclosure is shown.
[0036] Figure 8 A schematic diagram of a howling suppression model according to an embodiment of the present disclosure is shown;
[0037] Figure 9 A flowchart illustrating the training of a howling detection model according to an embodiment of the present disclosure is shown schematically;
[0038] Figure 10A flowchart illustrating the training of a howling suppression model according to an embodiment of the present disclosure is shown schematically;
[0039] Figure 11 A block diagram schematically illustrates a first audio processing module of a first device according to an embodiment of the present disclosure;
[0040] Figure 12 A block diagram of a howling suppression device according to an embodiment of the present disclosure is shown schematically;
[0041] Figure 13 A schematic diagram of a storage medium according to an embodiment of the present disclosure is shown; and
[0042] Figure 14 A block diagram of an electronic device according to a disclosed embodiment is shown schematically.
[0043] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation
[0044] The principles and spirit of this disclosure will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.
[0045] Those skilled in the art will recognize that embodiments of this disclosure can be implemented as a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0046] According to embodiments of this disclosure, a howling suppression method, a howling suppression device, a storage medium, and an electronic device are provided.
[0047] In this document, any number of elements in the accompanying figures is for illustrative purposes and not for limitation, and any naming is for distinction only and has no limiting meaning.
[0048] The principles and spirit of this disclosure are explained in detail below with reference to several representative embodiments. Invention Overview
[0050] The inventors discovered that in acoustic scenarios such as conference rooms, auditoriums, and KTVs, microphones pick up sound while speakers play it back. The signal played by the speakers is then picked up by the microphones, creating a loop. In these acoustic scenarios, feedback is often generated by the system itself, and the acoustic characteristics of this feedback are mostly single-frequency or multi-frequency continuous feedback. Figure 1 The spectrum of the multi-frequency full howling signal in the above scenario is shown. Figure 1 In the diagram, F11 is the time-domain waveform diagram, and F12 is the spectrum diagram. The horizontal axis of F11 and F12 represents time, the vertical axis of F11 represents amplitude, and the vertical axis of F12 represents frequency. The brightness of the graph in F12 represents the energy level of that frequency. Figure 2 The graph of the single-frequency howling signal with background noise in the above scenario is shown. Figure 2 The F21 time-domain waveform diagram and F22 spectrum diagram are shown in the figure. The horizontal axis of F21 and F22 is time, the vertical axis of F21 is amplitude, and the vertical axis of F22 is frequency. The brightness of the F22 graph represents the energy of that frequency. Figure 1 and Figure 2 For example, the howling in the above acoustic scenarios has relatively fixed and easily identifiable characteristics.
[0051] In the acoustic loop scenario of Real-Time Communication (RTC), factors such as differences in the built-in audio processing performance of different devices, variations in network transmission environments between devices, changes in device location, and differences in device frequency response can all introduce uncertainties to the transmission of audio signals in the acoustic loop. Furthermore, the changes and effects of these factors are not linear, making it impossible to quantitatively measure and analyze the transfer function of the acoustic loop in an RTC scenario. Simultaneously, these nonlinear factors also introduce many characteristics distinct from traditional feedback scenarios, such as the intermittency, multiple frequency points, frequency shifting, and frequency diffusion of feedback.
[0052] For example, built-in audio processing in devices includes noise reduction. Noise tracking in noise reduction might identify feedback as noise and eliminate it to some extent. However, because the acoustic loop still exists, feedback can recur due to external stimuli. On the other hand, if noise reduction only partially eliminates feedback, the audio signal will exhibit intermittent feedback, fluctuating in intensity. Simultaneously, other nonlinear processing can affect the system's phase amplitude characteristics, causing changes and spread in the feedback frequency. Furthermore, due to differences in the frequency response of different devices during acquisition and playback, the transfer functions of the acoustic loops are inherently inconsistent, resulting in different feedback frequencies and characteristics from different devices.
[0053] This shows that the howling signals generated in the acoustic loop of instant messaging are more complex and more uncertain.
[0054] Figure 3 The diagram shows the spectrum of the howling signal with complex characteristics in the above scenario. Figure 3 The F31 time-domain waveform diagram and F32 spectrum diagram are shown in the figure. The horizontal axis of F31 and F32 is time, the vertical axis of F31 is amplitude, and the vertical axis of F32 is frequency. The brightness of the F32 graph indicates the energy level of that frequency. Figure 3 The howling signal in the above scenario is compared to Figure 1 and Figure 2 It exhibits more complex howling characteristics. Specifically, Figure 1 F12 indicates that the howling is multi-frequency, and according to F11, the energy (amplitude) of the howling drowns out the background sound, so the energy display in F11 is relatively uniform. Figure 2 F21 shows the energy (amplitude) changing over time, thus the howling in F21 does not drown out the background sound. Meanwhile, the long straight line in F22, where the frequency remains constant over time, indicates a single-frequency howling. Therefore, Figure 1 The signal shown is a multi-frequency full howling signal, while Figure 2 The image shown is a single-frequency howling. See also... Figure 3 In F31, three instances with larger amplitudes indicate that the howling has masked the original background sound. Meanwhile, F32 does not clearly display multi-frequency or single-frequency howling as in F12 and F22. Figure 3 In such scenarios, the howling signal is compared to Figure 1 and Figure 2 It has more complex howling characteristics.
[0055] Current howling suppression technologies generally employ:
[0056] Option 1: Use frequency shifting and phase shifting, notch filtering, and adaptive filtering to suppress howling.
[0057] Frequency-shifting and phase-shifting methods: These methods alter the conditions that cause howling, thereby suppressing it. However, this method changes both the phase and frequency, altering signal characteristics—typically, the speaker's timbre—leading to distortion. Furthermore, in real-time communication scenarios, various nonlinearities exist, making it impossible to cover them all with frequency-shifting and phase-shifting alone, such as determining the appropriate frequency and phase changes. Generally, frequency-shifting and phase-shifting are more suitable for relatively fixed scenarios, requiring targeted optimization through system transfer function analysis. However, for real-time communication scenarios with numerous nonlinearities and uncertainties, their howling suppression effect often fails.
[0058] Notch filtering: First, the frequency of the howling is determined, and then a notch filter is applied to suppress the howling at that frequency, thus achieving the purpose of howling suppression. A crucial prerequisite for notch filtering is the accurate detection of the howling frequency. However, in instant messaging scenarios, howling exhibits characteristics such as intermittency, multiple frequency points, frequency shifting, and frequency diffusion, making frequency prediction extremely difficult and hindering the practical application of this method. Notch filtering is generally suitable for scenarios with fixed howling frequencies and continuous howling.
[0059] Adaptive filtering: This method dynamically updates the filter coefficients to filter the howling signal. It eliminates the need for howling frequency detection in notch filtering and estimates the acoustic feedback signal in real time. However, adaptive filtering is only suitable for filtering linear components and is unlikely to be effective in scenarios with many nonlinear factors, such as instant messaging.
[0060] Therefore, frequency-shifting and phase-shifting methods, notch filtering, and adaptive filtering are more suitable for acoustic scenarios where the conditions for howling are relatively fixed, such as traditional conferencing and hearing aid systems. These methods are less effective at suppressing howling in acoustic loop scenarios involving real-time communication with numerous nonlinearities and uncertainties.
[0061] Option 2: First, detect the howling frequency points, then remove the signals at the howling frequency points, and finally repair the signals near the howling frequency points using a neural network.
[0062] Analysis shows that Solution 2 essentially still uses a notch filter-like approach to suppress howling, but it also provides a signal restoration network compared to the traditional notch filter. While it improves sound quality compared to the traditional notch filter, like the traditional method, it is still not suitable for acoustic loop scenarios in instant messaging.
[0063] In view of the above, the technical solution of this disclosure is as follows: In the feedback suppression method of this disclosure, in the acoustic loop of instant messaging, the audio features of the audio signal to be processed are input to the feedback detection model for detection to obtain feedback feature parameters, and feedback suppression is performed based on the audio features and feedback feature parameters. Since the causes of feedback signals in the acoustic loop of instant messaging scenarios are uncertain, compared with common acoustic scenarios such as conference rooms, auditoriums, and KTVs, complex and non-fixed feedback signals are more likely to occur in instant messaging scenarios, and such feedback signals are difficult to detect and suppress. Based on this, in this disclosure, the feedback signal is first accurately detected by the feedback detection model, which is trained with richer feedback feature parameters in instant messaging scenarios; on this basis, the feedback suppression model is also used to learn the influence of these feedback feature parameters on feedback suppression, thereby using the feedback suppression model to suppress feedback in the audio signal to be processed.
[0064] After introducing the basic principles of this disclosure, various non-limiting embodiments of this disclosure will be described in detail below.
[0065] Exemplary methods
[0066] The following is combined Figure 4 This document describes a feedback suppression method according to an exemplary embodiment of the present disclosure. The feedback suppression method is applied to a first device, which is used for real-time communication with a second device. The first device and the second device belong to the same acoustic loop. The first device includes a first communication module, a first audio acquisition module, and a first audio playback module. The second device includes a second communication module, a second audio acquisition module, and a second audio playback module.
[0067] refer to Figure 4 As shown, the whistling suppression method may include the following steps:
[0068] Step S110: Extract the audio features of the audio signal to be processed. The audio signal to be processed is the audio signal acquired by the first device through its first audio acquisition module. The audio signal to be processed is the superposition of the acoustic signal emitted by the sound source and the second audio signal played by the second audio playback module.
[0069] Step S120: Input the audio features into the feedback detection model, and the feedback detection model outputs the feedback feature parameters of the audio signal to be processed;
[0070] Step S130: Input the howling feature parameters and the audio features into the howling suppression model;
[0071] Step S140: Obtain the howling suppression audio signal based on the output of the howling suppression model.
[0072] In the feedback suppression method of this disclosure, in the acoustic loop of instant messaging, the audio features of the audio signal to be processed are input into a feedback detection model for detection to obtain feedback feature parameters, and feedback suppression is performed based on the audio features and feedback feature parameters. Since the causes of feedback signals in the acoustic loop of instant messaging scenarios are uncertain, compared with common acoustic scenarios such as conference rooms, auditoriums, and KTVs, complex and non-fixed feedback signals are more likely to occur in instant messaging scenarios, making these types of feedback signals difficult to detect and suppress. Therefore, in this disclosure, the feedback signal is first accurately detected by a feedback detection model, which is trained using richer feedback feature parameters specific to instant messaging scenarios. Based on this, a feedback suppression model is used to learn the influence of these feedback feature parameters on feedback suppression, thereby using the feedback suppression model to suppress feedback in the audio signal to be processed.
[0073] The following is for reference. Figure 5 , Figure 5 A schematic diagram of the acoustic loop in an instant messaging scenario according to an embodiment of the present disclosure is shown. Figure 5 As shown, the first device 10, the second device 20, and the user 30 are located in the same physical space A. Physical space A can be, for example, a conference room or an office. The first device 10 includes a first audio acquisition module 11, a first communication module 12, and a first audio playback module 13. The second device 20 includes a second audio acquisition module 21, a second communication module 22, and a second audio playback module 23. The first device 10 and the second device 20 communicate in real-time.
[0074] Figure 5 The diagram shows two acoustic loops, C1 (solid arrow) and C2 (dashed arrow). Acoustic loop C1 is the audio signal transmission path based on the first audio acquisition module 11, the first communication module 12, the second communication module 22, and the second audio playback module 23. Acoustic loop C2 is the audio signal transmission path based on the second audio acquisition module 21, the second communication module 22, the first communication module 12, and the first audio playback module 13.
[0075] Taking acoustic loop C1 as an example, when sound source 30 emits an acoustic signal, the acoustic signal is acquired by the first audio acquisition module 11, transmitted via the first communication module 12 to the second communication module 22, and then transmitted by the second communication module 22 to the second audio playback module 23 for playback. The second audio signal played by the second audio playback module 23 is acquired by the first audio acquisition module 11 and superimposed on the acoustic signal emitted by sound source 30, entering the aforementioned acoustic loop C1 to complete closed-loop transmission. Furthermore, the second audio playback module 23 is located within the pickup distance of the first audio acquisition module 11, and the playback volume of the second audio playback module 23 is sufficient for the second audio signal to be picked up by the first audio acquisition module 11. Thus, the audio signal is transmitted within the acoustic loop C1. Similarly, acoustic loop C2 also completes closed-loop transmission of the audio signal in a similar manner. Therefore, since the second device 20 can also serve as the first device in its acoustic loop C2, the feedback suppression method of this disclosure can also be applied to the second device 20.
[0076] The following is for reference. Figure 6 , Figure 6 A schematic diagram of a model structure of a howling suppression method according to an embodiment of the present disclosure is shown.
[0077] like Figure 6As shown, in the feedback suppression method of this application, audio features are first extracted from the audio signal to be processed. After extraction, the audio features are input to the feedback detection model M1. The feedback feature parameters output by the feedback detection model M1 and the aforementioned audio features are then input to the feedback suppression model M2 for feedback suppression. Since the audio features are feature data extracted from the audio signal, and feature data cannot be played, it is necessary to restore the features output by the feedback suppression model M2 to obtain a playable feedback-suppressed audio signal.
[0078] In the exemplary embodiments of this disclosure, a Short-Time Fourier Transform (STFT) can be used to extract features from the audio signal to be processed. Considering the scenario of instant messaging, different sampling rates can be used in the STFT, such as 48kHz / 16kHz (music mode and voice mode). Furthermore, the number of points in the Fourier Transform (the more points in the Fourier Transform, the higher the frequency resolution) can be selected according to requirements (for example, in a music scenario, a 48kHz sampling rate would require more points in the Fourier Transform; correspondingly, in a voice scenario, a 16kHz sampling rate would require fewer points. Further adjustments can also be made based on the trade-off between overhead and accuracy). In some embodiments, 512 points can be selected. The frame length and frame shift of the audio signal to be processed can also be selected in conjunction with the audio acquisition module, audio playback module, or other processing modules that need to process the audio signal to be processed. For example, a frame length of 20 milliseconds and a frame shift of 10 milliseconds can be selected. The extracted audio features can be one or more of the spectral features, Bark spectrum features, Mel spectrum features, Mel cepstral features, and fundamental frequency features of the audio signal to be processed.
[0079] In an exemplary embodiment of this disclosure, the howling suppression model M2 can directly output suppressed audio features. After outputting the suppressed audio features, the howling suppression model M2 can perform a reverse operation relative to feature extraction to restore the suppressed audio features and obtain the howling suppressed audio signal. Thus, an audio signal for transmission or playback can be obtained through feature restoration.
[0080] In an exemplary embodiment of this disclosure, the howling suppression model M2 can also output a howling suppression mask, which characterizes the howling suppression frequency gain of the audio features of the reference sample signal compared to the audio features of the audio signal to be processed. In other words, the howling suppression mask provides the gain value at each frequency of the audio features, and the amplitude of each frequency of the audio features is multiplied by the gain value to suppress howling. Thus, during the training of the howling suppression model M2, the audio features of the howling sample signal can be used as the input of the howling suppression model M2. Based on the howling suppression frequency gain of the audio features of the reference sample information compared to the audio features of the howling sample signal, the howling suppression mask is obtained based on the howling suppression frequency gain, and the howling suppression mask is used as the output of the howling suppression model M2, thereby training the howling suppression model M2 to output a howling suppression mask capable of howling suppression. Since the audio features of the reference sample signal represented by the howling suppression mask are less than the howling suppression frequency gain of the audio features of the audio signal to be processed, the amount of information contained in the howling suppression mask is less than that of the audio features. Therefore, training the howling suppression model M2 based on the howling suppression mask can improve the training efficiency of the howling suppression model M2. Simultaneously, the howling suppression model M2 performs multiple calculations on the input audio features to obtain the howling suppression mask, resulting in less computation and higher computational efficiency. After the howling suppression model M2 outputs the howling suppression mask, it is multiplied by the audio features of the audio signal to be processed (the howling suppression frequency gain of the howling suppression mask is used to adjust the energy of each frequency point of the audio features of the audio signal to be processed, in order to remove / suppress howling), to obtain suppressed audio features. The suppressed audio features are then subjected to the inverse operation relative to feature extraction to restore the suppressed audio features, thus obtaining the howling-suppressed audio signal. Therefore, an audio signal for transmission or playback can be obtained through feature restoration.
[0081] See below. Figure 7 , Figure 7 A schematic diagram of a howling detection model according to an embodiment of the present disclosure is shown.
[0082] The feedback detection model sequentially includes an input processing layer 201, a first intermediate processing layer 204, and a classification output layer 206. The audio features are assigned to the input processing layer 201 and then input to the first intermediate processing layer 204. The first intermediate processing layer 204 is used to obtain local features of the audio features and uses these local features as feedback intermediate features. The classification output layer 206 is used to classify the feedback intermediate features to obtain feedback result features. The feedback intermediate features and the feedback result features are used as feedback feature parameters input to the feedback suppression model.
[0083] In an exemplary embodiment of this disclosure, the howling detection model further includes a backbone layer 202 and a recurrent layer 203 connected between the input processing layer 201 and the intermediate processing layer 204. The backbone layer 202 is used to perform convolution and / or pooling processing on the data input to the backbone layer 202 to further compress the data input to the backbone layer 203. The recurrent layer 203 is used to establish the correlation between the input data of the recurrent layer 203 and the output data of the recurrent layer. Since audio features are time-varying sequential features, the howling suppression of audio features at a certain moment is related to the audio features at adjacent moments. Therefore, the recurrent layer 203 can learn the correlation of sequential features, thereby improving the accuracy of howling suppression. The howling detection model also includes an attention layer 205 connected between the first intermediate processing layer 204 and the classification output layer 206. The attention layer 205 is used to perform weighted summation on the data input to the attention layer 205. The weights in attention layer 205 are also part of the learning process during the training of the howling detection model. Through training the howling detection model, the influence of the data output by the first intermediate processing layer 204 on the classification output layer 206 is learned, thereby improving the accuracy of the howling result features output by the classification output layer 206. The howling detection model can also have other structures, and this disclosure is not intended to limit it.
[0084] In the exemplary embodiments of this disclosure, because the propagation conditions of the audio signal propagation path in the acoustic loop suitable for instant messaging are complex and have high uncertainty, in order to enable the howling suppression model to obtain richer information about the howling signal, the howling feature parameters output by the howling detection model can include howling intermediate features and howling result features. The howling intermediate features are used to represent the spectral characteristics of the howling signal. The howling result features include one or more of the following: howling detection result, howling level, howling type, howling continuity, and frequency shift parameters. Therefore, on the one hand, by combining the intermediate features and the result features of howling, the howling characteristics are represented in a multidimensional way, which is then used by the howling suppression model for suppression. On the other hand, since the howling signal of the acoustic loop of instant messaging is different from common acoustic scenarios, it has the characteristics of discontinuity, multiple frequency points, frequency point movement, and frequency point diffusion. Therefore, the howling detection model detects the characteristics of howling signals that are different from common acoustic scenarios, so that the howling result features include one or more of the following: howling detection result, howling level, howling type, howling continuity, and frequency point movement parameters. During the training process, the howling suppression model can adjust the model parameters of the howling suppression model according to the howling result features, thereby removing howling signals with the above-mentioned howling result features from the audio signal to be processed.
[0085] Furthermore, the feedback detection result indicates whether feedback exists in the audio features input to the feedback detection model; the feedback level indicates the feedback intensity of the audio features input to the feedback detection model; the feedback type includes single-frequency feedback, multi-frequency feedback, and diffuse feedback; the feedback continuity includes continuous feedback and intermittent feedback; the frequency shift parameter includes a frequency shift type parameter and a frequency shift amplitude parameter. The frequency shift type parameter indicates whether frequency shift exists in the audio features input to the feedback detection model, and the frequency shift amplitude parameter indicates the magnitude of the frequency shift of the audio features input to the feedback detection model. Each of the above feedback result features can be a numerical label or a one-hot vector. Thus, feedback features can be described and vectorized from multiple different feature description methods.
[0086] See below. Figure 8 , Figure 8 A schematic diagram of a howling suppression model according to an embodiment of the present disclosure is shown.
[0087] The howling suppression model sequentially includes an encoder 214, a second intermediate processing layer 215, and a decoder 216. The encoder 214 performs feature encoding on the audio features and the howling result features to obtain encoded features. The second intermediate processing layer 215 performs feature filtering on the encoded features and the howling intermediate features. The decoder 216 performs feature decoding on the filtered encoded features. Since the howling intermediate features have already been processed by the intermediate layer in the howling detection model, they do not need to be encoded again in the howling suppression model. Therefore, the howling intermediate features can be input into the second intermediate processing layer 215 of the howling suppression model, which helps the howling suppression model learn the howling features better, achieving better suppression while avoiding redundant processing of the howling intermediate features.
[0088] In an exemplary embodiment of this disclosure, the second intermediate processing layer 216 may sequentially include a plurality of connected Long Short-Term Memory (LSTM) units and a fully connected layer. The LSM units are used to perform feature filtering on the encoded features, and the fully connected layer is used to perform a weighted summation of the outputs of the plurality of LSM units. Thus, the relationship between howling features and audio features is effectively learned through the plurality of LSM units and the fully connected layer.
[0089] In an exemplary embodiment of this disclosure, audio features can be convolved via convolutional layer 211 and input into encoder 214, thereby ensuring that the feature size of the input audio features can be adapted to encoder 214. Feedback result features can be learned via embedding layer 212 and input into encoder 214, thereby ensuring that the feature size of the input feedback result features can be adapted to encoder 214. Feedback intermediate features can be convolved via convolutional layer 213 and input into second intermediate processing layer 216, thereby ensuring that the feature size of the input feedback intermediate features can be adapted to second intermediate processing layer 216. Convolutional layer 217 can be, for example, a deconvolutional layer, used to restore the features output by decoder 215 to be consistent with the audio features of the input feedback suppression model.
[0090] The above is merely an illustrative representation of one network structure for a howling suppression model, and this disclosure is not intended to limit it.
[0091] See below. Figure 9 , Figure 9 A flowchart illustrating the training of a howling detection model according to an embodiment of the present disclosure is shown schematically. Figure 9 The following steps are shown:
[0092] Step S101: Obtain the first set of sample signals.
[0093] In an exemplary embodiment of this disclosure, the first sample signal set includes a plurality of first sample signals and howling characteristic parameters of the first sample signals. The howling characteristic parameters include the same parameter items as the howling result features. For example, if the howling result features include howling detection results, howling level, howling type, howling continuity, and frequency shift parameters, then the howling characteristic parameters also include howling detection results, howling level, howling type, howling continuity, and frequency shift parameters. The parameter values of each parameter item vary depending on the specific audio signal.
[0094] The first sample signal includes a reference sample signal and a howling sample signal. The reference sample signal is the audio signal played by the playback device in the acoustic loop. Specifically, the playback device is used during model training, such as... Figure 5 A playback device is installed at the location of the central sound source 30 to play a reference sample signal. The howling sample signal is the superposition of the audio signal acquired by the first audio acquisition module and played by the playback device, and the second audio signal played by the second audio playback module.
[0095] Step S102: Extract the first sample audio features of the first sample signal.
[0096] Specifically, the algorithm for extracting the first sample signal can be the same as the algorithm for extracting the audio features of the audio signal to be processed.
[0097] Step S103: Use the first sample audio features as input to the howling detection model, and adjust the model parameters of the howling detection model according to the difference between the output of the howling detection model and the corresponding howling characteristic parameters.
[0098] Specifically, by adjusting the model parameters of the howling detection model, the howling detection model can output howling result features consistent with the corresponding howling characteristic parameters, thereby improving the detection performance of the howling detection model.
[0099] See below. Figure 10 , Figure 10 A flowchart illustrating the training of a howling suppression model according to an embodiment of the present disclosure is shown schematically. Figure 10 The following steps are shown:
[0100] Step S104: Obtain the second set of sample signals.
[0101] In an exemplary embodiment of this disclosure, the second sample signal set includes a plurality of second sample signal pairs, each second sample signal pair including a howling sample signal and a reference sample signal. The reference sample signal is an audio signal played by the playback device in the acoustic loop. Specifically, the playback device, during model training, in... Figure 5 A playback device is installed at the location of the central sound source 30 to play a reference sample signal. The howling sample signal is the superposition of the audio signal acquired by the first audio acquisition module and played by the playback device, and the second audio signal played by the second audio playback module.
[0102] Step S105: Extract the second sample audio features of the howling sample signals in the second sample signal set.
[0103] Specifically, the algorithm for extracting howling sample signals can be the same as the algorithm for extracting audio features of the audio signal to be processed.
[0104] Step S106: Input the second sample audio features into the howling detection model to obtain the howling feature parameters.
[0105] Step S107: Input the second sample audio features and the howling feature parameters into the howling suppression model, and adjust the model parameters of the howling suppression model according to the difference between the output of the howling detection model and the corresponding reference sample signal.
[0106] Specifically, by adjusting the model parameters of the feedback suppression model, the feedback suppression model can output a suppressed audio signal that is closer to the reference sample signal, thereby improving the detection performance of the feedback suppression model.
[0107] In an exemplary embodiment of this disclosure, the howling detection model is trained before the howling suppression model. Therefore, during the training of the howling suppression model, accurate howling feature parameters output by the howling detection model can be obtained, thereby improving the training efficiency and performance of the howling suppression model.
[0108] In the exemplary embodiments of this disclosure, since the training of the howling detection model and the howling suppression model requires reference sample signals without howling signals, the aforementioned first sample signal set and second sample signal set need to include howling sample signals with howling signals and reference sample signals without howling signals. Further, the second sample signal set is used to train the howling suppression model, so its howling sample signals and reference sample signals need to be paired, i.e., using the same signal source, such as... Figure 5 The acoustic loop of the instant messaging system recorded both howling and non-howling signals, and time-aligned them to serve as paired howling sample signals and reference sample signals. The howling sample signals in the first and second sample signal sets can be the same or different signals; similarly, the reference sample signals in the first and second sample signal sets can be the same or different signals.
[0109] Due to the unique nature of instant messaging scenarios, conventional datasets and signal construction methods struggle to simulate realistic howling conditions. Furthermore, there are currently no open-source datasets of this kind. Therefore, both the first and second sample signal sets require actual data acquisition. In some implementations, the howling sample signals in the first and second sample signal sets are made identical, and the reference sample signals in the first and second sample signal sets are also made identical. This reduces the number of data acquisition steps for the sample signal sets and improves data acquisition efficiency.
[0110] Furthermore, considering the complexity of instant messaging scenarios, the acquisition schemes for the first and second sample signal sets can involve different audio content, different devices, different environments, and different communication parameters, thereby improving the robustness of the howling suppression algorithm.
[0111] In an exemplary embodiment of this disclosure, the reference sample signal may be generated based on different audio content. The audio content includes one or more of speech, music, ambient sound, ringing, bird calls, and whistling, so that the reference sample signal covers different audio content.
[0112] In an exemplary embodiment of this disclosure, the first device and the second device include an audio processing module having an audio processing algorithm. For different howling sample signals, the first device and the second device have different performance and different audio processing algorithms, thereby enabling the acquisition of howling sample signals to cover devices with different performance and devices with different audio processing algorithms.
[0113] In an exemplary embodiment of this disclosure, for different howling sample signals, the spatial region where the acoustic loop is located has different noise environments. Under the same acquisition conditions, the first audio signal acquired in a spatial region with a first noise environment and the second audio signal acquired in a spatial region with a second noise environment have different signal-to-noise ratios. The same acquisition conditions include the same equipment, the same spatial region, and the same sound source. This allows the acquisition of howling sample signals to cover environments with different signal-to-noise ratios. For example, for the same conference room, different background noises can be achieved, thereby acquiring different howling sample signals in the same conference room with different background noises.
[0114] In an exemplary embodiment of this disclosure, the audio transmission parameters between the first device and the second device differ for different feedback sample signals. These audio transmission parameters include one or more of the following: the relative position between the first device and the second device, network communication parameters between the first device and the second device, and the real-time volume of the first device and the second device. This ensures that the acquisition of feedback sample signals covers acoustic loops with different audio transmission parameters.
[0115] Therefore, by collecting the aforementioned reference sample signals and howling sample signals, it is possible to cover a variety of different real-time communication situations, thereby improving the robustness of the obtained howling detection model and howling suppression model.
[0116] In the exemplary embodiments of this disclosure, the loss function of the howling suppression model is any one of the error loss function, the audio quality loss function, and the adversarial loss function. The loss function of the howling suppression model can also be a weighted sum of any number of the error loss function, the audio quality loss function, and the adversarial loss function. Each loss function is only used during model training; when using the model for howling suppression, it is not necessary to calculate the loss function.
[0117] In an exemplary embodiment of this disclosure, the loss function of the howling suppression model includes an error loss function, which is calculated based on the error between the howling-suppressed audio signal and a reference sample signal. Specifically, the error loss function can be calculated based on the MSE (mean-square error) between the howling-suppressed audio signal and the reference sample signal.
[0118] In an exemplary embodiment of this disclosure, the loss function of the howling suppression model includes an audio quality loss function, which is calculated based on the audio quality of the howling-suppressed audio signal and the audio quality of the reference sample signal. Specifically, the mean opinion score (MOS) of the howling-suppressed audio signal and the mean opinion score of the reference sample signal can be obtained separately, and the audio quality loss function is calculated using the obtained mean opinion score. The mean opinion score can be obtained, for example, through a trained audio quality scoring network model, and this disclosure is not intended to limit it.
[0119] In an exemplary embodiment of this disclosure, the loss function of the howling suppression model includes an adversarial loss function. The adversarial loss function is calculated based on the probability that the discriminator will correctly identify the output of the howling suppression model as a first signal or a second signal. The first signal represents the corresponding output as the howling suppression signal, and the second signal represents the corresponding output as a reference sample signal. When the discriminator identifies the output of the howling suppression model as the first signal, the discriminator makes a correct identification. Specifically, the purpose of the adversarial loss function is to enable the discriminator to identify the output of the howling suppression model as the second signal, that is, an audio signal that does not inherently exhibit howling. Thus, through the adversarial interaction between the howling suppression model and the discriminator, the output of the howling suppression model achieves a better suppression effect.
[0120] In an exemplary embodiment of this disclosure, the howling suppression model can be reused for noise suppression. Therefore, the step of obtaining a howling-suppressed audio signal based on the output of the howling suppression model may further include: noise cancellation of the howling-suppressed audio signal, where the noise is ambient noise in the audio signal, and the ambient noise has a fixed frequency. Ambient noise, for example, is noise with a fixed frequency such as fan or air conditioner noise. In one illustrative embodiment, the howling suppression model may be a Deep Complex Convolution Recurrent Network (DCCRN). Since both the howling suppression model and the noise suppression model are designed to eliminate specific signals, their design factors are largely similar. Therefore, this disclosure allows the howling suppression model to be reused for noise suppression, thereby utilizing a single model to perform both howling suppression and noise suppression tasks simultaneously. To achieve the reuse of the howling suppression model, this disclosure processes the sample signal set of the howling suppression model. In an exemplary embodiment of this disclosure, the howling sample signal in the second sample signal pair can be obtained according to the following steps: The reference sample signal played by the sound source, acquired by the first audio acquisition module, is superimposed with the second audio signal played by the second audio playback module to obtain a quasi-howling sample signal, wherein the reference sample signal has no noise; the quasi-howling sample signal is superimposed with a noise audio signal to generate the howling sample signal. Thus, the howling sample signal in the second sample signal pair used to train the howling suppression model contains both howling and noise, while the reference sample signal contains neither howling nor noise. Therefore, according to the training method of the howling suppression model described above, the howling suppression model can be enabled to perform howling suppression and noise suppression.
[0121] See below. Figure 11 , Figure 11 A block diagram schematically illustrates a first audio processing module of a first device according to an embodiment of the present disclosure. The first device may include a first audio processing module 15. The first audio processing module 15 has an audio processing algorithm. The audio processing algorithm may include one or more of an acoustic echo cancellation algorithm, a noise suppression algorithm, and an automatic gain control algorithm. Figure 11 In the first audio processing module 15, there are echo cancellation module 151, howl suppression module 152, noise cancellation module 153 and automatic gain module 154.
[0122] Echo cancellation module 151 is used to execute an acoustic echo cancellation algorithm, which is used to eliminate acoustic echoes in the audio signal acquired by the first audio acquisition module 11, wherein the acoustic echoes include those from the first audio playback module (e.g., the first audio playback module). Figure 5The audio signal played by label 13 is the echo signal formed by the first audio acquisition module 11.
[0123] The howling suppression module 152 is used to perform Figure 4 The method for suppressing howling is shown in the figure.
[0124] The noise cancellation module 153 is used to execute a noise suppression algorithm, which is used to suppress noise in the audio signal acquired by the first audio acquisition module. The noise is the ambient noise of the audio signal acquired by the first audio acquisition module, and the ambient noise has a fixed frequency. When the howling suppression model in the howling suppression method can be reused for noise suppression, the noise suppression module 153 can also be omitted.
[0125] The automatic gain control module 154 is used to execute an automatic gain control algorithm, which is used to adjust the volume of the audio signal acquired by the first audio acquisition module 11 to within a set volume range.
[0126] In an exemplary embodiment of this disclosure, the first device may further include a built-in audio processing module 14. The built-in audio processing module 14 is integrated into the first device. The built-in audio processing module 14 may be a non-linear processing module, and because it is device-specific and customized by various manufacturers, its audio signal processing during real-time communication is not controllable. The built-in audio processing module 14 may have an on / off switch. The built-in audio processing module 14 may also execute one or more of acoustic echo cancellation algorithms, noise suppression algorithms, and automatic gain control algorithms.
[0127] In an exemplary embodiment of this disclosure, the first audio acquisition module 11 of the first device acquires the acquired audio signal, which is then processed by the built-in audio processing module 14 (if enabled) and enters the acoustic echo cancellation module 151 of the first audio processing module 15 to perform acoustic echo cancellation on the audio signal to be processed. The audio signal to be processed after acoustic echo cancellation enters the howling suppression module 152 for howling suppression. The audio signal to be processed after howling suppression enters the noise cancellation module 153 to perform noise suppression on the howling-suppressed audio signal using the noise suppression algorithm. The audio signal to be processed after noise suppression enters the automatic gain module 154 to perform automatic gain control on the noise-suppressed audio signal using the automatic gain control algorithm. The audio signal after automatic gain control can be output to the first communication module or played directly by the first audio playback module.
[0128] Therefore, in the first audio processing module, the feedback suppression module 152 performs feedback suppression after the acoustic echo cancellation module 151 performs echo cancellation to prevent interference from the echo signal. At the same time, the feedback suppression module 152 performs feedback suppression before the noise cancellation module 153 performs noise suppression to prevent the noise cancellation module 153 from causing further damage to the feedback signal, so as not to reduce the accuracy of feedback detection and the effect of feedback suppression.
[0129] The above is merely an illustrative description of various embodiments provided in this disclosure. This disclosure is not intended to limit the scope of the disclosure. Each embodiment can be used individually or in combination.
[0130] Exemplary device
[0131] After introducing the howling suppression method according to exemplary embodiments of this disclosure, the following will refer to... Figure 12 The following describes a feedback suppression device according to an exemplary embodiment of the present disclosure. The feedback suppression device is applied to a first device, which is used for real-time communication with a second device. The first device and the second device belong to the same acoustic loop. The first device includes a first communication module, a first audio acquisition module, and a first audio playback module. The second device includes a second communication module, a second audio acquisition module, and a second audio playback module.
[0132] refer to Figure 12 As shown, the feedback suppression device 300 of this exemplary embodiment may include: an audio feature extraction module 310, a feedback detection module 320, a feedback suppression input module 330, and a feedback suppression output module 340.
[0133] The audio feature extraction module 310 can be used to extract audio features of the audio signal to be processed. The audio signal to be processed is the audio signal acquired by the first device through its first audio acquisition module. The audio signal to be processed is the superposition of the acoustic signal emitted by the sound source and the second audio signal played by the second audio playback module.
[0134] The feedback detection module 320 can be used to input the audio features into the feedback detection model, and the feedback detection model outputs the feedback feature parameters of the audio signal to be processed;
[0135] The feedback suppression input module 330 can be used to input the feedback feature parameters and the audio features into the feedback suppression model;
[0136] The howling suppression output module 340 can be used to obtain a howling suppression audio signal based on the output of the howling suppression model.
[0137] According to an exemplary embodiment of this disclosure, the acoustic loop is an audio signal transmission path based on the first audio acquisition module, the first communication module, the second communication module, and the second audio playback module. The audio signal to be processed is transmitted in a closed loop in the acoustic loop via the first audio acquisition module, the first communication module, the second communication module, the second audio playback module, and the first audio acquisition module in sequence.
[0138] According to an exemplary embodiment of this disclosure, the second audio playback module is located within the pickup distance of the first audio acquisition module, and the playback volume of the second audio playback module is sufficient to enable the second audio signal to be picked up by the first audio acquisition module.
[0139] According to an exemplary embodiment of this disclosure, the audio feature is one of the following: spectral features, Bark spectrum features, Mel spectrum features, Mel cepstral features, and fundamental frequency features of the audio signal to be processed.
[0140] According to an exemplary embodiment of this disclosure, the howling detection model sequentially includes an input processing layer, a first intermediate processing layer, and a classification output layer. The audio features are assigned to the input processing layer and input to the first intermediate processing layer via the input processing layer. The first intermediate processing layer is used to obtain local features of the audio features and use the local features as howling intermediate features. The classification output layer is used to classify the howling intermediate features to obtain howling result features. The howling intermediate features and the howling result features are used as howling feature parameters input to the howling suppression model.
[0141] According to an exemplary embodiment of this disclosure, the howling detection model further includes a backbone layer and a recurrent layer connected between the input processing layer and the intermediate processing layer. The backbone layer is used to perform convolution and / or pooling processing on the data input to the backbone layer. The recurrent layer is used to establish the correlation between the input data of the recurrent layer and the output data of the recurrent layer. The howling detection model further includes an attention layer connected between the first intermediate processing layer and the classification output layer. The attention layer is used to perform weighted summation on the data input to the attention layer.
[0142] According to an exemplary embodiment of the present disclosure, the howling suppression model sequentially includes an encoder, a second intermediate processing layer, and a decoder. The encoder is used to perform feature encoding on the audio features and the howling result features to obtain encoded features. The second intermediate processing layer is used to perform feature filtering on the encoded features and the howling intermediate features. The decoder is used to perform feature decoding on the filtered encoded features.
[0143] According to an exemplary embodiment of the present disclosure, the second intermediate processing layer sequentially includes a plurality of connected long short-term memory units and a fully connected layer. The long short-term memory units are used to perform feature filtering on the encoded features, and the fully connected layer is used to perform weighted summation on the outputs of the plurality of long short-term memory units.
[0144] According to an exemplary embodiment of this disclosure, the howling detection model is trained through the following steps:
[0145] A first sample signal set is obtained, which includes multiple first sample signals and howling characteristic parameters of the first sample signals. The howling characteristic parameters include the same parameter items as the howling result features. The first sample signals include reference sample signals and howling sample signals. The reference sample signals are audio signals played by the playback device in the acoustic loop. The howling sample signals are the superposition of the audio signals acquired by the first audio acquisition module and played by the playback device and the second audio signals played by the second audio playback module. First sample audio features of the first sample signals are extracted. The first sample audio features are used as input to the howling detection model. The model parameters of the howling detection model are adjusted according to the difference between the output of the howling detection model and the corresponding howling characteristic parameters.
[0146] According to an exemplary embodiment of this disclosure, the howling suppression model is trained through the following steps: acquiring a second sample signal set, the second sample signal set including multiple second sample signal pairs, each second sample signal pair including a howling sample signal and a reference sample signal, the reference sample signal being an audio signal played by a playback device in the acoustic loop, the howling sample signal being the superposition of an audio signal acquired by the first audio acquisition module and played by the playback device and a second audio signal played by the second audio playback module; extracting second sample audio features from the howling sample signals in the second sample signal set; inputting the second sample audio features into the howling detection model to obtain the howling feature parameters; inputting the second sample audio features and the howling feature parameters into the howling suppression model, and adjusting the model parameters of the howling suppression model according to the difference between the output of the howling detection model and the corresponding reference sample signal.
[0147] According to an exemplary embodiment of this disclosure, the loss function of the howling suppression model includes an error loss function, which is calculated based on the error between the howling suppressed audio signal and the reference sample signal.
[0148] According to an exemplary embodiment of this disclosure, the loss function of the howling suppression model includes an audio quality loss function, which is calculated based on the audio quality of the howling suppressed audio signal and the audio quality of the reference sample signal.
[0149] According to an exemplary embodiment of this disclosure, the loss function of the howling suppression model includes an adversarial loss function, which is calculated based on the probability that the discriminator makes a correct judgment. The discriminator is used to classify the output result of the howling suppression model as a first signal or a second signal. The first signal represents that the corresponding output result is the howling suppression signal, and the second signal represents that the corresponding output result is a reference sample signal. When the discriminator classifies the output result of the howling suppression model as the first signal, the discriminator makes a correct judgment.
[0150] According to an exemplary embodiment of this disclosure, the loss function of the howling suppression model is a weighted sum of an error loss function, an audio quality loss function, and an adversarial loss function.
[0151] According to an exemplary embodiment of this disclosure, the howling detection model is trained prior to the howling suppression model.
[0152] According to an exemplary embodiment of this disclosure, the howling suppression output module further includes: a first noise cancellation module, configured to cancel noise in the howling suppression audio signal, wherein the noise is ambient noise in the audio signal, and the ambient noise has a fixed frequency.
[0153] According to an exemplary embodiment of this disclosure, the howling sample signal in the second sample signal pair is obtained by the following steps: superimposing the reference sample signal played by the sound source acquired by the first audio acquisition module with the second audio signal played by the second audio playback module as a quasi-howling sample signal, wherein the reference sample signal has no noise; superimposing the quasi-howling sample signal with the noise audio signal to generate the howling sample signal.
[0154] According to an exemplary embodiment of this disclosure, the howling suppression model is a deep complex convolutional recurrent network.
[0155] According to an exemplary embodiment of this disclosure, the howling result features include one or more of the following: howling detection result, howling level, howling type, howling continuity, and frequency point shift parameters.
[0156] According to an exemplary embodiment of this disclosure, the feedback detection result is used to indicate whether feedback exists in the audio features input to the feedback detection model, the feedback level is used to indicate the feedback intensity of the audio features input to the feedback detection model, the feedback type includes single-frequency feedback, multi-frequency feedback, and diffuse feedback, the feedback continuity includes continuous feedback and intermittent feedback, the frequency shift parameter includes a frequency shift type parameter and a frequency shift amplitude parameter, the frequency shift type parameter is used to indicate whether frequency shift exists in the audio features input to the feedback detection model, and the frequency shift amplitude parameter is used to indicate the amplitude of the frequency shift of the audio features input to the feedback detection model.
[0157] According to an exemplary embodiment of this disclosure, the reference sample signal is generated based on different audio content, including one or more of speech, music, ambient sound, ringing, bird calls, and whistling.
[0158] According to an exemplary embodiment of this disclosure, the first device and the second device include an audio processing module having an audio processing algorithm. For different howling sample signals, the first device and the second device have different performance and different audio processing algorithms.
[0159] According to an exemplary embodiment of this disclosure, for different howling sample signals, the spatial region where the acoustic loop is located has different noise environments. Under the same acquisition conditions, the first audio signal acquired in the spatial region with a first noise environment and the second audio signal acquired in the spatial region with a second noise environment have different signal-to-noise ratios. The same acquisition conditions include the same equipment, the same spatial region, and the same sound source.
[0160] According to an exemplary embodiment of this disclosure, the audio transmission parameters between the first device and the second device are different for different howling sample signals. The audio transmission parameters include one or more of the following: the relative position between the first device and the second device, the network communication parameters between the first device and the second device, and the real-time volume of the first device and the second device.
[0161] According to an exemplary embodiment of the present disclosure, the howling suppression output module includes: a first feature restoration module, configured to perform an inverse operation relative to feature extraction on the suppressed audio features output by the howling suppression model, and restore the suppressed audio features to obtain the howling suppression audio signal.
[0162] According to an exemplary embodiment of this disclosure, the howling suppression output module includes: a mask acquisition module, configured to acquire a howling suppression mask of the output of the howling suppression model, the howling suppression mask being used to characterize the howling suppression frequency gain of the audio features of a reference sample signal relative to the audio features of the audio signal to be processed; a suppression feature acquisition module, configured to multiply the howling suppression mask with the audio features of the audio signal to be processed to obtain suppressed audio features; and a second feature restoration module, configured to perform an inverse operation relative to feature extraction on the suppressed audio features to restore the suppressed audio features to obtain the howling suppression audio signal.
[0163] According to an exemplary embodiment of this disclosure, the first device includes an audio processing module, the audio processing module having an audio processing algorithm, the audio processing algorithm including one or more of an acoustic echo cancellation algorithm, a noise suppression algorithm, and an automatic gain control algorithm, wherein the acoustic echo cancellation algorithm is used to eliminate acoustic echoes in the audio signal acquired by the first audio acquisition module, the acoustic echoes including echo signals formed by the first audio playback module playing an audio signal and being acquired by the first audio acquisition module; the noise suppression algorithm is used to suppress noise in the audio signal acquired by the first audio acquisition module, the noise being ambient noise of the audio signal acquired by the first audio acquisition module, the ambient noise having a fixed frequency; the automatic gain control algorithm is used to adjust the volume of the audio signal acquired by the first audio acquisition module to within a set volume range.
[0164] According to an exemplary embodiment of the present disclosure, the audio processing module further includes: an echo cancellation module for performing acoustic echo cancellation on the audio signal to be processed using the acoustic echo cancellation algorithm; a noise suppression module for performing noise suppression on the feedback-suppressed audio signal using the noise suppression algorithm; and an automatic gain control module for performing automatic gain control on the noise-suppressed audio signal using the automatic gain control algorithm.
[0165] Since the functional modules of the howling suppression device in this embodiment are the same as those in the above-described howling suppression method, they will not be described again here.
[0166] Exemplary storage media
[0167] After introducing the howling suppression method and apparatus according to exemplary embodiments of the present disclosure, the following references are made. Figure 13 The storage medium of the exemplary embodiments of this disclosure will be described.
[0168] refer to Figure 13As shown, a program product 1000 for implementing the above-described method according to an embodiment of the present disclosure is described. This product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a device such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0169] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0170] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0171] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0172] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0173] Exemplary electronic devices
[0174] Having described the storage medium of exemplary embodiments of this disclosure, the following references are made. Figure 14 An electronic device according to an exemplary embodiment of the present disclosure will be described.
[0175] Figure 14 The electronic device 800 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0176] like Figure 14 As shown, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one processing unit 810, at least one storage unit 820, a bus 830 connecting different system components (including storage unit 820 and processing unit 810), and a display unit 840.
[0177] The storage unit stores program code that can be executed by the processing unit 810, causing the processing unit 810 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 810 can perform actions such as... Figure 4 The steps are shown in the figure.
[0178] Storage unit 820 may include volatile storage units, such as random access memory (RAM) 8201 and / or cache memory 8202, and may further include read-only memory (ROM) 8203.
[0179] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0180] Bus 830 may include a data bus, an address bus, and a control bus.
[0181] Electronic device 800 can also communicate with one or more external devices 900 (e.g., keyboard, pointing device, Bluetooth device, etc.) via input / output (I / O) interface 850. Electronic device 800 also includes a display unit 840 connected to input / output (I / O) interface 850 for display purposes. Furthermore, electronic device 800 can communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 860. As shown, network adapter 860 communicates with other modules of electronic device 800 via bus 830. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0182] It should be noted that although several modules or sub-modules of the whistling suppression device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0183] Furthermore, although the operations of the methods disclosed herein are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0184] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A method for suppressing howling, characterized in that, The method is applied to a first device for real-time communication with a second device, the first device and the second device belonging to the same acoustic loop, the first device including a first communication module, a first audio acquisition module and a first audio playback module, and the second device including a second communication module, a second audio acquisition module and a second audio playback module. Extract the audio features of the audio signal to be processed, wherein the audio signal to be processed is the audio signal acquired by the first device through its first audio acquisition module, and the audio signal to be processed is the superposition of the acoustic signal emitted by the sound source and the second audio signal played by the second audio playback module; The audio features are input into the feedback detection model, and the feedback detection model outputs the feedback feature parameters of the audio signal to be processed. The feedback feature parameters and the audio features are input into the feedback suppression model; Based on the output of the aforementioned howling suppression model, a howling suppression audio signal is obtained; The feedback detection model includes an input processing layer, a first intermediate processing layer, and a classification output layer. The audio features are assigned to the input processing layer and then input to the first intermediate processing layer. The first intermediate processing layer is used to obtain local features of the audio features and uses the local features as feedback intermediate features. The classification output layer is used to classify the feedback intermediate features to obtain feedback result features. The feedback intermediate features and feedback result features are used as feedback feature parameters input to the feedback suppression model. The feedback suppression model includes an encoder, a second intermediate processing layer, and a decoder. The encoder is used to perform feature encoding on the audio features and the feedback result features to obtain encoded features. The second intermediate processing layer is used to perform feature filtering on the encoded features and the feedback intermediate features. The decoder is used to perform feature decoding on the filtered encoded features.
2. The howling suppression method according to claim 1, characterized in that, The acoustic loop is an audio signal transmission path based on the first audio acquisition module, the first communication module, the second communication module, and the second audio playback module. The audio signal to be processed is transmitted in a closed loop through the first audio acquisition module, the first communication module, the second communication module, the second audio playback module, and the first audio acquisition module in sequence.
3. The howling suppression method according to claim 2, characterized in that, The second audio playback module is located within the pickup distance of the first audio acquisition module, and the playback volume of the second audio playback module is sufficient to enable the second audio signal to be picked up by the first audio acquisition module.
4. The howling suppression method according to claim 1, characterized in that, The audio feature is one of the following: Bark spectrum feature, Mel spectrum feature, Mel cepstral feature, and fundamental frequency feature of the audio signal to be processed.
5. The howling suppression method according to claim 1, characterized in that, The howling detection model further includes a backbone layer and a recurrent layer connected between the input processing layer and the intermediate processing layer. The backbone layer is used to perform convolution and / or pooling processing on the data input to the backbone layer, and the recurrent layer is used to establish the correlation between the input data and the output data of the recurrent layer. The howling detection model further includes an attention layer connected between the first intermediate processing layer and the classification output layer, the attention layer being used to perform weighted summation on the data input to the attention layer.
6. The howling suppression method according to claim 1, characterized in that, The second intermediate processing layer sequentially includes multiple connected long short-term memory units and a fully connected layer. The long short-term memory units are used to perform feature filtering on the encoded features, and the fully connected layer is used to perform weighted summation on the outputs of the multiple long short-term memory units.
7. The howling suppression method according to claim 1, characterized in that, The howling detection model is trained through the following steps: A first sample signal set is acquired, which includes multiple first sample signals and howling characteristic parameters of the first sample signals. The howling characteristic parameters include the same parameter items as the howling result features. The first sample signal includes a reference sample signal and a howling sample signal. The reference sample signal is the audio signal played by the playback device in the acoustic loop. The howling sample signal is the superposition of the audio signal acquired by the first audio acquisition module and played by the playback device and the second audio signal played by the second audio playback module. Extract the first sample audio features from the first sample signal; The first sample audio features are used as input to the howling detection model. Based on the difference between the output of the howling detection model and the corresponding howling characteristic parameters, the model parameters of the howling detection model are adjusted.
8. The howling suppression method according to claim 1, characterized in that, The howling suppression model is trained through the following steps: Acquire a second sample signal set, which includes multiple second sample signal pairs. Each second sample signal pair includes a howling sample signal and a reference sample signal. The reference sample signal is the audio signal played by the playback device in the acoustic loop. The howling sample signal is the superposition of the audio signal acquired by the first audio acquisition module and played by the playback device and the second audio signal played by the second audio playback module. Extract the second sample audio features of the howling sample signals from the second sample signal set; The second sample audio features are input into the howling detection model to obtain the howling feature parameters; The second sample audio features and the howling feature parameters are input into the howling suppression model. Based on the difference between the output of the howling detection model and the corresponding reference sample signal, the model parameters of the howling suppression model are adjusted.
9. The howling suppression method according to claim 8, characterized in that, The loss function of the howling suppression model includes an error loss function, which is calculated based on the error between the howling suppressed audio signal and the reference sample signal.
10. The howling suppression method according to claim 9, characterized in that, The loss function of the howling suppression model includes an audio quality loss function, which is calculated based on the audio quality of the howling suppressed audio signal and the audio quality of the reference sample signal.
11. The howling suppression method according to claim 9, characterized in that, The loss function of the howling suppression model includes an adversarial loss function, which is calculated based on the probability that the discriminator makes a correct judgment. The discriminator is used to distinguish the output result of the howling suppression model as a first signal or a second signal. The first signal represents that the corresponding output result is a howling suppression signal, and the second signal represents that the corresponding output result is a reference sample signal. When the discriminator distinguishes the output result of the howling suppression model as the first signal, the discriminator makes a correct judgment.
12. The howling suppression method according to claim 9, characterized in that, The loss function of the feedback suppression model is a weighted sum of the error loss function, the audio quality loss function, and the adversarial loss function.
13. The howling suppression method according to claim 9, characterized in that, The howling detection model is trained before the howling suppression model.
14. The howling suppression method according to claim 9, characterized in that, The step of obtaining the howling-suppressed audio signal based on the output of the howling suppression model further includes: The feedback suppression audio signal is subjected to noise cancellation, wherein the noise is ambient noise in the audio signal and the ambient noise has a fixed frequency.
15. The howling suppression method according to claim 14, characterized in that, The howling sample signal in the second sample signal pair is obtained according to the following steps: The reference sample signal played by the sound source and acquired by the first audio acquisition module is superimposed with the second audio signal played by the second audio playback module to form a quasi-feedback sample signal. The reference sample signal has no noise. The quasi-howling sample signal is superimposed with the noise audio signal to generate the howling sample signal.
16. The howling suppression method according to claim 15, characterized in that, The howling suppression model is a deep complex convolutional recurrent network.
17. The howling suppression method according to any one of claims 5 to 16, characterized in that, The characteristics of the howling result include one or more of the following: howling detection result, howling level, howling type, howling continuity, and frequency shift parameters.
18. The howling suppression method according to claim 17, characterized in that, The feedback detection result is used to indicate whether feedback exists in the audio features input to the feedback detection model. The feedback level is used to indicate the feedback intensity of the audio features input to the feedback detection model. The feedback type includes single-frequency feedback, multi-frequency feedback, and diffuse feedback. The feedback continuity includes continuous feedback and intermittent feedback. The frequency shift parameter includes a frequency shift type parameter and a frequency shift amplitude parameter. The frequency shift type parameter is used to indicate whether frequency shift exists in the audio features input to the feedback detection model. The frequency shift amplitude parameter is used to indicate the magnitude of the frequency shift of the audio features input to the feedback detection model.
19. The howling suppression method according to any one of claims 7 to 16, characterized in that, The reference sample signal is generated based on different audio content, which includes one or more of speech, music, and ambient sound.
20. The howling suppression method according to any one of claims 7 to 16, characterized in that, The first device and the second device include an audio processing module, which has an audio processing algorithm. For different howling sample signals, the first device and the second device have different performance and different audio processing algorithms.
21. The howling suppression method according to any one of claims 7 to 16, characterized in that, For different feedback sample signals, the spatial region where the acoustic loop is located has different noise environments. Under the same acquisition conditions, the first audio signal acquired in the spatial region with the first noise environment and the second audio signal acquired in the spatial region with the second noise environment have different signal-to-noise ratios. The same acquisition conditions include the same equipment, the same spatial region, and the same sound source.
22. The howling suppression method according to any one of claims 7 to 16, characterized in that, For different feedback sample signals, the audio transmission parameters between the first device and the second device are different. The audio transmission parameters include one or more of the following: the relative position between the first device and the second device, the network communication parameters between the first device and the second device, and the real-time volume of the first device and the second device.
23. The howling suppression method according to any one of claims 1 to 16, characterized in that, Obtaining the howling-suppressed audio signal based on the output of the howling suppression model includes: The suppressed audio features output by the howling suppression model are subjected to an inverse operation relative to feature extraction to restore the suppressed audio features and obtain the howling suppression audio signal.
24. The howling suppression method according to any one of claims 1 to 16, characterized in that, Obtaining the howling-suppressed audio signal based on the output of the howling suppression model includes: Obtain the howling suppression mask of the output of the howling suppression model. The howling suppression mask is used to characterize the howling suppression frequency gain of the audio features of the reference sample signal relative to the audio features of the audio signal to be processed. The feedback suppression mask is multiplied with the audio features of the audio signal to be processed to obtain suppressed audio features; The suppressed audio features are then subjected to an inverse operation relative to feature extraction to restore the suppressed audio features and obtain the howling suppressed audio signal.
25. The howling suppression method according to any one of claims 2 to 13, characterized in that, The first device includes an audio processing module, which has an audio processing algorithm, including one or more of acoustic echo cancellation algorithm, noise suppression algorithm, and automatic gain control algorithm. The acoustic echo cancellation algorithm is used to eliminate acoustic echoes in the audio signal acquired by the first audio acquisition module. The acoustic echoes include the echo signals formed by the first audio playback module playing the audio signal and being acquired by the first audio acquisition module. The noise suppression algorithm is used to suppress noise in the audio signal acquired by the first audio acquisition module. The noise is the environmental noise of the audio signal acquired by the first audio acquisition module, and the environmental noise has a fixed frequency. The automatic gain control algorithm is used to adjust the volume of the audio signal acquired by the first audio acquisition module to within a set volume range.
26. The howling suppression method according to claim 25, characterized in that, Before extracting the audio features of the audio signal to be processed, the process also includes: The acoustic echo cancellation algorithm described above is used to perform acoustic echo cancellation on the audio signal to be processed. After obtaining the howling-suppressed audio signal based on the output of the howling suppression model, the process further includes: The noise suppression algorithm described above is used to suppress noise in the feedback-suppressing audio signal; and The automatic gain control algorithm described above is used to perform automatic gain control on the noise-suppressed audio signal.
27. A whistling suppression device, characterized in that, An apparatus is applied to a first device for real-time communication with a second device, wherein the first device and the second device belong to the same acoustic loop. The first device includes a first communication module, a first audio acquisition module, and a first audio playback module, and the second device includes a second communication module, a second audio acquisition module, and a second audio playback module. The apparatus comprises: An audio feature extraction module is used to extract audio features of an audio signal to be processed. The audio signal to be processed is an audio signal acquired by the first device through its first audio acquisition module. The audio signal to be processed is the superposition of an acoustic signal emitted by a sound source and a second audio signal played by the second audio playback module. The feedback detection module is used to input the audio features into the feedback detection model, and the feedback detection model outputs the feedback feature parameters of the audio signal to be processed; The feedback suppression input module is used to input the feedback feature parameters and the audio features into the feedback suppression model; The feedback suppression output module is used to obtain the feedback suppression audio signal based on the output of the feedback suppression model; The feedback detection model includes an input processing layer, a first intermediate processing layer, and a classification output layer. The audio features are assigned to the input processing layer and then input to the first intermediate processing layer. The first intermediate processing layer is used to obtain local features of the audio features and uses the local features as feedback intermediate features. The classification output layer is used to classify the feedback intermediate features to obtain feedback result features. The feedback intermediate features and feedback result features are used as feedback feature parameters input to the feedback suppression model. The feedback suppression model includes an encoder, a second intermediate processing layer, and a decoder. The encoder is used to perform feature encoding on the audio features and the feedback result features to obtain encoded features. The second intermediate processing layer is used to perform feature filtering on the encoded features and the feedback intermediate features. The decoder is used to perform feature decoding on the filtered encoded features.
28. The whistling suppression device according to claim 27, characterized in that, The acoustic loop is an audio signal transmission path based on the first audio acquisition module, the first communication module, the second communication module, and the second audio playback module. The audio signal to be processed is transmitted in a closed loop through the first audio acquisition module, the first communication module, the second communication module, the second audio playback module, and the first audio acquisition module in sequence.
29. The whistling suppression device according to claim 28, characterized in that, The second audio playback module is located within the pickup distance of the first audio acquisition module, and the playback volume of the second audio playback module is sufficient to enable the second audio signal to be picked up by the first audio acquisition module.
30. The whistling suppression device according to claim 27, characterized in that, The audio feature is one of the following: Bark spectrum feature, Mel spectrum feature, Mel cepstral feature, and fundamental frequency feature of the audio signal to be processed.
31. The whistling suppression device according to claim 27, characterized in that, The howling detection model further includes a backbone layer and a recurrent layer connected between the input processing layer and the intermediate processing layer. The backbone layer is used to perform convolution and / or pooling processing on the data input to the backbone layer, and the recurrent layer is used to establish the correlation between the input data and the output data of the recurrent layer. The howling detection model further includes an attention layer connected between the first intermediate processing layer and the classification output layer, the attention layer being used to perform weighted summation on the data input to the attention layer.
32. The whistling suppression device according to claim 27, characterized in that, The second intermediate processing layer sequentially includes multiple connected long short-term memory units and a fully connected layer. The long short-term memory units are used to perform feature filtering on the encoded features, and the fully connected layer is used to perform weighted summation on the outputs of the multiple long short-term memory units.
33. The whistling suppression device according to claim 27, characterized in that, The howling detection model is trained through the following steps: A first sample signal set is acquired, which includes multiple first sample signals and howling characteristic parameters of the first sample signals. The howling characteristic parameters include the same parameter items as the howling result features. The first sample signal includes a reference sample signal and a howling sample signal. The reference sample signal is the audio signal played by the playback device in the acoustic loop. The howling sample signal is the superposition of the audio signal acquired by the first audio acquisition module and played by the playback device and the second audio signal played by the second audio playback module. Extract the first sample audio features from the first sample signal; The first sample audio features are used as input to the howling detection model. Based on the difference between the output of the howling detection model and the corresponding howling characteristic parameters, the model parameters of the howling detection model are adjusted.
34. The whistling suppression device according to claim 27, characterized in that, The howling suppression model is trained through the following steps: Acquire a second sample signal set, which includes multiple second sample signal pairs. Each second sample signal pair includes a howling sample signal and a reference sample signal. The reference sample signal is the audio signal played by the playback device in the acoustic loop. The howling sample signal is the superposition of the audio signal acquired by the first audio acquisition module and played by the playback device and the second audio signal played by the second audio playback module. Extract the second sample audio features of the howling sample signals from the second sample signal set; The second sample audio features are input into the howling detection model to obtain the howling feature parameters; The second sample audio features and the howling feature parameters are input into the howling suppression model. Based on the difference between the output of the howling detection model and the corresponding reference sample signal, the model parameters of the howling suppression model are adjusted.
35. The whistling suppression device according to claim 34, characterized in that, The loss function of the howling suppression model includes an error loss function, which is calculated based on the error between the howling suppressed audio signal and the reference sample signal.
36. The whistling suppression device according to claim 34, characterized in that, The loss function of the howling suppression model includes an audio quality loss function, which is calculated based on the audio quality of the howling suppressed audio signal and the audio quality of the reference sample signal.
37. The whistling suppression device according to claim 34, characterized in that, The loss function of the howling suppression model includes an adversarial loss function, which is calculated based on the probability that the discriminator makes a correct judgment. The discriminator is used to distinguish the output result of the howling suppression model as a first signal or a second signal. The first signal represents that the corresponding output result is a howling suppression signal, and the second signal represents that the corresponding output result is a reference sample signal. When the discriminator distinguishes the output result of the howling suppression model as the first signal, the discriminator makes a correct judgment.
38. The whistling suppression device according to claim 34, characterized in that, The loss function of the feedback suppression model is a weighted sum of the error loss function, the audio quality loss function, and the adversarial loss function.
39. The whistling suppression device according to claim 34, characterized in that, The howling detection model is trained before the howling suppression model.
40. The whistling suppression device according to claim 34, characterized in that, The howling suppression output module also includes: The first noise cancellation module is used to cancel noise in the howling suppression audio signal, wherein the noise is ambient noise in the audio signal and the ambient noise has a fixed frequency.
41. The whistling suppression device according to claim 40, characterized in that, The howling sample signal in the second sample signal pair is obtained according to the following steps: The reference sample signal played by the sound source and acquired by the first audio acquisition module is superimposed with the second audio signal played by the second audio playback module to form a quasi-feedback sample signal. The reference sample signal has no noise. The quasi-howling sample signal is superimposed with the noise audio signal to generate the howling sample signal.
42. The whistling suppression device according to claim 41, characterized in that, The howling suppression model is a deep complex convolutional recurrent network.
43. The whistling suppression device according to any one of claims 27 to 42, characterized in that, The characteristics of the howling result include one or more of the following: howling detection result, howling level, howling type, howling continuity, and frequency shift parameters.
44. The whistling suppression device according to claim 43, characterized in that, The feedback detection result is used to indicate whether feedback exists in the audio features input to the feedback detection model. The feedback level is used to indicate the feedback intensity of the audio features input to the feedback detection model. The feedback type includes single-frequency feedback, multi-frequency feedback, and diffuse feedback. The feedback continuity includes continuous feedback and intermittent feedback. The frequency shift parameter includes a frequency shift type parameter and a frequency shift amplitude parameter. The frequency shift type parameter is used to indicate whether frequency shift exists in the audio features input to the feedback detection model. The frequency shift amplitude parameter is used to indicate the magnitude of the frequency shift of the audio features input to the feedback detection model.
45. The whistling suppression device according to any one of claims 33 to 42, characterized in that, The reference sample signal is generated based on different audio content, which includes one or more of speech, music, and ambient sound.
46. The whistling suppression device according to any one of claims 33 to 42, characterized in that, The first device and the second device include an audio processing module, which has an audio processing algorithm. For different howling sample signals, the first device and the second device have different performance and different audio processing algorithms.
47. The whistling suppression device according to any one of claims 33 to 42, characterized in that, For different feedback sample signals, the spatial region where the acoustic loop is located has different noise environments. Under the same acquisition conditions, the first audio signal acquired in the spatial region with the first noise environment and the second audio signal acquired in the spatial region with the second noise environment have different signal-to-noise ratios. The same acquisition conditions include the same equipment, the same spatial region, and the same sound source.
48. The whistling suppression device according to any one of claims 33 to 42, characterized in that, For different feedback sample signals, the audio transmission parameters between the first device and the second device are different. The audio transmission parameters include one or more of the following: the relative position between the first device and the second device, the network communication parameters between the first device and the second device, and the real-time volume of the first device and the second device.
49. The whistling suppression device according to any one of claims 27 to 42, characterized in that, The howling suppression output module includes: The first feature restoration module is used to perform the inverse operation relative to feature extraction on the suppressed audio features output by the howling suppression model, and restore the suppressed audio features to obtain the howling suppression audio signal.
50. The whistling suppression device according to any one of claims 27 to 42, characterized in that, The howling suppression output module includes: The mask acquisition module is used to acquire the howling suppression mask of the output of the howling suppression model. The howling suppression mask is used to characterize the howling suppression frequency gain of the audio features of the reference sample signal compared with the audio features of the audio signal to be processed. The suppression feature acquisition module is used to multiply the howling suppression mask with the audio features of the audio signal to be processed to obtain suppressed audio features; The second feature restoration module is used to perform the inverse operation relative to feature extraction on the suppressed audio features, and restore the suppressed audio features to obtain the howling suppressed audio signal.
51. The whistling suppression device according to any one of claims 28 to 39, characterized in that, The first device includes an audio processing module, which has an audio processing algorithm, including one or more of acoustic echo cancellation algorithm, noise suppression algorithm, and automatic gain control algorithm. The acoustic echo cancellation algorithm is used to eliminate acoustic echoes in the audio signal acquired by the first audio acquisition module. The acoustic echoes include the echo signals formed by the first audio playback module playing the audio signal and being acquired by the first audio acquisition module. The noise suppression algorithm is used to suppress noise in the audio signal acquired by the first audio acquisition module. The noise is the environmental noise of the audio signal acquired by the first audio acquisition module, and the environmental noise has a fixed frequency. The automatic gain control algorithm is used to adjust the volume of the audio signal acquired by the first audio acquisition module to within a set volume range.
52. The whistling suppression device according to claim 51, characterized in that, The audio processing module further includes: The echo cancellation module is used to perform acoustic echo cancellation on the audio signal to be processed using the acoustic echo cancellation algorithm. A noise suppression module is used to suppress noise in the feedback-suppressing audio signal using the aforementioned noise suppression algorithm; and An automatic gain control module is used to perform automatic gain control on the noise-suppressed audio signal using the automatic gain control algorithm.
53. A storage medium having a computer program stored thereon, characterized in that, The computer program is executed by the processor to achieve the following: The howling suppression method according to any one of claims 1 to 26.
54. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the executable instructions: The howling suppression method according to any one of claims 1 to 26.
Citation Information
Patent Citations
Howling detection method and device, medium and computing equipment
CN114067837A