Howling suppression method, device and equipment and readable storage medium
By training and fine-tuning the neural network model online and using real-time data to generate augmented signals to expand the database, the problem of poor performance of traditional howling suppression methods is solved, and a highly efficient howling suppression effect is achieved.
Patent Information
- Application Number
- CN202410626767.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-11-21
AI Technical Summary
Existing methods for suppressing howling cannot completely eliminate howling. Traditional deep learning-based solutions use offline data that does not match the actual real-time process, resulting in poor howling suppression performance.
The neural network model is trained and fine-tuned using online real-time data. The initial model is trained by using sample howling signals from the first howling database, augmented signals are generated and the database is expanded, and the second neural network model is fine-tuned to obtain the model. Finally, the model is used for howling suppression of the target speech signal.
This greatly improves the effect of howling suppression, enabling the neural network model to match the actual real-time process and achieving efficient howling suppression.
Smart Images

Figure CN120998215A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of voice processing, and particularly relate to a howling suppression method, device, equipment and readable storage medium. BACKGROUND
[0002] In a sound amplification system, the voice signal collected by a microphone is transmitted to a loudspeaker for amplification and broadcasting, and the voice signal played by the loudspeaker is picked up by the microphone again. The transmission and feedback of the voice signal between the loudspeaker and the microphone form an acoustic loop.
[0003] During the transmission of the voice signal, when the volume is large, the feedback loop of the sound forms positive feedback, that is, the acoustic loop gain is greater than 1. The sound is amplified step by step in the continuous feedback and produces a piercing howling sound, which seriously affects the auditory experience of the user. To suppress the howling, common methods include frequency shifting and phase shifting method, notch method, adaptive feedback control method, etc. Among them, the frequency shifting and phase shifting method mainly changes the frequency or phase of the sound in real time during sound processing to destroy the phase characteristics required for positive feedback to occur. The notch method forcibly lowers the acoustic loop gain of the frequency point where the howling occurs through a notch filter to suppress the howling. The adaptive feedback control method needs to converge to a very deep depth to suppress the howling.
[0004] However, the above howling suppression methods have poor quality and cannot completely eliminate the howling. SUMMARY
[0005] Embodiments of the present application provide a howling suppression method, device, equipment and readable storage medium, the training and fine-tuning of the model both use online real-time data to achieve the purpose of improving the howling suppression performance of the second neural network model, thereby greatly improving the effect of howling suppression.
[0006] In a first aspect, embodiments of the present application provide a howling suppression method, comprising:
[0007] training an initial model using sample howling signals in a first howling database to obtain a first neural network model;
[0008] generating an augmented signal using the first neural network model;
[0009] fine-tuning the first neural network model using sample howling signals in a second howling database to obtain a second neural network model, the second howling database being obtained by adding the augmented signal to the first howling database;
[0010] suppressing howling of a target voice signal using the second neural network model.
[0011] In a second aspect, embodiments of the present application provide a howling suppression method, comprising:
[0012] obtaining a target speech signal;
[0013] processing the target speech signal by using a second neural network model to obtain a to-be-played signal, the second neural network model being obtained by fine-tuning a first neural network model by using sample howling signals in a second howling database, the first neural network model being obtained by training an initial model by using sample howling signals in a first howling database, the second howling database being obtained by adding an augmented signal to the first howling database, the augmented signal being a signal obtained by using the first neural network model;
[0014] playing the to-be-played signal.
[0015] In a third aspect, an embodiment of the present application provides a howling suppression apparatus, comprising:
[0016] a training module configured to train an initial model by using sample howling signals in a first howling database to obtain a first neural network model;
[0017] a processing module configured to generate an augmented signal by using the first neural network model;
[0018] a fine-tuning module configured to fine-tune the first neural network model by using sample howling signals in a second howling database to obtain a second neural network model, the second howling database being obtained by adding the augmented signal to the first howling database;
[0019] a howling suppression module configured to perform howling suppression on a target speech signal by using the second neural network model.
[0020] In a fourth aspect, an embodiment of the present application provides a howling suppression apparatus, comprising:
[0021] an obtaining module configured to obtain a target speech signal;
[0022] a processing module configured to process the target speech signal by using a second neural network model to obtain a to-be-played signal, the second neural network model being obtained by fine-tuning a first neural network model by using sample howling signals in a second howling database, the first neural network model being obtained by training an initial model by using sample howling signals in a first howling database, the second howling database being obtained by adding an augmented signal to the first howling database, the augmented signal being a signal obtained by using the first neural network model;
[0023] a playing module configured to play the to-be-played signal.
[0024] In a fifth aspect, an electronic device is provided, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the method of the first aspect or any possible implementation of the first aspect. Alternatively, the electronic device implements the method of the second aspect or any possible implementation of the second aspect.
[0025] In a sixth aspect, a computer readable storage medium is provided, which stores computer instructions. When the computer instructions are executed by a processor, the computer instructions are used to implement the method of the first aspect or any possible implementation of the first aspect. Alternatively, the computer instructions are used to implement the method of the second aspect or any possible implementation of the second aspect.
[0026] In a seventh aspect, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the computer program implements the method of the first aspect or any possible implementation of the first aspect. Alternatively, the computer program implements the method of the second aspect or any possible implementation of the second aspect.
[0027] The echo suppression method, device, and equipment provided in the embodiments of the present application and the readable storage medium. The electronic device trains an initial model by using a sample echo signal in a first echo database to obtain a first neural network model. Then, the electronic device generates an augmented signal by using the first neural network model, adds the augmented signal in the first database to obtain a second echo database, and fine tunes the first neural network model by using a sample echo signal in the second echo database to obtain a second neural network model. Finally, the electronic device performs echo suppression on a target voice signal by using the second neural network model. By using this scheme, the electronic device pre-trains the first neural network model, obtains the augmented signal based on the first neural network model and expands the echo database, fine tunes the first neural network model by using the expanded echo database to obtain the second neural network model, and performs echo suppression by using the second neural network model, which greatly improves the effect of echo suppression. Moreover, the sample echo signal in the model training stage depends on a clean voice signal and a real-time feedback signal, which matches the actual real-time process, so that the first neural network model trained by using the sample echo signal has a good echo suppression effect, thereby achieving the purpose of improving the echo suppression performance of the second neural network model. BRIEF DESCRIPTION OF DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0029] Figure 1 is a schematic diagram of the howling occurrence process;
[0030] Figure 2 is a flowchart of the howling suppression method provided by the embodiments of the present application;
[0031] Figure 3A is a schematic diagram of the first acoustic feedback path in the howling suppression method provided by the embodiments of the present application;
[0032] Figure 3B is a schematic diagram of the second acoustic feedback path in the howling suppression method provided by the embodiments of the present application;
[0033] Figure 3C is a schematic diagram of the third acoustic feedback path in the howling suppression method provided by the embodiments of the present application;
[0034] Figure 4 is a flowchart of the howling suppression method provided by the embodiments of the present application, in which the first neural network model is used to generate an augmented signal;
[0035] Figure 5 is a flowchart of the model training in the howling suppression method provided by the embodiments of the present application;
[0036] Figure 6 is a structural schematic diagram of the second neural network model in the howling suppression method provided by the embodiments of the present application;
[0037] Figure 7A is a structural schematic diagram of the encoding module in Figure 6 ;
[0038] Figure 7B is a structural schematic diagram of the decoding module in Figure 6 ;
[0039] Figure 8 is a flowchart of another howling suppression method provided by the embodiments of the present application;
[0040] Figure 9 is a schematic diagram of a howling suppression device provided by the embodiments of the present application;
[0041] Figure 10 is a schematic diagram of another howling suppression device provided by the embodiments of the present application;
[0042] Figure 11FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0043] Howling is an oscillation caused by positive feedback in a sound reinforcement system. The howling phenomenon refers to: a microphone in a sound reinforcement system collects a sound signal, converts the sound signal into an electric signal, the electric signal is amplified by a power amplifier, and the amplified electric signal is converted into a sound signal and played by a loudspeaker. Part of the sound signal played by the loudspeaker is collected by the microphone through various paths, forming an acoustic loop of “loudspeaker → microphone → amplifier → loudspeaker”, and thus repeatedly circulating to form positive feedback. When the amplitude balance condition and the phase balance condition are met at the same time, oscillation occurs, which is manifested as howling.
[0044] No matter in a local sound reinforcement system or an online conference, when there is positive feedback between a microphone for sound pickup and a loudspeaker for sound reinforcement, that is, the signal recorded by the microphone is played by the loudspeaker placed in the same space, and then picked up again by the same microphone to form a closed acoustic loop, a howling signal is easily generated. For example, in a conference room or a teaching scene, a microphone collects a speech signal of a speaker, the speech signal is subjected to analog-to-digital conversion, amplification, etc., and finally played by a loudspeaker on site. The speech signal played by the loudspeaker is collected again by the microphone, and thus circulates to generate howling.
[0045] For another example, in a conference scene, a device A is located in a remote conference room, and devices B and C are located in a local conference room, and the microphone of the device B and the device C are both turned on. The speech signal collected by the device A in the remote conference room is played by the loudspeaker of the device B, and the sound played by the loudspeaker of the device B is picked up by the microphone of the device C and played, and thus circulates to generate howling.
[0046] Figure 1 FIG. 2 is a schematic diagram of a howling occurrence process. Please refer to Figure 1 The speech signal collected by the microphone reaches the loudspeaker after being subjected to amplitude control and delay processing, the speech signal played by the loudspeaker is picked up again by the microphone, and thus repeatedly circulates to generate howling. Once the howling phenomenon occurs, a harsh and sharp sound is generated, which is easy to damage the electronic device. Moreover, it is easy to harm the auditory system of the user.
[0047] To suppress howling, common howling suppression algorithms include frequency-shifting phase-shifting method, notch method, adaptive feedback control method, etc. Among them, the frequency-shifting phase-shifting method is to destroy the positive feedback path, and the method has limited effect and can easily destroy the sound quality. The notch method detects the howling frequency point after howling occurs, and then performs notch filtering on the howling frequency point to remove the howling signal. The howling suppression effect depends heavily on whether the detection of the howling frequency point is accurate, and it is difficult to completely remove the howling signal. The adaptive feedback control method removes howling by canceling the feedback path. This method needs to converge very deeply to suppress howling, and when the relative position of the loudspeaker and the microphone changes, the filter needs to be re-converged.
[0048] Due to the limitations of the above howling suppression methods, the industry considers combining deep learning to suppress howling. However, the traditional deep learning-based howling suppression scheme uses offline data, which does not match the actual real-time process, resulting in poor howling suppression effect.
[0049] Based on this, the embodiments of the present application provide a howling suppression method, device, equipment and readable storage medium, the training and fine-tuning of the model both use online real-time data, achieving the purpose of improving the howling suppression performance of the second neural network model, and greatly improving the effect of howling suppression.
[0050] During model training, the execution subject of the embodiments of the present application is an electronic device, which is, for example, a server, a desktop computer, a notebook computer, etc. The electronic device pre-trains a first neural network model, obtains an augmented signal based on the first neural network model, and fine-tunes the first neural network model using the augmented signal to obtain a second neural network model. After obtaining the second neural network model, the second neural network model is deployed on the electronic device, such as a conference control device, a server, etc., so as to suppress howling of a target voice signal.
[0051] In the embodiments of the present application, the electronic device for model training and the electronic device for howling suppression of the target voice signal using the trained second neural network model can be the same electronic device or different electronic devices, which is not limited in the embodiments of the present application.
[0052] Figure 2 is a flowchart of the howling suppression method provided by the embodiments of the present application, and the present embodiment is described from the perspective of model training. The present embodiment includes:
[0053] 201, training an initial model using sample howling signals in a first howling database to obtain a first neural network model.
[0054] In the embodiments of the present application, a first howling database for storing sample howling signals is constructed in advance. The electronic device trains an initial model using the sample howling signals in the first howling database to obtain a first neural network model. The sample howling signals are generated according to a clean speech signal and a real-time feedback signal corresponding to the clean speech signal. Since the generation of the sample howling signals depends on the clean speech signal and the real-time feedback signal, which matches the actual real-time process, the first neural network model trained using the sample howling signals has good howling suppression effect.
[0055] 202. Generating an augmented signal using the first neural network model.
[0056] After the first neural network model is trained, the electronic device applies the first neural network model to a howling generation system to obtain an augmented signal. The howling generation system is also referred to as a speech processing system, a sound amplification system, etc.
[0057] 203. Fine-tuning the first neural network model using sample howling signals in a second howling database to obtain a second neural network model, the second howling database being obtained by adding the augmented signal to the first howling database.
[0058] After the electronic device generates the augmented signal using the first neural network model, the augmented signal is added to the first howling database to expand the first howling database. For the sake of clarity, the initial howling database is referred to as the first howling database, and the howling database with the augmented signal added is referred to as the second howling database. Therefore, it can be seen that the sample howling signals in the second howling database include the sample howling signals in the first howling database and the augmented signal.
[0059] 204. Performing howling suppression on a target speech signal using the second neural network model.
[0060] After the electronic device fine-tunes the first neural network model to obtain the second neural network model, the second neural network model is applied to a first acoustic feedback path of the howling generation system to perform howling suppression on a target speech signal. The target speech signal is a speech signal picked up by a microphone and acquired by the electronic device in the process of howling suppression.
[0061] For example, the second neural network model is deployed on the electronic device. In a teaching scenario or a conference scenario, the microphone and the loudspeaker are located in the same space. During the speech of a speaker, the microphone continuously picks up a target speech signal, and the electronic device acquires the target speech signal. Then, the electronic device performs howling suppression on the target speech signal using the second neural network model. The target speech signal after howling suppression is finally played by the loudspeaker.
[0062] For another example, in an online conference scenario, a device B in a local conference room plays a voice signal from a remote device A, a microphone of a device C picks up the sound played by the loudspeaker of the device B and other sounds in the live scene to obtain a target voice signal, and then the electronic device performs howling suppression on the target voice signal by using the second neural network model.
[0063] The howling suppression method provided by the embodiments of the present application includes that the electronic device trains an initial model by using sample howling signals in a first howling database to obtain a first neural network model. Then, the electronic device generates an augmented signal by using the first neural network model, adds the augmented signal in the first database to obtain a second howling database, and fine-tunes the first neural network model by using sample howling signals in the second howling database to obtain a second neural network model. Finally, the electronic device performs howling suppression on a target voice signal by using the second neural network model. By using this scheme, the electronic device pre-trains the first neural network model, expands the howling database based on the first neural network model to obtain the augmented signal, fine-tunes the first neural network model by using the expanded howling database to obtain the second neural network model, and performs howling suppression by using the second neural network model, which greatly improves the howling suppression effect. Moreover, the sample howling signals in the model training stage depend on clean voice signals and real-time feedback signals, which are matched with the actual real-time process, so that the first neural network model trained by using the sample howling signals has a good howling suppression effect, thereby achieving the purpose of improving the howling suppression performance of the second neural network model.
[0064] The howling suppression method provided by the embodiments of the present application is applied to an electronic device in a howling generation system, which is also called a voice processing system, a sound amplification system, etc., and includes a microphone, an electronic device, a loudspeaker, etc. which are independently arranged or integrally arranged. When howling does not occur, the voice signal picked up by the microphone is processed by the electronic device and played by the loudspeaker. If howling occurs, the sound played by the loudspeaker is picked up by the microphone again, and the cycle is repeated to cause howling. The signal transmission path between the microphone and the loudspeaker is called an acoustic feedback path. According to whether the first neural network model, the second neural network model, etc. are deployed in the acoustic feedback path, the acoustic feedback path can be divided into a first acoustic feedback path, a second acoustic feedback path and a third acoustic feedback path. The first acoustic feedback path is mainly used to generate sample howling signals in the first howling database, the second acoustic feedback path is mainly used to generate augmented signals in the second howling database, and the third acoustic feedback path is mainly used to perform howling suppression on a target voice signal. The acoustic feedback paths are described in detail as follows.
[0065] The first acoustic feedback path is: microphone→electronic device (without deploying the first neural network model)→loudspeaker, and the first acoustic feedback path is mainly used to generate sample howling signals in the first howling database. Figure 3Ais a schematic diagram of a first acoustic feedback path in a howling suppression method provided by an embodiment of the present application. Please refer to Figure 3A , the microphone, the electronic device, and the loudspeaker are all independent devices, the first neural network model or the second neural network model is not deployed on the electronic device, the electronic device has functions such as electrical signal conversion, amplitude control, delay processing, and loudspeaker nonlinear distortion processing, and the electronic device is, for example, a conference control or the like; or the electronic device is a remote server or the like.
[0066] The second acoustic feedback path is: microphone → electronic device (deploying the first neural network model) → loudspeaker. Figure 3B is a schematic diagram of a second acoustic feedback path in a howling suppression method provided by an embodiment of the present application. Please refer to Figure 3B , the microphone, the electronic device, and the loudspeaker are all independent devices, the first neural network model is deployed on the electronic device, and the electronic device also has functions such as electrical signal conversion, amplitude control, delay processing, and loudspeaker nonlinear distortion processing, and the electronic device is, for example, a conference control or the like; or the electronic device is a remote server or the like. After the first neural network model is applied to the first acoustic feedback path to obtain the second acoustic feedback path, the electronic device simulates a howling occurrence process based on the second acoustic feedback path, and takes the output signal of the loudspeaker in the second acoustic feedback path as an augmented signal.
[0067] The first acoustic feedback path and the second acoustic feedback path described above can be obtained by simulation, modeling, or the like, or according to a physical object, and embodiments of the present application are not limited.
[0068] The third acoustic feedback path is: microphone → electronic device (deploying the second neural network model) → loudspeaker. Figure 3C is a schematic diagram of a third acoustic feedback path in a howling suppression method provided by an embodiment of the present application. Please refer to Figure 3C , the microphone, the electronic device, and the loudspeaker are all independent devices, the second neural network model is deployed on the electronic device, and the electronic device also has functions such as electrical signal conversion, amplitude control, delay processing, and loudspeaker nonlinear distortion processing, and the electronic device is, for example, a conference control or the like; or the electronic device is a remote server or the like. After the second neural network model is applied to the first acoustic feedback path to obtain the third acoustic feedback path, the electronic device performs howling suppression on the target speech signal based on the third acoustic feedback path, so that the loudspeaker plays the speech signal after howling suppression.
[0069] It should be noted that although in the above various acoustic feedback paths, the microphone, the loudspeaker, and the electronic device are independently deployed. However, embodiments of the present application are not limited thereto. For example, in other possible implementation manners, the loudspeaker and the microphone are integrated on the electronic device, or any one of the loudspeaker and the microphone is integrated on the electronic device.
[0070] Optionally, in the above embodiment, in the process of generating the augmented signal by the electronic device using the first neural network model, first, the electronic device applies the first neural network model to a first acoustic feedback path to obtain a second acoustic feedback path, the first acoustic feedback path being a signal transmission path between a microphone and a loudspeaker. Then, a howling occurrence process is simulated based on the second acoustic feedback path to generate at least one augmented signal.
[0071] Please refer to Figure 3B After the first neural network model is applied to the first acoustic feedback path to obtain the second acoustic feedback path, the electronic device simulates a howling occurrence process based on the second acoustic feedback path, and takes the output signal of the loudspeaker in the second acoustic feedback path as the augmented signal.
[0072] With this scheme, the electronic device simulates a howling occurrence process based on the second acoustic feedback path to which the first neural network model is applied to obtain the augmented signal, achieving the purpose of quickly and accurately obtaining the augmented signal.
[0073] Figure 4 is a flowchart of generating an augmented signal by a first neural network model in the howling suppression method provided by the embodiments of the present application. The present embodiment includes:
[0074] 401. Obtain at least one signal pair, each signal pair including a clean speech signal and a feedback signal, the feedback signal being a simulated signal that the clean speech signal collected by a microphone is played by a loudspeaker and then collected by the microphone again.
[0075] In this step, the electronic device simulates or emulates a real-time feedback signal corresponding to the clean speech signal, and takes the clean speech signal and the real-time feedback signal as a signal pair. Wherein, the clean speech signal is x clean , and the simulated feedback signal x feedback is:
[0076] 402. For each signal pair, input the clean speech signal and the feedback signal in the signal pair into the microphone in the second acoustic feedback path, and take the output of the loudspeaker in the second acoustic feedback path as the augmented signal to obtain at least one augmented signal.
[0077] The electronic device simulates a howling occurrence process, and takes the sum of the clean speech signal x clean and the real-time feedback signal xfeedback as the input of the microphone. The sum signal picked up by the microphone is processed by gain processing, delay processing and the first neural network model, and then played by the loudspeaker. The electronic device takes the signal played by the loudspeaker as the augmented signal x′ howling . The augmented signal is the howling signal that the first neural network model fails to eliminate.
[0078] With this scheme, the first neural network model trained offline is applied to the acoustic feedback path to obtain the howling signal that the first neural network fails to eliminate, i.e., to obtain the augmented signal, the augmented signal is used to expand the howling database so as to fine-tune the model, and the purpose of improving the howling suppression effect is achieved.
[0079] Optionally, in the above embodiment, the electronic device further inputs the sum signal into the microphone in the first acoustic feedback path, and takes the output of the loudspeaker of the first acoustic feedback path as the sample howling signal to obtain at least one sample howling signal. Then, the electronic device constructs the first howling database according to the at least one sample howling signal.
[0080] For example, referring to Figure 3A , the electronic device simulates the howling generation process, and takes the sum of the clean speech signal x clean and the real-time feedback signal xfeedback as the input of the microphone. The sum signal picked up by the microphone is processed through non-linear distortion and the like, and then played by the loudspeaker. The electronic device takes the signal played by the loudspeaker as the sample howling signal x howling . The generation of the sample howling signal x howling is as shown in formula (1):
[0081]
[0082] wherein, represents the convolution operation, h(t) represents the room impulse response, G represents the gain, t represents the time, and Δt represents the time delay that changes. NL represents the non-linear processing. The non-linear processing NL includes a hard clipping function and a memoryless S-shaped function. The non-linear processing NL is as shown in formula (2):
[0083]
[0084] wherein, is a memoryless S-shaped function, and the factor b(t) in the memoryless S-shaped function is as shown in formula (3):
[0085] b(t) = 1.5 x x hard (t) - 0.3 x hard 2 (t) formula (3)
[0086] wherein, x hard (t) is a hard clipping function, and is as shown in formula (4):
[0087]
[0088] x(t) represents the signal value at time t, x max represents a threshold value.
[0089] In formula (2) to formula (4), the gain γ is set to 4 for example, x max is set to 1 for example, and a is an emulation coefficient, and a is randomly selected in the range of [0.2, 4] for example to emulate different degrees of nonlinear distortion.
[0090] According to formula (1) to formula (4), the electronic device simulates the occurrence process of each signal pair to obtain at least one sample howling signal. Then, the electronic device uses the sample howling signals x howling to construct a first howling database. Subsequently, the sample howling signals x howling in the first howling database are used to train a first neural network model.
[0091] With this scheme, the electronic device simulates the occurrence process of howling to construct a first howling database, and the sample howling signals are generated according to clean speech signals and real-time feedback signals, so that the training of the first neural network model uses online real-time data, which achieves the purposes of improving the howling suppression performance of the first neural network model and quickly constructing a high-quality first howling database. Moreover, the nonlinear distortion of the loudspeaker is considered, so that the sample howling signals are more realistic.
[0092] Next, the training process of the first neural network model, the process of fine-tuning the first neural network model to obtain a second neural network model, and the process of using the second neural network model for howling suppression are described in detail.
[0093] First, the training process of the first neural network model.
[0094] After the first howling database is constructed, the electronic device trains an initial model using the sample howling signals in the first howling database to obtain the first neural network model. In the training process, the sample howling signals are used as input, the clean speech signals corresponding to the sample howling signals are used as targets, and the model training is performed without considering acoustic feedback.
[0095] Figure 5 is a flowchart of model training in the howling suppression method provided by the embodiments of the present application. The present embodiment includes:
[0096] 501, determine the first complex frequency spectrum of the sample howling signal;
[0097] For each sample howling signal in the first howling database, the electronic device determines a time domain howling signal x(t) of the sample howling signal, performs a short-time Fourier transform analysis on the time domain howling signal x(t) to obtain a first complex frequency spectrum X1(k, f), which is an input feature of the initial model. Wherein, k represents a time frame, and f represents a frequency. The initial model is, for example, a convolutional recurrent neural (CRN) model.
[0098] 502、inputting the first complex frequency spectrum into the initial model, so that the initial model outputs a mask matrix.
[0099] The electronic device inputs the real part and the imaginary part of the first complex frequency spectrum X1(k, f) into the initial model as an input of the initial model, so that the initial model outputs a mask matrix M(k, f).
[0100] 503、determining a second complex frequency spectrum according to the mask matrix and the first complex frequency spectrum.
[0101] The electronic device performs a point multiplication processing on the mask matrix and the aforementioned first complex frequency spectrum X1(k, f) to obtain a second complex frequency spectrum X2(k, f).
[0102] 504、determining a loss function according to the second complex frequency spectrum and a third complex frequency spectrum of a reference signal, the reference signal being a clean speech signal corresponding to the sample howling signal.
[0103] The electronic device takes the clean speech signal corresponding to the sample howling signal as the reference signal, performs a short-time Fourier transform analysis on a time domain signal of the reference signal to obtain a third complex frequency spectrum, and takes a mean square error of the third complex frequency spectrum and the second complex frequency spectrum as the loss function.
[0104] 505、determining whether the initial model converges according to the loss function and the reference signal, and if the initial model converges, performing step 506; if the initial model does not converge, returning to step 501.
[0105] 506、when the initial model converges, determining the converged initial model as the first neural network model.
[0106] The electronic device continuously adjusts parameters of the initial model, and determines whether the initial model converges according to the loss function and the reference signal. If the initial model does not converge, the electronic device returns to step 501 to continue taking the sample howling signal from the first howling database and determining the first complex frequency spectrum. If the initial model converges, the electronic device takes the converged initial model as the first neural network model.
[0107] With this scheme, the electronic device constructs a loss function according to the mean square error of the third complex frequency spectrum and the second complex frequency spectrum of the reference signal, and determines whether the initial model converges by using the loss function and the reference signal, so as to quickly and accurately train the first neural network model. The third complex frequency spectrum is also called a target complex frequency spectrum.
[0108] Secondly, the process of fine-tuning the first neural network model to obtain the second neural network model.
[0109] In this process, the electronic device inputs the complex frequency spectrum of the sample howling signal in the second howling database into the first neural network model. Then, the electronic device fine-tunes the first neural network model according to a reference signal and a loss function, the reference signal is a clean speech signal corresponding to the sample howling signal in the second howling database, and the loss function is the mean square error of the target complex frequency spectrum of the reference signal and the output complex frequency spectrum of the first neural network model.
[0110] In the fine-tuning process, the sample howling signal in the second howling database is taken as input, the clean speech signal corresponding to the sample howling signal is taken as target, and the model is fine-tuned without considering acoustic feedback.
[0111] For example, after the second howling database is constructed, the electronic device randomly selects x howling or x′ howling from the second howling database, determines the time domain howling signal of the randomly selected sample howling signal, and then determines the complex frequency spectrum. Then, the electronic device inputs the complex frequency spectrum of the sample howling signal in the second howling database into the first neural network model to output a mask matrix, and obtains the output complex frequency spectrum of the first neural network model according to the mask matrix and the complex frequency spectrum determined based on the time domain howling signal. The loss function is determined according to the output complex frequency spectrum of the first neural network model and the target complex frequency spectrum of the reference signal.
[0112] In the process of continuously adjusting the parameters of the first neural network model by the electronic device, whether the first neural network model converges is determined according to the loss function and the reference signal. If the first neural network model does not converge, the sample howling signal is continuously taken from the second howling database, and the complex frequency spectrum is determined. If the first neural network model converges, the converged first neural network model is taken as the second neural network model.
[0113] With this scheme, the electronic device fine-tunes the first neural network model to obtain the second neural network model, so as to achieve the purpose of improving the howling suppression effect.
[0114] Finally, the process of using the second neural network model for howling suppression.
[0115] In the process, the electronic device inputs the fourth complex spectrum of the target voice signal into the second neural network model, so that the second neural network model outputs a target mask matrix. Then, the electronic device determines a fifth complex spectrum according to the target mask matrix and the fourth complex spectrum, generates a time domain signal according to the fifth complex spectrum, and generates a to-be-played signal according to the time domain signal, the to-be-played signal being a voice signal after howling suppression.
[0116] Figure 6 is a structural schematic diagram of the second neural network model in the howling suppression method provided by the embodiment of the present application. Please refer to Figure 6 , the second neural network model includes an encoding module (Encoder Block), a decoding module (Decoder Block), and an intermediate layer that plays a connecting role. The input of the second neural network model is the fourth complex spectrum of the target voice signal, and the output is a target mask matrix. Then, the electronic device performs point multiplication processing on the mask matrix and the aforementioned fourth complex spectrum to obtain a fifth complex spectrum, and performs inverse Fourier transform on the fifth complex spectrum to obtain an output time domain signal. According to the output time domain signal, a to-be-played signal is generated, and the to-be-played signal is a voice signal after howling suppression.
[0117] Please refer to Figure 6 , the encoding module and the decoding module are each, for example, 4, such as Figure 6 , the encoding module 2→16, the encoding module 16→32, the encoding module 32→64, and the encoding module 64→128; the decoding module 128→64, the decoding module 64→32, the decoding module 32→16, and the decoding module 16→2. The intermediate layer is, for example, a long short-term memory (LSTM) and the like. The decoding side also includes a convolution layer Conv2d8→8 and the like. The number of paths of each layer of the encoding module and the decoding module is as shown in Figure 6 .
[0118] It should be noted that, although Figure 6 is described by taking an example in which the second neural network model includes four encoding modules and four decoding modules. However, the embodiment of the present application is not limited, and in other feasible implementation manners, the number of encoding modules and decoding modules can also be set according to requirements.
[0119] Figure 7A is a structural schematic diagram of the encoding module in Figure 6 . Please refer to Figure 7A, one encoding module includes a convolution layer (Conv2d), a normalization layer (BatchNorm) and an activation layer (Leaky_relu), the convolution kernel size of the encoding module is (3x3), and the step length of the convolution kernel is, for example, (2, 1), that is, the frequency domain step length is 2, and the time domain step length is 1. The convolution is, for example, a causal convolution to meet the real-time processing requirement.
[0120] Figure 7B is Figure 6 A structural diagram of a decoding module in the method is shown in FIG. 8. Figure 7B , one decoding module includes a transposed convolution (ConvTranspose2d), a normalization layer (BatchNorm) and an activation layer (Leaky_relu), the convolution kernel size of the transposed layer of the decoding module is (3x3), and the step length of the convolution kernel is, for example, (2, 1), that is, the frequency domain step length is 2, and the time domain step length is 1. The convolution is, for example, a causal convolution to meet the real-time processing requirement.
[0121] With the scheme, the electronic device generates a first neural network model, fine-tunes the first neural network model to obtain a second neural network model, and uses the second neural network model for howling suppression, thereby achieving the purpose of improving the howling suppression effect.
[0122] Figure 8 is a flowchart of another howling suppression method provided by an embodiment of the present application. The embodiment includes the following steps.
[0123] 801, obtaining a target speech signal.
[0124] 802, processing the target speech signal by using a second neural network model to obtain a to-be-played signal, the second neural network model being obtained by fine-tuning a first neural network model by using sample howling signals in a second howling database, the first neural network model being obtained by training an initial model by using sample howling signals in a first howling database, the second howling database being obtained by adding an augmented signal to the first howling database, and the augmented signal being a signal obtained by using the first neural network model.
[0125] 803, playing the to-be-played signal.
[0126] The embodiment is explained from the perspective of howling suppression on a target speech signal by using a second neural network model, and specific details can be referred to the above-described embodiments of model training and fine-tuning, which will not be described herein again.
[0127] The following is an apparatus embodiment of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the apparatus embodiments of the present application, please refer to the method embodiments of the present application.
[0128] Figure 9A schematic diagram of a howling suppression device provided for an embodiment of the present application. The howling suppression device 900 comprises a training module 91, a processing module 92, a fine-tuning module 93 and a howling suppression module 94.
[0129] The training module 91 is configured to train an initial model using sample howling signals in a first howling database to obtain a first neural network model.
[0130] The processing module 92 is configured to generate augmented signals using the first neural network model.
[0131] The fine-tuning module 93 is configured to fine-tune the first neural network model using sample howling signals in a second howling database to obtain a second neural network model, the second howling database being obtained by adding the augmented signals to the first howling database.
[0132] The howling suppression module 94 is configured to suppress howling in a target speech signal using the second neural network model.
[0133] In an implementation, the processing module 92 is configured to apply the first neural network model to a first acoustic feedback path to obtain a second acoustic feedback path, the first acoustic feedback path being a signal transmission path between a microphone and a loudspeaker; and simulate a howling generation process based on the second acoustic feedback path to generate at least one augmented signal.
[0134] In an implementation, when the processing module 92 simulates the howling generation process based on the second acoustic feedback path to generate at least one augmented signal, the processing module 92 is configured to obtain at least one signal pair, each signal pair comprising a clean speech signal and a feedback signal, the feedback signal being a signal simulated by the microphone after the clean speech signal is played by the loudspeaker and then collected by the microphone again; and for each signal pair, input the clean speech signal and the feedback signal in the signal pair into the microphone in the second acoustic feedback path, and take the output of the loudspeaker in the second acoustic feedback path as an augmented signal to obtain at least one augmented signal.
[0135] In an implementation, the processing module 92 is further configured to input the clean speech signal and the feedback signal in the signal pair into the microphone in the first acoustic feedback path, take the output of the loudspeaker in the first acoustic feedback path as a sample howling signal to obtain at least one sample howling signal; and construct the first howling database according to the at least one sample howling signal.
[0136] In a possible implementation, the fine-tuning module 93 is configured to input a complex spectrum of a sample howling signal in the second howling database into the first neural network model, fine-tune the first neural network model according to a reference signal and a loss function, the reference signal being a clean speech signal corresponding to the sample howling signal in the second howling database, and the loss function being a mean square error of a target complex spectrum of the reference signal and an output complex spectrum of the first neural network model.
[0137] In a possible implementation, the training module 91 is configured to determine a first complex spectrum of the sample howling signal, input the first complex spectrum into the initial model to make the initial model output a mask matrix, determine a second complex spectrum according to the mask matrix and the first complex spectrum, determine a loss function according to a third complex spectrum of a reference signal and the second complex spectrum, the reference signal being a clean speech signal corresponding to the sample howling signal, determine whether the initial model converges according to the loss function and the reference signal, and when the initial model converges, determine the converged initial model as the first neural network model.
[0138] In a possible implementation, the howling suppression module 94 is configured to input a fourth complex spectrum of the target speech signal into the second neural network model to make the second neural network model output a target mask matrix, determine a fifth complex spectrum according to the target mask matrix and the fourth complex spectrum, generate a time domain signal according to the fifth complex spectrum, and generate a to-be-played signal according to the time domain signal, the to-be-played signal being a howling-suppressed speech signal.
[0139] In a possible implementation, the first neural network model includes at least one encoding module, at least one decoding module, and an intermediate layer, the intermediate layer being configured to connect the encoding module and the decoding module, the encoding module including a convolution layer, a normalization layer, and an activation layer, and the decoding module including a transposed convolution, a normalization layer, and an activation layer.
[0140] The howling suppression apparatus provided by the embodiments of the present application can perform the actions of the electronic device for training and fine-tuning the neural network model in the above embodiments, and has similar implementation principles and technical effects, which will not be described here again.
[0141] Figure 10 FIG. 1 shows a schematic diagram of another howling suppression apparatus provided by the embodiments of the present application. The howling suppression apparatus 1000 includes an acquisition module 101, a processing module 102, and a playing module 103.
[0142] The acquisition module 101 is configured to acquire a target speech signal.
[0143] The processing module 102 is configured to process the target voice signal by using a second neural network model to obtain a to-be-played signal, the second neural network model is obtained by fine-tuning the first neural network model by using sample howling signals in a second howling database, the first neural network model is obtained by training an initial model by using sample howling signals in a first howling database, and the second howling database is obtained by adding augmented signals to the first howling database, the augmented signals are signals obtained by using the first neural network model.
[0144] The playing module 103 is configured to play the to-be-played signal.
[0145] The howling suppression device provided in the embodiments of the present application can perform the action of the electronic device for suppressing howling by using the second neural network model in the above embodiments, and the implementation principle and technical effects are similar, which will not be described here.
[0146] Figure 11 FIG. 1 is a structural schematic diagram of an electronic device provided in the embodiments of the present application. Please refer to Figure 11 The electronic device 1100 provided in the embodiments of the present application includes at least one processor 111, at least one communication bus 112, a user interface 113, at least one network interface 114, and a memory 115.
[0147] The communication bus 112 is configured to realize the connection and communication between the components.
[0148] The user interface 113 can include a display screen (Display) and a camera (Camera), and the optional user interface 113 can further include a standard wired interface and a wireless interface. The display screen is configured to display an editing interface, a roaming interface, and the like.
[0149] The network interface 114 can optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0150] The processor 111 can include one or more processing cores. The processor 111 connects various parts within the entire electronic device 1100 by various interfaces and lines, and performs various functions of the electronic device 1100 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 115, and calling data stored in the memory 115. Alternatively, the processor 111 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA). The processor 111 can integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes an operating system, a user interface, and an application program; the GPU is responsible for rendering and drawing a panoramic sphere required to be displayed on a display screen; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 111, but can be implemented by a separate chip.
[0151] The memory 115 can include a random access memory (RAM) and can also include a read-only memory (ROM). Alternatively, the memory 115 includes a non-transitory computer-readable storage medium. The memory 115 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 115 can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store data involved in the above-mentioned various method embodiments, etc. The memory 115 can alternatively be at least one storage device located away from the aforementioned processor 111. As shown, the memory 115 as a computer storage medium can include an operating system, a network communication module, a user interface module, and an operating application program of the electronic device. Figure 11 As shown, the memory 115 as a computer storage medium can include an operating system, a network communication module, a user interface module, and an operating application program of the electronic device.
[0152] The embodiment of the present application further provides a computer readable storage medium, wherein the computer readable storage medium stores computer instructions, and the computer instructions are executed by a processor to implement the howling suppression method implemented by the electronic device.
[0153] The embodiment of the present application further provides a computer program product, which contains a computer program, and the computer program is executed by a processor to implement the howling suppression method implemented by the electronic device.
[0154] Those skilled in the art should understand that the embodiment of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can adopt a computer program product implemented on one or more computer readable storage media containing computer usable program code (including but not limited to disk memory, CD-ROM, optical memory, etc.).
[0155] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiment of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams and the combination of the flows and / or blocks can be implemented by computer program instructions. These computer program instructions can be provided to a general purpose computer, a special purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the computer or other programmable data processing device produce a device implemented in the flowcharts and / or block diagrams. Figure 1 The function specified in one flow or multiple flows and / or blocks. Figure 1 The function specified in one block or multiple blocks.
[0156] These computer program instructions can also be stored in a computer readable memory capable of guiding the computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer readable memory produce a product including instruction devices, which implement the flowcharts and / or block diagrams. Figure 1 The function specified in one flow or multiple flows and / or blocks. Figure 1 The function specified in one block or multiple blocks.
[0157] These computer program instructions can also be loaded into the computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to produce a computer implemented process, so that the instructions executed on the computer or other programmable device provide a process for implementing the flowcharts and / or block diagrams. Figure 1 The function specified in one flow or multiple flows and / or blocks. Figure 1 The function specified in one block or multiple blocks.
[0158] In one typical arrangement, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0159] Memory can include non-persistent memory and / or volatile memory, such as a random access memory (RAM) including a cache area for the temporary storage of data. A memory can also include non-volatile memory, such as a read only memory (ROM), EPROM, EEPROM, or flash memory. Memory can further include a data storage 112, which can include a disk drive, an optical memory, a solid-state memory, or other storage media. Memory can store computer readable instructions that, when processed by a processor, cause a computing device to perform operations.
[0160] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for the storage of information. Information can be computer readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile discs (DVDs) or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer readable media does not include transitory media, such as modulated data signals and carrier waves.
[0161] It should also be noted that the terms "comprising," "including," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements in the list, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without limitation, an element preceded by "comprises a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0162] The above merely provides an embodiment of the present application and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of claims of the present application.
Claims
1. A howling suppression method characterized by, The method comprises the following steps: training an initial model by using sample howling signals in a first howling database to obtain a first neural network model; generating augmented signals by using the first neural network model; fine-tuning the first neural network model by using sample howling signals in a second howling database to obtain a second neural network model, wherein the second howling database is obtained by adding the augmented signals to the first howling database; performing howling suppression on a target voice signal by using the second neural network model.
2. The method of claim 1, wherein, The method of generating augmented signals by using the first neural network model comprises the following steps: applying the first neural network model to a first acoustic feedback path to obtain a second acoustic feedback path, wherein the first acoustic feedback path is a signal transmission path between a microphone and a loudspeaker; simulating a howling generation process based on the second acoustic feedback path to generate at least one augmented signal.
3. The method of claim 2, wherein, The method of simulating a howling generation process based on the second acoustic feedback path to generate at least one augmented signal comprises the following steps: obtaining at least one signal pair, wherein each signal pair comprises a clean voice signal and a feedback signal, and the feedback signal is a signal simulated by a microphone after the clean voice signal is played by a loudspeaker and then collected by the microphone again; for each signal pair, inputting the clean voice signal and the feedback signal in the signal pair into a microphone in the second acoustic feedback path, and taking the output of a loudspeaker in the second acoustic feedback path as an augmented signal to obtain at least one augmented signal.
4. The method of claim 3, wherein, The method further comprises the following steps: inputting the clean speech signal x clean and the feedback signal xfeedback into a microphone in the first acoustic feedback path, outputting the loudspeaker of the first acoustic feedback path as a sample howling signal x howling to obtain at least one sample howling signal; constructing the first howling database according to the at least one sample howling signal; wherein: the feedback signal denotes a convolution operation, h(t) denotes a room impulse response, G denotes a gain, t denotes time, At denotes a time-varying delay, and NL denotes a non-linear processing.
5. The method according to any one of claims 1 to 4, characterized in that, the method of fine-tuning the first neural network model by using sample howling signals in the second howling database comprises the following steps: inputting the complex spectrum of the sample howling signal in the second howling database into the first neural network model; fine-tuning the first neural network model according to a reference signal and a loss function, wherein the reference signal is a clean voice signal corresponding to the sample howling signal in the second howling database, and the loss function is a mean square error between a target complex spectrum of the reference signal and an output complex spectrum of the first neural network model.
6. The method according to any one of claims 1 to 4, characterized in that, The method of training an initial model by using sample howling signals in a first howling database to obtain a first neural network model comprises the following steps: determining a first complex spectrum of the sample howling signal; inputting the first complex spectrum into the initial model to make the initial model output a mask matrix; determining a second complex spectrum according to the mask matrix and the first complex spectrum; determining a loss function according to the second complex spectrum and a third complex spectrum of a reference signal, wherein the reference signal is a clean voice signal corresponding to the sample howling signal; determining whether the initial model converges according to the loss function and the reference signal; when the initial model converges, determining that the converged initial model is the first neural network model.
7. The method according to any one of claims 1 to 3, characterized in that, The method of performing howling suppression on a target voice signal by using the second neural network model comprises the following steps: inputting a fourth complex spectrum of the target voice signal into the second neural network model to make the second neural network model output a target mask matrix; determining a fifth complex frequency spectrum according to the target mask matrix and the fourth complex frequency spectrum; generating a time domain signal according to the fifth complex frequency spectrum; generating a to-be-played signal according to the time domain signal, the to-be-played signal being a speech signal subjected to howling suppression.
8. The method of any one of claims 1-3, wherein the first neural network model comprises at least one encoding module, at least one decoding module, and an intermediate layer for connecting the encoding module and the decoding module, the encoding module comprising a convolution layer, a normalization layer, and an activation layer, and the decoding module comprising a transposed convolution, a normalization layer, and an activation layer.
9. A howling suppression method characterized by, comprising: obtaining a target speech signal; processing the target speech signal by using a second neural network model to obtain a to-be-played signal, the second neural network model being obtained by fine-tuning the first neural network model by using sample howling signals in a second howling database, the second howling database being obtained by adding an augmented signal to the first howling database, the augmented signal being a signal obtained by using the first neural network model; playing the to-be-played signal.
10. A howling suppressing device, characterized by comprising: comprising: a training module configured to train an initial model by using sample howling signals in a first howling database to obtain a first neural network model; a processing module configured to generate an augmented signal by using the first neural network model; a fine-tuning module configured to fine-tune the first neural network model by using sample howling signals in a second howling database to obtain a second neural network model, the second howling database being obtained by adding the augmented signal to the first howling database; a howling suppression module configured to suppress howling of a target speech signal by using the second neural network model.
11. A howling suppression device, characterized by comprising: an obtaining module configured to obtain a target speech signal; a processing module configured to process the target speech signal by using a second neural network model to obtain a to-be-played signal, the second neural network model being obtained by fine-tuning the first neural network model by using sample howling signals in a second howling database, the second howling database being obtained by adding an augmented signal to the first howling database, the augmented signal being a signal obtained by using the first neural network model; a playing module configured to play the to-be-played signal.
12. An electronic device comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to enable the electronic device to implement the method of any one of claims 1-9.
13. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1-9.