Acoustic feedback elimination method and device, equipment and storage medium

By adjusting filter parameters in real time using acoustic feedback suppression and register separation models, the problem of acoustic feedback howling in the public address system was solved, achieving real-time and accurate acoustic feedback elimination and improving sound quality.

CN121397401APending Publication Date: 2026-01-23IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511239047.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

In scenarios such as medium to large-scale conferences, classrooms, hearing aids, and in-vehicle karaoke without microphones, existing technologies cause feedback from microphones and speakers in the sound amplification system, resulting in howling that is difficult to eliminate accurately.

Method used

An acoustic feedback suppression model is used to predict the intensity distribution of the target feedback signal of the loudspeaker in the microphone signal, and beam processing is used to eliminate the feedback signal. The acoustic feedback suppression model and the sound zone separation model are used to adjust the filter parameters in real time to avoid the convergence delay of the adaptive filter when the path changes abruptly.

Benefits of technology

It achieves real-time and accurate acoustic feedback cancellation, improves sound quality, and adapts to signal changes in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397401A_ABST
    Figure CN121397401A_ABST
Patent Text Reader

Abstract

The invention discloses a sound feedback elimination method and device, equipment and a storage medium. The method comprises the following steps: acquiring at least one path of target microphone signal acquired by a microphone array; the target microphone signal comprises an original sound source signal and a target feedback signal of the loudspeaker; for each target microphone signal, predicting the target microphone signal by using an acoustic feedback suppression model to obtain first feedback intensity information of the target microphone signal, the first feedback intensity information being used for representing intensity distribution of the target feedback signal in the target microphone signal; and performing beam processing on the target microphone signal at least based on the first feedback intensity information to obtain a clean sound source signal corresponding to the target microphone signal. In this way, the accuracy of sound feedback elimination can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of sound processing technology, and in particular to a method, apparatus, device and storage medium for eliminating acoustic feedback. Background Technology

[0002] In applications such as medium to large-scale conferences, classrooms, hearing aids, and in-car karaoke systems, amplification systems are often used to improve the volume and sound quality of the audio. In these systems, microphones and speakers operate in the same space. The microphones pick up the signals played by the speakers, forming an acoustic closed loop, which can lead to acoustic feedback and consequently, howling.

[0003] Accurately eliminating acoustic feedback is crucial for improving sound quality. Summary of the Invention

[0004] The main technical problem addressed by this application is to provide a method, apparatus, device, and storage medium for acoustic feedback cancellation, which can improve the accuracy of acoustic feedback cancellation.

[0005] To address the aforementioned technical problems, this application provides a method for acoustic feedback cancellation, comprising: acquiring at least one target microphone signal collected by a microphone array; the target microphone signal includes an original sound source signal and a target feedback signal from a loudspeaker; for each target microphone signal, predicting the target microphone signal using an acoustic feedback suppression model to obtain first feedback intensity information of the target microphone signal, the first feedback intensity information being used to characterize the intensity distribution of the target feedback signal in the target microphone signal; and performing beamforming on the target microphone signal based at least on the first feedback intensity information to obtain a clean sound source signal corresponding to the target microphone signal.

[0006] To address the aforementioned technical problems, another technical solution adopted in this application is to provide an acoustic feedback cancellation device, comprising: an acquisition module, a prediction module, and a beam processing module. The acquisition module is used to acquire at least one target microphone signal collected by a microphone array; the target microphone signal includes an original sound source signal and a target feedback signal from a loudspeaker; the prediction module is used to predict each target microphone signal using an acoustic feedback suppression model to obtain first feedback intensity information of the target microphone signal, the first feedback intensity information being used to characterize the intensity distribution of the target feedback signal in the target microphone signal; the beam processing module is used to perform beam processing on the target microphone signal based at least on the first feedback intensity information to obtain a clean sound source signal corresponding to the target microphone signal.

[0007] To solve the above technical problems, the application adopts still another technical solution: providing an electronic device, comprising a memory and a processor coupled with each other, the memory storing program instructions; the processor is used to execute the program instructions stored in the memory to realize the above method.

[0008] To solve the above technical problems, the application adopts still another technical solution: providing a computer readable storage medium for storing program instructions, the program instructions can be executed to realize the above method.

[0009] The above scheme, for each target microphone signal, first uses the sound feedback suppression model to predict the intensity distribution of the target feedback signal of the loudspeaker in the target microphone signal, and then performs beam processing on the target microphone signal based on the intensity distribution of the target feedback signal to obtain the corresponding clean sound source signal. Among them, the existing filter method using adaptive filter for sound feedback elimination depends on the current signal transmission path. Once the signal transmission path changes, the adaptive filter must adjust the parameters to adapt to the new path. This process needs a certain time to complete, so in the filter convergence stage, the feedback elimination effect will be significantly worse. And the application predicts the intensity distribution of the target feedback signal in real time through the model, without waiting for the filter to converge. As long as the model outputs the intensity distribution in real time, beam processing can be done immediately, so it can not only eliminate sound feedback in real time, but also ensure the accuracy of the elimination effect. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1 is a flowchart of an embodiment of the sound feedback elimination method provided by the application;

[0011] Figure 2 is a layout diagram of an embodiment of the vehicle-mounted four sound area K song scene without microphone provided by the application;

[0012] Figure 3 is Figure 1 is a flowchart of an embodiment of the step S13 shown in the figure;

[0013] Figure 4 is a flowchart of an embodiment of the sound feedback suppression and interference provided by the application;

[0014] Figure 5 is a framework diagram of an embodiment of the sound feedback elimination device provided by the application;

[0015] Figure 6 is a framework diagram of an embodiment of the electronic device provided by the application;

[0016] Figure 7 is a framework diagram of the computer readable storage medium provided by the application. DETAILED DESCRIPTION

[0017] For the purposes of the present application, the technical solutions and effects are more clear and explicit, the following embodiments are further described with reference to the drawings.

[0018] In addition, if the description of "first", "second" and the like is involved in the embodiments of the present application, the description of "first", "second" and the like is only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second" can be explicitly or implicitly included at least one of the features. In addition, the technical solutions of each embodiment can be combined with each other, but it must be based on the realization of the ordinary skilled in the art, when the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, nor in the protection scope claimed by the present application.

[0019] It should be noted that the acoustic feedback elimination method provided by the present application is applicable to any scene that needs to eliminate the feedback signal of the loudspeaker, such as medium and large conference, classroom, hearing aid, vehicle-mounted K song and the like. Among them, the feedback signal of the loudspeaker is the loudspeaker playing signal collected by the microphone.

[0020] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the acoustic feedback elimination method provided by the present application. It should be noted that if there is substantially the same result, the embodiment is not limited to the flow order shown in Figure 1 . As shown in Figure 1 , the embodiment includes:

[0021] S11: obtaining at least one target microphone signal collected by a microphone array; the target microphone signal includes an original sound source signal and a target feedback signal of the loudspeaker.

[0022] The embodiment is used to predict the intensity distribution of the target feedback signal in the target microphone signal by using the acoustic feedback suppression model, so as to distinguish the original sound source signal and the target feedback signal of the loudspeaker in the microphone signal through the intensity distribution, and then suppress the target feedback signal of the loudspeaker through beam processing to obtain a relatively clean sound source signal.

[0023] Among them, the sound source signal is the voice signal of the speaker. In an implementation scenario, there is at least one speaker, and the original sound source signal includes the sound source signal of each speaker.

[0024] In order to facilitate the understanding of the scheme of the present application, the following takes the vehicle-mounted four sound area K song scene without microphone as shown in Figure 2 to explain the implementation of each scheme.

[0025] As shown in Figure 2As shown, the vehicle-mounted four-zone microphone-free K-song scene includes four sound zones, different sound zones correspond to different spatial regions, and the four sound zones from left to right and from top to bottom are the main driver, the co-driver, the rear of the main driver, and the rear of the co-driver. Each sound zone is provided with a microphone and a loudspeaker, and the positions of the speakers (sound sources) in the four sound zones are relatively fixed, wherein the sound signals of the speakers in each sound zone are sound source signals.

[0026] The microphone array is an array composed of multiple microphones. For example, Figure 2 As shown, the microphone array is an array composed of four microphones. Each microphone collects the sound source signals of each sound zone and the feedback signals played by the loudspeaker. The mixed signals of the sound source signals of the four sound zones and the feedback signals played by the loudspeaker collected by each microphone are used as microphone signals.

[0027] Specifically, the mixed signals of the original sound source signals collected by each microphone and the original feedback signals played by the loudspeaker are used as original microphone signals.

[0028] In an embodiment, each original microphone signal can be directly used as the target microphone signal of the corresponding microphone. At this time, each target microphone signal includes the original sound source signal and the original feedback signal of the loudspeaker.

[0029] In another embodiment, after obtaining at least one original microphone signal collected by the microphone array, adaptive feedback cancellation processing is first performed on each original microphone signal to obtain the corresponding target microphone signal. Since the adaptive feedback cancellation processing only suppresses the feedback signal and does not process the sound source signal, the obtained target microphone signal includes the original sound source signal and the target feedback signal (the residual feedback signal after the original feedback signal is preliminarily suppressed) of the loudspeaker.

[0030] It should be noted that the adaptive feedback cancellation processing performed above is performed by using an adaptive feedback cancellation algorithm (AFC). The algorithm first estimates the impulse response of the feedback path from the loudspeaker to the microphone (i.e., the signal transmission path) through an adaptive filter, then convolves the feedback signal played by the loudspeaker with the estimated feedback path to obtain a pre-estimated feedback signal, and finally subtracts the pre-estimated feedback signal from the microphone signal to retain the sound source signal. Ideally, using a filter to model the acoustic feedback path can completely eliminate the feedback signal, however, due to the correlation between the sound source signal and the feedback signal played by the loudspeaker (the sound collected by the microphone and the sound played by the loudspeaker are very similar), the filter cannot converge to the ideal state. In addition, during the sound processing process, there may be a case where the feedback path from the loudspeaker to the microphone changes suddenly, for example, when the car karaoke is turned on or off, the space changes, causing the feedback path to change suddenly. Once the signal transmission path changes, the adaptive filter must adjust the parameters to adapt to the new path, and this process takes a certain amount of time to complete, so during the filter convergence phase, the feedback cancellation effect will be significantly worse, so it is not possible to achieve good feedback cancellation effect through adaptive feedback cancellation processing, but it is possible to first use adaptive feedback cancellation processing to eliminate part of the feedback signal, and then further perform feedback cancellation processing, such as the acoustic feedback suppression model and beam processing of the present embodiment.

[0031] wherein the corresponding Figure 2 In the scenario shown, the original sound source signal of each target microphone signal includes the sound zone source signal of at least one sound zone, and the at least one sound zone includes at least one of the main driver area, the co-driver area, the main driver rear area, and the co-driver rear area of the vehicle.

[0032] S12: For each target microphone signal, the acoustic feedback suppression model is used to predict the target microphone signal to obtain first feedback intensity information of the target microphone signal, the first feedback intensity information being used to represent the intensity distribution of the target feedback signal in the target microphone signal.

[0033] In the present embodiment, the acoustic feedback suppression model is processed for each target microphone signal respectively.

[0034] In an embodiment, for each target microphone signal, the first feedback intensity information of the target microphone signal includes: an amplitude spectrum ratio of the original sound source signal to the target microphone signal at each time-frequency point. The amplitude spectrum ratio of the original sound source signal at each time-frequency point is used to represent the intensity proportion of the original sound source signal in the target microphone signal.

[0035] That is, in this embodiment, after inputting the target microphone signal into the acoustic feedback suppression model, the acoustic feedback suppression model processes the target microphone signal, and outputs the amplitude spectrum ratio of the original sound source signal to the target microphone signal at each time-frequency point. In this embodiment, the intensity proportion of the original sound source signal at each time-frequency point of the target microphone signal is used to represent the intensity distribution of the target feedback signal in the target microphone signal, wherein the intensity proportion of the original sound source signal in the target microphone signal is negatively correlated with the intensity proportion of the target feedback signal in the target microphone signal.

[0036] Of course, since the target microphone signal contains the original sound source signal and the target feedback signal, in other embodiments, the acoustic feedback suppression model can directly output the amplitude spectrum ratio of the target feedback signal to the target microphone signal at each time-frequency point.

[0037] Among them, the acoustic feedback suppression model can be trained to have the ability to directly output the amplitude spectrum ratio of the original sound source signal to the target microphone signal at each time-frequency point, or directly output the amplitude spectrum ratio of the target feedback signal to the target microphone signal at each time-frequency point.

[0038] The training process of the acoustic feedback suppression model is described below:

[0039] First, determine the training data, for example, generate the training data by simulation. By building a closed-loop simulation system, generate multiple sets of feedback signals and clean sound source signals. In order to cover more feedback scenarios, different parameters (such as different delays, gains, etc.) within a certain range are set to simulate different loudspeaker feedback signals and clean sound source signals. After generating the training data, in order to make the model learn the main information of the data more efficiently, the training data is generally subjected to feature extraction, such as log energy spectrum features.

[0040] Then, the model selection needs to consider the modeling ability of the network and the actual landing parameter amount and calculation amount demand. The model training framework adopts a masking-based or mapping-based manner. Then, compare the labeled information in the simulation data with the model output to calculate the estimation error with a loss function (such as mean square error). Update the model parameters by backpropagating the error to obtain the gradient. After multiple iterations, the error gradually converges, and a stable model for estimating signals is obtained. The model only performs forward inference when applied.

[0041] In an embodiment, the acoustic feedback suppression model is an AFS model, and the labeled information in the simulation data is an ideal amplitude mask (IAM) of the sound source signal, which is used to represent the amplitude spectrum ratio of the sound source signal and the mixed signal with feedback (the mixed signal of the signal with feedback and the sound source signal) at each time-frequency point.

[0042] For example, refer to the following formula:

[0043]

[0044] In the formula, k and n represent frequency and time frame respectively, |S(k, n)| represents the amplitude spectrum of the sound source signal at time-frequency point (k, n), |N(k, n)| represents the amplitude spectrum of the mixed signal with feedback at time-frequency point (k, n), and mask(k, n) represents the amplitude spectrum ratio of the sound source signal and the mixed signal (target microphone signal) at time-frequency point (k, n); wherein, the sound source signal herein is not distinguished as the sound source signal of which sound zone, and the sound source signal includes the sound source signal of at least one sound zone; |S(k, n)| and |N(k, n)| are both sound signals in the training data. AFS (k, n) represents the amplitude spectrum ratio of the sound source signal and the mixed signal (target microphone signal) at time-frequency point (k, n); wherein, the sound source signal herein is not distinguished as the sound source signal of which sound zone, and the sound source signal includes the sound source signal of at least one sound zone; |S(k, n)| and |N(k, n)| are both sound signals in the training data.

[0045] S13: performing beam processing on the target microphone signal based on at least the first feedback intensity information to obtain a clean sound source signal corresponding to the target microphone signal.

[0046] In an embodiment, in a scene containing only one sound zone, such as a hearing aid scene, or in a scene without distinguishing different sound sources, the first feedback intensity information can be directly used to perform beam processing on the target microphone signal to suppress the target feedback signal in the target microphone signal, so as to obtain a clean sound source signal corresponding to the target microphone signal. For example, the signal of the time-frequency point with a relatively high feedback intensity is suppressed to a relatively large extent, and the signal of the time-frequency point with a relatively small feedback intensity is suppressed to a relatively small extent or not suppressed.

[0047] In another embodiment, in a scene containing multiple sound zones, or in a scene needing to distinguish different sound sources, the original sound source signal includes the sound zone sound source signals of multiple sound zones.

[0048] For each microphone, the sound zone where the microphone is located is regarded as the near-end sound zone of the microphone, and for the near-end sound zone of each microphone, the sound zone sound source signal of other sound zones is an interference signal of the near-end sound zone sound source signal. For example, Figure 2For example, for the sound source signal collected by the driver area microphone, the sound source signals of other audio areas in the vehicle are interference signals of the sound source signal of the driver area. In order to reduce the interference of the sound source signals of other audio areas on the sound source signals of the near-end audio area, the sound source signals of different audio areas need to be distinguished, and then the sound source signals of other audio areas outside the near-end audio area are regarded as interference signals for suppression, so that the final output sound signal only contains the clean sound source signal of the microphone near-end audio area.

[0049] Therefore, in an embodiment, before the target microphone signal is beam processed based on at least the first feedback intensity information to obtain the clean sound source signal corresponding to the target microphone signal, the audio area distinguishing information of the target microphone signal needs to be obtained, so that the audio area distinguishing information can be used to distinguish the sound source signals of different audio areas and suppress the interference sound source signals of other audio areas.

[0050] The above scheme, for each target microphone signal, first uses the sound feedback suppression model to predict the intensity distribution of the target feedback signal of the loudspeaker in the target microphone signal, and then performs beam processing on the target microphone signal based on the intensity distribution of the target feedback signal to obtain the corresponding clean sound source signal. Among them, the existing filter method using adaptive filter for sound feedback elimination depends on the current signal transmission path. Once the signal transmission path changes, the adaptive filter must adjust the parameters to adapt to the new path. This process needs a certain time to complete, so during the filter convergence stage, the feedback elimination effect will be significantly worse. The present application predicts the intensity distribution of the target feedback signal in real time through the model, and does not need to wait for the filter to converge. As long as the model outputs the intensity distribution in real time, beam processing can be performed immediately, so that sound feedback can be eliminated in real time, and the accuracy of the elimination effect can be guaranteed.

[0051] In some embodiments, for a scene containing multiple audio areas or needing to distinguish different sound sources, the audio area distinguishing information of the target microphone signal needs to be obtained, so that the audio area distinguishing information can be used to distinguish the sound source signals of different audio areas and suppress the interference sound source signals of other audio areas to obtain the clean near-end sound source signal.

[0052] In a specific embodiment, a pre-trained audio area separation model can be used to predict the target microphone signal to obtain the audio area distinguishing information of the target microphone signal. For example, for each target microphone signal, the existence probability of each audio area sound source signal in the target microphone signal can be obtained by using the audio area separation model, and used as the audio area distinguishing information of the target microphone signal.

[0053] Then, for each target microphone signal, based on the first feedback intensity information and the sound zone distinguishing information of the target microphone signal, beamforming is performed on the target microphone signal to obtain a clean sound source signal of a near-end sound zone, wherein the near-end sound zone is a sound zone where a microphone for collecting the target microphone signal is located in the microphone array. For details, please refer to the following Figure 3 The related description of the embodiments shown in the drawings.

[0054] Specifically, the sound zone distinguishing information of the target microphone signal is obtained, including the following steps:

[0055] First, the target microphone signal is predicted using a sound zone separation model to obtain target information ratios of each sound zone sound source signal and the target microphone signal at each time-frequency point; wherein the target information ratio includes an energy spectrum ratio or an amplitude spectrum ratio, and each target information ratio is used to represent the existence probability of each sound zone sound source signal at each time-frequency point.

[0056] Second, the target information ratios of each sound zone sound source signal at each time-frequency point are integrated to obtain the sound zone distinguishing information of the target microphone signal.

[0057] For example, there are four sound zones, namely sound zone 1, sound zone 2, sound zone 3 and sound zone 4. The target information ratios of each sound zone sound source signal and the target microphone signal at each time-frequency point include: the target information ratios of the sound zone 1 sound source signal at each time-frequency point, the target information ratios of the sound zone 2 sound source signal at each time-frequency point, the target information ratios of the sound zone 3 sound source signal at each time-frequency point, and the target information ratios of the sound zone 4 sound source signal at each time-frequency point. The sound zone distinguishing information of the target microphone signal obtained by integrating the target information ratios of each sound zone sound source signal at each time-frequency point includes: the integrated data of the target information ratios of all time-frequency points of sound zone 1, the integrated data of the target information ratios of all time-frequency points of sound zone 2, the integrated data of the target information ratios of all time-frequency points of sound zone 3, and the integrated data of the target information ratios of all time-frequency points of sound zone 4.

[0058] It should be noted that by training the sound zone separation model, the sound zone separation model can directly output the target information ratios of each sound zone sound source signal and the target microphone signal at each time-frequency point, wherein each target information ratio is used to represent the existence probability of each sound zone sound source signal at each time-frequency point, that is, the greater the ratio value, the greater the probability of the corresponding sound zone sound source signal.

[0059] The training process of the sound zone separation model is described below:

[0060] First, determine the training data, for example, generate the training data by simulation. Taking Figure 2 For example, taking the vehicle-mounted multi-sound zone scene shown in the drawing as an example, the positions of the microphone and the speaker need to be considered, such as Figure 2As shown, the microphone position is fixed, the speaker moves within a certain range of four seats, the propagation path from the sound source position to the microphone is recorded by simulating the room impulse response, then the simulated impulse response is convolved with the sound source signal of the speaker to obtain the sound source signal received by the microphone. In addition, considering the scenario of multiple people speaking at the same time, the amplitude of the sound source signal at different positions is limited according to a certain signal-to-interference ratio, so that the amplitudes of the voices of speakers at different positions are different, thereby simulating various speaking scenarios. After obtaining the microphone signal and the label of the target speaker (near-end speaker) of each zone, the log energy spectrum features and the phase difference features of the microphone array signal are extracted.

[0061] Then, the extracted features are input into the sound zone separation model, and a feedforward sequence memory neural network is used to separate the sound source signal of each zone. Of course, the microphone signals of each zone can also be directly input into the sound zone separation model, and the sound zone separation model can be used to extract features and predict the sound source signal of each zone.

[0062] Among them, the difference is obtained by comparing the labeled information in the simulation data with the predicted information output by the model, and then the model parameter is updated based on the difference. The labeled information can be the energy spectrum proportion of the near-end sound source signal of each zone and the target microphone signal, or the amplitude spectrum proportion of the near-end sound source signal of each zone and the target microphone signal.

[0063] In a specific embodiment, the labeled information is an ideal ratio mask (IRM) of the near-end sound source signal corresponding to each microphone, which is used to represent the energy spectrum proportion of the near-end sound source signal of each zone and the target microphone signal.

[0064] For example, please refer to the following formula:

[0065]

[0066] In the formula, k and n respectively represent frequency and time frame, Ei(k,n) represents the energy spectrum of the i-th sound source (near-end sound source) at the time-frequency point (k, n), and mask i (k,n) represents the energy spectrum proportion of the i-th sound source signal and the target microphone signal at the time-frequency point (k, n). The training loss function is defined as the mean square error value between the model predicted mask i_pre and the real mask i_true .

[0067] Wherein, after obtaining the first feedback intensity information and the sound zone separation information, the target microphone signal can be further processed based on the first feedback intensity information and the sound zone separation information of the target microphone signal to obtain the clean sound source signal of the near-end sound zone.

[0068] Specifically, refer to Figure 3 , Figure 3 is Figure 1 a flowchart of an embodiment of step S13. In this embodiment, based on the first feedback intensity information of the target microphone signal and the sound zone distinguishing information, the target microphone signal is beamformed to obtain the clean sound source signal of the near-end sound zone, including:

[0069] S31: fuse the first feedback intensity information of the target microphone signal and the sound zone distinguishing information to obtain first fusion data; the first fusion data is used to represent the first confidence probability that the target microphone signal at each time-frequency point belongs to each sound zone sound source signal and is not the target feedback signal.

[0070] In an embodiment, the product between the first feedback intensity information of the target microphone signal and the sound zone distinguishing information can be obtained, and the product is taken as the first fusion data.

[0071] In a specific embodiment, the first feedback intensity information of the target microphone signal is the amplitude spectrum ratio of the original sound source signal and the target microphone signal at each time-frequency point, and the sound zone distinguishing information is the target information proportion (for example, the energy spectrum proportion) of each sound zone sound source signal and the target microphone signal at each time-frequency point, so the first fusion data is the product between the first feedback intensity information and the target information proportion of each sound zone sound source signal and the target microphone signal at each time-frequency point. The first fusion data includes the sub-fusion data of four sound zones. For each sound zone (for example, sound zone 1), the sub-fusion data of sound zone 1 is the product between the first feedback intensity information and the target information proportion of sound zone 1 sound source signal and the target microphone signal at each time-frequency point. The corresponding sub-fusion data of sound zone 1 is used to represent the first confidence probability that the target microphone signal at each time-frequency point belongs to the sound zone 1 sound source signal and is not the target feedback signal.

[0072] S32: beamforming the target microphone signal using the first fusion data to obtain the clean sound source signal of the near-end sound zone.

[0073] In an embodiment, step S32 further includes the following steps:

[0074] First, determine the target filter coefficient using the first fusion data; the target filter coefficient includes the coefficient value at each time-frequency point, and the first confidence probability of each time-frequency point is positively correlated with the corresponding coefficient value.

[0075] In one embodiment, a beamforming algorithm can be used to obtain target filtering coefficients based on the first fused data. Specifically, with the goal of maximizing the signal-to-noise ratio (the ratio between the near-end sound source signal and the noise signal), the feature vector corresponding to the near-end sound source signal is solved, and this feature vector is used as the target filtering coefficient. The noise signal includes sound source signals from other sound regions (interference signals) and the feedback signal from the loudspeaker. The solved target filtering coefficients include coefficient values ​​at each time-frequency point, and the first confidence probability at each time-frequency point is positively correlated with the corresponding coefficient value.

[0076] Second, the target microphone signal is beamformed using the target filtering coefficients to obtain a clean sound source signal in the near-end sound range.

[0077] In one embodiment, the target feedback signal and other sound source signals in the target microphone signal, except for the near-end sound region, can be suppressed first using the target filtering coefficients, and the sound source signals in the near-end sound region can be preserved or enhanced to obtain an intermediate processed signal; then, the intermediate processed signal can be used to obtain a clean sound source signal in the near-end sound region.

[0078] The microphones that collect signals from each target microphone are distributed in different frequency ranges, and the frequency range where the microphone that collects the target microphone signal is located is the near-end frequency range of the target microphone signal. For example... Figure 2 As shown, microphones are distributed in the driver's seat area, passenger seat area, area behind the driver's seat, and area behind the passenger seat. For the microphones in the driver's seat area, beamforming is applied to the target microphone signal acquired by the driver's seat microphone to obtain the clean sound source signal of the speaker in the driver's seat area. Similarly, for the microphones in the passenger seat area, beamforming is applied to the target microphone signal acquired by the passenger seat microphone to obtain the clean sound source signal of the speaker in the passenger seat area. Through this method, the clean sound source signals of the near-end sound range of each microphone can be obtained.

[0079] In this embodiment, beamforming of the target microphone signal using the target filtering coefficient is to suppress the target feedback signal and other sound source signals except for the near-end sound region, while preserving or enhancing the sound source signals of the near-end sound region, thus obtaining a cleaner near-end sound source signal (i.e., a clean sound source signal in the near-end sound region).

[0080] In one embodiment, the intermediate processed signal obtained after beam processing can be directly used as the clean sound source signal of the near-end sound zone corresponding to the microphone's sound zone.

[0081] In another embodiment, further sound feedback suppression and interference signal suppression processing can be performed on the basis of the intermediate processed signals, so as to obtain cleaner sound source signals of the near-end sound zones. Specifically, for each intermediate processed signal, a sound feedback suppression model can be used to predict the intermediate processed signal, to obtain second feedback intensity information; the second feedback intensity information and the sound zone distinguishing information are fused to obtain second fusion data; and based on the second fusion data, a clean sound source signal of the near-end sound zone is obtained; wherein the second fusion data is used to represent a second confidence probability of the intermediate processed signal belonging to the sound source signals of each sound zone at each time-frequency point.

[0082] wherein, based on the second confidence probability of each time-frequency point, a first time-frequency point with a second confidence probability lower than a certain threshold and a second time-frequency point with a second confidence probability not lower than a certain threshold can be found from each time-frequency point, the sound signal of the first time-frequency point is suppressed to a certain extent, and the sound signal of the second time-frequency point is retained, so as to obtain the clean sound source signal of the near-end sound zone.

[0083] In some embodiments, after obtaining the clean sound source signals corresponding to the target microphone signals of each channel, at least one of the following operations can be performed:

[0084] First, a superimposed signal of the clean sound source signals corresponding to the target microphone signals of each channel is obtained, and the superimposed signal is played.

[0085] This way can make the sound played by the loudspeaker include the sound of the speakers in each sound zone, which is more suitable for a multi-person chatting scenario.

[0086] Second, in response to the selection of at least one sound zone, the clean sound source signals corresponding to the target microphone signals of the near-end sound zone are played.

[0087] This way provides the user with more choices, and the user can choose to play only the sound signal of the speaker in a specific sound zone, for example, the sound signal of the main speaker, which is more suitable for a conference scenario and a karaoke scenario.

[0088] In a specific embodiment, please refer to Figure 4 , Figure 4 is a flow framework diagram of an embodiment of the sound feedback and interference suppression provided by the present application.

[0089] Taking the acquisition of two microphone signals by a microphone array (the microphone signals not processed are referred to as original microphone signals) as an example, for each original microphone signal (including an original sound source signal and an original feedback signal), an adaptive filtering process is first performed to suppress part of the feedback signal in the microphone signal, to obtain a target microphone signal of each channel.

[0090] Then, for each target microphone signal, the first feedback intensity information and the sound region separation information of the target microphone signal are fused to obtain first fusion information, and the first fusion information is used to guide beam processing to obtain an intermediate processing signal of the target microphone signal. In the intermediate processing signal, most of the feedback signals and the interference signals of other sound regions have been suppressed, and the intermediate processing signal is essentially a cleaner near-end sound region source signal.

[0091] Further, for each target microphone signal, further suppression of interference signals and feedback signals can be performed based on the cleaner intermediate processing signal to obtain a cleaner near-end sound region source signal. Specifically, the intermediate processing signal of each target microphone signal is input into a sound feedback suppression model to obtain second feedback intensity information of residual feedback signals in the intermediate processing signal, and then the second feedback intensity information and the sound region separation information output by the sound region separation model are fused to obtain second fusion data. The second fusion data is used to represent a second confidence probability of the intermediate processing signal belonging to each sound region source signal at each time-frequency point, and a cleaner near-end sound region source signal is obtained based on the second confidence probability.

[0092] Further, for each target microphone signal, further suppression of interference signals and feedback signals can be performed based on the cleaner intermediate processing signal to obtain a cleaner near-end sound region source signal. Specifically, the intermediate processing signal of each target microphone signal is input into a sound feedback suppression model to obtain second feedback intensity information of residual feedback signals in the intermediate processing signal, and then the second feedback intensity information and the sound region separation information output by the sound region separation model are fused to obtain second fusion data. The second fusion data is used to represent a second confidence probability of the intermediate processing signal belonging to each sound region source signal at each time-frequency point, and a cleaner near-end sound region source signal is obtained based on the second confidence probability.

[0093] Please refer to Figure 5 , Figure 5 is a framework schematic diagram of an embodiment of the sound feedback elimination device provided in the present application. In the embodiment, the sound feedback elimination device 50 includes an acquisition module 51, a prediction module 52, and a beam processing module 53. The acquisition module 51 is configured to acquire at least one target microphone signal collected by a microphone array; the target microphone signal includes an original sound source signal and a target feedback signal of a loudspeaker; the prediction module 52 is configured to, for each target microphone signal, use a sound feedback suppression model to predict the target microphone signal to obtain first feedback intensity information of the target microphone signal, the first feedback intensity information being used to represent an intensity distribution of the target feedback signal in the target microphone signal; and the beam processing module 53 is configured to perform beam processing on the target microphone signal based on at least the first feedback intensity information to obtain a clean sound source signal corresponding to the target microphone signal.

[0094] In some embodiments, the first feedback intensity information of the target microphone signal predicted by the prediction module 52 comprises a ratio of amplitudes of the original sound source signal and the target microphone signal at each time-frequency point; wherein the ratio of amplitudes of the original sound source signal at each time-frequency point is used to represent the intensity proportion of the original sound source signal in the target microphone signal.

[0095] In some embodiments, the original sound source signal comprises sound source signals of multiple sound zones; different sound zones correspond to different spatial regions; before the beam processing module 53 performs beam processing on the target microphone signal based on at least the first feedback intensity information to obtain the clean sound source signal corresponding to the target microphone signal, the acoustic feedback elimination device 50 is further configured to: for each target microphone signal, obtain the existence probability of each sound zone sound source signal in the target microphone signal as the sound zone distinguishing information of the target microphone signal; the beam processing module 53 performs beam processing on the target microphone signal based on at least the first feedback intensity information to obtain the clean sound source signal corresponding to the target microphone signal, comprising: based on the first feedback intensity information and the sound zone distinguishing information of the target microphone signal, performing beam processing on the target microphone signal to obtain the clean sound source signal of the near-end sound zone, the near-end sound zone being the sound zone where the microphone collecting the target microphone signal is located.

[0096] In some embodiments, obtaining the existence probability of each sound zone sound source signal in the target microphone signal as the sound zone distinguishing information of the target microphone signal comprises: predicting the target information proportion of each sound zone sound source signal and the target microphone signal at each time-frequency point by using a sound zone separation model; wherein the target information proportion comprises an energy spectrum proportion or an amplitude spectrum proportion, and each target information proportion is used to represent the existence probability of each sound zone sound source signal at each time-frequency point; and synthesizing the target information proportions of each sound zone sound source signal at each time-frequency point to obtain the sound zone distinguishing information of the target microphone signal.

[0097] In some embodiments, based on the first feedback intensity information and the sound zone distinguishing information of the target microphone signal, performing beam processing on the target microphone signal to obtain the clean sound source signal of the near-end sound zone comprises: fusing the first feedback intensity information and the sound zone distinguishing information of the target microphone signal to obtain first fusion data; the first fusion data is used to represent the first confidence probability that the target microphone signal at each time-frequency point belongs to each sound zone sound source signal and is not the target feedback signal; and using the first fusion data to perform beam processing on the target microphone signal to obtain the clean sound source signal of the near-end sound zone.

[0098] In some embodiments, fusing the first feedback intensity information and the sound zone distinguishing information of the target microphone signal to obtain the first fusion data comprises: obtaining the product between the first feedback intensity information and the sound zone distinguishing information of the target microphone signal as the first fusion data.

[0099] In some embodiments, the beamforming of the target microphone signal by using the first fusion data to obtain the clean sound source signal of the near-end sound region comprises: determining target filter coefficients by using the first fusion data; the target filter coefficients comprise coefficient values at each time-frequency point, and the first confidence probability of each time-frequency point is positively correlated with the corresponding coefficient value; and performing beamforming on the target microphone signal by using the target filter coefficients to obtain the clean sound source signal of the near-end sound region.

[0100] The beamforming of the target microphone signal by using the target filter coefficients to obtain the clean sound source signal of the near-end sound region comprises: performing suppression processing on the target feedback signal and the sound source signal of other sound regions except the near-end sound region in the target microphone signal by using the target filter coefficients, and performing reservation or enhancement processing on the sound region sound source signal of the near-end sound region to obtain an intermediate processing signal; and obtaining the clean sound source signal of the near-end sound region by using the intermediate processing signal.

[0101] In some embodiments, the obtaining of the clean sound source signal of the near-end sound region by using the intermediate processing signal comprises: taking the intermediate processing signal as the clean sound source signal of the near-end sound region; or predicting the intermediate processing signal by using a sound feedback suppression model to obtain second feedback intensity information; fusing the second feedback intensity information and the sound region distinguishing information to obtain second fusion data; and obtaining the clean sound source signal of the near-end sound region based on the second fusion data; wherein the second fusion data is used to represent the second confidence probability of the intermediate processing signal belonging to each sound region sound source signal at each time-frequency point.

[0102] In some embodiments, the microphones for collecting each target microphone signal are distributed in different sound regions, and the sound region where the microphone for collecting the target microphone signal is located is the near-end sound region of the target microphone signal; after obtaining the clean sound source signal corresponding to each target microphone signal, the sound feedback elimination device 50 is further configured to: obtain a superimposed signal of the clean sound source signals corresponding to each target microphone signal, and play the superimposed signal; and / or, in response to the selection of at least one sound region, play the clean sound source signal corresponding to the target microphone signal of the near-end sound region which is the selected sound region.

[0103] In some embodiments, the original sound source signal of the target microphone signal obtained by the obtaining module 51 comprises the sound region sound source signal of at least one sound region, and the at least one sound region comprises at least one of a main driver area, a co-driver area, a main driver rear area and a co-driver rear area of a vehicle; and / or the obtaining module 51 obtains at least one target microphone signal collected by a microphone array, comprising: obtaining at least one original microphone signal collected by the microphone array; the original microphone signal comprises an original sound source signal and an original feedback signal of a loudspeaker; and performing adaptive feedback elimination processing on each original microphone signal to obtain a corresponding target microphone signal.

[0104] Please refer to Figure 6 , Figure 6 is a framework schematic diagram of an embodiment of the electronic device provided in the present application. In the embodiment, the electronic device 60 comprises a memory 61 and a processor 62 coupled with each other.

[0105] The memory 61 stores program instructions, and the processor 62 is configured to execute the program instructions stored in the memory 61 to implement the steps of any of the above methods. In a specific implementation scenario, the electronic device 60 can include but is not limited to a microcomputer, a server, and in addition, the electronic device 60 can also include a notebook computer, a tablet computer and other mobile devices, which are not limited herein.

[0106] Specifically, the processor 62 is configured to control itself and the memory 61 to implement the steps of any of the above embodiments. The processor 62 can also be referred to as a CPU (Central Processing Unit). The processor 62 can be an integrated circuit chip having a processing capability of signals. The processor 62 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 62 can be jointly implemented by integrated circuit chips.

[0107] Please refer to Figure 7 , Figure 7 is a framework schematic diagram of a computer readable storage medium provided in the present application. The computer readable storage medium 70 of the embodiment of the present application stores program instructions 71, which when executed implement the method provided in any of the above embodiments and any non-conflicting combination. Wherein the program instructions 71 can form a program file and be stored in the above computer readable storage medium 70 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) executes all or part of the steps of the methods of various embodiments of the present application. And the aforementioned computer readable storage medium 70 includes: a U disk, a mobile hard disk, a ROM (Read-Only Memory), a RAM (Random Access Memory), a magnetic disk or an optical disk and various media that can store program codes, or a computer, a server, a mobile phone, a tablet and other terminal devices.

[0108] The above scheme, for each target microphone signal, first uses the acoustic feedback suppression model to predict the intensity distribution of the target feedback signal of the loudspeaker in the target microphone signal, and then performs beam processing on the target microphone signal based on the intensity distribution of the target feedback signal to obtain the corresponding clean sound source signal. Among them, the existing filter method using adaptive filter for acoustic feedback elimination depends on the current signal transmission path. Once the signal transmission path changes, the adaptive filter must adjust the parameters to adapt to the new path. This process needs a certain time to complete, so in the filter convergence stage, the feedback elimination effect will be significantly worse. The present application predicts the intensity distribution of the target feedback signal in real time through the model, without waiting for the filter to converge. As long as the model outputs the intensity distribution in real time, beam processing can be performed immediately, so that the acoustic feedback can be eliminated in real time, and the accuracy of the elimination effect can be guaranteed.

[0109] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiment descriptions. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0110] The above description of each embodiment tends to emphasize the differences between each embodiment, and the same or similar parts can be mutually referred to. For the sake of brevity, they will not be repeated here.

[0111] In several embodiments provided by the present application, it should be understood that the disclosed method and device can be implemented by other means. For example, the above-described device implementation is only schematic; for example, the division of the modules or units is only a logical function division, and there can be another division manner in actual implementation; for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed mutual elements can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other form.

[0112] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the present embodiment scheme.

[0113] In addition, each of the functional units in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0114] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (such as a personal computer, a server, or a network device) or a processor to perform all or part of the steps of the methods in the embodiments of the present application. The foregoing storage medium includes: various types of U disks, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical disks, and various other media that can store program codes.

[0115] The above only describes the embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation based on the content of the specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method of acoustic feedback cancellation, characterized by, The method comprises: acquiring at least one target microphone signal collected by a microphone array; the target microphone signal comprises an original sound source signal and a target feedback signal of a loudspeaker; for each target microphone signal, using an acoustic feedback suppression model to predict the target microphone signal to obtain first feedback intensity information of the target microphone signal, the first feedback intensity information being used to represent intensity distribution of the target feedback signal in the target microphone signal; based on at least the first feedback intensity information, performing beam processing on the target microphone signal to obtain a clean sound source signal corresponding to the target microphone signal.

2. The method of claim 1, wherein, The first feedback intensity information of the target microphone signal comprises an amplitude spectrum ratio of the original sound source signal and the target microphone signal at each time-frequency point; wherein the amplitude spectrum ratio of the original sound source signal at each time-frequency point is used to represent the intensity proportion of the original sound source signal in the target microphone signal.

3. The method of claim 1, wherein, The original sound source signal comprises sound zone sound source signals of multiple sound zones; different sound zones correspond to different spatial regions; Before the at least based on the first feedback intensity information, performing beam processing on the target microphone signal to obtain a clean sound source signal corresponding to the target microphone signal, the method further comprises: for each target microphone signal, acquiring an existence probability of each sound zone sound source signal in the target microphone signal as sound zone distinguishing information of the target microphone signal; The at least based on the first feedback intensity information, performing beam processing on the target microphone signal to obtain a clean sound source signal corresponding to the target microphone signal, comprises: based on the first feedback intensity information and the sound zone distinguishing information of the target microphone signal, performing beam processing on the target microphone signal to obtain a clean sound source signal of a near-end sound zone, the near-end sound zone being a sound zone where a microphone collecting the target microphone signal is located.

4. The method of claim 3, wherein, The acquiring an existence probability of each sound zone sound source signal in the target microphone signal as sound zone distinguishing information of the target microphone signal comprises: using a sound zone separation model to predict the target microphone signal to obtain a target information proportion of each sound zone sound source signal and the target microphone signal at each time-frequency point; wherein the target information proportion comprises an energy spectrum proportion or an amplitude spectrum proportion, and each target information proportion is used to represent an existence probability of each sound zone sound source signal at each time-frequency point; integrating the target information proportions of each sound zone sound source signal at each time-frequency point to obtain the sound zone distinguishing information of the target microphone signal.

5. The method of claim 3, wherein, The based on the first feedback intensity information and the sound zone distinguishing information of the target microphone signal, performing beam processing on the target microphone signal to obtain a clean sound source signal of a near-end sound zone comprises: fusing the first feedback intensity information and the sound zone distinguishing information of the target microphone signal to obtain first fusion data; the first fusion data is used to represent a first confidence probability that the target microphone signal at each time-frequency point belongs to each sound zone sound source signal and is not a target feedback signal; Beamforming the target microphone signal by using the first fusion data to obtain a clean sound source signal of the near-end sound region.

6. The method of claim 5, wherein, The first fusion data is obtained by fusing the first feedback intensity information and the sound region distinguishing information of the target microphone signal, and comprises: The product of the first feedback intensity information and the sound region distinguishing information of the target microphone signal is obtained as the first fusion data.

7. The method of claim 5, wherein, The beamforming the target microphone signal by using the first fusion data to obtain a clean sound source signal of the near-end sound region comprises: The target filter coefficient is determined by using the first fusion data; the target filter coefficient comprises a coefficient value at each time-frequency point, and the first confidence probability of each time-frequency point is positively correlated with the corresponding coefficient value; The target filter coefficient is used to perform beamforming on the target microphone signal to obtain a clean sound source signal of the near-end sound region.

8. The method of claim 7, wherein, The target filter coefficient is used to perform beamforming on the target microphone signal to obtain a clean sound source signal of the near-end sound region comprises: The target filter coefficient is used to perform suppression processing on the target feedback signal and the sound source signal of other sound regions except the near-end sound region in the target microphone signal, and perform reservation or enhancement processing on the sound region sound source signal of the near-end sound region to obtain an intermediate processing signal; The clean sound source signal of the near-end sound region is obtained by using the intermediate processing signal.

9. The method of claim 8, wherein, The clean sound source signal of the near-end sound region is obtained by using the intermediate processing signal comprises: The intermediate processing signal is taken as the clean sound source signal of the near-end sound region; Or, the intermediate processing signal is predicted by using the acoustic feedback suppression model to obtain second feedback intensity information; the second feedback intensity information and the sound region distinguishing information are fused to obtain second fusion data; and the clean sound source signal of the near-end sound region is obtained based on the second fusion data; The second fusion data is used to represent the second confidence probability of the intermediate processing signal belonging to each sound region sound source signal at each time-frequency point.

10. The method of claim 1, wherein, The microphones of each of the target microphone signals are distributed in different sound regions, and the sound region where the microphone of the target microphone signal is located is the near-end sound region of the target microphone signal; After obtaining the clean sound source signal corresponding to each of the target microphone signals, the method further comprises: An overlap signal of the clean sound source signals corresponding to each of the target microphone signals is obtained, and the overlap signal is played; And / or, In response to the selection of at least one sound region, the clean sound source signal corresponding to the target microphone signal of the selected sound region is played as the near-end sound region.

11. The method of claim 1, wherein, The original sound source signal of the target microphone signal comprises sound region sound source signals of at least one sound region, and the at least one sound region comprises at least one of a main driver area, a co-driver area, a main driver rear area and a co-driver rear area of a vehicle; And / or, the at least one target microphone signal collected by the microphone array comprises: At least one original microphone signal collected by the microphone array is obtained; the original microphone signal comprises an original sound source signal and an original feedback signal of a loudspeaker; Adaptive feedback cancellation processing is performed on each of the original microphone signals to obtain a corresponding target microphone signal.

12. An acoustic feedback cancelling apparatus, characterized by The apparatus comprises: An acquisition module configured to acquire at least one target microphone signal collected by a microphone array, the target microphone signal comprising an original sound source signal and a target feedback signal of a loudspeaker; A prediction module configured to, for each of the target microphone signals, predict the target microphone signal using a sound feedback suppression model to obtain first feedback intensity information of the target microphone signal, the first feedback intensity information being used to represent intensity distribution of the target feedback signal in the target microphone signal; A beam processing module configured to perform beam processing on the target microphone signal based on at least the first feedback intensity information to obtain a clean sound source signal corresponding to the target microphone signal.

13. An electronic device, comprising: comprise a memory and a processor coupled to each other, The memory stores program instructions; The processor is configured to execute the program instructions stored in the memory to implement the method of any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions capable of being executed by the processor, and the program instructions are capable of being executed by the processor to implement the method of any one of claims 1-11.