Deep learning-based echo cancellation method, device and readable storage medium
By employing a deep learning-based echo cancellation method, which utilizes neural networks to process the compressed complex spectrum of signals from both near and far microphones, the problem of echo suppression in virtual reality and augmented reality is solved, resulting in a more realistic auditory experience.
Patent Information
- Application Number
- PCT/CN2024/105520
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2026-01-22
AI Technical Summary
In virtual reality and augmented reality technologies, echo problems affect the user's auditory experience, especially in multi-room conversations where echoes are difficult to suppress effectively.
A deep learning-based echo cancellation method is adopted. By acquiring the compressed complex spectrum of the far-end and near-end microphone signals, the echo cancellation process is performed using a trained neural network model, and a short-time Fourier inverse transform is performed to restore the clean near-end speech signal.
It effectively suppresses echo components, enhances the auditory realism of virtual space sound systems, and improves user experience.
Smart Images

Figure CN2024105520_22012026_PF_FP_ABST
Abstract
Description
A deep learning-based echo cancellation method, device, and readable storage medium Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more particularly to the field of speech processing technology, applicable to echo cancellation scenarios in virtual space sound systems. Specifically, this application discloses an echo cancellation method, device, and readable storage medium based on deep learning. Background Technology
[0002] With the development of modern communication, virtual reality, and augmented reality technologies, users sometimes simultaneously communicate with different speakers in multiple rooms. To achieve the goal of "virtual reality," audio design sometimes requires reconstructing information from speakers in different distant rooms to different locations in the near room. Furthermore, the virtual location changes accordingly as the actual location of the distant speakers moves, thus satisfying the listener's auditory realism. Auditory realism is a crucial factor in metaverse design, combining visual and tactile elements to provide users with a more immersive virtual environment. However, reverberation is unavoidable in real acoustic environments, so echoes often occur during two-way conversations, affecting the user's auditory experience.
[0003] It is important to note that the techniques described in this section are not necessarily those previously conceived or adopted. Unless otherwise specified, no technique described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be recognized in any prior art. Technical issues
[0004] This application provides a deep learning-based echo cancellation method, device, and readable storage medium, aiming to at least partially solve one of the problems in related technologies. Technical solutions
[0005] To address the aforementioned technical problems, the first aspect of this application provides a deep learning-based echo cancellation method, comprising:
[0006] Acquire the corresponding remote microphone signal in the remote room, and acquire the corresponding near microphone signal in the near room.
[0007] Using the remote microphone signal as a reference signal, a first compressed complex spectrum corresponding to the reference signal is obtained, and a second compressed complex spectrum corresponding to the near microphone signal is obtained.
[0008] The first compressed complex spectrum and the second compressed complex spectrum are input into the trained neural network model for echo cancellation processing, and the compressed complex spectrum of near-end speech is output.
[0009] A clean near-end speech signal is obtained by performing a short-time inverse Fourier transform on the complex spectrum of the near-end speech compression.
[0010] A second aspect of this application provides a deep learning-based echo cancellation device, comprising:
[0011] The signal acquisition module is used to acquire the corresponding remote microphone signal in the remote room and the corresponding near microphone signal in the near room.
[0012] The complex spectrum acquisition module is used to use the far-end microphone signal as a reference signal to acquire the first compressed complex spectrum corresponding to the reference signal, and to acquire the second compressed complex spectrum corresponding to the near-end microphone signal.
[0013] The model processing module is used to input the first compressed complex spectrum and the second compressed complex spectrum into the trained neural network model for echo cancellation processing and output the near-end speech compressed complex spectrum;
[0014] The signal transformation module is used to perform short-time inverse Fourier transform on the complex spectrum of near-end speech compression to obtain a clean near-end speech signal.
[0015] A third aspect of this application provides an electronic device, including a memory and a processor, wherein the processor is configured to execute a computer program stored in the memory; when the processor executes the computer program, it implements the steps of the deep learning-based echo cancellation method provided in the first aspect of this application.
[0016] The fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the deep learning-based echo cancellation method provided in the first aspect of this application. Beneficial effects
[0017] As can be seen from the above, according to the deep learning-based echo cancellation method, device, and readable storage medium provided in this application, the corresponding far-end microphone signal of the far-end room and the corresponding near-end microphone signal of the near-end room are acquired; the far-end microphone signal is used as a reference signal to acquire the corresponding first compressed complex spectrum and the corresponding second compressed complex spectrum of the near-end microphone signal; the first and second compressed complex spectra are input into the trained neural network model for echo cancellation processing, and the near-end speech compressed complex spectrum is output; a short-time inverse Fourier transform is performed on the near-end speech compressed complex spectrum to obtain a clean near-end speech signal. Through the implementation of this application, by using the microphone signals of each far-end room as neural network reference information to recover the clean near-end speech signal, echo components can be effectively suppressed, thereby realizing echo suppression in a virtual space sound system.
[0018] It should be understood that the description in this section is not intended to identify key or important features of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0019] The accompanying drawings exemplify embodiments and form part of the specification, working together with the textual description to explain exemplary implementations of the embodiments. The drawings shown are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0020] Figure 1 is a schematic diagram of the basic process of a deep learning-based echo cancellation method provided in an embodiment of this application;
[0021] Figure 2 is a schematic diagram of an echo cancellation application scenario based on VBAP technology provided in an embodiment of this application;
[0022] Figure 3 is a schematic diagram of the structure of a neural network model provided in an embodiment of this application;
[0023] Figure 4 is a characterization diagram of test results corresponding to different speaker layouts provided in an embodiment of this application;
[0024] Figure 5 is a schematic diagram of an echo cancellation application scenario based on Ambisonic technology provided in an embodiment of this application;
[0025] Figure 6 is a characterization diagram of test results for processing recorded data using different algorithms under different speaker layouts according to an embodiment of this application;
[0026] Figure 7 is a schematic diagram of the functional modules of a deep learning-based echo cancellation device provided in an embodiment of this application;
[0027] Figure 8 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Embodiments of the present invention
[0028] To make the inventive objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0029] In the description of the embodiments of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. The term "multiple" means two or more, unless otherwise explicitly specified. The term "comprising" indicates the presence of the described feature, whole, step, operation, element, and / or component, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or sets thereof. The term "and / or" describes the relationship between related objects, indicating that three relationships may exist. For example, A and / or B may include three cases: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the preceding and following related objects belong to an "or" relationship.
[0030] To address the echo problem in virtual space sound systems provided by related technologies, this application provides an echo cancellation method based on deep learning in one embodiment. Figure 1 is a basic flowchart of the echo cancellation method based on deep learning provided in this embodiment. The echo cancellation method based on deep learning includes the following steps:
[0031] Step 101: Obtain the corresponding remote microphone signal in the remote room and the corresponding near microphone signal in the near room.
[0032] Specifically, in practical applications, it is necessary to reconstruct the voice of a speaker in a distant room to different locations in a near room. Notably, the number of distant rooms can be K, where K is greater than or equal to 1. The near rooms are equipped with speaker arrays, which can be arranged in a ring. In this embodiment, when a speaker in the distant room speaks, the speaker's voice can be reconstructed based on the speakers in the near rooms. Multiple virtual sound source locations can be set in the near rooms; each virtual sound source location can be the location of a specific speaker or a location between two adjacent speakers. It should also be noted that both the near and distant rooms can contain speakers and microphones, and the voice signal played by the speakers is recorded by the microphones through an acoustic echo path.
[0033] Vector Basis Amplitude Translation (VBAP) is a virtual sound image synthesis technique. Compared to sound field synthesis methods, it requires fewer speakers and offers greater flexibility in speaker placement. Furthermore, VBAP can be applied to various audio systems, such as stereo, surround sound, and spatial sound systems, and has been incorporated into the MPEG-H standard. In this embodiment, for a speaker array on a ring, when reproducing a virtual sound source using VBAP, typically two speakers on either side of the virtual sound source are used. However, in scenarios where the virtual sound source and speaker positions coincide, the speaker at the virtual sound source's location is used directly for reproduction.
[0034] Figure 2 shows a schematic diagram of an echo cancellation application scenario based on VBAP technology provided in this embodiment. In the figure, r... k (n) represents the speech signal of the speaker in the k-th remote room. This indicates the location of the virtual sound source. In some embodiments of this example, the step of obtaining the corresponding near-end microphone signal in the near-end room includes: obtaining the near-end speaker signal corresponding to the k-th speaker in the near-end room; and obtaining the corresponding near-end microphone signal based on the near-end speaker signal.
[0035] The near-end speaker signal is represented as:
[0036] Where K represents the number of remote rooms, δ p,k This represents the gain of the p-th loudspeaker relative to the sound source signal in the k-th far room, obtained using the VBAP algorithm. This represents the signal from the far-end speaker in one of the channels of the k-th far-end room;
[0037] The near-end microphone signal is represented as:
[0038] Where P is the number of loudspeakers arrayed in the near-end room, h p (n) represents the echo path, s(n) represents the near-end speech signal, and v(n) represents additive noise.
[0039] Step 102: Using the far-end microphone signal as a reference signal, obtain the first compressed complex spectrum corresponding to the reference signal, and obtain the second compressed complex spectrum corresponding to the near-end microphone signal.
[0040] Specifically, in this embodiment, before inputting the far-end microphone signal and the near-end microphone signal into the neural network model, a short-time Fourier transform is performed on the two signals, that is, the signals are converted from the time domain to the time-frequency domain, and the corresponding compressed complex spectra are obtained respectively.
[0041] In some embodiments of this example, the first compressed complex spectrum is represented as:
[0042] in, Represents the first compressed complex spectrum. θ and θ represent the amplitude and phase information of the corresponding remote microphone signal in the k-th remote room, respectively. β is a preset constant, and the preferred value of β can be 1 / 2.
[0043] Step 103: Input the first compressed complex spectrum and the second compressed complex spectrum into the trained neural network model for echo cancellation processing, and output the near-end speech compressed complex spectrum.
[0044] Figure 3 shows a schematic diagram of the structure of a neural network model provided in this embodiment. Specifically, the neural network model in this embodiment includes: a sequentially connected encoder network, a temporal modeling network, a decoder network, and a linear layer. The encoder network is also skip-connected to the decoder network. In a preferred embodiment, the temporal modeling network includes two parallel Long Short-Term Memory (LSTM) networks.
[0045] The compressed complex spectrum includes real and imaginary part information. In practical applications, this embodiment can compress the first complex spectrum. Second compressed complex spectrum After performing the concatenation operation, the input features are used as input features of the neural network model. The parameters of the input features are [10, T, 161]. In this embodiment, the encoding network may optionally include five cascaded gated convolutional layers (Conv_GLU), with parameters of [16, T, 80], [32, T, 39], [64, T, 19], [128, T, 9], and [256, T, 4], respectively. Next, the signal features extracted by the encoding network are used as input to a time-series modeling network that includes two parallel long short-term memory networks. The time-series modeling network is then used to obtain the temporal relationship of the signal. The parameters of the Long Short-Term Memory (LSTM) network are [T, 1024]. The output features of the temporal modeling network then serve as input to five cascaded gated deconvolutional layers (Deconv_GLU). The parameters of these five layers are [128, T, 9], [64, T, 19], [32, T, 39], [16, T, 80], and [1, T, 161]. The decoded features are processed by two parallel linear layers to obtain the overall output of the neural network model, which is the complex spectrum of near-end speech compression. The parameters of the linear layer are [T, 161].
[0046] In this embodiment, an initial neural network model is trained based on preset training samples. A preset loss function is used to determine whether the training loss value meets the model convergence condition. The loss function is used to measure the difference between the predicted output and the sample label during the training phase. If the model does not meet the convergence condition, the parameters of the neural network model are further adjusted and iterative training continues until the model converges, thus obtaining the trained neural network model.
[0047] Continuing with the neural network model in Figure 3 above, the decoding network includes two decoding network branches, each connected to a linear layer; the loss function of the neural network model is expressed as:
[0048] Where Loss represents the loss value. and Let represent the real and imaginary parts of the predicted speech signal output by the neural network model, respectively. and These represent the real and imaginary parts indicated by the labels of the training samples of the speech signal, respectively.
[0049] Step 104: Perform a short-time inverse Fourier transform on the complex spectrum of the near-end speech to obtain a clean near-end speech signal.
[0050] Finally, in this embodiment, the short-time inverse Fourier transform is performed on the complex spectrum of the near-end speech output by the model to reconstruct a clean near-end speech signal, which effectively suppresses echo components.
[0051] To better demonstrate the results of the echo cancellation scheme based on VBAP and deep learning technology provided in the foregoing embodiments of this application, a comparative verification experiment is also provided in one embodiment of this application, as follows:
[0052] Assuming the number of remote rooms in the practical application is 3, i.e., K=3, this also means that at most three speakers can be speaking simultaneously in multiple remote rooms within the same time period. The near-end rooms are equipped with 5.1 surround sound speakers; the subwoofer was not used in the experiment, hence P=5. The five speakers used in the experiment are at the same height, arranged on the same horizontal plane in a circular array, meaning the distance from each speaker to the center of the circle is the same. Furthermore, the database used is the DNS database, and the room impulse response is generated using the mirror method. The experimental data covers scenarios with one, two, and three speakers speaking simultaneously in the remote rooms. In addition, the experiment also compared echo denoising algorithms based on the Transformer model, testing metrics including Echo Return Loss Enhancement (ERLE) and Perceptual Evaluation of Speech Quality (PESQ).
[0053] The model performs well with a five-speaker layout in the near-end. To test the model for different speaker layouts in the near-end room, six speakers evenly distributed on a ring were selected, and a test condition with two speakers in the far-end room was chosen. This means the near-end room needs to reproduce two virtual sources, with the virtual sources positioned at 45° and 135°. For the six-speaker layout in the near-end room, the test results for this embodiment's model and the Transformer model, using the microphone signal as a reference, are shown in Figure 4. The test results provided in this embodiment correspond to different speaker layouts. Figure 4(a) is the speech spectrum of the microphone received signal (PESQ = 1.95), Figure 4(b) is the speech spectrum of the clean near-end speech signal, Figure 4(c) is the speech spectrum after processing based on the Transformer model (ERLE = 64.3dB, PESQ = 2.25), and Figure 4(d) is the speech spectrum after processing based on the model of this embodiment (ERLE = 63.5dB, PESQ = 2.54). Because for a six-speaker arrangement in the near-end room, the processing results in the figures show that both the model of this embodiment and the Transformer model have good suppression effects on echo components, but the Transformer model has more severe speech distortion. It is worth noting that when the number of far-end speakers is fixed at K = 3, the model has no limit on the number of speakers in the near-end room and can be applied to any number and orientation of speakers.
[0054] In some embodiments of this example, before using the remote microphone signal as a reference signal, the method further includes: obtaining the number of sound sources in all remote rooms; if the number of sound sources is greater than or equal to a preset number threshold (e.g., 3), then encoding the remote microphone signal into B-format.
[0055] In different practical application scenarios, the number of sound sources in the remote rooms will vary. For application scenarios with a small number of sound sources in the remote rooms, the echo cancellation scheme based on VBAP and deep learning technology in the previous embodiment can be referred to, using the single-channel microphone signals (i.e., remote microphone signals) of multiple remote rooms as reference signals. For application scenarios with a large number of sound sources in the remote rooms, this embodiment adopts an echo cancellation scheme based on Ambisonics and deep learning technology. Unlike traditional methods that usually use D-format signals as reference signals, this embodiment encodes the microphone signals of multiple channels in the remote rooms into B-format (i.e., first-order Ambisonics format). In this case, there is no upper limit on the number of remote rooms. Then, this format signal is used as a reference signal, and the compressed complex spectrum is extracted and input into the neural network model to achieve the purpose of echo cancellation.
[0056] It is worth mentioning that Ambisonic technology is a classic method in spatial sound field reproduction. It is based on the spherical harmonic analysis theory of plane waves and uses spherical harmonic functions to encode and reconstruct spatial sound fields. Depending on the order of the spherical harmonic function, Ambisonic technology can use different numbers and arrangements of loudspeakers to reproduce spatial sound fields, which has high flexibility and scalability.
[0057] Figure 5 shows a schematic diagram of an echo cancellation application scenario based on Ambisonic technology provided in this embodiment. The only difference between this embodiment and the schematic diagram of an echo cancellation application scenario based on VBAP technology shown in Figure 2 is that the remote microphone signal needs to be encoded into a B-format reference signal. The B-format reference signal has more channels than the ordinary reference signal. All other implementations of the two schemes can remain the same.
[0058] To better demonstrate the results of the echo cancellation scheme based on Ambisonic and deep learning technologies provided in the foregoing embodiments of this application, a comparative verification experiment is also provided in one embodiment of this application, as follows:
[0059] Assuming there are four speakers in the near-end room, the following three comparative experiments were set up. Experiment 1: The near-end room uses a fixed speaker layout, i.e., only one speaker type, and this model is labeled as D-format; Experiment 2: The speaker layout in the near-end room is not fixed. The decoding method is generated according to the actual layout during the near-end microphone signal generation process, but the reference signal only uses the W channel of the B-format signal. This model is labeled as Singlechn; Experiment 3: The speaker layout in the near-end room is not fixed. The decoding method is generated according to the actual layout during the near-end microphone signal generation process, and the reference signal only uses the B-format signal. This model is labeled as B-format.
[0060] To test the model's generalization ability for flexible speaker layouts, speaker layouts not present in the training sets of the three models were given, and experiments were conducted in real rooms. This experiment only included one far-end room and four speakers in the near-end room. Figure 6 shows the test results characterization diagrams of different algorithms used to process the recorded data under different speaker layouts provided in this embodiment. In Figure 6(a), the speech spectrum diagram of the microphone received signal (PESQ = 2.45) is shown; in Figure 6(b), the speech spectrum diagram of the clean near-end speech signal is shown; and in Figure 6(c), the speech spectrum diagram based on PB is shown. The speech spectrograms of the FDLMS algorithm (ERLE = 14.1dB, PESQ = 2.59), Figure 6(d) is the speech spectrogram of the D-format model based on this embodiment (ERLE = 17.7dB, PESQ = 2.76), Figure 6(e) is the speech spectrogram of the Singlechn model (ERLE = 52.6dB, PESQ = 2.85), and Figure 6(f) is the speech spectrogram of the B-format model (ERLE = 54.7dB, PESQ = 2.97).
[0061] The test results above show that if the signal of a single-layout speaker is used as a reference, the model's generalization ability for unknown speaker layouts is insufficient. Using one channel of D-format as a reference can also achieve a certain echo suppression effect, but its performance is not as good as the B-format model, which also demonstrates good model generalization ability.
[0062] In some embodiments of this example, the echo cancellation method further includes: determining the virtual sound source positions corresponding to multiple actual sound sources in the near room and the far room respectively; generating multiple voice reconstruction instructions based on multiple clean near-end speech signals corresponding to the multiple actual sound sources; and sending the voice reconstruction instructions to the speakers or speaker combinations in the near room corresponding to each virtual sound source position respectively, so as to instruct the speakers in different locations in the near room to play the corresponding clean near-end speech signals.
[0063] Specifically, in this embodiment, after recovering the clean near-end speech signal from the microphone received signal based on the neural network model, the corresponding loudspeaker / loudspeaker combination is determined according to the virtual sound source position in the near-end room. Then, the corresponding clean near-end speech signal is sent to loudspeakers in different directions, so that the speech signals of multiple speakers in the far-end room are reconstructed in the near-end room, and echo suppression of the virtual spatial sound system is also achieved. It is worth mentioning that multiple loudspeakers can be arranged in a ring in the near-end room. If the virtual source is located between two loudspeakers, the two loudspeakers on both sides of the virtual source are controlled to reconstruct the corresponding actual sound source. If the virtual source coincides with the position of a single loudspeaker, the single loudspeaker at the position of the virtual source is controlled to reconstruct the corresponding actual sound source.
[0064] It should be understood that the sequence number of each step in this embodiment does not imply the order in which the steps are executed. The execution order of each step should be determined by its function and internal logic, and should not constitute a unique limitation on the implementation process of this application embodiment.
[0065] Figure 7 illustrates a deep learning-based echo cancellation device according to an embodiment of this application. This echo cancellation device can be used to implement the echo cancellation method in the aforementioned embodiments, and mainly includes:
[0066] The signal acquisition module 701 is used to acquire the corresponding remote microphone signal in the remote room and the corresponding near microphone signal in the near room.
[0067] The complex spectrum acquisition module 702 is used to acquire the first compressed complex spectrum corresponding to the reference signal by taking the remote microphone signal as the reference signal, and to acquire the second compressed complex spectrum corresponding to the near microphone signal.
[0068] The model processing module 703 is used to input the first compressed complex spectrum and the second compressed complex spectrum into the trained neural network model for echo cancellation processing and output the near-end speech compressed complex spectrum.
[0069] The signal transformation module 704 is used to perform short-time inverse Fourier transform on the complex spectrum of near-end speech compression to obtain a clean near-end speech signal.
[0070] In one optional embodiment of this example, the echo cancellation device further includes: a signal encoding module, used to obtain the number of sound sources in all remote rooms before using the remote microphone signal as a reference signal; if the number of sound sources is greater than or equal to a preset number threshold, the remote microphone signal is encoded into B-format.
[0071] In one optional embodiment of this invention, the echo cancellation device further includes: a voice reconstruction module, used to determine the virtual sound source positions corresponding to multiple actual sound sources in the near room and the far room respectively; generate multiple voice reconstruction instructions based on multiple clean near-end voice signals corresponding to multiple actual sound sources; and send the voice reconstruction instructions to the speakers or speaker combinations in the near room corresponding to each virtual sound source position respectively, so as to instruct the speakers in different locations in the near room to play the corresponding clean near-end voice signals.
[0072] It should be noted that the echo cancellation methods in the foregoing method embodiments can all be implemented based on the echo cancellation device provided in this embodiment. Those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the echo cancellation device described in this embodiment can be implemented by referring to the corresponding working process in the foregoing method embodiments, and will not be repeated here.
[0073] Based on the technical solution of the above embodiments of this application, the corresponding far-end microphone signal of the far-end room and the corresponding near-end microphone signal of the near-end room are acquired; the far-end microphone signal is used as a reference signal to acquire the corresponding first compressed complex spectrum and the corresponding second compressed complex spectrum of the near-end microphone signal; the first and second compressed complex spectra are input into the trained neural network model for echo cancellation processing, and the near-end speech compressed complex spectrum is output; the near-end speech compressed complex spectrum is subjected to a short-time inverse Fourier transform to obtain a clean near-end speech signal. Through the implementation of this application, by using the microphone signals of each far-end room as neural network reference information to recover the clean near-end speech signal, echo components can be effectively suppressed, thereby realizing echo suppression of the virtual space sound system.
[0074] Please refer to Figure 8, which illustrates an electronic device according to an embodiment of this application. This electronic device can be used to implement the deep learning-based echo cancellation method described in the foregoing embodiments. As shown in Figure 8, the electronic device mainly includes:
[0075] The system includes a memory 801, a processor 802, and a bus 803, with the memory 801 and processor 802 connected via the bus 803. The memory 801 stores a computer program that can run on the processor 802. When the processor 802 executes the computer program, it implements the deep learning-based echo cancellation method described in the preceding embodiments. The number of processors can be one or more.
[0076] The memory 801 can be a high-speed random access memory (RAM) or a non-volatile memory, such as a disk storage device. The memory 801 is used to store executable program code, and the processor 802 is coupled to the memory 801.
[0077] Furthermore, embodiments of this application also provide a computer-readable storage medium, which may be disposed in the electronic device in the above embodiments, and the computer-readable storage medium may be the memory in the embodiment shown in FIG8 above.
[0078] The computer-readable storage medium stores a computer program that, when executed by a processor, implements the deep learning-based echo cancellation method described in the foregoing embodiments. Furthermore, the computer-readable storage medium can also be a USB flash drive, external hard drive, read-only memory (ROM), RAM, magnetic disk, or optical disk, or any other medium capable of storing program code.
[0079] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0080] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0081] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0082] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned readable storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0083] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0084] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0085] The above is a description of the echo cancellation method, device and readable storage medium based on deep learning provided in this application. For those skilled in the art, based on the ideas of the embodiments of this application, there will be changes in the specific implementation and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A deep learning based echo cancellation method, characterized by, The method comprises the following steps: acquiring a far-end microphone signal corresponding to a far-end room and acquiring a near-end microphone signal corresponding to a near-end room; taking the far-end microphone signal as a reference signal, acquiring a first compressed complex spectrum corresponding to the reference signal, and acquiring a second compressed complex spectrum corresponding to the near-end microphone signal; inputting the first compressed complex spectrum and the second compressed complex spectrum into a trained neural network model for echo cancellation processing, and outputting a near-end speech compressed complex spectrum; performing inverse short-time Fourier transform on the near-end speech compressed complex spectrum to obtain a pure near-end speech signal.
2. The echo cancellation method of claim 1, wherein, Before the step of taking the far-end microphone signal as a reference signal, the method further comprises the following steps: acquiring a number of sound sources of all the far-end rooms; if the number of sound sources is greater than or equal to a preset number threshold, encoding the far-end microphone signal into a B-format.
3. The echo cancellation method of claim 1, wherein, The step of acquiring a near-end microphone signal corresponding to a near-end room comprises the following steps: acquire a near-end room kth loudspeaker corresponding near-end loudspeaker signal; the near-end loudspeaker signal is expressed as: where K represents the number of remote rooms, δ p,k represents the pth loudspeaker for sound source signal gain of the kth remote room, a far-end loudspeaker signal of a channel of a kth far-end room; According to the near-end loudspeaker signal, a corresponding near-end microphone signal is obtained; the near-end microphone signal is represented as: where P is the number of loudspeakers of the proximal in-room array arrangement, h p (n) represents the echo path, s(n) represents the proximal speech signal, and v(n) represents additive noise.
4. The echo cancellation method of claim 1, wherein, The first compressed complex spectrum is represented as: wherein representing said first compressed complex spectrum, and θ respectively represent amplitude information and phase information of a far-end microphone signal corresponding to the kth far-end room, and β is a preset constant.
5. The echo cancellation method of claim 1, wherein, The neural network model comprises a sequentially connected encoding network, a time series modeling network, a decoding network, and a linear layer, and the encoding network is further connected to the decoding network in a skip connection manner.
6. The echo cancellation method of claim 5, wherein, The decoding network comprises two decoding network branches, and the two decoding network branches are connected with a linear layer respectively; and a loss function of the neural network model is represented as: wherein Loss represents a loss value, and respectively represent real and imaginary parts of a predicted speech signal output by the neural network model, and respectively represent real and imaginary parts indicated by a label of a speech signal training sample.
7. The echo cancellation method according to any one of claims 1 to 6, characterized by, The method further comprises the following steps: determining virtual sound source positions corresponding to a plurality of actual sound sources of the far-end rooms in the near-end room; generating a plurality of speech reconstruction instructions based on a plurality of pure near-end speech signals corresponding to the plurality of actual sound sources; respectively sending the speech reconstruction instructions to loudspeakers or loudspeaker combinations corresponding to the virtual sound source positions in the near-end room, so that the loudspeakers in different directions of the near-end room play the corresponding pure near-end speech signals.
8. A deep learning based echo cancellation apparatus, characterized by, The method comprises the following steps: a signal acquisition module, configured to acquire a far-end microphone signal corresponding to a far-end room and acquire a near-end microphone signal corresponding to a near-end room; a complex spectrum acquisition module, configured to take the far-end microphone signal as a reference signal, acquire a first compressed complex spectrum corresponding to the reference signal, and acquire a second compressed complex spectrum corresponding to the near-end microphone signal; a model processing module, configured to input the first compressed complex spectrum and the second compressed complex spectrum into a trained neural network model for echo cancellation processing, and output a near-end speech compressed complex spectrum; a signal transformation module, configured to perform inverse short-time Fourier transform on the near-end speech compressed complex spectrum to obtain a pure near-end speech signal.
9. An electronic device, comprising: The method comprises the following steps: a memory and a processor; the processor is configured to execute a computer program stored in the memory; when the processor executes the computer program, the steps in the deep learning-based echo cancellation method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by the processor, the steps in the deep learning-based echo cancellation method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Audio space rendering device and method
CN104010265A
Method and device for eliminating echo signal, computing equipment and storage medium
CN113763977A
Model training method, echo cancellation method, system and device and storage medium
CN114530160A
Method for training echo cancellation model, echo cancellation method and corresponding device
CN115762552A
Echo cancellation method and device, audio equipment and storage medium
CN117727317A