Electronic device and method for low latency speech enhancement using autoregressive adjustment-based neural network model
By adopting an autoregressive adjustment neural network model in speech enhancement technology, using teacher forced mode and model prediction to replace the real waveform, the problem of high algorithm delay in the existing technology is solved, and the low-latency voice enhancement effect is achieved.
Patent Information
- Application Number
- CN202380070118.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-10
- Filing Date
- 2023-10-10
- Publication Date
- 2025-05-13
AI Technical Summary
The existing voice enhancement technology has high algorithm delays, which is difficult to meet the needs of low latency (less than 10ms) applications.
The neural network model based on autoregressive regulation is adopted to reduce the algorithm delay by training the model in the teacher forced mode and replacing the real shift waveform in the autoregressive channel with the prediction of the model in subsequent training iterations.
Low-latency voice enhancement is achieved, and the algorithm delay is reduced to only 2ms, significantly improving the voice enhancement quality.
Smart Images

Figure CN119998876A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computing, and in particular to methods for processing and analyzing audio recordings. Background Art
[0002] The problem of real-time streaming (“live”) speech processing is of great practical importance for modern digital hearing aids, acoustically transparent hearing devices, and telecommunications. The undetectable lag limit for live, real-time processing is a subject of investigation and debate, but is estimated to be around 5–30 ms, depending on the application. Given that speech enhancement tools are typically deployed in a joint pipeline with other speech processing tools (e.g., echo cancellation) and within a signal transmission channel, the overall latency requirement is very stringent and is barely met for many applications by mainstream speech enhancement solutions, which typically rely on algorithmic (by model design) latency exceeding 30–60 ms. To address this problem, Defossez et al. (“Real time speech enhancement in the waveform domain”) [4] proposed a convolutional architecture with long short-term memory (LSTM) layers for real-time stream processing. However, this architecture still suffers from algorithmic latency exceeding 15 ms. Therefore, there is a significant need for research dedicated to low-latency (less than 10 ms) speech enhancement models.
[0003] Low-latency speech enhancement has recently attracted a lot of attention from the research community. Time-domain causal neural architectures have been explored for this task, as spectral domain approaches tend to be limited by the window size of the short-time Fourier transform, which is typically chosen to be longer than 20-30ms. Recent work argues that time-frequency domain architectures can also be exploited by using asymmetric analysis-synthesis pairs on the windowing of the direct short-time Fourier transform and its inverse transform. Summary of the invention
[0004] Technical Solution
[0005] An electronic device and method for low-latency speech enhancement using a neural network model based on autoregressive regulation are provided.
[0006] According to one aspect of the present disclosure, a method for training and operating a neural network model comprises: in an initial training iteration, training the neural network model in a teacher forcing mode and outputting a prediction of the neural network model, wherein in the teacher forcing mode, an autoregressive channel comprises a true shifted waveform; and in at least one additional training iteration, replacing the true shifted waveform in the autoregressive channel with a prediction of the neural network model obtained in a previous training iteration.
[0007] According to one aspect of the present disclosure, an electronic device includes: at least one memory storing at least one instruction; and at least one processor configured to execute the at least one instruction to perform the following operations: in an initial training iteration, training a neural network model in a teacher forcing mode and outputting a prediction of the neural network model, wherein in the teacher forcing mode, an autoregressive channel includes a true shifted waveform; and in at least one additional training iteration, replacing the true shifted waveform in the autoregressive channel with the prediction of the neural network model obtained in a previous training iteration.
[0008] According to one aspect of the present disclosure, a non-transitory computer-readable medium stores stored instructions, which, when executed by at least one processor, cause the at least one processor to perform the following operations: in an initial training iteration, train a neural network model in a teacher forcing mode and output a prediction of the neural network model, wherein in the teacher forcing mode, an autoregressive channel includes a true shifted waveform; and in at least one additional training iteration, replace the true shifted waveform in the autoregressive channel with the prediction of the neural network model obtained in a previous training iteration. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features and advantages of specific embodiments of the present disclosure are explained in the following description in conjunction with the accompanying drawings, in which:
[0010] Figure 1 is a block diagram illustrating an example architecture of a neural network model according to an embodiment of the present disclosure;
[0011] Figure 2 shows the adjustment and training of a neural network model according to an embodiment of the present disclosure; and
[0012] Figure 3 is a data graph showing the reduction in error rate of the enhanced model during iterative training according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0013] The following detailed description of example embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements. The embodiments are described below in order to explain the disclosed systems and methods with reference to the drawings of specific example embodiments for sample applications illustratively shown in the drawings.
[0014] The foregoing disclosure provides examples and descriptions, but is not intended to be exhaustive or to limit implementations to the disclosed precise forms. According to the above disclosure, modifications and variations are possible, or modifications and variations may be obtained from the practice of the embodiments. In addition, one or more features or components of an embodiment may be incorporated into another embodiment (or one or more features of another embodiment) or combined with another embodiment (or one or more features of another embodiment). In addition, in the description of the flow charts and operations provided below, it is understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be switched.
[0015] It will be apparent that the systems and / or methods described herein may be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code - it is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0016] Even though particular combinations of features are recited in the claims and / or disclosed herein, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed herein. Although each dependent claim listed below may be directly dependent on only one claim, the disclosure of possible implementations includes the combination of each dependent claim with every other claim in the claim set.
[0017] Unless explicitly described as such, the elements, actions or instructions used herein should not be interpreted as critical or essential. In addition, as used herein, the article "one" is intended to include one or more items, and can be used interchangeably with "one or more". In the case of only one item, the term "one" or similar language is used. In addition, as used herein, the terms "having", "including" etc. are intended to be open terms. In addition, unless otherwise explicitly stated, the phrase "based on" is intended to mean "based at least in part". In addition, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B or both A and B.
[0018] Streaming speech processing can be performed by processing discrete "blocks" of waveform samples in a sequential manner. The block size and total future context used for its processing determine the algorithmic delay, where the algorithmic delay can be defined as the total delay due to algorithmic reasons. The algorithmic delay can also be defined as the maximum duration of future context required to produce each time step of the processed waveform. In contrast, hardware delay can be defined as the delay imposed by the duration calculated by the hardware. The total delay is the sum of the algorithmic delay and the hardware delay. The algorithmic delay imposes a major constraint on the total delay, while the hardware delay can be reduced by manipulating the model size and hardware efficiency. The present disclosure mainly discusses improvements to the algorithmic delay.
[0019] Autoregressive models are a form of generative model used in a variety of applications, including but not limited to language modeling, text-to-speech translation, and image generation. For example, autoregressive models applied to conditional waveform generation are used in neural sound coding. An example is a fully convolutional autoregressive model, which uses causal dilated convolutions to model waveform sequences, producing highly realistic speech samples conditioned on language features. Dilated convolutions help increase the receptive field of the model, while causality enables samples to be generated in a sequential (autoregressive) manner. In one or more embodiments of the present disclosure, causal convolutions are similarly used for autoregressive conditioned generation, but using very different types of conditional information (degraded waveforms), and using waveform samples that are generated in blocks rather than one by one.
[0020] The CARGAN model combines autoregressive conditioning with the power of generative adversarial networks to mitigate artifacts during spectrogram inversion. In one or more embodiments of the present disclosure, autoregressive conditioning is similarly combined with adversarial training, but with a different task in mind and with a much smaller block size (<10ms) than the 92ms used in the CARGAN model.
[0021] Although teacher forcing was originally proposed for training recurrent neural networks, teacher forcing is a training process for autoregressive models. The method provides the model with previous real samples during training and then learns to predict the next sample. During the inference phase, the model uses its own samples for autoregressive conditioning (free-running mode) because the true value (ground-truth) is not available. The inventors have found that the use of real samples (teacher forcing) greatly improves the quality of speech enhancement in the training mechanism (see row 300GT of Table 1 below). However, due to the training-inference mismatch, the model trained with teacher forcing shows unsatisfactory results during inference (see row 300Honest of Table 1 below). One of the most characteristic artifacts that the inventors have observed is the silent area that appears in the predicted waveform. The model seems to rely heavily on true conditioning to detect speech areas and silent areas.
[0022]
Table 1
[0023] Results. The model with our proposed autoregressive adjustment is denoted as AR.
[0024]
[0025]
[0026] In short, the exemplary embodiments of the present disclosure provide a method and system in which a general algorithm enables efficient training of autoregressive speech enhancement models for low-delay applications. When implemented by a computer, these embodiments can be significantly improved over non-autoregressive baselines under different training losses and neural architectures. Compared with the related art previously discussed, the embodiments of the present disclosure consider domain-agnostic techniques for improving low-delay speech enhancement models, which can potentially be used with any low-delay causal neural architecture. The present disclosure demonstrates that such embodiments provide considerable improvements in particular for time-domain models, although the method is not limited to a specific domain. Although streaming low-delay models may be constrained by limited future context, the sequential nature of the generation process provides them with the benefits of autoregressive regulation. Since such models process waveforms in a block-by-block manner, they can use their own predictions for previous blocks when predicting the current block. This information can then be used to more accurately model clean waveforms and noise suppression. For example, given a denoised waveform from a previous time step, the model is more likely to understand the characteristics of the noise and the speaker's voice. Indeed, although in real life the true waveform cannot be achieved perfectly in any form, the predictions of the model can be used as a proxy for the true waveform, and thus an autoregressive model can be formed.
[0027] Under ideal conditions, when the model is conditioned on the true waveform from the previous time step, it can provide excellent results, outperforming its non-autoregressive counterpart by a considerable margin. For example, embodiments of the present disclosure advantageously achieved an algorithm delay of only 2 ms in testing, compared to an algorithm delay of over 15 ms in the related art that does not include autoregressive conditioning information.
[0028] Typically, low-latency speech enhancement models consist of causal neural layers (e.g., unidirectional LSTM, causal convolution, causal attention layers, etc.) operating in the time domain or frequency domain. Time domain architectures also tend to include strided convolution layers and downsampling / upsampling to facilitate context aggregation. In one or more embodiments of the present disclosure, these architectures can be modified to implement autoregressive conditioning. For example, additional input features containing information for autoregressive conditioning can be cascaded. In particular, for a time domain architecture whose first layer is typically a one-dimensional convolution, in addition to a channel containing a noisy waveform, a channel containing a waveform with past predictions can also be included (see Figure 2 ).
[0029] For most experiments conducted by the inventors, a simple time-domain architecture that can be referred to as WaveUNet+LSTM was used. The WaveUNet+LSTM model is a fully convolutional neural network that is enhanced with a long short-term memory (LSTM) layer at the bottleneck.
[0030] Figure 1 is a block diagram showing an example of a neural network model with a WaveUNet+LSTM architecture according to an embodiment of the present invention.
[0031] The architecture is based on a convolutional encoder-decoder UNet architecture, with downsampling layers receiving inputs (left column) and upsampling layers providing outputs (right column), and is enhanced with unidirectional LSTM layers at the bottleneck to enable the use of large receptive fields for past time steps. The UNet structure shown uses strided convolutional downsampling layers with kernel size 2 and stride 2, and nearest neighbor upsampling, but other parameters are also within the scope of the present disclosure. The parameter K adjusts the overall depth of the UNet structure, the parameter N determines the number of residual blocks within each layer, and the array C determines the number of channels at each level of the UNet structure. The algorithmic delay of the neural network is adjusted by the number of downsampling / upsampling layers K, and is equal to 2K. Note that the architecture shown is not limiting, and other suitable architectures and suitable modifications of the architecture shown may also be used.
[0032] As previously described, teacher forcing is a very convenient way to train autoregressive models in terms of training speed. When training in a teacher forcing mechanism, time-consuming sequential reasoning (free running mode) is not required. This is especially important for convolutional autoregressive models that can be effectively parallelized during the training phase. Without this parallelization, it would be difficult to train such a model in a meaningful time. For example, even when using an effective implementation with an activated cache, the duration of autoregressive reasoning for a two-second audio clip in free running mode by a WaveNet model can be up to 1000 times the duration of teacher-forced reasoning (forward pass in the training phase) for the same clip. Using the WaveUNet+LSTM architecture shown, this factor can be reduced to 75, but the result is still undesirable. However, the shorter duration of teacher forcing is offset by the training-inference mismatch, which can lead to significant quality degradation, as observed in Table 1. In the related art, the method of alleviating this mismatch explicitly relies on the possibility of performing autoregressive reasoning in free running mode during training. As mentioned above, the forward pass in free-running mode takes several orders of magnitude longer than teacher forcing, which complicates the use of this technique in practice and loses the advantage of faster processing.
[0033] Embodiments of the present disclosure provide an alternative way to reduce the gap between training and inference that does not require a time-consuming free-running mode during training. Embodiments of the present disclosure iteratively replace autoregressive conditioning with predictions of the model in a teacher-forcing mode.
[0034] Figure 2 The adjustment and training of a neural network model according to an embodiment of the present disclosure is shown. The model shown has an algorithm delay of 32 time steps (2ms at a 16kHz sampling rate), but the present disclosure is not limited thereto.
[0035] During conditioning, the predicted time steps from block 1 can be reused when making predictions for block 2. Then, during training, the model can use its own predictions to come up with higher-order predictions. The true waveform and predictions can be shifted before forming the channel with autoregressive conditioning to avoid leakage of future information.
[0036] In the initial stage of training, the model can be trained in the standard teacher forcing mode, where the autoregressive channel ( Figure 2 The top row of ) contains the real waveform (such as Figure 2 ). In the next stage, the true waveform in the autoregressive channel may be replaced by the prediction of the model obtained in teacher forcing mode (using true values in autoregressive conditioning). In each subsequent stage or iteration of training, the autoregressive input channel may contain the prediction of the model obtained in the previous stage. In general, during this training process, the model can be conditioned on its own predictions. As training proceeds, the order of the predictions of the model to be conditioned on can be gradually increased, for example, the number of forward passes performed before calculating the loss and performing the backward pass can be increased. Note that in embodiments of the present disclosure, gradients can be propagated through the last forward pass without also propagating gradients in previous forward passes.
[0037] An embodiment of the forward pass, and more specifically, the forward function of the model, will be further described in a standard training pipeline including a forward pass, computing loss, backpropagation, and weight optimization. A modified iterative forward function according to an embodiment of the present disclosure is summarized below and is presented in Figure 2 The lower part is shown schematically.
[0038] According to one or more embodiments described above, a method may include a model training phase and an interference phase.
[0039] The model training phase may iteratively replace the autoregressive adjustment with the prediction of the model in teacher-forcing mode. In training initialization, the model may be trained in standard teacher-forcing mode, wherein, in standard teacher-forcing mode, the autoregressive channel contains a true shifted waveform. At the end of training initialization (which may also be referred to as "iteration 0" or "initial training iteration"), the true shifted waveform may be used as the autoregressive channel to generate the output of the model. Then, in all subsequent iterations of multiple training iterations, the shifted waveform output by the previous training iteration may be used as the autoregressive channel to generate further outputs of the model. The output of the final training iteration may be used for backpropagation without also using the output of the previous iteration; for example, in some embodiments, only the output of the final iteration is backpropagated, without backpropagating the output of the previous iteration.
[0040] The inference phase may provide an additional channel containing past predictions (i.e., predictions output during the training phase). The inference phase may then use the obtained model to perform speech enhancement.
[0041] A sample pseudocode algorithm for autoregressive training forward functions is provided. The algorithm uses:
[0042] Noise frequency x;
[0043] Clean Audio y;
[0044] Model m, which takes x and y as input;
[0045] a plan {E,N} consisting of a list of integers E and a list of integers N (e.g., starting at round E[i] and doing N[i] iterations); and
[0046] An integer e which is the current round number.
[0047] algorithm:
[0048] start
[0049] i←0
[0050] While e>E[i]do
[0051] i←i+1 Note: Search from E to the left for the round closest to e
[0052] end while
[0053] for i=0to N[i]do with NOGRAD
[0054] y←m(x,y) Note: Adjust its own prediction
[0055] end for
[0056] y = m(x, y) Note: Make final prediction
[0057] returny
[0058] In a series of experiments, the inventors used a batch size of 16, an Adam optimizer with a learning rate of 0.0002 and a decay of 0.999, and betas of 0.8 and 0.9. The iterative autoregressive run was trained for 1000 rounds and the non-autoregressive run was trained for 2000 rounds, with each epoch consisting of 1000 batch iterations. The best rounds were selected based on the validation results of the UTMOS loss metric, as this metric showed the closest correlation to MOS (mean opinion score). UTMOS (UTokyo-SaruLab mean opinion score) is a state-of-the-art objective speech quality metric described in the following paper: Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari, "Utmos: Utokyo-sarulab system for voicemos challenge 2022," arXiv preprint arXiv:2022, 2204.02152.
[0059] When testing the autoregressive run, the following plan is used: a single iteration is performed for the first 300 rounds, and then additional iterations are added every 100 rounds. The detailed configurations of the WaveUNet+LSTM and ConvTasNet architectures are as follows. For the main configuration of WaveUNet, the number of levels (K) within the UNet hierarchy is fixed, and the number of residual blocks at each level and the number of channels within the residual block are 4, 7 and 16, 24, 32, 48, 64, 96, 128, respectively. The LSTM width is equal to 512. This configuration corresponds to an algorithm delay of 8ms. For ConvTasNet, the original architecture implementation is adjusted to match the algorithm delay to 8ms and the number of multiplication-accumulation operations per second to 2 billion. These configurations are incorporated herein by reference. For most experiments, the inventors used a voice cloning toolkit (VCTK) dataset with standard training verification segmentation.
[0060] Several series of experiments were conducted to test the idea of iterative autoregression:
[0061] 1. “Motivational experiments” that revealed the positive effects of teacher-forcing questions and iterated autoregression.
[0062] 2. Loss variation experiments, which measure other loss metrics such as SI-SNR and adversarial loss with L1Spec loss.
[0063] 3. Test iterative autoregression with a more challenging DNS dataset.
[0064] 4. Examining iterative autoregression using the ConvTasNet architecture.
[0065] 5. Experiments with different delays to test the generalizability of the results of embodiments of the disclosed method.
[0066] The conducted experiments consistently revealed significant improvements of the tested embodiments over the baseline, thus demonstrating high practical value and general applicability.
[0067] Figure 3 is a data graph showing the reduction in error rate of the enhanced model during iterative training according to an embodiment of the present disclosure. Figure 3 In FIG. 1 , the dependence of the difference between the training mode and the test mode for the same audio data (average of 100 audio inputs) with increasing number of iterations is shown. As described above, in the illustrated experiment, additional iterations are added starting at round 300. It can be seen that when training using iterative autoregression, the output of the training mode becomes close to the output of the test mode, which enables solving the training-inference mismatch and improving the quality without losing the speed of teacher enforcement.
[0068] One or more embodiments disclosed herein may be used in various devices that transmit, receive, and record audio to improve the user experience of listening to audio (e.g., speech) recordings. For example, example embodiments may be used to denoise speech recorded in a noisy environment. Example embodiments may also be employed in various devices that support floating point or fixed point calculations. Due to the strong preference for low algorithmic latency in digital hearing aid devices, embodiments may be of particular interest to such devices.
[0069] The embodiments of the present disclosure may be executed and / or implemented on any electronic device including computing tools, audio playback components, and memory (RAM, ROM, etc.). Some non-limiting examples of such devices include smart phones, tablet computers, headphones, speakers, navigation systems, vehicle-mounted equipment, notebooks, smart watches, etc. The computing device may include, but is not limited to, a central processing unit (CPU), an audio processing unit, a processor, a neural processing unit (NPU), a graphics processing unit (GPU). In a non-limiting embodiment, the computing device may be implemented as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a system on a chip (SoC). The electronic device may also include, but is not limited to, a (touch) screen, an I / O tool, a camera, a communication tool, a speaker, a microphone, etc.
[0070] Embodiments of the present disclosure may also be implemented as a non-transitory computer-readable medium having computer-executable instructions stored thereon, which, when executed by a processor of the device, cause the device to perform any steps and / or operations of the embodiments. Any type of data may be processed, stored, and transmitted by an intelligent system trained using the above method. The learning phase may be performed online or offline. The trained neural network may be transmitted to a user device, for example, in the form of weights and other parameters and / or computer-executable instructions, and the trained neural network may be stored on the user device to be used in the inference (in use) phase.
[0071] At least one of the multiple modules may be implemented by an AI model. Functions associated with AI may be performed by non-volatile memory, volatile memory, and a processor. The processor may include one or more processors. Each such processor may include, but is not limited to, a general-purpose processor (such as a central processing unit (CPU), an application processor (AP), etc.), a graphics processing unit (such as a graphics processing unit (GPU)), a visual processing unit (VPU)), and / or an AI-specific processor (such as a neural processing unit (NPU).
[0072] One or more processors may control the processing of input data according to predefined operating rules or artificial intelligence (AI) models stored in non-volatile memory and volatile memory. Predefined operating rules or artificial intelligence models are provided by training or learning. Here, providing by training or learning means making predefined operating rules or AI models with desired characteristics by applying a learning algorithm to multiple learning data. Learning can be performed in the device itself that executes AI according to the embodiment, and / or learning can be achieved by a separate server / system.
[0073] The AI model can be composed of multiple neural network layers. Each layer has multiple weight values, and the layer operation is performed by the calculation of the previous layer and the operation of multiple weights. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks, etc.
[0074] The learning algorithm is a method for training a predetermined target device using a plurality of learning data to enable, allow or control the target device to perform low-latency speech enhancement, determination or prediction. Examples of learning algorithms include but are not limited to supervised learning, unsupervised learning, semi-supervised learning or reinforcement learning, etc.
Claims
1. A method for training and operating a neural network model, the method comprising performing the following operations by at least one processor of an electronic device: In an initial training iteration, the neural network model is trained in a teacher-forcing mode, and predictions of the neural network model are output, wherein, In the teacher-forcing mode, the autoregressive channel includes a true shifted waveform; In at least one additional training iteration, the true shifted waveform in the autoregressive channel is replaced with the prediction of the neural network model obtained in a previous training iteration.
2. The method according to claim 1, wherein: The at least one additional training iteration comprises a plurality of training iterations, each training iteration outputting a corresponding prediction of the neural network model to the autoregressive channel for a next iteration of the plurality of training iterations.
3. The method according to claim 2, wherein: The neural network model is configured to perform at least one forward pass, compute a loss and perform at least one backward pass, and wherein during training, the number of forward passes performed before calculating the loss and performing the at least one backward pass gradually increases.
4. The method according to claim 2, wherein: Only the output of the final iteration of the plurality of training iterations is back-propagated.
5. The method of claim 1 , further comprising performing inference by the neural network model by: An additional channel is provided for the neural network model, wherein: The additional channel contains at least one prediction output by the neural network model during training; and Speech enhancement is performed using the neural network model.
6. The method according to claim 5, wherein: The neural network model includes a fully convolutional neural network.
7. The method according to claim 6, wherein: The fully convolutional neural network includes a WaveUNet architecture enhanced with a long short-term memory (LSTM) layer at its bottleneck.
8. An electronic device comprising: at least one memory storing at least one instruction; as well as At least one processor is configured to execute the at least one instruction to perform the following operations: In an initial training iteration, a neural network model is trained in a teacher forcing mode, and predictions of the neural network model are output, wherein in the teacher forcing mode, the autoregressive channel comprises a true shifted waveform, and In at least one additional training iteration, the true shifted waveform in the autoregressive channel is replaced with the prediction of the neural network model obtained in a previous training iteration.
9. The electronic device according to claim 8, wherein: The at least one additional training iteration comprises a plurality of training iterations, each training iteration outputting a corresponding prediction of the neural network model to the autoregressive channel for a next iteration of the plurality of training iterations.
10. The electronic device according to claim 9, wherein: The neural network model is configured to perform at least one forward pass, compute a loss and perform at least one backward pass, and wherein during training, the number of forward passes performed before calculating the loss and performing the at least one backward pass gradually increases.
11. The electronic device according to claim 9, wherein: Only the output of the final iteration of the plurality of training iterations is back-propagated.
12. The electronic device according to claim 8, wherein: The at least one processor is further configured to perform reasoning by: providing an additional channel to the neural network model, wherein the additional channel contains at least one prediction output by the neural network model during training, and Speech enhancement is performed using the neural network model.
13. The electronic device according to claim 12, wherein: The neural network model includes a fully convolutional neural network.
14. The electronic device according to claim 13, wherein: The fully convolutional neural network includes a WaveUNet architecture enhanced with a long short-term memory (LSTM) layer at its bottleneck.
15. A non-transitory computer-readable medium, wherein: The non-transitory computer-readable medium has instructions stored thereon, which when executed by at least one processor cause the at least one processor to perform the following operations: In an initial training iteration, training a neural network model in a teacher forcing mode, and outputting predictions of the neural network model, wherein in the teacher forcing mode, the autoregressive channel comprises a true shifted waveform; and In at least one additional training iteration, the true shifted waveform in the autoregressive channel is replaced with the prediction of the neural network model obtained in a previous training iteration.