Speech signal repair and enhancement using an integrated network based on progressive learning
The integrated repairer enhancer network (IREN) addresses voice communication degradations with reduced latency and resource usage by progressively training a single network to handle multiple degradations, enhancing speech quality in real-world systems.
Patent Information
- Application Number
- PCT/US2024/043991
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2024-08-27
- Publication Date
- 2025-12-04
AI Technical Summary
Conventional approaches to address voice communication degradations, such as echoes, noise, and artifacts, often introduce excessive latency and resource constraints, making them impractical for commercial applications.
An integrated repairer enhancer network (IREN) is trained progressively to simultaneously address multiple types of degradations using transfer learning, combining repair and enhancement networks into a single network to minimize resource usage.
The IREN effectively reduces latency and resource consumption while improving speech quality by integrating degradation repair and enhancement features, suitable for real-world communication systems.
Smart Images

Figure US2024043991_04122025_PF_FP_ABST
Abstract
Description
SPEECH SIGNAL REPAIR AND ENHANCEMENT USING AN INTEGRATED NETWORK BASED ON PROGRESSIVE LEARNINGCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of and priority to Indian Provisional Application No. 202441042495, filed on May 31, 2024, and titled “SPEECH SIGNAL REPAIR AND ENHANCEMENT USING AN INTEGRATED NETWORK BASED ON PROGRESSIVE LEARNING,” the content of which is herein incorporated by reference in its entirety for all purposes.BACKGROUND
[0002] Voice communications represent a critical means through which people connect with one another. As voice communications traverse communication channels, they suffer several types of degradations. For example, the degradations in voice communications can result from echoes due to microphone and speaker acoustic coupling, interfering noises due to the environment of the participants, artifacts introduced by codecs due to compression and processing, packet loss due to network deficiencies, play-out circuit processing and related hardware non-linearities, etc. Conventional approaches exist for addressing these and other types of degradations. However, such conventional approaches tend to add too much latency and / or other undesirable detriments that can render such approaches impractical or undesirable, particularly for applications that seek to concurrently address multiple of these degradations in a communication channel.SUMMARY
[0003] Systems and methods are described herein for speech signal repair and enhancement using an integrated network based on progressive learning. For example, a degraded speech signal is received on a speech channel and processed through an integrated repairer enhancer network (IREN) to generate a clean speech signal. Embodiments initially train a repairer network to ameliorate one or more of a first type of degradations (disrepair-related degradations).Embodiments then use transfer learning from the repairer network to train the IREN to ameliorate one or more of a second type of degradations (de-enhancement-related degradations). The resulting IREN is a single integrated machine learning network that ameliorates both types of degradations.BRIEF DESCRIPTION OF THE DRAWINGS
[0004] A further understanding of the nature and advantages of various embodiments may be realized by reference to the following figures. In the appended figures, similar components or features may have the same reference label. Further, various components of the same type may bedistinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.
[0005] FIG. 1 shows a representation of a conventional approach to addressing degradations in a voice communication channel.
[0006] FIG. 2 shows a high-level representation of novel approaches described herein.
[0007] FIGS. 3A and 3B show block diagrams of source model training systems, according to embodiments described herein.
[0008] FIG. 4 shows an illustrative integrated repairer enhancer network (IREN) training system that uses the repairing network to generate an IREN with integrated bandwidth expansion features.
[0009] FIGS. 5A and 5B show illustrative IREN training systems that use the repairing network to generate an IREN with integrated comfort enhancement features.
[0010] FIGS. 6 A - 6C show three illustrative training phases for an illustrative IREN training system that integrates burst loss prediction and repair into the IREN.
[0011] FIG. 7 shows an illustrative implementation environment for a progressively trained IREN, according to embodiments described herein.
[0012] FIG. 8 shows a conceptual circuit block diagram of a partial audio management system environment of a wearable audio component (WAC) with an automated attention handling system (AHS) having an integrated speech enhancer.
[0013] FIG. 9 provides a schematic illustration of an illustrative computational system that can implement various system components and / or perform various steps of methods provided by various embodiments relating to training of an IREN, as described herein.
[0014] FIG. 10 provides a schematic illustration of another illustrative computational system that can implement various system components and / or perform various steps of methods provided by various embodiments relating to operation of an IREN, as described herein.
[0015] FIG. 11 shows a flow diagram of an illustrative method for automated speech repair and enhancement for speech audio signals, according to various embodiments.
[0016] FIG. 12 shows a flow diagram of an illustrative implementation of progressive training of an IREN at stage of the method of FIG. 11.DETAILED DESCRIPTION
[0017] As voice communications traverse communication channels, they tend to suffer several types of degradations. For example, the degradations in voice communications can result from echoes due to microphone and speaker acoustic coupling, interfering noises due to the environment of the participants, artifacts introduced by codecs due to compression and processing, packet loss due to network deficiencies, play-out circuit processing and related hardware nonlinearities, etc. Conventional approaches typically address each degradation using a separate technique. Each technique is often based on tailoring digital signal processing to operate within computational budget limitations. More recently, attempts have been made to incorporate machine learning (ML) models into these techniques by tailoring models to address each type of degradation. However, introduction of these ML models has tended to appreciably increase latency, computation, and memory. Consequently, such attempts often do not operate within practical resource limitations (e.g., timing, computational, memory, and / or other budgets) of commercial applications.
[0018] For the sake of illustration, FIG. 1 shows a representation of a conventional approach to addressing degradations in a voice communication channel. As illustrated, the approach generally involves receiving a noisy speech signal 105 (e.g., including other voice sounds, ambient noises, artifacts, etc.) on the voice channel, processing the signal through a series of digital processing networks, and outputting a clean speech signal 155. For the sake of illustrated, the digital processing networks can include one or more of a denoising network 110, an artifact correction network 120, a bandwidth expansion network 130, a packet loss concealment network 140, and an equalization network 150.
[0019] For example, the denoising network 110 can address the degradation of speech channels caused by ambient noise. It can isolate and remove unwanted noise components from the audio signal to improve clarity and intelligibility. Utilizing techniques such as deep learning models (e.g., autoencoders) or spectral gating, the denoising network 110 can analyze the spectral characteristics of both noise and speech, distinguishing between them based on training on diverse noise conditions. The denoising network 110 can then suppress or eliminate the noise components while preserving the speech signal's integrity, thereby enhancing the overall quality of the audio.
[0020] The artifact correction network 120 can be designed to rectify artifacts in speech channels that arise from digital processing errors or system malfunctions, which can include jitter, clipping, and other distortions. For example, use of one or more codecs can cause artifacts, which can degrade the signal. The artifact correction network 120 can employ algorithms that identify and modify distorted speech segments, restoring them to their presumed original state, such as by advanced signal processing techniques that model normal speech patterns and use these models to predict and replace corrupted segments.
[0021] The bandwidth expansion network 130 can be designed to address degradations from the use of narrowband audio signals in speech channels, which lack higher frequency content. The bandwidth expansion network 130 can be designed to enhance the perceptual quality of speech by artificially extending the frequency range of the audio signal. Using psychoacoustic models and deep neural networks, it can estimate the missing high-frequency components based on the available narrowband signal. The process can involve learning the relationship between low and high frequencies from wideband speech data and then applying this knowledge to expand the bandwidth of narrowband signals.
[0022] In environments where speech transmission is subject to data packet loss (e.g., in voice over Internet protocol (VoIP) or streamed communications), a packet loss concealment (PLC) network 140 can mitigate the impact of missing packets on perceived audio quality. The PLC network 140 can use techniques, such as interpolation, pattern matching, and / or predictive coding, to estimate and synthesize the missing audio segments. By analyzing the received packets, the PLC network 140 can predict characteristics of the lost audio, thereby filling in gaps to minimize jarring silences or abrupt changes for a smoother auditory experience.
[0023] The equalization network 150 can be used to correct frequency response imbalances in a speech channel. These may be caused by the transmission medium, recording equipment, or environmental factors. The equalization network 150 can adjust the balance of frequencies to achieve a desired or more natural sounding frequency response. Utilizing parametric, graphic, convolutional, and / or other filters, the equalization network 150 can modify the amplitude of specific frequency bands. Some equalization networks 150 can use adaptive equalization techniques to dynamically adjust to changing conditions or different speech characteristics.
[0024] The term “repair networks” is generally used herein to refer to networks, such as the denoising network 110 and the artifact correction network 120, which seek to repair degradations due to noise and artifacts. The term “enhancement networks” is generally used herein to refer tonetworks, such as the bandwidth expansion network 130, the packet loss concealment network 140, and / or the equalization network 150, which seek to enhance the speech signal beyond what is achievable by the repair networks. Other types of enhancement networks can include, for example, a de-reverberation network to address reverberation-based degradations, an echo cancellation network to address echo-based degradations, a voice activity detection (VAD) network to help distinguish between human speech and non-speech segments of the signal, an audio restoration network to address distortion-based degradations (e.g., harmonic distortion, clipping, etc.), a dynamic range compression network to address degradations resulting from fluctuating signal levels, etc.
[0025] Many conventional systems have only one or more repair networks. Limited conventional systems also include one or more enhancement networks. Regardless, as illustrated, conventional approaches tend to decide which degradations will be addressed, dedicate and tailor a network to address each of those degradations, and place each of those networks serially in the signal processing path. Notably, each network consumes processing time, computational resources, memory resources, power resources, and / or other resources. For example, users tend to be disturbed by speech channel latencies in excess of approximately 200 milliseconds (ms). In a typical speech channel, delays unrelated to any digital signal processing (i.e., to any repair and / or enhancement networks) can be around 150 ms. Even if each conventional signal processing network can be designed to add an average of around 20 ms, there may be room to address only two or three types of degradation before the latency becomes bothersome to users. Further, many commercial applications have strict constraints on power, memory, processing, etc., which may only support incorporation of a limited number of these conventional signal processing networks.
[0026] Embodiments described herein seek to provide repair and enhancement for speech channels in a manner that involves relatively low resource usage, such as relatively low computing, latency, and memory. FIG. 2 shows a high-level representation of novel approaches described herein. As illustrated, the novel approaches generally involve receiving a noisy speech signal 105 on a speech channel, processing the signal through an integrated repairer enhancer network (IREN) 210, and outputting a clean speech signal 155. Rather than using a series of multiple digital processing networks, each tailored to address a particular type of degradation, the IREN 210 can be implemented as a single integrated network. As described herein, the IREN 210 is generated according to a progressive learning approach, by which the IREN 210 is initially trained on a first set of degradation types to produce a source network. The source network is used progressively to train (by transfer learning and / or other techniques) a target network until a desiredIREN 210 is produced. For example, the initial training produces a repair network (with features of one or more types of repair networks). The repair network is then further trained to integrate features of one or more enhancement networks until a single IREN 210 is achieved that simultaneously addressed several types of degradations.
[0027] FIGS. 3A and 3B show block diagrams of source model training systems 300, according to embodiments described herein. In each source model training system 300, a clean training speech signal 305 is received as an input to a disrepairer 310, which generates a corresponding disrepaired training speech signal 330. The clean training speech signal 305 is a corpus of clean speech samples produced according to a predefined maximum bandwidth coverage. In one embodiment, the clean training speech signal 305 is a corpus of wideband and / or super-wideband clean speech signal samples.
[0028] The disrepaired training speech signal 330 is generated to mimic the type of speech signal that might be received over a real-world speech channel, including simulated noise and / or artifacts. The disrepaired training speech signal 330 can be provided as an input to a repairing network 340. The repairing network 340 can be any feasible type of machine learning network that uses error back-propagation to learn how to reproduce the clean training speech signal 305 from the disrepaired training speech signal 330 (i.e., to learn to effectively remove the effects of the disrepairer 310). FIGS. 3 A and 3B show a simplified conceptualization of the back- propagation as a subtractor 345 that essentially generates an error as a difference between the disrepaired training speech signal 330 and the clean training speech signal 305 and propagates the error back into the repairing network 340 for learning.
[0029] In the source model training system 300-1 of FIG. 3 A, the disrepairer 310-1 includes a source-to-microphone response (STMR) simulator block 320, a mixer block 320, a signal-to-noise ratio (SNR) selection block 322, and a noise production block 324. The STMR simulator block 320 models and analyzes how sound travels from a source (e.g., speaker) to a microphone through one or more predicted environments and / or setups. The simulation can be implemented using digital signal processing techniques that incorporate acoustic modeling to mimic physical spaces, and / or using approaches like impulse response analysis and convolution with environmental noise models. One or more simulations is configured to capture and predict interactions between an emitted sound from a source and its reception by a microphone, considering factors like distance, intervening media, ambient noise, and channel properties.
[0030] The STMR simulator block 320 can be configured with vectors representing one or more longer channel cases, such as representing hands-free communication modes. In such cases, room impulse response (RIR) vectors can represent different types of rooms, different positions of a speaking individual relative to a hands-free microphone, etc. For example, different STMR convolutions are used to output vectors representing different simulated cases. The STMR simulator block 320 can also be configured with vectors that do not include STMR convolution for shorter channel cases, such as to represent handset, earbud, and / or other cases in which the source and microphone are in close proximity. In some implementations, even for shorter channel cases, the STMR simulator block 320 can be configured to model feedback control, body effect, microacoustic effects (e.g., occlusion or wind noise), etc.
[0031] As illustrated, the SNR selection block 322 can select among several predefined SNR schemes (e.g., sets of design constraints). For each SNR scheme, the noise production block 324 can produce one or more pre-modeled noise signals. For example, the noise signals can represent any suitable types of noises that are likely to be experienced on a speech channel, such as wind noise, traffic noise, ambient conversational noises, nature sounds, etc. The mixer block 320 outputs a signal that is essentially an adding together of the noise signals and the clean training speech signal 305 for the different SNR schemes. This output signal is the disrepaired training speech signal 330-1.
[0032] In the source model training system 300-2 of FIG. 3B, the disrepairer 310-2 includes the STMR simulator block 320, a mixer block 320, the SNR selection block 322, the noise production block 324, and a codec 350. The codec 350 can be implemented as a full-band codec. In a speech channel, a full-band codec typically designed to encode and decode audio covering a wide frequency range from about 20 Hz to 20 kHz, which encompasses the full spectrum of human hearing. Such a codec 350 typically employs a high sampling rate (e.g., 40 kHz) to adequately capture frequencies up to 20 kHz based on the Nyquist theorem. The codec 350 can perform quantization by converting the sampled analog audio signals into digital form with a sufficient bit depth (e.g., 16 bits or more per sample) to maintain high sound quality throughout the processes of compression and decompression. Some implementations of the codec 350 utilize advanced compression techniques to efficiently handle the large amount of data generated by the high sampling rate and bit depth while minimizing loss of quality.
[0033] In typical speech channels, codecs can generate several types of artifacts due, for example, to the compressing and decompressing of the audio data for reducing the data's bandwidth. Common artifacts in these cases can include pre-echo, where a sound is heard beforeits actual occurrence, and temporal smearing, where sharp speech sounds become blurred over time. Such artifacts can tend to result from inaccuracies in the psychoacoustic models used by audio codecs, which attempt to determine which sounds are audible in the presence of louder noises. In some cases, codecs can tend to over-compress speech, thereby stripping out natural tonal variations and dynamic range of human voices and resulting in a “robotic” or “metallic” sound. Lossy compression in some codecs can further manifest as a loss of clarity in consonant sounds. The codec 350 in the disrepairer 310-2 is designed to mimic expected types of codecs that may be seen in a speech channel (e.g., a worst-case codec), so that passing the signal through the codec 350 introduces codec-related artifacts into the signal.
[0034] As illustrated, the mixer 320 can output the disrepaired training speech signal 330-1 as described with reference to FIG. 3A. This disrepaired training speech signal 330-1 can be processed by the codec 350, which can generate a corresponding disrepaired training speech signal 330-2. The disrepaired training speech signal 330-2 thereby includes both introduced noise and introduced artifacts. Thus, in FIG. 3B embodiments, the repairing network 340 is trained to regenerate the clean training speech signal 305 in presence of both introduced noise and introduced artifacts.
[0035] In both of FIGS. 3A and 3B, the repairing network 340 can be implemented using any feasible machine learning model that can process the disrepaired training speech signal 330 to regenerate the clean training speech signal 305 by leveraging error back-propagation for training. One implementation of the repairing network 340 uses an autoencoder network (e.g., a denoising autoencoder). The autoencoder is configured to learn to encode the input into a lower-dimensional space (referred to herein as an “embedding layer”) and then to decode it back to the original dimensions, while filtering out noise. In this context, the embedding layer is a bottleneck layer, at which the disrepaired training speech signal 330 is represented in a most compressed representation.
[0036] Another implementation of the repairing network 340 uses a convolutional neural network (CNN). The CNN can be effective for processing the disrepaired training speech signal 330, particularly in cases where the noise and signal characteristics have spatial-like patterns or correlations that convolutional layers can exploit. In the context of the repairing network 340, the CNN can be adapted to include an embedding layer after convolutional layers, where the convolved features can be flattened and passed through a dense layer to form embeddings prior to reconstruction or classification layers.
[0037] Another implementation of the repairing network 340 uses a recurrent neural network (RNN). The RNN (e.g., or a variant, such as a long short-term memory (LSTM) network, a gated recurrent unit (GRU), etc.) can exploit sequential (e.g., time-series) aspects of the disrepaired training speech signal 330 to capture temporal dynamics. In one version of this implementation, an embedding layer is incorporated before the recurrent layers to transform raw features into a dense representation. In another version of this implementation, an embedding layer is incorporated within the recurrent structure by treating the outputs of the recurrent layers as embeddings.
[0038] Another implementation of the repairing network 340 uses a transformer model. Transformer models can exploit self-attention mechanisms to weigh the importance of different parts of the input data without the constraints of sequential processing. In such implementations, input signals are first converted into embeddings (i.e., ultimately in an embedding layer) prior to being processed by attention layers.
[0039] Another implementation of the repairing network 340 uses generative adversarial network (GAN). In such an implementation, a generator network can attempt to clean the noisy signal, while a discriminator network can try to distinguish between the cleaned signal and an actual clean signal (e.g., via the subtractor 345). The generator can include the embedding layer as part of its architecture, converting noisy input into a dense, meaningful representation before attempting to reconstruct a clean output.
[0040] In general, the repairing network 340 is trained to keep learning in feedback until its bottleneck layer is able to automatically and faithfully repair any disrepair caused by the disrepairer 310. In this way, the repairing network 340 can take the place of at least a denoising network and / or an artifact correction network, such as the denoising network 110 and / or artifact correction network 120 of the conventional approach illustrated in FIG. 1. In embodiments described herein, the repairing network 340 also becomes a baseline network (e.g., teacher network, source network, etc.) for integration of one or more enhancement network features by transfer learning and / or other techniques.
[0041] FIG. 4 shows an illustrative IREN training system 400 that uses the repairing network 340 to generate an IREN 210-1 with integrated bandwidth expansion features. Embodiments of the IREN training system 400 can include the disrepairer 310 that receives a clean training speech signal 305. The disrepairer 310 can be implemented as the disrepairer 310 of FIG. 3 A or 3B, or any feasible variant thereof, such that the output of the disrepairer 310 is the disrepaired trainingspeech signal 330 with at least introduced noise and / or codec artifacts. The disrepaired training speech signal 330 is fed to a bandwidth limiter 410. As described above, speech can often be degraded from the use of narrowband audio signals in speech channels, which lack higher frequency content. For example, compression techniques, lower fidelity microphones, and / or other factors can tend to narrow the band of frequencies used to transmit the audio signal.
[0042] The bandwidth limiter 410 can include any components feasible for mimicking such bandwidth limiting. In the illustrated implementation, the bandwidth limiter 410 includes a downsampler 415, a bandpass filter 420, and an up-sampler 425. The down-sampler 415 reduces the sampling rate of the audio signal, which effectively decreases the amount of data being processed. Such down-sampling can be typical in audio systems for reducing computational loads and storage requirements for handling the audio data, but it can lead to loss of higher frequency components of the signal (due to a lower Nyquist frequency), which can tend to degrade speech. The down- sampled signal can be passed through the bandpass filter 420, which selectively allows frequencies within a certain range to pass while attenuating frequencies outside this range. The bandpass filter 420 can be tuned to pass frequencies in the band that is most salient for speech intelligibility, such as between 300 Hz and 3400 Hz. The down-sampled, filtered signal is then passed through the up- sampler 425, which increases the sampling rate back to its original level (or to different, e.g., higher, resolution). Such up-sampling typically involves interpolation and estimation to fill-in data points between existing samples. Such interpolation and estimation tend not to produce a perfect recreation of the original signal, thereby manifesting degradations to the speech audio signal.
[0043] In some implementations, the bandwidth limiter 410 further includes a narrow-band codec 350’ between the down-sampler 415 and the up-sampler 425 (e.g., after the bandpass filter 420). The narrow-band codec 350’ encodes and compresses the down-sampled and filtered speech audio signal within a limited frequency range (the bandpass range, such as 300 Hz to 3400 Hz). The narrow-band codec 350’ can digitize the signal (if not already digital) and compresses the digital signal (e.g., by predictive coding and quantization). Such codecs tend to have lossy compression, such that some information is permanently discarded. While this reduces the data load on the system, it can also affect signal quality. The narrow-band codec 350’ can then encapsulates the signal into a format suitable for transmission or storage (to mimic what would happen in real-world conditions) and can then decompress the signal. The decompressed output is an approximation, such that additional degradations are likely introduced.
[0044] The output of the bandwidth limiter 410 is illustrated as a disrepaired de-enhanced training speech signal 430-1, indicating that the clean training speech signal 305 has been purposefully degraded in manner that mimic those addressed by both repair and enhancement networks. The disrepaired de-enhanced training speech signal 430-1 is loaded into the input of the IREN 210-1. Transfer learning is used to train the baseline of the IREN 210-1, and then error back-propagation is used to continue training the IREN 210-1 to compensate for the integration of bandwidth limiting de-enhancements. As illustrated, the error back-propagation is based on a comparison of the output of the IREN 210-1 against the clean training speech signal 305 (e.g., the error production and back-propagation is represented by subtractor 345).
[0045] Transfer learning in this context involves leveraging knowledge (features, weights, biases, etc.) from a pre-existing “teacher network” (i.e., the repairing network 340) to accelerate and enhance the learning process of a new “student network” (i.e., the IREN 210-1) on a related but extended task. Here, the repairing network 340 has already been successfully trained to address degradations due to disrepairing and possesses a fully trained bottleneck layer, which effectively captures the most salient information needed for repairing those degradations. The bottleneck layer acts as a compact representation or embedding of the learned features, encapsulating the distilled knowledge of the network about that problem. The trained bottleneck layer from the repairing network 340 can be directly used or fine-tuned in the IREN 210-1. This allows the IREN 210-1 to not start from scratch, but instead to build upon the learned representations of the repairing network 340, adapting and extending them to address the new complexities of degradations due to de-enhancements.
[0046] FIGS. 5A and 5B show illustrative IREN training systems 500 that use the repairing network 340 to generate an IREN 210-2 with integrated comfort enhancement features. Embodiments of the IREN training system 500 can include the disrepairer 310 that receives a clean training speech signal 305. The disrepairer 310 can be implemented as the disrepairer 310 of FIG. 3 A or 3B, or any feasible variant thereof, such that the output of the disrepairer 310 is the disrepaired training speech signal 330 with at least introduced noise and / or codec artifacts.
[0047] As illustrated, the clean training speech signal 305 can also be passed through a comfort enhancement processor 510. As used herein, “comfort enhancement” generally describes audio enhancement features, such as automatic gain control (AGC), volumetric control, and equalization. For example, AGC can automatically adjust the audio signal's volume, ensuring consistent loudness without manual intervention; and equalization can involve adjusting the balance between frequency components within an audio signal (e.g., to boost bass, enhance speech frequencybands, etc.). The comfort enhancement processor 510 can include any suitable components for implementing audio comfort enhancement features. For example, the comfort enhancement processor 510 can include an AGC block for maintaining consistent audio levels (e.g., by using a control loop to adjust the gain automatically based on the input signal's amplitude to ensure output stability), a volumetric control block for allowing manual adjustment of output volume through a user interface (e.g., using a digital potentiometer, or the like), an equalization block with multiple filters (e.g., band-pass, low-pass, and / or high-pass filters) that are configurable via software or hardware to modify specific frequency bands according to user preferences or environmental requirements, etc.
[0048] In the IREN training system 500-1 FIG. 5A, the disrepaired training speech signal 330 at the output of the disrepairer 310 is loaded into the input of the IREN 210-2. In the IREN training system 500-2 FIG. 5B, the input path to the IREN 210-2 further includes a de-enhancer 530 that introduces one or more types of degradations that are corrected by one or more corresponding types of enhancement networks. When the disrepaired training speech signal 330 is provided at the input of the de-enhancer 530, the de-enhancer 530 generates a disrepaired de-enhanced training speech signal 430 at its output. In one embodiment, the de-enhancer 530 is the bandwidth limiter 410 of FIG. 4, so that the output of the de-enhancer 530 is the disrepaired de-enhanced training speech signal 430-1. In embodiments of FIG. 5B, the disrepaired de-enhanced training speech signal 430 at the output of the de-enhancer 530 is loaded into the input of the IREN 210-2.
[0049] As illustrated, the comfort enhancement processor 510 can output a comfort-enhanced speech signal 520, which is a version of the clean training speech signal 305 with one or more comfort enhancement features applied. Transfer learning is used to train the baseline of the IREN 210-1 based on the repairing network 340 (as described above), and then error back-propagation is used to continue training the IREN 210-1 to compensate for the integration of comfort enhancements. In this case, rather than de-enhancing the signal at the input side of the IREN 210- 2, the error back-propagation is based on a comparison of the output of the IREN 210-2 against the comfort-enhanced speech signal 520 (e.g., the error production and back-propagation is represented by subtractor 345). Thus, the error accounts for the comfort enhancement at the output ide ofthe IREN 210-2.
[0050] In some embodiments, the IREN 210-2 is trained for a set of reference comfort enhancement configurations is available that correspond to a typical set of likely real -world configuration options. The options include one or more AGC target levels, automatic level control levels, equalization settings, and / or combinations thereof. As illustrated, the comfort enhancementprocessor 510 can produce two outputs: a configuration vector 515, and the comfort-enhanced speech signal 520. In one implementation, training iterates through a sequence of the reference comfort enhancement configurations. In each iteration, the configuration vector 515 is a row vector that represents a selected one of the reference comfort enhancement configurations for the iteration, and the comfort-enhanced speech signal 520 is the signal produced by that reference configuration. In another implementation, the configuration vector 515 is a row vector that represents multiple (e.g., all) of the reference comfort enhancement configurations concatenated.
[0051] FIGS. 6A - 6C show three illustrative training phases for an illustrative IREN training system 600 that integrates burst loss prediction and repair into the IREN 210-2. Turning first to FIG. 6 A, the IREN training system 600-1 initially performs base training for continuous embedding prediction even when there are burst packet losses. Some embodiments begin base training from the trained repairing network 340. Other embodiments begin base training after one or more progressive levels of training of the IREN 210. For example, the base training can begin with IREN 210-1 (trained according to FIG. 4), IREN 210-2 (trained according to FIG. 5 A or 5B), etc.
[0052] In some embodiments, the IREN training system 600-1 inputs the clean training speech signal 305 to a disrepairer 310. The disrepairer 310 can be implemented as the disrepairer 310 of FIG. 3 A or 3B, or any feasible variant thereof, such that the output of the disrepairer 310 is the disrepaired training speech signal 330 with at least introduced noise and / or codec artifacts. Some embodiments further include a de-enhancer 530, such as described with reference to FIG. 5B (e.g., implemented as the bandwidth limiter 410 of FIG. 4). As such, the input to the base training model (e.g., disrepairer 310, IREN 210-1, IREN 210-2, etc.) can be the disrepaired training speech signal 330 or a disrepaired de-enhanced training speech signal 430.
[0053] Because the base training is concerned with generating continuous embeddings, the IREN training system 600-1 pulls embedding outputs 605 from the bottleneck layer of the base training model. As described above, the embedding outputs 605 represent a highly compressed version of the speech signal with the most salient features needed to recreate the clean training speech signal 305 in a manner that mitigates or eliminates degradations caused by the disrepairer 310 and / or the de-enhancer 530. The embedding outputs 605 are passed through a frame delay block 610 into an embedding history repository 615.
[0054] The embedding outputs 605 can be high-dimensional vectors representing the salient features of the input data at a more abstract level. Each set of one or more vectors can beconsidered as a “frame” of embedding data. Each frame, then, can be a structured series of numerical values (e.g., floating-point numbers), where each value or set of values encodes some aspect of the input data's information content. The frame delay block 610 is designed to delay the output by K frames (K is a positive integer); when the frame delay block 610 receives frame T from the base training model, it outputs frame T-K (i.e., from K frames ago). In one implementation, the frame delay block 610 uses a software buffer, in which each incoming frame of embeddings is added to a queue, and frames are removed from the queue after a specified delay period (e.g., after K frames). In another implementation, the frame delay block 610 uses a hardware delay line (e.g., a digital delay line), which can shift data through registers or memory elements (e.g., using a field-programmable gate array, a digital signal processor, or other suitable components). In another implementation, the frame delay block 610 uses a circular buffer, in which a fixed-size array and two pointers manage the start and end of the data, and the buffer can rotate back to the beginning once the end of the array is reached.
[0055] The embedding history repository 615 can be implemented with any suitable components for storing at least a previous N frames of embeddings (N is a positive integer). N may or may not be equal to K. In one implementation, the embedding history repository 615 is a database. In another implementation, the embedding history repository 615 is an in-memory data structure, such as a list, dictionary, ring buffer, or custom data structure. In another implementation, the embedding history repository 615 is a file system (e.g., using flat files, or binary formats). For each of a sequence of frame times, the output of the embedding history repository 615 is the most recent N frames.
[0056] As illustrated, a most recent frame of embedding outputs 605 and the output of the embedding history repository 615 are used for base training of an embedding predictor 620-1. The embedding predictor 620-1 is a machine learning model that uses a sequence of previous embedding frames to forecast the next frame's embedding vector, effectively learning the temporal dynamics or patterns within the embeddings. In particular, the last N frames are taken from the embedding history repository 615 as an input to the embedding predictor 620-1, thereby providing a historical context that aids the model in predicting the subsequent (most recent) frame's embedding. The output of the embedding predictor 620-1 (the predicted current frame) is compared to the actual current frame generated at the bottleneck layer of the base training model. The comparison generates an error that is back-propagated for use in training the embedding predictor 620-1. For example, a subtractor 345 represents a loss function (e.g., mean squared error (MSE), cosine similarity, etc.) that results in a calculated error, and the calculated error is used toadjust the weights of the embedding predictor 620-1 through back-propagation. In one embodiment, the embedding predictor 620-1 is implemented by a recurrent neural network (RNN). In another embodiment, the embedding predictor 620-1 is implemented by a long short-term memory network (LSTM). In another embodiment, the embedding predictor 620-1 is implemented by a gated recurrent unit (GRU). In another embodiment, the embedding predictor 620-1 is implemented by a transformer network. In another embodiment, the embedding predictor 620-1 is implemented by a convolutional neural network (CNN). In another embodiment, the embedding predictor 620-1 is implemented by a feedforward neural network.
[0057] At the end of the base training phase, the embedding predictor 620-1 is trained to predict the current frame of embedding outputs 605 using the past N frames. Turning to FIG. 6B, transfer learning can be used to enhance the embedding predictor 620-1 to be able to predict burst losses. The embedding predictor 620-1 from the base training phase (FIG. 6A) is used as a teacher model to teach a second-phase embedding predictor 620-2. As described with reference to FIG. 6A, the disrepaired training speech signal 330 or disrepaired de-enhanced training speech signal 430 can continue to be input to the trained repairing network 340 or partially trained IREN 210, so that frames of embedding outputs 605 are generated at the bottleneck layer.
[0058] The second training phase introduces a packet loss simulator 625, which simulates a burst packet loss condition. The simulation is represented in a simplified manner as a switch that is in a “good” position when simulating that there is no burst loss, or in a “lost” position when simulating that there is burst loss. When the packet loss simulator 625 is simulating the good packet condition, the IREN training system 600-2 is essentially the same as the IREN training system 600-1 of FIG. 6A. For example, the embedding outputs 605 are passed through the frame delay block 610 into the embedding history repository 615, and the embedding history repository 615 generates the input to the embedding predictor 620-2. When the packet loss simulator 625 is simulating the lost packet condition, the output of the embedding predictor 620-2 is fed back to the embedding history repository 615. As such, while simulating that there is burst loss, the embedding predictor 620-2 learns to predict the current frame based on its own prior predictions.
[0059] At the end of the second training phase, the embedding predictor 620-2 is trained to predict a current frame of embedding outputs 605 when there is burst loss. Turning to FIG. 6C, transfer learning can be used to enhance the IREN 210 (the decoder portion of the IREN 210) for phase continuity during burst packet loss and for lost packet synthesis. As in the second training phase, a packet loss simulator 625 simulates either a “good” packet condition when there is no burst loss, or a “lost” packet condition when simulating that there is burst loss. In the good packetcondition, the IREN training system 600-3 can be similar to any of the systems described in FIGS. 4, 5A, or 5B. For example, either the repairing network 340 or an IREN 210 (e.g., IREN 210-1 or IREN 210-2) is used as a teacher model for IREN 210-3. Depending on what training is desired and what training has already occurred, the input to the IREN 210-3 can come from a disrepairer 310 (i.e., disrepaired training speech signal 330) or from a de-enhancer 530 (i.e., disrepaired deenhanced training speech signal 430), and error back-propagation can be based on the clean training speech signal 305 or comfort-enhanced training speech 520 from a comfort enhancement processor 510.
[0060] Essentially, in the good packet condition, the encoder (input) portion of the IREN 210-3 receives “good” packets. In this condition, the IREN 210-3 generates embedding outputs 605 at the bottleneck layer, and the decoder (output) portion of the IREN 210-3 decodes those embeddings to produce highly intelligible speech. The embedding outputs 605 are also stored in the embedding history repository 615 in anticipation of a possible future lost packet condition. When the lost packet condition occurs, the input to the IREN 210-3 is essentially disconnected, which effectively disables the encoding layers of the IREN 210-3. Instead, the trained embedding predictor 620-2 can inject predicted embeddings into the decoding layers of the IREN 210-3 from which the IREN 210-3 can generate an output by effectively synthesizing the lost packets. Errors continue to be computed as back-propagated to train the decoder layers of the IREN 210-3.
[0061] Some embodiments of the IREN 210-3 include “attention” between its encoder and decoder portions. In this context, “attention” is a mechanism that helps the decoder focus on relevant parts of the input sequence for each step of the output generation. For example, the encoder processes the input data and converts it into the embedding outputs 605. When the embedding outputs 605 are passed to the decoder, the attention mechanism dynamically selects which parts of the input sequence are most relevant for predicting each element in the output sequence, such as by creating a set of attention weights that are applied to vectors of the embedding outputs 605. The weights can be learned during training and can be adjusted during each step of the decoding process. The attention mechanism can be gated based on a lost packet indicator (e.g., a flag) output or controlled by the embedding predictor 620-2. The indicator can be used to disable the attention mechanism during the lost packet condition.
[0062] FIG. 7 shows an illustrative implementation environment 700 for a progressively trained IREN 210, according to embodiments described herein. The environment includes a speech audio device 710 with an integrated speech enhancer 720. The speech audio device 710 can be any type of device that receives speech audio at an input interface 715 and outputs speech audio via aspeaker 725. For example, the speech audio device can be a telephone headset, voice over Internet protocol (VoIP) phone, teleconferencing system, etc., and the speaker 725 can be integrated into a wired headset, wireless headset, teleconferencing hub, earbuds, etc.
[0063] As illustrated, a speech signal is received via a speech audio channel 705. The speech audio channel 705 can include one or more phone lines, wired and / or wireless Internet infrastructure, air interfaces, etc.; such that the input interface 715 can be a microphone, phone jack, Ethernet port, optical port, wireless port, etc. As described herein, by the time the speech signal reaches the speech enhancer 720, it can be considered a noisy speech signal 105 because of degradations due to characteristics of the speech audio channel 705 (e.g., noise, burst packet loss, etc.) and / or intervening components (e.g., codecs, bandwidth limiting devices, etc.). Within the speech enhancer, the trained IREN 210 mitigates and / or eliminates the effects of the degradations to output a clean speech signal 155 to the microphone 725.
[0064] Another type of implementation environment 700 for a progressively trained IREN 210 is for speech enhancement in an attention handling system. When a user is listening to music or other desired audio through a wearable audio component, or WAC (e.g., earbuds or on-ear headphones), active noise control (ANC) works to suppress any ambient sound. However, in some instances, ambient sound that is intended for the user can be very important to the user’s connectivity with others. For example, although the user desired to suppress undesirable ambient sound, the user may still desire to be able on occasion to enter into desired conversations. When entering such a conversation, speech enhancement described herein can be used to improve the intelligibility of the speech being received through the wearable audio components. In these contexts, the term “noisy speech” continues to be used to describe degraded received speech signals, the term “playback audio” is used to generally refer to any recorded or streaming audio signal that is being played to the user through the WAC (e.g., music, audiobook, podcast, radio broadcast, live event broadcast, etc.), and the term “ambient audio” is used to generally refer to any audio in the vicinity of the WAC other than the noisy speech and the playback audio.
[0065] FIG. 8 shows a conceptual circuit block diagram of a partial audio management system environment 800 of a wearable audio component (WAC) with an automated attention handling system (AHS) 830 having an integrated speech enhancer 720. The speech enhancer 720 includes a progressively trained IREN 210. As illustrated, the AHS 830 includes an attention seeking (AS) trigger detection block 810, a conversation end detection block 820, and the speech enhancer 720. The AHS 830 is illustrated in context of a playback audio 825 stream, a noisy speech signal 105, and an ambient audio signal 805. For example, in the context of earbuds with ANC, the playbackaudio 825 stream may be received via a connected device (e.g., a smartphone, computer, or other device connected via a wired or wireless connection), and the noisy speech signal 105 and ambient audio signal 805 can be received via a reference microphone 835 (e.g., part of the ANC system).
[0066] The role of the AHS 830 can be generally described as to toggle the audio environment between an active mode and a conversation mode based on whether a desired conversation is detected, as represented by a switch network 815. In the active mode, the user is listening to the desired audio 825 via a speaker 725, and an ANC system (not shown) is suppressing as much of the ambient audio 805 (including any noisy speech 105) as possible. This is conceptually represented by the switches of the switch network 815 being in the solid-line position, whereby the desired audio 825 passes through to the speaker 725, and the ambient audio 805 and noisy speech 105 do not.
[0067] At some point, the AHS system 830 switches to a conversation mode. In some embodiments (or in some cases) the user manually toggles the AHS system 830 to the conversation mode, such as via a particular user interface component and / or command. In other embodiments (or in other cases), switching is based on automatic detection of a conversation trigger by the AS trigger detection block 810. Some techniques for automatic conversation detections are described in International Application No. PCT / US2024 / 014606, titled “AUTOMATED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMS BASED ON LINGUISTIC NAME EMBEDDING, filed on February 6, 2024; International Application No. PCT / US2024 / 014788, titled “AUTOMATED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMS BASED ON UNIVERSAL SOUND CONVERSION”, filed on February 7, 2024; and International Application No. PCT / US2024 / 014820, titled “NAMEDETECTION BASED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMS”, filed on February 7, 2024.
[0068] Switching to the conversation mode is represented by moving the switch network 815 to the dashed-line position. In this position, the ambient audio 825 and the noisy speech 105 are allowed to pass through to the speaker 725, while the desired audio 825 is not. In this mode, at least the noisy speech 105 is passed through the speech enhancer 830 in line with the speaker 725. The IREN 210 in the speech enhancer 720 generates clean speech 155 from the noisy speech 105 and outputs the clean speech to the speaker 725. At some subsequent time, the conversation can end, and the AHS system 830 switches the switch network 815 back to the solid-line position. In some embodiments (or cases), the end of the conversation is automatically detected by theconversation end detection block 820. In other embodiments (or cases), the end of the conversation is detected based on a user interaction and / or command.
[0069] FIG. 9 provides a schematic illustration of an illustrative computational system 900 that can implement various system components and / or perform various steps of methods provided by various embodiments relating to training of an integrated repairer enhancer network (IREN), as described herein. FIG. 9 is meant only to provide a generalized illustration of various components, any or all of which may be utilized as appropriate. FIG. 9, therefore, broadly illustrates how individual system elements may be implemented in a relatively separated or relatively more integrated manner.
[0070] The computational system 900 is shown including hardware elements that can be electrically coupled via a bus 905 (or may otherwise be in communication, as appropriate). The hardware elements may include one or more processors 910, including, without limitation, one or more general-purpose processors and / or one or more special-purpose processors (such as digital signal processing chips, graphics acceleration processors, video decoders, and / or the like); one or more input devices 915; and one or more output devices 920. The input devices 915 can include wired and / or wireless ports, buttons, switches, microphones, touch interfaces, and / or any other suitable input device 915; and the output devices 920 can include indicator lights, displays, speakers, and / or any other suitable output devices 920.
[0071] The computational system 900 may further include (and / or be in communication with) one or more non-transitory storage devices 925, which can comprise, without limitation, local and / or network accessible storage, and / or can include, without limitation, a disk drive, a drive array, an optical storage device, a solid-state storage device, such as a random access memory (“RAM”), and / or a read-only memory (“ROM”), which can be programmable, flash-updateable and / or the like. Such storage devices may be configured to implement any appropriate data stores, including, without limitation, various file systems, database structures, and / or the like. In some embodiments, the storage devices 925 include one or more machine models, such as the repairer network 340, one or more IRENs 210 in one or more phases of progressive training, the embedding predictor 620 (e.g., including the embedding history repository 615), etc. Some embodiments can also include one or more reference sample databases 950, such as including corpuses of clean speech samples, noise samples and / or models, etc.
[0072] The computational system 900 can also include a communications subsystem 930, which can include, without limitation, a modem, a network card (wireless or wired), an infraredcommunication device, a wireless communication device, and / or a chipset (such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication device, etc.), and / or the like. As described herein, the communications subsystem 930 supports multiple communication technologies. Further, as described herein, the communications subsystem 930 can provide communications with one or more networks.
[0073] In many embodiments, the computational system 900 will further include a working memory 935, which can include a RAM or ROM device, as described herein. The computational system 900 also can include software elements, shown as currently being located within the working memory 935, including an operating system 940, device drivers, executable libraries, and / or other code, such as one or more application programs 945, which may include computer programs provided by various embodiments, and / or may be designed to implement methods, and / or configure systems, provided by other embodiments, as described herein. Merely by way of example, one or more procedures described with respect to the method(s) discussed herein can be implemented as code and / or instructions executable by a computer (and / or a processor within a computer); in an aspect, then, such code and / or instructions can be used to configure and / or adapt a general -purpose computer (or other device) to perform one or more operations in accordance with the described methods. In some embodiments, the operating system 940 and the working memory 935 are used in conjunction with the one or more processors 910 to implement some or all of the disrepairer 310, the de-enhancer 530, the comfort enhancement processor 510, and / or other components to support IREN training. While the machine learning models are illustrated as stored in the storage devices 925, aspects of the models can alternatively be considered as being part of working memory 935.
[0074] A set of these instructions and / or codes can be stored on a non-transitory computer- readable storage medium, such as the non-transitory storage device(s) 925 described above. In some cases, the storage medium can be incorporated within a computer system, such as computer system 900. In other embodiments, the storage medium can be separate from a computer system (e.g., a removable medium, such as a compact disc), and / or provided in an installation package, such that the storage medium can be used to program, configure, and / or adapt a general -purpose computer with the instructions / code stored thereon. These instructions can take the form of executable code, which is executable by the computational system 900 and / or can take the form of source and / or installable code, which, upon compilation and / or installation on the computational system 900 (e.g., using any of a variety of generally available compilers, installation programs, compression / decompression utilities, etc.), then takes the form of executable code.
[0075] It will be apparent to those skilled in the art that substantial variations may be made in accordance with specific requirements. For example, customized hardware can also be used, and / or particular elements can be implemented in hardware, software (including portable software, such as applets, etc.), or both. Further, connection to other computing devices, such as network input / output devices, may be employed.
[0076] As mentioned above, in one aspect, some embodiments may employ a computer system (such as the computer system 900) to perform methods in accordance with various embodiments of the invention. According to a set of embodiments, some or all of the procedures of such methods are performed by the computational system 900 in response to processor 910 executing one or more sequences of one or more instructions (which can be incorporated into the operating system 940 and / or other code, such as an application program 945) contained in the working memory 935. Such instructions may be read into the working memory 935 from another computer-readable medium, such as one or more of the non-transitory storage device(s) 925. Merely by way of example, execution of the sequences of instructions contained in the working memory 935 can cause the processor(s) 910 to perform one or more procedures of the methods described herein.
[0077] The terms “machine-readable medium,” “computer-readable storage medium” and “computer-readable medium,” as used herein, refer to any medium that participates in providing data that causes a machine to operate in a specific fashion. These mediums may be non-transitory. In an embodiment implemented using the computer system 900, various computer-readable media can be involved in providing instructions / code to processor(s) 910 for execution and / or can be used to store and / or carry such instructions / code. In many implementations, a computer-readable medium is a physical and / or tangible storage medium. Such a medium may take the form of a non-volatile media or volatile media. Non-volatile media include, for example, optical and / or magnetic disks, such as the non-transitory storage device(s) 925. Volatile media include, without limitation, dynamic memory, such as the working memory 935. Common forms of physical and / or tangible computer-readable media include, for example, a floppy disk, a flexible disk, hard disk, magnetic tape, or any other magnetic medium, a CD-ROM, any other optical medium, any other physical medium with patterns of marks, a RAM, a PROM, EPROM, a FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read instructions and / or code.
[0078] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to the processor(s) 910 for execution. Merely by way ofexample, the instructions may initially be carried on a magnetic disk and / or optical disc of a remote computer. A remote computer can load the instructions into its dynamic memory and send the instructions as signals over a transmission medium to be received and / or executed by the computer system 900. The communications subsystem 930 (and / or components thereof) generally will receive signals, and the bus 905 then can carry the signals (and / or the data, instructions, etc., carried by the signals) to the working memory 935, from which the processor(s) 910 retrieves and executes the instructions. The instructions received by the working memory 935 may optionally be stored on a non-transitory storage device 925 either before or after execution by the processor(s) 910.
[0079] FIG. 10 provides a schematic illustration of another illustrative computational system 1000 that can implement various system components and / or perform various steps of methods provided by various embodiments relating to operation of an integrated repairer enhancer network (IREN) 210, as described herein. In some cases, the same computational system 900 used for training the IREN is also used as the computational system 1000 that provides an operational environment for a trained IREN 210. In other cases, computational system 1000 is different from computational system 900. To avoid over-complicating the description, common reference designators are used for FIGS. 9 and 10, even though the respective computational systems may include different components and / or similar components that operate differently. For example, in general, the bus 905, processors 910, input devices 915, output devices 920, communications subsystem 930, operating system 940, and application programs 945 can provide similar features to those described with reference to FIG. 9.
[0080] In FIG. 10, the non-transitory storage devices 925 are illustrated as storing the trained IREN 210. The working memory 935 is illustrated as implementing a speech enhancer 720. For example, noisy speech signals are received via one or more input devices 915. The speech enhancer 720 executes in working memory 935 to process the noisy speech signals through the IREN 210 and to generate corresponding clean speech. The clean speech can then be output via one or more output device 920 (e.g., a speaker).
[0081] FIG. 11 shows a flow diagram of an illustrative method 1100 for automated speech repair and enhancement for speech audio signals, according to various embodiments. Embodiments of the method 1100 begin at stage 1104 by processing a corpus of clean speech signals by a disrepairer circuit to generate corresponding disrepaired speech signals having one or more disrepair-related degradations. As described herein, some embodiments of the disrepairer circuit add noise to the clean speech signals to generate noisy speech audio signals. For example,the processing in stage 1104 includes applying one or more noise models to the corpus of clean speech signals to add noise to the corresponding disrepaired speech signal. Other embodiments can additionally or alternatively add artifacts to the clean speech signals to mimic the types of artifacts that would result from passing the signal through compression, decompression, and / or other components. For example, the processing in stage 1104 includes applying one or more codec models to the corpus of clean speech signals to add codec-related artifacts to the corresponding disrepaired speech signals.
[0082] At stage 1108, embodiment can load the disrepaired speech signals to a repairer network to generate repairer output signals. For example, the disrepaired speech signals are input to encoding layers of a neural network or other machine learning network via a data loader. The repairer network generates increasingly compact representations that seek to preserve only the most salient features of the signal, until a bottleneck layer is reached (representing a most compact representation). Decoder layers of the repairer network generate increasingly expansive representations of the bottleneck layer representation in an attempt to ultimately output representations (i.e., the repairer output signals) that reproduce the disrepaired speech signals.
[0083] At stage 1112, embodiments can train the repairer network to ameliorate the one or more disrepair-related degradations by back-propagating a repairer error signal computed based on comparing the corpus of clean speech signals with the repairer output signals. The term “ameliorate” is used herein to denote remediate at least until an error loss function falls below a predetermined threshold. For example, the clean speech signals and the repairer output signals are passed to a loss function that computes an error indicating how accurately the repairer output signals represent the clean speech signals. The error is back-propagated to iteratively train the repairer network until the error is below a predetermined threshold. The disrepair-related degradations are considered ameliorated when they are removed from the repairer output signal by the repairer network to an extent such that the error is below the predetermined threshold.
[0084] At stage 1116, embodiments can apply transfer learning to initially train an integrated repairer enhancer network (IREN) from the repairer network. As described herein, the transfer learning involves using the repairer network as a teacher model and the IREN as a student model. For example, weights are copied from the repairer network to the IREN for at least some of the network layers, such that the copied layers maintain learned features from the repairer network’s training data. The architecture of the IREN can then be modified for further training, such as by freezing certain weights of the initial layers copied from the repairer network (i.e., those layers do not update during the training of the student model) and adding new layers or replacing final layersof the IREN network to tailor the output to the additional tasks. New or modified layer may be initialized with randomly initialized weights (i.e., trained from scratch). New training data is then applied to the IREN to effectively adapt the initial learning from the repairer network to the additional tasks.
[0085] At stage 1120, embodiments can progressively train the IREN to ameliorate each of one or more de-enhancement-related degradations. As described herein, the de-enhancement-related degradations can relate to bandwidth limitations, comfort enhancements, burst packet losses, and / or other real-world behaviors. The progressive training, for each de-enhancement-related degradation, can involve performing steps 1124 - 1136. For example, for N degradations, the method 1100 can iterate stages 1124 - 1136 N times (N is a positive integer).
[0086] At stage 1124, embodiments can generate IREN input signals from the disrepaired speech signals and generating reference signals from the corpus of clean speech signals, such that the IREN input signals are more degraded than the reference signals relative to the de- enhancement-related degradation. In some embodiments, this involves degrading (i.e., deenhancing) the IREN inputs signals relative to the reference signals. In other embodiments, this involves enhancing the reference signals relative to the IREN inputs signals.
[0087] At stage 1128, embodiments can load the IREN input signals to the IREN to generate IREN output signals. For example, as described above, the IREN input signals are loaded to input layers, the IREN iteratively compresses the IREN input signals through input layers until a bottleneck layer, and the IREN iteratively decompresses the bottleneck layer representation through output layers to generate the IREN output signals. At stage 1132, embodiments can train the IREN to ameliorate the de-enhancement-related degradation based on back-propagating an IREN error signal computed based on comparing the IREN output signals with the reference signals. As described above, the de-enhancement-related degradations are considered ameliorated when they are removed from the IREN output signals by the IREN to an extent such that the error is below a predetermined threshold.
[0088] At stage 1136, a determination can be made as to whether additional progressive training of the IREN is desired for additional de-enhancement-related degradations. If so, embodiments can return to stage 1124 for the next de-enhancement-related degradation. If not, the training can be considered complete, and the method 1100 can end. In some cases, the fully trained IREN can then be used to generate clean audio from degraded audio. For example, as illustrated, embodiments can continue at stage 1140 by receiving (by the IREN) a degraded speech audiosignal that is degraded by at least one of the disrepair-related degradations and by at least one of the de-enhancem ent-related degradations that the IREN was trained to ameliorate. At stage 1144, embodiments can output (by the IREN), automatically in response to the receiving at stage 1140, a clean speech signal with the at least one of the disrepair-related degradations and by at least one of the de-enhancement-related degradations ameliorated.
[0089] FIG. 12 shows a flow diagram of an illustrative implementation of progressive training of an IREN at stage 1120 of the method 1100 of FIG. 11. Embodiments can begin at stage 1204 by determining which type of de-enhancement-related degradation is being trained. The illustrated implementation shows three de-enhancement-related degradations: bandwidth limitation, comfort de-enhancement, and burst loss. As described herein, other embodiments can be trained to ameliorate other types of de-enhancement-related degradations.
[0090] In some embodiments, one of the one or more de-enhancement-related degradations is bandwidth limitation. In such embodiments, for the one of the one or more de-enhancement- related degradations, generating the IREN input signals from the disrepaired speech signals at stage 1124 (shown as stage 1124-1) involves passing the disrepaired speech signals through a bandwidth limiter. The reference signals can be the clean speech signals. As such, the IREN input signals are more degraded than the reference signals relative to the bandwidth limitation. In some such embodiments, passing the disrepaired speech signals through the bandwidth limiter involves down-sampling, bandpass filtering, and up-sampling the disrepaired speech signals. In some such embodiments, passing the disrepaired speech signals through the bandwidth limiter further involves applying a narrow-band codec subsequent to the down-sampling and prior to the up- sampling.
[0091] In other embodiments, one of the one or more de-enhancement-related degradations is comfort de-enhancement. In such embodiments, for the one of the one or more de-enhancement- related degradations, the generating the reference signals from the corpus of clean speech signals at stage 1124 (shown as stage 1124-2) involves passing the corpus of clean speech signals through a comfort enhancement processor, such that the IREN input signals are more degraded than the reference signals relative to the comfort de-enhancement (i.e., the reference signals are more enhanced than the IREN input signals). In some such embodiments, passing the corpus of clean speech signals through the comfort enhancement processor involves: applying each of a predefined plurality of comfort enhancement configuration options to the clean speech signals, each comfort enhancement configuration defining corresponding settings for one or more of automatic gaincontrol, volumetric control, or equalization control; and training the IREN comprises training the IREN for the predefined plurality of comfort enhancement configuration options.
[0092] In other embodiments, one of the one or more de-enhancement-related degradations is burst loss. In such embodiments, for the one of the one or more de-enhancement-related degradations, the progressive training can, at stage 1208, involve inputting a training speech audio signal to a base training model (i.e., the repairer network or a partially trained IREN) to cause a bottleneck layer of the base training model to generate a sequence of embedding frames. At stage 1212, embodiments can base train an embedding predictor to predict and output an embedding by, for each of a sequence of base training times, generating a present frame prediction based on preceding ones of the sequence of embedding frames, and back-propagating a frame error signal computed based on comparing the present frame prediction with a present one of the sequence of embedding frames.
[0093] At stage 1216, embodiments can further train the embedding predictor subsequent the base training at stage 1212, to predict and output lost embedding frames. This can be performed for each of a sequence of further training times, some of the further training times being good packet times and others of the further training times being lost packet times. The training at stage 1216 can involve generating the present frame prediction based on an embedding repository comprising a fixed number of most recently added frames and back-propagating the frame error signal computed based on comparing the present frame prediction with the present one of the sequence of embedding frames. In stage 1216, the embedding repository is updated differently based on whether the present frame is simulated as “good” or “lost.” If the further training time is one of the good frame times, the training at stage 1216 can further involve adding the present one of the sequence of embedding frames to the embedding repository. If the further training time is one of the good frame times, the training at stage 1216 can further involve adding the present frame prediction to the embedding repository.
[0094] As illustrated, stages 1208 - 1216 can be an implementation of stage 1124 (shown as stage 1124-3). In such an implementation, the IREN input signals are a combination of the historical embeddings received from the embedding repository and the embeddings injected into the bottleneck layer from the embedding predictor. As these signals are an estimation of the actual frame embeddings being generated by the base training model, the IREN input signals can be considered as more degraded than the reference signals with respect to burst loss (i.e., as recited in stage 1124). Further, loading of these signals causes the IREN to output IREN output signals.
[0095] At stage 1128, subsequent to an of stages 1124-1, 1124-2, or 1124-3, embodiments can load the IREN input signals to the IREN to generate IREN output signals, as described with reference to FIG. 11. Further as described with reference to FIG. 11, embodiments can then train the IREN to ameliorate the type of de-enhancement-related degradation at stage 1132 and can determine whether to iterate for further progressive training at stage 1136.
[0096] Having described several example configurations, various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the disclosure. For example, the above elements may be components of a larger system, wherein other rules may take precedence over or otherwise modify the application of the invention. Also, a number of steps may be undertaken before, during, or after the above elements are considered.
Claims
WHAT IS CLAIMED IS:
1. A method for automated speech repair and enhancement for speech audio signals, the method comprising: processing a corpus of clean speech signals by a disrepairer circuit to generate corresponding disrepaired speech signals having one or more disrepair-related degradations; loading the disrepaired speech signals to a repairer network to generate repairer output signals; training the repairer network to ameliorate the one or more disrepair-related degradations by back-propagating a repairer error signal computed based on comparing the corpus of clean speech signals with the repairer output signals; applying transfer learning to initially train an integrated repairer enhancer network (IREN) from the repairer network; and progressively training the IREN to ameliorate each of one or more de-enhancement- related degradations by: generating IREN input signals from the disrepaired speech signals and generating reference signals from the corpus of clean speech signals, such that the IREN input signals are more degraded than the reference signals relative to the de-enhancement- related degradation; loading the IREN input signals to the IREN to generate IREN output signals; and training the IREN to ameliorate the de-enhancement-related degradation based on back-propagating an IREN error signal computed based on comparing the IREN output signals with the reference signals.
2. The method of claim 1, further comprising, subsequent to the progressively training: receiving, by the IREN, a degraded speech audio signal that is degraded by at least one of the disrepair-related degradations and by at least one of the de-enhancement-related degradations; outputting, by the IREN, automatically in response to the receiving, a clean speech signal with the at least one of the disrepair-related degradations and by at least one of the de- enhancement-related degradations ameliorated.
3. The method of claim 1, wherein:one of the one or more de-enhancement-related degradations is bandwidth limitation; and for the one of the one or more de-enhancement-related degradations, the generating the IREN input signals from the disrepaired speech signals comprises passing the disrepaired speech signals through a bandwidth limiter, such that the IREN input signals are more degraded than the reference signals relative to the bandwidth limitation.
4. The method of claim 3, wherein passing the disrepaired speech signals through the bandwidth limiter comprises down-sampling, bandpass filtering, and up-sampling the disrepaired speech signals.
5. The method of claim 4, wherein passing the disrepaired speech signals through the bandwidth limiter further comprises applying a narrow-band codec subsequent to the down-sampling and prior to the up-sampling.
6. The method of claim 1, wherein: one of the one or more de-enhancement-related degradations is comfort deenhancement; and for the one of the one or more de-enhancement-related degradations, the generating the reference signals from the corpus of clean speech signals comprises passing the corpus of clean speech signals through a comfort enhancement processor, such that the IREN input signals are more degraded than the reference signals relative to the comfort de-enhancement.
7. The method of claim 6, wherein: passing the corpus of clean speech signals through the comfort enhancement processor comprises applying each of a predefined plurality of comfort enhancement configuration options to the clean speech signals, each comfort enhancement configuration defining corresponding settings for one or more of automatic gain control, volumetric control, or equalization control; and training the IREN comprises training the IREN for the predefined plurality of comfort enhancement configuration options.
8. The method of claim 1, wherein: processing the corpus of clean speech signals by the disrepairer circuit comprises applying one or more noise models to the corpus of clean speech signals to add noise to the corresponding disrepaired speech signal.
9. The method of claim 1, wherein: processing the corpus of clean speech signals by the disrepairer circuit comprises applying one or more codec models to the corpus of clean speech signals to add codec-related artifacts to the corresponding disrepaired speech signals.
10. The method of claim 1, further comprising: inputting a training speech audio signal to a base training model to cause a bottleneck layer of the base training model to generate a sequence of embedding frames, wherein the base training model is the repairer network or the IREN; base training an embedding predictor to predict an embedding by, for each of a sequence of base training times, generating a present frame prediction based on preceding ones of the sequence of embedding frames, and back-propagating a frame error signal computed based on comparing the present frame prediction with a present one of the sequence of embedding frames; and further training the embedding predictor subsequent the base training by, for each of a sequence of further training times, some of the further training times being good packet times and others of the further training times being lost packet times, generating the present frame prediction based on an embedding repository comprising a fixed number of most recently added frames, back-propagating the frame error signal computed based on comparing the present frame prediction with the present one of the sequence of embedding frames, and adding to the embedding repository either the present one of the sequence of embedding frames if the further training time is one of the good frame times or the present frame prediction if the further training time is one of the lost frame times, wherein one of the one or more de-enhancement-related degradations is burst loss, and progressively training the IREN to ameliorate the burst loss using the embedding predictor subsequent to the further training.
11. A system for automated speech repair and enhancement for speech audio signals, the system comprising: one or more processors: a non-transitory computer-readable memory having instructions stored thereon which, when executed, cause the one or more processors to perform steps comprising:processing a corpus of clean speech signals by a disrepairer circuit to generate corresponding disrepaired speech signals having one or more disrepair-related degradations; loading the disrepaired speech signals to a repairer network to generate repairer output signals; training the repairer network to ameliorate the one or more disrepair-related degradations by back-propagating a repairer error signal computed based on comparing the corpus of clean speech signals with the repairer output signals; applying transfer learning to initially train an integrated repairer enhancer network (IREN) from the repairer network; and progressively training the IREN to ameliorate each of one or more de- enhancem ent-related degradations by: generating IREN input signals from the disrepaired speech signals and generating reference signals from the corpus of clean speech signals, such that the IREN input signals are more degraded than the reference signals relative to the de-enhancement-related degradation; loading the IREN input signals to the IREN to generate IREN output signals; and training the IREN to ameliorate the de-enhancement-related degradation based on back-propagating an IREN error signal computed based on comparing the IREN output signals with the reference signals.
12. The system of claim 11, wherein the steps further comprise, subsequent to the progressively training: receiving, by the IREN, a degraded speech audio signal that is degraded by at least one of the disrepair-related degradations and by at least one of the de-enhancement-related degradations; outputting, by the IREN, automatically in response to the receiving, a clean speech signal with the at least one of the disrepair-related degradations and by at least one of the de- enhancement-related degradations ameliorated.
13. The system of claim 11, wherein: one of the one or more de-enhancement-related degradations is bandwidth limitation; andfor the one of the one or more de-enhancement-related degradations, the generating the IREN input signals from the disrepaired speech signals comprises passing the disrepaired speech signals through a bandwidth limiter, such that the IREN input signals are more degraded than the reference signals relative to the bandwidth limitation.
14. The system of claim 13, wherein passing the disrepaired speech signals through the bandwidth limiter comprises down-sampling, bandpass filtering, and up-sampling the disrepaired speech signals.
15. The system of claim 14, wherein passing the disrepaired speech signals through the bandwidth limiter further comprises applying a narrow-band codec subsequent to the down-sampling and prior to the up-sampling.
16. The system of claim 11, wherein: one of the one or more de-enhancement-related degradations is comfort deenhancement; and for the one of the one or more de-enhancement-related degradations, the generating the reference signals from the corpus of clean speech signals comprises passing the corpus of clean speech signals through a comfort enhancement processor, such that the IREN input signals are more degraded than the reference signals relative to the comfort de-enhancement.
17. The system of claim 16, wherein: passing the corpus of clean speech signals through the comfort enhancement processor comprises applying each of a predefined plurality of comfort enhancement configuration options to the clean speech signals, each comfort enhancement configuration defining corresponding settings for one or more of automatic gain control, volumetric control, or equalization control; and training the IREN comprises training the IREN for the predefined plurality of comfort enhancement configuration options.
18. The system of claim 11, wherein: processing the corpus of clean speech signals by the disrepairer circuit comprises applying one or more noise models to the corpus of clean speech signals to add noise to the corresponding disrepaired speech.
19. The system of claim 11, wherein:processing the corpus of clean speech signals by the disrepairer circuit comprises applying one or more codec models to the corpus of clean speech signals to add codec-related artifacts to the corresponding disrepaired speech signals.
20. The system of claim 11, wherein the steps further comprise: inputting a training speech audio signal to a base training model to cause a bottleneck layer of the base training model to generate a sequence of embedding frames, wherein the base training model is the repairer network or the IREN; base training an embedding predictor to predict an embedding by, for each of a sequence of base training times, generating a present frame prediction based on preceding ones of the sequence of embedding frames, and back-propagating a frame error signal computed based on comparing the present frame prediction with a present one of the sequence of embedding frames; and further training the embedding predictor subsequent the base training by, for each of a sequence of further training times, some of the further training times being good packet times and others of the further training times being lost packet times, generating the present frame prediction based on an embedding repository comprising a fixed number of most recently added frames, back-propagating the frame error signal computed based on comparing the present frame prediction with the present one of the sequence of embedding frames, and adding to the embedding repository either the present one of the sequence of embedding frames if the further training time is one of the good frame times or the present frame prediction if the further training time is one of the lost frame times, wherein one of the one or more de-enhancement-related degradations is burst loss, and progressively training the IREN to ameliorate the burst loss using the embedding predictor subsequent to the further training.
Citation Information
Patent Citations
Automated attention handling in active noise control systems based on linguistic name embedding
WO2025128138A1
Automated attention handling in active noise control systems based on universal sound conversion
WO2025128139A1
Name-detection based attention handling in active noise control systems
WO2025128140A1