Generative audio codec for signal synthesis based on spectral envelope features and pitch information
The generative audio codec system addresses inefficiencies in audio coding by separating spectral envelope and pitch features using feedback recurrent autoencoders and neural synthesizers, achieving efficient bit compression and real-time processing on resource-constrained devices.
Patent Information
- Application Number
- PCT/US2025/028260
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2025-05-07
- Publication Date
- 2025-11-20
AI Technical Summary
Existing audio coding technologies face challenges in efficiently compressing audio data while maintaining quality, particularly with the increased computational complexity of neural network-based decoders, which can lead to prolonged frame processing times on resource-constrained devices.
A generative audio codec system utilizing feedback recurrent autoencoders and neural network-based synthesizers separates spectral envelope and pitch features, employing filter banks and neural networks to generate encoded representations, which are then processed by a neural speech synthesizer to produce reconstructed audio.
The system achieves efficient bit compression of audio data by avoiding double-coding of information, reducing computational load, and enabling real-time processing on resource-constrained devices.
Smart Images

Figure US2025028260_20112025_PF_FP_ABST
Abstract
Description
Qualcomm Docket No.2403757WO GENERATIVE AUDIO CODEC FOR SIGNAL SYNTHESIS BASED ON SPECTRAL ENVELOPE FEATURES AND PITCH INFORMATION FIELD
[0001] Aspects of the present disclosure generally relate to audio coding (e.g., audio encoding and / or decoding). In some implementations, examples are described for performing audio coding using a generative audio codec system including one or more feedback recurrent autoencoders and a neural network-based synthesizer. BACKGROUND
[0002] Audio coding (also referred to as voice coding and / or speech coding) is a technique used to represent a digitized audio signal using as few bits as possible (thus compressing the speech data), while attempting to maintain a certain level of audio quality. An audio or voice encoder is used to encode (or compress) the digitized audio (e.g., speech, music, etc.) signal to a lower bit-rate stream of data. The lower bit-rate stream of data can be input to an audio or voice decoder, which decodes the stream of data and constructs an approximation or reconstruction of the original signal. The audio or voice encoder-decoder structure can be referred to as an audio coder (or voice coder or speech coder) or an audio / voice / speech coder-decoder (codec).
[0003] Audio coders exploit the fact that speech signals are highly correlated waveforms. Some speech coding techniques are based on a source-filter model of speech production, which assumes that the vocal cords are the source of spectrally flat sound (an excitation signal), and that the vocal tract acts as a filter to spectrally shape the various sounds of speech. The different phonemes (e.g., vowels, fricatives, and voice fricatives) can be distinguished by their excitation (source) and spectral shape (filter). SUMMARY
[0004] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary has the sole purpose to present certain concepts relating to one or more aspectsQualcomm Docket No.2403757WO relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.
[0005] Disclosed are systems, methods, apparatuses, and computer-readable media for audio coding (e.g., encoding and / or decoding audio data). According to at least one illustrative example, an apparatus for processing audio data is provided. The apparatus includes at least one memory and at least one processor coupled to the at least one memory and configured to: determine spectral envelope features of the audio, wherein the spectral envelope features are determined based on processing the audio using a filter bank including a plurality of filters; determine one or more pitch features of the audio; generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the spectral envelope features, wherein the first encoded representation is based on state feedback between a decoder portion and an encoder portion of the first neural network-based autoencoder; and generate a second encoded representation corresponding to the one or more pitch features.
[0006] In another example, a method for processing audio data is provided. The method includes: determining spectral envelope features of the audio, wherein the spectral envelope features are determined based on processing the audio using a filter bank including a plurality of filters; determining one or more pitch features of the audio; generating, using a first neural network-based autoencoder, a first encoded representation corresponding to the spectral envelope features, wherein the first encoded representation is based on state feedback between a decoder portion and an encoder portion of the first neural network-based autoencoder; and generating a second encoded representation corresponding to the one or more pitch feature;.
[0007] In another example, a non-transitory computer-readable medium is provided that includes instructions that, when executed by at least one processor, cause the at least one processor to: determine spectral envelope features of the audio, wherein the spectral envelope features are determined based on processing the audio using a filter bank including a plurality of filters; determine one or more pitch features of the audio; generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the spectral envelope features, wherein the first encoded representation is based on state feedback between a decoder portion and an encoder portion of the firstQualcomm Docket No.2403757WO neural network-based autoencoder; and generate a second encoded representation corresponding to the one or more pitch features.
[0008] In another example, an apparatus for processing audio data is provided. The apparatus includes: means for determining spectral envelope features of the audio, wherein the spectral envelope features are determined based on processing the audio using a filter bank including a plurality of filters; means for determining one or more pitch features of the audio; means for generating, using a first neural network-based autoencoder, a first encoded representation corresponding to the spectral envelope features, wherein the first encoded representation is based on state feedback between a decoder portion and an encoder portion of the first neural network-based autoencoder; and means for generating a second encoded representation corresponding to the one or more pitch features.
[0009] In another illustrative example, an apparatus for processing audio data (e.g., decoding audio data) is provided. The apparatus includes at least one memory and at least one processor coupled to the at least one memory and configured to: receive a first encoded representation corresponding to spectral envelope features of the audio; receive a second encoded representation corresponding to one or more pitch features of the audio; generate reconstructed spectral envelope features corresponding to the audio, based on processing the first encoded representation using a first neural network-based decoder; and generate a synthesized audio output signal based on processing the reconstructed spectral envelope features and the second encoded representation using a neural network- based signal synthesizer.
[0010] In another example, a method for processing audio data is provided. The method includes: receiving a first encoded representation corresponding to spectral envelope features of the audio; receiving a second encoded representation corresponding to one or more pitch features of the audio; generating reconstructed spectral envelope features corresponding to the audio, based on processing the first encoded representation using a first neural network-based decoder; and generating a synthesized audio output signal based on processing the reconstructed spectral envelope features and the second encoded representation using a neural network-based signal synthesizer.
[0011] In another example, a non-transitory computer-readable medium is provided that includes instructions that, when executed by at least one processor, cause the at leastQualcomm Docket No.2403757WO one processor to: receive a first encoded representation corresponding to spectral envelope features of the audio; receive a second encoded representation corresponding to one or more pitch features of the audio; generate reconstructed spectral envelope features corresponding to the audio, based on processing the first encoded representation using a first neural network-based decoder; and generate a synthesized audio output signal based on processing the reconstructed spectral envelope features and the second encoded representation using a neural network-based signal synthesizer.
[0012] In another example, an apparatus for processing audio data is provided. The apparatus includes: means for receiving a first encoded representation corresponding to spectral envelope features of the audio; means for receiving a second encoded representation corresponding to one or more pitch features of the audio; means for generating reconstructed spectral envelope features corresponding to the audio, based on processing the first encoded representation using a first neural network-based decoder; and means for generating a synthesized audio output signal based on processing the reconstructed spectral envelope features and the second encoded representation using a neural network-based signal synthesizer.
[0013] Aspects generally include a method, apparatus, system, computer program product, non-transitory computer-readable medium, user equipment, base station, wireless communication device, and / or processing system as substantially described herein with reference to and as illustrated by the drawings and specification.
[0014] The foregoing has outlined rather broadly the features and technical advantages of examples according to the disclosure in order that the detailed description that follows may be better understood. Additional features and advantages will be described hereinafter. The conception and specific examples disclosed may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. Characteristics of the concepts disclosed herein, both their organization and method of operation, together with associated advantages, will be better understood from the following description when considered in connection with the accompanying figures. Each of the figures is provided for the purposes of illustration and description, and not as a definition of the limits of the claims.
[0015] While aspects are described in the present disclosure by illustration to someQualcomm Docket No.2403757WO examples, those skilled in the art will understand that such aspects may be implemented in many different arrangements and scenarios. Techniques described herein may be implemented using different platform types, devices, systems, shapes, sizes, and / or packaging arrangements. For example, some aspects may be implemented via integrated chip implementations or other non-module-component based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, and / or artificial intelligence devices). Aspects may be implemented in chip-level components, modular components, non-modular components, non-chip-level components, device-level components, and / or system-level components. Devices incorporating described aspects and features may include additional components and features for implementation and practice of claimed and described aspects. For example, transmission and reception of wireless signals may include one or more components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders, and / or summers). It is intended that aspects described herein may be practiced in a wide variety of devices, components, systems, distributed arrangements, and / or end-user devices of varying size, shape, and constitution.
[0016] Other objects and advantages associated with the aspects disclosed herein will be apparent to those skilled in the art based on the accompanying drawings and detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.
[0017] The foregoing, together with other features and aspects, will become more apparent upon referring to the following specification, claims, and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are presented to aid in the description of various aspects of the disclosure and are provided solely for illustration of the aspects and not limitation thereof.Qualcomm Docket No.2403757WO
[0019] FIG.1 is a block diagram illustrating an example speech processing system, in accordance with some examples;
[0020] FIG. 2A is a block diagram illustrating an example feature generator, in accordance with some examples;
[0021] FIG.2B is a block diagram illustrating an example of a voice coding system, in accordance with some examples;
[0022] FIG. 2C is a block diagram illustrating an example of a code-excited linear prediction (CELP)-based voice coding system utilizing a fixed codebook (FCB), in accordance with some examples;
[0023] FIG. 3 is a block diagram illustrating an example of a voice coding signal synthesis system utilizing a linear time-varying filter generated using a neural network model and a separate linear predictive coding (LPC) filter, in accordance with some examples;
[0024] FIG.4 is a block diagram illustrating an example of an audio codec system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer, in accordance with some examples;
[0025] FIG. 5A is a block diagram illustrating an example of an audio codec system that can perform spectral envelope feature extraction from an audio signal based on truncation of discrete cosine transform (DCT) coefficients, in accordance with some examples;
[0026] FIG. 5B is a block diagram illustrating an example of an audio codec system that can perform spectral envelope feature extraction using a filter bank and a determination of energies per frequency band, in accordance with some examples;
[0027] FIG.6 is a block diagram illustrating an example of a decoder of an audio codec system that can be used to perform neural speech synthesis based on using an inverse DCT to determine companded filter bank energy information from a decoded cepstrum, in accordance with some examples;
[0028] FIG.7 is a block diagram illustrating an example of a decoder of an audio codec system that can be used to perform neural speech synthesis based on using an inverseQualcomm Docket No.2403757WO DCT and an inverse companding function to determine de-companded filter bank energy information from a decoded cepstrum, in accordance with some examples;
[0029] FIG.8 is a block diagram illustrating an example of a decoder of an audio codec system that can be used to perform neural speech synthesis based on generating reconstructed features with both spectral envelope and pitch information, in accordance with some examples;
[0030] FIG.9 is a block diagram illustrating an example of a decoder of a voice coding signal synthesis system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer comprising a neural homomorphic vocoder (NHV), in accordance with some examples;
[0031] FIG.10 is a block diagram illustrating an example of a decoder of a voice coding signal synthesis system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer comprising a linear predictive coding (LPC) network, in accordance with some examples;
[0032] FIG.11 is a block diagram illustrating an example of an audio codec system that encodes energy information using a first feedback recurrent autoencoder (FRAE) and encodes spectral envelope features using a second FRAE, in accordance with some examples;
[0033] FIG.12 is a block diagram illustrating an example of an audio codec system that can be used to perform neural speech synthesis based on an encoded vector of features including spectral energy features, energy features, and pitch information, in accordance with some examples;
[0034] FIG.13 is a block diagram illustrating an example of an audio codec system that can be used to perform neural speech synthesis based on encoded vectors of various combinations of spectral energy features, energy features, and pitch information, in accordance with some examples;
[0035] FIG.14 is a block diagram illustrating an example of an audio codec system that can be used to perform neural speech synthesis based on using a first denoiser associated with spectral envelope feature extraction and a second denoiser associated with pitch extraction, in accordance with some examples;Qualcomm Docket No.2403757WO
[0036] FIG.15 is a block diagram illustrating an example of an audio codec system that can be used to perform neural speech synthesis based on using a front-end non-speech detector (FNSD) to switch between a neural synthesizer audio codec path for encoding and / or decoding speech signals and an additional audio codec path for encoding and / or decoding non-speech signals, in accordance with some examples;
[0037] FIG. 16 is a flow chart illustrating an example of a process for processing one or more audio samples, in accordance with some examples;
[0038] FIG. 17 is a flow chart illustrating an example of a process for processing one or more audio samples, in accordance with some examples; and
[0039] FIG. 18 is a block diagram illustrating an example of a computing system for implementing certain aspects described herein. DETAILED DESCRIPTION
[0040] Certain aspects of this disclosure are provided below for illustration purposes. Alternate aspects may be devised without departing from the scope of the disclosure. Additionally, well-known elements of the disclosure will not be described in detail or will be omitted so as not to obscure the relevant details of the disclosure. Some of the aspects described herein may be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and description are not intended to be restrictive.
[0041] The ensuing description provides example aspects only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the example aspects will provide those skilled in the art with an enabling description for implementing an example aspect. It should be understood that various changes may be made in the function and arrangement of elements without departing from the scope of the application as set forth in the appended claims.
[0042] Audio coding (e.g., speech coding, music signal coding, or other type of audio coding) can be performed on a digitized audio signal (e.g., a speech signal) to compress the amount of data for storage, transmission, and / or various other uses. A voice codingQualcomm Docket No.2403757WO system can be an audio coding system that is used to perform audio coding on a speech signal that represents the speech (e.g., voiced words or sounds) uttered by a user of the system. A voice coding system can also be referred to as a voice or speech coder or a voice coder-decoder (codec). A voice coding system may include one or more encoders (e.g., voice encoders) and one or more decoders (e.g., voice decoders). For example, a voice encoder can be used to process an input speech signal. Input speech signals may include a digitized speech signal generated from an analog speech signal from a given source, where the resulting digitized speech signal is a discrete-time speech signal with sample values (e.g., also referred to as samples or audio samples) that are also discretized.
[0043] Voice coders (e.g., voice encoders and / or voice decoders) can be designed and implemented to exploit the fact that speech signals are highly correlated waveforms. The samples of an input speech signal can be divided into blocks of N samples each, where a block of N samples is referred to as a frame. In some examples, each frame can be 10-20 milliseconds (ms) in length. A voice encoder can generate a compressed signal corresponding to the input, digitized speech signal. The compressed signal generated by the voice encoder can include a lower bitrate stream of audio data that represents the speech signal using as few bits as possible, while attempting (e.g., by the voice encoder) to maintain a certain quality level for the speech. The compressed signal generated by the voice encoder can have a lower bitrate than the uncompressed digitized speech signal provided as input to the voice encoder. In some cases, voice encoders may be used to reduce the bitrate of an input speech signal. The bitrate of a signal is based on the sampling frequency (e.g., the sampling frequency of a microphone used to sample a user’s speech) and the number of bits per sample. For example, the bitrate of a signal (BR) can be equal to S*b, where S represents the sampling frequency and b represents the number of bits per sample.
[0044] The output of a voice encoder may be a compressed speech signal. The compressed speech signal can be stored and / or transmitted to and processed by a corresponding voice decoder. The corresponding voice decoder can be a voice decoder that is associated with the voice encoder and / or a voice decoder that is configured to decode the compressed speech signals generated by the voice encoder. In some cases, a voice decoder may communicate with a voice encoder, for example to request speech data, to transmit feedback information, and / or to provide other communications to or from the voice encoder and voice decoder. The voice decoder can be configured to decode theQualcomm Docket No.2403757WO data of the compressed speech signal and construct a reconstructed speech signal that approximates the original speech signal. For example, the reconstructed speech signal may include a digitized, discrete-time signal with a same or similar bitrate as the bitrate of the original speech signal. In some cases, the voice decoder may implement an inverse of a voice coding algorithm used by the voice encoder.
[0045] In some examples, one or more machine learning systems, networks, models, etc., can be used to perform audio coding (e.g., encoding and / or decoding of one or more audio samples). For example, machine learning systems using one or more neural network models can be used to implement a voice coder to encode and / or decode voice data (e.g., speech signals, etc.). In some cases, machine learning systems (e.g., using one or more neural network models) can be used to generate reconstructed voice or audio signals using a process of neural synthesis. For example, using features extracted from one or more frames of audio data, a neural network-based voice decoder can generate coefficients for one or more linear filters, learned filters, etc., that may be used to perform voice decoding. In some cases, a neural network-based voice decoder can be trained to estimate the parameters (e.g., linear filter coefficients, etc.) of a speech synthesis pipeline, where learned filters that are generated or tuned using a neural network-based voice decoder may subsequently be used to generate a reconstructed speech signal (e.g., a synthesized speech signal, a speech synthesis signal, etc.).
[0046] Neural network-based voice or speech decoders may also be referred to as neural decoders. Neural decoders can be more computationally complex to implement than non- neural or non-machine learning-based voice or speech decoders. For example, a neural decoder may be associated with a longer frame processing time (e.g., a longer frame decoding time) than a non-neural decoder, when the neural decoder and the non-neural decoder are implemented or executed on the same computational device, platform, hardware, etc. In some cases, the use of neural decoders with resource-constrained devices such as UEs, smartphones, mobile computing devices, wearables, etc., may also be associated with increased frame processing times by the neural decoder, including increased frame processing times beyond the real-time processing requirement associated with the frame length (e.g., 20ms speech frames, etc.).
[0047] Systems, apparatuses, methods (also referred to as processes), and computer- readable media (collectively referred to herein as “systems and techniques”) are describedQualcomm Docket No.2403757WO herein that can be used to perform neural network-based and / or machine learning-based encoding and decoding of audio data. For example, the systems and techniques can be used to provide a bit-efficient generative audio coding system (e.g., audio codec) that utilizes an encoder with one or more feedback recurrent autoencoders (FRAEs) to generate encoded representations of audio that separately represents pitch information and spectral envelope information of the audio data being encoded. The systems and techniques can utilize a decoder that receives the encoded representations of the audio from the encoder, and generates a reconstructed (e.g., recovered) audio signal based on the encoded representations received from the encoder. In some examples, the decoder can include one or more FRAE decoders and a neural synthesizer configured to perform signal synthesis and generate the reconstructed audio signal from the encoded representations received from the encoder.
[0048] In some examples, the systems and techniques can be used to implement a bit- efficient generative voice codec configured to process audio data comprising voice or speech signals. The encoder of the bit-efficient generative voice codec can be used to extract pitch and spectral envelope features from an input audio (e.g., one or more audio frames provided as input to the encoder). The encoder can split the input audio signal into respective spectral features (e.g., spectral envelope features associated with the input audio signal) and respective pitch features or pitch information, where the spectral features and the pitch information are non-overlapping information. For example, information represented in and / or indicated by the spectral features is not represented in or indicated by the pitch information, and information represented in and / or indicated by the pitch information is not represented in or indicated by the spectral features. In some examples, the spectral features can be spectral envelope features associated with the overall shape of the frequency spectrum of the audio signal (e.g., where the spectral envelope connects the peaks of the individual frequency components in the frequency spectrum of an audio signal).
[0049] In some aspects, the encoder of the generative voice codec system can be configured to extract spectral envelope features based on determining a cepstrum and / or cepstral coefficients corresponding to the input audio received by the encoder. The bit efficiency of the encoder can be improved based on generating and / or extracting the spectral envelope features without including pitch information. Removing pitch information from the extracted spectral envelope features can avoid the double-coding orQualcomm Docket No.2403757WO overlapping of information between the spectral features and the pitch information included in the encoded audio representations transmitted between the encoder and the decoder of the generative voice codec.
[0050] In some examples, pitch information can be removed from the spectral envelope features based on performing discrete cosine transform (DCT) truncation to remove or skip the computation of a subset of higher order DCT coefficients. For example, to generate the spectral envelope features, the encoder can process the audio signal using a filterbank of size N (e.g., N frequency bands), apply a DCT of size N, and perform truncation to the first M DCT coefficients (e.g., cepstral coefficients) where M is less than N. Based on performing the DCT truncation to remove pitch information from the spectral features, the extracted spectral envelope features can be generated with a relatively high resolution (e.g., corresponding to the relatively large filterbank size N), and truncated to the smaller size M without a reduction in the resolution. Based on utilizing a relatively large filterbank with a relatively large number of frequency bands, and subsequently truncating the DCT of the relatively large filterbank output, the systems and techniques can be used to code only spectral envelope information in the cepstral features, without including or double coding pitch information in the cepstral features.
[0051] The cepstral features used to code the spectral envelope information (e.g., the extracted spectral envelope features associated with the filterbank and DCT truncation) can be compressed (e.g., encoded) using one or more feedback recurrent autoencoders (FRAE) included in the encoder of the generative voice codec system. In some aspects, a first FRAE can be used to compress (e.g., encode) the cepstral features corresponding to the spectral envelope of the input audio signal, and a second FRAE can be used to compress (e.g., encode) extracted pitch information associated with the input audio signal. In some examples, an FRAE can be used to compress (e.g., encode) the cepstral features corresponding to the spectral envelope of the input audio signal, and the extracted pitch information can be quantized using a pitch quantizer or pitch quantization engine.
[0052] The systems and techniques can utilize one or more FRAEs in the encoder-side of the generative voice codec, and can include one or more FRAE decoders (e.g., the FRAE autoencoder architecture without the encoder portion thereof) in the decoder-side of the generative voice codec. The FRAE decoders can be used to decode the encoded cepstral features (e.g., encoded spectral envelope information) and / or to decode theQualcomm Docket No.2403757WO encoded pitch information (e.g., in examples where an FRAE is used to encode the pitch information at the encoder-side of the generative voice codec).
[0053] Autoencoders have gained popularity in recent years as they are able to learn efficient representations of input data without the need of labels (e.g., based on performing unsupervised learning). Various types of autoencoders exist and are well explained in “Autoencoder and its various variants” 2018 IEEE International Conference on Systems, Man, and Cybernetics , Zhang et. al. In speech coding, autoencoders conditioned on log Mel Cepstrum inputs and / or spectrogram inputs have been used for speech compression in "Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders" arXiv:2102.06610v1 [eess.AS] 12 Feb 2021, Casebeer et. al.
[0054] Speech coding is a lossy compression process. Autoencoders perform dimensionality reduction (e.g., an N-dimensional vector input vector is passed into the encoder of the autoencoder, and lossy compression is performed to represent the important aspects of the input vector in an M-dimensional vector, where M is smaller than N, e.g. by an order of magnitude). A drawback of using an autoencoder for speech coding is that it cannot necessarily exploit the temporal relationship between sets of input data. U.S. Patent No. 11,526,734 (“the ‘734 patent”, assigned to Qualcomm, Inc.), "Method and Apparatus for recurrent auto encoding", Yang et. al. improves upon an autoencoder to exploit temporal redundancies and correlations and describes a feedback recurrent autoencoder ("FRAE"). The FRAE in the '734 patent can be used for training and application of compression of sequential data with temporal correlation. The recurrent structure of the FRAE can be used to efficiently extract the redundancy embedded along the time-dimension of sequential data and enables compact discrete representation of the data at the bottleneck in a sequential fashion. For example, Table 1 of the '734 patent illustrates the MSE (Mean Squared Error) used as Mel-scale mean-square-error configured as a reconstruction loss for training of the FRAE, where the MSE of each frequency bin is scaled according to its weight at Mel-frequency both for latent feedback and output feedback.
[0055] The FRAE described in the ‘734 patent has two advantageous features not present in autoencoders: (1) recurrent layers, e.g. LSTM or GRU layers, have memory of the past; and (2) feedback from the decoder of the autoencoder to the encoder of theQualcomm Docket No.2403757WO autoencoder. The feedback connection 150 in Fig.2, Fig.3 and Fig. 7 of the ‘734 patent provides additional historical information from a state (ht) in the decoder to the encoder indicative of how reconstruction of a prior input has fared. This feedback loop is present during training and inference, which allows the encoder to be trained to respond to particular feedback conditions in a manner that improves reproduction of the input data by the decoder. Thus, the feedback loop is analogous to a mode switch input in that it is not encoded by the encoder, but it is used as an input that influences how the encoder operates on the next input vector.
[0056] In addition, the ‘734 patent describes a second feedback connection (152) from the state (ht) of the decoder for a first iteration of series of inputs Xt, to the next iteration of series of inputs Xt+1. This second feedback connection (152) enables the decoder to learn from its previous reconstruction attempts, providing additional historical context about how reconstruction of a prior input has fared. This second feedback loop is also present during both training and inference, which allows the decoder to be trained to respond to particular feedback conditions in a manner that improves reproduction of the input data by the decoder.
[0057] The ‘734 patent also describes other optional feedback connections. For example, an embedding vector (z) may be fed back via a third feedback connection (356 in Fig.3 of the ‘734 patent) to the encoder. As another optional example, the '734 patent describes that the FRAE can be a variational autoencoder. In this example, an output of the decoder (754 in Fig. 7 of the ‘734 patent) is sample and parameterized to a generate an autoregressive prior which can be used as a fourth feedback connection to condition a prior model for the next latent space embedded vector (zt+1). Similar to the previously described feedback connection, each of these optional connections is present both training and inference, and thus are trained as part of the training process thereby improving functionality of the FRAE.
[0058] The decoded spectral envelope features and the decoded and / or dequantized pitch information determined by the decoder of the generative voice codec can be processed by a neural speech synthesizer included in the decoder. For example, the neural speech synthesizer can be a neural homomorphic vocoder (NHV), among various other neural speech synthesizers. Based on implementing a neural speech synthesizer (e.g., NHV, etc.) configured to process the decoded spectral envelope features and / or pitchQualcomm Docket No.2403757WO information corresponding to the encoded audio signal, the systems and techniques can be used to provide a generative decoder (e.g., a signal synthesis decoder) for generating the reconstructed audio signal. For example, the neural speech synthesizer of the decoder can generate a waveform corresponding to a reconstructed audio signal, based on using one or more generative machine learning models (e.g., generative neural network(s)) included in the neural speech synthesizer to process the decoded spectral envelope features and pitch information received from the encoder.
[0059] Further aspects of the systems and techniques will be described with respect to the figures.
[0060] FIG.1 illustrates an example implementation of a system-on-a-chip (SoC) 100, which may include a central processing unit (CPU) 102 or a multi-core CPU, configured to perform one or more of the functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with a computational device (e.g., neural network with weights), delays, frequency bin information, task information, among other information may be stored in a memory block associated with a neural processing unit (NPU) 108, in a memory block associated with a CPU 102, in a memory block associated with a graphics processing unit (GPU) 104, in a memory block associated with a digital signal processor (DSP) 106, in a memory block 118, and / or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from a memory block 118.
[0061] The SoC 100 may also include additional processing blocks tailored to specific functions, such as a GPU 104, a DSP 106, a connectivity block 110, which may include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, and the like, and a multimedia processor 112 that may, for example, detect and recognize gestures, speech, and / or other interactive user action(s) or input(s). In one implementation, the NPU 108 is implemented in the CPU 102, DSP 106, and / or GPU 104. The SoC 100 may also include a sensor processor 114, image signal processors (ISPs) 116, and / or a signal synthesis system 120. For example, the signal synthesis system 120 can be implemented or configured as a speech synthesis system, including a neural speech decoder and / or a neural homomorphic vocoder (NHV) system, which can be used to generate speech (e.g.,Qualcomm Docket No.2403757WO perform speech synthesis), can be implemented in a text-to-speech (TTS) system, etc. In some examples, the sensor processor 114 can be associated with or connected to one or more sensors for providing sensor input(s) to sensor processor 114. For example, the one or more sensors and the sensor processor 114 can be provided in, coupled to, or otherwise associated with a same computing device.
[0062] In some examples, the one or more sensors can include one or more microphones for receiving sound (e.g., an audio input), including sound or audio inputs that can be used to perform various speech synthesis tasks and / or to generate a reconstructed speech signal from the audio input, etc. In some cases, the sound or audio input received by the one or more microphones (and / or other sensors) may be digitized into data packets for analysis and / or transmission. The audio input may include ambient sounds in the vicinity of a computing device associated with the SoC 100 and / or may include speech from a user of the computing device associated with the SoC 100. In some cases, a computing device associated with the SoC 100 can additionally, or alternatively, be communicatively coupled to one or more peripheral devices (not shown) and / or configured to communicate with one or more remote computing devices or external resources, for example using a wireless transceiver and a communication network, such as a cellular communication network.
[0063] SoC 100, DSP 106, NPU 108 and / or signal synthesis (e.g., neural speech decoder, NHV, etc.) system 120 may be configured to perform audio signal processing. For example, the signal synthesis system 120 may be configured to perform steps for speech synthesis and / or neural homomorphic vocoding, etc. As another example, one or more portions of the steps, such as feature generation, for speech synthesis and / or NHV may be performed by the signal synthesis system 120 while the DSP 106 / NPU 108 performs other steps, such as steps using one or more machine learning networks and / or machine learning techniques according to aspects of the present disclosure and as described herein.
[0064] FIG. 2A depicts an example of a feature generator 200, in accordance with aspects of the present disclosure. It should be understood that many techniques may be used to generate feature vectors for an audio input and that feature generator 200 is just a single example of a technique that may be used to generate feature vectors.
[0065] Feature generator 200 receives an audio signal at signal pre-processor 202. AsQualcomm Docket No.2403757WO above, the audio signal may be from an audio source of an electronic device, such a microphone. Signal pre-processor 202 may perform various pre-processing steps on the received audio signal. For example, signal pre-processor 202 may split the audio signal into parallel audio signals and delay one of the signals by a predetermined amount of time to prepare the audio signals for input into a Fast-Fourier Transform (FFT) circuit. As another example, signal pre-processor 202 may perform a windowing function, such as a Hamming, Hann, Blackman-Harris, Kaiser-Bessel window function, or other sine-based window function, which may improve the performance of further processing stages, such as signal domain transformer 204. Generally, a windowing (or window) function in may be used to reduce the amplitude of discontinuities at the boundaries of each finite sequence of received audio signal data to improve further processing. As another example, signal pre-processor 202 may convert the audio signal data from parallel to serial, or vice versa, for further processing. The pre-processed audio signal from the signal pre-processor 202 may be provided to signal domain transformer 204, which may transform the pre-processed audio signal from a first domain into a second domain, such as from a time domain into a frequency domain.
[0066] In some aspects, signal domain transformer 204 implements a Fourier transform, such as a Fast-Fourier transform (FFT). For example, in some cases, the Fast Fourier transform may be a 16-band (or bin, channel, or point) FFT, which generates a compact feature set that may be efficiency processed by a model. In some cases, a Fourier transform provides fine spectral domain information about the incoming audio signal as compared to conventional single channel processing, such as conventional hardware SNR threshold detection. The result of signal domain transformer 204 is a set of audio features, such as a set of voltages, powers, or energies per frequency band in the transformed data.
[0067] The set of audio features may then be provided to signal feature filter 206, which may reduce the size of or compress the feature set in the audio feature data. In some aspects, signal feature filter 206 may discard certain features from the audio feature set, such as symmetric or redundant features from multiple bands of a multi-band FFT. Discarding this data reduces the overall size of the data stream for further processing and may be referred to a compressing the data stream. For example, in some cases, a 16-band FFT may include 8 symmetric or redundant bands of after the powers are squared because audio signals are real. Thus, signal feature filter 206 may filter out the redundant or symmetric band information and output an audio feature vector 208. In some cases, outputQualcomm Docket No.2403757WO of the signal feature filter may be compressed or otherwise processed prior to output as the audio feature vector 208. The audio feature vector 208 may be provided to a speech synthesis system (e.g., such as the signal synthesis system 120 of FIG. 1) for processing by speech synthesis or NHV model.
[0068] Audio coding (e.g., speech coding, music signal coding, or other type of audio coding) can be performed on a digitized audio signal (e.g., a speech signal) to compress the amount of data for storage, transmission, and / or other use. FIG.2B is a block diagram illustrating an example of a voice coding system 250 (which can also be referred to as a voice or speech coder or a voice coder-decoder (codec)). A voice encoder 252 of the voice coding system 250 can use a voice coding algorithm to process a speech signal 251. The speech signal 251 can include a digitized speech signal generated from an analog speech signal from a given source. For instance, the digitized speech signal can be generated using a filter to eliminate aliasing, a sampler to convert to discrete-time, and an analog- to-digital converter for converting the analog signal to the digital domain. The resulting digitized speech signal (e.g., speech signal 251) is a discrete-time speech signal with sample values (referred to herein as samples) that are also discretized.
[0069] Using the voice coding algorithm, the voice encoder 252 can generate a compressed signal (including a lower bit-rate stream of data) that represents the speech signal 251 using as few bits as possible, while attempting to maintain a certain quality level for the speech. The voice encoder 252 can use any suitable voice coding algorithm, such as a linear prediction coding algorithm (e.g., Code-excited linear prediction (CELP), algebraic-CELP (ACELP), or other linear prediction technique) or other voice coding algorithm.
[0070] The voice encoder 252 can compress the speech signal 251 in an attempt to reduce the bit-rate of the speech signal 251. The bit-rate of a signal is based on the sampling frequency and the number of bits per sample. For instance, the bit-rate of aspeech signal can be determined as ^^^^ ൌ ^^ ∗ ^^, where BR is the bit-rate, S is the samplingfrequency, and b is the number of bits per sample. In one illustrative example, at a sampling frequency (S) of 8 kilohertz (kHz) and at 16 bits per sample (b), the bit-rate (BR) of a signal would be a bit-rate of 128 kilobits per second (kbps).
[0071] The compressed speech signal can then be stored and / or sent to and processed by a voice decoder 254. In some examples, the voice decoder 254 can communicate withQualcomm Docket No.2403757WO the voice encoder 252, such as to request speech data, send feedback information, and / or provide other communications to the voice encoder 252. In some examples, the voice encoder 252 or a channel encoder can perform channel coding on the compressed speech signal before the compressed speech signal is sent to the voice decoder 254. For instance, channel coding can provide error protection to the bitstream of the compressed speech signal to protect the bitstream from noise and / or interference that can occur during transmission on a communication channel.
[0072] The voice decoder 254 can decode the data of the compressed speech signal and construct a reconstructed speech signal 255 that approximates the original speech signal 251. The reconstructed speech signal 255 includes a digitized, discrete-time signal that can have the same or similar bit-rate as that of the original speech signal 251. The voice decoder 254 can use an inverse of the voice coding algorithm used by the voice encoder 252, which as noted above can include any suitable voice coding algorithm, such as a linear prediction coding algorithm (e.g., CELP, ACELP, or other suitable linear prediction technique) or other voice coding algorithm. In some cases, the reconstructed speech signal 255 can be converted to continuous-time analog signal, such as by performing digital-to- analog conversion and anti-aliasing filtering.
[0073] Voice coders can exploit the fact that speech signals are highly correlated waveforms. The samples of an input speech signal can be divided into blocks of N samples each, where a block of N samples is referred to as a frame. In one illustrative example, each frame can be 10-20 milliseconds (ms) in length.
[0074] Various voice coding algorithms can be used to encode a speech signal. For instance, code-excited linear prediction (CELP) is one example of a voice coding algorithm. The CELP model is based on a source-filter model of speech production, which assumes that the vocal cords are the source of spectrally flat sound (an excitation signal), and that the vocal tract acts as a filter to spectrally shape the various sounds of speech. The different phonemes (e.g., vowels, fricatives, and voice fricatives) can be distinguished by their excitation (source) and spectral shape (filter).
[0075] In general, CELP uses a linear prediction (LP) model to model the vocal tract, and uses entries of a fixed codebook (FCB) as input to the LP model. For instance, long- term linear prediction can be used to model pitch of a speech signal, and short-term linear prediction can be used to model the spectral shape (phoneme) of the speech signal. EntriesQualcomm Docket No.2403757WO in the FCB are based on coding of a residual signal that remains after the long-term and short-term linear prediction modeling is performed. For example, long-term linear prediction and short-term linear prediction models can be used for speech synthesis, and a fixed codebook (FCB) can be searched during encoding to locate the best residual for input to the long-term and short-term linear prediction models. The FCB provides the residual speech components not captured by the short-term and long-term linear prediction models. A residual, and a corresponding index, can be selected at the encoder based on an analysis-by-synthesis process that is performed to choose the best parameters so as to match the original speech signal as closely as possible. The index can be sent to the decoder, which can extract the corresponding LTP residual from the FCB based on the index.
[0076] FIG. 2C is a block diagram illustrating an example of a CELP-based voice coding system 270, including a voice encoder 272 and a voice decoder 274. The voice encoder 272 can obtain a speech signal 271 and can segment the samples of the speech signal into frames and sub-frames. For instance, a frame of N samples can be divided into sub-frames. In one illustrative example, a frame of 240 samples can be divided into four sub-frames each having 60 samples. For each frame, sub-frame, or sample, the voice encoder 272 chooses the parameters (e.g., gain, filter coefficients or linear prediction (LP) coefficients, etc.) for a synthetic speech signal so as to match as much as possible the synthetic speech signal with the original speech signal.
[0077] The voice encoder 272 can include a short-term linear prediction (LP) engine 280, a long-term linear prediction (LTP) engine 282, and a fixed codebook (FCB) 284. The short-term LP engine 280 models the spectral shape (phoneme) of the speech signal. For example, the short-term LP engine 280 can perform a short-term LP analysis on each frame to yield linear prediction (LP) coefficients. In some examples, the input to the short- term LP engine 280 can be the original speech signal or a pre-processed version of the original speech signal. In some implementations, the short-term LP engine 280 can perform linear prediction for each frame by estimating the value of a current speech sample based on a linear combination of past speech samples. For example, a speechsignal s(n) can be represented using an autoregressive (AR) model, such as ^^^^^^ ൌ∑^ ^ୀ^ ^^^^^^^^ െ ^^^ ^ ^^^^^^, where each sample is represented as a linear combination of m samples plus a prediction error term ^^^^^^. The weighting coefficients a1, a2, through am can be referred to as the LP coefficients. The prediction error term ^^^^^^Qualcomm Docket No.2403757WOcan be found as follows: ^^^^^^ ൌ ^^^^^^ െ ∑^ ^ୀ^ ^^^^^^^^ െ ^^^. By minimizing the mean square prediction error with respect to the filter coefficients, the short-term LP engine 280 can obtain the LP coefficients. The LP coefficients can be used to form an analysis filter as given in Equation 1, below: ^ ^^^^^^ ൌ 1 െ ^ ^^^^^ି^Eq. (1)
[0078] The short-term LP engine 280 can solve for ^^^^^^ (which can be referred to as a transfer function) by computing the LP coefficients (^^^) that minimize the error in the above AR model equation (^^^^^^) or other error metric. In some implementations, the LP coefficients can be determined using a Levinson–Durbin method, a Leroux–Gueguen algorithm, or other suitable technique. In some examples, the voice encoder 272 can send the LP coefficients to the voice decoder 274. In some examples, the voice decoder 274 can determine the LP coefficients, in which case the voice encoder 272 may not send the LP coefficients to the voice decoder 274. In some examples, Line Spectral Frequencies (LSFs) can be computed instead of or in addition to LP coefficients.
[0079] The LTP engine 282 models the pitch of the speech signal. Pitch is a feature that determines the spacing or periodicity of the impulses in a speech signal. For example, speech signals are generated when the airflow from the lungs is periodically interrupted by movements of the vocal cords. The time between successive vocal cord openings corresponds to the pitch period. The LTP engine 282 can be applied to each frame or each sub-frame of a frame after the short-term LP engine 280 is applied to the frame. The LTP engine 282 can predict a current signal sample from a past sample that is one or more pitch periods apart from a current sample (hence the term “long-term”). For instance, thecurrent signal sample can be predicted as ^^^^^^^ ൌ ^^^^^^^^ െ ^^^, where T denotes the pitchperiod, ^^^ denotes the pitch gain, and ^^^^^ െ ^^^ denotes an LP residual for a previoussample one or more pitch periods apart from a current sample. Pitch period can be estimated at every frame. By comparing a frame with past samples, it is possible to identify the period in which the signal repeats itself, resulting in an estimate of the actual pitch period. The LTP engine 282 can be applied separately to each sub-frame.
[0080] The FCB 284 can include a number (denoted as L) of long-term linear prediction (LTP) residuals. An LTP residual includes the speech signal components that remain afterQualcomm Docket No.2403757WO the long-term and short-term linear prediction modeling is performed. The LTP residuals can be, for example, fixed or adaptive and can contain deterministic pulses or random noise (e.g., white noise samples). The voice encoder 272 can pass through the number L of LTP residuals in the FCB 284 a number of times for each segment (e.g., each frame or other group of samples) of the input speech signal, and can calculate an error value (e.g., a mean-squared error value) after each pass. The LTP residuals can be represented using codevectors. The length of each codevector can be equal to the length of each sub-frame, in which case a search of the FCB 284 is performed once every sub-frame. The LTP residual providing the lowest error can be selected by the voice encoder 272. The voice encoder 272 can select an index corresponding to the LTP residual selected from the FCB 284 for a given sub-frame or frame. The voice encoder 272 can send the index to the voice decoder 274 indicating which LTP residual is selected from the FCB 284 for the given sub-frame or frame. A gain associated with the lowest error can also be selected, and send to the voice decoder 274.
[0081] The voice decoder 274 includes an FCB 294, an LTP engine 292, and a short- term LP engine 290. The FCB 294 has the same LTP residuals (e.g., codevectors) as the FCB 284. The voice decoder 274 can extract an LTP residual from the FCB 294 using the index transmitted to the voice decoder 274 from the voice encoder 272. The extracted LTP residual can be scaled to the appropriate level and filtered by the LTP engine 292 and the short-term LP engine 290 to generate a reconstructed speech signal 275. The LTP engine 292 creates periodicity in the signal associated with the fundamental pitch frequency, and the short-term LP engine 290 generates the spectral envelope of the signal.
[0082] Other linear predictive-based coding systems can also be used to code voice signals, including enhanced voice services (EVS), adaptive multi-rate (AMR) voice coding systems, mixed excitation linear prediction (MELP) voice coding systems, linear predictive coding-10 (LPC-10), among others.
[0083] A voice codec for some applications and / or devices (e.g., Internet-of-Things (IoT) applications and devices) may be needed to deliver higher quality coding of speech signals at low bit-rates, with low complexity, and with low memory requirements. Existing linear predictive-based codecs cannot meet such requirements. For example, ACELP-based coding systems provide high quality, but do not provide low bit-rate or low complexity / low memory. Other linear-predictive coding systems provide low bit-rate andQualcomm Docket No.2403757WO low complexity / low memory, but do not provide high quality. In some cases, machine learning systems (e.g., using a neural network model) can be used to generate reconstructed voice or audio signals. For example, using features extracted from a frame of audio data, a neural network-based voice decoder can generate coefficients for at least one linear filter. The linear filter can then be used to generate a reconstructed signal. However, such a neural network-based voice decoder can be highly complex and resource intensive. For instance, the neural network-based voice decoder will have to perform the operations of a linear predictive filter (LPC), such as the short-term LP engine 280 of FIG. 2C. Such LPC operations can include complex operations that require the use of a large amount of computing resources by the neural network-based voice decoder.
[0084] FIG.3 is a diagram illustrating an example of a voice decoding signal synthesis system 300 utilizing a linear time-varying filter 304 with coefficients generated using a neural network (NN) filter estimator 302 and a separate linear predictive coding (LPC) filter 306. The voice decoding signal synthesis system 300 is configured to decode data of the compressed speech signal to generate a reconstructed speech signal ^̂^^^^^ (also referred to as a synthesized speech sample) for a current time instant n that approximates an original speech signal that was previously compressed by a voice encoder (not shown). The voice encoder can be similar to and can perform some or all of the functions of the voice encoder 252 described above with respect to FIG. 2B, or other type of voice encoder. For example, the voice encoder can include a short-term LP engine, an LTP engine, and an FCB. In another example, the voice encoder can include a magnitude spectrum generator (e.g., Mel-scale magnitude spectrum or full spectrum magnitude), a short-term linear prediction (LP) engine, and a pitch tracker that detects a fundamental pitch harmonic frequency of the speech and pitch correlation.
[0085] The voice encoder can extract (and in some cases quantize) a set of features (referred to as a feature set) from the speech signal, and can send the extracted (and in some cases quantized) feature set to the voice decoding signal synthesis system 300. The features that are computed by the voice encoder can depend on a particular encoder implementation used. Various illustrative examples of feature sets are provided below according to different encoder implementations, which can be extracted by the voice encoder (and in some cases quantized), and sent to the voice decoding signal synthesis system 300. However, one of ordinary skill will appreciate that other feature sets can be extracted by the voice encoder. For example, the voice encoder can extract any set ofQualcomm Docket No.2403757WO features, can quantize that feature set, and can send the feature set to the voice decoding signal synthesis system 300.
[0086] As noted above, various combinations of features can be extracted as a feature set by the voice encoder. For example, a feature set can include one or any combination of the following features: Linear Prediction (LP) coefficients; Line Spectral Pairs (LSPs); Line Spectral Frequencies (LSFs); pitch lag with integer or fractional accuracy; pitch gain; pitch correlation; Mel-scale frequency cepstral coefficients (also referred to as Mel cepstrum) of the speech signal; Bark-scale frequency cepstral coefficients (also referred to as bark cepstrum) of the speech signal; Mel-scale frequency cepstral coefficients of the LTP residual; Bark-scale frequency cepstral coefficients of the LTP residual; a spectrum (e.g., Discrete Fourier Transform (DFT) or other spectrum) of the speech signal; and / or a spectrum (e.g., DFT or other spectrum) of the LTP residual; voicing level of each frequency band of each speech frame; fundamental frequency of pitch harmonics; pitch correlation of each speech frame; time domain pitch lag of each speech frame.
[0087] For any one or more of the other features listed above, the voice encoder can use any estimation and / or quantization method, such as an engine or algorithm from any suitable voice codec (e.g. EVS, AMR, or other voice codec) or a neural network-based estimation and / or quantization scheme (e.g., convolutional or fully-connected (dense) or recurrent Autoencoder, or other neural network-based estimation and / or quantization scheme). The voice encoder can also use any frame size, frame overlap, and / or update rate for each feature. The voice encoder can also include extra redundancies in the features to ensure robustness of operation against packet losses. Examples of estimation and quantization methods for each example feature are provided below for illustrative purposes, where other examples of estimation and quantization methods can be used by the voice encoder.
[0088] As noted above, one example of features that can be extracted from a voice signal by the voice encoder includes linear prediction (LP) coefficients and / or line spectral frequencies (LSFs). Various estimation techniques can be used to compute the LP coefficients and / or LSFs. For example, the voice encoder can estimate LP coefficients (and / or LSFs) from a speech signal using the Levinson-Durbin algorithm. In some examples, the LP coefficients and / or LSFs can be estimated using an autocovariance method for LP estimation. In some cases, the LP coefficients can be determined, and anQualcomm Docket No.2403757WO LP to LSF conversion algorithm can be performed to obtain the LSFs. Any other LP and / or LSF estimation engine or algorithm can be used, such as an LP and / or LSF estimation engine or algorithm from an existing codec (e.g., EVS, AMR, or other voice codec).
[0089] Various quantization techniques can be used to quantize the LP coefficients and / or LSFs. For example, the voice encoder can use a single stage vector quantization (SSVQ) technique, a multi-stage vector quantization (MSVQ), or other vector quantization technique to quantize the LP coefficients and / or LSFs. In some cases, a predictive or adaptive SSVQ or MSVQ (or other vector quantization technique) can be used to quantize the LP coefficients and / or LSFs. In another example, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the LP coefficients and / or LSFs. Any other LP and / or LSF quantization engine or algorithm can be used, such as an LP and / or LSF quantization engine or algorithm from an existing codec (e.g., EVS, AMR, or other voice codec).
[0090] Another example of features that can be extracted from a voice signal by the voice encoder includes pitch lag (integer and / or fractional), pitch gain, and / or pitch correlation. Various estimation techniques can be used to compute the pitch lag, pitch gain, and / or pitch correlation. For example, the voice encoder can estimate the pitch lag, pitch gain, and / or pitch correlation (or any combination thereof) from a speech signal using any pitch lag, gain, correlation estimation engine or algorithm (e.g. autocorrelation- based pitch lag estimation). For example, the voice encoder can use a pitch lag, gain, and / or correlation estimation engine (or algorithm) from any suitable voice codec (e.g. EVS, AMR, or other voice codec). Various quantization techniques can be used to quantize the pitch lag, pitch gain, and / or pitch correlation. For example, the voice encoder can quantize the pitch lag, pitch gain, and / or pitch correlation (or any combination thereof) from a speech signal using any pitch lag, gain, correlation quantization engine or algorithm from any suitable voice codec (e.g. EVS, AMR, or other voice codec). In some cases, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the pitch lag, pitch gain, and / or pitch correlation features.
[0091] Another example of features that can be extracted from a voice signal by the voice encoder includes the Mel cepstrum coefficients and / or Bark cepstrum coefficients of the speech signal, and / or the Mel cepstrum coefficients and / or Bark cepstrumQualcomm Docket No.2403757WO coefficients of the LTP residual. Various estimation techniques can be used to compute the Mel cepstrum coefficients and / or Bark cepstrum coefficients. For example, the voice encoder can use a Mel or Bark frequency cepstrum technique that includes Mel or Bark frequency filter banks computation, filter bank energy computation, logarithm application, and discrete cosine transform (DCT) or truncation of the DCT. Various quantization techniques can be used to quantize the Mel cepstrum coefficients and / or Bark cepstrum coefficients. For example, vector quantization (single stage or multistage) or predictive / adaptive vector quantization can be used. In some cases, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the Mel cepstrum coefficients and / or Bark cepstrum coefficients. Any other suitable cepstrum quantization methods can be used.
[0092] Another example of features that can be extracted from a voice signal by the voice encoder includes the spectrum of the speech signal and / or the spectrum of the LTP residual. Various estimation techniques can be used to compute the spectrum of the speech signal and / or the LTP residual. For example, a Discrete Fourier transform (DFT), a Fast Fourier Transform (FFT), or other transform of the speech signal can be determined. Quantization techniques that can be used to quantize the spectrum of the voice signal can include vector quantization (single stage or multistage) or predictive / adaptive vector quantization. In some cases, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the spectrum. Any other suitable spectrum quantization methods can be used.
[0093] As noted above, any one of the above-described features or any combination of the above-described features can be estimated, quantized, and sent by the voice encoder to the voice decoding signal synthesis system 300 depending on the particular encoder implementation that is used. In one illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, and pitch correlation. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum of the speech signal. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the speech signal. In another illustrative example, the voice encoder can estimate, quantize, and send pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum ofQualcomm Docket No.2403757WO the speech signal. In another illustrative example, the voice encoder can estimate, quantize, and send pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the speech signal. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum of the LTP residual. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the LTP residual.
[0094] The voice decoding signal synthesis system 300 includes a neural network filter estimator 302, a linear time-varying filter 304 generated by the neural network, and a linear predictive coding (LPC) filter 306. The LPC filter 306 can include a time-varying LPC filter. The neural network filter estimator 302 is trained to generate filter coefficients for the linear time-varying filter 304. The neural network model of the neural network filter estimator 302 can include any neural network architecture that can be trained to model the filter coefficients for the linear time-varying filter 304. Examples of neural network architectures that can be included in the neural network filter estimator 302 include a generative neural network (e.g., a generative-adversarial network (GAN)), convolutional neural networks (CNN), an autoencoder, and / or other type(s) of neural network architectures or models.
[0095] The voice decoding signal synthesis system 300 (e.g., the neural network model of the neural network filter estimator 302) can be trained using any suitable neural network training technique. In some examples, the neural network model of the neural network filter estimator 302 can be trained using supervised learning techniques based on backpropagation. For instance, corresponding input and target output pairs can be provided to the neural network filter estimator 302 for training. In one example, for each time instant n, the input to the neural network filter estimator 302 can include log-Mel-frequency spectrum features or coefficients (e.g., 80 log-Mel features ^^^^^, ^^^ 301 shownin FIG.3). In some examples, the target output (or label or ground truth) for training the neural network filter estimator 302 can include the target speech sample ^^^^^^ 305 for the current time instant n, as shown in FIG.3. In such examples, a loss 307 will be computed based on the reconstructed sample ^̂^^^^^ and the target output speech sample ^^^^^^ 305 for time instant n. In some examples, the target output can include a speech signal ^^^^^^ that is generated after passing the target speech through an LPC analysis filter (inverse of LPCQualcomm Docket No.2403757WO filter 306). In this case, the loss will be computed based on the output ^̂^^^^^ generated by the linear time-varying filter generated by the neural network 304 and the target output^^^^^^(e.g., where both are in speech residual domain).
[0096] Backpropagation can be performed to train the neural network filter estimator 302 using the inputs and the target output. Backpropagation can include a forward pass, a loss function, a backward pass, and a parameter update to update one or more parameters (e.g., weight, bias, or other parameter). The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. The process can be repeated for a certain number of iterations for each set of inputs until the neural network filter estimator 302 is trained well enough so that the weights (and / or other parameters) of the various layers are accurately tuned.
[0097] In some aspects, training of the neural network filter estimator 302 and / or training of one or more of the machine learning systems or neural networks described herein can be performed using online training (e.g., in some case on-device training), offline training, and / or various combinations of online and offline training.
[0098] In some cases, online may refer to time periods during which the input data is processed, for instance for performance of generative voice codec processing implemented by the systems and techniques described herein. In some examples, offline may refer to idle time periods or time periods during which input data is not being processed. Additionally, offline may be based on one or more time conditions (e.g., after a particular amount of time has expired, such as a day, a week, a month, etc.) and / or may be based on various other conditions such as network and / or server availability, etc., among various others. In some aspects, offline training of a machine learning model (e.g., a neural network model) can be performed by a first device (e.g., a server device) to generate a pre-trained model, and a second device can receive the trained model from the second device. In some cases, the second device (e.g., a mobile device, an XR device, a vehicle or system / component of the vehicle, or other device) can perform online (or on- device) training of the pre-trained model to further adapt or tune the parameters of the model.
[0099] In some cases, the forward pass can include passing the input data (e.g., the log-Mel-frequency spectrum features or coefficients, such as the 80 log-Mel features ^^^^^, ^^^301 shown in FIG.3) through the neural network filter estimator 302. The weights of theQualcomm Docket No.2403757WO neural network model are initially randomized before the neural network filter estimator 302 is trained. For a first training iteration for the neural network filter estimator 302, the output will likely include values that do not give preference to any particular output due to the weights being randomly selected at initialization. With the initial weights, the neural network filter estimator 302 is unable to determine low level features and thus cannot make an accurate estimation of the filter coefficients for the linear time-varying filter 304. A loss function can be used to analyze the loss 307 (or error) in the reconstructed or synthesized sample output ^̂^^^^^. Any suitable loss function definition can be used. One example of a loss function includes a mean squared error (MSE). The MSE is defined as^^௧^௧^^ ൌ ∑^ ଶ^^^^^^^^^^^^^ െ ^^^^^^^^^^^^^ଶ, which calculates the sum of one-half times the actualanswer the predicted (output) answer squared. The loss can be set to be equal tothe value Other loss functions may include a difference of magnitude spectrums between the target and output signals, where the difference may be computed as absolute difference, squared difference, or logarithmic difference between the magnitude spectrum of each speech frame, and then aggregated over all speech frames.
[0100] The loss (or error) will be high for the first training iterations since the actual values will be much different than the predicted output. The goal of training is to minimize the amount of loss so that the predicted output is the same as the training label. The neural network filter estimator 302 can perform a backward pass by determining which inputs (weights) most contributed to the loss of the network, and can adjust the weights so that the loss decreases and is eventually minimized. A derivative of the loss with respect to the weights (denoted as dL / dW, where W represents the weights at a particular layer) can be computed to determine the weights that contributed most to the loss of the network. After the derivative is computed, a weight update can be performed by updating all the weights of the filters. For example, the weights can be updated so that they change in theopposite direction of the gradient. The weight update can be denoted as ^^ ൌ ^^ௗ^ ^െ ^^ௗ^,where w denotes a weight, widenotes the initial weight, and η denotes alearning rate can be set to any suitable value, with a high learning rate including larger weight updates and a lower value indicating smaller weight updates. In some examples, to train the neural network filter estimator 302, a multi-resolution STFT loss ^^ோand adversarial losses ^^ீand ^^^can be computed from ^^^^^^ and ^^^^^^. Because linear time- varying filters are fully differentiable, gradients can propagate back to the neural networkQualcomm Docket No.2403757WO filter estimator 302.
[0101] Using the filter coefficients generated by the neural network filter estimator 302, the linear time-varying filter 304 can process an excitation signal 303 to generate another signal ^̂^^^^^. The signal ^̂^^^^^ can be used as an excitation signal to excite the LPC filter 306. The linear time-varying filter 304 is a linear filter, which preserves the linearity property between inputs and outputs. For instance, a linear filter is associated with amapping ℒ:ℝℤ → ℝℤ, ^^^^^^ → ^^^^^^ ൌ ℒ^^^^^^^^ that has the following property: for any^^^^^^^, ^^ଶ^^^^ any ^^, ^^ ∈ ℝ, ℒ^^^ ^^^^^^^ ^ ^^ ^^ଶ^^^^^ ൌ ^^ ℒ^^^^^^^^^ ^ ^^ ℒ^^^ଶ^^^^^.In one if input1 produces output1 and input2 produces output2, then a combined input of (input1+input2) will produce output = output1+output2. The time- varying nature of the linear time-varying filter 304 indicates that the filter response depends on the time of excitation of the linear time-varying filter 304 (e.g., a new set of coefficients used to filter each frame (block of time) of input at the time of excitation). In some cases, time-varying linear filters can be characterized by the set of impulseresponses at each time lag ℎ^^^^^ ൌ ℒ^^^^^^ െ ^^^^ for each ^^ ∈ ℤ. In some examples, theoutput of a time varying linear filter is ℒ^^^^^^^^ ൌ ∑^ ^^^^^^ ⋆ ℎ^^^^^ (where ⋆ is aconvolutional operator) or some heuristic combination of filter input and impulse responses, e.g. overlap-add on windowed and filtered signal segments, etc.
[0102] The LPC filter 306 can use the signal ^̂^^^^^ as input to generate the reconstructed or synthesized speech sample ^̂^^^^^ for the current time instant n. The LPC filter 306 is a linear filter and in some cases is time varying, as defined above with respect to the linear time-varying filter 304. The LPC filter 306 can be a form of time-varying filter used for processing of speech. The LPC filter 306 includes filter coefficients for each speech frame that can be computed using the autocorrelation of a speech or audio signal.
[0103] In some examples, the LPC filter 306 can be used to model the spectral shape (or phenome or envelope) of the speech signal. For example, at the voice encoder, a signal ^^^^^^ can be filtered by an autoregressive (AR) process synthesizer to obtain an AR signal. As described above, a linear predictor can be used to predict the AR signal (which can be denoted as prediction ^̂^^^^^) as a linear combination of the previous m samples as follows: ^ ^^^^^^^Qualcomm Docket No.2403757WO
[0104] where the ^^^^terms (^^^^, ^^^ଶ, … ^^^^) are estimates of the AR parameters (also referred to as LP coefficients). A residual signal can be the difference between the original AR signal and the predicted AR signal represented as prediction ^̂^^^^^ (e.g., the difference between the actual sample and the predicted sample). At the voice encoder, the linear prediction coding can be used to find the best linear prediction coefficients for minimizing a quadratic error function, and thus the error. The linear prediction process removes the short-term correlation from the speech signal. The linear prediction coefficients are an efficient way to represent the short-term spectrum of the speech signal. At the voice decoding signal synthesis system 300, the LPC filter 306 determines the prediction ^̂^^^^^ for the current sample n using computed or received coefficients and A transfer function ^^^^^^^. For instance, in some examples, the LPC filter coefficients are received from the encoder. In other examples, the voice decoding signal synthesis system 300 can derive the LPC filter coefficients, such as using other features (e.g., Mel spectrum features) sent by an encoder to the voice decoding signal synthesis system 300. For instance, the voice decoding signal synthesis system 300 can use Mel spectrum features 301 to derive the LPC filter coefficients for the LPC filter 306. The LPC filter 306 can determine the final reconstructed (or predicted) sample ^̂^^^^^ using the output ^̂^^^^^ from the linear time- varying filter 304 (for the current sample n).
[0105] In some examples, further components can be used along with the neural network filter estimator 302 and the linear time-varying filter 304, such as an impulse train generator 314 and a random noise generator 316. In some examples, the linear time- varying filter 304 can include a harmonic linear time-varying filter 318 and a noise linear time-varying filter 320. A voice encoder can include a pitch tracker 310 and a feature extraction engine 312. In some examples, the feature extraction engine 312 can be the same as or similar to the feature generator 200 of FIG. 2A. In some cases, the original speech signal ^^ and reconstructed signal ^^ are divided into non-overlapping frames with frame length L. The term ^^ can be defined as a frame index, the term ^^ can be defined as a discrete time index, and the term ^^ can be defined as a feature index. The total numberof frames ^^ and total number of sampling points ^^ may follow ^^ ൌ ^^ ൈ ^^. In ^^^, ^^, ℎ^,ℎ^, 0 ^ ^^ െ 1. The terms ^^, ^^, ^^, ^^, ^^^, ^^^ are finite duration signals, in which 0 ^ ^^ ^^^ െ 1. Impulse responses ℎ^, and ℎ^ may be infinitely long, in which ^^ ∈ ℤ. Impulseresponse h may be causal, in which ^^ ∈ ℤ E Z and ^^ ^ 0.Qualcomm Docket No.2403757WO
[0106] To perform the speech synthesis process, the impulse train generator 314 can generate an impulse train ^^^^^^ from a frame-wise fundamental frequency ^^^^^^^output by the pitch tracker 310. In one illustrative example, the impulse train generator 314 can generate alias-free discrete time impulse trains using additive synthesis. For instance, as illustrated in equation (1) below, the impulse train generator 314 can use a low-passed sum of sinusoids to generate an impulse train: ^2^^^^^^^^^ ^ ^^^^^ ൌ ^^^^^^ ^^௧ ^2^^^^ ^^^^^^^^^^^^ ,^^^^^^ ^ 1^^^^ ^
[0107] wherehold or linearinterpolation, ^^^^^^ ൌ ^^^^^ / ^^^^, and ^^^ is the sampling rate. In some cases, thecomputationally complexity of additive synthesis can be reduced with approximations. For example, the impulse train generator 314 or other component (e.g., a processor) of the voice decoding signal synthesis system 300 can round the fundamental periods to the nearest multiples of the sampling period. In such an example, the discrete impulse train is sparse. The impulse train generator 314 can then generate the impulse train sequentially (e.g., one pitch mark at a time).
[0108] The pitch tracker 310 can process the input ^^^^^^ for the time instant n to generate the frame-wise fundamental frequency ^^^^^^^ output, which is provided to and processed by the impulse train generator 314 of the voice decoding signal synthesis system 300. The random noise generator 316 of the voice decoding signal synthesis system 3400 can sample a noise signal ^^^^^^ from a Gaussian distribution.
[0109] The neural network filter estimator 302 can estimate impulse responsesℎ^ ^^^,^^^ and ℎ^ ^^^,^^^ for each frame, given the log-Mel spectrogram ^^^^^, ^^^ extractedfrom the input ^^^^^^ by the feature extraction engine 312 of the encoder. In some aspects,complex cepstrums (ℎ^^ and ℎ^^) can be used as the internal description of impulseresponses (ℎ^and ℎ^) for the neural network filter estimator 302. Complex cepstrums describe the magnitude response and the group delay of filters simultaneously. The group delay of filters affects the timbre of speech. In some cases, instead of using linear-phase or minimum-phase filters, the neural network filter estimator 302 can use mixed-phase filters, with phase characteristics learned from the dataset.Qualcomm Docket No.2403757WO
[0110] In some examples, the length of a complex cepstrum can be restricted, essentially restricting the levels of detail in the magnitude and phase response. Restricting the length of a complex cepstrum can be used to control the complexity of the filters. In some cases, the neural network filter estimator 302 can predicts low-frequency coefficients, in which the high-frequency cepstrum coefficients can be set to zero. In one illustrative example, two 10 millisecond (ms) long complex cepstrums are predicted in each frame. In some cases, the neural network filter estimator 302 can use a discrete Fourier transform (DFT) and an inverse-DFT (IDFT) to generate the impulse responses ℎ^and ℎ^. In some cases, the neural network filter estimator 302 can approximate an infinite impulse response (IIR) (ℎ^^^^,^^^ and ℎ^^^^,^^^) using finite impulse responses (FIRs). The DFT size can be set to at least a threshold size (e.g., N=1024) to avoid aliasing.
[0111] Using the impulse response ℎ^^^^,^^^, the harmonic LTV filter 318 can filter the impulse train ^^^^^^ from the impulse train generator 314 to generate a harmonic component ^^^^^^^. Using the impulse response ℎ^^^^,^^^, the noise LTV filter 320 can filter the noise signal ^^^^^^ to generate a noise component ^^^^^^^. The voice decoding signal synthesis system 300 can combine (e.g., by summing / adding or otherwise combining) the output of the harmonic LTV filter 318 (the harmonic component ^^^^^^^) and the output of the noise LTV filter 320 (the noise component ^^^^^^^) can be combined (e.g., summed or otherwise combined) to obtain the excitation signal ^^^^^^.
[0112] As noted previously, systems and techniques are described herein that can be used to provide a bit-efficient generative audio codec system including a machine learning (e.g., neural network)-based encoder and a machine learning (e.g., neural network)-based decoder. In some examples, the bit-efficient generative audio codec can be a hybrid DSP- ML codec, where the encoder and / or decoder are configured to perform combinations of non-ML-based DSP processing operations and ML-based processing operations to encode and decode audio data, respectively.
[0113] FIG.4 is a block diagram illustrating an example of an audio codec system 400 that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer. In some aspects, the audio codec system 400 can also be referred to as a voice coding system (e.g., vocoder), a signal synthesis system, a voice coding signal synthesis system, etc. In one illustrative example, the audio codec system 400 can be aQualcomm Docket No.2403757WO generative audio codec (e.g., a generative voice codec) including one or more feedback recurrent autoencoders (FRAEs) that may be used to generate encoded audio data, and one or more neural synthesizers that may be used to generate reconstructed audio data based on the encoded audio data.
[0114] For example, the audio codec system 400 (e.g., a generative voice codec) can include an encoder 405 (e.g., a transmitter) and a decoder 410 (e.g., a receiver). The encoder 405 can be used to generate encoded features that are transmitted to the decoder 410 over a channel 440. For example, the encoder 405 can receive audio 415 and generate encoded features corresponding to the audio 415. The encoded features can be generated using a feedback recurrent autoencoder (FRAE) 425 included in the encoder 405. In some examples, the encoder 405 can include one or more FRAEs, where each FRAE of the one or more FRAEs is configured to generate a corresponding one or more encoded features associated with the audio 415. In some aspects, the audio 415 may include and / or comprise a speech signal, a voice signal, etc.
[0115] The decoder 410 can receive encoded features (e.g., from the encoder 405, over the channel 440) and generate reconstructed audio 455 based at least in part on the encoded features received from the encoder 405. The reconstructed audio 455 may include and / or comprise a reconstructed speech signal, a reconstructed voice signal, etc. The reconstructed audio 455 can be generated using a neural network-based speech synthesizer 450 (e.g., also referred to as a neural speech synthesizer and / or a neural synthesizer) included in the decoder 410. The reconstructed audio 455 generated using the neural speech synthesizer 450 may also be referred to as synthesized speech. In some examples, the decoder 410 can include an FRAE decoder 445, which can be used to decode the encoded features received from the encoder 405. In some aspects, the FRAE decoder 445 can be the same as or similar to a decoder implemented by the FRAE autoencoder 425 included in the encoder 405. In some examples, the decoder 410 can include one or more FRAE decoders 445, which can correspond to one or more FRAE autoencoders 425 included in the encoder 405. For example, the number of FRAE decoders 445 included in the decoder 410 can be equal to the number of FRAE autoencoders 425 included in the encoder 405.
[0116] In some examples, the audio 415 includes speech. The audio 415 may be an example of the speech signal 251 of FIG. 2B, the speech signal 271 of FIG. 2C, and / orQualcomm Docket No.2403757WO the speech signal 303 of FIG.3, etc. The encoder 405 of the generative voice codec system 400 can be configured to generate one or more encoded representations of the input audio 415. For example, the one or more encoded representations can include spectral envelope information associated with the audio 415, and / or can include pitch information associated with the audio 415, etc. In some aspects, the encoder 405 can extract spectral envelope features from the audio 415 using spectral envelope feature extraction 420, and can encode the extracted spectral envelope features using the FRAE 425. The encoded spectral features z can be transmitted from the FRAE 425 to the decoder 410, using the channel 440. In some examples, the encoder 405 can extract pitch information from the audio 415 using pitch extraction engine 430, and can perform quantization 435 to generate quantized pitch information associated with the audio 415. The quantized pitch information can be transmitted from the pitch quantizer 435 to the decoder 410, using the channel 440.
[0117] In some aspects, the encoder 405 can be configured to extract spectral features (e.g., spectral envelope features) from the audio 415 using the spectral envelope feature extraction engine 420. In one illustrative example, the spectral envelope feature extraction engine 420 can perform a cepstrum computation to determine cepstrum information and / or one or more cepstral coefficients corresponding to the audio 415. In some cases, the spectral features extracted from the audio 415 using the spectral envelope feature extraction 420 can include mel-frequency cepstral coefficients (MFCC), such as MFCC- 24 features (e.g., 24-dimensional MFCC features). In some examples, the spectral envelope feature extraction 420 applies one or more mel-scaled filter(s) (e.g., a mel-scaled filterbank), a logarithmic compression, and / or a discrete cosine transform (DCT) to the audio 415 (e.g., and / or to a magnitude spectrum associated with the audio 415). The encoder 405 processes the extracted spectral features using a feedback recurrent autoencoder (FRAE) 425 to generate encoded features z. The FRAE 425 includes a decoder and an encoder. The FRAE 425 can implement feedback of state information h between the decoder and the encoder included in the FRAE 425. For example, the state information h can be determined or obtained at the decoder of the FRAE 425, and feedback of the state information h can be performed for the decoder of the FRAE 425 (e.g., the state information h is fed back to the decoder of the FRAE 425) and for the encoder of the FRAE 425 (e.g., the state information h is fed back from the decoder of the FRAE 425 to the encoder of the FRAE 425).Qualcomm Docket No.2403757WO
[0118] In some examples, the spectral envelope feature extraction engine 420 can be configured to generate and / or determine (e.g., extract) one or more types of spectral envelope features. For example, the extracted spectral envelope features may include cepstrum information and / or cepstral coefficients. In some cases, cepstral liftering can be performed based on DCT truncation to exclude or remove pitch information and only capture spectral envelope information in the extracted cepstrum or cepstral coefficients. In some examples, the extracted spectral envelope features can include companded (e.g., log) filterbank energies associated with the input audio 415. For example, companded filterbank energies can be determined based on applying an inverse DCT (e.g., IDCT) to a liftered cepstrum. In some cases, the extracted spectral envelope features can include filterbank energies determined based on uncompanding (e.g., exp) the companded filterbank energies. In some examples, the extracted spectral envelope features can include a full resolution spectrum (e.g., DFT domain with companded (e.g., log, linear amplitude, etc.) information, etc.). The full resolution spectrum can be smoothed to the envelope of the spectrum, for example based on interpolating the companded or uncompanded filterbank energies. In some aspects, the extracted spectral envelope features can include one or more linear prediction (LP) coefficients, for example determined based on processing one or more speech frames of the input audio 415 using the Levinson-Durbin algorithm (e.g., autocorrelation technique), and / or using a covariance technique. In some examples, the extracted spectral envelope features can include one or more of line spectral frequency (LSF) information and / or line spectral pair (LSP) information. In some examples, the spectral envelope feature extraction engine 420 can be configured to extract a first set of one or more spectral features comprising cepstrum information and / or cepstral coefficients, and to extract a second set of one or more spectral features comprising LP coefficients, LSF information, and / or LSP information. For example, the second set of one or more spectral features can include LP coefficients (LPCs) estimated by the spectral envelope feature extraction engine 420 based on the Levinson-Durbin algorithm or other speech autocorrelation techniques. In some examples, the spectral envelope feature extraction engine 420 can estimate one or more LP coefficients from spectral envelope features (e.g., the first set of one or more spectral features comprising cepstrum information and / or cepstral coefficients, etc.), using a neural network-based technique or a DSP-based technique. In some aspects, LSP coefficients can be determined by the spectral envelope feature extraction engine 420Qualcomm Docket No.2403757WO based on autocorrelation of the speech signal 415. In some examples, the spectral envelope feature extraction engine 420 can determine one or more LSPs, and can convert the one or more LSPs to a corresponding one or more LSFs.
[0119] In some aspects, the encoded features or encoded feature information transmitted from the generative voice codec system encoder 405 can comprise a latent representation z between the encoder of the FRAE 425 and the decoder of the FRAE 425. The encoder 405 can be configured to pass (e.g., transmit) the encoded features z through a channel 440 to the decoder 410.
[0120] The encoder 405 also processes the audio 415 using a pitch extraction engine 430 to extract pitch information from the audio 415. In some aspects, the audio 415 can be processed in parallel by the spectral envelope feature extraction engine 420 and the pitch extraction engine 430. In some cases, the encoder 405 can include a denoiser 418 that processes the input audio 415 and provides a de-noised audio to the spectral envelope feature extraction engine 420 and / or to the pitch extraction engine 430.
[0121] Based on the audio 415, the pitch extraction engine 430 can generate one or more types of pitch information. For example, the pitch extraction engine 430 can output pitch information in the frequency-domain (e.g., f0pitch information of the audio 415, in units of Hertz (Hz)) and / or can output pitch information in the time-domain (e.g., pitch lag in samples, with or without a fractional component, and / or pitch lag in milliseconds). In some cases, the pitch extraction engine 430 can be used to generate pitch estimation indicative of a pitch lag (or pitch delay) from the audio 415, and / or to identify a pitch correlation from the audio 415. In some aspects, the pitch information generated by the pitch extraction engine 430 can include a pitch lag and a pitch correlation.
[0122] The encoder 405 can be configured to processes the pitch information of the audio 415 (e.g., pitch, pitch lag, and / or pitch correlation, determined using the pitch extraction engine 430) using a quantizer 435 Q() to generate a quantized pitch signal. The encoder 405 passes the quantized pitch signal through the channel 440 to the decoder 410. In some examples, the quantizer 435 Q() may also be referred to as a pitch quantizer. In some aspects, the quantizer 435 Q() can perform vector quantization (VQ) to generate the quantized pitch signal for transmission to the decoder 410 over the channel 440. In some examples, the quantizer 435 Q() can perform single-stage VQ and / or can perform multi- stage VQ (e.g., MSVQ). In some cases, the quantizer 435 Q() can be implemented usingQualcomm Docket No.2403757WO one or more FRAEs. In some aspects, the quantizer 435 Q() can be configured to implement forward error correction (FEC) for the quantized pitch signal that is transmitted to the decoder 410 over the channel 440. For example, the quantizer 435 Q() can implement multiple description coding (MDC) for the transmission of the quantized pitch signal over the channel 440, can implement full-redundancy FEC for the transmission of the quantized pitch signal over the channel 440, etc.
[0123] The generative voice codec system decoder 410 can be configured to receive the encoded features z from the FRAE 425 included in the generative voice codec system encoder 405. For example, the decoder 410 can receive the encoded features z via the channel 440. The decoder 410 decodes the encoded features z using a FRAE 445 to generate decoded features. In some examples, the FRAE 445 of the decoder 410 includes only a decoder, without an encoder. In some examples, the FRAE 445 of the decoder 410 can include an encoder. In some examples, the decoded features generated by the FRAE 445 include mel-frequency cepstral coefficients (MFCC), such as MFCC-24 features (e.g., 24-dimensional MFCC features).
[0124] The decoded features determined using the FRAE decoder 445 can be provided to a neural speech synthesizer 450, which can be configured to generate a reconstructed audio 455 (e.g., synthesized speech) based at least in part on the decoded spectral envelope features from the FRAE decoder 445. In one illustrative example, the neural speech synthesizer 450 receives as input the decoded spectral envelope features (e.g., from the FRAE decoder 445) and the received pitch encoding information (e.g., the quantized pitch signal received by the decoder 410 over the channel 440 and from the pitch quantizer 435 of the encoder 405). In some aspects, the neural speech synthesizer 450 can generate the reconstructed audio 455 (e.g., synthesized speech) based on processing the decoded spectral envelope features along with a pitch signal (e.g., the quantized pitch signal or a reconstructed variant thereof).
[0125] For example, the decoder 410 can receive the quantized pitch signal (e.g., indicative of the pitch, the pitch lag, and / or the pitch correlation of the audio 415) from the encoder 405 via the channel 440. In some examples, the decoder 410 passes the quantized pitch signal to the neural speech synthesizer 450, and the neural speech synthesizer 450 processes the decoded spectral envelope features along with the quantized pitch signal to generate reconstructed audio 455 (e.g., synthesized speech).Qualcomm Docket No.2403757WO
[0126] In some cases, the decoder 410 can receive the quantized pitch signal (e.g., indicative of the pitch, the pitch lag, and / or the pitch correlation of the audio 415) from the encoder 405 via the channel 440, and can perform dequantization of the quantized pitch signal to obtain a reconstructed pitch signal. The reconstructed pitch signal can be passed to the neural speech synthesizer, and used in combination with the decoded spectral envelope features from the FRAE decoder 445 to generate the reconstructed audio 455 (e.g., synthesized speech). In some examples, the decoder 410 can include a reconstruction engine that is separate from the neural speech synthesizer 450. The reconstruction engine processes the quantized pitch signal to reconstruct the pitch signal (e.g., including the pitch, the pitch lag, and / or the pitch correlation) before the neural speech synthesizer 450 using the reconstructed pitch signal (e.g., the pitch, the pitch lag, and / or the pitch correlation) to generate the reconstructed audio 455.
[0127] In some examples, the output of the neural speech synthesizer 450 can be provided to one or more linear predictive coding (LPC) layers 452 included in the decoder 410. The LPC 452 can be used to perform linear prediction analysis and / or linear prediction synthesis, based on the output of the neural speech synthesizer 450. For example, the LPC 452 can perform linear prediction based on the output of the neural speech synthesizer 450, to generate the reconstructed audio 455 (e.g., synthesized speech). In some aspects, the decoder 410 does not include the LPC 452, and the neural speech synthesizer 450 can be trained and / or configured to generate the reconstructed audio 455 directly (e.g., the output of the neural speech synthesizer 450 can be the reconstructed audio 455).
[0128] In some aspects, the LPC layers 452 can be used to implement a linear prediction (LP) synthesis filter, based on: ^ ^^^^^^ ^^^ ^ ^^^^^^
[0129] Here, ^^^^^^filter (e.g., the input signal to the LPC layers 452), ^^^^^^represents the output signal (e.g., the output of the LPC layers 452, which can be the reconstructed audio 455), ^^^represents the linear prediction coefficients associated with implementing the LP synthesis filter, and ^^ is a value corresponding to the LP filter order. For example, the LP filter order p can have a value of 16 or 12, etc., among various other LP filter order values.Qualcomm Docket No.2403757WO
[0130] In some aspects, the LP synthesis filter associated with the LPC layers 452 can be implemented in the time domain or in the frequency domain. For example, in the time domain, the LP synthesis filter can be implemented based on a difference equation. In some cases, the LP synthesis filter can be implemented in the time domain by convolving the input signal x[n] with the LP filter impulse response or an approximation of the LP filter impulse response. For example, the LP filter can be associated with an infinite impulse response (IIR), which can be approximated using a finite segment of the IIR (e.g., such as the first N samples of the IIR for a configured integer value of N, etc.). The LP synthesis filter associated with the LPC layers 452 can be implemented in the frequency domain based on multiplying the FFT of the input signal x[n] with the frequency response of the LP filter, and then determining the IFFT of the result to convert the output of the frequency domain LP synthesis filter into a time domain signal corresponding to the reconstructed audio 455.
[0131] In some examples, the LP coefficients ^^^can be estimated from the decoded spectral envelope features obtained using the FRAE decoder 445. In some cases, the spectral envelope features can be LP coefficients (e.g., determined by the spectral envelope feature extraction engine 420 based on using the Levinson-Durbin algorithm or other autocorrelation technique, or using a covariance technique, to process the input speech frames of the audio 415), and the decoded LP coefficients on the decoder side can be used for the LP filter implemented by the LPC layers 452.
[0132] In some examples, the extracted spectral envelope features may be LSFs or LSPs, and the decoded LSFs or LSPs obtained using the FRAE decoder 445 can be converted to the corresponding LP coefficients ^^^. In some cases, the LP coefficients ^^^of the LPC layers 452 can be estimated from the spectral envelope features using a neural network-based and / or a DSP-based technique. For example, the extracted spectral envelope features may comprise a cepstrum, and a DSP-based technique can be used to convert the decoded cepstrum into filterbank energies, and subsequently interpolate the filterbank energies to estimate an FFT square magnitude at all FFT bins (e.g., including bins beyond the centers of filterbank filters). After estimating the FFT square magnitude based on the interpolated filterbank energies, the DSP-based technique can include applying an IFFT to obtain an estimate of the autocorrelation sequence, and utilizing the Levinson-Durbin algorithm on the estimated autocorrelation sequence to obtain the LP coefficients ^^^.Qualcomm Docket No.2403757WO
[0133] In some examples, the spectral envelope feature extraction 420 can be performed based at least in part on spectral feature learning. For example, spectral feature learning can be performed by one or more machine learning models configured and / or trained to generate as output the one or more extracted spectral envelope features. In some cases, a cepstrum or cepstrum information may be an optional input to a machine learning spectral feature learning model used to implement the spectral envelope feature extraction 420. In some cases, the spectral envelope feature extraction 420 can be implemented using one or more machine learning spectral feature learning models or engines, which can be jointly trained with the neural speech synthesizer 450, an NHV used to implement the neural speech synthesizer (e.g., such as the NHV-based neural speech synthesizer 950 of FIG. 9), an LPC network used to implement the neural speech synthesizer (e.g., such as the LPC network-based neural speech synthesizer 1050 of FIG.10), etc.
[0134] In some cases, joint training of the spectral envelope feature extraction 420 and the neural speech synthesizer 450 can be performed based on a short-time Fourier transform (STFT) loss and / or STFT loss function. In some aspects, joint training of the spectral envelope feature extraction 420 and the neural speech synthesizer 450 can be performed based on an adversarial and / or generative adversarial network (GAN) loss. For example, joint training can be performed using a multi-resolution STFT loss, based on calculating STFT amplitude spectrograms from ground truth speech information and the synthesized speech information 455. The multi-resolution STFT loss can then be determined as a sum of mean absolute error and mean absolute error in the log domain. The STFT amplitude spectrograms can be calculated at different window lengths to capture representations of the error and / or loss at different time and / or frequency resolutions. In some aspects, joint training based on adversarial or GAN loss can be used to learn temporal fine structures in speech signals. An adversarial or GAN loss can be used to match the distribution of real speech and the synthesized speech 455. For example, the neural speech synthesizer 450 (and / or an NHV-based neural speech synthesizer) can be used as the generator network. A separate discriminator network can be used to train based on the adversarial loss, with both the NHV and the discriminator jointly trained using adversarial training techniques. In some examples, the discriminator network can be implemented as a classifier configured to classify whether an input audio represents real speech or synthesized speech. The generator network can attempt to fool the discriminator network to classify synthesized speech as real speech. Over the course ofQualcomm Docket No.2403757WO training, the generator network improves and begins producing (e.g., generating or synthesizing) speech that is very similar to the real speech. In some cases, the discriminator network can be implemented using WaveNet.
[0135] The neural speech synthesizer 450 can be implemented using one or more trained neural networks. For example, the one or more trained neural networks can be trained to perform speech synthesis and / or signal synthesis. In some examples, the neural speech synthesizer 450 can be implemented using or based on an LPCNet machine learning architecture, a WaveNet machine learning architecture, a WaveRNN machine learning architecture, etc. In one illustrative example, the neural speech synthesizer 450 can be provided as a neural homomorphic vocoder (NHV).
[0136] For example, FIG.9 illustrates an example of an audio codec system 900 (e.g., generative voice codec) that includes a decoder 910 configured to implement an NHV- based neural speech synthesizer. In some examples, an NHV is a type of neural vocoder that can synthesize speech with source-filter models controlled by one or more neural networks. An NHV-based neural vocoder may include one or more neural networks in a source-filter model that can synthesize speech based on filtering impulse trains and noise with linear time-varying (LTV) filters, with the one or more neural networks used to control the LTV filters by estimating complex cepstrums of time-varying impulse responses given acoustic features. Traditional or non-neural vocoders may operate based on decomposing speech into various parameters such as pitch, timbre, rhythm, etc., which can subsequently be manipulated and resynthesized to generate a desired output audio or voice signal. Neural vocoders (e.g., such as NHVs, etc.) can apply transformations and manipulations to speech signals directly within a learned feature space of the neural network. The learned feature space used by neural vocoders may capture more complex relationships and characteristics of speech than non-neural network-based signal processing and / or vocoder techniques. For example, neural vocoders can be trained to learn a mapping between raw speech waveforms and the spectral or cepstral representations of the speech. An NHV system can apply one or more homomorphic processing techniques within the learned space corresponding to the mapping.
[0137] As noted above, FIG. 9 is a block diagram illustrating an example of an audio codec system 900 (e.g., a generative voice codec, a voice coding signal synthesis system, etc.) that includes a decoder 910 that can be used to generate reconstructed audio (e.g.,Qualcomm Docket No.2403757WO synthesized speech) using an NHV-based neural speech synthesizer 950, in accordance with some examples. In some aspects, the decoder 910 may also be referred to as an NHV synthesizer decoder. In some examples, the audio codec system 900 of FIG.9 can be the same as or similar to the audio codec system 400 of FIG. 4. In some cases, the channel 940 of FIG. 9 can be the same as or similar to the channel 440 of FIG. 4, and may be associated with an encoder that is the same as or similar to the encoder 405 of FIG.4, etc.
[0138] In some aspects, the NHV speech synthesizer 950 of FIG.9 can be the same as or similar to the NHV system 300 of FIG.3. For example, the neural filter estimator 952 can be the same as or similar to the NN filter estimator 302 of FIG.3. In some cases, the noise generator 970 can be the same as or similar to the random noise generator 316 of FIG.3, and the noise LTVF 975 can be the same as or similar to the noise LTV filter 320 of FIG. 3. In some examples, the pulse train generator 960 can be the same as or similar to the impulse train generator 314 of FIG.3, and the harmonic LTVF 965 can be the same as or similar to the harmonic LTV filter 318 of FIG.3.
[0139] In some aspects, the decoder 910 can include a neural speech synthesizer 950 that is the same as or similar to the neural speech synthesizer 450 included in the decoder 410 of FIG. 4. For example, the neural speech synthesizer 950 can receive a first input comprising decoded spectral envelope features (e.g., determined by an FRAE decoder 945 that is the same as or similar to the FRAE decoder 445 of FIG.4). The neural speech synthesizer 950 can receive a second input comprising a received pitch encoding (e.g., received pitch information), that is the same as or similar to the received pitch encoding provided to the neural speech synthesizer 450 of FIG.4.
[0140] In one illustrative example, the NHV synthesizer decoder 910 can include the FRAE decoder 945, which may be the same as or similar to the FRAE decoder 445 of FIG.4, and may include a pitch de-quantization engine 938 (e.g., de Q()). The pitch de- quantization engine 938 can be used to process a received pitch encoding obtained by the decoder 910 over the channel 940. For example, the received pitch encoding information can be quantized pitch information or a quantized pitch signal transmitted to the decoder 910 over the channel 940 by a corresponding encoder (e.g., an encoder associated with the decoder 910, which may be the same as or similar to the encoder 405 of FIG.4, etc.). In some aspects, the pitch dequantization engine 938 can generate reconstructed pitch information based on performing a codebook lookup to dequantize the received pitchQualcomm Docket No.2403757WO encoding obtained from the channel 940 (e.g., to dequantize the quantized pitch encoding received by the decoder 910 over the channel 940).
[0141] The dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 938 can include pitch information in the frequency-domain (e.g., f0 pitch frequency information in Hz), pitch information in the time-domain (e.g., pitch lag or pitch delay information in samples, milliseconds, etc.), and / or pitch correlation information. In some aspects, the dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 938 can include information indicative of a voiced or unvoiced (V / UV) classification.
[0142] In one illustrative example, the neural speech synthesizer 950 can be implemented as a neural homomorphic vocoder (NHV) speech synthesizer and / or an NHV-based neural speech synthesizer. For example, the neural speech synthesizer 950 can include a neural filter estimator 952 configured to generate respective filter specification or filter configuration information to parameterize one or more linear time- varying (LTV) filters of the NHV speech synthesizer 950. In some aspects, the neural filter estimator 952 can be trained to generate respective filter coefficients for a noise linear time-varying filter (LTVF) 975 and to generate respective filter coefficients for a harmonic LTVF 965.
[0143] The neural network model of the neural filter estimator 952 can include any neural network architecture that can be trained to model the filter coefficients for the noise LTVF 975 and the harmonic LTVF 965 (e.g., and / or that can be trained to model the filter coefficients for one or more additional filters implemented by the NHV speech synthesizer 950, in either the time-domain, the frequency-domain, or combinations thereof). Examples of neural network architectures that may be included in the neural network filter estimator 952 can include a generative neural network (e.g., a generative- adversarial network (GAN)), convolutional neural networks (CNN), an autoencoder, and / or other type(s) of neural networks, etc.
[0144] In some examples, the neural filter estimator 952 can be configured to receive a stream of decoded spectral envelope features from the FRAE decoder 945. Based on the decoded spectral envelope features, the neural filter estimator 952 can generate corresponding filter characterization parameters for each filter of one or more filters included in the NHV speech synthesizer 950. The respective filter characterizationQualcomm Docket No.2403757WO parameters can be output from the neural filter estimator 952 and used to parameterize and / or configure corresponding learned filters (e.g., learned linear filters, etc.) for each respective set of filter characterization parameters. For example, the neural filter estimator 952 can generate a first set of filters corresponding to the noise LTVF 975, can generate a second set of filter characterization parameters corresponding to the harmonic LTVF 965, etc. The one or more filters included in the NHV speech synthesizer 950 (e.g., the noise LTVF 975, the harmonic LTVF 965, etc.) can be specified in various different forms. For example, the noise LTVF 975, the harmonic LTVF 965, and / or various other filters that may be included in the NHV speech synthesizer 950 can be specified or characterized (e.g., using the filter characterization parameters determined by the neural filter estimator 952) as cepstrums, as frequency responses, as time-domain impulse responses, as difference equation coefficients, etc.
[0145] In some aspects, the neural filter estimator 952 can be configured to receive as input the decoded spectral envelope features from the FRAE decoder 945, and may additionally receive as input at least a portion of the dequantized pitch information generated by the pitch dequantization engine 938. For example, the neural filter estimator 952 may receive as input the decoded spectral envelope features and pitch dequantization information (e.g., such as f0 pitch information in the frequency domain, pitch lag or pitch delay in the time domain (e.g., in units of samples or milliseconds, etc.), etc.). In some examples, the pitch dequantization information provided as input to the neural filter estimator 952 may include voiced / unvoiced (V / UV) classification information indicating whether the underlying audio represented in the encoded information received by the decoder 910 over the channel 940 corresponds to voiced or unvoiced sounds, speech, etc.
[0146] In examples where the neural filter estimator 952 receives spectral envelope features and dequantized pitch information as inputs, the neural filter estimator 952 can generate the corresponding filter characterization parameters for each NHV filter (e.g., noise LTVF 975, harmonic LTVF 965, etc.) based on the decoded spectral envelope features and the dequantized pitch information.
[0147] The noise LTVF 975 and the harmonic LTVF 965 can be implemented as time domain filters or frequency domain filters. In some examples, the noise LTVF 975 and the harmonic LTVF 965 can be implemented in the time domain or the frequency domain, independent of whether the neural filter estimator 952 is configured to generate theQualcomm Docket No.2403757WO corresponding filter characterization parameters in the time domain or the frequency domain. For example, in some aspects, the neural filter estimator 952 can output impulse response-based filter characterization parameters, and the operation of noise LTVF 975 and / or harmonic LTVF 965 can be implemented in the time domain as convolutions with the impulse response. Using the same impulse response-based filter characterization parameters, the operation of noise LTVF 975 and / or harmonic LTVF 965 may be implemented in the frequency domain based on converting the impulse response to frequency response (e.g., using an FFT transform) and multiplying with the FFT of the input signal, and subsequently converting back to the time domain using an IFFT transform.
[0148] In one illustrative example, the NHV speech synthesizer 950 can include a pulse train generator 960. The dequantized pitch information (e.g., generated using the pitch dequantization engine 938) can be provided to the pulse train generator 960. In some aspects, the pulse train generator 960 can generate a pulse train based at least in part on the pitch frequency (e.g., f0) and / or pitch lag or pitch delay information obtained from the pitch dequantization engine 938 of the decoder 910. In some examples, the pulse train generator 960 can be implemented as a cosine sum pulse generator, which can be configured to process the dequantized pitch information (e.g., obtained from the pitch dequantization engine 938) to generate a pulse train p[n]. The pulse train may also be referred to as an impulse train. For example, the pulse train generator 960 can be a differentiable cosine sum pulse generator, and / or can be a non-differentiable cosine sum pulse generator (e.g., among various other pulse generators).
[0149] The pulse train generated by the pulse train generator 960 can be processed by the harmonic LTVF 965, using the corresponding harmonic filter characterization parameters determined by the neural filter estimator 952 for the harmonic LTVF 965. In some aspects, the harmonic LTVF 965 can generate a harmonic output based on processing the pulse train from the pulse train generator 960.
[0150] In some examples, the pulse train generator 960 may receive an additional input from the pitch dequantization engine 938, indicative of a voice or unvoiced (e.g., V / UV) classification. For example, based on receiving an unvoiced (UV) indication or classification from the pitch dequantization engine 938, the pulse train generator 960 can be configured to generate a 0 output or a noise output that is provided to the harmonicQualcomm Docket No.2403757WO LTVF 965 instead of the pulse train p[n] (e.g., instead of the pulse train p[n] provided from the pulse train generator 960 to the harmonic LTVF 965 in response to a voiced (V) indication or classification from the pitch dequantization engine 938).
[0151] The NHV speech synthesizer 950 can include a noise generator 970 that is configured to generate a noise signal (e.g., white noise, etc.) for processing by the noise LTVF 975. For example, the noise LTVF 975 can be parameterized based on the respective noise filter characterization parameters generated by the neural filter estimator 952 for the noise LTVF 975, and can subsequently be used to process the noise signal generated by the noise generator 970. Based on processing the noise signal from the noise generator 970, the noise LTVF 975 can generate a noise-filtered output.
[0152] The NHV speech synthesizer 950 can include a combination function 980 to combine the harmonic output (e.g., from the harmonic LTVF 965) and the noise-filtered output (e.g., from the noise LTVF 975) to generate the reconstructed audio 955. In some examples, the reconstructed audio 955 includes synthesized speech. In some aspects, the reconstructed audio 955 can be the same as or similar to the reconstructed audio 455 of FIG. 4. In some examples, the combined noise LTVF 975 filter output and harmonic LTVF 965 filter output (e.g., from the combination operation 980) can comprise the reconstructed audio 955 (e.g., synthesized speech). In some aspects, the output of the combination operation 980 can be provided to an LPC 954, which may be the same as or similar to the LPC 452 of FIG.4. For example, the LPC 954 can perform linear predictive coding on the LTVF filter output from the NHV speech synthesizer 950, and may generate as output the reconstructed audio 955 (e.g., synthesized speech). In some examples, a linear time invariant (LTI) post-filter can be included between the output of the combination operation 980 and the input to the LPC 954.
[0153] In some examples, the combination function 980 can include an adder, a multiplier, a divider, weighted sum, a weighted product, a weighted ratio, an average, a weighted average, a weighted mean, a weighted median, a weighted mode, or a combination thereof. The reconstructed audio 955 may be an example of the reconstructed speech signal 255 of FIG. 2B, the reconstructed speech signal 275 of FIG. 2C, the reconstructed audio 455 of FIG.4, etc., or vice versa.
[0154] FIG.5A is a block diagram illustrating an example of a spectral envelope feature extraction branch 500 that can be included in a generative voice codec system. ForQualcomm Docket No.2403757WO example, the spectral envelope feature extraction branch 500 can be the same as or similar to the spectral envelope feature extraction 420 of the generative voice codec system 400 of FIG. 4, and / or can be used to implement the spectral envelope feature extraction 420 of the generative voice codec system 400 of FIG.4.
[0155] In some aspects, the spectral envelope feature extraction branch 500 can be used to generate (e.g., extract) one or more features 565 based on an audio signal 515. The features 565 can also be referred to as extracted features and / or spectral features (e.g., spectral envelope features), and can correspond to a spectral envelope of the audio signal 515. As noted previously, in some examples, the spectral features 565 can be spectral envelope features associated with the overall shape of the frequency spectrum of the audio signal 515 (e.g., where the spectral envelope connects the peaks of the individual frequency components in the frequency spectrum of an audio signal).
[0156] In some examples, the audio signal 515 of FIG.5A can be the same as or similar to the audio signal 415 of FIG. 4. The extracted spectral features 565 of FIG. 5A can be the same as or similar to a plurality of spectral envelope features generated by the spectral envelope feature extraction 420 and provided to the FRAE 425 included in the encoder 405 of FIG. 4. In one illustrative example, the audio signal 515 of FIG. 5A can be the same as or similar to the input provided to the spectral envelope feature extraction 420 of FIG.4.
[0157] For example, in cases where a denoiser (e.g., such as the denoiser 418 of FIG. 4) is not included before the spectral envelope feature extraction branch 500, the audio signal 515 can be the input speech or input audio provided to an encoder that implements the spectral envelope feature extraction branch 500 (e.g., such as the input audio 415 provided to the encoder 405 of FIG.4).
[0158] In cases where a denoiser (e.g., such as the denoiser 418 of FIG.4) is included before the spectral envelope feature extraction branch 500, the audio signal 515 can be the same as or similar to the de-noised audio output provided by the denoiser (e.g., the same as or similar to the output of the denoiser 418 of FIG.4, etc.).
[0159] In some aspects, the spectral envelope feature extraction branch 500 can include a frame splitting operation 520 configured to split the input audio signal 520 into a plurality of audio frames (e.g., a plurality of frames of audio data). In some examples, the frame splitting operation 520 can generate a plurality of equal length audio frames fromQualcomm Docket No.2403757WO the input audio signal 520. In some cases, the plurality of frames generated by the frame splitting operation 520 may be non-overlapping (e.g., each respective portion of the input audio signal 520 is included in only one audio frame of the plurality of audio frames). In some examples, the plurality of frames generated by the frame splitting operation 520 may be overlapping. For example, portions of the input audio signal 520 can be included in multiple audio frames (e.g., an overlapping portion or common subset between adjacent or consecutive audio frames of the plurality of audio frames generated using the frame splitting operation 520).
[0160] The spectral envelope feature extraction branch 500 can implement (e.g., apply) a filter bank 532 and an energy-per-band computation 534 to process the input audio signal 515 (e.g., in examples where the frame splitting operation 520 is not included or not utilized) or to process the plurality of frames split from the input audio signal 515 by the frame splitting operation 520 (e.g., in examples where the frame splitting operation 520 is included and utilized by the spectral envelope feature extraction branch 500).
[0161] In some aspects, the filter bank 532 and the energy-per-band computation 534 can be implemented as a filtering and energy computation branch 530 (e.g., a combined processing block 530), which can be the same as or similar to the example filtering and energy computation branch 570 of FIG. 5B. In some cases, the filter bank 532 can be applied prior to performing the energy-per-band computation 534. In some examples, the energy-per-band computation 534 can be performed prior to applying the filter bank 532 (e.g., such as in some cases where the filter bank 532 is computed in the frequency domain).
[0162] In one illustrative example, the filter bank 532 and the energy-per-band computation 534 of FIG. 5A can be implemented using the filtering and energy computation branch 570 of FIG.5B. In some aspects, the combined processing block 530 of FIG.5A can be implemented using the filtering and energy computation branch 570 of FIG.5B.
[0163] The filtering and energy computation branch 570 of FIG. 5B can optionally include a windowing operation 572, which may be the same as or similar to the frame splitting operation 520 of FIG.5A. In some aspects, the frame splitting operation 520 and the windowing operation 572 can be different. For example, the windowing operation 572 may be utilized in examples where the frame splitting operation 520 is not included in theQualcomm Docket No.2403757WO decoder or is not utilized by the decoder, etc. In some cases, the windowing operation 572 is not included or not utilized in examples where the frame splitting operation 520 is utilized.
[0164] The filtering and energy computation branch 570 of FIG. 5B can apply a fast Fourier Transform (FFT) 574 to convert or transform between the time domain and the frequency domain, or vice versa. The FFT 574 can be of various sizes or dimensions, and for example, may be configured as an FFT 574 with a size that is different from (e.g., does not match) the frame length (e.g., frame length associated with one or more of the frame splitting operation 520 of FIG.5A and / or the windowing function 572 of FIG.5B).
[0165] The output of the FFT 574 can be provided as input to a magnitude function 576. The magnitude function 576 can be configured to calculate a magnitude of the input received by the magnitude function 576. For example, the magnitude function 576 can be implemented as a squared modulus function (e.g., |·|2). In some examples, the magnitude function 576 being (or including) a squared modulus function (e.g., |·|2) can assist with bringing the input to the magnitude function 576 (e.g., the output from the FFT 574) into the real-valued domain. In some aspects, the magnitude function 576 can be implemented using other powers besides a power of two (e.g., other powers besides |·|2can be used for the magnitude function 576). For example, the magnitude function 576 can be implemented as |·| (e.g., a power of one). In some cases, the magnitude function 576 can be implemented as |·|afor any real-valued a > 0, including both fractional and integer values of a. In one illustrative example, the output of the magnitude function 576 can be referred to as a magnitude spectrum.
[0166] A filter bank 580 can be applied to the magnitude spectrum generated as output by the magnitude function 576. In one illustrative example, the filter bank 580 can be a filter bank of size N and can include a first filter 582-1 associated with a first frequency range, …, and an Nthfilter 582-N associated with an Nthfrequency range. In some aspects, the filter bank 580 of FIG.5B can be the same as or similar to the filter bank 532 of FIG. 5A (e.g., filter bank 532 and filter bank 580 can be filter banks of size N).
[0167] In some examples, the filter bank 580 is applied to the magnitude spectrum generated by the magnitude function 576 (e.g., the filter bank 580 is applied after the magnitude function 576). In some cases, the filter bank 580 can be applied before the magnitude function 576, for example by applying the filter bank 580 to the output of theQualcomm Docket No.2403757WO FFT 574. In such examples, the magnitude function 576 may subsequently be applied to the output of the filter bank 580. In one illustrative example, the order of the magnitude function 576 and the filter bank 580 can be reversed (e.g., such that the filter bank 580 is applied before the magnitude function 576 is applied) based on the filter bank 580 having a frequency response that is real-valued in the FFT domain associated with the FFT 574. If the filter bank 580 frequency response is not real-valued in the FFT domain associated with the FFT 574, the magnitude function 576 can be applied prior to the filter bank 580, and the input to the filter bank 580 can be the magnitude spectrum generated as output by the magnitude function 576.
[0168] In some aspects, the filter bank 530 of FIG.5A and / or the filter bank 580 of FIG. 5B (e.g., which can be the same as one another) may be computed in the time domain or the frequency domain. In one illustrative example, the filter bank 580 is a filter bank of size N and includes N different filters (e.g., filter 1582-1, a filter 2, …, filter N 582-N). Each filter of the N different filters within the filter bank 580 can be associated with a respective filter center frequency and / or a respective filter structure information. In some aspects, each filter of the N different filters can be used to perform filtering for a different frequency band or frequency range that is associated with the particular filter 582-1, …, 582-N.
[0169] In one illustrative example, the filters 582-1, …, 582-N of the filter bank 580 can have respective filter center frequencies and / or filter structure that are based on or configured based on a perceptual structure. For example, in some aspects, the filter bank 580 can implement the plurality of filter 582-1, …, 582-N as a Mel filter bank perceptual structure, a Bark filter bank perceptual structure, etc. In some aspects, the filter bank 580 can be a Mel filter bank including N Mel-scaled filters 582-1, …, 582-N that can be used to produce a set of coefficients. In some cases, the N Mel-scaled filters 582-1, …, 582-N can include a bank of triangular bandpass filters spaced on a logarithmic scale (e.g., the Mel scale) that is configured to replicate the non-linear human perception of pitch.
[0170] In examples where the filter bank 580 is implemented as a Mel filter bank, the output of the filtering and energy computation branch 570 of FIG.5B (e.g., the output of the filtering and energy computation block 530 of the spectral envelope feature extraction branch 500 of FIG.5A) can be referred to as Mel frequency cepstral coefficients (MFCC) or a Mel cepstrum.Qualcomm Docket No.2403757WO
[0171] In examples where the filter bank 580 is implemented as a Bark filter bank, the output of the filtering and energy computation branch 570 of FIG.5B (e.g., the output of the filtering and energy computation block 530 of the spectral envelope feature extraction branch 500 of FIG.5A) can be referred to as Bark frequency cepstral coefficients (BFCC) or a Bark cepstrum.
[0172] In one illustrative example, the operations of the filtering and energy computation branch 570 as illustrated in FIG. 5B may correspond to an example where the filter bank 580 is computed in the frequency domain. The output of the filter bank 580 can comprise a respective filter frequency response determined based on processing the magnitude spectrum generated by the magnitude function 576 with each respective one of the N filters 582-1, …, 582-N (e.g., the respective filter frequency response for filters 582-1, …, 582-N corresponding to the N sums 582-1, …, 584-N of FIG.5B).
[0173] To obtain a respective energy information 590-1, …, 590-N indicative of the energies per frequency band (e.g., per frequency band of the filter bank 580 filters 582-1, …, 582-N) the filter frequency response can be multiplied, elementwise for each frequency range or frequency bin corresponding to the sums 584-1, …, 584-N, with a signal FFT or |FFT|2, based on whether the filter bank 580 is applied before or after the magnitude function 576.
[0174] The output of the filtering and energy computation branch 570 of FIG. 5B can be the energies per band information indicative of the respective energy 590-1, …, 590- N within each of the N frequency bands associated with the filter bank 580. In one illustrative example, the energies per band information 590-1, …, 590-N can be the same as the output of the energy-per-band determination 534 of the spectral envelope feature extraction branch 500 of FIG.5A.
[0175] In some aspects, the energies per band information 590-1, …, 590-N can be provided to an optional companding function 540. The companding function 540 can be implemented by a compander, which may compress and / or expand a dynamic range of the input to the companding function 540. In some aspects, the companding function 540 can be implemented as Log(x) using any logarithm base (e.g., natural log ln(x), log10, etc.). In some cases, the companding function 540 can be implemented as (x)afor a ≤ 1. In some examples, the companding function 540 can be implemented as a log approximation function (e.g., such as ^^൫^√^^െ 1൯ for an integer n > 0). In some cases,Qualcomm Docket No.2403757WO multiplicative and additive constants can be disregarded, and the companding function540 may be implemented as ^√^^, etc.
[0176] In some examples, the companding function 540 is not included or is not utilized, and the energies per band information 590-1, …, 590-N may be provided directly to a Discrete Cosine Transform (DCT) function 555.
[0177] The DCT function 555 can be of size N (e.g., the same size N as the filter bank 532 and / or filter bank 580). The input to the DCT function 555 can be the companded energies per band of the filter bank 532 (e.g., in examples where the companding function 540 is included and utilized in the spectral envelope feature extraction branch 500 of FIG. 5A), or can be the un-companded (e.g., non-companded) energies per band of the filter bank 532 (e.g., in examples where the companding function 540 is not included or not utilized).
[0178] Applying or computing the DCT function 555 of size N can be associated with computing N DCT coefficients (e.g., where the DCT coefficients are coefficients for a sum of a series of cosine functions that approximates and / or represents the input sequence or signal in the frequency domain). The domain to which the input audio signal 515 is transformed after the DCT function 555 is referred to as “cepstrum.”
[0179] In one illustrative example, the spectral envelope feature extraction branch 500 of FIG.5A can implement DCT truncation 560, to truncate the output of the DCT function 555 to include only the first M DCT coefficients, where M < N. For example, the output of the DCT function 555 may be the N DCT coefficients 1, 2, …, N. The DCT truncation 560 can truncate to the first M DCT coefficients 1, 2, …, M, based on removing or skipping the computation (e.g., at the time of computation of the DCT function 555) of the last (N-M) DCT coefficients M+1, M+2, …, N.
[0180] In some aspects, the DCT truncation 560 can be implemented based on configuring the DCT computation 555 to implement a DCT of size N but only compute the first M coefficients of the expected N coefficients (with M<N). In such examples, the last (N-M) DCT coefficients M+1, M+2, …, N are neither computed nor output by the DCT function 555.
[0181] In another example, the DCT truncation 560 can be implemented based on removing or deleting the last (N-M) DCT coefficients M+1, M+2, …, N, after the DCTQualcomm Docket No.2403757WO function 555 generates as output the entire series of N DCT coefficients 1, 2, …, N.
[0182] In some aspects, the DCT truncation to the first M DCT coefficients where M<N (e.g., the DCT truncation 560 of FIG. 5A) can also be referred to as “cepstral liftering.” In one illustrative example, the DCT truncation 560 (e.g., cepstral liftering) implemented by the spectral envelope feature extraction branch 500 of FIG. 5A can be utilized to remove pitch information from the spectral envelope features 565.
[0183] For example, the pitch information can be removed from the extracted spectral envelope features 565 based on performing the DCT truncation (e.g., cepstral liftering) 560 to remove or skip the computation of a subset of higher order DCT coefficients. Based on performing the DCT truncation 560 to remove pitch information from the spectral features output included in the DCT function 555 cepstrum output, the extracted spectral envelope features 565 can be generated with a relatively high resolution (e.g., corresponding to the relatively large filter bank 532 size N), and then truncated to the smaller size M without a reduction in the resolution. Based on utilizing a relatively large filter bank 532 with a relatively large number of frequency bands (e.g., N, where N>M), and subsequently truncating the DCT of the relatively large filter bank output or cepstrum, the systems and techniques can be used to code only spectral envelope information in the cepstral feature output 565 of the spectral envelope feature extraction branch 500, without including or double coding pitch information in the extracted cepstral features 565.
[0184] In some aspects, the DCT truncation 560 (e.g., cepstral liftering) can be configured using a value of M (e.g., the truncation point corresponding to the first M DCT or cepstral coefficients out of the total N coefficients associated with the DCT 555 of size N>M) that is selected to be small enough to eliminate the pitch information from the cepstrum. For example, for a signal (e.g., audio signal 515) with a 16 kilohertz (kHz) sampling rate, in one example, the filter bank 532 and DCT function 555 can be configured with a size N = 80, and the DCT truncation (e.g. cepstral liftering) 560 can be configured to truncate at M = 24. The result of the DCT truncation 560 with M = 24 can be a set of features 565 comprising MFCC-24 Mel Frequency Cepstral Coefficients (MFCC) features.
[0185] Configuring the DCT truncation or cepstral liftering 560 with a value of M<N that is small enough to eliminate pitch information from the cepstrum can improve the bit efficiency associated with the spectral envelope feature extraction branch 500 of FIG.5AQualcomm Docket No.2403757WO and / or the bit efficiency associated with encoded performed using the extracted spectral features 565 (e.g., encoding performed using the encoder 405 of the generative voice codec system 400 of FIG.4, etc.). In one illustrative example, the systems and techniques can be configured to perform the DCT truncation 560 to remove any pitch information from the cepstrum (e.g., from the spectral features or cepstrum 565) based on the encoder of the generative voice codec system (e.g., encoder 405 of generative voice codec system 400 of FIG.4) being configured to extract, code, and transmit pitch information separately from the spectral envelope feature information. For example, the encoder 405 of FIG. 4 includes separate pitch extraction engines 430 and spectral envelope feature extraction engines 420, which are used to extract, code, and transmit respective pitch information and respective spectral envelope feature information that are separate from each other. In some aspects, performing the DCT truncation 560 (e.g., cepstral liftering) of the spectral envelope feature extraction branch 500 of FIG. 5A can be used to prevent double coding of pitch information in the extracted spectral features 565 that are coded and transmitted from the encoder 405 to the decoder 410 of FIG.4. Preventing the double coding of pitch information can improve the bit efficiency of the generative video codec system described herein, and can allow the encoder 405 to extract and send separate pitch and spectral envelope features only (e.g., the encoder 405 does not send explicit per-frame absolute phase information and / or pitch pulse locations, as in existing techniques for audio coding).
[0186] In some aspects, the size N of the filter bank 532 and DCT function 555 can be configured to be larger than the truncation size M, and to additionally be large enough to implement a relatively large (e.g., relatively high resolution) filter bank 532. At relatively small sizes of N, the filter bank 532 may comprise a coarse filter bank. Relatively small sizes of N (e.g., associated with coarse filter bank implementations) do not preserve spectral envelope information as well as larger values of N. Utilizing the relatively large value of N to implement a relatively high resolution, non-coarse filter bank 532 may improve the preservation of spectral envelope information in the output of the filter bank 532 and subsequent downstream operations of the spectral envelope feature extraction branch 500 of FIG. 5A. Subsequently applying the DCT truncation (e.g., cepstral liftering) 560 to the non-coarse filter bank 532 output can preserve the high resolution of the extracted spectral envelope features 565. For example, DCT function 555 can be associated with energy compaction properties that may be used to compress more spectralQualcomm Docket No.2403757WO information in fewer DCT coefficients. Utilizing a relatively small value of N corresponding to a coarse filter bank implementation for the filter bank 532 (e.g., instead of performing truncation to decrease a large N to a relatively small M, using the DCT truncation or cepstral liftering 560 of FIG. 5A) can decrease the spectral envelope resolution from the beginning of the processing performed by the spectral envelope feature extraction branch 500 of FIG. 5A, making it difficult or impossible to recover spectral envelope resolution at later stages of the encoder or decoder (e.g., encoder 405 or decoder 410 of the generative voice codec 400 of FIG.4, etc.).
[0187] FIG. 6 is a block diagram illustrating an example of a decoder 610 of an audio codec system 600 (e.g., generative voice codec) that can be the same as or similar to the audio codec system 400 of FIG.4. For example, in some aspects, the decoder 610 can be the same as or similar to the decoder 410 of FIG. 4, the channel 640 can be the same as or similar to the channel 440 of FIG. 4, the FRAE decoder 645 can be the same as or similar to the FRAE decoder 445 of FIG.4, the neural speech synthesizer 650 can be the same as or similar to the neural speech synthesizer 450 of FIG.4, the LPC 652 can be the same as or similar to the LPC 452 of FIG. 4, and / or the reconstructed audio 655 (e.g., synthesized speech) can be the same as or similar to the reconstructed audio 455 of FIG. 4.
[0188] In some aspects, the decoder 610 can receive encoded spectral envelope features over the channel 640 from a corresponding encoder (e.g., such as the encoder 405 of FIG. 4), and may utilize the FRAE decoder 645 to decode the encoded spectral features received over the channel 640. The output of the decoding performed using the FRAE decoder 645 can be an M-dimensional cepstrum 620, where the dimension M of the cepstrum 620 corresponds to a truncation dimension applied by the encoder (e.g., the cepstrum 620 can have dimension M that is the same as the truncation size M associated with the DCT truncation (e.g. , cepstral liftering) 560 of FIG.5A, etc.).
[0189] In one illustrative example, the decoder 610 can include an inverse DCT function 626 of size N, that is configured to generate filter bank energy information 628 as output, based on receiving the decoded M-dimensional cepstrum 620 as input from the FRAE decoder 645. In some aspects, the inverse DCT function 626 can correspond to a DCT function implemented by the encoder associated with the decoder 610. For example, the inverse DCT function 626 can be an inverse of the DCT function 555 of FIG. 5A,Qualcomm Docket No.2403757WO which may be implemented to generate extracted spectral envelope features of the cepstrum 620 at the encoder associated with the decoder 610 of FIG.6 (e.g., such as the encoder 405 of FIG.4, etc.).
[0190] In some aspects, the inverse DCT function 626 can be applied to the M- dimensional cepstrum 620 to generate companded or un-companded filter bank energy information 628 that is then provided as input to the neural speech synthesizer 650. For example, if the encoded spectral features received over the channel 640 by the decoder 610 are companded at the encoder-side (e.g., if the encoder associated with decoder 610 includes and utilizes a companding function, such as the companding function 540 of FIG. 5A, to generate the encoded spectral features), the output of the inverse DCT function626 can be companded filter bank energy information 628.
[0191] In examples where the encoded spectral features received over the channel 640 by the decoder 610 are not companded at the encoder-side (e.g., are un-companded, based on the encoder associated with decoder 610 not including or not utilizing a companding function such as the companding function 540 of FIG.5A to generate the encoded spectral features), the output of the inverse DCT function 626 can be un-companded filter bank energy information 628. In some aspects, the neural speech synthesizer 650 can be configured to generate the reconstructed audio signal 655 using either companded or un- companded filter bank energy information 628 from the inverse DCT function 626 as input.
[0192] In some examples, the inverse DCT function 626 can receive as input (e.g., and the FRAE decoder 645 can generate as output) an M-dimensional cepstrum 620 comprising 24-MFCC information. The 24-MFCC information of the M-dimensional (e.g., M = 24) cepstrum 620 can be converted back to log Mel energies as companded filter bank energy information 628 generated by the inverse DCT function 626, or can be converted to the reconstructed filter bank energies (e.g., non-log, non-companded / un- companded filter bank energies) as un-companded filter bank energy information 628 generated by the inverse DCT function 626.
[0193] In some aspects, the decoder 610 can include a pitch encoding dequantization engine 638 (e.g., de Q()), configured to perform dequantization of quantized pitch information received over the channel 640 from the corresponding encoder associated with the decoder 610. For example, the pitch encoding dequantization engine 638 canQualcomm Docket No.2403757WO dequantize the quantized pitch information generated by the pitch quantization engine 435 Q() of the encoder 405 of FIG.4, etc. In some cases, the received pitch encoding can be dequantized by the pitch encoding dequantization engine 638, based on performing lookup before being fed to the neural speech synthesizer 650. For example, the neural speech synthesizer 650 can be configured and / or trained to generate the reconstructed audio signal 655 utilizing quantized pitch information (e.g., in examples where the decoder 610 does not include the pitch encoding dequantization engine 638 or does not perform pitch dequantization) and / or utilizing dequantized (e.g., recovered) pitch information in examples where the decoder includes and utilizes the pitch encoding dequantizing engine 638.
[0194] FIG. 7 is a block diagram illustrating an example of a decoder 710 of an audio codec system 700 (e.g., generative voice codec) that can be the same as or similar to the audio codec system 400 of FIG.4, the audio codec system 600 of FIG.6, etc. For example, in some aspects, the decoder 710 can be the same as or similar to the decoder 410 of FIG. 4, the decoder 610 of FIG. 6, etc. In some cases, the channel 740 can be the same as or similar to the channel 440 of FIG.4 and / or 640 of FIG.6, the FRAE decoder 745 can be the same as or similar to the FRAE decoder 445 of FIG.4 and / or 645 of FIG.6, the neural speech synthesizer 750 can be the same as or similar to the neural speech synthesizer 450 of FIG.4 and / or 650 of FIG.6, the LPC 752 can be the same as or similar to the LPC 452 of FIG. 4 and / or 652 of FIG. 6, and / or the reconstructed audio 755 (e.g., synthesized speech) can be the same as or similar to the reconstructed audio 455 of FIG.4 and / or 655 of FIG.6, etc.
[0195] In some examples, the pitch encoding dequantization engine 738 can be the same as or similar to the pitch encoding dequantization engine 638 of FIG. 6. The M- dimensional cepstrum 720 can be the same as or similar to the M-dimensional cepstrum 620 of FIG. 6. The inverse DCT function 726 of size N can be the same as or similar to the inverse DCT function 626 of size N of FIG. 6. In some aspects, the decoder 710 can receive encoded spectral envelope features over the channel 740 from a corresponding encoder (e.g., such as the encoder 405 of FIG.4, etc.), and may utilize the FRAE decoder 745 to decode the encoded spectral features received over the channel 740. The output of the decoding performed using the FRAE decoder 745 can be the M-dimensional cepstrum 720, where the dimension M of the cepstrum 720 corresponds to a truncation dimension applied by the encoder (e.g., the cepstrum 720 can have dimension M that is the same asQualcomm Docket No.2403757WO the truncation size M associated with the DCT truncation (e.g. , cepstral liftering) 560 of FIG.5A, etc.).
[0196] In one illustrative example, the M-dimensional cepstrum 720 can be processed by the inverse DCT function 726 to generate companded filter bank energies 728, which may be the same as or similar to the filter bank energies 628 of FIG.6 (e.g., in examples where the decoder 610 of FIG. 6 receives companded encoded spectral features). The companded filter bank energies 728 can be processed by an inverse companding function 745 included in the decoder 710 to generate de-companded (e.g., un-companded) filter bank energies 748.
[0197] The de-companded filter bank energies 748 can be recovered filter bank energy information corresponding to the filter bank energy information determined by the encoder associated with decoder 710. The inverse companding function 745 included in the decoder 710 can be an inverse of a companding function included in and applied by the associated encoder in generating the encoded, companded spectral envelope features transmitted to the decoder 710 over the channel 740. For example, the inverse companding function 745 can be an inverse of the companding function 540 of FIG.5A, etc.
[0198] In one illustrative example, the inverse companding function 745 can be implemented as exp(x), based on the encoder associated with the decoder 710 implementing a natural log companding function during the computation of the extracted spectral envelope features that are encoded by the associated encoder of the generative voice codec system 700 and transmitted to the decoder 710 over the channel 740.
[0199] The neural speech synthesizer 750 can be configured to generate the reconstructed audio signal 755 based on the de-companded filter bank energies 748 and recovered pitch information (e.g., either quantized or dequantized) also received by the decoder 710 over the channel 740 and from the corresponding encoder associated with decoder 710.
[0200] In some aspects, a decoder of the generative voice codec systems described herein can utilize decoded spectral envelope features and decoded and / or dequantized pitch information as separate inputs to a neural speech synthesizer of the decoder. For example, the neural speech synthesizer 450 of the decoder 410 of FIG.4 can utilize a first input comprising decoded spectral envelope features and a second input comprising pitchQualcomm Docket No.2403757WO information to generate the reconstructed audio 455. The neural speech synthesizer 650 of the decoder 610 of FIG. 6 can utilize a first input comprising filter bank energy information 628 of decoded spectral envelope features, and a second input comprising received pitch encoding information to generate the reconstructed audio 655. The neural speech synthesizer 750 of the decoder 710 of FIG. 7 can utilize a first input comprising de-companded filter bank energies 748 of the decoded spectral envelope features, and a second input comprising the received pitch encoding information to generate the reconstructed audio 755.
[0201] In one illustrative example, the systems and techniques can utilize a generative voice codec system with a decoder that is configured to combine pitch information and spectral envelope features (e.g., received and decoded separately by the decoder) into a combined feature or combined representation that includes both pitch and spectral features in a single feature vector or a single representation. The single, combined representation of pitch and spectral features can be generated by the decoder and provided as input to a neural speech synthesizer of the decoder to generate a reconstructed audio signal from the combined pitch and spectral feature.
[0202] For example, FIG.8 is a block diagram illustrating an example of a generative voice codec system 800 with a decoder 810 that can be configured to perform feature reconstruction 880 to generate a combined feature or representation that includes both spectral features (e.g., based on the M-dimensional cepstrum 820) and pitch information 839.
[0203] In some aspects, the generative voice codec system 800 of FIG. 8 can be the same as or similar to the generative voice codec system 400 of FIG.4, 600 of FIG.6, 700 of FIG.7, etc. The decoder 810 of FIG.8 can be the same as or similar to the decoder 410 of FIG.4, the decoder 610 of FIG.6, and / or the decoder 710 of FIG.7, etc. In some cases, the channel 840 of FIG.8 can be the same as or similar to the channel 440 of FIG.4, the channel 640 of FIG.6, the channel 740 of FIG.7, etc. The FRAE decoder 845 of FIG.8 can be the same as or similar to the FRAE decoder 445 of FIG.4, 645 of FIG. 6, and / or 745 of FIG. 7, etc. The neural speech synthesizer 850 of FIG. 8 can be the same as or similar to the neural speech synthesizer 450 of FIG.4, 650 of FIG.6, and / or 750 of FIG. 7, etc. The LPC 852 of FIG. 8 can be the same as or similar to the LPC 452 of FIG. 4, 652 of FIG. 6, and / or 752 of FIG. 7, etc. The reconstructed audio 855 (e.g., synthesizedQualcomm Docket No.2403757WO speech) of FIG.8 can be the same as or similar to the reconstructed audio 455 of FIG.4, 655 of FIG.6, 755 of FIG.7, etc.
[0204] In some aspects, the M-dimensional cepstrum 820 of FIG.8 can be the same as or similar to the M-dimensional cepstrum 620 of FIG. 6 and / or 720 of FIG. 7, etc. The pitch dequantization engine 838 of FIG. 8 can be the same as or similar to the pitch dequantization engine 638 of FIG.6 and / or 738 of FIG.7, etc. The inverse DCT function 826 of size N can be the same as or similar to the inverse DCT function (of size N) 626 of FIG. 6 and / or the inverse DCT function (of size N) 726 of FIG. 7. The inverse companding function 845 of FIG. 8 can be the same as or similar to the inverse companding function 745 of FIG.7.
[0205] As noted above, the decoder 810 of FIG.8 can be configured to perform feature reconstruction 880 to generate a combined feature or representation that includes both spectral features (e.g., based on the M-dimensional cepstrum 820) and pitch information 839. In one illustrative example, the combined spectral and pitch feature generated by the feature reconstruction engine 880 can be fed directly to the neural speech synthesizer 850 for generating the reconstructed audio 855 based on the combined spectral and pitch feature. In examples where the neural speech synthesizer 850 receives the combined spectral and pitch feature generated by the feature reconstruction engine 880, the inverse DCT function 826 and the inverse companding function 845 may be left out from the decoder 810 or may be included but deactivated within the decoder 810.
[0206] In one illustrative example, the feature reconstruction engine 880 can generate a combined feature or representation of both spectral envelope features 820 and pitch information 839, based on the feature reconstruction engine 880 being configured to add a pitch peak (e.g., based on the pitch information 839) to the high quefrency of the spectral envelope cepstrum 820.
[0207] As noted previously, the cepstrum is a representation that can be generated from taking the inverse Fourier transform of the logarithm of the magnitude spectrum of a signal. The quefrency of a cepstrum can be a frequency of frequencies, or a measure of time in the cepstral domain (e.g., analogous to frequency in the spectral domain). In some aspects, the feature reconstruction engine 880 can add a pitch peak determined from the pitch information 839 to the cepstrum 820, where the added pitch peak is at a quefrency corresponding to the fundamental frequency (e.g., pitch) of the voice or speech beingQualcomm Docket No.2403757WO represented.
[0208] In one illustrative example, the feature reconstruction engine 880 can be configured to add a pitch peak to the high quefrency of the cepstrum 820 between the M- th and N-th cepstral coefficients (e.g., where the M-th through N-th cepstral coefficients of the input cepstrum 820 provided to the feature reconstruction engine 880 are 0 based on the DCT truncation (e.g., cepstral liftering) performed previously at the encoder, such as the DCT truncation 560 of FIG.5A). In some aspects, the feature reconstruction engine 880 can calculate the pitch peak and the corresponding location for insertion or addition of the pitch peak into the high quefrency of the cepstrum 820 in order to reconstruct or add back to the cepstrum 820 the pitch information that was previously removed by the DCT truncation (e.g., cepstral liftering) implemented at the encoder-side of the generative voice codec system 800. In some aspects, the output of the feature reconstruction engine 880 can be an augmented or reconstructed cepstrum that is based on the truncated M- dimensional cepstrum 820 and the calculated pitch peak inserted into the high quefrency of the truncated M-dimensional cepstrum 820 by the feature reconstruction engine 880.
[0209] In some examples, the augmented or reconstructed cepstrum (e.g., generated and output by the feature reconstruction engine 880, based on insertion of the pitch peak determined from the pitch information 839 into the high quefrency of the M-dimensional truncated cepstrum 820) can be provided to an inverse DCT function 826 of size N. In examples where the inverse DCT function 826 is utilized, the input to the inverse DCT function 826 can be the augmented cepstrum from the feature reconstruction engine 880 and the output of the inverse DCT function 826 can be 80-log Mel feature information that includes spectral envelope or energy information and the additional pitch information associated with the pitch peaks added by the feature reconstruction engine 880. In some cases, the input, truncated M-dimensional cepstrum 820 has been smoothed to the spectral envelope features only (e.g., based on the spectral envelope feature extraction and DCT truncation performed at the encoder-side of the generative video codec system 800). After the feature reconstruction engine 880, the augmented or reconstructed cepstrum generated as output by the feature reconstruction engine 880 can be a cepstrum or other representation of a full spectral structure (e.g., including pitch peaks and no longer smoothed to envelope only).
[0210] In some cases, the output of the inverse DCT function 826 (e.g., 80-log MelQualcomm Docket No.2403757WO feature information that includes spectral information and the additional pitch information) can be provided as input to the neural speech synthesizer 850. In some aspects, the output of the inverse DCT function 826 can be provided as input to an inverse companding function 845, which can apply the inverse of a companding function implemented on the encoder-side of the generative voice codec system 800.
[0211] In some aspects, the feature reconstruction engine 880 can be configured to generate or reconstruct a reconstructed full N-dimensional cepstrum, from inputs comprising the truncated M-dimensional cepstrum 820 decoded by the FRAE decoder 845, and the dequantized pitch information 839 from the pitch dequantization engine 838. The feature reconstruction engine 880 can be implemented using one or more machine learning networks, one or more DSPs or DSP processing techniques, and / or various combinations thereof.
[0212] In some examples, the feature reconstruction engine 880 is configured to generate an output of reconstructed features with both spectral envelope and pitch information (e.g., an output representation or features that includes a combination of spectral envelope and pitch information). In some cases, the reconstructed combined feature representation output by the feature reconstruction engine 880 can comprise a cepstrum or log Mel representation.
[0213] In some examples, the reconstructed combined feature representation output by the feature reconstruction engine 880 can comprise representations other than a cepstrum or log Mel representation (e.g., the reconstructed combined feature representation can be generated in any feature domain by the feature reconstruction engine 880).
[0214] In some aspects, the systems and techniques can include a generative voice codec encoder that is configured to generate, encode, and transmit (e.g., to a corresponding generative voice codec decoder) one or more energy features and / or energy information based on the input audio signal that is also used to determine the encoded spectral envelope features and pitch information. For example, the encoder 405 of FIG.4 and / or various other encoders described herein can be configured to generate, encode, and transmit one or more energy features and / or energy information to a corresponding decoder (e.g., the decoder 410 of FIG. 4, and / or any one of the decoders of FIGS. 1-15 described herein, etc.).
[0215] In some examples, the encoder can include an energy feature computation toQualcomm Docket No.2403757WO compute one or more energy features or energy information indicative of an audio frame energy. For example, the frame energy can be determined by the encoder as a sum of squares of the signal amplitudes for the audio samples within a particular frame. In some cases, the encoder can compute, encode, and transmit the respective frame energy information or features for each frame of audio data included in a plurality of frames of audio data of the input audio signal to the encoder (e.g., such as input audio signal 415 to the encoder 405 of FIG.4, etc.).
[0216] In some aspects, the energy feature(s) determined by the encoder can replace the 0-th coordinate or the cepstral features in the cepstral feature vector generated, encoded, and transmitted from the encoder to the decoder of the generative voice codec implemented using the systems and techniques described herein. For example, the energy features determined by the decoder can replace the 0-th coordinate of the cepstral features in a cepstral feature vector generated by the encoder (e.g., generated based on the DCT truncation or cepstral liftering performed by the encoder, such as the DCT truncation or cepstral liftering 560 of FIG.5A, etc.). The cepstral feature vector with the inserted energy feature information at the 0-th coordinate of the cepstral features can subsequently be coded by the FRAE autoencoder of the encoder (e.g., such as the FRAE autoencoder 425 included in the encoder 405 of FIG.4, etc.).
[0217] In some examples, the energy feature(s) and / or energy information determined by the encoder can be coded and transmitted from the encoder to the decoder of the generative voice codec system separately from the spectral envelope feature information and / or the pitch information that are determined from the same input audio signal by the encoder. For example, the energy feature(s) may be encoded using an additional FRAE autoencoder implemented by the encoder (e.g., a separate instance of an FRAE autoencoder from the FRAE used for encoded the spectral envelope features). In some examples, the 0-th coordinate of the cepstral features may optionally be dropped or removed, with only the remaining cepstral features after removal or dropping of the 0-th coordinate of the cepstral features being coded and transmitted from the FRAE autoencoder of the encoder-side to the FRAE decoder of the decoder-side.
[0218] FIG.10 is a block diagram illustrating an example of a decoder 1010 of a voice coding signal synthesis system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer comprising a linear predictionQualcomm Docket No.2403757WO coding (LPC) network. The decoder 1010 can be included in an audio codec system (e.g., generative voice codec system) 1000 that can be used to generate reconstructed audio (e.g., synthesized speech) 1055.
[0219] In some aspects, the generative voice codec system 1000 of FIG. 10 can be the same as or similar to the generative voice codec system 400 of FIG.4, 600 of FIG.6, 700 of FIG.7, 800 of FIG.8, 900 of FIG.9, etc. In some cases, the channel 1040 of FIG.10 can be the same as or similar to the channel 440 of FIG.4, the channel 640 of FIG.6, the channel 740 of FIG.7, the channel 840 of FIG.8, the channel 940 of FIG.9, etc.
[0220] The decoder 1010 of FIG. 10 can be the same as or similar to the decoder 410 of FIG.4, the decoder 610 of FIG.6, the decoder 710 of FIG.7, the decoder 810 of FIG. 8, the decoder 910 of FIG. 9, etc. For example, the decoder 1010 can include an FRAE decoder 1045 the same as or similar to the FRAE decoder 445 of FIG. 4, 645 of FIG. 6, 745 of FIG. 7, 845 of FIG.8, 945 of FIG. 9, etc. The decoder 1010 can include a neural speech synthesizer 1050 that can be the same as or similar to the neural speech synthesizer 450 of FIG.4, 650 of FIG.6, 750 of FIG.7, 850 of FIG.8, 950 of FIG.9, etc. The decoder 1010 can generate reconstructed audio 1055 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 455 of FIG.4, 655 of FIG.6, 755 of FIG.7, 855 of FIG.8, 955 of FIG.9, etc.
[0221] In some examples, the neural speech synthesizer 1050 can receive a first input comprising decoded spectral envelope features (e.g., determined by the FRAE decoder 1045). The neural speech synthesizer 1050 can receive a second input comprising a received pitch encoding (e.g., received pitch information), that is the same as or similar to the received pitch encoding provided to the neural speech synthesizer 450 of FIG. 4, etc. For example, the decoder 1010 may include a pitch de-quantization engine 1038 (e.g., de Q()). The pitch de-quantization engine 1038 can be used to process a received pitch encoding obtained by the decoder 1010 over the channel 1040. For example, the received pitch encoding information can be quantized pitch information or a quantized pitch signal transmitted to the decoder 1010 over the channel 1040 by a corresponding encoder (e.g., an encoder associated with the decoder 1010, which may be the same as or similar to the encoder 405 of FIG. 4, etc.). In some aspects, the pitch dequantization engine 1038 can generate reconstructed pitch information based on performing a codebook lookup to dequantize the received pitch encoding obtained from the channel 1040 (e.g., toQualcomm Docket No.2403757WO dequantize the quantized pitch encoding received by the decoder 1010 over the channel 1040). The dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 1038 can include pitch information in the frequency-domain (e.g., f0 pitch frequency information in Hz), pitch information in the time-domain (e.g., pitch lag or pitch delay information in samples, milliseconds, etc.), and / or pitch correlation information. In some aspects, the dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 1038 can include information indicative of a voiced or unvoiced (V / UV) classification.
[0222] In one illustrative example, the neural speech synthesizer 1050 can be implemented as a linear predictive coding (LPC) network. For example, the LPC network-based neural speech synthesizer 1050 can be used to implement a linearprediction (LP) synthesis filter, based on ^^^^^^ ൌ ∑^ ^ୀ^ ^^^^^^^^ െ ^^^ ^ ^^^^^^. Here, ^^^^^^represents the input signal to the LP filter, ^^^^^^ represents the output signal, ^^^represents the linear prediction coefficients associated with implementing the LP synthesis filter, and ^^ is a value corresponding to the LP filter order. For example, the LP filter order p can have a value of 16 or 12, etc., among various other LP filter order values.
[0223] For example, the LPC-based neural speech synthesizer 1050 can include a frame rate network 1052 configured to process inputs comprising the decoded features obtained using the FRAE decoder 1045 and the decoded pitch information obtained using the pitch dequantization engine 1038. An LPC estimation engine 1062 can process the decoded features from the FRAE decoder 1045 to determine one or more estimated LP coefficients. In some aspects, the LPC estimation engine 1062 can generate estimated LP coefficients ^^^, based on an input comprising the decoded features determined using the FRAE decoder 1045.
[0224] In some examples, the LP coefficients estimated using the LPC estimation engine 1062 can be provided to an LP prediction engine 1064, configured to generate asoutput a prediction p[n], where ^^^^^^ ൌ ∑^ ^ୀ^ ^^^^^^^^ െ ^^^. In some aspects, the LP prediction engine 1064 can generate the LP prediction p[n] based on a first input comprising the LP coefficients ^^^estimated by the LPC estimation engine 1062, and a feedback input s[n-1].
[0225] The estimated or predicted LP coefficients p[n] can be provided as input to a sample rate network 1072 and a downstream combination or summation operation 1080.Qualcomm Docket No.2403757WO The sample rate network 1072 can receive additional inputs comprising the output of the frame rate network 1052, the feedback s[n-1] from the previous step n-1, and an intermediate feedback value e[n-1] from the same previous step n-1.
[0226] The output of the sample rate network 1072 can be the probability distribution P(e[n]), which is a probability distribution for e[n]. A sampling engine 1074 can perform sampling from the probability distribution P(e[n]) to obtain a realization or representation of e[n].
[0227] The representation of e[n] determined by the sampling engine 1074 can be combined with the LP prediction p[n] (e.g., determined by the LP prediction engine 1064), using the combination or summation operation 1080 to thereby generate as output the signal s[n]. In some aspects, the output signal s[n] can be the same as the reconstructed audio signal 1055 of the LPC network-based neural speech synthesizer 1050 and / or decoder 1010.
[0228] The representation of e[n] determined by the sampling engine 1074 can additionally be provided to a first feedback calculation 1078, which generates as output e[n-1] provided as an additional input to the sample rate network 1072.
[0229] The output signal s[n] of the combination or summation operation 1080 can be output as the reconstructed audio 1055 and may additionally be provided to a second feedback calculation 1079, which generates the representation s[n-1] based on the input s[n]. The representation s[n-1] can be provided as a feedback input to the LP prediction engine 1064 and to the sample rate network 1072.
[0230] FIG. 11 is a block diagram illustrating an example of an audio codec system (e.g., generative voice codec system) 1100 that encodes energy information 1110 using a first FRAE autoencoder 1115 and encodes spectral envelope features 1120 using a second FRAE autoencoder 1125, in accordance with some examples.
[0231] In some aspects, the generative voice codec system 1100 of FIG.11 can be the same as or similar to the generative voice codec system 400 of FIG.4, 600 of FIG.6, 700 of FIG. 7, 800 of FIG.8, 900 of FIG. 9, 1000 of FIG. 10, etc. The encoder 1105 of FIG. 11 can be the same as or similar to the encoder 405 of FIG. 4, etc. For example, the encoder 1105 can include spectral envelope feature extraction 1120 the same as or similar to the spectral envelope feature extraction 420 of FIG.4, can include pitch extraction 1130Qualcomm Docket No.2403757WO the same as or similar to the pitch extraction 430 of FIG.4, can include a spectral feature FRAE autoencoder 1125 the same as or similar to the FRAE autoencoder 425 of FIG.4, can include a first pitch quantization engine 1135 and a second pitch quantization engine 1137 each of which may be the same as or similar to the pitch quantization engine 435 of FIG.4, etc.
[0232] In some cases, the channel 1140 of FIG.11 can be the same as or similar to the channel 440 of FIG.4, the channel 640 of FIG.6, the channel 740 of FIG.7, the channel 840 of FIG.8, the channel 940 of FIG. 9, the channel 1040 of FIG.10, etc. The decoder 1110 of FIG. 11 can be the same as or similar to the decoder 410 of FIG. 4, the decoder 610 of FIG.6, the decoder 710 of FIG.7, the decoder 810 of FIG.8, the decoder 910 of FIG. 9, the decoder 1010 of FIG. 10, etc. For example, the decoder 1110 can include a spectral feature FRAE decoder 1145 the same as or similar to the FRAE decoder 445 of FIG.4, 645 of FIG.6, 745 of FIG.7, 845 of FIG.8, 945 of FIG.9, 1045 of FIG.10, etc. The decoder 1110 of FIG. 11 may further include a second FRAE decoder 1147 that is the same as or similar to the spectral feature FRAE decoder 1125. The decoder 1110 can include a neural speech synthesizer 1150 that can be the same as or similar to the neural speech synthesizer 450 of FIG. 4, 650 of FIG. 6, 750 of FIG. 7, 850 of FIG. 8, 950 of FIG. 9, 1050 of FIG. 10, etc. The decoder 1110 can generate reconstructed audio 1155 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 455 of FIG. 4, 655 of FIG. 6, 755 of FIG.7, 855 of FIG. 8, 955 of FIG. 9, 1055 of FIG. 10, etc.
[0233] In some examples, the energy features 1110 and extracted pitch information 1130 (e.g., each of pitch lag, V / UV classification, and / or pitch correlation, etc.) can be coded (e.g., encoded) by the encoder 1105 using non-machine learning techniques, such as a single stage vector quantization (SSVQ) technique, a multi-stage vector quantization (MSVQ), other vector quantization technique, gain-shape technique, etc. The energy features 1110 and extracted pitch information 1130 (e.g., pitch lag, V / UV classification, and / or pitch correlation, etc.) can be coded with or without forward error correction (FEC) prior to transmission over the channel 1140 between the encoder 1105 and decoder 1110 of the generative voice codec system 1100.
[0234] In some examples, the energy features 1110 and extracted pitch information 1130 (e.g., each of pitch lag, V / UV classification, and / or pitch correlation, etc.) can beQualcomm Docket No.2403757WO coded (e.g., encoded) by the encoder 1105 using one or more machine learning coding (e.g., encoding) techniques.
[0235] For example, the energy features 1110 and / or pitch information (e.g., pitch lag, V / UV classification, and / or pitch correlation) can be encoded using a corresponding one or more FRAE autoencoders. A separate FRAE autoencoder instance can be included in the encoder 1105 for coding of the energy feature(s) and for coding the pitch information, where the separate FRAE autoencoder instances for coding energy features and / or pitch information is additionally separate from the spectral feature FRAE autoencoder 1125 used by the encoder 1105 to code the extracted spectral envelope features 1120. The FRAE autoencoder instances implemented by the encoder 1105 for coding of the energy features 1110 and / or for coding the pitch information 1130 (e.g., pitch lag, V / UV classification, and / or pitch correlation) can be implemented with FEC, and / or can be implemented without FEC.
[0236] In some aspects, energy features 1110 can be coded using a corresponding separate FRAE autoencoder 1115 that is separate from the FRAE autoencoder 1125 used to code the extracted spectral envelope features 1120. For example, the extracted spectral envelope features 1120 can be coded and transmitted from the encoder 1105 as the latent representation z generated from the parameter h of the FRAE autoencoder 1125. The energy features 1110 can be coded and transmitted from the encoder 1105 as the latent representation z2generated from the parameter h2of the FRAE autoencoder 1115.
[0237] In some examples, the pitch information determined by pitch extraction 1130 can be coded using non-machine learning techniques (e.g., without a corresponding FRAE autoencoder instance for pitch). For example, the pitch extraction engine 1130 of the encoder 1105 can output pitch information (e.g., pitch lag, V / UV classification, and / or pitch correlation) to a first quantization engine 1135 and a second quantization engine 1137, which can be the same as or similar to one another. The first and second quantization engines 1135 and 1137, respectively, can be used to implement FEC techniques to improve the error concealment of the quantized pitch encoding information transmitted over the channel 1140 from the encoder 1105 to the decoder 1110. In some cases, FEC techniques implemented by the encoder 1105 for non-ML coded information (e.g., energy features 1110 and / or pitch information 1130) can include coding with multiple codebooks using multiple description coding (MDC). In some examples, FEC techniquesQualcomm Docket No.2403757WO implemented by the encoder 1105 for non-ML coded information can include full redundancy FEC.
[0238] For example, full redundancy FEC can be implemented for the extracted pitch information 1130 (e.g., pitch lag, V / UV classification, pitch correlation, etc.) using the first and second pitch quantization engines 1135 and 1137, respectively, of the encoder 1105. The encoder 1105 can generate and transmit two identical copies of the encoding, where the two identical copies (e.g., from the first quantization engine 1135 and the second quantization engine 1137) are transmitted over the channel 1140 and to the decoder 1110 using two separate packets. In another example, the encoder 1105 can generate two redundant encodings of the pitch information 1130, for example by using two different codebooks with different sizes. In some aspects, the secondary encoding is of a lower bit rate than the primary encoding (e.g., the secondary encoding can be generated using a smaller codebook, and the primary encoding can be generated using a larger codebook). The primary and secondary encodings of the pitch information 1130 can be transmitted in different (e.g., separate) packets from the encoder 1105 to the decoder 1110, over the channel 1140.
[0239] In some cases, the pitch information 1130 can be coded and transmitted from the encoder 1105 to the decoder 1110 using ML-based coding techniques (e.g., using a dedicated FRAE autoencoder instance for the pitch information 1130, such as the FRAE autoencoder 1115), and the energy feature(s) 1110 can be coded and transmitted from the encoder 1105 to the decoder 1110 using non-ML based coding techniques (e.g., such as the quantization engines 1135, 1137 and one or more FEC and / or MDC techniques, etc.).
[0240] Each FRAE autoencoder 1115, 1125 included in the encoder 1105 and used to generate a corresponding encoded representation or latent (e.g., the spectral feature latent z and the energy feature latent z2) can be associated with a corresponding FRAE decoder instance implemented in the decoder 1110. For example, the FRAE autoencoder 1115 of the encoder 1105 can be used to generate and transmit the latent representation z2 of the energy features 1110 to a corresponding FRAE decoder 1147 included in the decoder 1110. The FRAE autoencoder 1125 of the encoder 1105 can be used to generate and transmit the latent representation z of the spectral envelope features 1120 to a corresponding FRAE decoder 1145 included in the decoder 1110. The decoded features generated by the FRAE decoder instances 1145 and 1147 of the decoder 1110 can beQualcomm Docket No.2403757WO provided as separate inputs to the neural speech synthesizer 1150. For example, the neural speech synthesizer 1150 can be configured to generate the reconstructed audio 1155 based on decoded (e.g., reconstructed) energy features determined by the energy FRAE decoder instance 1147, decoded (e.g., reconstructed) spectral features determined the spectral FRAE decoder instance 1145, and pitch information obtained from the pitch extraction engine 1130 of the encoder 1105.
[0241] In some aspects, the encoder can be configured to combine and / or concatenate the different types of features extracted from an input audio signal, and perform coding of the combined features (e.g., energy features, spectral features, pitch features) using a single FRAE autoencoder instance, with or without utilizing MDC. For example, FIG.12 is a block diagram illustrating an example of an audio codec system (e.g., generative voice codec system) 1200 that can be used to perform neural speech synthesis based on a combined vector of encoded features including energy features 1210, spectral features 1220, and pitch information 1230, in accordance with some examples.
[0242] In some aspects, the generative voice codec system 1200 of FIG. 12 can be the same as or similar to the generative voice codec system 400 of FIG.4, 600 of FIG.6, 700 of FIG. 7, 800 of FIG. 8, 900 of FIG. 9, 1000 of FIG. 10, 1100 of FIG. 11, etc. For example, the encoder 1205 of FIG.12 can be the same as or similar to the encoder 405 of FIG.4, 1105 of FIG.11, etc. For example, the encoder 1205 can include spectral envelope feature extraction 1220 the same as or similar to the spectral envelope feature extraction 420 of FIG. 4, 1120 of FIG. 11, etc. Energy determination 1210 of FIG. 12 can be the same as or similar to the energy determination 1110 of FIG.11. Pitch extraction 1230 of FIG. 12 can be the same as or similar to the pitch extraction 430 of FIG. 4, and / or 1130 of FIG.11, etc. A spectral feature FRAE autoencoder 1225 can be the same as or similar to the FRAE autoencoder 425 of FIG.4, 1125 of FIG.11 and / or 1115 of FIG.11, etc.
[0243] In some cases, the channel 1240 of FIG.12 can be the same as or similar to the channel 440 of FIG.4, the channel 640 of FIG.6, the channel 740 of FIG.7, the channel 840 of FIG.8, the channel 940 of FIG.9, the channel 1040 of FIG.10, the channel 1140 of FIG.11, etc. The decoder 1210 of FIG.12 can be the same as or similar to the decoder 410 of FIG.4, the decoder 610 of FIG.6, the decoder 710 of FIG.7, the decoder 810 of FIG.8, the decoder 910 of FIG.9, the decoder 1010 of FIG.10, the decoder 1110 of FIG. 11, etc. For example, the decoder 1210 can include a spectral feature FRAE decoder 1245Qualcomm Docket No.2403757WO the same as or similar to the FRAE decoder 445 of FIG.4, 645 of FIG.6, 745 of FIG.7, 845 of FIG.8, 945 of FIG. 9, 1045 of FIG. 10, 1145 of FIG. 11 and / or 1147 of FIG. 11, etc. The decoder 1210 can include a neural speech synthesizer 1250 that can be the same as or similar to the neural speech synthesizer 450 of FIG.4, 650 of FIG.6, 750 of FIG.7, 850 of FIG. 8, 950 of FIG. 9, 1050 of FIG. 10, 1150 of FIG. 11, etc. The decoder 1210 can generate reconstructed audio 1255 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 455 of FIG.4, 655 of FIG.6, 755 of FIG. 7, 855 of FIG.8, 955 of FIG.9, 1055 of FIG.10, 1155 of FIG.11, etc.
[0244] The energy features 1210 of FIG.12 can be the same as or similar to the energy features 1110 of FIG.11. The spectral envelope features 1220 of FIG.12 can be the same as or similar to the spectral envelope features 1120 of FIG.11. The pitch information 1230 of FIG.12 can be the same as or similar to the pitch information 1130 of FIG.11.
[0245] In one illustrative example, the encoder 1205 can include a concatenation engine 1203 configured to generate a combined vector of features from the three separate inputs of the energy features 1210, spectral envelope features 1220, and pitch information 1230. The combined vector of features generated by the concatenation engine 1203 can be provided to a single FRAE autoencoder 1225, which can be used to generate and transmit a single latent representation z of the encoded combined feature vector that includes the energy, spectral, and pitch features. For example, the combined vector of features generated by the concatenation engine 1203 can include spectral envelope features 1220 (e.g., cepstrum, with potential omission of the cepstrum 0-th bin or replacement of the cepstrum 0-th bin with one or more energy features if present), energy features and / or energy correction features 1210 (e.g., in some cases, in addition to the cepstrum 0-th bin or as a replacement of the cepstrum 0-th bin), pitch information 1230 (e.g., f0 pitch information, pitch lag), V / UV classification, pitch correlation (if present), etc. In some cases, the encoder 1205 can use the concatenation engine 1203 to generate a combined feature vector with one or more additional features extracted from the input audio (e.g., in addition to the energy features 1210, spectral envelope features 1220, and pitch information 1230).
[0246] The decoder 1210 can include a single FRAE decoder 1245 that decodes the combined encoded feature vector received from the encoder 1205. In some aspects, the decoded combined feature vector can be provided directly to the neural speechQualcomm Docket No.2403757WO synthesizer 1250 as input for generating the reconstructed audio 1255. In some examples, the decoder 1210 can include a de-concatenation engine 1207 to separate the reconstructed energy, spectral envelope, and pitch features from the decoded combined feature vector output by the FRAE decoder 1245, prior to providing the separated features of the decoded combined feature vector as respective inputs to the neural speech synthesizer 1250 for generating the reconstructed audio 1255.
[0247] FIG. 13 is a block diagram illustrating an example of an audio codec system (e.g., generative voice codec system) 1300 that can be used to perform neural speech synthesis based on encoded vectors of various combinations of spectral energy features, energy features, and pitch information, in accordance with some examples. For example, a concatenation engine 1303 can be the same as the concatenation engine 1203 of FIG. 12, and can generate as output a combined feature vector that includes the energy features 1310, the spectral envelope features 1320, and the pitch information 1330. The encoder 1305 of FIG. 13 can further include a sub-vector splitting operation 1307 configured to split the combined feature vector from the concatenation engine 1303 into various different subsets, combinations or sub-combinations, etc., of the input feature types (e.g., energy features 1310, spectral envelope features 1320, and pitch information 1330). Each sub-vector split from the combined feature vector by the sub-vector splitting operation 1307 can be provided to a corresponding FRAE autoencoder instance 1315, 1325, …, etc. of the encoder 1305, which may generate a corresponding latent representation z, z2, …, etc., for transmission to a corresponding FRAE decoder instance 1347, 1345, …, etc. of the decoder 1310.
[0248] In some aspects, the generative voice codec system 1300 of FIG.13 can be the same as or similar to the generative voice codec system 400 of FIG.4, 600 of FIG.6, 700 of FIG.7, 800 of FIG.8, 900 of FIG.9, 1000 of FIG.10, 1100 of FIG.11, 1200 of FIG. 12, etc. The encoder 1305 of FIG.13 can be the same as or similar to the encoder 405 of FIG.4, 1105 of FIG.11, 1205 of FIG.12, etc. For example, the encoder 1305 can include spectral envelope feature extraction 1320 the same as or similar to the spectral envelope feature extraction 420 of FIG. 4, 1120 of FIG. 11, 1220 of FIG. 12, etc. Energy determination 1310 of FIG.13 can be the same as or similar to the energy determination 1110 of FIG.11 and / or 1210 of FIG.12. Pitch extraction 1330 of FIG.13 can be the same as or similar to the pitch extraction 430 of FIG.4, 1130 of FIG.11, 1230 of FIG.12, etc. A pitch quantization engine 1335 of FIG. 13 can be the same as or similar to the pitchQualcomm Docket No.2403757WO quantization engine 435 of FIG.4, 1135 and / or 1137 of FIG.11, etc.
[0249] The encoder 1305 of FIG.13 can include a first FRAE autoencoder 1325 and a second FRAE autoencoder 1315, which may be the same as or similar to one another. In some aspects, the FRAE autoencoder 1325 and / or the FRAE autoencoder 1315 can be the same as or similar to one or more of the FRAE autoencoder 425 of FIG.4, 1125 of FIG. 11 and / or 1115 of FIG.11, 1225 of FIG.12, etc.
[0250] The encoder 1305 of FIG. 13 can include a concatenation engine 1303 that is the same as or similar to the concatenation engine 1205 of FIG.12.
[0251] In some cases, the channel 1340 of FIG.13 can be the same as or similar to the channel 440 of FIG.4, the channel 640 of FIG.6, the channel 740 of FIG.7, the channel 840 of FIG.8, the channel 940 of FIG.9, the channel 1040 of FIG.10, the channel 1140 of FIG.11, the channel 1240 of FIG.12, etc.
[0252] The decoder 1310 of FIG. 13 can be the same as or similar to the decoder 410 of FIG.4, the decoder 610 of FIG.6, the decoder 710 of FIG.7, the decoder 810 of FIG. 8, the decoder 910 of FIG.9, the decoder 1010 of FIG.10, the decoder 1110 of FIG. 11, the decoder 1210 of FIG.12, etc. For example, the decoder 1310 can include a first FRAE decoder 1345 and a second FRAE decoder 1347, which may be the same as or similar to one another. In some aspects, the FRAE decoder 1345 and / or the FRAE decoder 1347 of FIG.13 can be the same as or similar to one or more of the FRAE decoder 445 of FIG.4, 645 of FIG.6, 745 of FIG.7, 845 of FIG.8, 945 of FIG.9, 1045 of FIG.10, 1145 of FIG. 11 and / or 1147 of FIG. 11, 1245 of FIG. 12, etc. The decoder 1310 can include a pitch dequantization engine 1338 that is the same as or similar to the pitch dequantization engine 638 of FIG.6, 738 of FIG.7, 838 of FIG.8, 938 of FIG.9, etc. The decoder 1310 can include a neural speech synthesizer 1350 that can be the same as or similar to the neural speech synthesizer 450 of FIG.4, 650 of FIG.6, 750 of FIG.7, 850 of FIG.8, 950 of FIG.9, 1050 of FIG.10, 1150 of FIG.11, 1250 of FIG.12, etc. The decoder 1310 can generate reconstructed audio 1355 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 455 of FIG.4, 655 of FIG.6, 755 of FIG.7, 855 of FIG. 8, 955 of FIG.9, 1055 of FIG.10, 1155 of FIG.11, 1255 of FIG.12, etc.
[0253] The decoder 1310 of FIG.13 can include a first de-concatenation engine 1349- 1 (e.g., associated with the first FRAE decoder 1345) and a second de-concatenation engine 1349-2 (e.g., associated with the second FRAE decoder 1347), which may be theQualcomm Docket No.2403757WO same as or similar to one another and / or may be the same as or similar to the de- concatenation engine 1207 of FIG.12, etc.
[0254] In some aspects, various combinations of the encoder features (e.g., energy features 1310, spectral envelope features 1320, pitch features 1330, etc.) can be grouped and / or maintained as separate feature sets, each group or feature set being coded with a corresponding separate FRAE autoencoder instance of the encoder 1305 and decoded with a corresponding separate FRAE decoder instance of the decoder 1310. The FRAE autoencoder-based coding performed by the FRAE autoencoder instances 1315, 1325 of the encoder 1305 can be implemented with or without MDC, and / or can be implemented with non-ML coding techniques with or without FEC and / or MDC.
[0255] In one illustrative example, the FRAE 1325 can be used to code and / or quantize a feature group (e.g. sub-vector of features obtained from the sub-vector splitting operation 1307) including pitch and pitch correlation information 1330. The FRAE 1315 can be used to code and / or quantize a second feature group (e.g., a second sub-vector of features obtained from the sub-vector splitting operation 1307) comprising spectral envelope features 1320. The quantization engine 1335 can be used to code and quantize a third feature group (e.g., a third sub-vector of features obtained from the sub-vector splitting operation 1307) including the energy features 1310 using classical non-ML coding techniques.
[0256] In another example, the sub-vector splitting operation 1307 can be applied to generate a first feature group or sub-vector comprising a first half of the spectral envelope features 1320 and the energy features 1320 (e.g., with the first sub-vector of spectral and energy features coded by the classical quantizer 1335), and to generate a second feature group or sub-vector comprising the remaining half of the spectral envelope features 1320 and the pitch information 1330 (e.g., with the second sub-vector of spectral and pitch features coded by one of the FRAE autoencoder instances 1315 or 1325 of the encoder 1305).
[0257] In some aspects, the systems and techniques can implement joint coding of features corresponding to multiple frames of audio data (e.g., multiple frames of a plurality of frames corresponding to the input audio signal to the encoder of the generative voice codec system, etc.). The joint coding can be performed with classical or non-ML- based quantization (e.g., using a quantization engine to code extracted features at theQualcomm Docket No.2403757WO encoder) and / or can be performed with ML-based quantization (e.g., using an FRAE autoencoder to code extracted features at the encoder). Joint coding of features corresponding to multiple frames of audio data can correspond to increased bit efficiency of the generative voice codec system.
[0258] FIG. 14 is a block diagram illustrating an example of an audio codec system (e.g., generative voice codec system) 1400 that can be used to perform neural speech synthesis based on using a first denoiser 1418 associated with spectral envelope feature extraction 1420 and a second denoiser 1419 associated with pitch extraction 1430, in accordance with some examples. In some aspects, the generative voice codec system 1400 of FIG.14 can be the same as or similar to the generative voice codec system 400 of FIG. 4, 600 of FIG. 6, 700 of FIG.7, 800 of FIG. 8, 900 of FIG. 9, 1000 of FIG. 10, 1100 of FIG.11, 1200 of FIG.12, 1300 of FIG.13, etc.
[0259] The encoder 1405 of FIG. 14 can be the same as or similar to the encoder 405 of FIG. 4, 1105 of FIG. 11, 1205 of FIG. 12, 1305 of FIG. 13, etc. For example, the encoder 1405 can include a first denoiser (e.g., spectral branch denoiser) 1418 that is the same as or similar to the denoiser 418 of FIG. 4, etc. The encoder 1405 can include a second denoiser (e.g., a pitch denoiser) 1419 that is the same as or similar to the denoiser 418 of FIG.4, and / or the spectral branch denoiser 1418, etc. The encoder 1405 can include spectral envelope feature extraction 1420 the same as or similar to the spectral envelope feature extraction 420 of FIG.4, 1120 of FIG.11, 1220 of FIG.12, 1320 of FIG.13, etc. The encoder 1405 can include pitch extraction 1430 the same as or similar to the pitch extraction 430 of FIG.4, 1130 of FIG.11, 1230 of FIG.12, 1330 of FIG.13, etc. A pitch quantization engine 1435 of FIG.14 can be the same as or similar to the pitch quantization engine 435 of FIG. 4, 1135 and / or 1137 of FIG. 11, 1335 of FIG. 13, etc. The encoder 1405 can include an FRAE autoencoder 1425 the same as or similar to the FRAE autoencoder 425 of FIG. 4, 1125 and / or 1115 of FIG. 11, 1225 of FIG. 12, 1315 and / or 1325 of FIG.13, etc.
[0260] In some cases, the channel 1440 of FIG.14 can be the same as or similar to the channel 440 of FIG.4, the channel 640 of FIG.6, the channel 740 of FIG.7, the channel 840 of FIG.8, the channel 940 of FIG.9, the channel 1040 of FIG.10, the channel 1140 of FIG. 11, the channel 1240 of FIG. 12, the channel 1340 of FIG. 13, etc. The decoder 1410 of FIG. 14 can be the same as or similar to the decoder 410 of FIG.4, the decoderQualcomm Docket No.2403757WO 610 of FIG.6, the decoder 710 of FIG.7, the decoder 810 of FIG.8, the decoder 910 of FIG. 9, the decoder 1010 of FIG. 10, the decoder 1110 of FIG. 11, the decoder 1210 of FIG.12, the decoder 1310 of FIG.13, etc. For example, the decoder 1410 can include a spectral feature FRAE decoder 1445 the same as or similar to the FRAE decoder 445 of FIG.4, 645 of FIG.6, 745 of FIG.7, 845 of FIG.8, 945 of FIG.9, 1045 of FIG.10, 1145 of FIG. 11 and / or 1147 of FIG. 11, 1245 of FIG. 12, 1345 and / or 1347 of FIG. 13, etc. The decoder 1410 can include a neural speech synthesizer 1450 that can be the same as or similar to the neural speech synthesizer 450 of FIG. 4, 650 of FIG. 6, 750 of FIG. 7, 850 of FIG.8, 950 of FIG.9, 1050 of FIG.10, 1150 of FIG.11, 1250 of FIG.12, 1350 of FIG.13, etc. The decoder 1410 can generate reconstructed audio 1455 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 455 of FIG. 4, 655 of FIG.6, 755 of FIG.7, 855 of FIG.8, 955 of FIG.9, 1055 of FIG.10, 1155 of FIG.11, 1255 of FIG.12, 1355 of FIG.13, etc. In some aspects, the decoder 1410 of FIG.14 can include an LPC 1452 between the output of the neural speech synthesizer 1450 and the reconstructed audio 1455 output from the decoder 1410. In some examples, the LPC 1452 of FIG.14 can be the same as or similar to the LPC 452 of FIG.4, 652 of FIG.6, 752 of FIG.7, 852 of FIG.8, 954 of FIG.9, etc.
[0261] In some cases, the generative voice codec system 1400 of FIG. 14 can be the same as the generative voice codec system 400 of FIG. 4, with the addition of the additional, second denoiser 1419 (e.g., pitch denoiser, denoiser for pitch, etc.) between the input audio signal 1415 and the pitch extraction engine 1430.
[0262] For example, in some aspects, separate denoisers 1418 and 1419 can be included in and utilized by the encoder 1405 to perform denoising of the input audio signal 1415 prior to processing for the spectral envelope feature extraction 1420 and the pitch extraction 1430. In some cases, the pitch denoiser 1419 can be configured to perform more aggressive denoising than the denoiser 1418 associated with the spectral envelope feature extraction 1420 and / or the spectral processing branch of the encoder 1405. The more aggressive denoising for pitch, applied by the pitch denoiser 1419, can improve the robustness of pitch estimation performed by the pitch extraction engine 1430 based on the denoised audio signal output from the pitch denoiser 1419.
[0263] In some aspects, the pitch denoiser 1419 can operate on the original input audio signal 1415 (e.g., the denoiser 1418 and the pitch denoiser 1419 may both receive theQualcomm Docket No.2403757WO same audio signal 1415 as input, and may operate in parallel within the encoder 1405). In another example, the pitch denoiser 1419 can operate in series (e.g., in sequence) with the denoiser 1418. For example, the input to the pitch denoiser 1419 can be the output of the denoiser 1418, and the pitch denoiser 1419 does not receive the encoder input audio signal 1415 (e.g., the output of the denoiser 1418 can be provided to the spectral envelope feature extraction 1420 and the pitch denoiser 1419 in parallel). In some aspects, the pitch extraction engine 1430 can receive as input the more aggressively de-noised version of the input audio signal 1415 that is generated using the pitch denoiser 1419 and / or the denoiser 1418. The pitch extraction engine 1430 can be configured to perform pitch estimation using one or more DSP techniques, one or more ML or neural network-based techniques, and / or various combinations thereof of DSP and ML or neural network techniques for pitch estimation.
[0264] FIG. 15 is a block diagram illustrating an example of an audio codec system (e.g., generative voice codec system) 1550 that can be used to perform neural speech synthesis based on using a front-end non-speech detector (FNSD) 1514 to switch between a neural synthesizer audio codec path for encoding and / or decoding speech signals and an additional audio codec path for encoding and / or decoding non-speech signals, in accordance with some examples. In some aspects, the generative voice codec system 1500 of FIG.15 can be the same as or similar to the generative voice codec system 400 of FIG. 4, 600 of FIG. 6, 700 of FIG.7, 800 of FIG. 8, 900 of FIG. 9, 1000 of FIG. 10, 1100 of FIG. 11, 1200 of FIG. 12, 1300 of FIG. 13, 1400 of FIG. 14, etc. The encoder 1505 of FIG.15 can be the same as or similar to the encoder 405 of FIG.4, 1105 of FIG.11, 1205 of FIG. 12, 1305 of FIG. 13, 1405 of FIG. 14, etc. For example, the encoder 1505 can include a denoiser (e.g., spectral branch denoiser) 1518 that is the same as or similar to the denoiser 418 of FIG. 4, 1418 and / or 1419 of FIG. 14, etc. The encoder 1505 can include spectral envelope feature extraction 1520 the same as or similar to the spectral envelope feature extraction 420 of FIG.4, 1120 of FIG.11, 1220 of FIG.12, 1320 of FIG. 13, 1420 of FIG.14, etc. The encoder 1505 can include pitch extraction 1530 the same as or similar to the pitch extraction 430 of FIG. 4, 1130 of FIG. 11, 1230 of FIG. 12, 1330 of FIG.13, 1430 of FIG.14, etc. A pitch quantization engine 1535 of FIG.15 can be the same as or similar to the pitch quantization engine 435 of FIG. 4, 1135 and / or 1137 of FIG. 11, 1335 of FIG.13, 1435 of FIG. 14, etc. The encoder 1505 can include an FRAE autoencoder 1525 the same as or similar to the FRAE autoencoder 425 of FIG. 4, 1125Qualcomm Docket No.2403757WO and / or 1115 of FIG.11, 1225 of FIG.12, 1315 and / or 1325 of FIG.13, 1425 of FIG.14, etc.
[0265] In some cases, the channel 1540 of FIG.15 can be the same as or similar to the channel 440 of FIG.4, the channel 640 of FIG.6, the channel 740 of FIG.7, the channel 840 of FIG.8, the channel 940 of FIG.9, the channel 1040 of FIG.10, the channel 1140 of FIG. 11, the channel 1240 of FIG. 12, the channel 1340 of FIG. 13, the channel 1440 of FIG.14, etc. The decoder 1510 of FIG.15 can be the same as or similar to the decoder 410 of FIG.4, the decoder 610 of FIG.6, the decoder 710 of FIG.7, the decoder 810 of FIG.8, the decoder 910 of FIG.9, the decoder 1010 of FIG.10, the decoder 1110 of FIG. 11, the decoder 1210 of FIG.12, the decoder 1310 of FIG. 13, the decoder 1410 of FIG. 14, etc. For example, the decoder 1510 can include a spectral feature FRAE decoder 1545 that is the same as or similar to the FRAE decoder 445 of FIG.4, 645 of FIG. 6, 745 of FIG. 7, 845 of FIG. 8, 945 of FIG. 9, 1045 of FIG. 10, 1145 of FIG. 11 and / or 1147 of FIG.11, 1245 of FIG.12, 1345 and / or 1347 of FIG.13, 1445 of FIG.14, etc. The decoder 1510 can include a neural speech synthesizer 1550 that can be the same as or similar to the neural speech synthesizer 450 of FIG.4, 650 of FIG.6, 750 of FIG.7, 850 of FIG.8, 950 of FIG.9, 1050 of FIG.10, 1150 of FIG.11, 1250 of FIG.12, 1350 of FIG.13, 1450 of FIG. 14, etc. The decoder 1510 can generate reconstructed audio 1555 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 455 of FIG.4, 655 of FIG.6, 755 of FIG.7, 855 of FIG.8, 955 of FIG.9, 1055 of FIG.10, 1155 of FIG.11, 1255 of FIG.12, 1355 of FIG. 13, 1455 of FIG.14, etc. In some aspects, the decoder 1510 of FIG.15 can include an LPC 1552 between the output of the neural speech synthesizer 1550 and the reconstructed audio 1555 output from the decoder 1510. In some examples, the LPC 1552 of FIG.15 can be the same as or similar to the LPC 452 of FIG. 4, 652 of FIG.6, 752 of FIG.7, 852 of FIG.8, 954 of FIG.9, 1452 of FIG.14, etc.
[0266] In one illustrative example, the encoder 1505 of FIG.15 can include a front-end non-speech detector (FNSD) 1514. For example, the FNSD 1514 can receive as input the audio signal 1515 obtained or received by the encoder 1505, and can perform speech and / or non-speech detection to determine whether the input audio signal 1515 includes speech or does not include speech (e.g., includes non-speech audio).
[0267] Input audio signals 1515 that are detected as speech by the FNSD 1514 can follow a speech coding path of the encoder 1505. For example, the speech coding pathQualcomm Docket No.2403757WO can be utilized based on the FNSD 1514 outputting an FNSD decision indicative of detecting the input audio signal 1515 is a speech signal. The FNSD decision indicative of the FNSD 1514 detecting speech within the input audio 1515 can be transmitted over the channel 1540 to the decoder 1510, and can be used to control a switch 1504 included in the encoder 1505.
[0268] The switch 1504 can be used to switch the input audio signal 1515 between the speech coding path of the encoder 1505 (e.g., in response to an FNSD decision indicative of the FNSD 1514 detecting speech) and a non-speech or general audio coding path of the encoder 1505 (e.g., in response to an FNSD decision indicative of the FNSD 1514 detecting non-speech or general audio).
[0269] The speech coding path of the encoder 1505 can be the same as or similar to the coding path of the encoder 405 of FIG. 4 and / or the encoder 1405 of FIG. 14 (e.g., including a denoiser 1518, spectral envelope feature extraction 1520, spectral FRAE 1525, pitch extraction 1530, and pitch quantization 1535, etc.).
[0270] When the FNSD 1514 detects speech signals, and generates a corresponding FNSD decision indicative of speech, the switch 1504 can couple the input audio signal 1515 to the denoiser 1518 at the beginning of the speech coding path of the encoder 1505.
[0271] When the FNSD 1514 detects non-speech or general audio signals, and generates a corresponding FNSD decision indicative of non-speech, the switch 1504 can couple the input audio signal 1515 to the non-speech or general audio encoder 1590. When the FNSD decision of the FNSD 1514 is indicative of non-speech detected in the input audio signal 1515, the denoiser 1518 and remaining components of the speech coding path of the encoder 1505 do not receive the input audio signal 1515, are not used to process or code (e.g., encode) the audio signal 1515 or any extracted features thereof, and do not transmit information over the channel 1540 to the decoder 1510.
[0272] Input audio signals 1515 that are detected as non-speech signals by the FNSD 1514 can follow an alternate audio coding path using a non-speech audio encoder 1590 of the encoder 1505 and a non-speech audio decoder 1595 of the decoder 1510. The non- speech audio encoder 1590 can be implemented as an alternate audio codec for coding non-speech and / or general audio signals. The non-speech audio encoder 1590 can be implemented as a non-ML audio encoder, a DSP, an ML or neural network-based audio encoder, and / or various combinations thereof, etc.Qualcomm Docket No.2403757WO
[0273] The coded output generated by the non-speech audio encoder 1590 on the non- speech coding path of the encoder 1505 can be quantized and transmitted over the channel 1540 from the encoder 1505 to the decoder 1510. The decoder 1510 can include a switch 1585 that can be controlled based on receiving the FNSD decision (e.g., indicative of speech or non-speech detection by FNSD 1514, and indicative of a corresponding use of either the speech coding path or non-speech coding path of the encoder 1505, decoder 1510, and generative voice codec system 1500).
[0274] For example, the received FNSD decision received over the channel 1540 by the decoder 1510 and from the FNSD 1514 can be decoded and used to switch the decoder 1510 between a speech coding (e.g., decoding) path of the decoder 1510, utilizing the FRAE decoder 1545, neural speech synthesizer 1550, and LPC 1552 to generate the reconstructed audio (e.g., synthesized speech) 1555, and a non-speech or general audio coding (e.g., decoding) path of the decoder 1510, utilizing a non-speech audio decoder 1595 associated with the non-speech audio encoder 1590 of the encoder 1505. The non- speech audio decoder 1595 can be implemented as an alternate audio codec for coding (e.g., decoding) non-speech and / or general audio signals that are encoded by the non- speech audio encoder 1590. The non-speech audio decoder 1595 can be implemented as a non-ML audio decoder, a DSP, an ML or neural network-based audio decoder, and / or various combinations thereof, etc. The non-speech audio decoder 1595 can generate as output a reconstructed non-speech audio 1598, based on decoding the encoded non- speech audio received over the channel 1540 from the non-speech audio encoder 1590 included in the non-speech coding path of the encoder 1505.
[0275] In one illustrative example, audio codec data from the non-speech audio encoder 1590 can be transmitted over the channel 1540 to the decoder 1510 only if the current audio frame (e.g., current frame of the input audio signal 1515) is detected by the FNSD 1514 as non-speech (e.g., the FNSD decision for the current frame is indicative of non- speech). The remaining portion(s) of the data may be transmitted only if the current frame is detected by the FNSD 1514 as speech (e.g., the FNSD decision for the current frame is indicative of speech, and the encoder 1505 uses the speech coding audio path to generate FRAE 1525 encoded spectral envelope features and pitch information 1530, where both the encoded spectral envelope features and the encoded pitch information comprise the remaining portion(s) of the data that are transmitted over the channel 1540 to the decoder 1510- based on the FNSD 1514 detecting speech in the input audio signal 1515).Qualcomm Docket No.2403757WO
[0276] In some examples, the encoder 1505 can be configured to process the input audio signal 151 using the denoiser 1518, VAD, and pitch extraction engine 1530 prior to processing the input audio signal 1515 with the FNSD 1514 to generate a speech or non- speech FNSD decision. For example, in some cases, the respective outputs of the speech coding path of the encoder 1505 can be provided as additional inputs to the FNSD 1514. In some aspects, the FNSD 1514 can be configured to generate the speech or non-speech FNSD decision based on inputs comprising the input audio signal 1515, a denoised version of the input audio signal 1515 generated by the denoiser 1518, extracted spectral envelope features from the spectral envelope feature extraction 1520, extracted pitch features or pitch information from the pitch extraction 1530, etc.
[0277] In some examples, the systems and techniques can implement a generative voice codec system with an encoder that includes a voice activity detection (VAD) engine on the speech path. For example, the VAD engine can be configured to analyze the input audio signal to the encoder and distinguish silence from active speech.
[0278] In some cases, the VAD engine can be implemented by or within the pitch extraction engine. For example, the pitch extraction engine 1530 (or any other pitch extraction engine described in FIGS. 1-15) can include the VAD engine and / or can perform or implement VAD to distinguish silence from active speech. In some examples, a separate VAD engine can be included in the encoder and the separate VAD engine can utilize the extracted pitch features or pitch information (e.g., from the pitch extraction engine) as additional inputs for generating a VAD decision indicative of silence or active speech. For example, a separate VAD engine can use pitch correlation information from the pitch extraction engine to determine a VAD decision indicative of silence or active speech in the input audio signal or frames thereof.
[0279] In some examples, a VAD decision (e.g., silence or active speech) can be used by the encoder of the generative voice codec system to enable discontinuous transmission with optional silence encoding. In some cases, based on the VAD decision indicating that the current frame of input audio is silence or inactive speech, the encoder 1505 can be configured not to transmit over the channel 1540 to the decoder 1510. In some examples, based on the VAD decision indicating that the current frame of input audio is silence or inactive speech, the encoder 1505 can be configured to generate, code, and transmit a coarse, low bit-rate silence encoding to the decoder 1510 via the channel 1540.Qualcomm Docket No.2403757WO
[0280] In examples where the encoder 1505 uses a VAD engine and / or VAD decision to transmit silence encodings over the channel 1540 to the decoder 1510, the encoder 1505 can be configured to transmit silence encodings sparsely (e.g., every N-th frame, not every frame or not in consecutive frames, etc., among various other discontinuous transmission schemes or techniques that can be implemented for silence encodings transmitted between the encoder 1505 and the decoder 1510 based on a VAD decision determined by the encoder 1505).
[0281] In some examples, the encoder 1505 can transmit over the channel 1540 to the decoder 1510 silence encodings that include coarse spectral representations of the background noise characteristics, such as low-order LPC coefficients, coarsely quantized (e.g., quantized with low bit-rate) spectral envelope features, etc.
[0282] In some cases, the decoder 1510 can include a silence decoder. In some cases, a silence decoder included in the decoder 1510 of the generative voice codec system 1500 can be implemented as a comfort noise generator. The silence decoder included in the decoder 1510 can use the received silence encoding (e.g., received over the channel 1540 from the encoder 1505, in response to a VAD decision indicative of silence or inactive speech). In some cases, the silence decoder included in the decoder 1510 can use the received silence encoding to update one or more comfort noise generation (CNG) parameters for implementing the comfort noise generator. In some examples, the silence decoder and / or CNG included in the decoder 1510 can be used to approximate or approximately replicate coarse background noise characteristics of the silence or inactive speech detected by the VAD engine or VAD decision for the currently coded frame of the input audio signal 151. For example, the silence decoder and / or CNG can approximate coarse background noise characteristics such as spectral envelope, and can avoid having complete digital silence in the reconstructed audio output signal 1555 during periods of silence or inactive speech in the input audio signal 1515 (e.g., as complete digital silence in the reconstructed audio output signal 1555 may be perceptually detrimental).
[0283] FIG. 16 is a flowchart diagram illustrating an example of a process 1600 for processing one or more audio samples. For example, the process 1600 can correspond to a process for encoding one or more audio samples.
[0284] In some examples, the process 1600 can be performed by a computing device or apparatus or a component or system (e.g., one or more chipsets, one or more processorsQualcomm Docket No.2403757WO such as one or more CPUs, DSPs, NPUs, NSPs, microcontrollers, ASICs, FPGAs, programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc., any combination thereof, and / or other component or system) of the computing device or apparatus. The operations of the process 1600 may be implemented as software components that are executed and run on one or more processors (e.g., processor 1810 of FIG. 18 or other processor(s)). In some examples, the process 1600 can be performed by an audio coding system or component thereof, including any of the audio coding systems and / or components thereof of FIGS.1-15. For example, the process 1600 can be performed by one or more of the voice encoder 252 of FIG.2B, the voice encoder 272 of FIG. 2C, the encoder of FIG. 3, the encoder 405 of the generative voice codec 400 of FIG. 4, the audio codec system(s) 500 and / or 570 of FIGS. 5A-5B, the encoder 1005 of the generative voice codec 1000 of FIG.10, the encoder 1105 of the generative voice codec 1100 of FIG.11, the encoder 1205 of the generative voice codec 1200 of FIG. 12, the encoder 1305 of the generative voice codec 1300 of FIG. 13, the encoder 1405 of the generative voice codec 1400 of FIG.14, and / or the encoder 1505 of the generative voice codec 1500 of FIG.15, etc.
[0285] In some aspects, the process 1600 can be performed by a UE, smartphone, mobile computing device, user computing device, etc. The process 1600 may be performed by an apparatus that may be a mobile device (e.g., a mobile phone), a network- connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, or other type of computing device. The operations of the process 1600 may be implemented as software components that are executed and run on one or more processors (e.g., processor 1810 of FIG.18, and / or other processor(s)).
[0286] At block 1602, the apparatus (or component thereof) can determine spectral envelope features of the audio, wherein the spectral envelope features are determined based on processing the audio using a filter bank including a plurality of filters.
[0287] In some examples, the spectral envelope features can be determined using one or more of the spectral envelope feature extraction engines 420 of FIG. 4, 1120 of FIG. 11, 1220 of FIG. 12, 1320 of FIG. 13, 1420 of FIG. 14, 1520 of FIG. 15, etc. The filter bank can be the same as or similar to the filter bank 532 of FIG.5A, and the plurality of filters can be the same as or similar to the plurality of filters 580 of FIG. 5B, includingQualcomm Docket No.2403757WO the filters 582-1, …, 582-N of FIG. 5B, etc. In some examples, the spectral envelope features can include one or more linear prediction coefficients (LPCs) corresponding to the audio.
[0288] In some cases, the spectral envelope features do not include pitch information of the audio. In some examples, to determine the spectral envelope features, the apparatus (or component thereof) can be configured to determine a plurality of cepstral coefficients based on processing the audio using the plurality of filters included in the filter bank. For example, the plurality of filters can comprise a first number of filters, and a number of cepstral coefficients included in the plurality of cepstral coefficients can be equal to the first number. For example, N cepstral coefficients can be generated corresponding to the N filters included in the plurality of filters 580 of the filter bank 532 of FIGS.5A-5B. The apparatus (or component thereof) can truncate the plurality of cepstral coefficients to obtain a set of truncated cepstral coefficients comprising a subset of the plurality of cepstral coefficients. In some examples, the truncation can be performed using the DCT 555 of FIG. 5A and the truncation operation 560 of FIG. 5A. For example, the set of truncated cepstral coefficients can comprise the M coefficients that are a subset of the N coefficients of FIG. 5A, where M < N. In some cases, the set of truncated cepstral coefficients includes a second number of cepstral coefficients, the second number less than the first number. In some examples, the set of truncated cepstral coefficients does not include pitch information of the audio, based on a difference between the first number and the second number.
[0289] In some cases, to truncate the plurality of cepstral coefficients, the apparatus (or component thereof) can be configured to compute a discrete cosine transform (DCT) of energy per band information associated with the plurality of filters of the filter bank. For example, the DCT can be the same as or similar to the DCT 555 of FIG.5. The apparatus (or component thereof) can skip computation of higher order DCT coefficients of the DCT corresponding to the difference between the first number and the second number.
[0290] In some examples, the apparatus (or component thereof) can be configured to remove pitch information of the audio based on truncating the plurality of cepstral coefficients to obtain the set of truncated cepstral coefficients.
[0291] At block 1604, the apparatus (or component thereof) can determine one or more pitch features of the audio. For example, the one or more pitch features can be determinedQualcomm Docket No.2403757WO using one or more of the pitch extraction engine 430 of FIG.4, 1130 of FIG.11, 1230 of FIG.12, 1330 of FIG.13, 1430 of FIG.14, 1530 of FIG.15, etc.
[0292] In some examples, the one or more pitch features are indicative of a pitch frequency associated with the audio or a pitch lag associated with the audio. In some examples, the one or more pitch features are indicative of one or more of pitch correlation information associated with the audio, or a voiced or unvoiced classification associated with the audio.
[0293] At block 1606, the apparatus (or component thereof) can generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the spectral envelope features, wherein the first encoded representation is based on state feedback between a decoder portion and an encoder portion of the first neural network- based autoencoder.
[0294] For example, the first neural network-based autoencoder can be the same as or similar to any one of the feedback recurrent autoencoders (FRAE) autoencoder instances of FIGS. 4-15. In some cases, the first neural network-based autoencoder is a feedback recurrent autoencoder (FRAE) and the first encoded representation comprises a latent associated with the first neural network-based autoencoder and the state feedback. For example, the latent can be the same as or similar to the latent z associated with the FRAE autoencoder instances of FIGS.4-15 and the state feedback can be the same as or similar to the state feedback information h associated with the FRAE autoencoder instances of FIGS.4-15.
[0295] At block 1608, the apparatus (or component thereof) can generate a second encoded representation corresponding to the one or more pitch features.
[0296] In some examples, the first encoded representation and the second encoded representation do not include absolute phase information or pitch pulse location information associated with the audio. In some cases, to generate the second encoded representation corresponding to the one or more pitch features, the apparatus (or component thereof) can be configured to generate quantized pitch information based on performing vector quantization of the one or more pitch features. For example, the quantized pitch information can be generated using a pitch quantization engine the same as or similar to one or more of the pitch quantization engine 435 of FIG.4, 1135 or 1137 of FIG.11, 1335 of FIG.13, 1435 of FIG.14, 1535 of FIG.15, etc.Qualcomm Docket No.2403757WO
[0297] In some cases, to generate the second encoded representation corresponding to the one or more pitch features, the apparatus (or component thereof) can be configured to process the one or more pitch features using a second neural network-based autoencoder, wherein the second encoded representation is generated based on state feedback between a respective decoder portion and a respective encoder portion of the second neural network-based autoencoder.
[0298] In some examples, to generate the second encoded representation corresponding to the one or more pitch features, the apparatus (or component thereof) can be configured to process the one or more pitch features using the first neural network-based autoencoder.
[0299] In some examples, the apparatus (or component thereof) can transmit the first encoded representation corresponding to the spectral envelope features and the second encoded representation corresponding to the one or more pitch features.
[0300] For example, the apparatus (or component thereof) can transmit the first and second encoded representations over a channel the same as or similar to one or more of the channels of FIGS. 4-15. The first and second encoded representations can be transmitted to a decoder, including a decoder of a generative voice codec system the same as or similar to one or more of the generative voice codec systems of FIGS.4-15.
[0301] FIG. 17 is a flowchart diagram illustrating an example of a process 1700 for processing one or more audio samples. For example, the process 1700 can correspond to a process for decoding one or more audio samples (e.g., decoding encoded audio samples and / or encoded audio data, etc.).
[0302] In some examples, the process 1700 can be performed by a computing device or apparatus or a component or system (e.g., one or more chipsets, one or more processors such as one or more CPUs, DSPs, NPUs, NSPs, microcontrollers, ASICs, FPGAs, programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc., any combination thereof, and / or other component or system) of the computing device or apparatus. The operations of the process 1700 may be implemented as software components that are executed and run on one or more processors (e.g., processor 1810 of FIG. 18 or other processor(s)). In some examples, the process 1700 can be performed by an audio coding system or component thereof, including any of the audio coding systems and / or components thereof of FIGS.1-15.Qualcomm Docket No.2403757WO
[0303] For example, the process 1700 can be performed by one or more of the voice decoder 254 of FIG. 2B, the voice decoder 274 of FIG. 2C, the decoder of FIG. 3, the decoder 410 of the generative voice codec 400 of FIG. 4, the audio codec system(s) 500 and / or 570 of FIGS. 5A-5B, the decoder 610 of the generative voice codec 600 of FIG. 6, the decoder 710 of the generative voice codec 700 of FIG. 7, the decoder 810 of the generative voice codec 800 of FIG.8, the decoder 910 of the generative voice codec 900 of FIG. 9, the decoder 1010 of the generative voice codec 1000 of FIG. 10, the decoder 1110 of the generative voice codec 1100 of FIG. 11, the decoder 1210 of the generative voice codec 1200 of FIG. 12, the decoder 1310 of the generative voice codec 1300 of FIG. 13, the decoder 1410 of the generative voice codec 1400 of FIG. 14, and / or the decoder 1510 of the generative voice codec 1500 of FIG.15, etc.
[0304] In some aspects, the process 1700 can be performed by a UE, smartphone, mobile computing device, user computing device, etc. The process 1700 may be performed by an apparatus that may be a mobile device (e.g., a mobile phone), a network- connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, or other type of computing device. The operations of the process 1700 may be implemented as software components that are executed and run on one or more processors (e.g., processor 1810 of FIG.18, and / or other processor(s)).
[0305] At block 1702, the apparatus (or component thereof) can receive a first encoded representation corresponding to spectral envelope features of the audio. For example, the first encoded representation can be received over a channel associated with an encoder of a generative voice codec system, including any one of the channels of FIGS. 4-15 and corresponding to the generative voice codec systems of FIGS. 4-15. In some examples, the spectral envelope features of the audio can be associated with a spectral envelope feature extraction engine included in the encoder, such as one or more of the spectral envelope feature extraction engines 420 of FIG.4, 1120 of FIG.11, 1220 of FIG.12, 1320 of FIG.13, 1420 of FIG.14, 1520 of FIG.15, etc.
[0306] In some examples, the first encoded representation corresponding to the spectral envelope features does not include pitch information associated with the audio.
[0307] At block 1704, the apparatus (or component thereof) can receive a second encoded representation corresponding to one or more pitch features of the audio. TheQualcomm Docket No.2403757WO second encoded representation can be received over the same channel as the first encoded representation. The one or more pitch features can be determined by a pitch extraction engine included in the encoder of the generative voice codec system, such as one or more of the pitch extraction engine 430 of FIG. 4, 1130 of FIG. 11, 1230 of FIG. 12, 1330 of FIG.13, 1430 of FIG.14, 1530 of FIG.15, etc.
[0308] In some examples, the one or more pitch features are indicative of a pitch frequency associated with the audio or a pitch lag associated with the audio. In some examples, the one or more pitch features are indicative of one or more of pitch correlation information associated with the audio, or a voiced or unvoiced classification associated with the audio. In some cases, the second encoded representation corresponding to the one or more pitch features can be quantized pitch information generated by a pitch quantization engine of the encoder of the generative voice codec system, which may be the same as or similar to one or more of the pitch quantization engine 435 of FIG.4, 1135 or 1137 of FIG.11, 1335 of FIG.13, 1435 of FIG.14, 1535 of FIG.15, etc.
[0309] In some cases, the second encoded representation comprises a quantized pitch encoding indicative of one or more of a pitch frequency, a pitch lag, a voiced or unvoiced classification, or pitch correlation information associated with the audio.
[0310] At block 1706, the apparatus (or component thereof) can generate reconstructed spectral envelope features corresponding to the audio, based on processing the first encoded representation using a first neural network-based decoder.
[0311] For example, the reconstructed spectral envelope features can be generated by a feedback recurrent autoencoder (FRAE) decoder included in the apparatus. In some cases, the first neural network-based decoder can be the same as or similar to the FRAE decoder 445 of FIG. 4, 645 of FIG. 6, 745 of FIG.7, 845 of FIG. 8, 945 of FIG. 9, 1045 of FIG. 10, 1115 and / or 1125 of FIG.11, 1225 of FIG.12, 1315 and / or 1325 of FIG.13, 1425 of FIG.14, 1525 of FIG.15, etc.
[0312] At block 1708, the apparatus (or component thereof) can generate a synthesized audio output signal based on processing the reconstructed spectral envelope features and the second encoded representation using a neural network-based signal synthesizer.
[0313] For example, the neural network-based signal synthesizer can be the same as or similar to one or more of the neural network-based signal synthesizer 450 of FIG.4, 650Qualcomm Docket No.2403757WO of FIG.6, 750 of FIG.7, 850 of FIG.8, 950 of FIG.9, 1050 of FIG.10, 1150 of FIG.11, 1250 of FIG.12, 1350 of FIG.13, 1450 of FIG.14, 1550 of FIG.15, etc. In some cases, the neural network-based signal synthesizer can be an NHV-based neural synthesizer such as the NHV-based neural synthesizer 950 of FIG. 9. In some examples, the neural network-based signal synthesizer can be an LPC network-based neural synthesizer, such as the LPC network-based neural synthesizer 1050 of FIG. 10. In some examples, the neural network-based signal synthesizer is a neural homomorphic vocoder (NHV).
[0314] In some examples, the synthesized audio output signal is a reconstructed signal corresponding to the audio. For example, the synthesized audio output signal can be the same as or similar to any one of the reconstructed audio signals of FIGS. 4-15. In some cases, the synthesized audio output is generated based on processing the reconstructed spectral envelope features using a neural network filter estimator of the NHV. For example, the neural network filter estimator can be the same as or similar to the neural filter estimator 952 of the NHV-based neural synthesizer 950 of FIG.9.
[0315] In some cases, the synthesized audio output is generated based on processing the reconstructed spectral envelope features and dequantized pitch features using the neural network filter estimator of the NHV, the dequantized pitch features based on the second encoded representation. For example, the dequantized pitch features can be generated using a pitch dequantization engine, such as the pitch dequantization engine 938 of the decoder 910 of FIG.9. In some examples, the reconstructed spectral envelope features can be obtained using an FRAE decoder, such as the FRAE decoder 945 of the decoder 910 of FIG. 9. In some cases, the synthesized audio output generated using the neural network filter estimator of the NHV can be the same as or similar to the reconstructed audio (e.g., synthesized speech) 955 of FIG. 9, generated using the neural filter estimator 952 of the NHV-based neural synthesizer 950 of FIG.9.
[0316] In some examples, to generate the synthesized audio output signal, the apparatus (or component thereof) can be configured to compute an inverse discrete cosine transform (DCT) of the reconstructed spectral envelope features to obtain reconstructed filter bank energies corresponding to the audio. For example, the inverse DCT can be the same as or similar to the inverse DCT 626 of FIG.6, 726 of FIG.7, 826 of FIG.8, etc. In some cases, the apparatus (or component thereof) can process the reconstructed filter bank energiesQualcomm Docket No.2403757WO and the second encoded representation using the neural network-based signal synthesizer to generate the synthesized audio output signal.
[0317] In some examples, to generate the synthesized audio output signal, the apparatus (or component thereof) can be configured to compute an inverse discrete cosine transform (DCT) of the reconstructed spectral envelope features to obtain reconstructed filter bank energies corresponding to the audio, and can further be configured to apply an inverse companding function to the reconstructed filter bank energies to obtain un-companded filter bank energies corresponding to the audio. For example, the inverse companding function can be the same as or similar to the inverse companding function 745 of FIG.7, 845 of FIG.8, etc. The apparatus (or component thereof) can process the un-companded filter bank energies and the second encoded representation using the neural network- based signal synthesizer to generate the synthesized audio output signal.
[0318] In some examples, the inverse DCT is an inverse of a DCT function associated with the first encoded representation. For example, the inverse DCT can be an inverse of the DCT function 555 of FIG.5A. In some examples, the inverse companding function is an inverse of a companding function associated with the first encoded representation. For example, the inverse companding function can be an inverse of the companding function 540 of FIG.5A.
[0319] In some cases, an apparatus used to implement the process 1600 and / or the process 1700 can include one or more microphones configured to obtain the one or more audio samples. In some cases, the apparatus further comprises one or more microphones configured to capture the one or more audio samples for speech synthesis.
[0320] In some cases, the processes described herein (e.g., the process 1600, process 1700, and / or any other process described herein) may be performed by a computing device or apparatus. In one example, the process 1600, the process 1700, and / or other technique or process described herein can be performed by a computing system having an architecture according to any of FIGS.1-15. In another example, the process 1600, the process 1700, and / or other technique or process described herein can be performed by the computing system 1800 shown in FIG. 18. For instance, a computing device with the computing device architecture of the computing system 1800 shown in FIG. 18 can implement the operations of the process 1600, can implement the operations of theQualcomm Docket No.2403757WO process 1700, and / or can implement one or more of the components and / or operations described herein with respect to any of FIGS.1-15.
[0001] In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component(s) that are configured to carry out the steps of processes described herein. In some examples, the computing device may include a display, one or more network interfaces configured to communicate and / or receive the data, any combination thereof, and / or other component(s). The one or more network interfaces may be configured to communicate and / or receive wired and / or wireless data, including data according to the 3G, 4G, 5G, and / or other cellular standard, data according to the WiFi (802.11x) standards, data according to the BluetoothTMstandard, data according to the Internet Protocol (IP) standard, and / or other types of data.
[0321] The components of the computing device may be implemented in circuitry. For example, the components may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein.
[0322] The process 1600 and the process 1700 are illustrated as logical flow diagrams, the operation of which represents a sequence of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the processes.Qualcomm Docket No.2403757WO
[0323] Additionally, the process 1600, the process 1700, and / or other process described herein, may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine- readable storage medium may be non-transitory.
[0324] FIG. 18 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. In particular, FIG.18 illustrates an example of computing system 1800, which may be for example any computing device making up internal computing system, a remote computing system, a camera, or any component thereof in which the components of the system are in communication with each other using connection 1805. Connection 1805 may be a physical connection using a bus, or a direct connection into processor 1810, such as in a chipset architecture. Connection 1805 may also be a virtual connection, networked connection, or logical connection.
[0325] In some aspects, computing system 1800 is a distributed system in which the functions described in this disclosure may be distributed within a datacenter, multiple data centers, a peer network, etc. In some aspects, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some aspects, the components may be physical or virtual devices.
[0326] Example system 1800 includes at least one processing unit (CPU or processor) 1810 and connection 1805 that communicatively couples various system components including system memory 1815, such as read-only memory (ROM) 1820 and random access memory (RAM) 1825 to processor 1810. Computing system 1800 may include a cache 1815 of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1810.
[0327] Processor 1810 may include any general-purpose processor and a hardware service or software service, such as services 1832, 1834, and 1836 stored in storage device 1830, configured to control processor 1810 as well as a special-purpose processor whereQualcomm Docket No.2403757WO software instructions are incorporated into the actual processor design. Processor 1810 may essentially be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.
[0328] To enable user interaction, computing system 1800 includes an input device 1845, which may represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc. Computing system 1800 may also include output device 1835, which may be one or more of a number of output mechanisms. In some instances, multimodal systems may enable a user to provide multiple types of input / output to communicate with computing system 1800.
[0329] Computing system 1800 may include communications interface 1840, which may generally govern and manage the user input and system output. The communication interface may perform or facilitate receipt and / or transmission wired or wireless communications using wired and / or wireless transceivers, including those making use of an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, an AppleTMLightningTMport / plug, an Ethernet port / plug, a fiber optic port / plug, a proprietary wired port / plug, 3G, 4G, 5G and / or other cellular data network wireless signal transfer, a BluetoothTMwireless signal transfer, a BluetoothTMlow energy (BLE) wireless signal transfer, an IBEACONTMwireless signal transfer, a radio-frequency identification (RFID) wireless signal transfer, near-field communications (NFC) wireless signal transfer, dedicated short range communication (DSRC) wireless signal transfer, 802.11 Wi-Fi wireless signal transfer, wireless local area network (WLAN) signal transfer, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), Infrared (IR) communication wireless signal transfer, Public Switched Telephone Network (PSTN) signal transfer, Integrated Services Digital Network (ISDN) signal transfer, ad-hoc network signal transfer, radio wave signal transfer, microwave signal transfer, infrared signal transfer, visible light signal transfer, ultraviolet light signal transfer, wireless signal transfer along the electromagnetic spectrum, or some combination thereof. The communications interface 1840 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers that are used to determine a location of the computing system 1800 based on receipt of one or more signals from one or more satellites associated with one or more GNSS systems. GNSSQualcomm Docket No.2403757WO systems include, but are not limited to, the US-based Global Positioning System (GPS), the Russia-based Global Navigation Satellite System (GLONASS), the China-based BeiDou Navigation Satellite System (BDS), and the Europe-based Galileo GNSS. There is no restriction on operating on any particular hardware arrangement, and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0330] Storage device 1830 may be a non-volatile and / or non-transitory and / or computer-readable memory device and may be a hard disk or other types of computer readable media which may store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, a floppy disk, a flexible disk, a hard disk, magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid-state memory, a compact disc read only memory (CD-ROM) optical disc, a rewritable compact disc (CD) optical disc, digital video disk (DVD) optical disc, a blu- ray disc (BDD) optical disc, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a Memory Stick® card, a smartcard chip, a EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (e.g., Level 1 (L1) cache, Level 2 (L2) cache, Level 3 (L3) cache, Level 4 (L4) cache, Level 5 (L5) cache, or other (L#) cache), resistive random-access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.
[0331] The storage device 1830 may include software services, servers, services, etc., that when the code that defines such software is executed by the processor 1810, it causes the system to perform a function. In some aspects, a hardware service that performs a particular function may include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1810, connection 1805, output device 1835, etc., to carry out the function. The term “computer- readable medium” includes, but is not limited to, portable or non-portable storage devices,Qualcomm Docket No.2403757WO optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and / or data. A computer-readable medium may include a non- transitory medium in which data may be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, memory or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.
[0332] Specific details are provided in the description above to provide a thorough understanding of the aspects and examples provided herein, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative aspects of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, aspects may be utilized in any number of environments and applications beyond those described herein without departing from the broader scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive. For the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate aspects, the methods may be performed in a different order than that described.
[0333] For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks comprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes,Qualcomm Docket No.2403757WO and other components may be shown as components in block diagram form in order not to obscure the aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects.
[0334] Further, those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0335] Individual aspects may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed, but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.
[0336] Processes and methods according to the above-described examples may be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions may include, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used may be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code. Examples of computer-Qualcomm Docket No.2403757WO readable media that may be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.
[0337] In some aspects the computer-readable storage devices, mediums, and memories may include a cable or wireless signal containing a bitstream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0338] Those of skill in the art will appreciate that information and signals may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof, in some cases depending in part on the particular application, in part on the desired design, in part on the corresponding technology, etc.
[0339] The various illustrative logical blocks, modules, and circuits described in connection with the aspects disclosed herein may be implemented or performed using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor(s) may perform the necessary tasks. Examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also may be embodied in peripherals or add-in cards. Such functionality may also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.
[0340] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.
[0341] The techniques described herein may also be implemented in electronicQualcomm Docket No.2403757WO hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer- readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that may be accessed, read, and / or executed by a computer, such as propagated signals or waves.
[0342] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, an application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.Qualcomm Docket No.2403757WO
[0343] One of ordinary skill will appreciate that the less than (“<”) and greater than (“>”) symbols or terminology used herein may be replaced with less than or equal to (“^”) and greater than or equal to (“^”) symbols, respectively, without departing from the scope of this description.
[0344] Where components are described as being “configured to” perform certain operations, such configuration may be accomplished, for example, by designing electronic circuits or other hardware to perform the operation, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform the operation, or any combination thereof.
[0345] The phrase “coupled to” or “communicatively coupled to” refers to any component that is physically connected to another component either directly or indirectly, and / or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and / or other suitable communication interface) either directly or indirectly.
[0346] Claim language or other language reciting “at least one of” a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and B, C and C, A and A and B, and so on), or any other ordering, duplication, or combination of A, B, and C. The language “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases “at least one” and “one or more” are used interchangeably herein.
[0347] Claim language or other language reciting “at least one processor configured to,” “at least one processor being configured to,” “one or more processors configured to,” “one or more processors being configured to,” or the like indicates that one processor or multiple processors (in any combination) can perform the associated operation(s). For example, claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operations X, Y, and Z; or that multipleQualcomm Docket No.2403757WO processors are each tasked with a certain subset of operations X, Y, and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language reciting “at least one processor configured to: X, Y, and Z” can mean that any single processor may only perform at least a subset of operations X, Y, and Z.
[0348] Where reference is made to one or more elements performing functions (e.g., steps of a method), one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function). Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.
[0349] Where reference is made to an entity (e.g., any entity or device described herein) performing functions or being configured to perform functions (e.g., steps of a method), the entity may be configured to cause one or more elements (individually or collectively) to perform the functions. The one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and / or any combination thereof. Where reference to the entity performing functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions. When the entity is configured to cause more than one component to collectively perform the functions, each function need not be performed by each of those components (e.g., different functions may be performed by different components) and / or each function need not be performed in whole by only one component (e.g., different components may perform different sub-functions of a function).
[0350] Illustrative aspects of the disclosure include:Qualcomm Docket No.2403757WO
[0351] Aspect 1. An apparatus configured to process audio, the apparatus comprising: one or more memories configured to store the audio; and one or more processors coupled to the one or more memories, the one or more processors being configured to: determine spectral envelope features of the audio, wherein the spectral envelope features are determined based on processing the audio using a filter bank including a plurality of filters; determine one or more pitch features of the audio; generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the spectral envelope features, wherein the first encoded representation is based on state feedback between a decoder portion and an encoder portion of the first neural network-based autoencoder; and generate a second encoded representation corresponding to the one or more pitch features.
[0352] Aspect 2. The apparatus of Aspect 1, wherein the first neural network-based autoencoder is a feedback recurrent autoencoder (FRAE).
[0353] Aspect 3. The apparatus of any of Aspects 1 to 2, wherein the first encoded representation comprises a latent associated with the first neural network-based autoencoder and the state feedback.
[0354] Aspect 4. The apparatus of any of Aspects 1 to 3, wherein the spectral envelope features do not include pitch information of the audio.
[0355] Aspect 5. The apparatus of any of Aspects 1 to 4, wherein, to determine the spectral envelope features, the one or more processors are configured to: determine a plurality of cepstral coefficients based on processing the audio using the plurality of filters included in the filter bank, wherein the plurality of filters comprises a first number of filters, and wherein a number of cepstral coefficients included in the plurality of cepstral coefficients is equal to the first number; and truncate the plurality of cepstral coefficients to obtain a set of truncated cepstral coefficients comprising a subset of the plurality of cepstral coefficients.
[0356] Aspect 6. The apparatus of Aspect 5, wherein: the set of truncated cepstral coefficients includes a second number of cepstral coefficients, the second number less than the first number; and the set of truncated cepstral coefficients does not include pitch information of the audio, based on a difference between the first number and the second numberQualcomm Docket No.2403757WO
[0357] Aspect 7. The apparatus of Aspect 6, wherein, to truncate the plurality of cepstral coefficients, the one or more processors are configured to: compute a discrete cosine transform (DCT) of energy per band information associated with the plurality of filters of the filter bank; and skip computation of higher order DCT coefficients of the DCT corresponding to the difference between the first number and the second number.
[0358] Aspect 8. The apparatus of any of Aspects 5 to 7, wherein the one or more processors are configured to remove pitch information of the audio based on truncating the plurality of cepstral coefficients to obtain the set of truncated cepstral coefficients.
[0359] Aspect 9. The apparatus of any of Aspects 1 to 8, wherein the first encoded representation and the second encoded representation do not include absolute phase information or pitch pulse location information associated with the audio.
[0360] Aspect 10. The apparatus of any of Aspects 1 to 9, wherein the one or more pitch features are indicative of a pitch frequency associated with the audio or a pitch lag associated with the audio.
[0361] Aspect 11. The apparatus of any of Aspects 1 to 10, wherein the one or more pitch features are indicative of one or more of pitch correlation information associated with the audio, or a voiced or unvoiced classification associated with the audio.
[0362] Aspect 12. The apparatus of any of Aspects 1 to 11, wherein, to generate the second encoded representation corresponding to the one or more pitch features, the one or more processors are configured to: generate quantized pitch information based on performing vector quantization of the one or more pitch features.
[0363] Aspect 13. The apparatus of any of Aspects 1 to 12, wherein, to generate the second encoded representation corresponding to the one or more pitch features, the one or more processors are configured to: process the one or more pitch features using a second neural network-based autoencoder, wherein the second encoded representation is generated based on state feedback between a respective decoder portion and a respective encoder portion of the second neural network-based autoencoder.
[0364] Aspect 14. The apparatus of any of Aspects 1 to 13, wherein the one or more processors are configured to: transmit the first encoded representation corresponding to the spectral envelope features and the second encoded representation corresponding to the one or more pitch features.Qualcomm Docket No.2403757WO
[0365] Aspect 15. The apparatus of any of Aspects 1 to 14, wherein the apparatus comprises an audio encoder or a voice encoder of a generative voice codec.
[0366] Aspect 16. The apparatus of any of Aspects 1 to 15, wherein: the apparatus comprises a voice encoder of a generative voice codec; and the one or more processors are configured to transmit the first encoded representation and the second encoded representation to a neural synthesizer included in a voice decoder of the generative voice codec.
[0367] Aspect 17. The apparatus of any of Aspects 1 to 16, further comprising one or more microphones configured to obtain the audio.
[0368] Aspect 18. An apparatus configured to process audio, the apparatus comprising: one or more memories configured to store the audio; and one or more processors coupled to the one or more memories, the one or more processors being configured to: receive a first encoded representation corresponding to spectral envelope features of the audio; receive a second encoded representation corresponding to one or more pitch features of the audio; generate reconstructed spectral envelope features corresponding to the audio, based on processing the first encoded representation using a first neural network-based decoder; and generate a synthesized audio output signal based on processing the reconstructed spectral envelope features and the second encoded representation using a neural network-based signal synthesizer.
[0369] Aspect 19. The apparatus of Aspect 18, wherein the neural network-based signal synthesizer is a neural homomorphic vocoder (NHV).
[0370] Aspect 20. The apparatus of Aspect 19, wherein the synthesized audio output signal is a reconstructed signal corresponding to the audio, and wherein the synthesized audio output is generated based on processing the reconstructed spectral envelope features using a neural network filter estimator of the NHV.
[0371] Aspect 21. The apparatus of Aspect 20, wherein the synthesized audio output is generated based on processing the reconstructed spectral envelope features and dequantized pitch features using the neural network filter estimator of the NHV, the dequantized pitch features based on the second encoded representation.
[0372] Aspect 22. The apparatus of any of Aspects 18 to 21, wherein, to generate the synthesized audio output signal, the one or more processors are configured to: computeQualcomm Docket No.2403757WO an inverse discrete cosine transform (DCT) of the reconstructed spectral envelope features to obtain reconstructed filter bank energies corresponding to the audio; and process the reconstructed filter bank energies and the second encoded representation using the neural network-based signal synthesizer to generate the synthesized audio output signal.
[0373] Aspect 23. The apparatus of any of Aspects 18 to 22, wherein, to generate the synthesized audio output signal, the one or more processors are configured to: compute an inverse discrete cosine transform (DCT) of the reconstructed spectral envelope features to obtain reconstructed filter bank energies corresponding to the audio; apply an inverse companding function to the reconstructed filter bank energies to obtain un-companded filter bank energies corresponding to the audio; and process the un-companded filter bank energies and the second encoded representation using the neural network-based signal synthesizer to generate the synthesized audio output signal.
[0374] Aspect 24. The apparatus of Aspect 23, wherein: the inverse DCT is an inverse of a DCT function associated with the first encoded representation; and the inverse companding function is an inverse of a companding function associated with the first encoded representation.
[0375] Aspect 25. The apparatus of any of Aspects 18 to 24, wherein the first encoded representation corresponding to the spectral envelope features does not include pitch information associated with the audio.
[0376] Aspect 26. The apparatus of any of Aspects 18 to 25, wherein the second encoded representation comprises a quantized pitch encoding indicative of one or more of a pitch frequency, a pitch lag, a voiced or unvoiced classification, or pitch correlation information associated with the audio.
[0377] Aspect 27. The apparatus of any of Aspects 18 to 26, wherein: the apparatus comprises an audio decoder or a voice decoder of a generative voice codec; and the first encoded representation and the second encoded representation are received from an audio encoder or a voice encoder of the generative voice codec.
[0378] Aspect 28. The apparatus of any of Aspects 18 to 27, further comprising one or more speakers configured to output the synthesized audio output signal.
[0379] Aspect 29. A method of processing audio, the method comprising: determining spectral envelope features of audio, wherein the spectral envelope features are determinedQualcomm Docket No.2403757WO based on processing the audio using a filter bank including a plurality of filters; determining one or more pitch features of the audio; generating, using a first neural network-based autoencoder, a first encoded representation corresponding to the spectral envelope features, wherein the first encoded representation is based on state feedback between a decoder portion and an encoder portion of the first neural network-based autoencoder; and generating a second encoded representation corresponding to the one or more pitch features.
[0380] Aspect 30. The method of Aspect 29, wherein the first neural network-based autoencoder is a feedback recurrent autoencoder (FRAE).
[0381] Aspect 31. The method of any of Aspects 29 to 30, wherein the first encoded representation comprises a latent associated with the first neural network-based autoencoder and the state feedback.
[0382] Aspect 32. The method of any of Aspects 29 to 31, wherein the spectral envelope features do not include pitch information of the audio.
[0383] Aspect 33. The method of any of Aspects 29 to 32, wherein determining the spectral envelope features comprises: determining a plurality of cepstral coefficients based on processing the audio using the plurality of filters included in the filter bank, wherein the plurality of filters comprises a first number of filters, and wherein a number of cepstral coefficients included in the plurality of cepstral coefficients is equal to the first number; and truncating the plurality of cepstral coefficients to obtain a set of truncated cepstral coefficients comprising a subset of the plurality of cepstral coefficients.
[0384] Aspect 34. The method of Aspect 33, wherein: the set of truncated cepstral coefficients includes a second number of cepstral coefficients, the second number less than the first number; and the set of truncated cepstral coefficients does not include pitch information of the audio, based on a difference between the first number and the second number
[0385] Aspect 35. The method of Aspect 34, wherein truncating the plurality of cepstral coefficients comprises: computing a discrete cosine transform (DCT) of energy per band information associated with the plurality of filters of the filter bank; and skipping computation of higher order DCT coefficients of the DCT corresponding to the difference between the first number and the second number.Qualcomm Docket No.2403757WO
[0386] Aspect 36. The method of any of Aspects 33 to 35, further comprising removing pitch information of the audio based on truncating the plurality of cepstral coefficients to obtain the set of truncated cepstral coefficients.
[0387] Aspect 37. The method of any of Aspects 29 to 36, wherein the first encoded representation and the second encoded representation do not include absolute phase information or pitch pulse location information associated with the audio.
[0388] Aspect 38. The method of any of Aspects 29 to 37, wherein the one or more pitch features are indicative of a pitch frequency associated with the audio or a pitch lag associated with the audio.
[0389] Aspect 39. The method of any of Aspects 29 to 38, wherein the one or more pitch features are indicative of one or more of pitch correlation information associated with the audio, or a voiced or unvoiced classification associated with the audio.
[0390] Aspect 40. The method of any of Aspects 29 to 39, wherein generating the second encoded representation corresponding to the one or more pitch features comprises generating quantized pitch information based on performing vector quantization of the one or more pitch features.
[0391] Aspect 41. The method of any of Aspects 29 to 40, wherein generating the second encoded representation corresponding to the one or more pitch features comprises: processing the one or more pitch features using a second neural network-based autoencoder, wherein the second encoded representation is generated based on state feedback between a respective decoder portion and a respective encoder portion of the second neural network-based autoencoder.
[0392] Aspect 42. The method of any of Aspects 29 to 41, wherein the method further includes transmitting the first encoded representation corresponding to the spectral envelope features and the second encoded representation corresponding to the one or more pitch features.
[0393] Aspect 43. A method of processing audio, the method comprising: receiving a first encoded representation corresponding to spectral envelope features of audio; receiving a second encoded representation corresponding to one or more pitch features of the audio; generating reconstructed spectral envelope features corresponding to the audio, based on processing the first encoded representation using a first neural network-basedQualcomm Docket No.2403757WO decoder; and generating a synthesized audio output signal based on processing the reconstructed spectral envelope features and the second encoded representation using a neural network-based signal synthesizer.
[0394] Aspect 44. The method of Aspect 43, wherein the neural network-based signal synthesizer is a neural homomorphic vocoder (NHV).
[0395] Aspect 45. The method of Aspect 44, wherein the synthesized audio output signal is a reconstructed signal corresponding to the audio, and wherein the synthesized audio output is generated based on processing the reconstructed spectral envelope features using a neural network filter estimator of the NHV.
[0396] Aspect 46. The method of Aspect 45, wherein the synthesized audio output is generated based on processing the reconstructed spectral envelope features and dequantized pitch features using the neural network filter estimator of the NHV, the dequantized pitch features based on the second encoded representation.
[0397] Aspect 47. The method of any of Aspects 43 to 46, wherein generating the synthesized audio output signal comprises: computing an inverse discrete cosine transform (DCT) of the reconstructed spectral envelope features to obtain reconstructed filter bank energies corresponding to the audio; and processing the reconstructed filter bank energies and the second encoded representation using the neural network-based signal synthesizer to generate the synthesized audio output signal.
[0398] Aspect 48. The method of any of Aspects 43 to 47, wherein generating the synthesized audio output signal comprises: computing an inverse discrete cosine transform (DCT) of the reconstructed spectral envelope features to obtain reconstructed filter bank energies corresponding to the audio; applying an inverse companding function to the reconstructed filter bank energies to obtain un-companded filter bank energies corresponding to the audio; and processing the un-companded filter bank energies and the second encoded representation using the neural network-based signal synthesizer to generate the synthesized audio output signal.
[0399] Aspect 49. The method of Aspect 48, wherein: the inverse DCT is an inverse of a DCT function associated with the first encoded representation; and the inverse companding function is an inverse of a companding function associated with the first encoded representation.Qualcomm Docket No.2403757WO
[0400] Aspect 50. The method of any of Aspects 43 to 49, wherein the first encoded representation corresponding to the spectral envelope features does not include pitch information associated with the audio.
[0401] Aspect 51. The method of any of Aspects 43 to 50, wherein the second encoded representation comprises a quantized pitch encoding indicative of one or more of a pitch frequency, a pitch lag, a voiced or unvoiced classification, or pitch correlation information associated with the audio.
[0402] Aspect 52. A method for processing one or more audio samples, comprising performing operations according to any of Aspects 1 to 17 or 29 to 42.
[0403] Aspect 53. A non-transitory computer-readable storage medium comprising instructions stored thereon which, when executed by at least one processor, causes the at least one processor to perform operations according to any of Aspects 1 to 17 or 29 to 42.
[0404] Aspect 54. An apparatus for processing one or more audio samples, the apparatus comprising one or more means for performing operations according to any of Aspects 1 to 17 or 29 to 42.
[0405] Aspect 55. A method for processing one or more audio samples, comprising performing operations according to any of Aspects 18 to 28 or 43 to 51.
[0406] Aspect 56. A non-transitory computer-readable storage medium comprising instructions stored thereon which, when executed by at least one processor, causes the at least one processor to perform operations according to any of Aspects 18 to 28 or 43 to 51.
[0407] Aspect 57. An apparatus for processing one or more audio samples, the apparatus comprising one or more means for performing operations according to any of Aspects 18 to 28 or 43 to 51.
Claims
Qualcomm Docket No.2403757WO CLAIMS What is claimed is:
1. An apparatus configured to process audio, the apparatus comprising: one or more memories configured to store the audio; and one or more processors coupled to the one or more memories, the one or more processors being configured to: determine spectral envelope features of the audio, wherein the spectral envelope features are determined based on processing the audio using a filter bank including a plurality of filters; determine one or more pitch features of the audio; generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the spectral envelope features, wherein the first encoded representation is based on state feedback between a decoder portion and an encoder portion of the first neural network-based autoencoder; and generate a second encoded representation corresponding to the one or more pitch features.
2. The apparatus of claim 1, wherein the first neural network-based autoencoder is a feedback recurrent autoencoder (FRAE).
3. The apparatus of claim 1, wherein the first encoded representation comprises a latent associated with the first neural network-based autoencoder and the state feedback.
4. The apparatus of claim 1, wherein the spectral envelope features do not include pitch information of the audio.
5. The apparatus of claim 1, wherein, to determine the spectral envelope features, the one or more processors are configured to: determine a plurality of cepstral coefficients based on processing the audio using the plurality of filters included in the filter bank, wherein the plurality of filters comprises a first number of filters, and wherein a number of cepstral coefficients included in the plurality of cepstral coefficients is equal to the first number; andQualcomm Docket No.2403757WO truncate the plurality of cepstral coefficients to obtain a set of truncated cepstral coefficients comprising a subset of the plurality of cepstral coefficients.
6. The apparatus of claim 5, wherein: the set of truncated cepstral coefficients includes a second number of cepstral coefficients, the second number less than the first number; and the set of truncated cepstral coefficients does not include pitch information of the audio, based on a difference between the first number and the second number.
7. The apparatus of claim 6, wherein, to truncate the plurality of cepstral coefficients, the one or more processors are configured to: compute a discrete cosine transform (DCT) of energy per band information associated with the plurality of filters of the filter bank; and skip computation of higher order DCT coefficients of the DCT corresponding to the difference between the first number and the second number.
8. The apparatus of claim 5, wherein the one or more processors are configured to remove pitch information of the audio based on truncating the plurality of cepstral coefficients to obtain the set of truncated cepstral coefficients.
9. The apparatus of claim 1, wherein the first encoded representation and the second encoded representation do not include absolute phase information or pitch pulse location information associated with the audio.
10. The apparatus of claim 1, wherein the one or more pitch features are indicative of a pitch frequency associated with the audio or a pitch lag associated with the audio.
11. The apparatus of claim 1, wherein the one or more pitch features are indicative of one or more of pitch correlation information associated with the audio, or a voiced or unvoiced classification associated with the audio.Qualcomm Docket No.2403757WO 12. The apparatus of claim 1, wherein, to generate the second encoded representation corresponding to the one or more pitch features, the one or more processors are configured to: generate quantized pitch information based on performing vector quantization of the one or more pitch features.
13. The apparatus of claim 1, wherein, to generate the second encoded representation corresponding to the one or more pitch features, the one or more processors are configured to: process the one or more pitch features using a second neural network-based autoencoder, wherein the second encoded representation is generated based on state feedback between a respective decoder portion and a respective encoder portion of the second neural network-based autoencoder.
14. The apparatus of claim 1, wherein the one or more processors are configured to: transmit the first encoded representation corresponding to the spectral envelope features and the second encoded representation corresponding to the one or more pitch features.
15. The apparatus of claim 1, wherein the spectral envelope features include one or more linear prediction coefficients (LPCs) corresponding to the audio.
16. The apparatus of claim 1, wherein: the apparatus comprises a voice encoder of a generative voice codec; and the one or more processors are configured to transmit the first encoded representation and the second encoded representation to a neural synthesizer included in a voice decoder of the generative voice codec.
17. The apparatus of claim 1, further comprising one or more microphones configured to obtain the audio.
18. An apparatus configured to process audio, the apparatus comprising: one or more memories configured to store the audio; andQualcomm Docket No.2403757WO one or more processors coupled to the one or more memories, the one or more processors being configured to: receive a first encoded representation corresponding to spectral envelope features of the audio; receive a second encoded representation corresponding to one or more pitch features of the audio; generate reconstructed spectral envelope features corresponding to the audio, based on processing the first encoded representation using a first neural network-based decoder; and generate a synthesized audio output signal based on processing the reconstructed spectral envelope features and the second encoded representation using a neural network-based signal synthesizer.
19. The apparatus of claim 18, wherein the neural network-based signal synthesizer is a neural homomorphic vocoder (NHV).
20. The apparatus of claim 19, wherein the synthesized audio output signal is a reconstructed signal corresponding to the audio, and wherein the synthesized audio output is generated based on processing the reconstructed spectral envelope features using a neural network filter estimator of the NHV.
21. The apparatus of claim 20, wherein the synthesized audio output is generated based on processing the reconstructed spectral envelope features and dequantized pitch features using the neural network filter estimator of the NHV, the dequantized pitch features based on the second encoded representation.
22. The apparatus of claim 18, wherein, to generate the synthesized audio output signal, the one or more processors are configured to: compute an inverse discrete cosine transform (DCT) of the reconstructed spectral envelope features to obtain reconstructed filter bank energies corresponding to the audio; andQualcomm Docket No.2403757WO process the reconstructed filter bank energies and the second encoded representation using the neural network-based signal synthesizer to generate the synthesized audio output signal.
23. The apparatus of claim 18, wherein, to generate the synthesized audio output signal, the one or more processors are configured to: compute an inverse discrete cosine transform (DCT) of the reconstructed spectral envelope features to obtain reconstructed filter bank energies corresponding to the audio; apply an inverse companding function to the reconstructed filter bank energies to obtain un-companded filter bank energies corresponding to the audio; and process the un-companded filter bank energies and the second encoded representation using the neural network-based signal synthesizer to generate the synthesized audio output signal.
24. The apparatus of claim 23, wherein: the inverse DCT is an inverse of a DCT function associated with the first encoded representation; and the inverse companding function is an inverse of a companding function associated with the first encoded representation.
25. The apparatus of claim 18, wherein the first encoded representation corresponding to the spectral envelope features does not include pitch information associated with the audio.
26. The apparatus of claim 18, wherein the second encoded representation comprises a quantized pitch encoding indicative of one or more of a pitch frequency, a pitch lag, a voiced or unvoiced classification, or pitch correlation information associated with the audio.
27. The apparatus of claim 18, wherein: the apparatus comprises an audio decoder or a voice decoder of a generative voice codec; andQualcomm Docket No.2403757WO the first encoded representation and the second encoded representation are received from an audio encoder or a voice encoder of the generative voice codec.
28. The apparatus of claim 18, further comprising one or more speakers configured to output the synthesized audio output signal.
29. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: determine spectral envelope features of audio, wherein the spectral envelope features are determined based on processing the audio using a filter bank including a plurality of filters; determine one or more pitch features of the audio; generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the spectral envelope features, wherein the first encoded representation is based on state feedback between a decoder portion and an encoder portion of the first neural network-based autoencoder; generate a second encoded representation corresponding to the one or more pitch features; and transmit the first encoded representation corresponding to the spectral envelope features and the second encoded representation corresponding to the one or more pitch features.
30. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: receive a first encoded representation corresponding to spectral envelope features of audio; receive a second encoded representation corresponding to one or more pitch features of the audio; generate reconstructed spectral envelope features corresponding to the audio, based on processing the first encoded representation using a first neural network-based decoder; andQualcomm Docket No.2403757WO generate a synthesized audio output signal based on processing the reconstructed spectral envelope features and the second encoded representation using a neural network-based signal synthesizer.
Citation Information
Patent Citations
Method and apparatus for recurrent auto-encoding
US11526734B2
Method and apparatus for recurrent auto-encoding
US20210089863A1
Audio coding using machine learning based linear filters and non-linear neural sources
WO2023064735A1