Generative audio codec for signal synthesis based on groupwise joint coding of spectral envelope features and pitch information
The generative audio codec system addresses inefficiencies in audio coding by employing feedback recurrent autoencoders and neural synthesizers for joint coding of spectral and pitch features, enhancing bit efficiency and reducing computational demands for high-quality audio reconstruction.
Patent Information
- Application Number
- PCT/US2025/028518
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2025-05-08
- Publication Date
- 2025-11-20
AI Technical Summary
Existing audio coding technologies face challenges in efficiently compressing audio data while maintaining quality, particularly with the increased computational complexity of neural network-based decoders and the need for improved bit efficiency in encoding and decoding processes.
A generative audio codec system utilizing feedback recurrent autoencoders and neural network-based synthesizers performs groupwise joint coding of spectral envelope features, pitch information, and energy information, enabling efficient encoding and decoding of audio data through separate processing and combined vector generation.
The system achieves improved bit efficiency and reduced computational complexity by separating and encoding audio features, allowing for high-quality audio reconstruction with reduced processing time, suitable for resource-constrained devices.
Smart Images

Figure US2025028518_20112025_PF_FP_ABST
Abstract
Description
Qualcomm Docket No.2404227WO GENERATIVE AUDIO CODEC FOR SIGNAL SYNTHESIS BASED ON GROUPWISE JOINT CODING OF SPECTRAL ENVELOPE FEATURES AND PITCH INFORMATION FIELD
[0001] Aspects of the present disclosure generally relate to audio coding (e.g., audio encoding and / or decoding). In some implementations, examples are described for performing audio coding using a generative audio codec system including one or more feedback recurrent autoencoders and a neural network-based synthesizer. BACKGROUND
[0002] Audio coding (also referred to as voice coding and / or speech coding) is a technique used to represent a digitized audio signal using as few bits as possible (thus compressing the speech data), while attempting to maintain a certain level of audio quality. An audio or voice encoder is used to encode (or compress) the digitized audio (e.g., speech, music, etc.) signal to a lower bit-rate stream of data. The lower bit-rate stream of data can be input to an audio or voice decoder, which decodes the stream of data and constructs an approximation or reconstruction of the original signal. The audio or voice encoder-decoder structure can be referred to as an audio coder (or voice coder or speech coder) or an audio / voice / speech coder-decoder (codec).
[0003] Audio coders exploit the fact that speech signals are highly correlated waveforms. Some speech coding techniques are based on a source-filter model of speech production, which assumes that the vocal cords are the source of spectrally flat sound (an excitation signal), and that the vocal tract acts as a filter to spectrally shape the various sounds of speech. The different phonemes (e.g., vowels, fricatives, and voice fricatives) can be distinguished by their excitation (source) and spectral shape (filter). SUMMARY
[0004] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary has the sole purpose to present certain concepts relating to one or more aspectsQualcomm Docket No.2404227WO relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.
[0005] Disclosed are systems, methods, apparatuses, and computer-readable media for audio coding (e.g., encoding and / or decoding audio data). According to at least one illustrative example, an apparatus for processing audio data is provided. The apparatus includes at least one memory and at least one processor coupled to the at least one memory and configured to: generate a first sub-vector of audio features, wherein the first sub- vector corresponds to a first combination of one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; generate a second sub-vector of audio features, wherein the second sub-vector corresponds to a second combination of one or more of the spectral envelope features of the audio, the energy features of the audio, or the pitch features of the audio; generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the first sub- vector of audio features; and generate, using a second neural network-based autoencoder, a second encoded representation corresponding to the second sub-vector of audio features.
[0006] In another example, a method for processing audio data is provided. The method includes: generating a first sub-vector of audio features, wherein the first sub-vector corresponds to a first combination of one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; generating a second sub-vector of audio features, wherein the second sub-vector corresponds to a second combination of one or more of the spectral envelope features of the audio, the energy features of the audio, or the pitch features of the audio; generating, using a first neural network-based autoencoder, a first encoded representation corresponding to the first sub- vector of audio features; and generating, using a second neural network-based autoencoder, a second encoded representation corresponding to the second sub-vector of audio features.
[0007] In another example, a non-transitory computer-readable medium is provided that includes instructions that, when executed by at least one processor, cause the at least one processor to: generate a first sub-vector of audio features, wherein the first sub-vector corresponds to a first combination of one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; generate a second sub-Qualcomm Docket No.2404227WO vector of audio features, wherein the second sub-vector corresponds to a second combination of one or more of the spectral envelope features of the audio, the energy features of the audio, or the pitch features of the audio; generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the first sub- vector of audio features; and generate, using a second neural network-based autoencoder, a second encoded representation corresponding to the second sub-vector of audio features.
[0008] In another example, an apparatus for processing audio data is provided. The apparatus includes: means for generating a first sub-vector of audio features, wherein the first sub-vector corresponds to a first combination of one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; means for generating a second sub-vector of audio features, wherein the second sub-vector corresponds to a second combination of one or more of the spectral envelope features of the audio, the energy features of the audio, or the pitch features of the audio; means for generating, using a first neural network-based autoencoder, a first encoded representation corresponding to the first sub-vector of audio features; and means for generating, using a second neural network-based autoencoder, a second encoded representation corresponding to the second sub-vector of audio features.
[0009] In another illustrative example, an apparatus for processing audio data (e.g., decoding audio data) is provided. The apparatus includes at least one memory and at least one processor coupled to the at least one memory and configured to: receive a first encoded representation corresponding to a first combination of features selected from a set of features comprising one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; receive a second encoded representation corresponding to a second combination of features selected from the set of features; generate a first reconstructed sub-vector of features based on the first encoded representation, wherein the first reconstructed sub-vector includes the first combination of features; generate a second reconstructed sub-vector of features based on the second encoded representation, wherein the second reconstructed sub-vector includes the second combination of features; and generate a synthesized audio output signal based on processing the first reconstructed sub-vector of features and the second reconstructed sub- vector of features using a neural network-based signal synthesizer.Qualcomm Docket No.2404227WO
[0010] In another example, a method for processing audio data is provided. The method includes: receiving a first encoded representation corresponding to a first combination of features selected from a set of features comprising one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; receiving a second encoded representation corresponding to a second combination of features selected from the set of features; generating a first reconstructed sub-vector of features based on the first encoded representation, wherein the first reconstructed sub-vector includes the first combination of features; generating a second reconstructed sub-vector of features based on the second encoded representation, wherein the second reconstructed sub-vector includes the second combination of features; and generating a synthesized audio output signal based on processing the first reconstructed sub-vector of features and the second reconstructed sub-vector of features using a neural network-based signal synthesizer.
[0011] In another example, a non-transitory computer-readable medium is provided that includes instructions that, when executed by at least one processor, cause the at least one processor to: receive a first encoded representation corresponding to a first combination of features selected from a set of features comprising one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; receive a second encoded representation corresponding to a second combination of features selected from the set of features; generate a first reconstructed sub-vector of features based on the first encoded representation, wherein the first reconstructed sub- vector includes the first combination of features; generate a second reconstructed sub- vector of features based on the second encoded representation, wherein the second reconstructed sub-vector includes the second combination of features; and generate a synthesized audio output signal based on processing the first reconstructed sub-vector of features and the second reconstructed sub-vector of features using a neural network-based signal synthesizer.
[0012] In another example, an apparatus for processing audio data is provided. The apparatus includes: means for receiving a first encoded representation corresponding to a first combination of features selected from a set of features comprising one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; means for receiving a second encoded representation corresponding to a second combination of features selected from the set of features; means for generating a firstQualcomm Docket No.2404227WO reconstructed sub-vector of features based on the first encoded representation, wherein the first reconstructed sub-vector includes the first combination of features; means for generating a second reconstructed sub-vector of features based on the second encoded representation, wherein the second reconstructed sub-vector includes the second combination of features; and means for generating a synthesized audio output signal based on processing the first reconstructed sub-vector of features and the second reconstructed sub-vector of features using a neural network-based signal synthesizer.
[0013] Aspects generally include a method, apparatus, system, computer program product, non-transitory computer-readable medium, user equipment, base station, wireless communication device, and / or processing system as substantially described herein with reference to and as illustrated by the drawings and specification.
[0014] The foregoing has outlined rather broadly the features and technical advantages of examples according to the disclosure in order that the detailed description that follows may be better understood. Additional features and advantages will be described hereinafter. The conception and specific examples disclosed may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. Characteristics of the concepts disclosed herein, both their organization and method of operation, together with associated advantages, will be better understood from the following description when considered in connection with the accompanying figures. Each of the figures is provided for the purposes of illustration and description, and not as a definition of the limits of the claims.
[0015] While aspects are described in the present disclosure by illustration to some examples, those skilled in the art will understand that such aspects may be implemented in many different arrangements and scenarios. Techniques described herein may be implemented using different platform types, devices, systems, shapes, sizes, and / or packaging arrangements. For example, some aspects may be implemented via integrated chip implementations or other non-module-component based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, and / or artificial intelligence devices). Aspects may be implemented in chip-level components, modular components, non-modular components, non-chip-level components, device-level components, and / or system-levelQualcomm Docket No.2404227WO components. Devices incorporating described aspects and features may include additional components and features for implementation and practice of claimed and described aspects. For example, transmission and reception of wireless signals may include one or more components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders, and / or summers). It is intended that aspects described herein may be practiced in a wide variety of devices, components, systems, distributed arrangements, and / or end-user devices of varying size, shape, and constitution.
[0016] Other objects and advantages associated with the aspects disclosed herein will be apparent to those skilled in the art based on the accompanying drawings and detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.
[0017] The foregoing, together with other features and aspects, will become more apparent upon referring to the following specification, claims, and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are presented to aid in the description of various aspects of the disclosure and are provided solely for illustration of the aspects and not limitation thereof.
[0019] FIG.1 is a block diagram illustrating an example speech processing system, in accordance with some examples;
[0020] FIG. 2A is a block diagram illustrating an example feature generator, in accordance with some examples;
[0021] FIG.2B is a block diagram illustrating an example of a voice coding system, in accordance with some examples;
[0022] FIG. 2C is a block diagram illustrating an example of a code-excited linear prediction (CELP)-based voice coding system utilizing a fixed codebook (FCB), in accordance with some examples;Qualcomm Docket No.2404227WO
[0023] FIG. 3 is a block diagram illustrating an example of a voice coding signal synthesis system utilizing a linear time-varying filter generated using a neural network model and a separate linear predictive coding (LPC) filter, in accordance with some examples;
[0024] FIG.4 is a block diagram illustrating an example of an audio codec system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer, in accordance with some examples;
[0025] FIG. 5A is a block diagram illustrating an example of an audio codec system that can perform spectral envelope feature extraction from an audio signal based on truncation of discrete cosine transform (DCT) coefficients, in accordance with some examples;
[0026] FIG. 5B is a block diagram illustrating an example of an audio codec system that can perform spectral envelope feature extraction using a filter bank and a determination of energies per frequency band, in accordance with some examples;
[0027] FIG.6 is a block diagram illustrating an example of a decoder of a voice coding signal synthesis system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer comprising a neural homomorphic vocoder (NHV), in accordance with some examples;
[0028] FIG.7 is a block diagram illustrating an example of a decoder of a voice coding signal synthesis system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer comprising a linear predictive coding (LPC) network, in accordance with some examples;
[0029] FIG.8 is a block diagram illustrating an example of an audio codec system that encodes energy information using a first feedback recurrent autoencoder (FRAE) and encodes spectral envelope features using a second FRAE, in accordance with some examples;
[0030] FIG.9 is a block diagram illustrating an example of an audio codec system that can be used to perform neural speech synthesis based on groupwise joint coding to generate encoded vectors of features corresponding to different combinations of spectral energy features, energy features, and pitch information, in accordance with some examples;Qualcomm Docket No.2404227WO
[0031] FIG.10 is a block diagram illustrating an example of an audio codec system that can be used to perform neural speech synthesis based on using a first denoiser associated with spectral envelope feature extraction and a second denoiser associated with pitch extraction, in accordance with some examples;
[0032] FIG.11 is a block diagram illustrating an example of an audio codec system that can be used to perform neural speech synthesis based on using a front-end non-speech detector (FNSD) to switch between a neural synthesizer audio codec path for encoding and / or decoding speech signals and an additional audio codec path for encoding and / or decoding non-speech signals, in accordance with some examples;
[0033] FIG. 12 is a flow chart illustrating an example of a process for processing one or more audio samples, in accordance with some examples;
[0034] FIG. 13 is a flow chart illustrating an example of a process for processing one or more audio samples, in accordance with some examples; and
[0035] FIG. 14 is a block diagram illustrating an example of a computing system for implementing certain aspects described herein. DETAILED DESCRIPTION
[0036] Certain aspects of this disclosure are provided below for illustration purposes. Alternate aspects may be devised without departing from the scope of the disclosure. Additionally, well-known elements of the disclosure will not be described in detail or will be omitted so as not to obscure the relevant details of the disclosure. Some of the aspects described herein may be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and description are not intended to be restrictive.
[0037] The ensuing description provides example aspects only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the example aspects will provide those skilled in the art with an enabling description for implementing an example aspect. It should be understood that various changes may be made in the function and arrangement of elements without departingQualcomm Docket No.2404227WO from the scope of the application as set forth in the appended claims.
[0038] Audio coding (e.g., speech coding, music signal coding, or other type of audio coding) can be performed on a digitized audio signal (e.g., a speech signal) to compress the amount of data for storage, transmission, and / or various other uses. A voice coding system can be an audio coding system that is used to perform audio coding on a speech signal that represents the speech (e.g., voiced words or sounds) uttered by a user of the system. A voice coding system can also be referred to as a voice or speech coder or a voice coder-decoder (codec). A voice coding system may include one or more encoders (e.g., voice encoders) and one or more decoders (e.g., voice decoders). For example, a voice encoder can be used to process an input speech signal. Input speech signals may include a digitized speech signal generated from an analog speech signal from a given source, where the resulting digitized speech signal is a discrete-time speech signal with sample values (e.g., also referred to as samples or audio samples) that are also discretized.
[0039] Voice coders (e.g., voice encoders and / or voice decoders) can be designed and implemented to exploit the fact that speech signals are highly correlated waveforms. The samples of an input speech signal can be divided into blocks of N samples each, where a block of N samples is referred to as a frame. In some examples, each frame can be 10-20 milliseconds (ms) in length. A voice encoder can generate a compressed signal corresponding to the input, digitized speech signal. The compressed signal generated by the voice encoder can include a lower bitrate stream of audio data that represents the speech signal using as few bits as possible, while attempting (e.g., by the voice encoder) to maintain a certain quality level for the speech. The compressed signal generated by the voice encoder can have a lower bitrate than the uncompressed digitized speech signal provided as input to the voice encoder. In some cases, voice encoders may be used to reduce the bitrate of an input speech signal. The bitrate of a signal is based on the sampling frequency (e.g., the sampling frequency of a microphone used to sample a user’s speech) and the number of bits per sample. For example, the bitrate of a signal (BR) can be equal to S*b, where S represents the sampling frequency and b represents the number of bits per sample.
[0040] The output of a voice encoder may be a compressed speech signal. The compressed speech signal can be stored and / or transmitted to and processed by a corresponding voice decoder. The corresponding voice decoder can be a voice decoderQualcomm Docket No.2404227WO that is associated with the voice encoder and / or a voice decoder that is configured to decode the compressed speech signals generated by the voice encoder. In some cases, a voice decoder may communicate with a voice encoder, for example to request speech data, to transmit feedback information, and / or to provide other communications to or from the voice encoder and voice decoder. The voice decoder can be configured to decode the data of the compressed speech signal and construct a reconstructed speech signal that approximates the original speech signal. For example, the reconstructed speech signal may include a digitized, discrete-time signal with a same or similar bitrate as the bitrate of the original speech signal. In some cases, the voice decoder may implement an inverse of a voice coding algorithm used by the voice encoder.
[0041] In some examples, one or more machine learning systems, networks, models, etc., can be used to perform audio coding (e.g., encoding and / or decoding of one or more audio samples). For example, machine learning systems using one or more neural network models can be used to implement a voice coder to encode and / or decode voice data (e.g., speech signals, etc.). In some cases, machine learning systems (e.g., using one or more neural network models) can be used to generate reconstructed voice or audio signals using a process of neural synthesis. For example, using features extracted from one or more frames of audio data, a neural network-based voice decoder can generate coefficients for one or more linear filters, learned filters, etc., that may be used to perform voice decoding. In some cases, a neural network-based voice decoder can be trained to estimate the parameters (e.g., linear filter coefficients, etc.) of a speech synthesis pipeline, where learned filters that are generated or tuned using a neural network-based voice decoder may subsequently be used to generate a reconstructed speech signal (e.g., a synthesized speech signal, a speech synthesis signal, etc.).
[0042] Neural network-based voice or speech decoders may also be referred to as neural decoders. Neural decoders can be more computationally complex to implement than non- neural or non-machine learning-based voice or speech decoders. For example, a neural decoder may be associated with a longer frame processing time (e.g., a longer frame decoding time) than a non-neural decoder, when the neural decoder and the non-neural decoder are implemented or executed on the same computational device, platform, hardware, etc. In some cases, the use of neural decoders with resource-constrained devices such as UEs, smartphones, mobile computing devices, wearables, etc., may also be associated with increased frame processing times by the neural decoder, includingQualcomm Docket No.2404227WO increased frame processing times beyond the real-time processing requirement associated with the frame length (e.g., 20ms speech frames, etc.).
[0043] Systems, apparatuses, methods (also referred to as processes), and computer- readable media (collectively referred to herein as “systems and techniques”) are described herein that can be used to perform neural network-based and / or machine learning-based encoding and decoding of audio data. For example, the systems and techniques can be used to provide a bit-efficient generative audio coding system (e.g., audio codec) that utilizes an encoder with one or more feedback recurrent autoencoders (FRAEs) to generate encoded representations of audio based on pitch information, spectral envelope information, and / or energy information associated with the audio data being encoded. In some examples, the pitch information, spectral envelope information, and energy information may each be generated separately at the encoder, based on corresponding processing of one or more portions of the input audio. The separately generated or determined pitch information, spectral envelope features, and energy features may be combined or concatenated for joint coding by the encoder. In some aspects, the encoder can perform group-wise joint coding of the pitch information, the spectral envelope features, and the energy features.
[0044] For example, the groupwise joint coding can correspond to generating a single, combined vector of features corresponding to the pitch information, the spectral envelope features, and the energy features. The single combined vector of features can be based on concatenation of the pitch information, the spectral envelope features, and the energy features. The single combined vector of features can be split and / or used to generate various sub-vectors, subsets, combinations, etc., of the different types of features generated or obtained by the encoder. For example, the single combined vector of features can be split into different sub-vectors and combinations of the pitch information (e.g., pitch features), spectral envelope features, and energy features.
[0045] Each sub-vector can include, represent, and / or indicate a different combination or subset of audio features determined at the encoder. Each sub-vector can be separately processed and encoded by the encoder. For example, a corresponding latent or other encoded representation can be generated by a respective FRAE included in the encoder for each sub-vector generated from the set of input audio features. The separate FRAE utilized to generate an encoded representation of a particular sub-vector of features canQualcomm Docket No.2404227WO be used to generate a corresponding encoded representation of the particular sub-vector, with or without multiple description coding (MDC) implemented for the encoded representation. In some cases, one or more sub-vectors of combined audio features can be coded using an FRAE of the encoder, and one or more sub-vectors of combined audio features can be coded using a non-machine learning coding technique implemented by the encoder. Sub-vectors of combined audio features that are coded with non-machine learning coding techniques can be coded with or without MDC and / or with or without forward error correction (FEC).
[0046] Each sub-vector of combined audio features that is coded using an FRAE of the encoder to generate the encoded representation can be provided to a corresponding decoder (e.g., a corresponding FRAE decoder) to generate a reconstructed representation of the encoded sub-vector of combined audio features. The decoded or reconstructed representation of each FRAE-encoded sub-vector of combined audio features can be provided as input to a neural network-based synthesizer configured to generate a reconstructed speech signal. Sub-vectors of combined audio features that are coded using non-machine learning coding techniques at the encoder can be dequantized and provided as inputs to the neural network-based synthesizer, and / or can be provided directly to the neural network-based synthesizer without being dequantized.
[0047] The systems and techniques can utilize a decoder that receives the encoded representations of the audio from the encoder, and generates a reconstructed (e.g., recovered) audio signal based on the encoded representations received from the encoder. In some examples, the decoder can include one or more FRAE decoders and a neural synthesizer configured to perform signal synthesis and generate the reconstructed audio signal from the encoded representations received from the encoder.
[0048] In some examples, the systems and techniques can be used to implement a bit- efficient generative voice codec configured to process audio data comprising voice or speech signals. The encoder of the bit-efficient generative voice codec can be used to extract pitch and spectral envelope features from an input audio (e.g., one or more audio frames provided as input to the encoder). The encoder can split the input audio signal into respective spectral features (e.g., spectral envelope features associated with the input audio signal) and respective pitch features or pitch information, where the spectral features and the pitch information are non-overlapping information. For example,Qualcomm Docket No.2404227WO information represented in and / or indicated by the spectral features is not represented in or indicated by the pitch information, and information represented in and / or indicated by the pitch information is not represented in or indicated by the spectral features. In some examples, the spectral features can be spectral envelope features associated with the overall shape of the frequency spectrum of the audio signal (e.g., where the spectral envelope connects the peaks of the individual frequency components in the frequency spectrum of an audio signal).
[0049] In some aspects, the encoder of the generative voice codec system can be configured to extract spectral envelope features based on determining a cepstrum and / or cepstral coefficients corresponding to the input audio received by the encoder. The bit efficiency of the encoder can be improved based on generating and / or extracting the spectral envelope features without including pitch information. Removing pitch information from the extracted spectral envelope features can avoid the double-coding or overlapping of information between the spectral features and the pitch information included in the encoded audio representations transmitted between the encoder and the decoder of the generative voice codec.
[0050] In some examples, pitch information can be removed from the spectral envelope features based on performing discrete cosine transform (DCT) truncation to remove or skip the computation of a subset of higher order DCT coefficients. For example, to generate the spectral envelope features, the encoder can process the audio signal using a filterbank of size N (e.g., N frequency bands), apply a DCT of size N, and perform truncation to the first M DCT coefficients (e.g., cepstral coefficients) where M is less than N. Based on performing the DCT truncation to remove pitch information from the spectral features, the extracted spectral envelope features can be generated with a relatively high resolution (e.g., corresponding to the relatively large filterbank size N), and truncated to the smaller size M without a reduction in the resolution. Based on utilizing a relatively large filterbank with a relatively large number of frequency bands, and subsequently truncating the DCT of the relatively large filterbank output, the systems and techniques can be used to code only spectral envelope information in the cepstral features, without including or double coding pitch information in the cepstral features.
[0051] The cepstral features used to code the spectral envelope information (e.g., the extracted spectral envelope features associated with the filterbank and DCT truncation)Qualcomm Docket No.2404227WO can be compressed (e.g., encoded) using one or more feedback recurrent autoencoders (FRAE) included in the encoder of the generative voice codec system. The systems and techniques can utilize one or more FRAEs in the encoder-side of the generative voice codec, and can include one or more FRAE decoders (e.g., the FRAE autoencoder architecture without the encoder portion thereof) in the decoder-side of the generative voice codec. The FRAE decoders can be used to decode the encoded representation of the combined feature vector of spectral envelope features, pitch features, and / or energy features.
[0052] Autoencoders have gained popularity in recent years as they are able to learn efficient representations of input data without the need of labels (e.g., based on performing unsupervised learning). Various types of autoencoders exist and are well explained in “Autoencoder and its various variants” 2018 IEEE International Conference on Systems, Man, and Cybernetics, Zhang et. al. In speech coding, autoencoders conditioned on log Mel Cepstrum inputs and / or spectrogram inputs have been used for speech compression in "Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders" arXiv:2102.06610v1 [eess.AS] 12 Feb 2021, Casebeer et. al.
[0053] Speech coding is a lossy compression process. Autoencoders perform dimensionality reduction (e.g., an N-dimensional vector input vector is passed into the encoder of the autoencoder, and lossy compression is performed to represent the important aspects of the input vector in an M-dimensional vector, where M is smaller than N, e.g. by an order of magnitude). A drawback of using an autoencoder for speech coding is that it cannot necessarily exploit the temporal relationship between sets of input data. U.S. Patent No. 11,526,734 (“the ‘734 patent”, assigned to Qualcomm, Inc.), "Method and Apparatus for recurrent auto encoding", Yang et. al. improves upon an autoencoder to exploit temporal redundancies and correlations and describes a feedback recurrent autoencoder ("FRAE"). The FRAE in the '734 patent can be used for training and application of compression of sequential data with temporal correlation. The recurrent structure of the FRAE can be used to efficiently extract the redundancy embedded along the time-dimension of sequential data and enables compact discrete representation of the data at the bottleneck in a sequential fashion. For example, Table 1 of the '734 patent illustrates the MSE (Mean Squared Error) used as Mel-scale mean-square-error configured as a reconstruction loss for training of the FRAE, where the MSE of eachQualcomm Docket No.2404227WO frequency bin is scaled according to its weight at Mel-frequency both for latent feedback and output feedback.
[0054] The FRAE described in the ‘734 patent has two advantageous features not present in autoencoders: (1) recurrent layers, e.g. LSTM or GRU layers, have memory of the past; and (2) feedback from the decoder of the autoencoder to the encoder of the autoencoder. The feedback connection 150 in Fig.2, Fig.3 and Fig. 7 of the ‘734 patent provides additional historical information from a state (ht) in the decoder to the encoder indicative of how reconstruction of a prior input has fared. This feedback loop is present during training and inference, which allows the encoder to be trained to respond to particular feedback conditions in a manner that improves reproduction of the input data by the decoder. Thus, the feedback loop is analogous to a mode switch input in that it is not encoded by the encoder, but it is used as an input that influences how the encoder operates on the next input vector.
[0055] In addition, the ‘734 patent describes a second feedback connection (152) from the state (ht) of the decoder for a first iteration of series of inputs Xt, to the next iteration of series of inputs Xt+1. This second feedback connection (152) enables the decoder to learn from its previous reconstruction attempts, providing additional historical context about how reconstruction of a prior input has fared. This second feedback loop is also present during both training and inference, which allows the decoder to be trained to respond to particular feedback conditions in a manner that improves reproduction of the input data by the decoder.
[0056] The ‘734 patent also describes other optional feedback connections. For example, an embedding vector (z) may be fed back via a third feedback connection (356 in Fig.3 of the ‘734 patent) to the encoder. As another optional example, the '734 patent describes that the FRAE can be a variational autoencoder. In this example, an output of the decoder (754 in Fig. 7 of the ‘734 patent) is sample and parameterized to a generate an autoregressive prior which can be used as a fourth feedback connection to condition a prior model for the next latent space embedded vector (zt+1). Similar to the previously described feedback connection, each of these optional connections is present both training and inference, and thus are trained as part of the training process thereby improving functionality of the FRAE.
[0057] The decoded spectral envelope features and the decoded and / or dequantizedQualcomm Docket No.2404227WO pitch information determined by the decoder of the generative voice codec can be processed by a neural speech synthesizer included in the decoder. For example, the neural speech synthesizer can be a neural homomorphic vocoder (NHV), among various other neural speech synthesizers. Based on implementing a neural speech synthesizer (e.g., NHV, etc.) configured to process the decoded spectral envelope features and / or pitch information corresponding to the encoded audio signal, the systems and techniques can be used to provide a generative decoder (e.g., a signal synthesis decoder) for generating the reconstructed audio signal. For example, the neural speech synthesizer of the decoder can generate a waveform corresponding to a reconstructed audio signal, based on using one or more generative machine learning models (e.g., generative neural network(s)) included in the neural speech synthesizer to process the decoded spectral envelope features and pitch information received from the encoder.
[0058] Further aspects of the systems and techniques will be described with respect to the figures.
[0059] FIG.1 illustrates an example implementation of a system-on-a-chip (SoC) 100, which may include a central processing unit (CPU) 102 or a multi-core CPU, configured to perform one or more of the functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with a computational device (e.g., neural network with weights), delays, frequency bin information, task information, among other information may be stored in a memory block associated with a neural processing unit (NPU) 108, in a memory block associated with a CPU 102, in a memory block associated with a graphics processing unit (GPU) 104, in a memory block associated with a digital signal processor (DSP) 106, in a memory block 118, and / or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from a memory block 118.
[0060] The SoC 100 may also include additional processing blocks tailored to specific functions, such as a GPU 104, a DSP 106, a connectivity block 110, which may include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, and the like, and a multimedia processor 112 that may, for example, detect and recognize gestures, speech, and / or other interactive user action(s) or input(s). In one implementation, the NPUQualcomm Docket No.2404227WO 108 is implemented in the CPU 102, DSP 106, and / or GPU 104. The SoC 100 may also include a sensor processor 114, image signal processors (ISPs) 116, and / or a signal synthesis system 120. For example, the signal synthesis system 120 can be implemented or configured as a speech synthesis system, including a neural speech decoder and / or a neural homomorphic vocoder (NHV) system, which can be used to generate speech (e.g., perform speech synthesis), can be implemented in a text-to-speech (TTS) system, etc. In some examples, the sensor processor 114 can be associated with or connected to one or more sensors for providing sensor input(s) to sensor processor 114. For example, the one or more sensors and the sensor processor 114 can be provided in, coupled to, or otherwise associated with a same computing device.
[0061] In some examples, the one or more sensors can include one or more microphones for receiving sound (e.g., an audio input), including sound or audio inputs that can be used to perform various speech synthesis tasks and / or to generate a reconstructed speech signal from the audio input, etc. In some cases, the sound or audio input received by the one or more microphones (and / or other sensors) may be digitized into data packets for analysis and / or transmission. The audio input may include ambient sounds in the vicinity of a computing device associated with the SoC 100 and / or may include speech from a user of the computing device associated with the SoC 100. In some cases, a computing device associated with the SoC 100 can additionally, or alternatively, be communicatively coupled to one or more peripheral devices (not shown) and / or configured to communicate with one or more remote computing devices or external resources, for example using a wireless transceiver and a communication network, such as a cellular communication network.
[0062] SoC 100, DSP 106, NPU 108 and / or signal synthesis (e.g., neural speech decoder, NHV, etc.) system 120 may be configured to perform audio signal processing. For example, the signal synthesis system 120 may be configured to perform steps for speech synthesis and / or neural homomorphic vocoding, etc. As another example, one or more portions of the steps, such as feature generation, for speech synthesis and / or NHV may be performed by the signal synthesis system 120 while the DSP 106 / NPU 108 performs other steps, such as steps using one or more machine learning networks and / or machine learning techniques according to aspects of the present disclosure and as described herein.Qualcomm Docket No.2404227WO
[0063] FIG. 2A depicts an example of a feature generator 200, in accordance with aspects of the present disclosure. It should be understood that many techniques may be used to generate feature vectors for an audio input and that feature generator 200 is just a single example of a technique that may be used to generate feature vectors.
[0064] Feature generator 200 receives an audio signal at signal pre-processor 202. As above, the audio signal may be from an audio source of an electronic device, such a microphone. Signal pre-processor 202 may perform various pre-processing steps on the received audio signal. For example, signal pre-processor 202 may split the audio signal into parallel audio signals and delay one of the signals by a predetermined amount of time to prepare the audio signals for input into a Fast-Fourier Transform (FFT) circuit. As another example, signal pre-processor 202 may perform a windowing function, such as a Hamming, Hann, Blackman-Harris, Kaiser-Bessel window function, or other sine-based window function, which may improve the performance of further processing stages, such as signal domain transformer 204. Generally, a windowing (or window) function in may be used to reduce the amplitude of discontinuities at the boundaries of each finite sequence of received audio signal data to improve further processing. As another example, signal pre-processor 202 may convert the audio signal data from parallel to serial, or vice versa, for further processing. The pre-processed audio signal from the signal pre-processor 202 may be provided to signal domain transformer 204, which may transform the pre-processed audio signal from a first domain into a second domain, such as from a time domain into a frequency domain.
[0065] In some aspects, signal domain transformer 204 implements a Fourier transform, such as a Fast-Fourier transform (FFT). For example, in some cases, the Fast Fourier transform may be a 16-band (or bin, channel, or point) FFT, which generates a compact feature set that may be efficiency processed by a model. In some cases, a Fourier transform provides fine spectral domain information about the incoming audio signal as compared to conventional single channel processing, such as conventional hardware SNR threshold detection. The result of signal domain transformer 204 is a set of audio features, such as a set of voltages, powers, or energies per frequency band in the transformed data.
[0066] The set of audio features may then be provided to signal feature filter 206, which may reduce the size of or compress the feature set in the audio feature data. In some aspects, signal feature filter 206 may discard certain features from the audio feature set,Qualcomm Docket No.2404227WO such as symmetric or redundant features from multiple bands of a multi-band FFT. Discarding this data reduces the overall size of the data stream for further processing and may be referred to a compressing the data stream. For example, in some cases, a 16-band FFT may include 8 symmetric or redundant bands of after the powers are squared because audio signals are real. Thus, signal feature filter 206 may filter out the redundant or symmetric band information and output an audio feature vector 208. In some cases, output of the signal feature filter may be compressed or otherwise processed prior to output as the audio feature vector 208. The audio feature vector 208 may be provided to a speech synthesis system (e.g., such as the signal synthesis system 120 of FIG. 1) for processing by speech synthesis or NHV model.
[0067] Audio coding (e.g., speech coding, music signal coding, or other type of audio coding) can be performed on a digitized audio signal (e.g., a speech signal) to compress the amount of data for storage, transmission, and / or other use. FIG.2B is a block diagram illustrating an example of a voice coding system 250 (which can also be referred to as a voice or speech coder or a voice coder-decoder (codec)). A voice encoder 252 of the voice coding system 250 can use a voice coding algorithm to process a speech signal 251. The speech signal 251 can include a digitized speech signal generated from an analog speech signal from a given source. For instance, the digitized speech signal can be generated using a filter to eliminate aliasing, a sampler to convert to discrete-time, and an analog- to-digital converter for converting the analog signal to the digital domain. The resulting digitized speech signal (e.g., speech signal 251) is a discrete-time speech signal with sample values (referred to herein as samples) that are also discretized.
[0068] Using the voice coding algorithm, the voice encoder 252 can generate a compressed signal (including a lower bit-rate stream of data) that represents the speech signal 251 using as few bits as possible, while attempting to maintain a certain quality level for the speech. The voice encoder 252 can use any suitable voice coding algorithm, such as a linear prediction coding algorithm (e.g., Code-excited linear prediction (CELP), algebraic-CELP (ACELP), or other linear prediction technique) or other voice coding algorithm.
[0069] The voice encoder 252 can compress the speech signal 251 in an attempt to reduce the bit-rate of the speech signal 251. The bit-rate of a signal is based on the sampling frequency and the number of bits per sample. For instance, the bit-rate of aQualcomm Docket No.2404227WOspeech signal can be determined as ^^^^ ൌ ^^ ∗ ^^, where BR is the bit-rate, S is the samplingfrequency, and b is the number of bits per sample. In one illustrative example, at a sampling frequency (S) of 8 kilohertz (kHz) and at 16 bits per sample (b), the bit-rate (BR) of a signal would be a bit-rate of 128 kilobits per second (kbps).
[0070] The compressed speech signal can then be stored and / or sent to and processed by a voice decoder 254. In some examples, the voice decoder 254 can communicate with the voice encoder 252, such as to request speech data, send feedback information, and / or provide other communications to the voice encoder 252. In some examples, the voice encoder 252 or a channel encoder can perform channel coding on the compressed speech signal before the compressed speech signal is sent to the voice decoder 254. For instance, channel coding can provide error protection to the bitstream of the compressed speech signal to protect the bitstream from noise and / or interference that can occur during transmission on a communication channel.
[0071] The voice decoder 254 can decode the data of the compressed speech signal and construct a reconstructed speech signal 255 that approximates the original speech signal 251. The reconstructed speech signal 255 includes a digitized, discrete-time signal that can have the same or similar bit-rate as that of the original speech signal 251. The voice decoder 254 can use an inverse of the voice coding algorithm used by the voice encoder 252, which as noted above can include any suitable voice coding algorithm, such as a linear prediction coding algorithm (e.g., CELP, ACELP, or other suitable linear prediction technique) or other voice coding algorithm. In some cases, the reconstructed speech signal 255 can be converted to continuous-time analog signal, such as by performing digital-to- analog conversion and anti-aliasing filtering.
[0072] Voice coders can exploit the fact that speech signals are highly correlated waveforms. The samples of an input speech signal can be divided into blocks of N samples each, where a block of N samples is referred to as a frame. In one illustrative example, each frame can be 10-20 milliseconds (ms) in length.
[0073] Various voice coding algorithms can be used to encode a speech signal. For instance, code-excited linear prediction (CELP) is one example of a voice coding algorithm. The CELP model is based on a source-filter model of speech production, which assumes that the vocal cords are the source of spectrally flat sound (an excitation signal), and that the vocal tract acts as a filter to spectrally shape the various sounds of speech.Qualcomm Docket No.2404227WO The different phonemes (e.g., vowels, fricatives, and voice fricatives) can be distinguished by their excitation (source) and spectral shape (filter).
[0074] In general, CELP uses a linear prediction (LP) model to model the vocal tract, and uses entries of a fixed codebook (FCB) as input to the LP model. For instance, long- term linear prediction can be used to model pitch of a speech signal, and short-term linear prediction can be used to model the spectral shape (phoneme) of the speech signal. Entries in the FCB are based on coding of a residual signal that remains after the long-term and short-term linear prediction modeling is performed. For example, long-term linear prediction and short-term linear prediction models can be used for speech synthesis, and a fixed codebook (FCB) can be searched during encoding to locate the best residual for input to the long-term and short-term linear prediction models. The FCB provides the residual speech components not captured by the short-term and long-term linear prediction models. A residual, and a corresponding index, can be selected at the encoder based on an analysis-by-synthesis process that is performed to choose the best parameters so as to match the original speech signal as closely as possible. The index can be sent to the decoder, which can extract the corresponding LTP residual from the FCB based on the index.
[0075] FIG. 2C is a block diagram illustrating an example of a CELP-based voice coding system 270, including a voice encoder 272 and a voice decoder 274. The voice encoder 272 can obtain a speech signal 271 and can segment the samples of the speech signal into frames and sub-frames. For instance, a frame of N samples can be divided into sub-frames. In one illustrative example, a frame of 240 samples can be divided into four sub-frames each having 60 samples. For each frame, sub-frame, or sample, the voice encoder 272 chooses the parameters (e.g., gain, filter coefficients or linear prediction (LP) coefficients, etc.) for a synthetic speech signal so as to match as much as possible the synthetic speech signal with the original speech signal.
[0076] The voice encoder 272 can include a short-term linear prediction (LP) engine 280, a long-term linear prediction (LTP) engine 282, and a fixed codebook (FCB) 284. The short-term LP engine 280 models the spectral shape (phoneme) of the speech signal. For example, the short-term LP engine 280 can perform a short-term LP analysis on each frame to yield linear prediction (LP) coefficients. In some examples, the input to the short- term LP engine 280 can be the original speech signal or a pre-processed version of theQualcomm Docket No.2404227WO original speech signal. In some implementations, the short-term LP engine 280 can perform linear prediction for each frame by estimating the value of a current speech sample based on a linear combination of past speech samples. For example, a speechsignal s(n) can be represented using an autoregressive (AR) model, such as ^^^^^^ ൌ∑^ ^ୀ^ ^^^^^^^^ െ ^^^ ^ ^^^^^^, where each sample is represented as a linear combination of the previous m samples plus a prediction error term ^^^^^^. The weighting coefficients a1, a2, through amcan be referred to as the LP coefficients. The prediction error term ^^^^^^can be found as follows: ^^^^^^ ൌ ^^^^^^ െ ∑^ ^ୀ^ ^^^^^^^^ െ ^^^. By minimizing the mean square prediction error with respect to the filter coefficients, the short-term LP engine 280 can obtain the LP coefficients. The LP coefficients can be used to form an analysis filter as given in Equation 1, below: ^ ^^^^^^ ൌ 1 െ ^ ^^ ି^^^^Eq. (1)
[0077] The short-term LP engine 280 can solve for ^^^^^^ (which can be referred to as a transfer function) by computing the LP coefficients (^^^) that minimize the error in the above AR model equation (^^^^^^) or other error metric. In some implementations, the LP coefficients can be determined using a Levinson–Durbin method, a Leroux–Gueguen algorithm, or other suitable technique. In some examples, the voice encoder 272 can send the LP coefficients to the voice decoder 274. In some examples, the voice decoder 274 can determine the LP coefficients, in which case the voice encoder 272 may not send the LP coefficients to the voice decoder 274. In some examples, Line Spectral Frequencies (LSFs) can be computed instead of or in addition to LP coefficients.
[0078] The LTP engine 282 models the pitch of the speech signal. Pitch is a feature that determines the spacing or periodicity of the impulses in a speech signal. For example, speech signals are generated when the airflow from the lungs is periodically interrupted by movements of the vocal cords. The time between successive vocal cord openings corresponds to the pitch period. The LTP engine 282 can be applied to each frame or each sub-frame of a frame after the short-term LP engine 280 is applied to the frame. The LTP engine 282 can predict a current signal sample from a past sample that is one or more pitch periods apart from a current sample (hence the term “long-term”). For instance, thecurrent signal sample can be predicted as ^^^^^^^ ൌ ^^^^^^^^ െ ^^^, where T denotes the pitchQualcomm Docket No.2404227WOperiod, ^^^ denotes the pitch gain, and ^^^^^ െ ^^^ denotes an LP residual for a previoussample one or more pitch periods apart from a current sample. Pitch period can be estimated at every frame. By comparing a frame with past samples, it is possible to identify the period in which the signal repeats itself, resulting in an estimate of the actual pitch period. The LTP engine 282 can be applied separately to each sub-frame.
[0079] The FCB 284 can include a number (denoted as L) of long-term linear prediction (LTP) residuals. An LTP residual includes the speech signal components that remain after the long-term and short-term linear prediction modeling is performed. The LTP residuals can be, for example, fixed or adaptive and can contain deterministic pulses or random noise (e.g., white noise samples). The voice encoder 272 can pass through the number L of LTP residuals in the FCB 284 a number of times for each segment (e.g., each frame or other group of samples) of the input speech signal, and can calculate an error value (e.g., a mean-squared error value) after each pass. The LTP residuals can be represented using codevectors. The length of each codevector can be equal to the length of each sub-frame, in which case a search of the FCB 284 is performed once every sub-frame. The LTP residual providing the lowest error can be selected by the voice encoder 272. The voice encoder 272 can select an index corresponding to the LTP residual selected from the FCB 284 for a given sub-frame or frame. The voice encoder 272 can send the index to the voice decoder 274 indicating which LTP residual is selected from the FCB 284 for the given sub-frame or frame. A gain associated with the lowest error can also be selected, and send to the voice decoder 274.
[0080] The voice decoder 274 includes an FCB 294, an LTP engine 292, and a short- term LP engine 290. The FCB 294 has the same LTP residuals (e.g., codevectors) as the FCB 284. The voice decoder 274 can extract an LTP residual from the FCB 294 using the index transmitted to the voice decoder 274 from the voice encoder 272. The extracted LTP residual can be scaled to the appropriate level and filtered by the LTP engine 292 and the short-term LP engine 290 to generate a reconstructed speech signal 275. The LTP engine 292 creates periodicity in the signal associated with the fundamental pitch frequency, and the short-term LP engine 290 generates the spectral envelope of the signal.
[0081] Other linear predictive-based coding systems can also be used to code voice signals, including enhanced voice services (EVS), adaptive multi-rate (AMR) voice coding systems, mixed excitation linear prediction (MELP) voice coding systems, linearQualcomm Docket No.2404227WO predictive coding-10 (LPC-10), among others.
[0082] A voice codec for some applications and / or devices (e.g., Internet-of-Things (IoT) applications and devices) may be needed to deliver higher quality coding of speech signals at low bit-rates, with low complexity, and with low memory requirements. Existing linear predictive-based codecs cannot meet such requirements. For example, ACELP-based coding systems provide high quality, but do not provide low bit-rate or low complexity / low memory. Other linear-predictive coding systems provide low bit-rate and low complexity / low memory, but do not provide high quality. In some cases, machine learning systems (e.g., using a neural network model) can be used to generate reconstructed voice or audio signals. For example, using features extracted from a frame of audio data, a neural network-based voice decoder can generate coefficients for at least one linear filter. The linear filter can then be used to generate a reconstructed signal. However, such a neural network-based voice decoder can be highly complex and resource intensive. For instance, the neural network-based voice decoder will have to perform the operations of a linear predictive filter (LPC), such as the short-term LP engine 280 of FIG. 2C. Such LPC operations can include complex operations that require the use of a large amount of computing resources by the neural network-based voice decoder.
[0083] FIG.3 is a diagram illustrating an example of a voice decoding signal synthesis system 300 utilizing a linear time-varying filter 304 with coefficients generated using a neural network (NN) filter estimator 302 and a separate linear predictive coding (LPC) filter 306. The voice decoding signal synthesis system 300 is configured to decode data of the compressed speech signal to generate a reconstructed speech signal ^̂^^^^^ (also referred to as a synthesized speech sample) for a current time instant n that approximates an original speech signal that was previously compressed by a voice encoder (not shown). The voice encoder can be similar to and can perform some or all of the functions of the voice encoder 252 described above with respect to FIG. 2B, or other type of voice encoder. For example, the voice encoder can include a short-term LP engine, an LTP engine, and an FCB. In another example, the voice encoder can include a magnitude spectrum generator (e.g., Mel-scale magnitude spectrum or full spectrum magnitude), a short-term linear prediction (LP) engine, and a pitch tracker that detects a fundamental pitch harmonic frequency of the speech and pitch correlation.
[0084] The voice encoder can extract (and in some cases quantize) a set of featuresQualcomm Docket No.2404227WO (referred to as a feature set) from the speech signal, and can send the extracted (and in some cases quantized) feature set to the voice decoding signal synthesis system 300. The features that are computed by the voice encoder can depend on a particular encoder implementation used. Various illustrative examples of feature sets are provided below according to different encoder implementations, which can be extracted by the voice encoder (and in some cases quantized), and sent to the voice decoding signal synthesis system 300. However, one of ordinary skill will appreciate that other feature sets can be extracted by the voice encoder. For example, the voice encoder can extract any set of features, can quantize that feature set, and can send the feature set to the voice decoding signal synthesis system 300.
[0085] As noted above, various combinations of features can be extracted as a feature set by the voice encoder. For example, a feature set can include one or any combination of the following features: Linear Prediction (LP) coefficients; Line Spectral Pairs (LSPs); Line Spectral Frequencies (LSFs); pitch lag with integer or fractional accuracy; pitch gain; pitch correlation; Mel-scale frequency cepstral coefficients (also referred to as Mel cepstrum) of the speech signal; Bark-scale frequency cepstral coefficients (also referred to as bark cepstrum) of the speech signal; Mel-scale frequency cepstral coefficients of the LTP residual; Bark-scale frequency cepstral coefficients of the LTP residual; a spectrum (e.g., Discrete Fourier Transform (DFT) or other spectrum) of the speech signal; and / or a spectrum (e.g., DFT or other spectrum) of the LTP residual; voicing level of each frequency band of each speech frame; fundamental frequency of pitch harmonics; pitch correlation of each speech frame; time domain pitch lag of each speech frame.
[0086] For any one or more of the other features listed above, the voice encoder can use any estimation and / or quantization method, such as an engine or algorithm from any suitable voice codec (e.g. EVS, AMR, or other voice codec) or a neural network-based estimation and / or quantization scheme (e.g., convolutional or fully-connected (dense) or recurrent Autoencoder, or other neural network-based estimation and / or quantization scheme). The voice encoder can also use any frame size, frame overlap, and / or update rate for each feature. The voice encoder can also include extra redundancies in the features to ensure robustness of operation against packet losses. Examples of estimation and quantization methods for each example feature are provided below for illustrative purposes, where other examples of estimation and quantization methods can be used by the voice encoder.Qualcomm Docket No.2404227WO
[0087] As noted above, one example of features that can be extracted from a voice signal by the voice encoder includes linear prediction (LP) coefficients and / or line spectral frequencies (LSFs). Various estimation techniques can be used to compute the LP coefficients and / or LSFs. For example, the voice encoder can estimate LP coefficients (and / or LSFs) from a speech signal using the Levinson-Durbin algorithm. In some examples, the LP coefficients and / or LSFs can be estimated using an autocovariance method for LP estimation. In some cases, the LP coefficients can be determined, and an LP to LSF conversion algorithm can be performed to obtain the LSFs. Any other LP and / or LSF estimation engine or algorithm can be used, such as an LP and / or LSF estimation engine or algorithm from an existing codec (e.g., EVS, AMR, or other voice codec).
[0088] Various quantization techniques can be used to quantize the LP coefficients and / or LSFs. For example, the voice encoder can use a single stage vector quantization (SSVQ) technique, a multi-stage vector quantization (MSVQ), or other vector quantization technique to quantize the LP coefficients and / or LSFs. In some cases, a predictive or adaptive SSVQ or MSVQ (or other vector quantization technique) can be used to quantize the LP coefficients and / or LSFs. In another example, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the LP coefficients and / or LSFs. Any other LP and / or LSF quantization engine or algorithm can be used, such as an LP and / or LSF quantization engine or algorithm from an existing codec (e.g., EVS, AMR, or other voice codec).
[0089] Another example of features that can be extracted from a voice signal by the voice encoder includes pitch lag (integer and / or fractional), pitch gain, and / or pitch correlation. Various estimation techniques can be used to compute the pitch lag, pitch gain, and / or pitch correlation. For example, the voice encoder can estimate the pitch lag, pitch gain, and / or pitch correlation (or any combination thereof) from a speech signal using any pitch lag, gain, correlation estimation engine or algorithm (e.g. autocorrelation- based pitch lag estimation). For example, the voice encoder can use a pitch lag, gain, and / or correlation estimation engine (or algorithm) from any suitable voice codec (e.g. EVS, AMR, or other voice codec). Various quantization techniques can be used to quantize the pitch lag, pitch gain, and / or pitch correlation. For example, the voice encoder can quantize the pitch lag, pitch gain, and / or pitch correlation (or any combination thereof) from a speech signal using any pitch lag, gain, correlation quantization engine orQualcomm Docket No.2404227WO algorithm from any suitable voice codec (e.g. EVS, AMR, or other voice codec). In some cases, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the pitch lag, pitch gain, and / or pitch correlation features.
[0090] Another example of features that can be extracted from a voice signal by the voice encoder includes the Mel cepstrum coefficients and / or Bark cepstrum coefficients of the speech signal, and / or the Mel cepstrum coefficients and / or Bark cepstrum coefficients of the LTP residual. Various estimation techniques can be used to compute the Mel cepstrum coefficients and / or Bark cepstrum coefficients. For example, the voice encoder can use a Mel or Bark frequency cepstrum technique that includes Mel or Bark frequency filter banks computation, filter bank energy computation, logarithm application, and discrete cosine transform (DCT) or truncation of the DCT. Various quantization techniques can be used to quantize the Mel cepstrum coefficients and / or Bark cepstrum coefficients. For example, vector quantization (single stage or multistage) or predictive / adaptive vector quantization can be used. In some cases, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the Mel cepstrum coefficients and / or Bark cepstrum coefficients. Any other suitable cepstrum quantization methods can be used.
[0091] Another example of features that can be extracted from a voice signal by the voice encoder includes the spectrum of the speech signal and / or the spectrum of the LTP residual. Various estimation techniques can be used to compute the spectrum of the speech signal and / or the LTP residual. For example, a Discrete Fourier transform (DFT), a Fast Fourier Transform (FFT), or other transform of the speech signal can be determined. Quantization techniques that can be used to quantize the spectrum of the voice signal can include vector quantization (single stage or multistage) or predictive / adaptive vector quantization. In some cases, an autoencoder or other neural network based technique can be used by the voice encoder to quantize the spectrum. Any other suitable spectrum quantization methods can be used.
[0092] As noted above, any one of the above-described features or any combination of the above-described features can be estimated, quantized, and sent by the voice encoder to the voice decoding signal synthesis system 300 depending on the particular encoder implementation that is used. In one illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, and pitchQualcomm Docket No.2404227WO correlation. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum of the speech signal. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the speech signal. In another illustrative example, the voice encoder can estimate, quantize, and send pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum of the speech signal. In another illustrative example, the voice encoder can estimate, quantize, and send pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the speech signal. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the Bark cepstrum of the LTP residual. In another illustrative example, the voice encoder can estimate, quantize, and send LP coefficients, pitch lag with fractional accuracy, pitch gain, pitch correlation, and the spectrum (e.g., DFT, FFT, or other spectrum) of the LTP residual.
[0093] The voice decoding signal synthesis system 300 includes a neural network filter estimator 302, a linear time-varying filter 304 generated by the neural network, and a linear predictive coding (LPC) filter 306. The LPC filter 306 can include a time-varying LPC filter. The neural network filter estimator 302 is trained to generate filter coefficients for the linear time-varying filter 304. The neural network model of the neural network filter estimator 302 can include any neural network architecture that can be trained to model the filter coefficients for the linear time-varying filter 304. Examples of neural network architectures that can be included in the neural network filter estimator 302 include a generative neural network (e.g., a generative-adversarial network (GAN)), convolutional neural networks (CNN), an autoencoder, and / or other type(s) of neural network architectures or models.
[0094] The voice decoding signal synthesis system 300 (e.g., the neural network model of the neural network filter estimator 302) can be trained using any suitable neural network training technique. In some examples, the neural network model of the neural network filter estimator 302 can be trained using supervised learning techniques based on backpropagation. For instance, corresponding input and target output pairs can be provided to the neural network filter estimator 302 for training. In one example, for each time instant n, the input to the neural network filter estimator 302 can include log-Mel-Qualcomm Docket No.2404227WOfrequency spectrum features or coefficients (e.g., 80 log-Mel features ^^^^^, ^^^ 301 shownin FIG.3). In some examples, the target output (or label or ground truth) for training the neural network filter estimator 302 can include the target speech sample ^^^^^^305 for the current time instant n, as shown in FIG.3. In such examples, a loss 307 will be computed based on the reconstructed sample ^̂^^^^^ and the target output speech sample ^^^^^^ 305 for time instant n. In some examples, the target output can include a speech signal ^^^^^^ that is generated after passing the target speech through an LPC analysis filter (inverse of LPC filter 306). In this case, the loss will be computed based on the output ^̂^^^^^ generated by the linear time-varying filter generated by the neural network 304 and the target output ^^^^^^ (e.g., where both are in speech residual domain).
[0095] Backpropagation can be performed to train the neural network filter estimator 302 using the inputs and the target output. Backpropagation can include a forward pass, a loss function, a backward pass, and a parameter update to update one or more parameters (e.g., weight, bias, or other parameter). The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. The process can be repeated for a certain number of iterations for each set of inputs until the neural network filter estimator 302 is trained well enough so that the weights (and / or other parameters) of the various layers are accurately tuned.
[0096] In some aspects, training of the neural network filter estimator 302 and / or training of one or more of the machine learning systems or neural networks described herein can be performed using online training (e.g., in some case on-device training), offline training, and / or various combinations of online and offline training.
[0097] In some cases, online may refer to time periods during which the input data is processed, for instance for performance of the generative voice codec processing implemented by the systems and techniques described herein. In some examples, offline may refer to idle time periods or time periods during which input data is not being processed. Additionally, offline may be based on one or more time conditions (e.g., after a particular amount of time has expired, such as a day, a week, a month, etc.) and / or may be based on various other conditions such as network and / or server availability, etc., among various others. In some aspects, offline training of a machine learning model (e.g., a neural network model) can be performed by a first device (e.g., a server device) to generate a pre-trained model, and a second device can receive the trained model from theQualcomm Docket No.2404227WO second device. In some cases, the second device (e.g., a mobile device, an XR device, a vehicle or system / component of the vehicle, or other device) can perform online (or on- device) training of the pre-trained model to further adapt or tune the parameters of the model.
[0098] In some cases, the forward pass can include passing the input data (e.g., the log-Mel-frequency spectrum features or coefficients, such as the 80 log-Mel features ^^^^^, ^^^301 shown in FIG.3) through the neural network filter estimator 302. The weights of the neural network model are initially randomized before the neural network filter estimator 302 is trained. For a first training iteration for the neural network filter estimator 302, the output will likely include values that do not give preference to any particular output due to the weights being randomly selected at initialization. With the initial weights, the neural network filter estimator 302 is unable to determine low level features and thus cannot make an accurate estimation of the filter coefficients for the linear time-varying filter 304. A loss function can be used to analyze the loss 307 (or error) in the reconstructed or synthesized sample output ^̂^^^^^. Any suitable loss function definition can be used. One example of a loss function includes a mean squared error (MSE). The MSE is defined as^^ ൌ ∑^^^^ ଶ௧^௧^^ଶ^^^^^^^^^^ െ ^^^^^^^^^^^^^ , which calculates the sum of one-half times the actualanswer minus the predicted (output) answer squared. The loss can be set to be equal to the value of ^^௧^௧^^. Other loss functions may include a difference of magnitude spectrums between the target and output signals, where the difference may be computed as absolute difference, squared difference, or logarithmic difference between the magnitude spectrum of each speech frame, and then aggregated over all speech frames.
[0099] The loss (or error) will be high for the first training iterations since the actual values will be much different than the predicted output. The goal of training is to minimize the amount of loss so that the predicted output is the same as the training label. The neural network filter estimator 302 can perform a backward pass by determining which inputs (weights) most contributed to the loss of the network, and can adjust the weights so that the loss decreases and is eventually minimized. A derivative of the loss with respect to the weights (denoted as dL / dW, where W represents the weights at a particular layer) can be computed to determine the weights that contributed most to the loss of the network. After the derivative is computed, a weight update can be performed by updating all the weights of the filters. For example, the weights can be updated so that they change in theQualcomm Docket No.2404227WOopposite direction of the gradient. The weight update can be denoted as ^^ ൌ ^^ௗ^ ^െ ^^ௗ^, where w denotes a weight, widenotes the initial weight, and η denotes a learningThe learning rate can be set to any suitable value, with a high learning rateweight updates and a lower value indicating smaller weight updates. In some examples, to train the neural network filter estimator 302, a multi-resolution STFT loss ^^ோand adversarial losses ^^ீand ^^^can be computed from ^^^^^^ and ^^^^^^. Because linear time- varying filters are fully differentiable, gradients can propagate back to the neural network filter estimator 302.
[0100] Using the filter coefficients generated by the neural network filter estimator 302, the linear time-varying filter 304 can process an excitation signal 303 to generate another signal ^̂^^^^^. The signal ^̂^^^^^ can be used as an excitation signal to excite the LPC filter 306. The linear time-varying filter 304 is a linear filter, which preserves the linearity property between inputs and outputs. For instance, a linear filter is associated with amapping ℒ:ℝℤ → ℝℤ, ^^^^^^ → ^^^^^^ ൌ ℒ^^^^^^^^ that has the following property: for any^^^^^^^, ^^ଶ^^^^ ∈ ℝℤ and any ^^, ^^ ∈ ℝ, ℒ^^^ ^^^^^^^ ^ ^^ ^^ଶ^^^^^ ൌ ^^ ℒ^^^^^^^^^ ^ ^^ ℒ^^^ଶ^^^^^.In one illustrative example, if input1 produces output1 and input2 produces output2, then a combined input of (input1+input2) will produce output = output1+output2. The time- varying nature of the linear time-varying filter 304 indicates that the filter response depends on the time of excitation of the linear time-varying filter 304 (e.g., a new set of coefficients used to filter each frame (block of time) of input at the time of excitation). In some cases, time-varying linear filters can be characterized by the set of impulseresponses at each time lag ℎ^^^^^ ൌ ℒ^^^^^^ െ ^^^^ for each ^^ ∈ ℤ. In some examples, theoutput of a time varying linear filter is ℒ^^^^^^^^ ൌ ∑^ ^^^^^^ ⋆ ℎ^^^^^ (where ⋆ is aconvolutional operator) or some heuristic combination of filter input and impulse responses, e.g. overlap-add on windowed and filtered signal segments, etc.
[0101] The LPC filter 306 can use the signal ^̂^^^^^ as input to generate the reconstructed or synthesized speech sample ^̂^^^^^ for the current time instant n. The LPC filter 306 is a linear filter and in some cases is time varying, as defined above with respect to the linear time-varying filter 304. The LPC filter 306 can be a form of time-varying filter used for processing of speech. The LPC filter 306 includes filter coefficients for each speech frame that can be computed using the autocorrelation of a speech or audio signal.
[0102] In some examples, the LPC filter 306 can be used to model the spectral shapeQualcomm Docket No.2404227WO (or phenome or envelope) of the speech signal. For example, at the voice encoder, a signal ^^^^^^ can be filtered by an autoregressive (AR) process synthesizer to obtain an AR signal. As described above, a linear predictor can be used to predict the AR signal (which can be denoted as prediction ^̂^^^^^) as a linear combination of the previous m samples as follows: ^ ^^^^^^^ ൌ ^െ^^^^^^ି^ ൌ 1 െ ^^^^^^^,
[0103] where the ^^^^of the AR parameters (also referred to as LP coefficients). A residual signal can be the difference between the original AR signal and the predicted AR signal represented as prediction ^̂^^^^^ (e.g., the difference between the actual sample and the predicted sample). At the voice encoder, the linear prediction coding can be used to find the best linear prediction coefficients for minimizing a quadratic error function, and thus the error. The linear prediction process removes the short-term correlation from the speech signal. The linear prediction coefficients are an efficient way to represent the short-term spectrum of the speech signal. At the voice decoding signal synthesis system 300, the LPC filter 306 determines the prediction ^̂^^^^^ for the current sample n using computed or received coefficients and A transfer function ^^^^^^^. For instance, in some examples, the LPC filter coefficients are received from the encoder. In other examples, the voice decoding signal synthesis system 300 can derive the LPC filter coefficients, such as using other features (e.g., Mel spectrum features) sent by an encoder to the voice decoding signal synthesis system 300. For instance, the voice decoding signal synthesis system 300 can use Mel spectrum features 301 to derive the LPC filter coefficients for the LPC filter 306. The LPC filter 306 can determine the final reconstructed (or predicted) sample ^̂^^^^^ using the output ^̂^^^^^ from the linear time- varying filter 304 (for the current sample n).
[0104] In some examples, further components can be used along with the neural network filter estimator 302 and the linear time-varying filter 304, such as an impulse train generator 314 and a random noise generator 316. In some examples, the linear time- varying filter 304 can include a harmonic linear time-varying filter 318 and a noise linear time-varying filter 320. A voice encoder can include a pitch tracker 310 and a feature extraction engine 312. In some examples, the feature extraction engine 312 can be the same as or similar to the feature generator 200 of FIG. 2A. In some cases, the original speech signal ^^ and reconstructed signal ^^ are divided into non-overlapping frames withQualcomm Docket No.2404227WO frame length L. The term ^^ can be defined as a frame index, the term ^^ can be defined as a discrete time index, and the term ^^ can be defined as a feature index. The total numberof frames ^^ and total number of sampling points ^^ may follow ^^ ൌ ^^ ൈ ^^. In ^^^, ^^, ℎ^,ℎ^, 0 ^ ^^ െ 1. The terms ^^, ^^, ^^, ^^, ^^^, ^^^ are finite duration signals, in which 0 ^ ^^ ^^^ െ 1. Impulse responses ℎ^, and ℎ^ may be infinitely long, in which ^^ ∈ ℤ. Impulseresponse h may be causal, in which ^^ ∈ ℤ E Z and ^^ ^ 0.
[0105] To perform the speech synthesis process, the impulse train generator 314 can generate an impulse train ^^^^^^ from a frame-wise fundamental frequency ^^^^^^^ output by the pitch tracker 310. In one illustrative example, the impulse train generator 314 can generate alias-free discrete time impulse trains using additive synthesis. For instance, as illustrated in equation (1) below, the impulse train generator 314 can use a low-passed sum of sinusoids to generate an impulse train: ^2^^^^^^^^^ ^ ^^^௧ ^^ ^^^^^^ ^^^2^^^^ ^^^^^^^^^^^^ ,^^^^^^ ^ ൌ 1
[0106] wherehold or linearinterpolation, ^^^^^^ ൌ ^^^^^ / ^^^^, and ^^^ is the sampling rate. In some cases, thecomputationally complexity of additive synthesis can be reduced with approximations. For example, the impulse train generator 314 or other component (e.g., a processor) of the voice decoding signal synthesis system 300 can round the fundamental periods to the nearest multiples of the sampling period. In such an example, the discrete impulse train is sparse. The impulse train generator 314 can then generate the impulse train sequentially (e.g., one pitch mark at a time).
[0107] The pitch tracker 310 can process the input ^^^^^^ for the time instant n to generate the frame-wise fundamental frequency ^^^^^^^ output, which is provided to and processed by the impulse train generator 314 of the voice decoding signal synthesis system 300. The random noise generator 316 of the voice decoding signal synthesis system 3400 can sample a noise signal ^^^^^^ from a Gaussian distribution.
[0108] The neural network filter estimator 302 can estimate impulse responsesℎ^ ^^^,^^^ and ℎ^ ^^^,^^^ for each frame, given the log-Mel spectrogram ^^^^^, ^^^ extractedfrom the input ^^^^^^ by the feature extraction engine 312 of the encoder. In some aspects,Qualcomm Docket No.2404227WOcomplex cepstrums (ℎ^^ and ℎ^^) can be used as the internal description of impulseresponses (ℎ^and ℎ^) for the neural network filter estimator 302. Complex cepstrums describe the magnitude response and the group delay of filters simultaneously. The group delay of filters affects the timbre of speech. In some cases, instead of using linear-phase or minimum-phase filters, the neural network filter estimator 302 can use mixed-phase filters, with phase characteristics learned from the dataset.
[0109] In some examples, the length of a complex cepstrum can be restricted, essentially restricting the levels of detail in the magnitude and phase response. Restricting the length of a complex cepstrum can be used to control the complexity of the filters. In some cases, the neural network filter estimator 302 can predicts low-frequency coefficients, in which the high-frequency cepstrum coefficients can be set to zero. In one illustrative example, two 10 millisecond (ms) long complex cepstrums are predicted in each frame. In some cases, the neural network filter estimator 302 can use a discrete Fourier transform (DFT) and an inverse-DFT (IDFT) to generate the impulse responses ℎ^and ℎ^. In some cases, the neural network filter estimator 302 can approximate an infinite impulse response (IIR) (ℎ^^^^,^^^ and ℎ^^^^,^^^) using finite impulse responses (FIRs). The DFT size can be set to at least a threshold size (e.g., N=1024) to avoid aliasing.
[0110] Using the impulse response ℎ^^^^,^^^, the harmonic LTV filter 318 can filter the impulse train ^^^^^^ from the impulse train generator 314 to generate a harmonic component ^^^^^^^. Using the impulse response ℎ^^^^,^^^, the noise LTV filter 320 can filter the noise signal ^^^^^^ to generate a noise component ^^^^^^^. The voice decoding signal synthesis system 300 can combine (e.g., by summing / adding or otherwise combining) the output of the harmonic LTV filter 318 (the harmonic component ^^^^^^^) and the output of the noise LTV filter 320 (the noise component ^^^^^^^) can be combined (e.g., summed or otherwise combined) to obtain the excitation signal ^^^^^^.
[0111] As noted previously, systems and techniques are described herein that can be used to provide a bit-efficient generative audio codec system including a machine learning (e.g., neural network)-based encoder and a machine learning (e.g., neural network)-based decoder. In some examples, the bit-efficient generative audio codec can be a hybrid DSP- ML codec, where the encoder and / or decoder are configured to perform combinations of non-ML-based DSP processing operations and ML-based processing operations toQualcomm Docket No.2404227WO encode and decode audio data, respectively.
[0112] FIG.4 is a block diagram illustrating an example of an audio codec system 400 that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer. In some aspects, the audio codec system 400 can also be referred to as a voice coding system (e.g., vocoder), a signal synthesis system, a voice coding signal synthesis system, etc. In one illustrative example, the audio codec system 400 can be a generative audio codec (e.g., a generative voice codec) including one or more feedback recurrent autoencoders (FRAEs) that may be used to generate encoded audio data, and one or more neural synthesizers that may be used to generate reconstructed audio data based on the encoded audio data.
[0113] For example, the audio codec system 400 (e.g., a generative voice codec) can include an encoder 405 (e.g., a transmitter) and a decoder 410 (e.g., a receiver). The encoder 405 can be used to generate encoded features that are transmitted to the decoder 410 over a channel 440. For example, the encoder 405 can receive audio 415 and generate encoded features corresponding to the audio 415. The encoded features can be generated using a feedback recurrent autoencoder (FRAE) 425 included in the encoder 405. In some examples, the encoder 405 can include one or more FRAEs, where each FRAE of the one or more FRAEs is configured to generate a corresponding one or more encoded features associated with the audio 415. In some aspects, the audio 415 may include and / or comprise a speech signal, a voice signal, etc.
[0114] The decoder 410 can receive encoded features (e.g., from the encoder 405, over the channel 440) and generate reconstructed audio 455 based at least in part on the encoded features received from the encoder 405. The reconstructed audio 455 may include and / or comprise a reconstructed speech signal, a reconstructed voice signal, etc. The reconstructed audio 455 can be generated using a neural network-based speech synthesizer 450 (e.g., also referred to as a neural speech synthesizer and / or a neural synthesizer) included in the decoder 410. The reconstructed audio 455 generated using the neural speech synthesizer 450 may also be referred to as synthesized speech. In some examples, the decoder 410 can include an FRAE decoder 445, which can be used to decode the encoded features received from the encoder 405. In some aspects, the FRAE decoder 445 can be the same as or similar to a decoder implemented by the FRAE autoencoder 425 included in the encoder 405. In some examples, the decoder 410 canQualcomm Docket No.2404227WO include one or more FRAE decoders 445, which can correspond to one or more FRAE autoencoders 425 included in the encoder 405. For example, the number of FRAE decoders 445 included in the decoder 410 can be equal to the number of FRAE autoencoders 425 included in the encoder 405.
[0115] In some examples, the audio 415 includes speech. The audio 415 may be an example of the speech signal 251 of FIG. 2B, the speech signal 271 of FIG. 2C, and / or the speech signal 303 of FIG.3, etc. The encoder 405 of the generative voice codec system 400 can be configured to generate one or more encoded representations of the input audio 415. For example, the one or more encoded representations can include spectral envelope information associated with the audio 415, and / or can include pitch information associated with the audio 415, etc. In some aspects, the encoder 405 can extract spectral envelope features from the audio 415 using spectral envelope feature extraction 420, and can encode the extracted spectral envelope features using the FRAE 425. The encoded spectral features z can be transmitted from the FRAE 425 to the decoder 410, using the channel 440. In some examples, the encoder 405 can extract pitch information from the audio 415 using pitch extraction engine 430, and can perform quantization 435 to generate quantized pitch information associated with the audio 415. The quantized pitch information can be transmitted from the pitch quantizer 435 to the decoder 410, using the channel 440.
[0116] In some aspects, the encoder 405 can be configured to extract spectral features (e.g., spectral envelope features) from the audio 415 using the spectral envelope feature extraction engine 420. In one illustrative example, the spectral envelope feature extraction engine 420 can perform a cepstrum computation to determine cepstrum information and / or one or more cepstral coefficients corresponding to the audio 415. In some cases, the spectral features extracted from the audio 415 using the spectral envelope feature extraction 420 can include mel-frequency cepstral coefficients (MFCC), such as MFCC- 24 features (e.g., 24-dimensional MFCC features). In some examples, the spectral envelope feature extraction 420 applies one or more mel-scaled filter(s) (e.g., a mel-scaled filterbank), a logarithmic compression, and / or a discrete cosine transform (DCT) to the audio 415 (e.g., and / or to a magnitude spectrum associated with the audio 415). The encoder 405 processes the extracted spectral features using a feedback recurrent autoencoder (FRAE) 425 to generate encoded features z. The FRAE 425 includes a decoder and an encoder. The FRAE 425 can implement feedback of state information hQualcomm Docket No.2404227WO between the decoder and the encoder included in the FRAE 425. For example, the state information h can be determined or obtained at the decoder of the FRAE 425, and feedback of the state information h can be performed for the decoder of the FRAE 425 (e.g., the state information h is fed back to the decoder of the FRAE 425) and for the encoder of the FRAE 425 (e.g., the state information h is fed back from the decoder of the FRAE 425 to the encoder of the FRAE 425).
[0117] In some examples, the spectral envelope feature extraction engine 420 can be configured to generate and / or determine (e.g., extract) one or more types of spectral envelope features. For example, the extracted spectral envelope features may include cepstrum information and / or cepstral coefficients. In some cases, cepstral liftering can be performed based on DCT truncation to exclude or remove pitch information and only capture spectral envelope information in the extracted cepstrum or cepstral coefficients. In some examples, the extracted spectral envelope features can include companded (e.g., log) filterbank energies associated with the input audio 415. For example, companded filterbank energies can be determined based on applying an inverse DCT (e.g., IDCT) to a liftered cepstrum. In some cases, the extracted spectral envelope features can include filterbank energies determined based on uncompanding (e.g., exp) the companded filterbank energies. In some examples, the extracted spectral envelope features can include a full resolution spectrum (e.g., DFT domain with companded (e.g., log, linear amplitude, etc.) information, etc.). The full resolution spectrum can be smoothed to the envelope of the spectrum, for example based on interpolating the companded or uncompanded filterbank energies. In some aspects, the extracted spectral envelope features can include one or more linear prediction (LP) coefficients, for example determined based on processing one or more speech frames of the input audio 415 using the Levinson-Durbin algorithm (e.g., autocorrelation technique), and / or using a covariance technique. In some examples, the extracted spectral envelope features can include one or more of line spectral frequency (LSF) information and / or line spectral pair (LSP) information.
[0118] In some aspects, the encoded features or encoded feature information transmitted from the generative voice codec system encoder 405 can comprise a latent representation z between the encoder of the FRAE 425 and the decoder of the FRAE 425. The encoder 405 can be configured to pass (e.g., transmit) the encoded features z through a channel 440 to the decoder 410.Qualcomm Docket No.2404227WO
[0119] The encoder 405 also processes the audio 415 using a pitch extraction engine 430 to extract pitch information from the audio 415. In some aspects, the audio 415 can be processed in parallel by the spectral envelope feature extraction engine 420 and the pitch extraction engine 430. In some cases, the encoder 405 can include a denoiser 418 that processes the input audio 415 and provides a de-noised audio to the spectral envelope feature extraction engine 420 and / or to the pitch extraction engine 430.
[0120] Based on the audio 415, the pitch extraction engine 430 can generate one or more types of pitch information. For example, the pitch extraction engine 430 can output pitch information in the frequency-domain (e.g., f0pitch information of the audio 415, in units of Hertz (Hz)) and / or can output pitch information in the time-domain (e.g., pitch lag in samples, with or without a fractional component, and / or pitch lag in milliseconds). In some cases, the pitch extraction engine 430 can be used to generate pitch estimation indicative of a pitch lag (or pitch delay) from the audio 415, and / or to identify a pitch correlation from the audio 415. In some aspects, the pitch information generated by the pitch extraction engine 430 can include a pitch lag and a pitch correlation.
[0121] The encoder 405 can be configured to processes the pitch information of the audio 415 (e.g., pitch, pitch lag, and / or pitch correlation, determined using the pitch extraction engine 430) using a quantizer 435 Q() to generate a quantized pitch signal. The encoder 405 passes the quantized pitch signal through the channel 440 to the decoder 410. In some examples, the quantizer 435 Q() may also be referred to as a pitch quantizer. In some aspects, the quantizer 435 Q() can perform vector quantization (VQ) to generate the quantized pitch signal for transmission to the decoder 410 over the channel 440. In some examples, the quantizer 435 Q() can perform single-stage VQ and / or can perform multi- stage VQ (e.g., MSVQ). In some cases, the quantizer 435 Q() can be implemented using one or more FRAEs. In some aspects, the quantizer 435 Q() can be configured to implement forward error correction (FEC) for the quantized pitch signal that is transmitted to the decoder 410 over the channel 440. For example, the quantizer 435 Q() can implement multiple description coding (MDC) for the transmission of the quantized pitch signal over the channel 440, can implement full-redundancy FEC for the transmission of the quantized pitch signal over the channel 440, etc.
[0122] The generative voice codec system decoder 410 can be configured to receive the encoded features z from the FRAE 425 included in the generative voice codec systemQualcomm Docket No.2404227WO encoder 405. For example, the decoder 410 can receive the encoded features z via the channel 440. The decoder 410 decodes the encoded features z using a FRAE 445 to generate decoded features. In some examples, the FRAE 445 of the decoder 410 includes only a decoder, without an encoder. In some examples, the FRAE 445 of the decoder 410 can include an encoder. In some examples, the decoded features generated by the FRAE 445 include mel-frequency cepstral coefficients (MFCC), such as MFCC-24 features (e.g., 24-dimensional MFCC features).
[0123] The decoded features determined using the FRAE decoder 445 can be provided to a neural speech synthesizer 450, which can be configured to generate a reconstructed audio 455 (e.g., synthesized speech) based at least in part on the decoded spectral envelope features from the FRAE decoder 445. In one illustrative example, the neural speech synthesizer 450 receives as input the decoded spectral envelope features (e.g., from the FRAE decoder 445) and the received pitch encoding information (e.g., the quantized pitch signal received by the decoder 410 over the channel 440 and from the pitch quantizer 435 of the encoder 405). In some aspects, the neural speech synthesizer 450 can generate the reconstructed audio 455 (e.g., synthesized speech) based on processing the decoded spectral envelope features along with a pitch signal (e.g., the quantized pitch signal or a reconstructed variant thereof).
[0124] For example, the decoder 410 can receive the quantized pitch signal (e.g., indicative of the pitch, the pitch lag, and / or the pitch correlation of the audio 415) from the encoder 405 via the channel 440. In some examples, the decoder 410 passes the quantized pitch signal to the neural speech synthesizer 450, and the neural speech synthesizer 450 processes the decoded spectral envelope features along with the quantized pitch signal to generate reconstructed audio 455 (e.g., synthesized speech).
[0125] In some cases, the decoder 410 can receive the quantized pitch signal (e.g., indicative of the pitch, the pitch lag, and / or the pitch correlation of the audio 415) from the encoder 405 via the channel 440, and can perform dequantization of the quantized pitch signal to obtain a reconstructed pitch signal. The reconstructed pitch signal can be passed to the neural speech synthesizer, and used in combination with the decoded spectral envelope features from the FRAE decoder 445 to generate the reconstructed audio 455 (e.g., synthesized speech). In some examples, the decoder 410 can include a reconstruction engine that is separate from the neural speech synthesizer 450. TheQualcomm Docket No.2404227WO reconstruction engine processes the quantized pitch signal to reconstruct the pitch signal (e.g., including the pitch, the pitch lag, and / or the pitch correlation) before the neural speech synthesizer 450 using the reconstructed pitch signal (e.g., the pitch, the pitch lag, and / or the pitch correlation) to generate the reconstructed audio 455.
[0126] In some examples, the output of the neural speech synthesizer 450 can be provided to one or more linear predictive coding (LPC) layers 452 included in the decoder 410. The LPC 452 can be used to perform linear prediction analysis and / or linear prediction synthesis, based on the output of the neural speech synthesizer 450. For example, the LPC 452 can perform linear prediction based on the output of the neural speech synthesizer 450, to generate the reconstructed audio 455 (e.g., synthesized speech). In some aspects, the decoder 410 does not include the LPC 452, and the neural speech synthesizer 450 can be trained and / or configured to generate the reconstructed audio 455 directly (e.g., the output of the neural speech synthesizer 450 can be the reconstructed audio 455).
[0127] In some aspects, the LPC layers 452 can be used to implement a linear prediction (LP) synthesis filter, based on: ^ ^^^^^^ ^ ^^^^^^^^ ^^^ ^ ^^^^^^
[0128] Here, ^^^^^^filter (e.g., the input signal to the LPC layers 452), ^^^^^^ represents the output signal (e.g., the output of the LPC layers 452, which can be the reconstructed audio 455), ^^^represents the linear prediction coefficients associated with implementing the LP synthesis filter, and ^^ is a value corresponding to the LP filter order. For example, the LP filter order p can have a value of 16 or 12, etc., among various other LP filter order values.
[0129] In some aspects, the LP synthesis filter associated with the LPC layers 452 can be implemented in the time domain or in the frequency domain. For example, in the time domain, the LP synthesis filter can be implemented based on a difference equation. In some cases, the LP synthesis filter can be implemented in the time domain by convolving the input signal x[n] with the LP filter impulse response or an approximation of the LP filter impulse response. For example, the LP filter can be associated with an infinite impulse response (IIR), which can be approximated using a finite segment of the IIR (e.g.,Qualcomm Docket No.2404227WO such as the first N samples of the IIR for a configured integer value of N, etc.). The LP synthesis filter associated with the LPC layers 452 can be implemented in the frequency domain based on multiplying the FFT of the input signal x[n] with the frequency response of the LP filter, and then determining the IFFT of the result to convert the output of the frequency domain LP synthesis filter into a time domain signal corresponding to the reconstructed audio 455.
[0130] In some examples, the LP coefficients ^^^can be estimated from the decoded spectral envelope features obtained using the FRAE decoder 445. In some cases, the spectral envelope features can be LP coefficients (e.g., determined by the spectral envelope feature extraction engine 420 based on using the Levinson-Durbin algorithm or other autocorrelation technique, or using a covariance technique, to process the input speech frames of the audio 415), and the decoded LP coefficients on the decoder side can be used for the LP filter implemented by the LPC layers 452.
[0131] In some examples, the extracted spectral envelope features may be LSFs or LSPs, and the decoded LSFs or LSPs obtained using the FRAE decoder 445 can be converted to the corresponding LP coefficients ^^^. In some cases, the LP coefficients ^^^of the LPC layers 452 can be estimated from the spectral envelope features using a neural network-based and / or a DSP-based technique. For example, the extracted spectral envelope features may comprise a cepstrum, and a DSP-based technique can be used to convert the decoded cepstrum into filterbank energies, and subsequently interpolate the filterbank energies to estimate an FFT square magnitude at all FFT bins (e.g., including bins beyond the centers of filterbank filters). After estimating the FFT square magnitude based on the interpolated filterbank energies, the DSP-based technique can include applying an IFFT to obtain an estimate of the autocorrelation sequence, and utilizing the Levinson-Durbin algorithm on the estimated autocorrelation sequence to obtain the LP coefficients ^^^.
[0132] In some examples, the spectral envelope feature extraction 420 can be performed based at least in part on spectral feature learning. For example, spectral feature learning can be performed by one or more machine learning models configured and / or trained to generate as output the one or more extracted spectral envelope features. In some cases, a cepstrum or cepstrum information may be an optional input to a machine learning spectral feature learning model used to implement the spectral envelope feature extractionQualcomm Docket No.2404227WO 420. In some cases, the spectral envelope feature extraction 420 can be implemented using one or more machine learning spectral feature learning models or engines, which can be jointly trained with the neural speech synthesizer 450, an NHV used to implement the neural speech synthesizer (e.g., such as the NHV-based neural speech synthesizer 650 of FIG. 6), an LPC network used to implement the neural speech synthesizer (e.g., such as the LPC network-based neural speech synthesizer 750 of FIG.7), etc.
[0133] In some cases, joint training of the spectral envelope feature extraction 420 and the neural speech synthesizer 450 can be performed based on a short-time Fourier transform (STFT) loss and / or STFT loss function. In some aspects, joint training of the spectral envelope feature extraction 420 and the neural speech synthesizer 450 can be performed based on an adversarial and / or generative adversarial network (GAN) loss. For example, joint training can be performed using a multi-resolution STFT loss, based on calculating STFT amplitude spectrograms from ground truth speech information and the synthesized speech information 455. The multi-resolution STFT loss can then be determined as a sum of mean absolute error and mean absolute error in the log domain. The STFT amplitude spectrograms can be calculated at different window lengths to capture representations of the error and / or loss at different time and / or frequency resolutions. In some aspects, joint training based on adversarial or GAN loss can be used to learn temporal fine structures in speech signals. An adversarial or GAN loss can be used to match the distribution of real speech and the synthesized speech 455. For example, the neural speech synthesizer 450 (and / or an NHV-based neural speech synthesizer) can be used as the generator network. A separate discriminator network can be used to train based on the adversarial loss, with both the NHV and the discriminator jointly trained using adversarial training techniques. In some examples, the discriminator network can be implemented as a classifier configured to classify whether an input audio represents real speech or synthesized speech. The generator network can attempt to fool the discriminator network to classify synthesized speech as real speech. Over the course of training, the generator network improves and begins producing (e.g., generating or synthesizing) speech that is very similar to the real speech. In some cases, the discriminator network can be implemented using WaveNet.
[0134] The neural speech synthesizer 450 can be implemented using one or more trained neural networks. For example, the one or more trained neural networks can be trained to perform speech synthesis and / or signal synthesis. In some examples, the neuralQualcomm Docket No.2404227WO speech synthesizer 450 can be implemented using or based on an LPCNet machine learning architecture, a WaveNet machine learning architecture, a WaveRNN machine learning architecture, etc. In one illustrative example, the neural speech synthesizer 450 can be provided as a neural homomorphic vocoder (NHV).
[0135] For example, FIG.6 illustrates an example of an audio codec system 600 (e.g., generative voice codec) that includes a decoder 610 configured to implement an NHV- based neural speech synthesizer. In some examples, an NHV is a type of neural vocoder that can synthesize speech with source-filter models controlled by one or more neural networks. An NHV-based neural vocoder may include one or more neural networks in a source-filter model that can synthesize speech based on filtering impulse trains and noise with linear time-varying (LTV) filters, with the one or more neural networks used to control the LTV filters by estimating complex cepstrums of time-varying impulse responses given acoustic features. Traditional or non-neural vocoders may operate based on decomposing speech into various parameters such as pitch, timbre, rhythm, etc., which can subsequently be manipulated and resynthesized to generate a desired output audio or voice signal. Neural vocoders (e.g., such as NHVs, etc.) can apply transformations and manipulations to speech signals directly within a learned feature space of the neural network. The learned feature space used by neural vocoders may capture more complex relationships and characteristics of speech than non-neural network-based signal processing and / or vocoder techniques. For example, neural vocoders can be trained to learn a mapping between raw speech waveforms and the spectral or cepstral representations of the speech. An NHV system can apply one or more homomorphic processing techniques within the learned space corresponding to the mapping.
[0136] As noted above, FIG. 6 is a block diagram illustrating an example of an audio codec system 600 (e.g., a generative voice codec, a voice coding signal synthesis system, etc.) that includes a decoder 610 that can be used to generate reconstructed audio (e.g., synthesized speech) using an NHV-based neural speech synthesizer 650, in accordance with some examples. In some aspects, the decoder 610 may also be referred to as an NHV synthesizer decoder. In some examples, the audio codec system 600 of FIG.6 can be the same as or similar to the audio codec system 400 of FIG. 4. In some cases, the channel 640 of FIG. 6 can be the same as or similar to the channel 440 of FIG. 4, and may be associated with an encoder that is the same as or similar to the encoder 405 of FIG.4, etc.Qualcomm Docket No.2404227WO
[0137] In some aspects, the NHV speech synthesizer 650 of FIG.6 can be the same as or similar to the NHV system 300 of FIG.3. For example, the neural filter estimator 652 can be the same as or similar to the NN filter estimator 302 of FIG.3. In some cases, the noise generator 670 can be the same as or similar to the random noise generator 316 of FIG.3, and the noise LTVF 675 can be the same as or similar to the noise LTV filter 320 of FIG. 3. In some examples, the pulse train generator 660 can be the same as or similar to the impulse train generator 314 of FIG.3, and the harmonic LTVF 665 can be the same as or similar to the harmonic LTV filter 318 of FIG.3.
[0138] In some aspects, the decoder 610 can include a neural speech synthesizer 650 that is the same as or similar to the neural speech synthesizer 450 included in the decoder 410 of FIG. 4. For example, the neural speech synthesizer 650 can receive a first input comprising decoded spectral envelope features (e.g., determined by an FRAE decoder 645 that is the same as or similar to the FRAE decoder 445 of FIG.4). The neural speech synthesizer 650 can receive a second input comprising a received pitch encoding (e.g., received pitch information), that is the same as or similar to the received pitch encoding provided to the neural speech synthesizer 450 of FIG.4.
[0139] In one illustrative example, the NHV synthesizer decoder 610 can include the FRAE decoder 645, which may be the same as or similar to the FRAE decoder 445 of FIG.4, and may include a pitch de-quantization engine 638 (e.g., de Q()). The pitch de- quantization engine 638 can be used to process a received pitch encoding obtained by the decoder 610 over the channel 640. For example, the received pitch encoding information can be quantized pitch information or a quantized pitch signal transmitted to the decoder 610 over the channel 640 by a corresponding encoder (e.g., an encoder associated with the decoder 610, which may be the same as or similar to the encoder 405 of FIG.4, etc.). In some aspects, the pitch dequantization engine 638 can generate reconstructed pitch information based on performing a codebook lookup to dequantize the received pitch encoding obtained from the channel 640 (e.g., to dequantize the quantized pitch encoding received by the decoder 610 over the channel 640).
[0140] The dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 638 can include pitch information in the frequency-domain (e.g., f0 pitch frequency information in Hz), pitch information in the time-domain (e.g., pitch lag or pitch delay information in samples, milliseconds, etc.), and / or pitch correlationQualcomm Docket No.2404227WO information. In some aspects, the dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 638 can include information indicative of a voiced or unvoiced (V / UV) classification.
[0141] In one illustrative example, the neural speech synthesizer 650 can be implemented as a neural homomorphic vocoder (NHV) speech synthesizer and / or an NHV-based neural speech synthesizer. For example, the neural speech synthesizer 650 can include a neural filter estimator 652 configured to generate respective filter specification or filter configuration information to parameterize one or more linear time- varying (LTV) filters of the NHV speech synthesizer 650. In some aspects, the neural filter estimator 652 can be trained to generate respective filter coefficients for a noise linear time-varying filter (LTVF) 675 and to generate respective filter coefficients for a harmonic LTVF 665.
[0142] The neural network model of the neural filter estimator 652 can include any neural network architecture that can be trained to model the filter coefficients for the noise LTVF 675 and the harmonic LTVF 665 (e.g., and / or that can be trained to model the filter coefficients for one or more additional filters implemented by the NHV speech synthesizer 650, in either the time-domain, the frequency-domain, or combinations thereof). Examples of neural network architectures that may be included in the neural network filter estimator 652 can include a generative neural network (e.g., a generative- adversarial network (GAN)), convolutional neural networks (CNN), an autoencoder, and / or other type(s) of neural networks, etc.
[0143] In some examples, the neural filter estimator 652 can be configured to receive a stream of decoded spectral envelope features from the FRAE decoder 645. Based on the decoded spectral envelope features, the neural filter estimator 652 can generate corresponding filter characterization parameters for each filter of one or more filters included in the NHV speech synthesizer 650. The respective filter characterization parameters can be output from the neural filter estimator 652 and used to parameterize and / or configure corresponding learned filters (e.g., learned linear filters, etc.) for each respective set of filter characterization parameters. For example, the neural filter estimator 652 can generate a first set of filters corresponding to the noise LTVF 675, can generate a second set of filter characterization parameters corresponding to the harmonic LTVF 665, etc. The one or more filters included in the NHV speech synthesizer 650 (e.g., theQualcomm Docket No.2404227WO noise LTVF 675, the harmonic LTVF 665, etc.) can be specified in various different forms. For example, the noise LTVF 675, the harmonic LTVF 665, and / or various other filters that may be included in the NHV speech synthesizer 650 can be specified or characterized (e.g., using the filter characterization parameters determined by the neural filter estimator 652) as cepstrums, as frequency responses, as time-domain impulse responses, as difference equation coefficients, etc.
[0144] In some aspects, the neural filter estimator 652 can be configured to receive as input the decoded spectral envelope features from the FRAE decoder 645, and may additionally receive as input at least a portion of the dequantized pitch information generated by the pitch dequantization engine 638. For example, the neural filter estimator 652 may receive as input the decoded spectral envelope features and pitch dequantization information (e.g., such as f0 pitch information in the frequency domain, pitch lag or pitch delay in the time domain (e.g., in units of samples or milliseconds, etc.), etc.). In some examples, the pitch dequantization information provided as input to the neural filter estimator 652 may include voiced / unvoiced (V / UV) classification information indicating whether the underlying audio represented in the encoded information received by the decoder 610 over the channel 640 corresponds to voiced or unvoiced sounds, speech, etc.
[0145] In examples where the neural filter estimator 652 receives spectral envelope features and dequantized pitch information as inputs, the neural filter estimator 652 can generate the corresponding filter characterization parameters for each NHV filter (e.g., noise LTVF 675, harmonic LTVF 665, etc.) based on the decoded spectral envelope features and the dequantized pitch information.
[0146] The noise LTVF 675 and the harmonic LTVF 665 can be implemented as time domain filters or frequency domain filters. In some examples, the noise LTVF 675 and the harmonic LTVF 665 can be implemented in the time domain or the frequency domain, independent of whether the neural filter estimator 652 is configured to generate the corresponding filter characterization parameters in the time domain or the frequency domain. For example, in some aspects, the neural filter estimator 652 can output impulse response-based filter characterization parameters, and the operation of noise LTVF 675 and / or harmonic LTVF 665 can be implemented in the time domain as convolutions with the impulse response. Using the same impulse response-based filter characterization parameters, the operation of noise LTVF 675 and / or harmonic LTVF 665 may beQualcomm Docket No.2404227WO implemented in the frequency domain based on converting the impulse response to frequency response (e.g., using an FFT transform) and multiplying with the FFT of the input signal, and subsequently converting back to the time domain using an IFFT transform.
[0147] In one illustrative example, the NHV speech synthesizer 650 can include a pulse train generator 660. The dequantized pitch information (e.g., generated using the pitch dequantization engine 638) can be provided to the pulse train generator 660. In some aspects, the pulse train generator 660 can generate a pulse train based at least in part on the pitch frequency (e.g., f0) and / or pitch lag or pitch delay information obtained from the pitch dequantization engine 638 of the decoder 610. In some examples, the pulse train generator 660 can be implemented as a cosine sum pulse generator, which can be configured to process the dequantized pitch information (e.g., obtained from the pitch dequantization engine 638) to generate a pulse train p[n]. The pulse train may also be referred to as an impulse train. For example, the pulse train generator 660 can be a differentiable cosine sum pulse generator, and / or can be a non-differentiable cosine sum pulse generator (e.g., among various other pulse generators).
[0148] The pulse train generated by the pulse train generator 660 can be processed by the harmonic LTVF 665, using the corresponding harmonic filter characterization parameters determined by the neural filter estimator 652 for the harmonic LTVF 665. In some aspects, the harmonic LTVF 665 can generate a harmonic output based on processing the pulse train from the pulse train generator 660.
[0149] In some examples, the pulse train generator 660 may receive an additional input from the pitch dequantization engine 638, indicative of a voice or unvoiced (e.g., V / UV) classification. For example, based on receiving an unvoiced (UV) indication or classification from the pitch dequantization engine 638, the pulse train generator 660 can be configured to generate a 0 output or a noise output that is provided to the harmonic LTVF 665 instead of the pulse train p[n] (e.g., instead of the pulse train p[n] provided from the pulse train generator 660 to the harmonic LTVF 665 in response to a voiced (V) indication or classification from the pitch dequantization engine 638).
[0150] The NHV speech synthesizer 650 can include a noise generator 670 that is configured to generate a noise signal (e.g., white noise, etc.) for processing by the noise LTVF 675. For example, the noise LTVF 675 can be parameterized based on theQualcomm Docket No.2404227WO respective noise filter characterization parameters generated by the neural filter estimator 652 for the noise LTVF 675, and can subsequently be used to process the noise signal generated by the noise generator 670. Based on processing the noise signal from the noise generator 670, the noise LTVF 675 can generate a noise-filtered output.
[0151] The NHV speech synthesizer 650 can include a combination function 680 to combine the harmonic output (e.g., from the harmonic LTVF 665) and the noise-filtered output (e.g., from the noise LTVF 675) to generate the reconstructed audio 655. In some examples, the reconstructed audio 655 includes synthesized speech. In some aspects, the reconstructed audio 655 can be the same as or similar to the reconstructed audio 455 of FIG. 4. In some examples, the combined noise LTVF 675 filter output and harmonic LTVF 665 filter output (e.g., from the combination operation 680) can comprise the reconstructed audio 655 (e.g., synthesized speech). In some aspects, the output of the combination operation 680 can be provided to an LPC 654, which may be the same as or similar to the LPC 452 of FIG.4. For example, the LPC 654 can perform linear predictive coding on the LTVF filter output from the NHV speech synthesizer 650, and may generate as output the reconstructed audio 655 (e.g., synthesized speech). In some examples, a linear time invariant (LTI) post-filter can be included between the output of the combination operation 680 and the input to the LPC 654.
[0152] In some examples, the combination function 680 can include an adder, a multiplier, a divider, weighted sum, a weighted product, a weighted ratio, an average, a weighted average, a weighted mean, a weighted median, a weighted mode, or a combination thereof. The reconstructed audio 655 may be an example of the reconstructed speech signal 255 of FIG. 2B, the reconstructed speech signal 275 of FIG. 2C, the reconstructed audio 455 of FIG.4, etc., or vice versa.
[0153] FIG.5A is a block diagram illustrating an example of a spectral envelope feature extraction branch 500 that can be included in a generative voice codec system. For example, the spectral envelope feature extraction branch 500 can be the same as or similar to the spectral envelope feature extraction 420 of the generative voice codec system 400 of FIG. 4, and / or can be used to implement the spectral envelope feature extraction 420 of the generative voice codec system 400 of FIG.4.
[0154] In some aspects, the spectral envelope feature extraction branch 500 can be used to generate (e.g., extract) one or more features 565 based on an audio signal 515. TheQualcomm Docket No.2404227WO features 565 can also be referred to as extracted features and / or spectral features (e.g., spectral envelope features), and can correspond to a spectral envelope of the audio signal 515. As noted previously, in some examples, the spectral features 565 can be spectral envelope features associated with the overall shape of the frequency spectrum of the audio signal 515 (e.g., where the spectral envelope connects the peaks of the individual frequency components in the frequency spectrum of an audio signal).
[0155] In some examples, the audio signal 515 of FIG.5A can be the same as or similar to the audio signal 415 of FIG. 4. The extracted spectral features 565 of FIG. 5A can be the same as or similar to a plurality of spectral envelope features generated by the spectral envelope feature extraction 420 and provided to the FRAE 425 included in the encoder 405 of FIG. 4. In one illustrative example, the audio signal 515 of FIG. 5A can be the same as or similar to the input provided to the spectral envelope feature extraction 420 of FIG.4.
[0156] For example, in cases where a denoiser (e.g., such as the denoiser 418 of FIG. 4) is not included before the spectral envelope feature extraction branch 500, the audio signal 515 can be the input speech or input audio provided to an encoder that implements the spectral envelope feature extraction branch 500 (e.g., such as the input audio 415 provided to the encoder 405 of FIG.4).
[0157] In cases where a denoiser (e.g., such as the denoiser 418 of FIG. 4) is included before the spectral envelope feature extraction branch 500, the audio signal 515 can be the same as or similar to the de-noised audio output provided by the denoiser (e.g., the same as or similar to the output of the denoiser 418 of FIG.4, etc.).
[0158] In some aspects, the spectral envelope feature extraction branch 500 can include a frame splitting operation 520 configured to split the input audio signal 520 into a plurality of audio frames (e.g., a plurality of frames of audio data). In some examples, the frame splitting operation 520 can generate a plurality of equal length audio frames from the input audio signal 520. In some cases, the plurality of frames generated by the frame splitting operation 520 may be non-overlapping (e.g., each respective portion of the input audio signal 520 is included in only one audio frame of the plurality of audio frames). In some examples, the plurality of frames generated by the frame splitting operation 520 may be overlapping. For example, portions of the input audio signal 520 can be included in multiple audio frames (e.g., an overlapping portion or common subset between adjacentQualcomm Docket No.2404227WO or consecutive audio frames of the plurality of audio frames generated using the frame splitting operation 520).
[0159] The spectral envelope feature extraction branch 500 can implement (e.g., apply) a filter bank 532 and an energy-per-band computation 534 to process the input audio signal 515 (e.g., in examples where the frame splitting operation 520 is not included or not utilized) or to process the plurality of frames split from the input audio signal 515 by the frame splitting operation 520 (e.g., in examples where the frame splitting operation 520 is included and utilized by the spectral envelope feature extraction branch 500).
[0160] In some aspects, the filter bank 532 and the energy-per-band computation 534 can be implemented as a filtering and energy computation branch 530 (e.g., a combined processing block 530), which can be the same as or similar to the example filtering and energy computation branch 570 of FIG. 5B. In some cases, the filter bank 532 can be applied prior to performing the energy-per-band computation 534. In some examples, the energy-per-band computation 534 can be performed prior to applying the filter bank 532 (e.g., such as in some cases where the filter bank 532 is computed in the frequency domain).
[0161] In one illustrative example, the filter bank 532 and the energy-per-band computation 534 of FIG. 5A can be implemented using the filtering and energy computation branch 570 of FIG.5B. In some aspects, the combined processing block 530 of FIG.5A can be implemented using the filtering and energy computation branch 570 of FIG.5B.
[0162] The filtering and energy computation branch 570 of FIG. 5B can optionally include a windowing operation 572, which may be the same as or similar to the frame splitting operation 520 of FIG.5A. In some aspects, the frame splitting operation 520 and the windowing operation 572 can be different. For example, the windowing operation 572 may be utilized in examples where the frame splitting operation 520 is not included in the decoder or is not utilized by the decoder, etc. In some cases, the windowing operation 572 is not included or not utilized in examples where the frame splitting operation 520 is utilized.
[0163] The filtering and energy computation branch 570 of FIG. 5B can apply a fast Fourier Transform (FFT) 574 to convert or transform between the time domain and the frequency domain, or vice versa. The FFT 574 can be of various sizes or dimensions, andQualcomm Docket No.2404227WO for example, may be configured as an FFT 574 with a size that is different from (e.g., does not match) the frame length (e.g., frame length associated with one or more of the frame splitting operation 520 of FIG.5A and / or the windowing function 572 of FIG.5B).
[0164] The output of the FFT 574 can be provided as input to a magnitude function 576. The magnitude function 576 can be configured to calculate a magnitude of the input received by the magnitude function 576. For example, the magnitude function 576 can be implemented as a squared modulus function (e.g., |·|2). In some examples, the magnitude function 576 being (or including) a squared modulus function (e.g., |·|2) can assist with bringing the input to the magnitude function 576 (e.g., the output from the FFT 574) into the real-valued domain. In some aspects, the magnitude function 576 can be implemented using other powers besides a power of two (e.g., other powers besides |·|2can be used for the magnitude function 576). For example, the magnitude function 576 can be implemented as |·| (e.g., a power of one). In some cases, the magnitude function 576 can be implemented as |·|afor any real-valued a > 0, including both fractional and integer values of a. In one illustrative example, the output of the magnitude function 576 can be referred to as a magnitude spectrum.
[0165] A filter bank 580 can be applied to the magnitude spectrum generated as output by the magnitude function 576. In one illustrative example, the filter bank 580 can be a filter bank of size N and can include a first filter 582-1 associated with a first frequency range, …, and an Nthfilter 582-N associated with an Nthfrequency range. In some aspects, the filter bank 580 of FIG.5B can be the same as or similar to the filter bank 532 of FIG. 5A (e.g., filter bank 532 and filter bank 580 can be filter banks of size N).
[0166] In some examples, the filter bank 580 is applied to the magnitude spectrum generated by the magnitude function 576 (e.g., the filter bank 580 is applied after the magnitude function 576). In some cases, the filter bank 580 can be applied before the magnitude function 576, for example by applying the filter bank 580 to the output of the FFT 574. In such examples, the magnitude function 576 may subsequently be applied to the output of the filter bank 580. In one illustrative example, the order of the magnitude function 576 and the filter bank 580 can be reversed (e.g., such that the filter bank 580 is applied before the magnitude function 576 is applied) based on the filter bank 580 having a frequency response that is real-valued in the FFT domain associated with the FFT 574. If the filter bank 580 frequency response is not real-valued in the FFT domain associatedQualcomm Docket No.2404227WO with the FFT 574, the magnitude function 576 can be applied prior to the filter bank 580, and the input to the filter bank 580 can be the magnitude spectrum generated as output by the magnitude function 576.
[0167] In some aspects, the filter bank 530 of FIG.5A and / or the filter bank 580 of FIG. 5B (e.g., which can be the same as one another) may be computed in the time domain or the frequency domain. In one illustrative example, the filter bank 580 is a filter bank of size N and includes N different filters (e.g., filter 1582-1, a filter 2, …, filter N 582-N). Each filter of the N different filters within the filter bank 580 can be associated with a respective filter center frequency and / or a respective filter structure information. In some aspects, each filter of the N different filters can be used to perform filtering for a different frequency band or frequency range that is associated with the particular filter 582-1, …, 582-N.
[0168] In one illustrative example, the filters 582-1, …, 582-N of the filter bank 580 can have respective filter center frequencies and / or filter structure that are based on or configured based on a perceptual structure. For example, in some aspects, the filter bank 580 can implement the plurality of filter 582-1, …, 582-N as a Mel filter bank perceptual structure, a Bark filter bank perceptual structure, etc. In some aspects, the filter bank 580 can be a Mel filter bank including N Mel-scaled filters 582-1, …, 582-N that can be used to produce a set of coefficients. In some cases, the N Mel-scaled filters 582-1, …, 582-N can include a bank of triangular bandpass filters spaced on a logarithmic scale (e.g., the Mel scale) that is configured to replicate the non-linear human perception of pitch.
[0169] In examples where the filter bank 580 is implemented as a Mel filter bank, the output of the filtering and energy computation branch 570 of FIG.5B (e.g., the output of the filtering and energy computation block 530 of the spectral envelope feature extraction branch 500 of FIG.5A) can be referred to as Mel frequency cepstral coefficients (MFCC) or a Mel cepstrum.
[0170] In examples where the filter bank 580 is implemented as a Bark filter bank, the output of the filtering and energy computation branch 570 of FIG.5B (e.g., the output of the filtering and energy computation block 530 of the spectral envelope feature extraction branch 500 of FIG.5A) can be referred to as Bark frequency cepstral coefficients (BFCC) or a Bark cepstrum.
[0171] In one illustrative example, the operations of the filtering and energyQualcomm Docket No.2404227WO computation branch 570 as illustrated in FIG. 5B may correspond to an example where the filter bank 580 is computed in the frequency domain. The output of the filter bank 580 can comprise a respective filter frequency response determined based on processing the magnitude spectrum generated by the magnitude function 576 with each respective one of the N filters 582-1, …, 582-N (e.g., the respective filter frequency response for filters 582-1, …, 582-N corresponding to the N sums 582-1, …, 584-N of FIG.5B).
[0172] To obtain a respective energy information 590-1, …, 590-N indicative of the energies per frequency band (e.g., per frequency band of the filter bank 580 filters 582-1, …, 582-N) the filter frequency response can be multiplied, elementwise for each frequency range or frequency bin corresponding to the sums 584-1, …, 584-N, with a signal FFT or |FFT|2, based on whether the filter bank 580 is applied before or after the magnitude function 576.
[0173] The output of the filtering and energy computation branch 570 of FIG. 5B can be the energies per band information indicative of the respective energy 590-1, …, 590- N within each of the N frequency bands associated with the filter bank 580. In one illustrative example, the energies per band information 590-1, …, 590-N can be the same as the output of the energy-per-band determination 534 of the spectral envelope feature extraction branch 500 of FIG.5A.
[0174] In some aspects, the energies per band information 590-1, …, 590-N can be provided to an optional companding function 540. The companding function 540 can be implemented by a compander, which may compress and / or expand a dynamic range of the input to the companding function 540. In some aspects, the companding function 540 can be implemented as Log(x) using any logarithm base (e.g., natural log ln(x), log10, etc.). In some cases, the companding function 540 can be implemented as (x)afor a ≤ 1. In some examples, the companding function 540 can be implemented as a logapproximation function (e.g., such as ^^൫ ^√^^െ 1൯ for an integer n > 0). In some cases,multiplicative and additive constants can be disregarded, and the companding function540 may be implemented as ^√^^, etc.
[0175] In some examples, the companding function 540 is not included or is not utilized, and the energies per band information 590-1, …, 590-N may be provided directly to a Discrete Cosine Transform (DCT) function 555.Qualcomm Docket No.2404227WO
[0176] The DCT function 555 can be of size N (e.g., the same size N as the filter bank 532 and / or filter bank 580). The input to the DCT function 555 can be the companded energies per band of the filter bank 532 (e.g., in examples where the companding function 540 is included and utilized in the spectral envelope feature extraction branch 500 of FIG. 5A), or can be the un-companded (e.g., non-companded) energies per band of the filter bank 532 (e.g., in examples where the companding function 540 is not included or not utilized).
[0177] Applying or computing the DCT function 555 of size N can be associated with computing N DCT coefficients (e.g., where the DCT coefficients are coefficients for a sum of a series of cosine functions that approximates and / or represents the input sequence or signal in the frequency domain). The domain to which the input audio signal 515 is transformed after the DCT function 555 is referred to as “cepstrum.”
[0178] In one illustrative example, the spectral envelope feature extraction branch 500 of FIG.5A can implement DCT truncation 560, to truncate the output of the DCT function 555 to include only the first M DCT coefficients, where M < N. For example, the output of the DCT function 555 may be the N DCT coefficients 1, 2, …, N. The DCT truncation 560 can truncate to the first M DCT coefficients 1, 2, …, M, based on removing or skipping the computation (e.g., at the time of computation of the DCT function 555) of the last (N-M) DCT coefficients M+1, M+2, …, N.
[0179] In some aspects, the DCT truncation 560 can be implemented based on configuring the DCT computation 555 to implement a DCT of size N but only compute the first M coefficients of the expected N coefficients (with M<N). In such examples, the last (N-M) DCT coefficients M+1, M+2, …, N are neither computed nor output by the DCT function 555.
[0180] In another example, the DCT truncation 560 can be implemented based on removing or deleting the last (N-M) DCT coefficients M+1, M+2, …, N, after the DCT function 555 generates as output the entire series of N DCT coefficients 1, 2, …, N.
[0181] In some aspects, the DCT truncation to the first M DCT coefficients where M<N (e.g., the DCT truncation 560 of FIG. 5A) can also be referred to as “cepstral liftering.” In one illustrative example, the DCT truncation 560 (e.g., cepstral liftering) implemented by the spectral envelope feature extraction branch 500 of FIG. 5A can be utilized to remove pitch information from the spectral envelope features 565.Qualcomm Docket No.2404227WO
[0182] For example, the pitch information can be removed from the extracted spectral envelope features 565 based on performing the DCT truncation (e.g., cepstral liftering) 560 to remove or skip the computation of a subset of higher order DCT coefficients. Based on performing the DCT truncation 560 to remove pitch information from the spectral features output included in the DCT function 555 cepstrum output, the extracted spectral envelope features 565 can be generated with a relatively high resolution (e.g., corresponding to the relatively large filter bank 532 size N), and then truncated to the smaller size M without a reduction in the resolution. Based on utilizing a relatively large filter bank 532 with a relatively large number of frequency bands (e.g., N, where N>M), and subsequently truncating the DCT of the relatively large filter bank output or cepstrum, the systems and techniques can be used to code only spectral envelope information in the cepstral feature output 565 of the spectral envelope feature extraction branch 500, without including or double coding pitch information in the extracted cepstral features 565.
[0183] In some aspects, the DCT truncation 560 (e.g., cepstral liftering) can be configured using a value of M (e.g., the truncation point corresponding to the first M DCT or cepstral coefficients out of the total N coefficients associated with the DCT 555 of size N>M) that is selected to be small enough to eliminate the pitch information from the cepstrum. For example, for a signal (e.g., audio signal 515) with a 16 kilohertz (kHz) sampling rate, in one example, the filter bank 532 and DCT function 555 can be configured with a size N = 80, and the DCT truncation (e.g. cepstral liftering) 560 can be configured to truncate at M = 24. The result of the DCT truncation 560 with M = 24 can be a set of features 565 comprising MFCC-24 Mel Frequency Cepstral Coefficients (MFCC) features.
[0184] Configuring the DCT truncation or cepstral liftering 560 with a value of M<N that is small enough to eliminate pitch information from the cepstrum can improve the bit efficiency associated with the spectral envelope feature extraction branch 500 of FIG.5A and / or the bit efficiency associated with encoded performed using the extracted spectral features 565 (e.g., encoding performed using the encoder 405 of the generative voice codec system 400 of FIG.4, etc.). In one illustrative example, the systems and techniques can be configured to perform the DCT truncation 560 to remove any pitch information from the cepstrum (e.g., from the spectral features or cepstrum 565) based on the encoder of the generative voice codec system (e.g., encoder 405 of generative voice codec system 400 of FIG.4) being configured to extract, code, and transmit pitch information separatelyQualcomm Docket No.2404227WO from the spectral envelope feature information. For example, the encoder 405 of FIG. 4 includes separate pitch extraction engines 430 and spectral envelope feature extraction engines 420, which are used to extract, code, and transmit respective pitch information and respective spectral envelope feature information that are separate from each other. In some aspects, performing the DCT truncation 560 (e.g., cepstral liftering) of the spectral envelope feature extraction branch 500 of FIG. 5A can be used to prevent double coding of pitch information in the extracted spectral features 565 that are coded and transmitted from the encoder 405 to the decoder 410 of FIG.4. Preventing the double coding of pitch information can improve the bit efficiency of the generative video codec system described herein, and can allow the encoder 405 to extract and send separate pitch and spectral envelope features only (e.g., the encoder 405 does not send explicit per-frame absolute phase information and / or pitch pulse locations, as in existing techniques for audio coding).
[0185] In some aspects, the size N of the filter bank 532 and DCT function 555 can be configured to be larger than the truncation size M, and to additionally be large enough to implement a relatively large (e.g., relatively high resolution) filter bank 532. At relatively small sizes of N, the filter bank 532 may comprise a coarse filter bank. Relatively small sizes of N (e.g., associated with coarse filter bank implementations) do not preserve spectral envelope information as well as larger values of N. Utilizing the relatively large value of N to implement a relatively high resolution, non-coarse filter bank 532 may improve the preservation of spectral envelope information in the output of the filter bank 532 and subsequent downstream operations of the spectral envelope feature extraction branch 500 of FIG. 5A. Subsequently applying the DCT truncation (e.g., cepstral liftering) 560 to the non-coarse filter bank 532 output can preserve the high resolution of the extracted spectral envelope features 565. For example, DCT function 555 can be associated with energy compaction properties that may be used to compress more spectral information in fewer DCT coefficients. Utilizing a relatively small value of N corresponding to a coarse filter bank implementation for the filter bank 532 (e.g., instead of performing truncation to decrease a large N to a relatively small M, using the DCT truncation or cepstral liftering 560 of FIG. 5A) can decrease the spectral envelope resolution from the beginning of the processing performed by the spectral envelope feature extraction branch 500 of FIG. 5A, making it difficult or impossible to recover spectral envelope resolution at later stages of the encoder or decoder (e.g., encoder 405Qualcomm Docket No.2404227WO or decoder 410 of the generative voice codec 400 of FIG.4, etc.).
[0186] FIG. 7 is a block diagram illustrating an example of a decoder 710 of a voice coding signal synthesis system that can be used to generate reconstructed audio (e.g., synthesized speech) using a neural speech synthesizer comprising a linear prediction coding (LPC) network. The decoder 710 can be included in an audio codec system (e.g., generative voice codec system) 700 that can be used to generate reconstructed audio (e.g., synthesized speech) 755.
[0187] In some aspects, the generative voice codec system 700 of FIG. 7 can be the same as or similar to the generative voice codec system 400 of FIG.4, 600 of FIG.6, etc. In some cases, the channel 740 of FIG.7 can be the same as or similar to the channel 440 of FIG.4, the channel 640 of FIG.6, etc.
[0188] The decoder 710 of FIG. 7 can be the same as or similar to the decoder 410 of FIG.4, the decoder 610 of FIG.6, etc. For example, the decoder 710 can include an FRAE decoder 745 the same as or similar to the FRAE decoder 445 of FIG. 4, 645 of FIG. 6, etc. The decoder 710 can include a neural speech synthesizer 750 that can be the same as or similar to the neural speech synthesizer 450 of FIG.4, 650 of FIG.6, etc. The decoder 710 can generate reconstructed audio 755 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 455 of FIG.4, 655 of FIG.6, etc.
[0189] In some examples, the neural speech synthesizer 750 can receive a first input comprising decoded spectral envelope features (e.g., determined by the FRAE decoder 745). The neural speech synthesizer 750 can receive a second input comprising a received pitch encoding (e.g., received pitch information), that is the same as or similar to the received pitch encoding provided to the neural speech synthesizer 450 of FIG.4, etc. For example, the decoder 710 may include a pitch de-quantization engine 738 (e.g., de Q()). The pitch de-quantization engine 738 can be used to process a received pitch encoding obtained by the decoder 710 over the channel 740. For example, the received pitch encoding information can be quantized pitch information or a quantized pitch signal transmitted to the decoder 710 over the channel 740 by a corresponding encoder (e.g., an encoder associated with the decoder 710, which may be the same as or similar to the encoder 405 of FIG. 4, etc.). In some aspects, the pitch dequantization engine 738 can generate reconstructed pitch information based on performing a codebook lookup to dequantize the received pitch encoding obtained from the channel 740 (e.g., to dequantizeQualcomm Docket No.2404227WO the quantized pitch encoding received by the decoder 710 over the channel 740). The dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 738 can include pitch information in the frequency-domain (e.g., f0 pitch frequency information in Hz), pitch information in the time-domain (e.g., pitch lag or pitch delay information in samples, milliseconds, etc.), and / or pitch correlation information. In some aspects, the dequantized (e.g., reconstructed) pitch information determined by the pitch dequantization engine 738 can include information indicative of a voiced or unvoiced (V / UV) classification.
[0190] In one illustrative example, the neural speech synthesizer 750 can be implemented as a linear predictive coding (LPC) network. For example, the LPC network-based neural speech synthesizer 750 can be used to implement a linear prediction(LP) synthesis filter, based on ^^^^^^ ൌ ∑^ ^ୀ^ ^^^^^^^^ െ ^^^ ^ ^^^^^^. Here, ^^^^^^ represents theinput signal to the LP filter, ^^^^^^ the output signal, ^^^represents the linearprediction coefficients associated the LP synthesis filter, and ^^ is a value corresponding to the LP filter order. For example, the LP filter order p can have a value of 16 or 12, etc., among various other LP filter order values.
[0191] For example, the LPC-based neural speech synthesizer 750 can include a frame rate network 752 configured to process inputs comprising the decoded features obtained using the FRAE decoder 745 and the decoded pitch information obtained using the pitch dequantization engine 738. An LPC estimation engine 762 can process the decoded features from the FRAE decoder 745 to determine one or more estimated LP coefficients. In some aspects, the LPC estimation engine 762 can generate estimated LP coefficients ^^^, based on an input comprising the decoded features determined using the FRAE decoder 745.
[0192] In some examples, the LP coefficients estimated using the LPC estimation engine 762 can be provided to an LP prediction engine 764, configured to generate asoutput a prediction p[n], where ^^^^^^ ൌ ∑^ ^ୀ^ ^^^^^^^^ െ ^^^. In some aspects, the LP prediction engine 764 can generate the LP prediction p[n] based on a first input comprising the LP coefficients ^^^estimated by the LPC estimation engine 762, and a feedback input s[n-1].
[0193] The estimated or predicted LP coefficients p[n] can be provided as input to a sample rate network 772 and a downstream combination or summation operation 780.Qualcomm Docket No.2404227WO The sample rate network 772 can receive additional inputs comprising the output of the frame rate network 752, the feedback s[n-1] from the previous step n-1, and an intermediate feedback value e[n-1] from the same previous step n-1.
[0194] The output of the sample rate network 772 can be the probability distribution P(e[n]), which is a probability distribution for e[n]. A sampling engine 774 can perform sampling from the probability distribution P(e[n]) to obtain a realization or representation of e[n].
[0195] The representation of e[n] determined by the sampling engine 774 can be combined with the LP prediction p[n] (e.g., determined by the LP prediction engine 764), using the combination or summation operation 780 to thereby generate as output the signal s[n]. In some aspects, the output signal s[n] can be the same as the reconstructed audio signal 755 of the LPC network-based neural speech synthesizer 750 and / or decoder 710.
[0196] The representation of e[n] determined by the sampling engine 774 can additionally be provided to a first feedback calculation 778, which generates as output e[n-1] provided as an additional input to the sample rate network 772.
[0197] The output signal s[n] of the combination or summation operation 780 can be output as the reconstructed audio 755 and may additionally be provided to a second feedback calculation 779, which generates the representation s[n-1] based on the input s[n]. The representation s[n-1] can be provided as a feedback input to the LP prediction engine 764 and to the sample rate network 772.
[0198] In some aspects, the systems and techniques can include a generative voice codec encoder that is configured to generate, encode, and transmit (e.g., to a corresponding generative voice codec decoder) one or more energy features and / or energy information based on the input audio signal that is also used to determine the encoded spectral envelope features and pitch information. For example, the encoder 405 of FIG.4 and / or various other encoders described herein can be configured to generate, encode, and transmit one or more energy features and / or energy information to a corresponding decoder (e.g., the decoder 410 of FIG. 4, and / or any one of the decoders of FIGS. 1-11 described herein, etc.).
[0199] In some examples, the encoder can include an energy feature computation toQualcomm Docket No.2404227WO compute one or more energy features or energy information indicative of an audio frame energy. For example, the frame energy can be determined by the encoder as a sum of squares of the signal amplitudes for the audio samples within a particular frame. In some cases, the encoder can compute, encode, and transmit the respective frame energy information or features for each frame of audio data included in a plurality of frames of audio data of the input audio signal to the encoder (e.g., such as input audio signal 415 to the encoder 405 of FIG.4, etc.).
[0200] In some aspects, the energy feature(s) determined by the encoder can replace the 0-th coordinate or the cepstral features in the cepstral feature vector generated, encoded, and transmitted from the encoder to the decoder of the generative voice codec implemented using the systems and techniques described herein. For example, the energy features determined by the decoder can replace the 0-th coordinate of the cepstral features in a cepstral feature vector generated by the encoder (e.g., generated based on the DCT truncation or cepstral liftering performed by the encoder, such as the DCT truncation or cepstral liftering 560 of FIG.5A, etc.). The cepstral feature vector with the inserted energy feature information at the 0-th coordinate of the cepstral features can subsequently be coded by the FRAE autoencoder of the encoder (e.g., such as the FRAE autoencoder 425 included in the encoder 405 of FIG.4, etc.).
[0201] In some examples, the energy feature(s) and / or energy information determined by the encoder can be coded and transmitted from the encoder to the decoder of the generative voice codec system separately from the spectral envelope feature information and / or the pitch information that are determined from the same input audio signal by the encoder. For example, the energy feature(s) may be encoded using an additional FRAE autoencoder implemented by the encoder (e.g., a separate instance of an FRAE autoencoder from the FRAE used for encoded the spectral envelope features). In some examples, the 0-th coordinate of the cepstral features may optionally be dropped or removed, with only the remaining cepstral features after removal or dropping of the 0-th coordinate of the cepstral features being coded and transmitted from the FRAE autoencoder of the encoder-side to the FRAE decoder of the decoder-side.
[0202] For example, FIG.8 is a block diagram illustrating an example of an audio codec system (e.g., generative voice codec system) 800 that encodes energy information 810 using a first FRAE autoencoder 815 and encodes spectral envelope features 820 using aQualcomm Docket No.2404227WO second FRAE autoencoder 825, in accordance with some examples.
[0203] In some aspects, the generative voice codec system 800 of FIG. 8 can be the same as or similar to the generative voice codec system 400 of FIG.4, 600 of FIG.6, 700 of FIG.7, etc. The encoder 805 of FIG.8 can be the same as or similar to the encoder 405 of FIG. 4, etc. For example, the encoder 805 can include spectral envelope feature extraction 820 the same as or similar to the spectral envelope feature extraction 420 of FIG.4, can include pitch extraction 830 the same as or similar to the pitch extraction 430 of FIG.4, can include a spectral feature FRAE autoencoder 825 the same as or similar to the FRAE autoencoder 425 of FIG. 4, can include a first pitch quantization engine 835 and a second pitch quantization engine 837 each of which may be the same as or similar to the pitch quantization engine 435 of FIG.4, etc.
[0204] In some cases, the channel 840 of FIG. 8 can be the same as or similar to the channel 440 of FIG. 4, the channel 640 of FIG. 6, the channel 740 of FIG. 7, etc. The decoder 810 of FIG. 8 can be the same as or similar to the decoder 410 of FIG. 4, the decoder 610 of FIG.6, the decoder 710 of FIG.7, etc. For example, the decoder 810 can include a spectral feature FRAE decoder 845 the same as or similar to the FRAE decoder 445 of FIG.4, 645 of FIG. 6, 745 of FIG.7, etc. The decoder 810 of FIG.8 may further include a second FRAE decoder 847 that is the same as or similar to the spectral feature FRAE decoder 825. The decoder 810 can include a neural speech synthesizer 850 that can be the same as or similar to the neural speech synthesizer 450 of FIG.4, 650 of FIG. 6, 750 of FIG. 7, etc. The decoder 810 can generate reconstructed audio 855 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 455 of FIG.4, 655 of FIG.6, 755 of FIG.7, etc.
[0205] In some examples, the energy features 810 and extracted pitch information 830 (e.g., each of pitch lag, V / UV classification, and / or pitch correlation, etc.) can be coded (e.g., encoded) by the encoder 805 using non-machine learning techniques, such as a single stage vector quantization (SSVQ) technique, a multi-stage vector quantization (MSVQ), other vector quantization technique, gain-shape technique, etc. The energy features 810 and extracted pitch information 830 (e.g., pitch lag, V / UV classification, and / or pitch correlation, etc.) can be coded with or without forward error correction (FEC) prior to transmission over the channel 840 between the encoder 805 and decoder 810 of the generative voice codec system 800.Qualcomm Docket No.2404227WO
[0206] In some examples, the energy features 810 and extracted pitch information 830 (e.g., each of pitch lag, V / UV classification, and / or pitch correlation, etc.) can be coded (e.g., encoded) by the encoder 805 using one or more machine learning coding (e.g., encoding) techniques.
[0207] For example, the energy features 810 and / or pitch information (e.g., pitch lag, V / UV classification, and / or pitch correlation) can be encoded using a corresponding one or more FRAE autoencoders. A separate FRAE autoencoder instance can be included in the encoder 805 for coding of the energy feature(s) and for coding the pitch information, where the separate FRAE autoencoder instances for coding energy features and / or pitch information is additionally separate from the spectral feature FRAE autoencoder 825 used by the encoder 805 to code the extracted spectral envelope features 820. The FRAE autoencoder instances implemented by the encoder 805 for coding of the energy features 810 and / or for coding the pitch information 830 (e.g., pitch lag, V / UV classification, and / or pitch correlation) can be implemented with FEC, and / or can be implemented without FEC.
[0208] In some aspects, energy features 810 can be coded using a corresponding separate FRAE autoencoder 815 that is separate from the FRAE autoencoder 825 used to code the extracted spectral envelope features 820. For example, the extracted spectral envelope features 820 can be coded and transmitted from the encoder 805 as the latent representation z generated from the parameter h of the FRAE autoencoder 825. The energy features 810 can be coded and transmitted from the encoder 805 as the latent representation z2generated from the parameter h2of the FRAE autoencoder 815.
[0209] In some examples, the pitch information determined by pitch extraction 830 can be coded using non-machine learning techniques (e.g., without a corresponding FRAE autoencoder instance for pitch). For example, the pitch extraction engine 830 of the encoder 805 can output pitch information (e.g., pitch lag, V / UV classification, and / or pitch correlation) to a first quantization engine 835 and a second quantization engine 837, which can be the same as or similar to one another. The first and second quantization engines 835 and 837, respectively, can be used to implement FEC techniques to improve the error concealment of the quantized pitch encoding information transmitted over the channel 840 from the encoder 805 to the decoder 810. In some cases, FEC techniques implemented by the encoder 805 for non-ML coded information (e.g., energy features 810Qualcomm Docket No.2404227WO and / or pitch information 830) can include coding with multiple codebooks using multiple description coding (MDC). In some examples, FEC techniques implemented by the encoder 805 for non-ML coded information can include full redundancy FEC.
[0210] For example, full redundancy FEC can be implemented for the extracted pitch information 830 (e.g., pitch lag, V / UV classification, pitch correlation, etc.) using the first and second pitch quantization engines 835 and 837, respectively, of the encoder 805. The encoder 805 can generate and transmit two identical copies of the encoding, where the two identical copies (e.g., from the first quantization engine 835 and the second quantization engine 837) are transmitted over the channel 840 and to the decoder 810 using two separate packets. In another example, the encoder 805 can generate two redundant encodings of the pitch information 830, for example by using two different codebooks with different sizes. In some aspects, the secondary encoding is of a lower bit rate than the primary encoding (e.g., the secondary encoding can be generated using a smaller codebook, and the primary encoding can be generated using a larger codebook). The primary and secondary encodings of the pitch information 830 can be transmitted in different (e.g., separate) packets from the encoder 805 to the decoder 810, over the channel 840.
[0211] In some cases, the pitch information 830 can be coded and transmitted from the encoder 805 to the decoder 810 using ML-based coding techniques (e.g., using a dedicated FRAE autoencoder instance for the pitch information 830, such as the FRAE autoencoder 815), and the energy feature(s) 810 can be coded and transmitted from the encoder 805 to the decoder 810 using non-ML based coding techniques (e.g., such as the quantization engines 835, 837 and one or more FEC and / or MDC techniques, etc.).
[0212] Each FRAE autoencoder 815, 825 included in the encoder 805 and used to generate a corresponding encoded representation or latent (e.g., the spectral feature latentz and the energy feature latent z2) can be associated with a corresponding FRAE decoderinstance implemented in the decoder 810. For example, the FRAE autoencoder 815 of the encoder 805 can be used to generate and transmit the latent representation z2of the energy features 810 to a corresponding FRAE decoder 847 included in the decoder 810. The FRAE autoencoder 825 of the encoder 805 can be used to generate and transmit the latent representation z of the spectral envelope features 820 to a corresponding FRAE decoder 845 included in the decoder 810. The decoded features generated by the FRAE decoderQualcomm Docket No.2404227WO instances 845 and 847 of the decoder 810 can be provided as separate inputs to the neural speech synthesizer 850. For example, the neural speech synthesizer 850 can be configured to generate the reconstructed audio 855 based on decoded (e.g., reconstructed) energy features determined by the energy FRAE decoder instance 847, decoded (e.g., reconstructed) spectral features determined the spectral FRAE decoder instance 845, and pitch information obtained from the pitch extraction engine 830 of the encoder 805.
[0213] In one illustrative example, the systems and techniques can be used to implement an audio codec system (e.g., generative voice codec system) that can be used to perform neural speech synthesis based on joint coding of spectral envelope features and pitch information. In some aspects, the systems and techniques can be used to implement an audio codec system (e.g., generative voice codec system) that can be used to perform neural speech synthesis based on joint coding of spectral envelope features, pitch information, and energy features.
[0214] For example, an encoder of the generative voice codec system can be configured to combine and / or concatenate the different types of features extracted from an input audio signal, and perform coding of the combined features (e.g., energy features, spectral features, pitch features) using a one or multiple FRAE autoencoder instances, with or without utilizing multiple description coding (MDC).
[0215] FIG.9 is a block diagram illustrating an example of an audio codec system (e.g., generative voice codec system) 900 that can be used to perform neural speech synthesis based on multiple combined vectors corresponding to various combinations of encoded features including energy features 910, spectral features 920, and pitch information 930. For example, the generative voice codec system 900 can be used to perform neural speech synthesis based on performing groupwise joint coding to generate encoded vectors of various combinations of spectral energy features, energy features, and pitch information, in accordance with some examples. For example, a concatenation engine 903 can generate as output a combined feature vector that includes the energy features 910, the spectral envelope features 920, and the pitch information 930. The encoder 905 of FIG. 9 can further include a sub-vector splitting operation 907 configured to split the combined feature vector from the concatenation engine 903 into various different subsets, combinations or sub-combinations, etc., of the input feature types (e.g., energy features 910, spectral envelope features 920, and pitch information 930). Each sub-vector splitQualcomm Docket No.2404227WO from the combined feature vector by the sub-vector splitting operation 907 can be provided to a corresponding FRAE autoencoder instance 915, 925, …, etc. of the encoder 905, which may generate a corresponding latent representation z, z2, …, etc., for transmission to a corresponding FRAE decoder instance 947, 945, …, etc. of the decoder 910.
[0216] In some aspects, the generative voice codec system 900 of FIG. 9 can be the same as or similar to the generative voice codec system 400 of FIG.4, 600 of FIG.6, 700 of FIG.7, 800 of FIG.8, , etc. The encoder 905 of FIG.9 can be the same as or similar to the encoder 405 of FIG. 4, 805 of FIG.8, etc. For example, the encoder 905 can include spectral envelope feature extraction 920 the same as or similar to the spectral envelope feature extraction 420 of FIG.4, 820 of FIG.8, etc. Energy determination 910 of FIG.9 can be the same as or similar to the energy determination 534 of FIG. 5A and / or can be associated with the energy-per-band information 590-1, …, 590-N of FIG.5B, etc. Pitch extraction 930 of FIG.9 can be the same as or similar to the pitch extraction 430 of FIG. 4, 830 of FIG. 8, etc. A pitch quantization engine 935 of FIG. 9 can be the same as or similar to the pitch quantization engine 435 of FIG.4, 835 and / or 837 of FIG.8, etc. The encoder 905 of FIG. 9 can include a first FRAE autoencoder 925 and a second FRAE autoencoder 915, which may be the same as or similar to one another. In some aspects, the FRAE autoencoder 925 and / or the FRAE autoencoder 915 can be the same as or similar to one or more of the FRAE autoencoder 425 of FIG.4, 825 of FIG.8 and / or 815 of FIG.8, etc. In some cases, the channel 940 of FIG.9 can be the same as or similar to the channel 440 of FIG. 4, the channel 640 of FIG. 6, the channel 740 of FIG. 7, the channel 840 of FIG.8, etc.
[0217] The decoder 910 of FIG. 9 can be the same as or similar to the decoder 410 of FIG.4, the decoder 610 of FIG.6, the decoder 710 of FIG.7, the decoder 810 of FIG.8, etc. For example, the decoder 910 can include a first FRAE decoder 945 and a second FRAE decoder 947, which may be the same as or similar to one another. In some aspects, the FRAE decoder 945 and / or the FRAE decoder 947 of FIG. 9 can be the same as or similar to one or more of the FRAE decoder 445 of FIG.4, 645 of FIG.6, 745 of FIG.7, 845 of FIG. 8 and / or 847 of FIG. 8, etc. The decoder 910 can include a pitch dequantization engine 938 that is the same as or similar to the pitch dequantization engine 638 of FIG. 6, 738 of FIG. 7, , etc. The decoder 910 can include a neural speech synthesizer 950 that can be the same as or similar to the neural speech synthesizer 450 ofQualcomm Docket No.2404227WO FIG.4, 650 of FIG.6, 750 of FIG.7, 850 of FIG.8, 950 of FIG.9, etc. The decoder 910 can generate reconstructed audio 955 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 455 of FIG.4, 655 of FIG.6, 755 of FIG. 7, 855 of FIG.8, etc.
[0218] The decoder 910 of FIG. 9 can include a first de-concatenation engine 949-1 (e.g., associated with the first FRAE decoder 945) and a second de-concatenation engine 949-2 (e.g., associated with the second FRAE decoder 947), configured to reverse the concatenation operations performed by the encoder 905 concatenation engine 903 and / or the sub-vector splitting operations performed by the encoder 905 sub-vector splitting engine 907.
[0219] As noted previously, the generative voice codec system 900 of FIG. 9 can be used to implement groupwise joint coding to generate various sub-vector combinations of features corresponding to an input audio data (e.g., a speech signal for encoding, etc.). For example, groupwise joint coding can be performed to generate various sub-vector combinations of features corresponding to an input audio data, where the features include one or more energy features 910 (e.g., which can be the same as or similar to the energy features 810 of FIG. 8), one or more spectral envelope features 920 (e.g., which can be the same as or similar to the spectral envelope features 820 of FIG.8), and pitch features 930 (e.g., which can be the same as or similar to the pitch features 830 of FIG.8). In some examples, the pitch features 930 may also be referred to as pitch information, and vice versa. The pitch features 930 can include pitch information in the frequency domain (e.g., f0pitch frequency), pitch information in the time domain, pitch lag, voiced / unvoiced (V / UV) classification information, pitch correlation information, etc.
[0220] In one illustrative example, groupwise joint coding can be performed by the encoder 905, to generate multiple sub-vectors of features selected from the set of energy features 910, spectral envelope features 920, and pitch extraction information (e.g., pitch features) 930. For example, the encoder 905 can generate and / or obtain the energy features 910, can generate and / or obtain the spectral envelope features 920, and can generate and / or obtain the pitch features 930. The encoder 905 can include a concatenation engine 903 configured to generate a single combined vector of features from the three separate inputs of the energy features 910, the spectral envelope features 920, and the pitch features 930.Qualcomm Docket No.2404227WO
[0221] In some aspects, the output of the concatenation engine 903 can correspond to groupwise joint coding with the number of groups equal to one (e.g., where the one group corresponds to the single, combined vector of energy features 910, spectral envelope features 920, and pitch features 930 generated by using the concatenation engine 903 to concatenate the three individual features inputs together in a single vector of features). In some cases, the single combined vector of features generated by the concatenation engine 903 can include spectral envelope features 920 (e.g., cepstrum, with potential omission of the cepstrum 0-th bin or replacement of the cepstrum 0-th bin with one or more energy features if present), energy features and / or energy correction features 910 (e.g., in some cases, in addition to the cepstrum 0-th bin or as a replacement of the cepstrum 0-th bin), pitch information 930 (e.g., f0 pitch information, pitch lag), V / UV classification, pitch correlation (if present), etc. In some cases, the encoder 905 can use the concatenation engine 903 to generate a combined feature vector with one or more additional features extracted from the input audio (e.g., in addition to the energy features 910, spectral envelope features 920, and pitch information 930).
[0222] The combined vector of features generated as output by the concatenation engine 903 can be used to generate the multiple sub-vectors of audio features that can be used by the encoder 905 (and / or decoder 910) to implement the groupwise joint coding. For example, the concatenation engine 903 can be used to generate a combined vector of features, which can be provided a sub-vector splitting engine 907 (e.g., also referred to as a sub-vector splitting operation).
[0223] The sub-vector splitting engine 907 can generate sub-vectors that include various combinations of one or more portions of the energy features 910, the spectral envelope features 920, and / or the pitch features 930. For example, sub-vectors can be generated to correspond to the discrete feature types obtained at the input of the encoder 905 (e.g., a first sub-vector including the energy features 910, a second sub-vector including the spectral envelope features 920, a third sub-vector including the pitch features 930). In some examples, the audio feature sub-vectors can correspond to multiple types of features (e.g., different feature types can be included in the same sub-vector). For example, a first audio feature sub-vector can include energy features 910 and spectral envelope features 920, and a second audio feature sub-vector can include pitch features 930, etc.Qualcomm Docket No.2404227WO
[0224] In some aspects, the input feature sets or types (e.g., energy features 910, spectral envelope features 920, pitch features 930) can be split across and included in multiple different sub-vectors. For example, a first sub-vector can include a first portion of the energy features 910, a first portion of the spectral envelope features 920, and a first portion of the pitch features 930. A second sub-vector can include a second portion of the energy features 910, a second portion of the spectral envelope features 920, and a second portion of the pitch features 930. A third sub-vector can include a third portion of the energy features 910, a third portion of the spectral envelope features 920, and a third portion of the pitch features 930.
[0225] In some examples, one or more sub-vectors can include multiple types of audio features and one or more sub-vectors can include a single type of audio features. For example, a first sub-vector can include a first portion of the energy features 910 and a first portion of the spectral envelope features 920. A second sub-vector can include a second portion of the energy features 910 and a second portion of the spectral envelope features 920. A third sub-vector can include the pitch features 930.
[0226] The plurality of audio feature sub-vectors generated as output by the sub-vector splitting engine 907 can include all of the features (or representations thereof) of the input features comprising the set of energy features 910, the spectral envelope features 920, and the pitch features 930. In some cases, the plurality of audio feature sub-vectors can include duplicated information, where a particular feature or subset of features are included in multiple different sub-vectors. In some examples, the plurality of audio feature sub- vectors can be non-overlapping, where each particular feature in the input set of energy features 910, spectral envelope features 920, and pitch features 930 is included in only one sub-vector of the plurality of sub-vectors generated as output by the sub-vector splitting engine 907.
[0227] In some aspects, various combinations of the encoder 905 features (e.g., energy features 910, spectral envelope features 920, pitch features 930, etc.) can be grouped and / or maintained as separate feature sets within the plurality of sub-vectors generated using the sub-vector splitting engine 917. In some examples, each group or feature set corresponding to a different sub-vector of audio features can be coded with a corresponding separate FRAE autoencoder instance of the encoder 905 and decoded with a corresponding separate FRAE decoder instance of the decoder 910. The FRAEQualcomm Docket No.2404227WO autoencoder-based coding performed by the FRAE autoencoder instances 915, 925 of the encoder 905 can be implemented with or without multiple description coding (MDC), and / or can be implemented with non-ML coding techniques with or without forward error correction (FEC) and / or MDC.
[0228] In one illustrative example, the FRAE 925 can be used to code and / or quantize a feature group (e.g. sub-vector of features obtained from the sub-vector splitting operation 907) including pitch and pitch correlation information 930. The FRAE 915 can be used to code and / or quantize a second feature group (e.g., a second sub-vector of features obtained from the sub-vector splitting operation 907) comprising spectral envelope features 920. Various other sub-vector combinations of the input set of audio features (e.g., energy features 910, spectral envelope features 920, pitch features 930, etc.) can also be provided as the respective inputs to the first and second FRAE autoencoder instances 915, 925 of the encoder 905.
[0229] In some aspects, the encoder 905 can code (e.g., encode) a first portion of the plurality of audio feature sub-vectors using a respective FRAE autoencoder instance (e.g., 915, 925, etc.) for each individual sub-vector. The encoder 905 can code (e.g., encode) a second portion or a remaining portion of the plurality of audio feature sub-vectors using classical or non-ML-based encoding techniques. For example, the encoder 905 can include one or more quantization engines 935 configured to quantize sub-vectors of audio features that are not provided to a respective FRAE autoencoder instance. In one illustrative example, a quantization engine 935 can be used to code and quantize a third feature group (e.g., a third sub-vector of features obtained from the sub-vector splitting operation 907) including the energy features 910 using classical non-ML coding techniques.
[0230] In another example, the sub-vector splitting operation 907 can be applied to generate a first feature group or sub-vector comprising a first half of the spectral envelope features 920 and the energy features 920 (e.g., with the first sub-vector of spectral and energy features coded by the classical quantizer 935), and to generate a second feature group or sub-vector comprising the remaining half of the spectral envelope features 920 and the pitch information 930 (e.g., with the second sub-vector of spectral and pitch features coded by one of the FRAE autoencoder instances 915 or 925 of the encoder 905).
[0231] In some aspects, the systems and techniques can implement joint coding ofQualcomm Docket No.2404227WO features corresponding to multiple frames of audio data (e.g., multiple frames of a plurality of frames corresponding to the input audio signal to the encoder of the generative voice codec system, etc.). The joint coding can be performed with classical or non-ML- based quantization (e.g., using a quantization engine to code extracted features at the encoder) and / or can be performed with ML-based quantization (e.g., using an FRAE autoencoder to code extracted features at the encoder). Joint coding of features corresponding to multiple frames of audio data can correspond to increased bit efficiency of the generative voice codec system.
[0232] Each sub-vector of combined audio features that is coded using an FRAE autoencoder instance (e.g., FRAE 915, FRAE 925, etc.) of the encoder 905 can be used to generate an encoded representation (e.g., a latent of the FRAE autoencoder instance) that can be provided to a corresponding FRAE decoder instance included in the decoder 910 of the generative voice codec system 900.
[0233] For example, a first sub-vector of audio features can be processed by the corresponding first FRAE instance 925, and used to generate a corresponding encoded representation of the first sub-vector comprising the latent z. A second sub-vector of audio features can be processed by the corresponding second FRAE instance 915, and used to generate a corresponding encoded representation of the second sub-vector comprising the latent z2. A third sub-vector of audio features can be processed by the quantization engine 935, which generates a quantized bitstream corresponding to an encoded representation of the third sub-vector.
[0234] Each encoded representation of sub-vectors of audio features can be transmitted over the channel 940 from the encoder 905 to the decoder 910. For example, the latent representations z and z2 corresponding to the first and second sub-vectors of audio features, respectively, and the quantized bitstream generated by the quantization engine 935 for the third sub-vector of audio features can each be transmitted over the channel 940 from the encoder 905 to the decoder 910. The decoder 910 can use the encoded representations of each sub-vector of audio features to generate the reconstructed audio 955 (e.g., synthesized speech).
[0235] The first sub-vector of audio features is encoded by the first FRAE autoencoder instance 925 at the encoder 905, and can be provided to a corresponding first FRAE decoder instance 945 at the decoder 910. The first FRAE decoder 945 can generate aQualcomm Docket No.2404227WO reconstructed representation of the first sub-vector of audio features, based on decoding the encoded representation of the first sub-vector of audio features received over the channel 940 from the first FRAE autoencoder 925. The second FRAE decoder 947 can generate a reconstructed representation of the second sub-vector of audio features, based on decoding the encoded representation of the second sub-vector of audio features received over the channel 940 from the second FRAE autoencoder 915.
[0236] In some aspects, the reconstructed representations of sub-vectors of audio features generated by the FRAE decoders 945 and / or 947 may be provided directly as input to the neural speech synthesizer 950. The decoded or reconstructed representations of audio feature sub-vectors generated by the FRAE decoders 945, 947 can represent the same combinations of the different types of audio features as was used by the encoder 905. In some examples, the reconstructed representations of sub-vectors of audio features generated as output by the FRAE decoders 945, 947 can be separated into the individual or discrete audio feature types using a corresponding de-concatenation engine 949-1, 949- 2 associated with each FRAE decoder 945, 947 (respectively).
[0237] For example, the FRAE decoder 945 can be associated with a corresponding de- concatenation engine 949-1. If the reconstructed sub-vector of audio features output by the FRAE decoder 945 includes a portion of energy features 910 and a portion of spectral envelope features 920, the de-concatenation engine 949-1 can separate the reconstructed sub-vector features and provide the neural speech synthesizer 950 with the separated portion of energy features 910 and the separated portion of spectral envelope features 920 obtained from the FRAE decoder 945. The FRAE decoder 947 can be associated with a corresponding de-concatenation engine 949-2, which can be used to perform the same or similar operations to split or separate the different types of audio features represented in the reconstructed sub-vector of features output by the FRAE decoder 947, prior to the input to the neural speech synthesizer 950.
[0238] In some cases, the decoded or reconstructed representation of each FRAE- encoded sub-vector of combined audio features can be provided as input to the neural network-based synthesizer 950, as a single sub-vector (e.g., without processing by a corresponding de-concatenation engine 949-1, 949-2) or can be provided to the neural network-based synthesizer 950 as separated or de-concatenated features split from the sub-vector reconstruction using the corresponding de-concatenation engine 949-1, 949-Qualcomm Docket No.2404227WO 2, etc. Sub-vectors of combined audio features that are coded using non-machine learning coding techniques at the encoder 905 (e.g., coded using the quantization engine 935) can be dequantized (e.g., using a corresponding dequantization engine 938 of the decoder 910) and provided as inputs to the neural network-based synthesizer 950, and / or can be provided directly to the neural network-based synthesizer without being dequantized.
[0239] FIG. 10 is a block diagram illustrating an example of an audio codec system (e.g., generative voice codec system) 1000 that can be used to perform neural speech synthesis based on using a first denoiser 1018 associated with spectral envelope feature extraction 1020 and a second denoiser 1019 associated with pitch extraction 1030, in accordance with some examples. In some aspects, the generative voice codec system 1000 of FIG.10 can be the same as or similar to the generative voice codec system 400 of FIG. 4, 600 of FIG.6, 700 of FIG.7, 800 of FIG.8, 900 of FIG.9, etc.
[0240] The encoder 1005 of FIG. 10 can be the same as or similar to the encoder 405 of FIG.4, 805 of FIG.8, 905 of FIG.9, etc. For example, the encoder 1005 can include a first denoiser (e.g., spectral branch denoiser) 1018 that is the same as or similar to the denoiser 418 of FIG.4, etc. The encoder 1005 can include a second denoiser (e.g., a pitch denoiser) 1019 that is the same as or similar to the denoiser 418 of FIG. 4, and / or the spectral branch denoiser 1018, etc. The encoder 1005 can include spectral envelope feature extraction 1020 the same as or similar to the spectral envelope feature extraction 420 of FIG. 4, 820 of FIG. 8, 920 of FIG. 9, etc. The encoder 1005 can include pitch extraction 1030 the same as or similar to the pitch extraction 430 of FIG. 4, 830 of FIG. 8, 930 of FIG.9, etc. A pitch quantization engine 1035 of FIG. 10 can be the same as or similar to the pitch quantization engine 435 of FIG.4, 835 and / or 837 of FIG.8, etc. The encoder 1005 can include an FRAE autoencoder 1025 the same as or similar to the FRAE autoencoder 425 of FIG.4, 825 and / or 815 of FIG.8, 925 of FIG.9, etc.
[0241] In some cases, the channel 1040 of FIG.10 can be the same as or similar to the channel 440 of FIG.4, the channel 640 of FIG.6, the channel 740 of FIG.7, the channel 840 of FIG. 8, the channel 940 of FIG. 9, etc. The decoder 1010 of FIG. 10 can be the same as or similar to the decoder 410 of FIG. 4, the decoder 610 of FIG. 6, the decoder 710 of FIG. 7, the decoder 810 of FIG. 8, the decoder 910 of FIG. 9, etc. For example, the decoder 1010 can include a spectral feature FRAE decoder 1045 the same as or similar to the FRAE decoder 445 of FIG. 4, 645 of FIG. 6, 745 of FIG. 7, 845 of FIG. 8 and / orQualcomm Docket No.2404227WO 847 of FIG. 8, 945 of FIG. 9, etc. The decoder 1010 can include a neural speech synthesizer 1050 that can be the same as or similar to the neural speech synthesizer 450 of FIG. 4, 650 of FIG. 6, 750 of FIG. 7, 850 of FIG. 8, 950 of FIG. 9, etc. The decoder 1010 can generate reconstructed audio 1055 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 455 of FIG.4, 655 of FIG.6, 755 of FIG.7, 855 of FIG.8, 955 of FIG.9, etc. In some aspects, the decoder 1010 of FIG.10 can include an LPC 1052 between the output of the neural speech synthesizer 1050 and the reconstructed audio 1055 output from the decoder 1010. In some examples, the LPC 1052 of FIG.10 can be the same as or similar to the LPC 452 of FIG.4, 654 of FIG.6, etc.
[0242] In some cases, the generative voice codec system 1000 of FIG. 10 can be the same as the generative voice codec system 400 of FIG. 4, with the addition of the additional, second denoiser 1019 (e.g., pitch denoiser, denoiser for pitch, etc.) between the input audio signal 1015 and the pitch extraction engine 1030.
[0243] For example, in some aspects, separate denoisers 1018 and 1019 can be included in and utilized by the encoder 1005 to perform denoising of the input audio signal 1015 prior to processing for the spectral envelope feature extraction 1020 and the pitch extraction 1030. In some cases, the pitch denoiser 1019 can be configured to perform more aggressive denoising than the denoiser 1018 associated with the spectral envelope feature extraction 1020 and / or the spectral processing branch of the encoder 1005. The more aggressive denoising for pitch, applied by the pitch denoiser 1019, can improve the robustness of pitch estimation performed by the pitch extraction engine 1030 based on the denoised audio signal output from the pitch denoiser 1019.
[0244] In some aspects, the pitch denoiser 1019 can operate on the original input audio signal 1015 (e.g., the denoiser 1018 and the pitch denoiser 1019 may both receive the same audio signal 1015 as input, and may operate in parallel within the encoder 1005). In another example, the pitch denoiser 1019 can operate in series (e.g., in sequence) with the denoiser 1018. For example, the input to the pitch denoiser 1019 can be the output of the denoiser 1018, and the pitch denoiser 1019 does not receive the encoder input audio signal 1015 (e.g., the output of the denoiser 1018 can be provided to the spectral envelope feature extraction 1020 and the pitch denoiser 1019 in parallel). In some aspects, the pitch extraction engine 1030 can receive as input the more aggressively de-noised version of the input audio signal 1015 that is generated using the pitch denoiser 1019 and / or theQualcomm Docket No.2404227WO denoiser 1018. The pitch extraction engine 1030 can be configured to perform pitch estimation using one or more DSP techniques, one or more ML or neural network-based techniques, and / or various combinations thereof of DSP and ML or neural network techniques for pitch estimation.
[0245] FIG. 11 is a block diagram illustrating an example of an audio codec system (e.g., generative voice codec system) 1150 that can be used to perform neural speech synthesis based on using a front-end non-speech detector (FNSD) 1114 to switch between a neural synthesizer audio codec path for encoding and / or decoding speech signals and an additional audio codec path for encoding and / or decoding non-speech signals, in accordance with some examples. In some aspects, the generative voice codec system 1100 of FIG.11 can be the same as or similar to the generative voice codec system 400 of FIG. 4, 600 of FIG. 6, 700 of FIG. 7, 800 of FIG. 8, 900 of FIG. 9, 1000 of FIG. 10, etc. The encoder 1105 of FIG. 11 can be the same as or similar to the encoder 405 of FIG.4, 805 of FIG.8, 905 of FIG.9, 1005 of FIG.10, etc. For example, the encoder 1105 can include a denoiser (e.g., spectral branch denoiser) 1118 that is the same as or similar to the denoiser 418 of FIG.4, 1018 and / or 1019 of FIG.10, etc. The encoder 1105 can include spectral envelope feature extraction 1120 the same as or similar to the spectral envelope feature extraction 420 of FIG.4, 820 of FIG.8, 920 of FIG.9, 1020 of FIG.10, etc. The encoder 1105 can include pitch extraction 1130 the same as or similar to the pitch extraction 430 of FIG. 4, 830 of FIG. 8, 930 of FIG. 9, 1030 of FIG. 10, etc. A pitch quantization engine 1135 of FIG.11 can be the same as or similar to the pitch quantization engine 435 of FIG.4, 835 and / or 837 of FIG.8, 1035 of FIG. 10, etc. The encoder 1105 can include an FRAE autoencoder 1125 the same as or similar to the FRAE autoencoder 425 of FIG.4, 825 and / or 815 of FIG.8, 925 of FIG.9, 1025 of FIG.10, etc.
[0246] In some cases, the channel 1140 of FIG.11 can be the same as or similar to the channel 440 of FIG.4, the channel 640 of FIG.6, the channel 740 of FIG.7, the channel 840 of FIG.8, the channel 940 of FIG. 9, the channel 1040 of FIG. 10, etc. The decoder 1110 of FIG. 11 can be the same as or similar to the decoder 410 of FIG. 4, the decoder 610 of FIG. 6, the decoder 710 of FIG. 7, the decoder 810 of FIG.8, the decoder 910 of FIG. 9, the decoder 1010 of FIG. 10, etc. For example, the decoder 1110 can include a spectral feature FRAE decoder 1145 that is the same as or similar to the FRAE decoder 445 of FIG. 4, 645 of FIG. 6, 745 of FIG.7, 845 of FIG. 8 and / or 847 of FIG. 8, 945 of FIG. 9, 1045 of FIG. 10, etc. The decoder 1110 can include a neural speech synthesizerQualcomm Docket No.2404227WO 1150 that can be the same as or similar to the neural speech synthesizer 450 of FIG. 4, 650 of FIG. 6, 750 of FIG. 7, 850 of FIG. 8, 950 of FIG. 9, 1050 of FIG. 10, etc. The decoder 1110 can generate reconstructed audio 1155 (e.g., synthesized speech) that can be the same as or similar to the reconstructed audio 455 of FIG.4, 655 of FIG.6, 755 of FIG.7, 855 of FIG.8, 955 of FIG. 9, 1055 of FIG. 10, etc. In some aspects, the decoder 1110 of FIG. 11 can include an LPC 1152 between the output of the neural speech synthesizer 1150 and the reconstructed audio 1155 output from the decoder 1110. In some examples, the LPC 1152 of FIG.11 can be the same as or similar to the LPC 452 of FIG. 4, 654 of FIG.6, 1052 of FIG.10, etc.
[0247] In one illustrative example, the encoder 1105 of FIG.11 can include a front-end non-speech detector (FNSD) 1114. For example, the FNSD 1114 can receive as input the audio signal 1115 obtained or received by the encoder 1105, and can perform speech and / or non-speech detection to determine whether the input audio signal 1115 includes speech or does not include speech (e.g., includes non-speech audio).
[0248] Input audio signals 1115 that are detected as speech by the FNSD 1114 can follow a speech coding path of the encoder 1105. For example, the speech coding path can be utilized based on the FNSD 1114 outputting an FNSD decision indicative of detecting the input audio signal 1115 is a speech signal. The FNSD decision indicative of the FNSD 1114 detecting speech within the input audio 1115 can be transmitted over the channel 1140 to the decoder 1110, and can be used to control a switch 1104 included in the encoder 1105.
[0249] The switch 1104 can be used to switch the input audio signal 1115 between the speech coding path of the encoder 1105 (e.g., in response to an FNSD decision indicative of the FNSD 1114 detecting speech) and a non-speech or general audio coding path of the encoder 1105 (e.g., in response to an FNSD decision indicative of the FNSD 1114 detecting non-speech or general audio).
[0250] The speech coding path of the encoder 1105 can be the same as or similar to the coding path of the encoder 405 of FIG. 4 and / or the encoder 1005 of FIG. 10 (e.g., including a denoiser 1118, spectral envelope feature extraction 1120, spectral FRAE 1125, pitch extraction 1130, and pitch quantization 1135, etc.).
[0251] When the FNSD 1114 detects speech signals, and generates a corresponding FNSD decision indicative of speech, the switch 1104 can couple the input audio signalQualcomm Docket No.2404227WO 1115 to the denoiser 1118 at the beginning of the speech coding path of the encoder 1105.
[0252] When the FNSD 1114 detects non-speech or general audio signals, and generates a corresponding FNSD decision indicative of non-speech, the switch 1104 can couple the input audio signal 1115 to the non-speech or general audio encoder 1190. When the FNSD decision of the FNSD 1114 is indicative of non-speech detected in the input audio signal 1115, the denoiser 1118 and remaining components of the speech coding path of the encoder 1105 do not receive the input audio signal 1115, are not used to process or code (e.g., encode) the audio signal 1115 or any extracted features thereof, and do not transmit information over the channel 1140 to the decoder 1110.
[0253] Input audio signals 1115 that are detected as non-speech signals by the FNSD 1114 can follow an alternate audio coding path using a non-speech audio encoder 1190 of the encoder 1105 and a non-speech audio decoder 1195 of the decoder 1110. The non- speech audio encoder 1190 can be implemented as an alternate audio codec for coding non-speech and / or general audio signals. The non-speech audio encoder 1190 can be implemented as a non-ML audio encoder, a DSP, an ML or neural network-based audio encoder, and / or various combinations thereof, etc.
[0254] The coded output generated by the non-speech audio encoder 1190 on the non- speech coding path of the encoder 1105 can be quantized and transmitted over the channel 1140 from the encoder 1105 to the decoder 1110. The decoder 1110 can include a switch 1185 that can be controlled based on receiving the FNSD decision (e.g., indicative of speech or non-speech detection by FNSD 1114, and indicative of a corresponding use of either the speech coding path or non-speech coding path of the encoder 1105, decoder 1110, and generative voice codec system 1100).
[0255] For example, the received FNSD decision received over the channel 1140 by the decoder 1110 and from the FNSD 1114 can be decoded and used to switch the decoder 1110 between a speech coding (e.g., decoding) path of the decoder 1110, utilizing the FRAE decoder 1145, neural speech synthesizer 1150, and LPC 1152 to generate the reconstructed audio (e.g., synthesized speech) 1155, and a non-speech or general audio coding (e.g., decoding) path of the decoder 1110, utilizing a non-speech audio decoder 1195 associated with the non-speech audio encoder 1190 of the encoder 1105. The non- speech audio decoder 1195 can be implemented as an alternate audio codec for coding (e.g., decoding) non-speech and / or general audio signals that are encoded by the non-Qualcomm Docket No.2404227WO speech audio encoder 1190. The non-speech audio decoder 1195 can be implemented as a non-ML audio decoder, a DSP, an ML or neural network-based audio decoder, and / or various combinations thereof, etc. The non-speech audio decoder 1195 can generate as output a reconstructed non-speech audio 1198, based on decoding the encoded non- speech audio received over the channel 1140 from the non-speech audio encoder 1190 included in the non-speech coding path of the encoder 1105.
[0256] In one illustrative example, audio codec data from the non-speech audio encoder 1190 can be transmitted over the channel 1140 to the decoder 1110 only if the current audio frame (e.g., current frame of the input audio signal 1115) is detected by the FNSD 1114 as non-speech (e.g., the FNSD decision for the current frame is indicative of non- speech). The remaining portion(s) of the data may be transmitted only if the current frame is detected by the FNSD 1114 as speech (e.g., the FNSD decision for the current frame is indicative of speech, and the encoder 1105 uses the speech coding audio path to generate FRAE 1125 encoded spectral envelope features and pitch information 1130, where both the encoded spectral envelope features and the encoded pitch information comprise the remaining portion(s) of the data that are transmitted over the channel 1140 to the decoder 1110- based on the FNSD 1114 detecting speech in the input audio signal 1115).
[0257] In some examples, the encoder 1105 can be configured to process the input audio signal 111 using the denoiser 1118, VAD, and pitch extraction engine 1130 prior to processing the input audio signal 1115 with the FNSD 1114 to generate a speech or non- speech FNSD decision. For example, in some cases, the respective outputs of the speech coding path of the encoder 1105 can be provided as additional inputs to the FNSD 1114. In some aspects, the FNSD 1114 can be configured to generate the speech or non-speech FNSD decision based on inputs comprising the input audio signal 1115, a denoised version of the input audio signal 1115 generated by the denoiser 1118, extracted spectral envelope features from the spectral envelope feature extraction 1120, extracted pitch features or pitch information from the pitch extraction 1130, etc.
[0258] In some examples, the systems and techniques can implement a generative voice codec system with an encoder that includes a voice activity detection (VAD) engine on the speech path. For example, the VAD engine can be configured to analyze the input audio signal to the encoder and distinguish silence from active speech.
[0259] In some cases, the VAD engine can be implemented by or within the pitchQualcomm Docket No.2404227WO extraction engine. For example, the pitch extraction engine 1130 (or any other pitch extraction engine described in FIGS. 1-11) can include the VAD engine and / or can perform or implement VAD to distinguish silence from active speech. In some examples, a separate VAD engine can be included in the encoder and the separate VAD engine can utilize the extracted pitch features or pitch information (e.g., from the pitch extraction engine) as additional inputs for generating a VAD decision indicative of silence or active speech. For example, a separate VAD engine can use pitch correlation information from the pitch extraction engine to determine a VAD decision indicative of silence or active speech in the input audio signal or frames thereof.
[0260] In some examples, a VAD decision (e.g., silence or active speech) can be used by the encoder of the generative voice codec system to enable discontinuous transmission with optional silence encoding. In some cases, based on the VAD decision indicating that the current frame of input audio is silence or inactive speech, the encoder 1105 can be configured not to transmit over the channel 1140 to the decoder 1110. In some examples, based on the VAD decision indicating that the current frame of input audio is silence or inactive speech, the encoder 1105 can be configured to generate, code, and transmit a coarse, low bit-rate silence encoding to the decoder 1110 via the channel 1140.
[0261] In examples where the encoder 1105 uses a VAD engine and / or VAD decision to transmit silence encodings over the channel 1140 to the decoder 1110, the encoder 1105 can be configured to transmit silence encodings sparsely (e.g., every N-th frame, not every frame or not in consecutive frames, etc., among various other discontinuous transmission schemes or techniques that can be implemented for silence encodings transmitted between the encoder 1105 and the decoder 1110 based on a VAD decision determined by the encoder 1105).
[0262] In some examples, the encoder 1105 can transmit over the channel 1140 to the decoder 1110 silence encodings that include coarse spectral representations of the background noise characteristics, such as low-order LPC coefficients, coarsely quantized (e.g., quantized with low bit-rate) spectral envelope features, etc.
[0263] In some cases, the decoder 1110 can include a silence decoder. In some cases, a silence decoder included in the decoder 1110 of the generative voice codec system 1100 can be implemented as a comfort noise generator. The silence decoder included in the decoder 1110 can use the received silence encoding (e.g., received over the channel 1140Qualcomm Docket No.2404227WO from the encoder 1105, in response to a VAD decision indicative of silence or inactive speech). In some cases, the silence decoder included in the decoder 1110 can use the received silence encoding to update one or more comfort noise generation (CNG) parameters for implementing the comfort noise generator. In some examples, the silence decoder and / or CNG included in the decoder 1110 can be used to approximate or approximately replicate coarse background noise characteristics of the silence or inactive speech detected by the VAD engine or VAD decision for the currently coded frame of the input audio signal 111. For example, the silence decoder and / or CNG can approximate coarse background noise characteristics such as spectral envelope, and can avoid having complete digital silence in the reconstructed audio output signal 1155 during periods of silence or inactive speech in the input audio signal 1115 (e.g., as complete digital silence in the reconstructed audio output signal 1155 may be perceptually detrimental).
[0264] FIG. 12 is a flowchart diagram illustrating an example of a process 1200 for processing one or more audio samples. For example, the process 1200 can correspond to a process for encoding one or more audio samples.
[0265] In some examples, the process 1200 can be performed by a computing device or apparatus or a component or system (e.g., one or more chipsets, one or more processors such as one or more CPUs, DSPs, NPUs, NSPs, microcontrollers, ASICs, FPGAs, programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc., any combination thereof, and / or other component or system) of the computing device or apparatus. The operations of the process 1200 may be implemented as software components that are executed and run on one or more processors (e.g., processor 1410 of FIG. 14 or other processor(s)). In some examples, the process 1200 can be performed by an audio coding system or component thereof, including any of the audio coding systems and / or components thereof of FIGS. 1-11. For example, the process 1200 can be performed by one or more of the voice encoder 252 of FIG.2B, the voice encoder 272 of FIG. 2C, the encoder of FIG. 3, the encoder 405 of the generative voice codec 400 of FIG. 4, the audio codec system(s) 500 and / or 570 of FIGS. 5A-5B, the encoder 705 of the generative voice codec 700 of FIG. 7, the encoder 805 of the generative voice codec 800 of FIG.8, the encoder 905 of the generative voice codec 900 of FIG. 9, the encoder 1005 of the generative voice codec 1000 of FIG. 10, and / or the encoder 1105 of the generative voice codec 1100 of FIG.11, etc.Qualcomm Docket No.2404227WO
[0266] In some aspects, the process 1200 can be performed by a UE, smartphone, mobile computing device, user computing device, etc. The process 1200 may be performed by an apparatus that may be a mobile device (e.g., a mobile phone), a network- connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, or other type of computing device. The operations of the process 1200 may be implemented as software components that are executed and run on one or more processors (e.g., processor 1410 of FIG.14, and / or other processor(s)).
[0267] At block 1202, the apparatus (or component thereof) can generate a first sub- vector of audio features, wherein the first sub-vector corresponds to a first combination of one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio.
[0268] For example, the first sub-vector can be generated using the concatenation engine 903 and the sub-vector splitting operation 907 included in the encoder 905 of FIG. 9. In some cases, the first sub-vector can be generated based on using the sub-vector splitting operation 907 of FIG. 9 to split a combined feature vector generated by the concatenation engine 903 into various different subsets, combinations or sub- combinations, etc., of the input feature types (e.g., energy features 910, spectral envelope features 920, and pitch information 930). Each sub-vector split from the combined feature vector by the sub-vector splitting operation 907 can be provided to a corresponding FRAE autoencoder instance 915, 925, …, etc. of the encoder 905, which may generate a corresponding latent representation z, z2, …, etc., for transmission to a corresponding FRAE decoder instance 947, 945, …, etc. of the decoder 910.
[0269] In some cases, the spectral envelope features and the energy features are determined based on a cepstrum of the audio. For example, the spectral envelope features can be the same as or similar to the spectral envelope features 920 of FIG. 9. In some cases, the energy features can be the same as or similar to the energy features 910 of FIG. 9. In some examples, the pitch features can be the same as or similar to the pitch features 930 of FIG.9.
[0270] In some cases, the pitch features are indicative of a pitch frequency associated with the audio or a pitch lag associated with the audio. In some examples, the pitch features are indicative of a voiced or unvoiced classification associated with the audio. InQualcomm Docket No.2404227WO some examples, the pitch features are indicative of pitch correlation information associated with the audio.
[0271] In some cases, the spectral envelope features can be determined based on processing the audio using a filter bank including a plurality of filters. In some examples, the spectral envelope features can be determined using one or more of the spectral envelope feature extraction engines 420 of FIG.4, 530 of FIG.5A, 820 of FIG., 8, 920 of FIG. 9, 1020 of FIG. 10, 1120 of FIG. 11, etc. In some cases, the spectral envelope features do not include pitch information of the audio. In some examples, to determine the spectral envelope features, the apparatus (or component thereof) can be configured to determine a plurality of cepstral coefficients based on processing the audio using the plurality of filters included in the filter bank. For example, the plurality of filters can comprise a first number of filters, and a number of cepstral coefficients included in the plurality of cepstral coefficients can be equal to the first number. For example, N cepstral coefficients can be generated corresponding to the N filters included in the plurality of filters 580 of the filter bank 532 of FIGS. 5A-5B. The apparatus (or component thereof) can truncate the plurality of cepstral coefficients to obtain a set of truncated cepstral coefficients comprising a subset of the plurality of cepstral coefficients. In some examples, the truncation can be performed using the DCT 555 of FIG. 5A and the truncation operation 560 of FIG. 5A. For example, the set of truncated cepstral coefficients can comprise the M coefficients that are a subset of the N coefficients of FIG. 5A, where M < N. In some cases, the set of truncated cepstral coefficients includes a second number of cepstral coefficients, the second number less than the first number. In some examples, the set of truncated cepstral coefficients does not include pitch information of the audio, based on a difference between the first number and the second number.
[0272] In some cases, to determine the one or more energy features, the apparatus (or component thereof) can be configured to determine the cepstrum based on processing a frame of audio data using a filter bank including a plurality of filters, wherein the frame of audio data is included in the audio, and determine the one or more energy features as a frame energy corresponding to the frame of audio data. In some examples, the apparatus (or component thereof) can be configured to determine the frame energy based on a sum of squares of signal amplitudes within the frame of audio data. In some cases, the spectral envelope features comprise respective cepstral features associated with a correspondingQualcomm Docket No.2404227WO cepstral bin included in a plurality of cepstral bins of the cepstrum. In some examples, to generate the combined feature vector, the apparatus (or component thereof) can be configured to replace the respective cepstral features associated with a first cepstral bin included in the plurality of cepstral bins with the one or more energy features.
[0273] In some cases, the spectral envelope features comprise respective cepstral features associated with a corresponding cepstral bin included in a plurality of cepstral bins of the cepstrum. In some examples, to generate the combined feature vector the apparatus (or component thereof) can be configured to append an additional cepstral bin to the plurality of cepstral bins, wherein the additional cepstral bin includes information indicative of the one or more energy features.
[0274] In some cases, each respective filter of the plurality of filters is associated with a corresponding frequency range, and the frame energy is indicative of a respective per- band energy determined for the corresponding frequency range of each respective filter of the plurality of filters.
[0275] At block 1204, the apparatus (or component thereof) can generate a second sub- vector of audio features, wherein the second sub-vector corresponds to a second combination of one or more of the spectral envelope features of the audio, the energy features of the audio, or the pitch features of the audio.
[0276] In some examples, the second sub-vector can be generated using the same concatenation engine 903 and sub-vector splitting operation 907 used to generate the first sub-vector at block 902 of the process 900. In some cases, each respective feature of a plurality of features included in a set comprising the spectral envelope features, the energy features, and the pitch features is included in at least one of the first sub-vector or the second sub-vector.
[0277] In some examples, the first combination is non-overlapping with the second combination, and features included in the first sub-vector are not included in the second sub-vector. In some cases, the first sub-vector includes a first portion of the spectral envelope features, and the second sub-vector includes a second portion of the spectral envelope features. In some examples, the first sub-vector further includes the pitch features, and wherein the second sub-vector further includes the energy features. In some cases, the first sub-vector further includes a first portion of the energy features, and wherein the second sub-vector further includes a second portion of the energy features.Qualcomm Docket No.2404227WO
[0278] At block 1206, the apparatus (or component thereof) can generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the first sub-vector of audio features.
[0279] For example, the first neural network-based autoencoder can be a first feedback recurrent autoencoder (FRAE). In some cases, the first neural network-based autoencoder can be the same as or similar to the first FRAE 925 of FIG. 9, and can generate a first encoded representation comprising a latent of the first FRAE 925, where the latent of the first FRAE 925 corresponds to the first sub-vector of audio features from the sub-vector splitting operation 907 of FIG.9.
[0280] At block 1208, the apparatus (or component thereof) can generate, using a second neural network-based autoencoder, a second encoded representation corresponding to the second sub-vector of audio features.
[0281] For example, the second neural network-based autoencoder can be a second feedback recurrent autoencoder (FRAE). In some cases, the second neural network-based autoencoder can be the same as or similar to the second FRAE 915 of FIG. 9, and can generate a second encoded representation comprising a latent of the second FRAE 915, where the latent of the second FRAE 915 corresponds to the second sub-vector of audio features from the sub-vector splitting operation 907 of FIG.9.
[0282] In some cases, the apparatus (or component thereof) can be configured to transmit the first encoded representation and the second encoded representation to a decoder. In some examples, the apparatus (or component thereof) can be configured to generate a third sub-vector of audio features, wherein the third sub-vector corresponds to a third combination of the spectral envelope features of the audio, the energy features of the audio, and the pitch features of the audio. The apparatus (or component thereof) can quantize the third sub-vector of audio features to obtain a third encoded representation corresponding to the third sub-vector of audio features. In some cases, the apparatus (or component thereof) can be configured to generate the first encoded representation, the second encoded representation, and the third encoded representation in parallel.
[0283] In some examples, the apparatus comprises an audio encoder or a voice encoder of a generative voice codec. In some examples, the apparatus comprises a voice encoder of a generative voice codec. In some cases, the apparatus (or component thereof) can be configured to transmit the first encoded representation and the second encodedQualcomm Docket No.2404227WO representation to a neural synthesizer included in a voice decoder of the generative voice codec. In some examples, the apparatus further comprises one or more microphones configured to obtain the audio.
[0284] FIG. 13 is a flowchart diagram illustrating an example of a process 1300 for processing one or more audio samples. For example, the process 1300 can correspond to a process for decoding one or more audio samples (e.g., decoding encoded audio samples and / or encoded audio data, etc.).
[0285] In some examples, the process 1300 can be performed by a computing device or apparatus or a component or system (e.g., one or more chipsets, one or more processors such as one or more CPUs, DSPs, NPUs, NSPs, microcontrollers, ASICs, FPGAs, programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc., any combination thereof, and / or other component or system) of the computing device or apparatus. The operations of the process 1300 may be implemented as software components that are executed and run on one or more processors (e.g., processor 1410 of FIG. 14 or other processor(s)). In some examples, the process 1300 can be performed by an audio coding system or component thereof, including any of the audio coding systems and / or components thereof of FIGS.1-11.
[0286] For example, the process 1300 can be performed by one or more of the voice decoder 254 of FIG. 2B, the voice decoder 274 of FIG. 2C, the decoder of FIG. 3, the decoder 410 of the generative voice codec 400 of FIG. 4, the audio codec system(s) 500 and / or 570 of FIGS. 5A-5B, the decoder 610 of the generative voice codec 600 of FIG. 6, the decoder 710 of the generative voice codec 700 of FIG. 7, the decoder 810 of the generative voice codec 800 of FIG.8, the decoder 910 of the generative voice codec 900 of FIG. 9, the decoder 1010 of the generative voice codec 1000 of FIG. 10, and / or the decoder 1110 of the generative voice codec 1100 of FIG.11, etc.
[0287] In some aspects, the process 1300 can be performed by a UE, smartphone, mobile computing device, user computing device, etc. The process 1300 may be performed by an apparatus that may be a mobile device (e.g., a mobile phone), a network- connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, or other type of computing device. The operations of the process 1300 mayQualcomm Docket No.2404227WO be implemented as software components that are executed and run on one or more processors (e.g., processor 1410 of FIG.14, and / or other processor(s)).
[0288] At block 1302, the apparatus (or component thereof) can receive a first encoded representation corresponding to a first combination of features selected from a set of features comprising one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio.
[0289] For example, the first encoded representation can correspond to a first combination of features generated by the concatenation engine 903 and sub-vector splitting operation 907 associated with the energy features 910, spectral envelope features 920, and pitch features 930 of the encoder 905 of FIG.9. In some cases, the first encoded representation can comprise a latent generated by a feedback recurrent autoencoder (FRAE) and corresponding to the first combination of features. For example, the first encoded representation can be a latent generated by the first FRAE 925 of the encoder 905 of FIG.9.
[0290] At block 1304, the apparatus (or component thereof) can receive a second encoded representation corresponding to a second combination of features selected from the set of features.
[0291] For example, the second encoded representation can correspond to a second combination of features generated by the concatenation engine 903 and sub-vector splitting operation 907 associated with the energy features 910, spectral envelope features 920, and pitch features 930 of the encoder 905 of FIG. 9. In some cases, the second encoded representation can comprise a latent generated by a second feedback recurrent autoencoder (FRAE) and corresponding to the second combination of features. For example, the second encoded representation can be a latent generated by the second FRAE 915 of the encoder 905 of FIG.9.
[0292] The first and second encoded representations can correspond to latents of associated first and second FRAEs of an encoder. For example, the first and second encoded representations can correspond to latents of the associated first FRAE 925 and the second FRAE 915 of FIG.9.Qualcomm Docket No.2404227WO
[0293] At block 1306, the apparatus (or component thereof) can generate a first reconstructed sub-vector of features based on the first encoded representation, wherein the first reconstructed sub-vector includes the first combination of features.
[0294] For example, to generate the first reconstructed sub-vector of features, the apparatus (or component thereof) can be configured to decode the first encoded representation using a first neural network-based decoder. In some cases, the first neural network-based decoder comprises a decoder of a feedback recurrent autoencoder (FRAE). For example, the first reconstructed sub-vector of features can be generated by the FRAE decoder 945 of the decoder 910 of FIG.9, based on the first encoded representation from the corresponding first FRAE 925 of the encoder 905 of FIG.9. In some cases, the first encoded representation is generated as a latent representation associated with an FRAE included in an encoder, and the first neural network-based decoder is associated with the FRAE included in the encoder.
[0295] At block 1308, the apparatus (or component thereof) can generate a second reconstructed sub-vector of features based on the second encoded representation, wherein the second reconstructed sub-vector includes the second combination of features.
[0296] For example, to generate the second reconstructed sub-vector of features, the apparatus (or component thereof) can be configured to decode the second encoded representation using a second neural network-based decoder. In some cases, the second neural network-based decoder comprises a decoder of a feedback recurrent autoencoder (FRAE). For example, the second reconstructed sub-vector of features can be generated by the FRAE decoder 947 of the decoder 910 of FIG. 9, based on the second encoded representation from the corresponding second FRAE 915 of the encoder 905 of FIG. 9. In some cases, the second encoded representation is generated as a latent representation associated with an FRAE included in an encoder, and the second neural network-based decoder is associated with the FRAE included in the encoder. In some examples, the first FRAE decoder 945 and the second FRAE decoder 947 can be the same as or similar to one another. In some cases, the first FRAE decoder 945 and the second FRAE decoder 947 can be associated with one another.
[0297] At block 1310, the apparatus (or component thereof) can generate a synthesized audio output signal based on processing the first reconstructed sub-vector of features andQualcomm Docket No.2404227WO the second reconstructed sub-vector of features using a neural network-based signal synthesizer.
[0298] For example, the synthesized audio output signal can be the same as or similar to the neural speech synthesizer 950 output of FIG. 9, the reconstructed audio (e.g., synthesized speech) 955 of FIG.9.
[0299] In some cases, the apparatus (or component thereof) can be configured to separate the first reconstructed sub-vector of features into separate groups based on feature types, the separate groups including one or more of a group of reconstructed spectral envelope features, a group of reconstructed energy features, or a group of reconstructed pitch features. For example, the first reconstructed sub-vector can be obtained from the first FRAE decoder 945, and can be separated into groups based on feature types using the corresponding de-concatenation engine 949-1 of FIG.9.
[0300] In some cases, the apparatus (or component thereof) can be configured to separate the second reconstructed sub-vector of features into separate groups based on feature types, the separate groups including one or more of a group of reconstructed spectral envelope features, a group of reconstructed energy features, or a group of reconstructed pitch features. For example, the second reconstructed sub-vector can be obtained from the second FRAE decoder 947, and can be separated into groups based on feature types using the corresponding de-concatenation engine 949-2 of FIG.9.
[0301] In some cases, the neural network-based signal synthesizer can be the same as or similar to the neural speech synthesizer 950 of FIG. 9. In some examples, the neural network-based signal synthesizer can be used to generate the synthesized audio output signal based on the second reconstructed sub-vector of features, and the separate groups separated from the first reconstructed sub-vector of features. For example, the neural speech synthesizer 950 of FIG. 9 can be used to generate the reconstructed audio (e.g., synthesized speech) 955 based on separated sub-vector features groups obtained from the respective first de-concatenation engine 949-1 and the second de-concatenation engine 949-2 of FIG.9.
[0302] In some cases, the apparatus (or component thereof) can be configured to receive a third encoded representation corresponding to a third combination of features selected from the set of features, wherein the third encoded representation comprises a quantized bit stream. The apparatus (or component thereof) can be configured to generateQualcomm Docket No.2404227WO the synthesized audio output signal further based on processing the third encoded representation.
[0303] In some cases, the apparatus (or component thereof) can be configured to generate a dequantized representation of the third combination of features based on a codebook lookup performed for the quantized bit stream. The apparatus (or component thereof) can generate the synthesized audio output signal further based on processing the dequantized representation of the third combination of features.
[0304] In some examples, the neural network-based signal synthesizer is a neural homomorphic vocoder (NHV). For example, the neural network-based signal synthesizer can be the same as or similar to the NHV of FIG. 6 used to implement the neural speech synthesizer 650 of the decoder 610. In some examples, the apparatus comprises an audio decoder or a voice decoder of a generative voice codec. In some cases, the encoded representation of the combined feature vector is received from an audio encoder or a voice encoder of the generative voice codec. In some cases, the apparatus further comprises one or more speakers configured to output the synthesized audio output signal.
[0305] In some cases, an apparatus used to implement the process 1200 and / or the process 1300 can include one or more microphones configured to obtain the one or more audio samples. In some cases, the apparatus further comprises one or more microphones configured to capture the one or more audio samples for speech synthesis.
[0306] In some cases, the processes described herein (e.g., the process 1200, process 1300, and / or any other process described herein) may be performed by a computing device or apparatus. In one example, the process 1200, the process 1300, and / or other technique or process described herein can be performed by a computing system having an architecture according to any of FIGS.1-11. In another example, the process 1200, the process 1300, and / or other technique or process described herein can be performed by the computing system 1400 shown in FIG. 14. For instance, a computing device with the computing device architecture of the computing system 1400 shown in FIG. 14 can implement the operations of the process 1200, can implement the operations of the process 1300, and / or can implement one or more of the components and / or operations described herein with respect to any of FIGS.1-11.
[0001] In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or moreQualcomm Docket No.2404227WO processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component(s) that are configured to carry out the steps of processes described herein. In some examples, the computing device may include a display, one or more network interfaces configured to communicate and / or receive the data, any combination thereof, and / or other component(s). The one or more network interfaces may be configured to communicate and / or receive wired and / or wireless data, including data according to the 3G, 4G, 5G, and / or other cellular standard, data according to the WiFi (802.11x) standards, data according to the BluetoothTMstandard, data according to the Internet Protocol (IP) standard, and / or other types of data.
[0307] The components of the computing device may be implemented in circuitry. For example, the components may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein.
[0308] The process 1200 and the process 1300 are illustrated as logical flow diagrams, the operation of which represents a sequence of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the processes.
[0309] Additionally, the process 1200, the process 1300, and / or other process described herein, may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As notedQualcomm Docket No.2404227WO above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine- readable storage medium may be non-transitory.
[0310] FIG. 14 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. In particular, FIG.14 illustrates an example of computing system 1400, which may be for example any computing device making up internal computing system, a remote computing system, a camera, or any component thereof in which the components of the system are in communication with each other using connection 1405. Connection 1405 may be a physical connection using a bus, or a direct connection into processor 1410, such as in a chipset architecture. Connection 1405 may also be a virtual connection, networked connection, or logical connection.
[0311] In some aspects, computing system 1400 is a distributed system in which the functions described in this disclosure may be distributed within a datacenter, multiple data centers, a peer network, etc. In some aspects, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some aspects, the components may be physical or virtual devices.
[0312] Example system 1400 includes at least one processing unit (CPU or processor) 1410 and connection 1405 that communicatively couples various system components including system memory 1415, such as read-only memory (ROM) 1420 and random access memory (RAM) 1425 to processor 1410. Computing system 1400 may include a cache 1415 of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1410.
[0313] Processor 1410 may include any general-purpose processor and a hardware service or software service, such as services 1432, 1434, and 1436 stored in storage device 1430, configured to control processor 1410 as well as a special-purpose processor where software instructions are incorporated into the actual processor design. Processor 1410 may essentially be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.Qualcomm Docket No.2404227WO
[0314] To enable user interaction, computing system 1400 includes an input device 1445, which may represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc. Computing system 1400 may also include output device 1435, which may be one or more of a number of output mechanisms. In some instances, multimodal systems may enable a user to provide multiple types of input / output to communicate with computing system 1400.
[0315] Computing system 1400 may include communications interface 1440, which may generally govern and manage the user input and system output. The communication interface may perform or facilitate receipt and / or transmission wired or wireless communications using wired and / or wireless transceivers, including those making use of an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, an AppleTMLightningTMport / plug, an Ethernet port / plug, a fiber optic port / plug, a proprietary wired port / plug, 3G, 4G, 5G and / or other cellular data network wireless signal transfer, a BluetoothTMwireless signal transfer, a BluetoothTMlow energy (BLE) wireless signal transfer, an IBEACONTMwireless signal transfer, a radio-frequency identification (RFID) wireless signal transfer, near-field communications (NFC) wireless signal transfer, dedicated short range communication (DSRC) wireless signal transfer, 802.11 Wi-Fi wireless signal transfer, wireless local area network (WLAN) signal transfer, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), Infrared (IR) communication wireless signal transfer, Public Switched Telephone Network (PSTN) signal transfer, Integrated Services Digital Network (ISDN) signal transfer, ad-hoc network signal transfer, radio wave signal transfer, microwave signal transfer, infrared signal transfer, visible light signal transfer, ultraviolet light signal transfer, wireless signal transfer along the electromagnetic spectrum, or some combination thereof. The communications interface 1440 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers that are used to determine a location of the computing system 1400 based on receipt of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the US-based Global Positioning System (GPS), the Russia-based Global Navigation Satellite System (GLONASS), the China-based BeiDou Navigation Satellite System (BDS), and the Europe-based Galileo GNSS. There is no restriction on operating on any particular hardware arrangement, and therefore theQualcomm Docket No.2404227WO basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0316] Storage device 1430 may be a non-volatile and / or non-transitory and / or computer-readable memory device and may be a hard disk or other types of computer readable media which may store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, a floppy disk, a flexible disk, a hard disk, magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid-state memory, a compact disc read only memory (CD-ROM) optical disc, a rewritable compact disc (CD) optical disc, digital video disk (DVD) optical disc, a blu- ray disc (BDD) optical disc, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a Memory Stick® card, a smartcard chip, a EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (e.g., Level 1 (L1) cache, Level 2 (L2) cache, Level 3 (L3) cache, Level 4 (L4) cache, Level 5 (L5) cache, or other (L#) cache), resistive random-access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.
[0317] The storage device 1430 may include software services, servers, services, etc., that when the code that defines such software is executed by the processor 1410, it causes the system to perform a function. In some aspects, a hardware service that performs a particular function may include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1410, connection 1405, output device 1435, etc., to carry out the function. The term “computer- readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and / or data. A computer-readable medium may include a non- transitory medium in which data may be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections.Qualcomm Docket No.2404227WO Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, memory or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.
[0318] Specific details are provided in the description above to provide a thorough understanding of the aspects and examples provided herein, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative aspects of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, aspects may be utilized in any number of environments and applications beyond those described herein without departing from the broader scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive. For the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate aspects, the methods may be performed in a different order than that described.
[0319] For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks comprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects.Qualcomm Docket No.2404227WO
[0320] Further, those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0321] Individual aspects may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed, but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.
[0322] Processes and methods according to the above-described examples may be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions may include, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used may be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code. Examples of computer- readable media that may be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.Qualcomm Docket No.2404227WO
[0323] In some aspects the computer-readable storage devices, mediums, and memories may include a cable or wireless signal containing a bitstream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0324] Those of skill in the art will appreciate that information and signals may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof, in some cases depending in part on the particular application, in part on the desired design, in part on the corresponding technology, etc.
[0325] The various illustrative logical blocks, modules, and circuits described in connection with the aspects disclosed herein may be implemented or performed using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor(s) may perform the necessary tasks. Examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also may be embodied in peripherals or add-in cards. Such functionality may also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.
[0326] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.
[0327] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices.Qualcomm Docket No.2404227WO Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer- readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that may be accessed, read, and / or executed by a computer, such as propagated signals or waves.
[0328] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, an application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.
[0329] One of ordinary skill will appreciate that the less than (“<”) and greater than (“>”) symbols or terminology used herein may be replaced with less than or equal to (“^”) and greater than or equal to (“^”) symbols, respectively, without departing from the scope of this description.Qualcomm Docket No.2404227WO
[0330] Where components are described as being “configured to” perform certain operations, such configuration may be accomplished, for example, by designing electronic circuits or other hardware to perform the operation, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform the operation, or any combination thereof.
[0331] The phrase “coupled to” or “communicatively coupled to” refers to any component that is physically connected to another component either directly or indirectly, and / or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and / or other suitable communication interface) either directly or indirectly.
[0332] Claim language or other language reciting “at least one of” a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and B, C and C, A and A and B, and so on), or any other ordering, duplication, or combination of A, B, and C. The language “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases “at least one” and “one or more” are used interchangeably herein.
[0333] Claim language or other language reciting “at least one processor configured to,” “at least one processor being configured to,” “one or more processors configured to,” “one or more processors being configured to,” or the like indicates that one processor or multiple processors (in any combination) can perform the associated operation(s). For example, claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each tasked with a certain subset of operations X, Y, and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claimQualcomm Docket No.2404227WO language reciting “at least one processor configured to: X, Y, and Z” can mean that any single processor may only perform at least a subset of operations X, Y, and Z.
[0334] Where reference is made to one or more elements performing functions (e.g., steps of a method), one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function). Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.
[0335] Where reference is made to an entity (e.g., any entity or device described herein) performing functions or being configured to perform functions (e.g., steps of a method), the entity may be configured to cause one or more elements (individually or collectively) to perform the functions. The one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and / or any combination thereof. Where reference to the entity performing functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions. When the entity is configured to cause more than one component to collectively perform the functions, each function need not be performed by each of those components (e.g., different functions may be performed by different components) and / or each function need not be performed in whole by only one component (e.g., different components may perform different sub-functions of a function).
[0336] Illustrative aspects of the disclosure include:
[0337] Aspect 1. An apparatus configured to process audio, the apparatus comprising: one or more memories configured to store the audio; and one or more processors coupled to the one or more memories, the one or more processors being configured to: generate a first sub-vector of audio features, wherein the first sub-vector corresponds to a firstQualcomm Docket No.2404227WO combination of one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; generate a second sub-vector of audio features, wherein the second sub-vector corresponds to a second combination of one or more of the spectral envelope features of the audio, the energy features of the audio, or the pitch features of the audio; generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the first sub-vector of audio features; and generate, using a second neural network-based autoencoder, a second encoded representation corresponding to the second sub-vector of audio features.
[0338] Aspect 2. The apparatus of Aspect 1, wherein the one or more processors are further configured to: generate a third sub-vector of audio features, wherein the third sub- vector corresponds to a third combination of the spectral envelope features of the audio, the energy features of the audio, and the pitch features of the audio; and quantize the third sub-vector of audio features to obtain a third encoded representation corresponding to the third sub-vector of audio features.
[0339] Aspect 3. The apparatus of Aspect 2, wherein the one or more processors are further configured to generate the first encoded representation, the second encoded representation, and the third encoded representation in parallel.
[0340] Aspect 4. The apparatus of any of Aspects 1 to 3, wherein: the one or more processors are configured to transmit the first encoded representation and the second encoded representation to a decoder.
[0341] Aspect 5. The apparatus of any of Aspects 1 to 4, wherein each respective feature of a plurality of features included in a set comprising the spectral envelope features, the energy features, and the pitch features is included in at least one of the first sub-vector or the second sub-vector.
[0342] Aspect 6. The apparatus of any of Aspects 1 to 5, wherein the first combination is non-overlapping with the second combination, and wherein features included in the first sub-vector are not included in the second sub-vector.
[0343] Aspect 7. The apparatus of any of Aspects 1 to 6, wherein the first neural network-based autoencoder comprises a first feedback recurrent autoencoder (FRAE), and wherein the second neural network-based autoencoder comprises a second FRAE.Qualcomm Docket No.2404227WO
[0344] Aspect 8. The apparatus of any of Aspects 1 to 7, wherein the first sub-vector includes a first portion of the spectral envelope features, and the second sub-vector includes a second portion of the spectral envelope features.
[0345] Aspect 9. The apparatus of Aspect 8, wherein the first sub-vector further includes the pitch features, and wherein the second sub-vector further includes the energy features.
[0346] Aspect 10. The apparatus of any of Aspects 8 to 9, wherein the first sub-vector further includes a first portion of the energy features, and wherein the second sub-vector further includes a second portion of the energy features.
[0347] Aspect 11. The apparatus of any of Aspects 1 to 10, wherein the spectral envelope features and the energy features are determined based on a cepstrum of the audio.
[0348] Aspect 12. The apparatus of any of Aspects 1 to 11, wherein the pitch features are indicative of a pitch frequency associated with the audio or a pitch lag associated with the audio.
[0349] Aspect 13. The apparatus of any of Aspects 1 to 12, wherein the pitch features are indicative of a voiced or unvoiced classification associated with the audio.
[0350] Aspect 14. The apparatus of any of Aspects 1 to 13, wherein the pitch features are indicative of pitch correlation information associated with the audio.
[0351] Aspect 15. The apparatus of any of Aspects 1 to 14, wherein the apparatus comprises an audio encoder or a voice encoder of a generative voice codec.
[0352] Aspect 16. The apparatus of any of Aspects 1 to 15, wherein: the apparatus comprises a voice encoder of a generative voice codec; and the one or more processors are configured to transmit the first encoded representation and the second encoded representation to a neural synthesizer included in a voice decoder of the generative voice codec.
[0353] Aspect 17. The apparatus of any of Aspects 1 to 16, further comprising one or more microphones configured to obtain the audio.
[0354] Aspect 18. An apparatus configured to process audio, the apparatus comprising: one or more memories configured to store the audio; and one or more processors coupledQualcomm Docket No.2404227WO to the one or more memories, the one or more processors being configured to: receive a first encoded representation corresponding to a first combination of features selected from a set of features comprising one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; receive a second encoded representation corresponding to a second combination of features selected from the set of features; generate a first reconstructed sub-vector of features based on the first encoded representation, wherein the first reconstructed sub-vector includes the first combination of features; generate a second reconstructed sub-vector of features based on the second encoded representation, wherein the second reconstructed sub-vector includes the second combination of features; and generate a synthesized audio output signal based on processing the first reconstructed sub-vector of features and the second reconstructed sub- vector of features using a neural network-based signal synthesizer.
[0355] Aspect 19. The apparatus of Aspect 18, wherein the one or more processors are further configured to: separate the first reconstructed sub-vector of features into separate groups based on feature types, the separate groups including one or more of a group of reconstructed spectral envelope features, a group of reconstructed energy features, or a group of reconstructed pitch features.
[0356] Aspect 20. The apparatus of Aspect 19, wherein the neural network-based signal synthesizer generates the synthesized audio output signal based on the second reconstructed sub-vector of features, and the separate groups separated from the first reconstructed sub-vector of features.
[0357] Aspect 21. The apparatus of any of Aspects 18 to 20, wherein, to generate the first reconstructed sub-vector of features, the one or more processors are configured to decode the first encoded representation using a first neural network-based decoder.
[0358] Aspect 22. The apparatus of Aspect 21, wherein the first neural network-based decoder comprises a decoder of a feedback recurrent autoencoder (FRAE).
[0359] Aspect 23. The apparatus of Aspect 22, wherein: the first encoded representation is generated as a latent representation associated with an FRAE included in an encoder; and the first neural network-based decoder is associated with the FRAE included in the encoder.Qualcomm Docket No.2404227WO
[0360] Aspect 24. The apparatus of any of Aspects 18 to 23, wherein the first encoded representation comprises a latent generated by a feedback recurrent autoencoder (FRAE) and corresponding to the first combination of features.
[0361] Aspect 25. The apparatus of any of Aspects 18 to 24, wherein the one or more processors are configured to: receive a third encoded representation corresponding to a third combination of features selected from the set of features, wherein the third encoded representation comprises a quantized bit stream; and generate the synthesized audio output signal further based on processing the third encoded representation.
[0362] Aspect 26. The apparatus of Aspect 25, wherein the one or more processors are configured to: generate a dequantized representation of the third combination of features based on a codebook lookup performed for the quantized bit stream; and generate the synthesized audio output signal further based on processing the dequantized representation of the third combination of features.
[0363] Aspect 27. The apparatus of any of Aspects 18 to 26, wherein the neural network-based signal synthesizer is a neural homomorphic vocoder (NHV).
[0364] Aspect 28. The apparatus of any of Aspects 18 to 27, wherein the apparatus comprises an audio decoder or a voice decoder of a generative voice codec.
[0365] Aspect 29. A method for processing audio, the method comprising: generating a first sub-vector of audio features, wherein the first sub-vector corresponds to a first combination of one or more of spectral envelope features of a frame of audio data, energy features of the frame of audio data, or pitch features of the frame of audio data; generating a second sub-vector of audio features, wherein the second sub-vector corresponds to a second combination of one or more of the spectral envelope features of the frame of audio data, the energy features of the frame of audio data, and the pitch features of the frame of audio data; generating, using a first neural network-based autoencoder, a first encoded representation corresponding to the first sub-vector of audio features; and generating, using a second neural network-based autoencoder, a second encoded representation corresponding to the second sub-vector of audio features.
[0366] Aspect 30. The method of Aspect 29, further comprising: generating a third sub- vector of audio features, wherein the third sub-vector corresponds to a third combination of the spectral envelope features of the frame of audio data, the energy features of theQualcomm Docket No.2404227WO frame of audio data, and the pitch features of the frame of audio data; and quantizing the third sub-vector of audio features to obtain a third encoded representation corresponding to the third sub-vector of audio features.
[0367] Aspect 31. The method of Aspect 30, further comprising generating the first encoded representation, the second encoded representation, and the third encoded representation in parallel.
[0368] Aspect 32. The method of any of Aspects 30 to 31, further comprising transmitting the first encoded representation and the second encoded representation to a decoder.
[0369] Aspect 33. The method of any of Aspects 29 to 32, wherein each respective feature of a plurality of features included in a set comprising the spectral envelope features, the energy features, and the pitch features is included in at least one of the first sub-vector or the second sub-vector.
[0370] Aspect 34. The method of any of Aspects 29 to 33, wherein the first combination is non-overlapping with the second combination, and wherein features included in the first sub-vector are not included in the second sub-vector.
[0371] Aspect 35. The method of any of Aspects 29 to 34, wherein the first neural network-based autoencoder comprises a first feedback recurrent autoencoder (FRAE), and wherein the second neural network-based autoencoder comprises a second FRAE.
[0372] Aspect 36. The method of any of Aspects 29 to 35, wherein the first sub-vector includes a first portion of the spectral envelope features, and the second sub-vector includes a second portion of the spectral envelope features.
[0373] Aspect 37. The method of Aspect 36, wherein the first sub-vector further includes the pitch features, and wherein the second sub-vector further includes the energy features.
[0374] Aspect 38. The method of any of Aspects 36 to 37, wherein the first sub-vector further includes a first portion of the energy features, and wherein the second sub-vector further includes a second portion of the energy features.Qualcomm Docket No.2404227WO
[0375] Aspect 39. The method of any of Aspects 29 to 38, wherein the spectral envelope features and the energy features are determined based on a cepstrum of the frame of audio data.
[0376] Aspect 40. The method of any of Aspects 29 to 39, wherein the pitch features are indicative of a pitch frequency associated with the frame of audio data or a pitch lag associated with the frame of audio data.
[0377] Aspect 41. The method of any of Aspects 29 to 40, wherein the pitch features are indicative of a voiced or unvoiced classification associated with the frame of audio data.
[0378] Aspect 42. The method of any of Aspects 29 to 41, wherein the pitch features are indicative of pitch correlation information associated with the frame of audio data.
[0379] Aspect 43. A method for processing audio, the method comprising: receiving a first encoded representation corresponding to a first combination of features selected from a set of features comprising one or more of spectral envelope features of a frame of audio data, energy features of the frame of audio data, or pitch features of the frame of audio data; receiving a second encoded representation corresponding to a second combination of features selected from the set of features; generating a first reconstructed sub-vector of features based on the first encoded representation, wherein the first reconstructed sub- vector includes the first combination of features; generating a second reconstructed sub- vector of features based on the second encoded representation, wherein the second reconstructed sub-vector includes the second combination of features; and generating a synthesized audio output signal based on processing the first reconstructed sub-vector of features and the second reconstructed sub-vector of features using a neural network-based signal synthesizer.
[0380] Aspect 44. The method of Aspect 43, further comprising: separating the first reconstructed sub-vector of features into separate groups based on feature types, the separate groups including one or more of a group of reconstructed spectral envelope features, a group of reconstructed energy features, or a group of reconstructed pitch features.
[0381] Aspect 45. The method of Aspect 44, wherein the neural network-based signal synthesizer generates the synthesized audio output signal based on the secondQualcomm Docket No.2404227WO reconstructed sub-vector of features, and the separate groups separated from the first reconstructed sub-vector of features.
[0382] Aspect 46. The method of any of Aspects 43 to 45, wherein generating the first reconstructed sub-vector of features is based on decoding the first encoded representation using a first neural network-based decoder.
[0383] Aspect 47. The method of Aspect 46, wherein the first neural network-based decoder comprises a decoder of a feedback recurrent autoencoder (FRAE).
[0384] Aspect 48. The method of Aspect 47, wherein: the first encoded representation is generated as a latent representation associated with an FRAE included in an encoder; and the first neural network-based decoder is associated with the FRAE included in the encoder.
[0385] Aspect 49. The method of any of Aspects 43 to 48, wherein the first encoded representation comprises a latent generated by a feedback recurrent autoencoder (FRAE) and corresponding to the first combination of features.
[0386] Aspect 50. The method of any of Aspects 43 to 49, further comprising: receiving a third encoded representation corresponding to a third combination of features selected from the set of features, wherein the third encoded representation comprises a quantized bit stream; and generating the synthesized audio output signal further based on processing the third encoded representation.
[0387] Aspect 51. The method of Aspect 50, further comprising: generating a dequantized representation of the third combination of features based on a codebook lookup performed for the quantized bit stream; and generating the synthesized audio output signal further based on processing the dequantized representation of the third combination of features.
[0388] Aspect 52. The method of any of Aspects 43 to 51, wherein the neural network- based signal synthesizer is a neural homomorphic vocoder (NHV).
[0389] Aspect 53. A method for processing one or more audio samples, comprising performing operations according to any of Aspects 1 to 17 or 29 to 42.Qualcomm Docket No.2404227WO
[0390] Aspect 54. A non-transitory computer-readable storage medium comprising instructions stored thereon which, when executed by at least one processor, causes the at least one processor to perform operations according to any of Aspects 1 to 17 or 29 to 42.
[0391] Aspect 55. An apparatus for processing one or more audio samples, the apparatus comprising one or more means for performing operations according to any of Aspects 1 to 17 or 29 to 42.
[0392] Aspect 56. A method for processing one or more audio samples, comprising performing operations according to any of Aspects 18 to 28 or 43 to 51.
[0393] Aspect 57. A non-transitory computer-readable storage medium comprising instructions stored thereon which, when executed by at least one processor, causes the at least one processor to perform operations according to any of Aspects 18 to 28 or 43 to 51.
[0394] Aspect 58. An apparatus for processing one or more audio samples, the apparatus comprising one or more means for performing operations according to any of Aspects 18 to 28 or 43 to 51.
Claims
Qualcomm Docket No.2404227WO CLAIMS What is claimed is:
1. An apparatus configured to process audio, the apparatus comprising: one or more memories configured to store the audio; and one or more processors coupled to the one or more memories, the one or more processors being configured to: generate a first sub-vector of audio features, wherein the first sub-vector corresponds to a first combination of one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; generate a second sub-vector of audio features, wherein the second sub- vector corresponds to a second combination of one or more of the spectral envelope features of the audio, the energy features of the audio, or the pitch features of the audio; generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the first sub-vector of audio features; and generate, using a second neural network-based autoencoder, a second encoded representation corresponding to the second sub-vector of audio features.
2. The apparatus of claim 1, wherein the one or more processors are further configured to: generate a third sub-vector of audio features, wherein the third sub-vector corresponds to a third combination of the spectral envelope features of the audio, the energy features of the audio, and the pitch features of the audio; and quantize the third sub-vector of audio features to obtain a third encoded representation corresponding to the third sub-vector of audio features.
3. The apparatus of claim 2, wherein the one or more processors are further configured to generate the first encoded representation, the second encoded representation, and the third encoded representation in parallel.
4. The apparatus of claim 1, wherein:Qualcomm Docket No.2404227WO the one or more processors are configured to transmit the first encoded representation and the second encoded representation to a decoder.
5. The apparatus of claim 1, wherein each respective feature of a plurality of features included in a set comprising the spectral envelope features, the energy features, and the pitch features is included in at least one of the first sub-vector or the second sub- vector.
6. The apparatus of claim 1, wherein the first combination is non-overlapping with the second combination, and wherein features included in the first sub-vector are not included in the second sub-vector.
7. The apparatus of claim 1, wherein the first neural network-based autoencoder comprises a first feedback recurrent autoencoder (FRAE), and wherein the second neural network-based autoencoder comprises a second FRAE.
8. The apparatus of claim 1, wherein the first sub-vector includes a first portion of the spectral envelope features, and the second sub-vector includes a second portion of the spectral envelope features.
9. The apparatus of claim 8, wherein the first sub-vector further includes the pitch features, and wherein the second sub-vector further includes the energy features.
10. The apparatus of claim 8, wherein the first sub-vector further includes a first portion of the energy features, and wherein the second sub-vector further includes a second portion of the energy features.
11. The apparatus of claim 1, wherein the spectral envelope features and the energy features are determined based on a cepstrum of the audio.
12. The apparatus of claim 1, wherein the pitch features are indicative of a pitch frequency associated with the audio or a pitch lag associated with the audio.Qualcomm Docket No.2404227WO 13. The apparatus of claim 1, wherein the pitch features are indicative of a voiced or unvoiced classification associated with the audio.
14. The apparatus of claim 1, wherein the pitch features are indicative of pitch correlation information associated with the audio.
15. The apparatus of claim 1, wherein the apparatus comprises an audio encoder or a voice encoder of a generative voice codec.
16. The apparatus of claim 1, wherein: the apparatus comprises a voice encoder of a generative voice codec; and the one or more processors are configured to transmit the first encoded representation and the second encoded representation to a neural synthesizer included in a voice decoder of the generative voice codec.
17. The apparatus of claim 1, further comprising one or more microphones configured to obtain the audio.
18. An apparatus configured to process audio, the apparatus comprising: one or more memories configured to store the audio; and one or more processors coupled to the one or more memories, the one or more processors being configured to: receive a first encoded representation corresponding to a first combination of features selected from a set of features comprising one or more of spectral envelope features of the audio, energy features of the audio, or pitch features of the audio; receive a second encoded representation corresponding to a second combination of features selected from the set of features; generate a first reconstructed sub-vector of features based on the first encoded representation, wherein the first reconstructed sub-vector includes the first combination of features;Qualcomm Docket No.2404227WO generate a second reconstructed sub-vector of features based on the second encoded representation, wherein the second reconstructed sub-vector includes the second combination of features; and generate a synthesized audio output signal based on processing the first reconstructed sub-vector of features and the second reconstructed sub-vector of features using a neural network-based signal synthesizer.
19. The apparatus of claim 18, wherein the one or more processors are further configured to: separate the first reconstructed sub-vector of features into separate groups based on feature types, the separate groups including one or more of a group of reconstructed spectral envelope features, a group of reconstructed energy features, or a group of reconstructed pitch features.
20. The apparatus of claim 19, wherein the neural network-based signal synthesizer generates the synthesized audio output signal based on the second reconstructed sub- vector of features, and the separate groups separated from the first reconstructed sub- vector of features.
21. The apparatus of claim 18, wherein, to generate the first reconstructed sub- vector of features, the one or more processors are configured to decode the first encoded representation using a first neural network-based decoder.
22. The apparatus of claim 21, wherein the first neural network-based decoder comprises a decoder of a feedback recurrent autoencoder (FRAE).
23. The apparatus of claim 22, wherein: the first encoded representation is generated as a latent representation associated with an FRAE included in an encoder; and the first neural network-based decoder is associated with the FRAE included in the encoder.Qualcomm Docket No.2404227WO 24. The apparatus of claim 18, wherein the first encoded representation comprises a latent generated by a feedback recurrent autoencoder (FRAE) and corresponding to the first combination of features.
25. The apparatus of claim 18, wherein the one or more processors are configured to: receive a third encoded representation corresponding to a third combination of features selected from the set of features, wherein the third encoded representation comprises a quantized bit stream; and generate the synthesized audio output signal further based on processing the third encoded representation.
26. The apparatus of claim 25, wherein the one or more processors are configured to: generate a dequantized representation of the third combination of features based on a codebook lookup performed for the quantized bit stream; and generate the synthesized audio output signal further based on processing the dequantized representation of the third combination of features.
27. The apparatus of claim 18, wherein the neural network-based signal synthesizer is a neural homomorphic vocoder (NHV).
28. The apparatus of claim 18, wherein the apparatus comprises an audio decoder or a voice decoder of a generative voice codec.
29. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: generate a first sub-vector of audio features, wherein the first sub-vector corresponds to a first combination of one or more of spectral envelope features of a frame of audio data, energy features of the frame of audio data, or pitch features of the frame of audio data; generate a second sub-vector of audio features, wherein the second sub-vector corresponds to a second combination of one or more of the spectral envelope features ofQualcomm Docket No.2404227WO the frame of audio data, the energy features of the frame of audio data, or the pitch features of the frame of audio data; generate, using a first neural network-based autoencoder, a first encoded representation corresponding to the first sub-vector of audio features; and generate, using a second neural network-based autoencoder, a second encoded representation corresponding to the second sub-vector of audio features.
30. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: receive a first encoded representation corresponding to a first combination of features selected from a set of features comprising spectral envelope features of a frame of audio data, energy features of the frame of audio data, and pitch features of the frame of audio data; receive a second encoded representation corresponding to a second combination of features selected from the set of features; generate a first reconstructed sub-vector of features based on the first encoded representation, wherein the first reconstructed sub-vector includes the first combination of features; generate a second reconstructed sub-vector of features based on the second encoded representation, wherein the second reconstructed sub-vector includes the second combination of features; and generate a synthesized audio output signal based on processing the first reconstructed sub-vector of features and the second reconstructed sub-vector of features using a neural network-based signal synthesizer.
Citation Information
Patent Citations
Method and apparatus for recurrent auto-encoding
US11526734B2
Method and apparatus for recurrent auto-encoding
US20210089863A1
Audio coding using machine learning based linear filters and non-linear neural sources
WO2023064735A1
Sample generation based on joint probability distribution
WO2023133001A1