Speech encoding using analysis-by-synthesis with machine learning speech synthesizer

By employing an ML-based speech synthesis model for analysis-by-synthesis in speech encoding, the limitations of traditional encoders are addressed, achieving efficient and high-quality speech coding with quality guarantees.

WO2025128230A1PCT designated stage expired Publication Date: 2025-06-19QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/054362
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-11-04
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing speech encoding technologies face challenges in achieving high-quality speech coding while maintaining efficient coding and ensuring quality guarantees, as traditional closed-loop encoders are limited by linear filters and open-loop machine learning encoders struggle with quality detection and complexity limitations.

Method used

The implementation of a device with a machine learning (ML)-based speech synthesis model that performs an analysis-by-synthesis operation, generating synthesized versions of the input speech signal to determine the best codebook entry and feature encoding, thereby overcoming the limitations of traditional encoders.

Benefits of technology

This approach enhances coding efficiency and quality by using smaller codebooks at lower bit rates, reducing transmission bandwidth and memory usage, while ensuring quality through local synthesis and codebook compensation for reconstruction issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024054362_19062025_PF_FP_ABST
    Figure US2024054362_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A device includes a memory configured to store data associated with a machine learning (ML)-based speech synthesis model. The device also includes a speech encoder that includes the ML-based speech synthesis model. The speech encoder is configured to perform an analysis-by-synthesis operation of an input speech signal that includes generation, by the ML-based speech synthesis model, of synthesized versions of the input speech signal.
Need to check novelty before this filing date? Find Prior Art

Description

SPEECH ENCODING USING ANALYSIS-BY-SYNTHESIS WITH MACHINE LEARNING SPEECH SYNTHESIZERI. Cross-Reference to Related Applications

[0001] The present application claims the benefit of priority from the commonly owned Greece Patent Application No. 20230101025, filed December 12, 2023, the contents of which are expressly incorporated herein by reference in their entirety.IL Field

[0002] The present disclosure is generally related to speech encoding.III. Description of Related Art

[0003] Advances in technology have resulted in smaller and more powerful computing devices. For example, there currently exist a variety of portable personal computing devices, including wireless telephones such as mobile and smart phones, tablets and laptop computers that are small, lightweight, and easily carried by users. These devices can communicate voice and data packets over wireless networks. Further, many such devices incorporate additional functionality such as a digital still camera, a digital video camera, a digital recorder, and an audio file player. Also, such devices can process executable instructions, including software applications, such as a web browser application, that can be used to access the Internet. As such, these devices can include significant computing capabilities.

[0004] Such computing devices often incorporate functionality to perform speech encoding. For example, traditional closed-loop speech encoders perform feature extraction and quantization of a speech signal. A closed-loop encoding operation includes an iterative local synthesis in which codewords are read from a codebook and used as an input to a linear filter-based speech synthesizer, such as a long-term prediction (LTP) filter and a linear prediction coding (LPC) filter, to generate candidates for a synthesized signal. Each candidate is compared to the speech signal to determine a loss, and the index of the codeword that resulted in the smallest (minimized) loss from among the candidates is sent, along with the encoded features of the speech signal, via a channel to a decoder. The decoder performs a synthesis operation using a codebook andlinear filter-based speech synthesizer that match the codebook and linear filter-based speech synthesizer at the encoder. Traditional closed-loop speech encoders can enable high quality speech coding and decoding, and although larger codebooks can be used to enhance the quality of the synthesized speech, such use of larger codebooks also increases the bit rate due to transmission of larger codebook indices to the decoder. In addition, the modelling power of such encoders can be limited by the use of linear filters.

[0005] In more recent approaches, open-loop machine learning (ML) speech coding systems have been introduced in which an encoder transmits a set of encoded features without using a codebook, and the decoder uses a neural network (NN) speech synthesizer that has been trained to generate a synthesized signal based on the set of features received from the encoder. As compared to closed-loop encoding in which an encoder generates a plurality of candidates for a synthesized signal and a selection is made as to which of these candidates should be used, in open-loop encoding the encoder does not generate a plurality of candidates for the synthesized signal. However, because the encoder does not synthesize speech locally, it cannot detect encoding issues and therefore cannot guarantee quality. Also, at the decoder, the quality of the synthesized signal is limited by the complexity of the NN speech synthesizer. To illustrate, although updating the system to use a larger bitrate (e.g., a larger amount of feature data) can increase the quality of the synthesized signal, at higher bitrates the quality of the synthesized signal becomes limited by the complexity of the synthesizer itself and is not improved by further increases in bitrate. Thus, there is a need for an improved speech encoder.IV. Summary

[0006] According to a particular implementation of the techniques disclosed herein, a device includes a memory configured to store data associated with a machine learning (ML)-based speech synthesis model. The device also includes a speech encoder that includes the ML-based speech synthesis model. The speech encoder is configured to perform an analysis-by-synthesis operation of an input speech signal that includes generation, by the ML-based speech synthesis model, of synthesized versions of the input speech signal.

[0007] According to a particular implementation of the techniques disclosed herein, a method includes obtaining, at a speech encoder, an input speech signal. The method also includes performing, at the speech encoder, an analysis-by-synthesis operation that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal.

[0008] According to a particular implementation of the techniques disclosed herein, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors that include a speech encoder, cause the one or more processors to obtain an input speech signal. The instructions, when executed by the one or more processors, also cause the one or more processors to perform an analysis-by- synthesis operation of the speech encoder that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal.

[0009] According to a particular implementation of the techniques disclosed herein, an apparatus includes means for obtaining an input speech signal. The apparatus also includes means for performing an analysis-by-synthesis operation that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal.

[0010] Other implementations, advantages, and features of the present disclosure will become apparent after review of the entire application, including the following sections: Brief Description of the Drawings, Detailed Description, and the Claims.V. Brief Description of the Drawings

[0011] FIG. l is a block diagram illustrating an example of a system operable to perform speech encoding using analysis-by-synthesis with an ML-based speech synthesis model, in accordance with some examples of the present disclosure.

[0012] FIG. 2 is a block diagram illustrating an example of components that can be implemented in the system of FIG. 1, in accordance with some examples of the present disclosure.

[0013] FIGS. 3 A-3F are block diagrams illustrating examples of components that can be implemented in the system of FIG. 1, in accordance with some examples of the present disclosure.

[0014] FIG. 4 is a block diagram illustrating an example of components that can be implemented in the system of FIG. 1, in accordance with some examples of the present disclosure.

[0015] FIG. 5 is a block diagram of an integrated circuit operable to perform speech encoding using analysis-by-synthesis with an ML-based speech synthesis model, in accordance with some examples of the present disclosure.

[0016] FIG. 6 is a diagram of a portable electronic device operable to perform speech encoding using analysis-by-synthesis with an ML-based speech synthesis model, in accordance with some examples of the present disclosure.

[0017] FIG. 7 is a diagram of a headset operable to perform speech encoding using analysis-by-synthesis with an ML-based speech synthesis model, in accordance with some examples of the present disclosure.

[0018] FIG. 8 is a diagram of a wearable electronic device operable to perform speech encoding using analysis-by-synthesis with an ML-based speech synthesis model, in accordance with some examples of the present disclosure.

[0019] FIG. 9 is a diagram of an extended reality device, such as augmented reality glasses, operable to perform speech encoding using analysis-by-synthesis with an ML- based speech synthesis model, in accordance with some examples of the present disclosure.

[0020] FIG. 10 is a diagram of a headset, such as a virtual reality, mixed reality, or augmented reality headset, operable to perform speech encoding using analysis-by- synthesis with an ML-based speech synthesis model, in accordance with some examples of the present disclosure.

[0021] FIG. 11 is a diagram of earbuds operable to perform speech encoding using analysis-by-synthesis with an ML-based speech synthesis model, in accordance with some examples of the present disclosure.

[0022] FIG. 12 is a diagram of a camera operable to perform speech encoding using analysis-by-synthesis with an ML-based speech synthesis model, in accordance with some examples of the present disclosure.

[0023] FIG. 13 is a diagram of a wireless speaker and voice activated device operable to perform speech encoding using analysis-by-synthesis with an ML-based speech synthesis model, in accordance with some examples of the present disclosure.

[0024] FIG. 14 is a diagram of a first example of a vehicle operable to perform speech encoding using analysis-by-synthesis with an ML-based speech synthesis model, in accordance with some examples of the present disclosure.

[0025] FIG. 15 is a diagram of a second example of a vehicle operable to perform speech encoding using analysis-by-synthesis with an ML-based speech synthesis model in accordance with some examples of the present disclosure.

[0026] FIGS. 16A-16B are diagrams illustrating examples of methods of speech encoding using analysis-by-synthesis with an ML-based speech synthesis model, in accordance with some examples of the present disclosure.

[0027] FIG. 17 is a block diagram of a particular illustrative example of a device that is operable to perform speech encoding using analysis-by-synthesis with an ML- based speech synthesis model, in accordance with some examples of the present disclosure.VI Detailed Description

[0028] Systems and methods to perform speech encoding using analysis-by- synthesis with an ML-based speech synthesis model are disclosed. Conventional closed-loop speech encoders can enable high quality speech coding and decoding; however, use of larger codebooks to improve quality decreases the coding efficiency due to transmission of larger codebook indices to the decoder, and the modelling powerof such encoders can be limited by the use of linear filters. Open-loop ML speech coding systems can improve the coding efficiency but cannot detect encoding issues and therefore cannot guarantee quality, and the quality of the synthesized signal is also limited by the complexity of the NN speech synthesizer.

[0029] The systems and methods described herein introduce improved speech coding that addresses problems associated with conventional closed-loop coding and open-loop ML coding. In one aspect, in the disclosed techniques, closed-loop coding is performed with a ML speech synthesizer to generate a feature encoding associated with input speech. In another aspect, closed-loop coding is performed with a ML speech synthesizer to also generate a codebook index. The codebook index is determined during an analysis-by-synthesis operation in which a series of codebook entries are used in conjunction with an ML-based speech synthesis model to generate synthesized speech candidates. The codebook entry that results in the best reconstructed speech from among the generated candidates is selected, and an indication (e.g., a codebook index) of the selected codebook entry and the feature encoding are used during decoding to generate a reconstructed version of the input speech.

[0030] Moreover, generating the synthesized speech candidates during the analysis- by-synthesis operation may include providing a sequence of the codebook entries as inputs to a neural -network based speech synthesizer during an iterative loop over the codebook entries. Each codebook entry is processed at the neural network speech synthesizer, and the outputs of the neural network speech synthesizer may be used as the synthesized speech candidates. Alternatively, the synthesized speech candidates may be generated by combining the output of the neural network speech synthesizer for each particular codebook entry with outputs of one or more filters that process that codebook entry, by combining the output of the neural network speech synthesizer for each particular codebook entry with the codebook entry, or both, as illustrative, non-limiting examples.

[0031] Thus, the problem that the modelling power in traditional closed-loop encoders can be limited by the use of linear filters is solved by using the analysis-by- synthesis operation with the ML-based speech synthesis model that provides increasedmodelling power in the speech synthesizer as compared to the linear filters of the traditional encoders. Additionally, smaller codebooks can be used to achieve the same encoding quality using a lower bit rate than traditional closed-loop encoders. Speech encoding using a lower bit rate reduces the amount of transmission bandwidth, memory usage, etc., associated with storing and transmitting the encoded speech.

[0032] Further, the problem of reconstruction issues, such as related to vocal fry, energy mismatches as compared to the input speech, and anomalous behavior associated with non-speech inputs that are not detected in open-loop ML encoders, is solved by performing local synthesis at the ML speech encoder. The local synthesis at the speech encoder can detect such reconstruction issues to improve robustness and provide a quality guarantee. The codebook entries can compensate for deficiencies in ML synthesis (e.g., by bridging differences to the original speech), and since the quality can be improved with increased codebook size, performing local synthesis at the speech encoder enables the encoding quality to scale with bit rate, in contrast to open-loop ML encoders that have an upper limit of encoding quality that is determined by the ML synthesizer complexity.

[0033] Particular aspects of the present disclosure are described below with reference to the drawings. In the description, common features are designated by common reference numbers. As used herein, various terminology is used for the purpose of describing particular implementations only and is not intended to be limiting of implementations. For example, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Further, some features described herein are singular in some implementations and plural in other implementations. To illustrate, FIG. 1 depicts a device 102 including one or more processors (“processor(s)” 190 of FIG. 1), which indicates that in some embodiments the device 102 includes a single processor 190 and in other embodiments the device 102 includes multiple processors 190. For ease of reference herein, such features are generally introduced as “one or more” features and are subsequently referred to in the singular or optional plural (as indicated by “(s)” in the name of the feature) unless aspects related to multiple of the features are being described.

[0034] In some drawings, multiple instances of a particular type of feature are used. Although these features are physically and / or logically distinct, the same reference number is used for each, and the different instances are distinguished by addition of a letter to the reference number. When the features as a group or a type are referred to herein e.g., when no particular one of the features is being referenced, the reference number is used without a distinguishing letter. However, when one particular feature of multiple features of the same type is referred to herein, the reference number is used with the distinguishing letter. For example, FIG. 1 depicts codebook entries associated with reference numbers 146A, 146B, and 146N. When referring to a particular one of these codebook entries, such as a codebook entry 146A, the distinguishing letter "A" is used. However, when referring to any arbitrary one of these codebook entries or to these codebook entries as a group, the reference number 146 is used without a distinguishing letter.

[0035] As used herein, the terms “comprise,” “comprises,” and “comprising” may be used interchangeably with “include,” “includes,” or “including.” Additionally, it will be understood that the term “wherein” may be used interchangeably with “where.” As used herein, “exemplary” may indicate an example, an embodiment, and / or an aspect, and should not be construed as limiting or as indicating a preference or a preferred embodiment. As used herein, an ordinal term (e.g., “first,” “second,” “third,” etc.) used to modify an element, such as a structure, a component, an operation, etc., does not by itself indicate any priority or order of the element with respect to another element, but rather merely distinguishes the element from another element having a same name (but for use of the ordinal term). As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to multiple (e.g., two or more) of a particular element.

[0036] As used herein, “coupled” may include “communicatively coupled,” “electrically coupled,” or “physically coupled,” and may also (or alternatively) include any combinations thereof. Two devices (or components) may be coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) directly or indirectly via one or more other devices, components, wires, buses, networks (e.g., a wired network, a wireless network, or a combination thereof), etc. Two devices (orcomponents) that are electrically coupled may be included in the same device or in different devices and may be connected via electronics, one or more connectors, or inductive coupling, as illustrative, non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled, such as in electrical communication, may send and receive signals (e.g., digital signals or analog signals) directly or indirectly, via one or more wires, buses, networks, etc. As used herein, “directly coupled” may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without intervening components.

[0037] In the present disclosure, terms such as “obtaining,” “determining,” “calculating,” “estimating,” “shifting,” “adjusting,” etc. may be used to describe how one or more operations are performed. It should be noted that such terms are not to be construed as limiting and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, “obtaining,” “generating,” “calculating,” “estimating,” “using,” “selecting,” “accessing,” and “determining” may be used interchangeably. For example, “obtaining,” “generating,” “calculating,” “estimating,” or “determining” a parameter (or a signal) may refer to actively generating, estimating, calculating, or determining the parameter (or the signal) or may refer to using, selecting, retrieving, receiving, or accessing the parameter (or signal) that is already generated, such as by another component or device.

[0038] As used herein, the term "machine learning" should be understood to have any of its usual and customary meanings within the fields of computers science and data science, such meanings including, for example, processes or techniques by which one or more computers can learn to perform some operation or function without being explicitly programmed to do so. As a typical example, machine learning can be used to enable one or more computers to analyze data to identify patterns in data and generate a result based on the analysis. For certain types of machine learning, the results that are generated include data that indicates an underlying structure or pattern of the data itself. Such techniques, for example, include so called "clustering" techniques, which identify clusters (e.g., groupings of data elements of the data).

[0039] For certain types of machine learning, the results that are generated include a data model (also referred to as a "machine-learning model" or simply a "model"). Typically, a model is generated using a first data set to facilitate analysis of a second data set. For example, a first portion of a large body of data may be used to generate a model that can be used to analyze the remaining portion of the large body of data. As another example, a set of historical data can be used to generate a model that can be used to analyze future data.

[0040] Since a model can be used to evaluate a set of data that is distinct from the data used to generate the model, the model can be viewed as a type of software (e.g., instructions, parameters, or both) that is automatically generated by the computer(s) during the machine learning process. As such, the model can be portable (e.g., can be generated at a first computer, and subsequently moved to a second computer for further training, for use, or both). Additionally, a model can be used in combination with one or more other models to perform a desired analysis. To illustrate, first data can be provided as input to a first model to generate first model output data, which can be provided (alone, with the first data, or with other data) as input to a second model to generate second model output data indicating a result of a desired analysis. Depending on the analysis and data involved, different combinations of models may be used to generate such results. In some examples, multiple models may provide model output that is input to a single model. In some examples, a single model provides model output to multiple models as input.

[0041] Examples of machine-learning models include, without limitation, perceptrons, neural networks, support vector machines, regression models, decision trees, Bayesian models, Boltzmann machines, adaptive neuro-fuzzy inference systems, as well as combinations, ensembles and variants of these and other types of models. Variants of neural networks include, for example and without limitation, prototypical networks, autoencoders, transformers, self-attention networks, convolutional neural networks, deep neural networks, deep belief networks, etc. Variants of decision trees include, for example and without limitation, random forests, boosted decision trees, etc.

[0042] Since machine-learning models are generated by computer(s) based on input data, machine-learning models can be discussed in terms of at least two distinct time windows - a creation / training phase and a runtime phase. During the creation / training phase, a model is created, trained, adapted, validated, or otherwise configured by the computer based on the input data (which in the creation / training phase, is generally referred to as "training data"). Note that the trained model corresponds to software that has been generated and / or refined during the creation / training phase to perform particular operations, such as classification, prediction, encoding, or other data analysis or data synthesis operations. During the runtime phase (or "inference" phase), the model is used to analyze input data to generate model output. The content of the model output depends on the type of model. For example, a model can be trained to perform classification tasks or regression tasks, as non-limiting examples. In some implementations, a model may be continuously, periodically, or occasionally updated, in which case training time and runtime may be interleaved or one version of the model can be used for inference while a copy is updated, after which the updated copy may be deployed for inference.

[0043] In some implementations, a previously generated model is trained (or retrained) using a machine-learning technique. In this context, "training" refers to adapting the model or parameters of the model to a particular data set. Unless otherwise clear from the specific context, the term "training" as used herein includes "re-training" or refining a model for a specific data set. For example, training may include so called "transfer learning." In transfer learning a base model may be trained using a generic or typical data set, and the base model may be subsequently refined (e.g., re-trained or further trained) using a more specific data set.

[0044] A data set used during training is referred to as a "training data set" or simply "training data". The data set may be labeled or unlabeled. "Labeled data" refers to data that has been assigned a categorical label indicating a group or category with which the data is associated, and "unlabeled data" refers to data that is not labeled. Typically, "supervised machine-learning processes" use labeled data to train a machine-learning model, and "unsupervised machine-learning processes" use unlabeled data to train a machine-learning model; however, it should be understood that a label associated withdata is itself merely another data element that can be used in any appropriate machinelearning process. To illustrate, many clustering operations can operate using unlabeled data; however, such a clustering operation can use labeled data by ignoring labels assigned to data or by treating the labels the same as other data elements.

[0045] Training a model based on a training data set generally involves changing parameters of the model with a goal of causing the output of the model to have particular characteristics based on data input to the model. To distinguish from model generation operations, model training may be referred to herein as optimization or optimization training. In this context, "optimization" refers to improving a metric, and does not mean finding an ideal (e.g., global maximum or global minimum) value of the metric. Examples of optimization trainers include, without limitation, backpropagation trainers, derivative free optimizers (DFOs), and extreme learning machines (ELMs). As one example of training a model, during supervised training of a neural network, an input data sample is associated with a label. When the input data sample is provided to the model, the model generates output data, which is compared to the label associated with the input data sample to generate an error value. Parameters of the model are modified in an attempt to reduce (e.g., optimize) the error value. As another example of training a model, during unsupervised training of an autoencoder, a data sample is provided as input to the autoencoder, and the autoencoder reduces the dimensionality of the data sample (which is a lossy operation) and attempts to reconstruct the data sample as output data. In this example, the output data is compared to the input data sample to generate a reconstruction loss, and parameters of the autoencoder are modified in an attempt to reduce (e.g., optimize) the reconstruction loss.

[0046] FIG. 1 shows a block diagram of a system 100 that includes a device 102 that is configured to perform speech encoding using analysis-by-synthesis at a speech encoder 140 with an ML-based speech synthesis model 144. The device 102 includes a memory 120 that is coupled to one or more processors 190 and configured to store ML- based speech synthesis model data 160 corresponding to the ML-based speech synthesis model 144. In a particular embodiment, the memory 120 corresponds to a dynamic random access memory (DRAM) of a double data rate (DDR) memory subsystem, static random access memory (SRAM), one or more other types of devices configured to storedata and / or instructions that are executable by the processor 190, or a combination thereof.

[0047] The device 102 is coupled to one or more audio sources 110 that are configured to provide audio data 114 (e.g., pulse code modulated (PCM) data) to the processor 190. In some embodiments, one or more of the audio source(s) 110 are integrated within the device 102. For example, the audio source(s) 110 can include media files that include the audio data 114 and that are stored in the memory 120. As another example, the audio source(s) 110 can include one or more microphones integrated within or coupled to the device 102.

[0048] Speech that is included in the audio data 114 is provided to the speech encoder 140 as an input speech signal 118. For example, the processor(s) 190 can include one or more components that are configured to process the audio data 114 (e.g., by performing noise reduction, echo cancellation, source separation, etc.) to obtain the input speech signal 118. In a particular embodiment, the audio source(s) 110 include a microphone configured to provide the input speech signal 118.

[0049] The processor 190 is configured to perform operations associated with encoding the input speech signal 118 at the speech encoder 140. In various embodiments, some or all of the functionality associated with the speech encoder 140 is performed via execution of instructions by the processor 190, performed by processing circuitry of the processor 190 in a hardware implementation, or a combination thereof.

[0050] The speech encoder 140 includes the ML-based speech synthesis model 144 and is configured to perform an analysis-by-synthesis operation 142 of the input speech signal 118. In particular, the analysis-by-synthesis operation 142 generates multiple synthesized versions 148 of the input speech signal 118 to determine a particular codebook entry 146 that results in a best quality match (of the multiple synthesized versions 148) to the input speech signal 118.

[0051] According to an aspect, the analysis-by-synthesis operation 142 includes generation, at the ML-based speech synthesis model 144, one or more of the synthesized versions 148 of the input speech signal 118 based on a set of features 156 of the inputspeech signal 118 and further based on one or more codebook entries (or “codewords”) 146 from one or more codebooks. In an example, the speech encoder 140 extracts the features 156, such as pitch lag or fundamental frequency, a spectrogram (e.g., a log-Mel spectrogram), one or more other features, or a combination thereof, for each frame of the input speech signal 118.

[0052] In a particular embodiment, the analysis-by-synthesis operation 142 includes performing, for each frame of the input speech signal 118, an iterative loop operation that iterates over the codebook entries 146 to generate corresponding synthesized versions 148 of the input speech signal 118. In an illustrative example, each particular codebook entry 146 is sequentially selected, the ML-based speech synthesis model 144 receives a representation of the set of features 156 and the particular codebook entry 146 as inputs, and a synthesized version 148 of the input speech signal 118 for the particular codebook entry 146 is generated based on an output of the ML-based speech synthesis model 144.

[0053] To illustrate, iteration over the codebook entries 146 can include selecting a first codebook entry 146A as input to the ML-based speech synthesis model 144 to generate a first synthesized version 148 A, selecting a second codebook entry 146B as input to the ML-based speech synthesis model 144 to generate a second synthesized version 148B, and continuing until a final codebook entry, illustrated as an Nth codebook entry 146N (where N a positive integer) is input to the ML-based speech synthesis model 144 to generate an Nth synthesized version 148N.

[0054] For each of the one or more of the synthesized versions 148A-148N of the input speech signal, an associated loss function, such as an error metric, is generated based on comparison of the one or more synthesized versions of the input speech signal to the input speech signal. For example, the error metric can include a vector distance measurement, such as a Multi Scale Spectral Loss (MSSL). The analysis-by-synthesis operation 142 can select a particular codebook entry 146 based on the error metrics (e.g., by identifying the codebook entry 146 associated with the least error) for output as the selected codebook entry 154.

[0055] Although in the above example the generation of each of the one or more of the synthesized versions of the input speech signal and the generation of the associated error metrics are performed in an iterative loop over all codebook entry indices, according to other embodiments, one or more other codebook search techniques are used to locate candidate codebook entries 146 to be input to the ML-based speech synthesis model 144. For example, a depth-first tree search method or other focused search method can be used to reduce computational complexity.

[0056] According to an embodiment, the ML-based speech synthesis model 144 is a machine-learning model (or a set of machine-learning models) configured and trained to generate low-bit rate, high quality representations of speech and can include components associated with a neural homomorphic vocoder (NHV), such as described in further detail with reference to FIG. 4. In some embodiments, an NHV is based on a two-state excitation model of the human vocal tract, which enables the NHV to encode audio data representing speech with high-fidelity and a low-bit rate. As one example, the NHV is configured to extract features 156 representing a segment (e.g., a frame or sub-frame) of the input speech signal 118 and provide the features 156 as input to a filter estimator. The filter estimator includes a neural network that is configured and trained to generate filter parameters for a noise filter (e.g., a first linear time varying (LTV) filter) and a harmonic filter (e.g., a second LTV filter). The noise filter is configured to modify a random noise signal based on the noise filter parameters from the filter estimator to generate data representing unvoiced speech components. The harmonic filter is configured to modify, based on the harmonic filter parameters from the filter estimator, an impulse train representing pitch of the audio segment to generate data representing voice speech components.

[0057] In addition, as also described with reference to FIG. 4, the filter estimator is further configured and trained to generate filter parameters of a codebook filter. The codebook filter is configured to modify, based on the codebook filter parameters from the filter estimator, a codeword (e.g., a codebook entry 146) to generate additional data representing speech components. The ML-based speech synthesis model 144 can be coupled to a codebook, such as a fixed codebook or an adaptive codebook, that can be accessed to obtain each of the codebook entries 146. In a particular embodiment, thespeech encoder 140 includes a ML-based codebook generation model configured to generate the codebook based on the features 156 for the current segment and / or one or more prior segments, based on one or more prior selected codebook entries, based on one or more other parameters, or a combination thereof.

[0058] Thus, the filter parameters (e.g., the noise filter parameters, the harmonic filter parameters, and the codebook filter parameters), and possibly other data, are used to generate a representation of the segment in an embodiment in which the ML-based speech synthesis model 144 corresponds to a modified NHV speech synthesis model. In other embodiments, the ML-based speech synthesis model 144 can correspond to one or more other types of speech synthesis models, such as low complexity parametric neural network (LPCNet), LTPNet (e.g., as described in U.S. Patent No. 11,437,050), WaveNet, Lyra, or EnCodec, as illustrative, non-limiting examples, modified to enable the analysis-by-synthesis operation 142 to be performed using the codebook entries 146 for enhanced quality assurance.

[0059] In the example illustrated in FIG. 1, the device 102 optionally includes a modem 170 coupled to the processor(s) 190 and configured to send a bitstream output that includes an output of the speech encoder 140, e.g., an indication of the selected codebook entry 154 from the analysis-by-synthesis operation 142 and an indication of the features 156. For example, the bitstream 172 can be sent, via a communication channel, to one or more remote devices 180. In this example, the remote device 180 includes a decoder system 182 including a speech decoder that uses a ML-based speech synthesis model 184 that matches the ML-based speech synthesis model 144 to generate a reproduced speech signal 188.

[0060] Performing the analysis-by-synthesis operation 142 enables the speech encoder 140 to perform encoding using the efficiency and accuracy of the ML-based speech synthesis model 144 while ensuring the quality of the reproduced speech signal in a manner that conventional open-loop ML-based speech encoders are unable to provide.

[0061] Although the input speech signal 118 (and / or the audio data 114) is described as being provided by the audio source 110, such as from one or moremicrophones or from the memory 120, in other embodiments the input speech signal 118 (and / or the audio data 114) can instead be generated by the one or more processors 190 (e.g., a digital signal processor (DSP), such as audio including speech corresponding to an output of a game engine or other speech generation application), an output of another component of the device 102, or received from another device (e.g., the remote device 180).

[0062] In some embodiments, the device 102 corresponds to or is included in one of various types of devices. In an illustrative example, the speech encoder 140 (e.g., the processor 190) is integrated in a mobile phone or tablet as depicted in FIG. 6, a headset device as depicted in FIG. 7, a wearable electronic device as depicted in FIG. 8, a mixed reality or augmented reality glasses device as depicted in FIG. 9, a virtual reality, mixed reality, or augmented reality headset as depicted in FIG. 10, earbuds, as described with reference to FIG. 11, a camera as depicted in FIG. 12, or a wireless speaker and voice activated device as depicted in FIG. 13. In another illustrative example, the processor 190 is integrated into a vehicle, such as described further with reference to FIG. 14 and FIG. 15.

[0063] FIG. 2 depicts an illustrative example 200 including components that may be implemented in the system 100, such as the speech encoder 140 of the device 102 and the decoder system 182 of the remote device 180. The speech encoder 140 includes an ML-based speech synthesis model 144 and is configured to perform an analysis-by- synthesis operation 142 that includes generation, by the ML-based speech synthesis model 144, of synthesized versions 148 (also referred to as “s[n] candidates”) of an input speech signal s[n] 118.

[0064] The speech encoder 140 performs feature extraction, feature quantization, and closed-loop encoding. As illustrated, a feature extraction operation 210 processes the input speech signal (s[n]) 118 to generate the set of features 156. A feature quantization operation 212 processes the features 156 to generate encoded features 256, such as via a feature quantization codebook lookup operation.

[0065] The closed-loop encoding operation 214 includes the analysis-by-synthesis operation 142 that generates, at the ML-based speech synthesis model 144, one or moreof the synthesized versions (s[n] candidates) 148 of the input speech signal 118 based on the set of features 156 of the input speech signal 118 and further based on one or more codebook entries 146 (also referred to as “codewords”) from one or more codebooks 220. To illustrate, the analysis-by-synthesis operation can operate in a loop that iteratively inputs each codeword 146 from the codebook 220 into the ML-based speech synthesis model 144, and the ML-based speech synthesis model 144 generates a corresponding s[n] candidate 148 for each codeword 146.

[0066] The codebook 220 is coupled to the ML-based speech synthesis model 144 and includes multiple codebook entries 146, each of which is accessible via a codeword index to retrieve the corresponding codeword (e.g., an excitation signal). In some examples, the codebook 220 is a stochastic codebook with regular tracks. The codebook 220 may include pulses, waves, waveform segments, etc. In some embodiments, the codebook 220 is an algebraic codebook, which in some cases may be structured as tracks and pulses (e.g. an entry / codeword is specified as track indices and pulse location indices). In some embodiments, the codebook 220 is a fixed codebook, such as a fixed codebook that is an innovation codebook, i.e., containing vectors that were obtained by randomly sampling from a chosen distribution, e.g. random uniform or random normal. In some embodiments, the codebook 220 is an adaptive codebook. The codebook 220 can be a multi-stage codebook or single-stage codebook. The vectors / codewords in the codebook may be, in some implementations, speech domain segments, excitation domain segments, noise sequences, pulse trains, or a mix / sum thereof, depending on the embodiment. In a particular embodiment, the speech encoder 140 includes a ML-based codebook generation model (not shown) that is configured to generate the codebook 220 based on the features 156 (e.g., based on the encoded features 256 or a dequantized version of the encoded features 256) for a current segment and / or one or more prior segments of the input speech signal s[n] 118, based on one or more prior selected codebook entries, based on one or more other parameters, or a combination thereof.

[0067] The ML-based speech synthesis model 144 includes a NN speech synthesizer 224 and optionally includes a filter 226 and a combiner 228. The ML-based speech synthesis model 144 generates each synthesized version (s[n] candidate) 148 based on an output of the neural network speech synthesizer 224 that receives arepresentation of the set of features 156 and the particular codebook entry 146 as inputs. To illustrate, the codeword 146 and the encoded features 256 (or a dequantized version of the encoded features 256) are provided as inputs to the NN speech synthesizer 224, and the output of the NN speech synthesizer 224 is input to the combiner 228, which outputs the s[n] candidate 148.

[0068] The codeword 146 is also provided as an input to the filter 226 (e.g., as an excitation signal to one or more linear time varying (LTV) filters with filter parameters that are based on the set of features), directly to the combiner 228 via a direct path 222, or a combination thereof. The combiner 228 combines the outputs of the NN speech synthesizer 224, the filter 226, and the direct path 222 to generate an s[n] candidate 148 for each iteration of the loop.

[0069] Optionally, the ML-based speech synthesis model 144 can omit the filter 226, the direct path 222, or both. For example, the synthesized version 148 of the input speech signal 118 can correspond to the output of the neural network speech synthesizer 224 combined with one or both of: the particular codebook entry 146; or an output of the filter 226 that receives the particular codebook entry 146 as an excitation signal. If both the filter 226 and the direct path 222 are omitted, the combiner 228 can also be omitted, and the s[n] candidates 148 correspond to the output of the NN speech synthesizer 224. Various examples of the ML-based speech synthesis model 144 are described with reference to FIGS. 3A-3F.

[0070] The analysis-by-synthesis operation 142 also includes performing a loss determination operation 232 to generate, for each of the one or more of the synthesized versions (s[n] candidates) 148, an associated error metric 234 based on comparison of the one or more synthesized versions (s[n] candidates) 148 to the input speech signal s[n] 118. To illustrate, the loss determination operation 232 can compute each error metric 234 as a loss function (e.g., a vector distance measurement) associated with each of the s[n] candidates 148.

[0071] The generation of each of the one or more of the synthesized versions (s[n] candidates) 148 of the input speech signal 118 and the generation of the associated error metrics 234 are performed in an iterative loop over codebook entry indices for thecodebook 220. For example, after generating an s[n] candidate 148 based on a particular codeword 146 having a particular codeword index, the speech encoder 140 can initiate a next iteration of the analysis-by-synthesis operation 142 by selecting, generating, or otherwise obtaining a next codeword index 238 that is used to retrieve a next codebook entry 146 to be used in generating a next s[n] candidate 148.

[0072] The analysis-by-synthesis operation 142 also includes performing an error minimization operation 236 to select a particular codebook entry 154 based on the error metrics 234, such as by identifying the s[n] candidate 148 with the smallest of the error metrics 234 and selecting the codebook entry 154 that was used to generate that s[n] candidate 148.

[0073] The speech encoder 140 outputs a bitstream 172 that includes an indication of the quantized features and also includes an indication (e.g., a codebook index) of the selected codebook entry 154. For example, the bitstream 172 output of the speech encoder 140 includes the encoded features 256 and also includes a selected codebook index 254 as an indication of the selected codebook entry 154. The bitstream 172 is provided, via a channel 202, to the decoder system 182 for generation of the reproduced speech signal 188.

[0074] The decoder system 182 is configured to perform a synthesis operation 280 using one or more codebooks 260 and an ML-based speech synthesis model 184. The ML-based speech synthesis model 184 includes a neural network speech synthesizer 264 and optionally includes a filter 266 and / or a direct path 262 coupled to a combiner 268.

[0075] Components used in the synthesis operation 280 match corresponding components of the analysis-by-synthesis operation 142 so that, when the decoder system 182 performs the synthesis operation 280 based on the encoded features 256 and the selected codebook index 254 from the bitstream 172, the resulting s[n] candidate (e.g., the reproduced speech signal 188) matches the s[n] candidate 148 that was selected by the speech encoder 140 as having the smallest of the error metrics 234. For example, the codebook 260 matches the codebook 220, so that the selected codebook entry 154 is retrieved using the selected codebook index 254. The ML-based speech synthesismodel 184 matches the ML-based speech synthesis model 144, including a neural network speech synthesizer 264 that matches the neural network speech synthesizer 224, and optionally a filter 266 that matches the filter 226 and a combiner 268 that matches the combiner 228.

[0076] The speech encoder 140, the feature extraction operation 210, the feature quantization operation 212, the closed-loop encoding operation 214, the codebook 220, the ML-based speech synthesis model 144, the neural network speech synthesizer 224, the filter 226, the combiner 228, the loss determination operation 232, the error minimization operation 236, the decoder system 182, the neural network speech synthesizer 264, the filter 266, the combiner 268, or any combination thereof, may be implemented in one or more processors or in processing circuitry. According to an aspect, the various components shown in FIG. 2 are illustrated to assist with understanding the operations performed by processor(s) 190 in accordance with some embodiments. The components may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide particular functionality and are preset on the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in the operations that can be performed. For instance, programmable circuits may execute software or firmware that cause the programmable circuits to operate in the manner defined by instructions of the software or firmware. Fixed-function circuits may execute software instructions (e.g., to receive parameters or output parameters), but the types of operations that the fixed-function circuits perform are generally immutable. In some examples, one or more of the units may be distinct circuit blocks (fixed-function or programmable), and in some examples, the one or more units may be integrated circuits.

[0077] FIGS. 3 A-3F depict various examples of components that can be implemented in the system 100 of FIG. 1, such as in the device 102. In particular, FIGS. 3A-3F depict components of the speech encoder 140, including the codebook 220 coupled to various embodiments of the ML-based speech synthesis model 144 of FIG. 2.

[0078] In FIG. 3A, the ML-based speech synthesis model 144 includes the neural network speech synthesizer 224, the filter 226, the direct path 222, and the combiner 228. A codeword from the codebook 220 is provided as input to the neural network speech synthesizer 224 and to the filter 226. The combiner 228 generates an s[n] candidate 148 by combining (e.g., mixing or summing) an output of the neural network speech synthesizer 224, an output of the filter 226, and the codeword, which is provided to the combiner 228 via the direct path 222.

[0079] In FIG. 3B, the ML-based speech synthesis model 144 includes the neural network speech synthesizer 224, the filter 226, and the combiner 228 but omits the direct path 222. The combiner 228 generates an s[n] candidate 148 by combining (e.g., mixing or summing) the output of the neural network speech synthesizer 224 and the output of the filter 226.

[0080] In FIG. 3C, the ML-based speech synthesis model 144 includes the neural network speech synthesizer 224, the direct path 222, and the combiner 228 but omits the filter 226. The combiner 228 generates an s[n] candidate 148 by combining (e.g., mixing or summing) the output of the neural network speech synthesizer 224 and the codeword received via the direct path 222.

[0081] In FIG. 3D, the ML-based speech synthesis model 144 includes the neural network speech synthesizer 224 and omits the filter 226, the direct path 222, and the combiner 228. The output generated by the neural network speech synthesizer 224 is used as the s[n] candidate 148.

[0082] In FIG. 3E, the ML-based speech synthesis model 144 includes the neural network speech synthesizer 224, the direct path 222, and the combiner 228 but omits the filter 226. The neural network speech synthesizer 224 processes the encoded features 256 but, in contrast to the example of FIG. 3C, the neural network speech synthesizer 224 does not receive codewords from the codebook 220 as input. The combiner 228 generates an s[n] candidate 148 by combining (e.g., mixing or summing) the output of the neural network speech synthesizer 224 and the codeword received via the direct path

[0083] In FIG. 3F, the ML-based speech synthesis model 144 includes the neural network speech synthesizer 224, the filter 226, and the combiner 228 but omits the direct path 222. The neural network speech synthesizer 224 processes the encoded features 256 but, in contrast to the example of FIG. 3B, the neural network speech synthesizer 224 does not receive codewords from the codebook 220 as input. The combiner 228 generates an s[n] candidate 148 by combining (e.g., mixing or summing) the output of the neural network speech synthesizer 224 and the output of the filter 226.

[0084] Although FIGS. 3A-3F illustrate that the output of the ML-based speech synthesis model 144 (e.g., the output of the combiner 228 in FIGS. 3A-C and FIGS. 3E- F or the output of the neural network speech synthesizer 224 in FIG. 3D) corresponds to the s[n] candidate 148, in other embodiments one or more additional filters or other processes may operate on the output of the ML-based speech synthesis model 144 to generate the s[n] candidate 148, such as described further with reference to the embodiment depicted in FIG. 4.

[0085] FIG. 4 depicts an illustrative example of components 400 that can be implemented in the speech encoder 140 and includes the ML-based speech synthesis model 144 coupled to one or more codebooks 420, such as a stochastic codebook.

[0086] The ML-based speech synthesis model 144 includes a neural network filter estimator 450 configured to process a set of features 402 of the input speech signal 118, such as a set of 80 log-Mel features, to generate filter parameters 462 (e.g., estimated impulse responses) for one or more filters of the ML-based speech synthesis model 144. The one or more filters include a harmonic filter 458 (e.g., a first LTV filter) and a noise filter 460 (e.g., a second LTV filter), in addition to a codebook filter 456. In some embodiments, one or more (or all) of the illustrated filters may be implemented in the time domain, as a convolution operation between the input signal and the filter impulse response. In other embodiments, one or more (or all) of the illustrated filters may be implemented in the frequency domain by multiplying the filter frequency response (e.g., spectrum or fast Fourier transform (FFT) of the impulse response) with the spectrum (e.g., FFT) of the filter input signal. Those skilled in the art will appreciate that other possibilities exist for the filter implementations, such as cepstral domain and otherrepresentations. All such possibilities are encompassed within the scope of the present disclosure.

[0087] The codebook filter 456 is configured to process a codeword 422 selected from the codebook 420 and using codebook filter parameters 462A received from the neural network filter estimator 450. The harmonic filter 458 is configured to process a pulse train that is generated by an impulse train generator 452 based on a pitch lag 404 of the input speech signal 118. The harmonic filter 458 uses harmonic filter parameters 462B from the neural network filter estimator 450 to process the pulse train. The noise filter 460 is configured to receive a random noise signal generated by a random noise generator 454 and to process the random noise signal using noise filter parameters 462C from the neural network filter estimator 450. In a particular embodiment, the features 402 and the pitch lag 404 correspond to the features 156 of FIG. 1, such as a dequantized version of the encoded features 256.

[0088] Outputs of the codebook filter 456, the harmonic filter 458, and the noise filter 460 are combined at a combiner 428. A filter 464 (e.g., a trainable causal finite impulse response (FIR) filter) filters the output of the combiner 428 to generate synthesized versions (also referred to as “s[n] candidates”) 430 of the input speech signal 118.

[0089] A loss determination operation 432 determines a loss function for each of the s[n] candidates 430 with respect to the input speech signal s[n] 118. To illustrate, error metrics associated with the s[n] candidates 430 can be determined as Multi Scale Spectral Losses (MSSLs). In some examples, MSSL can be computed as the mean absolute or mean squared difference of the linear scale or log scale magnitude spectra of input speech and synthesized speech, the magnitude spectra computed at multiple time and frequency resolutions (e.g., multiple settings of window lengths, hops, and FFT sizes). According to an aspect, mean squared error (MSE) loss is not used; there is no waveform matching like in traditional code excited linear prediction (CELP). The MSSLs can be processed at a minimization operation 436 to identify a smallest loss associated with the s[n] candidates 430, which is used to determine the selected codebook entry 154 that resulted in the smallest MSSL across all of the codewords 422.

[0090] The components 400 can correspond to a neural homomorphic vocoder (NHV) that has been modified to use the codebook 420 and the codebook filter 456 to enhance operation using the closed-loop analysis-by-synthesis operation 142 of FIG. 1. For example, during the analysis-by-synthesis operation 142, a next codeword index 438 is used during each iteration of the loop to select another codeword 422 from the codebook 420, and inference is performed for each codeword in order to locate the codeword that results in the most accurate synthesized version of the input speech signal 118. However, the present techniques are not limited to NHV and may be applied to a variety of open-loop speech coders that use a ML synthesizer.

[0091] In a particular embodiment, training of the NN filter estimator 450 and the codebook filter 456 is performed with the codebook 420 in the training loop, so that the codebook filter 456 is optimized for the best codeword 422. According to an aspect, during training, the MSSL loss is determined for each of the codewords 422, and the smallest MSSL loss is used to train the codebook filter 456. Training can be simplified by using fixed orthogonal codewords.

[0092] According to an aspect, the codewords 422 in the codebook 420 can also be determined via training. However, because joint training of the codewords 422 and the codebook filter 456 may be impractical due to high complexity, in some embodiments training is performed by first fixing the codewords 422 (e.g., as random vectors) and training the codebook filter 456, followed by freezing all parts of the network other than the codewords 422 and training the codewords 422 to minimize the MSSL losses. During training of the codewords 422, one or more sparsity constraints may be used to ensure that the trained codewords 422 correspond to pulse trains.

[0093] Those skilled in the art will appreciate that other embodiments similar to the one depicted in FIG. 4 are possible, for example by omitting or making optional certain components. All such embodiments are covered in the scope of this the present disclosure. In one example, the random noise generator 454 and random noise LTV filters 460 may be omitted or made optional. In this case, the codebook 420 and codebook filter 456 assume the responsibility of noise modeling. In another example, the impulse train generator 452 and the harmonic LTV filter 458 may be omitted ormade optional, in which case the codebook 420 and codebook filter 456 assume the responsibility of modeling the harmonic component(s) of the speech signal.

[0094] FIG. 5 depicts an embodiment 500 of the device 102 as an integrated circuit 502 that includes the one or more processors 190. The integrated circuit 502 also includes a signal input 504, such as one or more bus interfaces, to enable the input speech signal 118 (or the audio data 114) to be received for processing. The integrated circuit 502 also includes a signal output 506, such as a bus interface, to enable sending of an output signal, such as encoded audio data 508. In this example, the encoded audio data 508 can correspond to or include the selected codebook entry 154, the selected codebook index 254, the features 156, the encoded features 256, the bitstream 172, etc. The integrated circuit 502 including the speech encoder 140 enables implementation of speech encoding using analysis-by-synthesis with an ML-based speech synthesis model as a component in a system, such as a mobile phone or tablet as depicted in FIG. 6, a headset as depicted in FIG. 7, a wearable electronic device as depicted in FIG. 8, a mixed reality or augmented reality glasses device as depicted in FIG. 9, a virtual reality, mixed reality, or augmented reality headset as depicted in FIG. 10, earbuds as depicted in FIG. 11, a camera as depicted in FIG. 12, a wireless speaker and voice activated device as depicted in FIG. 13, or a vehicle as depicted in FIG. 14 or FIG. 15.

[0095] FIG. 6 depicts an embodiment 600 in which the device 102 includes a mobile device 602, such as a phone or tablet, as illustrative, non-limiting examples. The mobile device 602 includes one or more microphones 606, one or more speakers 608, and a display screen 604. Components of the processor 190, including the speech encoder 140, are integrated in the mobile device 602 and are illustrated using dashed lines to indicate internal components that are not generally visible to a user of the mobile device 602. In a particular example, the speech encoder 140 is operable to obtain audio data representing speech captured by the microphone(s) 606 and encode the speech using analysis-by-synthesis with an ML-based speech synthesis model. Using the analysis-by-synthesis operation with the ML-based speech synthesis model enables the speech encoder 140 to combine the efficiency and accuracy of the ML-based speech synthesis model with the quality assurance provided by the analysis-by-synthesis operation. As a result, the mobile device 602 can generate a representation of thespeech with relatively high coding efficiency that is suitable for high quality reproduction.

[0096] FIG. 7 depicts an embodiment 700 in which the device 102 includes a headset device 702. The headset device 702 includes one or more microphones 706 and one or more speakers 708. Components of the processor 190, including the speech encoder 140, are integrated in the headset device 702. In a particular example, the speech encoder 140 is operable to obtain audio data representing speech captured by the microphone(s) 706 and encode the speech using analysis-by-synthesis with an ML- based speech synthesis model. Using the analysis-by-synthesis operation with the ML- based speech synthesis model enables the speech encoder 140 to combine the efficiency and accuracy of the ML-based speech synthesis model with the quality assurance provided by the analysis-by-synthesis operation. As a result, the headset device 702 can generate a representation of the speech with relatively high coding efficiency that is suitable for high quality reproduction.

[0097] FIG. 8 depicts an embodiment 800 in which the device 102 includes a wearable electronic device 802, illustrated as a “smart watch.” The wearable electronic device 802 includes a display screen 804, one or more microphones 806, and one or more speakers 808. Components of the processor 190, including the speech encoder 140, are integrated in the wearable electronic device 802. In a particular example, the speech encoder 140 is operable to obtain audio data representing speech captured by the microphone(s) 606 and encode the speech using analysis-by-synthesis with an ML- based speech synthesis model. Using the analysis-by-synthesis operation with the ML- based speech synthesis model enables the speech encoder 140 to combine the efficiency and accuracy of the ML-based speech synthesis model with the quality assurance provided by the analysis-by-synthesis operation. As a result, the wearable electronic device 802 can generate a representation of the speech with relatively high coding efficiency that is suitable for high quality reproduction. In some embodiments, the wearable electronic device 802 is configured to generate a notification based on content of the speech. For example, the display screen 804 can generate visual information based on the content of the speech. As another example, the wearable electronic device802 can include a haptic device that provides a haptic notification (e.g., vibrates) based on content of the speech.

[0098] FIG. 9 depicts an embodiment 900 in which the device 102 includes a portable electronic device that corresponds to augmented reality or mixed reality glasses 902. The glasses 902 include a holographic projection unit 904 configured to project visual data onto a surface of a lens 906 or to reflect the visual data off of a surface of the lens 906 and onto the wearer's retina. The glasses 902 also include one or more microphones 908 and one or more speakers 910. Components of the processor 190, including the speech encoder 140, are integrated in the glasses 902. In a particular example, the speech encoder 140 is operable to obtain audio data representing speech captured by the microphone(s) 908 and encode the speech using analysis-by-synthesis with an ML-based speech synthesis model. Using the analysis-by-synthesis operation with the ML-based speech synthesis model enables the speech encoder 140 to combine the efficiency and accuracy of the ML-based speech synthesis model with the quality assurance provided by the analysis-by-synthesis operation. As a result, the glasses 902 can generate a representation of the speech with relatively high coding efficiency that is suitable for high quality reproduction. In a particular example, the holographic projection unit 904 is configured to display a notification corresponding to the speech, such as a text representation of the content of the speech.

[0099] FIG. 10 depicts an embodiment 1000 in which the device 102 includes a portable electronic device that corresponds to a virtual reality, mixed reality, or augmented reality headset 1002. A visual interface device 1004 is positioned in front of the user's eyes to enable display of augmented reality, mixed reality, or virtual reality images or scenes to the user while the headset 1002 is worn. The headset 1002 also includes one or more microphones 1006 and one or more speakers 1008. Components of the processor 190, including the speech encoder 140, are integrated in the headset 1002. In a particular example, the speech encoder 140 is operable to obtain audio data representing speech captured by the microphone(s) 1006 and encode the speech using analysis-by-synthesis with an ML-based speech synthesis model. Using the analysis- by-synthesis operation with the ML-based speech synthesis model enables the speech encoder 140 to combine the efficiency and accuracy of the ML-based speech synthesismodel with the quality assurance provided by the analysis-by-synthesis operation. As a result, the headset 1002 can generate a representation of the speech with relatively high coding efficiency that is suitable for high quality reproduction.

[0100] FIG. 11 depicts an embodiment 1100 in which the device 102 includes a portable electronic device that corresponds to a pair of earbuds 1106 that includes a first earbud 1102 and a second earbud 1104. Although earbuds are described, it should be understood that the present technology can be applied to other in-ear or over-ear audio devices.

[0101] The first earbud 1102 includes a first microphone 1120, such as a high signal-to-noise microphone positioned to capture the voice of a wearer of the first earbud 1102, an array of one or more other microphones configured to detect ambient sounds and spatially distributed to support beamforming, illustrated as microphones 1122 A, 1122B, and 1122C, an “inner” microphone 1124 proximate to the wearer’s ear canal (e.g., to assist with active noise cancelling), and a self-speech microphone 1126, such as a bone conduction microphone configured to convert sound vibrations of the wearer’s ear bone or skull into an audio signal.

[0102] The second earbud 1104 can be configured in a substantially similar manner as the first earbud 1102. In some embodiments, the first earbud 1102 is also configured to receive one or more audio signals generated by one or more microphones of the second earbud 1104, such as via wireless transmission between the earbuds 1102, 1104, or via wired transmission in embodiments in which the earbuds 1102, 1104 are coupled via a transmission line.

[0103] In some embodiments, the earbuds 1102, 1104 are configured to automatically switch between various operating modes, such as a passthrough mode in which ambient sound is played via a speaker 1130, a playback mode in which nonambient sound (e.g., streaming audio corresponding to a phone conversation, media playback, video game, etc.) is played back through the speaker 1130, and an audio zoom mode or beamforming mode in which one or more ambient sounds are emphasized and / or other ambient sounds are suppressed for playback at the speaker 1130. In otherembodiments, the earbuds 1102, 1104 may support fewer modes or may support one or more other modes in place of, or in addition to, the described modes.

[0104] In an illustrative example, the earbuds 1102, 1104 can automatically transition from the playback mode to the passthrough mode in response to detecting the wearer’s voice and may automatically transition back to the playback mode after the wearer has ceased speaking. In some examples, the earbuds 1102, 1104 can operate in two or more of the modes concurrently, such as by performing audio zoom on a particular ambient sound (e.g., a dog barking) and playing out the audio zoomed sound superimposed on the sound being played out while the wearer is listening to music (which can be reduced in volume while the audio zoomed sound is being played). In this example, the wearer can be alerted to the ambient sound associated with the audio event without halting playback of the music.

[0105] In FIG. 11, components of the processor 190, including the speech encoder 140, are integrated in the earbuds 1102, 1104. In a particular example, the speech encoder 140 is operable to obtain audio data representing sound including speech captured by one or more of the microphone(s) 1120, 1122, 1124, 1126 and encode the speech using analysis-by-synthesis with an ML-based speech synthesis model. Using the analysis-by-synthesis operation with the ML-based speech synthesis model enables the speech encoder 140 to combine the efficiency and accuracy of the ML-based speech synthesis model with the quality assurance provided by the analysis-by-synthesis operation. As a result, the earbuds 1102, 1104 can generate a representation of the speech with relatively high coding efficiency that is suitable for high quality reproduction.

[0106] FIG. 12 depicts an embodiment 1200 in which the device 102 includes a portable electronic device that corresponds to a camera device 1202. The camera device 1202 includes one or more microphones 1206 and one or more speakers 1208. Components of the processor 190, including the speech encoder 140, are integrated in the camera device 1202. In a particular example, the speech encoder 140 is operable to obtain audio data representing speech captured by the microphone(s) 1206 and encode the speech using analysis-by-synthesis with an ML-based speech synthesis model.Using the analysis-by-synthesis operation with the ML-based speech synthesis model enables the speech encoder 140 to combine the efficiency and accuracy of the ML-based speech synthesis model with the quality assurance provided by the analysis-by-synthesis operation. As a result, the camera device 1202 can generate a representation of the speech with relatively high coding efficiency that is suitable for high quality reproduction.

[0107] FIG. 13 is an embodiment 1300 in which the device 102 includes a wireless speaker and voice activated device 1302. The wireless speaker and voice activated device 1302 can have wireless network connectivity and is configured to execute an assistant operation. The wireless speaker and voice activated device 1302 includes one or more microphones 1306 and one or more speakers 1308. Components of the processor 190, including the speech encoder 140, are integrated in the wireless speaker and voice activated device 1302. In a particular example, the speech encoder 140 is operable to obtain audio data representing speech captured by the microphone(s) 1306 and encode the speech using analysis-by-synthesis with an ML-based speech synthesis model. Using the analysis-by-synthesis operation with the ML-based speech synthesis model enables the speech encoder 140 to combine the efficiency and accuracy of the ML-based speech synthesis model with the quality assurance provided by the analysis- by-synthesis operation. As a result, the wireless speaker and voice activated device 1302 can generate a representation of the speech with relatively high coding efficiency that is suitable for high quality reproduction. In some embodiments, the encoded speech is transmitted to a remote device, such as a remote server, in which at least a portion of a speech recognition process is performed to interpret and execute one or more of the assistant operations detected in the speech of a user of the wireless speaker and voice activated device 1302.

[0108] FIG. 14 depicts an embodiment 1400 in which the device 102 corresponds to, or is integrated within, a vehicle 1402, illustrated as a manned or unmanned aerial device (e.g., a package delivery drone). The vehicle 1402 includes one or more microphones 1406, and one or more speakers 1408. Components of the processor 190, including the speech encoder 140, are integrated in the vehicle 1402. In a particular example, the speech encoder 140 is operable to obtain audio data representing speechcaptured by the microphone(s) 1406 and encode the speech using analysis-by-synthesis with an ML-based speech synthesis model. Using the analysis-by-synthesis operation with the ML-based speech synthesis model enables the speech encoder 140 to combine the efficiency and accuracy of the ML-based speech synthesis model with the quality assurance provided by the analysis-by-synthesis operation. As a result, the vehicle 1402 can generate a representation of the speech with relatively high coding efficiency that is suitable for high quality reproduction. For example, a spoken instruction can be captured by the microphone(s) 1406 and transmitted as a bitstream (e.g., the bitstream 172 of FIG. 1) to a remote device for processing (e.g., to detect delivery instructions or to determine whether the spoken instructions were from an authorized user).

[0109] FIG. 15 depicts another embodiment 1500 in which the device 102 corresponds to, or is integrated within, a vehicle 1502, illustrated as a car. The vehicle 1502 includes a display screen 1520, one or more microphones 1506, and one or more speakers 1508. Components of the processor 190, including the speech encoder 140, are integrated in the vehicle 1502. In a particular example, the speech encoder 140 is operable to obtain audio data representing speech captured by the microphone(s) 1506 and encode the speech using analysis-by-synthesis with an ML-based speech synthesis model. Using the analysis-by-synthesis operation with the ML-based speech synthesis model enables the speech encoder 140 to combine the efficiency and accuracy of the ML-based speech synthesis model with the quality assurance provided by the analysis- by-synthesis operation. As a result, the vehicle 1502 can generate a representation of the speech with relatively high coding efficiency that is suitable for high quality reproduction. For example, a spoken instruction can be captured by the microphone(s) 1406 and transmitted as a bitstream (e.g., the bitstream 172 of FIG. 1) to a remote device for processing (e.g., to obtain navigation data for display at the display screen 1520).

[0110] Although FIGS. 5-15 include examples of devices that include the speech encoder 140, in other embodiments the devices may each include speech decoding functionality associated with the decoder system 182 instead of, or in addition to, the speech encoding functionality associated with the speech encoder 140. For example, the mobile device 602 of FIG. 6 may be configured to receive the bitstream 172 (e.g.,from a remote device during a voice call) including the encoded features 256 and the neural network speech synthesizer 264 and generate the reproduced speech signal 188 using the codebook 260 and the ML-based speech synthesis model 184, such as described previously with reference to FIG. 2.[OHl] FIG. 16A illustrates an example of a method 1600 of speech encoding. One or more operations of the method 1600 may be performed by the system 100 of FIG. 1 (e.g., the device 102, the processor 190, and / or the speech encoder 140), as an illustrative, non-limiting example.

[0112] The method 1600 includes, at block 1602, obtaining, at a speech encoder, an input speech signal. For example, the speech encoder 140 receives the input speech signal 118.

[0113] The method 1600 includes, at block 1604, performing, at the speech encoder, an analysis-by-synthesis operation that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal. For example, the speech encoder 140 performs the analysis-by-synthesis operation 142 that includes the ML-based speech synthesis model 144 generating the synthesized versions 148 of the input speech signal 118.

[0114] In some embodiments, the method 1600 includes, during the analysis-by- synthesis operation, generating, at the ML-based speech synthesis model, one or more of the synthesized versions of the input speech signal based on a set of features of the input speech signal and further based on one or more codebook entries from one or more codebooks. For example, the ML-based speech synthesis model 144 generates the synthesized versions 148 based on the set of features 156 (e.g., the encoded features 256, or a dequantized version of the encoded features 256) and further based on the codebook entries 146 from the one or more codebooks 220. The method 1600 can also include generating, for each of the one or more of the synthesized versions of the input speech signal, an associated error metric based on comparison of the one or more synthesized versions of the input speech signal to the input speech signal, such as the error metrics 234 generated at the loss determination operation 232. The method 1600 can further include selecting a particular codebook entry based on the error metrics,such as described with reference to selected codebook entry 154 at the error minimization operation 236. The method 1600 can also include generating a bitstream output of the speech encoder that includes an indication of the selected codebook entry, such as the bitstream 172 that includes the selected codebook index 254 of the selected codebook entry 154.

[0115] Performing the analysis-by-synthesis operation enables speech encoding using the efficiency and accuracy of the ML-based speech synthesis model while ensuring the quality of the reproduced speech signal in a manner that conventional open-loop ML-based speech encoders are unable to provide.

[0116] The method 1600 of FIG. 16A may be implemented by a field- programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a digital signal processor (DSP), a controller, another hardware device, firmware device, or any combination thereof. As an example, the method 1600 of FIG. 16A may be performed by a processor that executes instructions, such as described with reference to FIG. 17.

[0117] FIG. 16B illustrates an example of a method 1690 of speech encoding. One or more operations of the method 1690 may be performed by the system 100 of FIG. 1 (e.g., the device 102, the processor 190, and / or the speech encoder 140), as an illustrative, non-limiting example.

[0118] The method 1690 includes, at block 1602, obtaining, at a speech encoder, an input speech signal. For example, the speech encoder 140 receives the input speech signal 118.

[0119] The method 1690 includes, at block 1604, performing, at the speech encoder, an analysis-by-synthesis operation that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal. For example, the speech encoder 140 performs the analysis-by-synthesis operation 142 that includes the ML-based speech synthesis model 144 generating the synthesized versions 148 of the input speech signal 118.

[0120] The method 1690 includes, during the analysis-by-synthesis operation, generating, at the ML-based speech synthesis model, a plurality of synthesized versions of the input speech signal, each version based on a set of features of the input speech signal and further based on a respective codebook entry from one or more codebooks, at 1606. For example, the ML-based speech synthesis model 144 generates the synthesized versions 148 based on the set of features 156 (e.g., the encoded features 256, or a dequantized version of the encoded features 256) and further based on the codebook entries 146 from the one or more codebooks 220.

[0121] The method 1690 includes, during the analysis-by-synthesis operation, comparing each of the synthesized versions of the input speech signal to the input speech signal, at 1608. For example, the speech encoder may generate, for each of the plurality of synthesized versions of the input speech signal, an associated error metric based on comparison of the synthesized version of the input speech signal to the input speech signal, such as the error metrics 234 generated at the loss determination operation 232.

[0122] The method 1690 includes, during the analysis-by-synthesis operation, selecting one of the plurality of the synthesized versions based on the comparison, at 1610, and selecting the codebook entry used to generate the selected one of the plurality of synthesized versions of the input speech signal, at 1612. For example, the speech encoder may select a particular synthesized version of the input signal and the codeword corresponding to the selected synthesized version as described with reference to determining the selected codebook entry 154 at the error minimization operation 236.

[0123] The method 1690 also includes generating a bitstream output of the speech encoder that includes an indication of the selected codebook entry, such as the bitstream 172 that includes the selected codebook index 254 of the selected codebook entry 154.

[0124] Performing the analysis-by-synthesis operation enables speech encoding using the efficiency and accuracy of the ML-based speech synthesis model while ensuring the quality of the reproduced speech signal in a manner that conventional open-loop ML-based speech encoders are unable to provide.

[0125] The method 1690 of FIG. 16B may be implemented by a field- programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a digital signal processor (DSP), a controller, another hardware device, firmware device, or any combination thereof. As an example, the method 1690 of FIG. 16B may be performed by a processor that executes instructions, such as described with reference to FIG. 17.

[0126] Referring to FIG. 17, a block diagram of a particular illustrative embodiment of a device is depicted and generally designated 1700. In various embodiments, the device 1700 may have more or fewer components than illustrated in FIG. 17. In an illustrative embodiment, the device 1700 may correspond to the device 102. In an illustrative embodiment, the device 1700 may perform one or more operations described with reference to FIGS. 1-16B.

[0127] In a particular embodiment, the device 1700 includes a processor 1706 (e.g., a central processing unit (CPU)). The device 1700 may include one or more additional processors 1710 (e.g., one or more DSPs). In a particular aspect, the processor 190 of FIG. 1 corresponds to the processor 1706, the processors 1710, or a combination thereof. The processors 1710 may include a speech and music coder-decoder (CODEC) 1708 that includes a voice coder (“vocoder”) encoder 1736, a vocoder decoder 1738, the speech encoder 140, or a combination thereof.

[0128] In this context, the term “processor” refers to an integrated circuit consisting of logic cells, interconnects, input / output blocks, clock management components, memory, and optionally other special purpose hardware components, designed to execute instructions and perform various computational tasks. Examples of processors include, without limitation, central processing units (CPUs), digital signal processors (DSPs), neural processing units (NPU), graphics processing units (GPUs), field programmable gate arrays (FPGAs), microcontrollers, quantum processors, coprocessors, vector processors, other similar circuits, and variants and combinations thereof. In some cases, a processor can be integrated with other components, such as communication components, input / output components, etc. to form a system on a chip (SOC) device or a packaged electronic device.

[0129] Taking CPUs as a starting point, a CPU typically includes one or more processor cores, each of which includes a complex, interconnected network of transistors and other circuit components defining logic gates, memory elements, etc. A core is responsible for executing instructions to, for example, perform arithmetic and logical operations. Typically, a CPU includes an Arithmetic Logic Unit (ALU) that handles mathematical operations and a Control Unit that generates signals to coordinate the operation of other CPU components, such as to manage operations of a fetch- decode-execute cycle.

[0130] CPUs and / or individual processor cores generally include local memory circuits, such as registers and cache to temporarily store data during operations. Registers include high-speed, small-sized memory units intimately connected to the logic cells of a CPU. Often registers include transistors arranged as groups of flip-flops, which are configured to store binary data. Caches include fast, on-chip memory circuits used to store frequently accessed data. Caches can be implemented, for example, using Static Random-Access Memory (SRAM) circuits.

[0131] Operations of a CPU (e.g., arithmetic operations, logic operations, and flow control operations) are directed by software and firmware. At the lowest level, the CPU includes an instruction set architecture (ISA) that specifies how individual operations are performed using hardware resources (e.g., registers, arithmetic units, etc.). Higher level software and firmware is translated into various combinations of ISA operations to cause the CPU to perform specific higher-level operations. For example, an ISA typically specifies how the hardware components of the CPU move and modify data to perform operations such as addition, multiplication, and subtraction, and high-level software is translated into sets of such operations to accomplish larger tasks, such as adding two columns in a spreadsheet. Generally, a CPU operates on various levels of software, including a kernel, an operating system, applications, and so forth, with each higher level of software generally being more abstracted from the ISA and usually more readily understandable by human users.

[0132] GPUs, NPUs, DSPs, microcontrollers, coprocessors, FPGAs, ASICS, and vector processors include components similar to those described above for CPUs. Thedifferences among these various types of processors are generally related to the use of specialized interconnection schemes and ISAs to improve a processor’s ability to perform particular types of operations. For example, the logic gates, local memory circuits, and the interconnects therebetween of a GPU are specifically designed to improve parallel processing, sharing of data between processor cores, and vector operations, and the ISA of the GPU may define operations that take advantage of these structures. As another example, ASICs are highly specialized processors that include similar circuitry arranged and interconnected for a particular task, such as encryption or signal processing. As yet another example, FPGAs are programmable devices that include an array of configurable logic blocks (e.g., interconnected sets of transistors and memory elements) that can be configured (often on the fly) to perform customizable logic functions.

[0133] A processor can be configured to perform a specific task by including, within the processor, specialized hardware to perform the task. Additionally, or alternatively, the processor can be configured to perform a specific task by loading and / or executing instructions (e.g., computer code) that, when executed, cause the processor to perform the specific task. Loading executable instructions to perform the task causes an internal configuration change in the processor that transforms what may otherwise be a general-purpose processor into a special purpose processor for performing the task.

[0134] The device 1700 may include a memory 1786 and a CODEC 1734. The memory 1786 may include instructions 1756, that are executable by the one or more additional processors 1710 (or the processor 1706) to implement the functionality described with reference to the speech encoder 140. The device 1700 may include the modem 170 coupled, via a transceiver 1750, to an antenna 1752.

[0135] The device 1700 may include a display 1728 coupled to a display controller 1726. One or more speakers 1792 and one or more microphones 1794 may be coupled to the CODEC 1734. The CODEC 1734 may include a digital -to-analog converter (DAC) 1702, an analog-to-digital converter (ADC) 1704, or both. In a particular embodiment, the CODEC 1734 may receive analog signals from the microphone(s)1794, convert the analog signals to digital signals using the analog-to-digital converter 1704, and provide the digital signals to the speech and music codec 1708. The speech and music codec 1708 may process the digital signals, and the digital signals may further be processed by the speech encoder 140. In a particular embodiment, the speech and music codec 1708 may provide digital signals to the CODEC 1734. The CODEC 1734 may convert the digital signals to analog signals using the digital -to-analog converter 1702 and may provide the analog signals to the speaker 1792.

[0136] In a particular embodiment, the device 1700 may be included in a system-in- package or system-on-chip device 1722. In a particular embodiment, the memory 1786, the processor 1706, the processors 1710, the display controller 1726, the CODEC 1734, and the modem 170 are included in the system -in-package or system-on-chip device 1722. In a particular embodiment, an input device 1730 and a power supply 1744 are coupled to the system-in-package or the system-on-chip device 1722. Moreover, in a particular embodiment, as illustrated in FIG. 17, the display 1728, the input device 1730, the speaker(s) 1792, the microphone(s) 1794, the antenna 1752, and the power supply 1744 are external to the system-in-package or the system-on-chip device 1722. In a particular embodiment, each of the display 1728, the input device 1730, the speaker(s) 1792, the microphone(s) 1794, the antenna 1752, and the power supply 1744 may be coupled to a component of the system-in-package or the system-on-chip device 1722, such as an interface or a controller.

[0137] The device 1700 may include a smart speaker, a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aerial vehicle, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a car, a computing device, a communication device, an internet-of- things (loT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.

[0138] In conjunction with the described embodiments, an apparatus includes means for obtaining an input speech signal. For example, the means for obtaining the input speech signal can include the system 100, the device 102, the audio source(s) 110, the processor(s) 190, the speech encoder 140, the integrated circuit 502, the microphone(s) 1794, the processor 1706, the processor(s) 1710, the system-in-package or the system- on-chip device 1722, the device 1700, other circuitry configured to obtain an input speech signal, or a combination thereof.

[0139] The apparatus also includes means for performing an analysis-by-synthesis operation that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal. For example, the means for performing an analysis-by-synthesis operation that includes generation, by a machine learning ML-based speech synthesis model, of synthesized versions of the input speech signal can include the system 100, the device 102, the processor(s) 190, the speech encoder 140, the ML-based speech synthesis model 144, the codebook 220, the neural network speech synthesizer 224, the filter 226, the neural network filter estimator 450, the impulse train generator 452, the random noise generator 454, the codebook filter 456, the harmonic filter 458, the noise filter 460, the integrated circuit 502, the processor 1706, the processor(s) 1710, the system-in-package or the system-on-chip device 1722, the device 1700, other circuitry configured to perform an analysis-by- synthesis operation that includes generation, by a machine learning ML-based speech synthesis model, of synthesized versions of the input speech signal, or a combination thereof.

[0140] In some embodiments, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as the memory 1786) includes instructions (e.g., the instructions 1756) that, when executed by one or more processors (e.g., the one or more processors 1710 or the processor 1706) that include a speech encoder (e.g., the speech encoder 140), cause the one or more processors to perform operations corresponding to at least a portion of any of the techniques described with reference to FIGS. 1-15, the methods of FIGS. 16A-16B, or any combination thereof. For example, the instructions, when executed by the one or more processors, cause the one or more processors to obtain an input speech signal (e.g., the input speech signal 118). Theinstructions also cause the one or more processors to perform an analysis-by-synthesis operation (e.g., the analysis-by-synthesis operation 142) of the speech encoder that includes generation, by a machine learning (ML)-based speech synthesis model (e.g., the ML-based speech synthesis model 144), of synthesized versions (e.g., the corresponding synthesized versions 148) of the input speech signal.

[0141] Particular aspects of the disclosure are described below in the following sets of interrelated Examples:

[0142] According to Example 1, a device includes a memory configured to store data associated with a machine learning (ML)-based speech synthesis model; and a speech encoder that includes the ML-based speech synthesis model and that is configured to perform an analysis-by-synthesis operation of an input speech signal that includes generation, by the ML-based speech synthesis model, of synthesized versions of the input speech signal.

[0143] Example 2 includes the device of Example 1, wherein the speech encoder further includes a codebook coupled to the ML-based speech synthesis model.

[0144] Example 3 includes the device of Example 2, wherein the codebook is a fixed codebook.

[0145] Example 4 includes the device of Example 2, wherein the codebook is an adaptive codebook.

[0146] Example 5 includes the device of Example 2 or Example 4 and further includes a ML-based codebook generation model configured to generate the codebook.

[0147] Example 6 includes the device of any of Examples 1 to 5, wherein the ML- based speech synthesis model includes a neural network speech synthesizer.

[0148] Example 7 includes the device of any of Examples 1 to 6, wherein the speech encoder is configured to, during the analysis-by-synthesis operation: generate, at the ML-based speech synthesis model, one or more of the synthesized versions of the input speech signal based on a set of features of the input speech signal and further based onone or more codebook entries from one or more codebooks; generate, for each of the one or more of the synthesized versions of the input speech signal, an associated error metric based on comparison of the one or more synthesized versions of the input speech signal to the input speech signal; and select a particular codebook entry based on the error metrics, wherein a bitstream output of the speech encoder includes an indication of the selected codebook entry.

[0149] Example 8 includes the device of Example 7, wherein a synthesized version of the input speech signal is based on an output of a neural network speech synthesizer that receives a representation of the set of features and the particular codebook entry as inputs.

[0150] Example 9 includes the device of Example 8, wherein the synthesized version of the input speech signal corresponds to the output of the neural network speech synthesizer combined with one or both of: the particular codebook entry; or an output of a filter that receives the particular codebook entry as an excitation signal.

[0151] Example 10 includes the device of any of Examples 7 to 9, wherein the generation of each of the one or more of the synthesized versions of the input speech signal and the generation of the associated error metrics are performed in an iterative loop over codebook entry indices.

[0152] Example 11 includes the device of any of Examples 1 to 10, wherein the ML-based speech synthesis model includes a neural network filter estimator configured to process a set of features of the input speech signal to generate filter parameters for one or more filters of the ML-based speech synthesis model.

[0153] Example 12 includes the device of Examples 11, wherein the one or more filters of the ML-based speech synthesis model include a harmonic filter configured to process a pulse train based on a pitch lag of the input speech signal, a noise filter configured to process a random noise signal, or a codebook filter configured to process codebook entries of a stochastic codebook, wherein the stochastic codebook is coupled to the ML-based speech synthesis model.

[0154] Example 13 includes the device of Example 11 or Example 12 and further includes a combiner configured to combine outputs of the one or more filters to generate a synthesized version of the input speech signal.

[0155] Example 14 includes the device of any of Examples 1 to 13 and further includes a microphone configured to provide the input speech signal.

[0156] Example 15 includes the device of any of Examples 1 to 14 and further includes one or more processors that include the speech encoder.

[0157] Example 16 includes the device of Example 15 and further includes a modem coupled to the one or more processors and configured to send a bitstream output that includes an indication of a codebook entry selection from the analysis-by-synthesis operation, wherein the codebook entry is related to a codebook coupled to the ML-based speech synthesis model.

[0158] Example 17 includes the device of Example 15 or Example 16, wherein the one or more processors are integrated in a headset device, the headset device further including one or more microphones configured to provide the input speech signal.

[0159] Example 18 includes the device of Example 15 or Example 16, wherein the one or more processors are integrated in at least one of a mobile phone, a tablet computer device, or a wearable electronic device.

[0160] Example 19 includes the device of Example 15 or Example 16, wherein the one or more processors are integrated in a vehicle, the vehicle further including one or more microphones configured to provide the input speech signal.

[0161] Example 20 includes the device of any of Examples 15 to 19, wherein the one or more processors are included in an integrated circuit.

[0162] Example 21 includes the device of any of Examples 1 to 20, further including a codebook that includes a plurality of codebook entries, and wherein the speech encoder is configured to perform the analysis-by-synthesis operation including: generate, at the ML-based speech synthesis model, a plurality of synthesized versions ofthe input speech signal, each version based on a set of features of the input speech signal and a codebook entry from the codebook; select one of the plurality of synthesized versions of the input speech signal; select the codebook entry used to generate the selected one of the plurality of synthesized versions of the input speech signal; and include an indication of the selected codebook entry as part of an encoded speech signal.

[0163] According to Example 22, a method includes obtaining, at a speech encoder, an input speech signal; and performing, at the speech encoder, an analysis-by-synthesis operation that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal.

[0164] Example 23 includes the method of Example 22, wherein the speech encoder includes a codebook coupled to the ML-based speech synthesis model.

[0165] Example 24 includes the method of Example 23, wherein the codebook is a fixed codebook.

[0166] Example 25 includes the method of Example 23, wherein the codebook is an adaptive codebook.

[0167] Example 26 includes the method of Example 23 or Example 25 and further includes generating the codebook using an ML-based codebook generation model.

[0168] Example 27 includes the method of any of Examples 22 to 26, wherein the ML-based speech synthesis model includes a neural network speech synthesizer.

[0169] Example 28 includes the method of any of Examples 22 to 27, wherein the analysis-by-synthesis operation includes generating, at the ML-based speech synthesis model, one or more of the synthesized versions of the input speech signal based on a set of features of the input speech signal and further based on one or more codebook entries from one or more codebooks.

[0170] Example 29 includes the method of any of Examples 22 to 28, wherein the analysis-by-synthesis operation includes generating, for each of the one or more of the synthesized versions of the input speech signal, an associated error metric based oncomparison of the one or more synthesized versions of the input speech signal to the input speech signal.

[0171] Example 30 includes the method of Example 29, wherein the analysis-by- synthesis operation includes selecting a particular codebook entry based on the error metrics.

[0172] Example 31 includes the method of Example 30, wherein a bitstream output of the speech encoder includes an indication of the selected codebook entry.

[0173] Example 32 includes the method of any of Examples 28 to 31, wherein a synthesized version of the input speech signal is based on an output of a neural network speech synthesizer that receives a representation of the set of features and the particular codebook entry as inputs.

[0174] Example 33 includes the method of Example 32, wherein the synthesized version of the input speech signal corresponds to the output of the neural network speech synthesizer combined with one or both of: the particular codebook entry; or an output of a filter that receives the particular codebook entry as an excitation signal.

[0175] Example 34 includes the method of any of Examples 29 to 33, wherein the generating of each of the one or more of the synthesized versions of the input speech signal and the generating of the associated error metrics are performed in an iterative loop over codebook entry indices.

[0176] Example 35 includes the method of any of Examples 22 to 34 and further includes processing a set of features of the input speech signal at a neural network filter estimator of the ML-based speech synthesis model to generate filter parameters for one or more filters of the ML-based speech synthesis model.

[0177] Example 36 includes the method of Example 35, wherein the one or more filters of the ML-based speech synthesis model include a harmonic filter configured to process a pulse train based on a pitch lag of the input speech signal, a noise filter configured to process a random noise signal, or a codebook filter configured to processcodebook entries of a stochastic codebook, wherein the stochastic codebook is coupled to the ML-based speech synthesis model.

[0178] Example 37 includes the method of Example 35 or Example 36 and further includes combining outputs of the one or more filters to generate a synthesized version of the input speech signal.

[0179] Example 38 includes the method of any of Examples 22 to 37 and further includes generating the input speech signal at a microphone.

[0180] Example 39 includes the method of any of Examples 22 to 38 and further includes sending a bitstream output that includes an indication of a codebook entry selection from the analysis-by-synthesis operation, wherein the codebook entry is related to a codebook coupled to the ML-based speech synthesis model.

[0181] According to Example 40, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors that include a speech encoder, cause the one or more processors to obtain an input speech signal; and perform an analysis-by-synthesis operation of the speech encoder that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal.

[0182] Example 41 includes the non-transitory computer-readable medium of Example 40, wherein the speech encoder includes a codebook coupled to the ML-based speech synthesis model.

[0183] Example 42 includes the non-transitory computer-readable medium of Example 41, wherein the codebook is a fixed codebook.

[0184] Example 43 includes the non-transitory computer-readable medium of Example 41, wherein the codebook is an adaptive codebook.

[0185] Example 44 includes the non-transitory computer-readable medium of Example 41 or Example 43, wherein the instructions further cause the one or more processors to generate the codebook using a ML-based codebook generation model.

[0186] Example 45 includes the non-transitory computer-readable medium of any of Examples 40 to 44, wherein the ML-based speech synthesis model includes a neural network speech synthesizer.

[0187] Example 46 includes the non-transitory computer-readable medium of any of Examples 40 to 45, wherein the analysis-by-synthesis operation includes: generation, at the ML-based speech synthesis model, of one or more of the synthesized versions of the input speech signal based on a set of features of the input speech signal and further based on one or more codebook entries from one or more codebooks; generation, for each of the one or more of the synthesized versions of the input speech signal, of an associated error metric based on comparison of the one or more synthesized versions of the input speech signal to the input speech signal; and selection of a particular codebook entry based on the error metrics, wherein a bitstream output of the speech encoder includes an indication of the selected codebook entry.

[0188] Example 47 includes the non-transitory computer-readable medium of Example 46, wherein a synthesized version of the input speech signal is based on an output of a neural network speech synthesizer that receives a representation of the set of features and the particular codebook entry as inputs.

[0189] Example 48 includes the non-transitory computer-readable medium of Example 47, wherein the synthesized version of the input speech signal corresponds to the output of the neural network speech synthesizer combined with one or both of: the particular codebook entry; or an output of a filter that receives the particular codebook entry as an excitation signal.

[0190] Example 49 includes the non-transitory computer-readable medium of any of Examples 46 to 48, wherein the generation of each of the one or more of the synthesized versions of the input speech signal and the generation of the associated error metrics are performed in an iterative loop over codebook entry indices.

[0191] Example 50 includes the non-transitory computer-readable medium of any of Examples 40 to 49, wherein the instructions further cause the one or more processors to process a set of features of the input speech signal at a neural network filter estimator ofthe ML-based speech synthesis model to generate filter parameters for one or more filters of the ML-based speech synthesis model.

[0192] Example 51 includes the non-transitory computer-readable medium of Example 50, wherein the one or more filters of the ML-based speech synthesis model include a harmonic filter configured to process a pulse train based on a pitch lag of the input speech signal, a noise filter configured to process a random noise signal, or a codebook filter configured to process codebook entries of a stochastic codebook, wherein the stochastic codebook is coupled to the ML-based speech synthesis model.

[0193] Example 52 includes the non-transitory computer-readable medium of Example 50 or Example 51, wherein the instructions further cause the one or more processors to combine outputs of the one or more filters to generate a synthesized version of the input speech signal.

[0194] Example 53 includes the non-transitory computer-readable medium of any of Examples 40 to 52, wherein the instructions further cause the one or more processors to obtain the input speech signal from a microphone.

[0195] Example 54 includes the non-transitory computer-readable medium of any of Examples 40 to 53, wherein the instructions further cause the one or more processors to send a bitstream output that includes an indication of a codebook entry selection from the analysis-by-synthesis operation, wherein the codebook entry is related to a codebook coupled to the ML-based speech synthesis model.

[0196] According to Example 55, an apparatus includes means for obtaining an input speech signal; and means for performing an analysis-by-synthesis operation that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal.

[0197] Example 56 includes the apparatus of Example 55, wherein the means for performing an analysis-by-synthesis operation includes a codebook coupled to the ML- based speech synthesis model.

[0198] Example 57 includes the apparatus of Example 56, wherein the codebook is a fixed codebook.

[0199] Example 58 includes the apparatus of Example 56, wherein the codebook is an adaptive codebook.

[0200] Example 59 includes the apparatus of Example 56 or Example 58 and further includes means for generating the codebook.

[0201] Example 60 includes the apparatus of any of Examples 55 to 59, wherein the ML-based speech synthesis model includes a neural network speech synthesizer.

[0202] Example 61 includes the apparatus of any of Examples 55 to 60, wherein the means for performing an analysis-by-synthesis operation includes: means for generating, at the ML-based speech synthesis model, one or more of the synthesized versions of the input speech signal based on a set of features of the input speech signal and further based on one or more codebook entries from one or more codebooks; means for generating, for each of the one or more of the synthesized versions of the input speech signal, an associated error metric based on comparison of the one or more synthesized versions of the input speech signal to the input speech signal; and means for selecting a particular codebook entry based on the error metrics, wherein a bitstream output of the means for performing an analysis-by-synthesis operation includes an indication of the selected codebook entry.

[0203] Example 62 includes the apparatus of Example 61, wherein a synthesized version of the input speech signal is based on an output of a neural network speech synthesizer that receives a representation of the set of features and the particular codebook entry as inputs.

[0204] Example 63 includes the apparatus of Example 62, wherein the synthesized version of the input speech signal corresponds to the output of the neural network speech synthesizer combined with one or both of: the particular codebook entry; or an output of a filter that receives the particular codebook entry as an excitation signal.

[0205] Example 64 includes the apparatus of any of Examples 61 to 63, wherein generation of the one or more of the synthesized versions of the input speech signal and generation of the associated error metrics are performed in an iterative loop over codebook entry indices.

[0206] Example 65 includes the apparatus of any of Examples 55 to 64, wherein the ML-based speech synthesis model includes means for processing a set of features of the input speech signal at a neural network filter estimator to generate filter parameters for one or more filters of the ML-based speech synthesis model.

[0207] Example 66 includes the apparatus of Example 61 Example 65, wherein the one or more filters of the ML-based speech synthesis model include means for harmonic filtering of a pulse train based on a pitch lag of the input speech signal, means for noise filtering of a random noise signal, or means for filtering codebook entries of a stochastic codebook, wherein the stochastic codebook is coupled to the ML-based speech synthesis model.

[0208] Example 67 includes the apparatus of Example 65 or Example 66 and further includes means for combining outputs of the one or more filters to generate a synthesized version of the input speech signal.

[0209] Example 68 includes the apparatus of any of Examples 55 to 67 and further includes means for providing the input speech signal.

[0210] Example 69 includes the apparatus of any of Examples 55 to 68 and further includes means for processing that includes the means for performing an analysis-by- synthesis operation, wherein the codebook entry is related to a codebook coupled to the ML-based speech synthesis model.

[0211] Example 70 includes the apparatus of Example 69 and further includes means for sending a bitstream output that includes an indication of a codebook entry selection from the analysis-by-synthesis operation.

[0212] Example 71 includes the apparatus of Example 69 or Example 70, wherein the means for processing is integrated in a headset apparatus, the headset apparatusfurther including one or more microphones configured to provide the input speech signal.

[0213] Example 72 includes the apparatus of Example 69 or Example 70, wherein the means for processing is integrated in at least one of a mobile phone, a tablet computer apparatus, or a wearable electronic apparatus.

[0214] Example 73 includes the apparatus of Example 69 or Example 70, wherein the means for processing is integrated in a vehicle, the vehicle further including one or more microphones configured to provide the input speech signal.

[0215] Example 74 includes the apparatus of any of Examples 69 to 73, wherein the means for processing is included in an integrated circuit.

[0216] Those of skill would further appreciate that the various illustrative logical blocks, configurations, circuits, and algorithm steps described in connection with the implementations disclosed herein may be implemented as electronic hardware, computer software executed by a processing device such as a hardware processor, or combinations of both. Various illustrative components, blocks, configurations, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or executable software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0217] The steps of a method or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module may reside in a memory device, such as random access memory (RAM), magnetoresistive random access memory (MRAM), spin-torque transfer MRAM (STT- MRAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, a removable disk, acompact disc read-only memory (CD-ROM), or any other form of non-transient storage medium known in the art. An exemplary memory device is coupled to the processor such that the processor can read information from, and write information to, the memory device. In the alternative, the memory device may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a computing device or a user terminal.

[0218] The previous description of the disclosed implementations is provided to enable a person skilled in the art to make or use the disclosed implementations. Various modifications to these implementations will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other implementations without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the implementations shown herein but is to be accorded the widest scope possible consistent with the principles and novel features as defined by the following claims.

Claims

WHAT IS CLAIMED IS:

1. A device comprising: a memory configured to store data associated with a machine learning (ML)-based speech synthesis model; and a speech encoder that includes the ML-based speech synthesis model and that is configured to perform an analysis-by-synthesis operation of an input speech signal that includes generation, by the ML-based speech synthesis model, of synthesized versions of the input speech signal.

2. The device of claim 1, wherein the speech encoder further includes a codebook coupled to the ML-based speech synthesis model.

3. The device of claim 2, wherein the codebook is a fixed codebook.

4. The device of claim 2, wherein the codebook is an adaptive codebook.

5. The device of claim 2, further comprising a ML-based codebook generation model configured to generate the codebook.

6. The device of claim 1, wherein the ML-based speech synthesis model includes a neural network speech synthesizer.

7. The device of claim 1, further comprising a codebook that includes a plurality of codebook entries, and wherein the speech encoder is configured to perform the analysis- by-synthesis operation including: generate, at the ML-based speech synthesis model, a plurality of synthesized versions of the input speech signal, each version based on a set of features of the input speech signal and a codebook entry from the codebook; select one of the plurality of synthesized versions of the input speech signal; select the codebook entry used to generate the selected one of the plurality of synthesized versions of the input speech signal; andinclude an indication of the selected codebook entry as part of an encoded speech signal.

8. The device of claim 1, wherein the speech encoder is configured to, during the analysis- by-synthesis operation: generate, at the ML-based speech synthesis model, one or more of the synthesized versions of the input speech signal based on a set of features of the input speech signal and further based on one or more codebook entries from one or more codebooks; generate, for each of the one or more of the synthesized versions of the input speech signal, an associated error metric based on comparison of the one or more synthesized versions of the input speech signal to the input speech signal; and select a particular codebook entry based on the error metrics, wherein a bitstream output of the speech encoder includes an indication of the selected codebook entry.

9. The device of claim 8, wherein a synthesized version of the input speech signal is based on an output of a neural network speech synthesizer that receives a representation of the set of features and the particular codebook entry as inputs.

10. The device of claim 9, wherein the synthesized version of the input speech signal corresponds to the output of the neural network speech synthesizer combined with one or both of: the particular codebook entry; or an output of a filter that receives the particular codebook entry as an excitation signal.

11. The device of claim 8, wherein the generation of each of the one or more of the synthesized versions of the input speech signal and the generation of the associated error metrics are performed in an iterative loop over codebook entry indices.

12. The device of claim 1, wherein the ML-based speech synthesis model includes a neural network filter estimator configured to process a set of features of the input speechsignal to generate filter parameters for one or more filters of the ML-based speech synthesis model.

13. The device of claim 12, wherein the one or more filters of the ML-based speech synthesis model include a harmonic filter configured to process a pulse train based on a pitch lag of the input speech signal, a noise filter configured to process a random noise signal, or a codebook filter configured to process codebook entries of a stochastic codebook, wherein the stochastic codebook is coupled to the ML-based speech synthesis model.

14. The device of claim 12, further comprising a combiner configured to combine outputs of the one or more filters to generate a synthesized version of the input speech signal.

15. The device of claim 1, further comprising a microphone configured to provide the input speech signal.

16. The device of claim 1, further comprising one or more processors that include the speech encoder.

17. The device of claim 16, further comprising a modem coupled to the one or more processors and configured to send a bitstream output that includes an indication of a codebook entry selection from the analysis-by-synthesis operation, wherein the codebook entry is related to a codebook coupled to the ML-based speech synthesis model.

18. The device of claim 16, wherein the one or more processors are integrated in a headset device, the headset device further including one or more microphones configured to provide the input speech signal.

19. The device of claim 16, wherein the one or more processors are integrated in at least one of a mobile phone, a tablet computer device, or a wearable electronic device.

20. The device of claim 16, wherein the one or more processors are integrated in a vehicle, the vehicle further including one or more microphones configured to provide theinput speech signal.

21. The device of claim 16, wherein the one or more processors are included in an integrated circuit.

22. A method comprising: obtaining, at a speech encoder, an input speech signal; and performing, at the speech encoder, an analysis-by-synthesis operation that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal.

23. The method of claim 22, wherein the speech encoder includes a codebook coupled to the ML-based speech synthesis model.

24. The method of claim 22, wherein the analysis-by-synthesis operation includes: generating, at the ML-based speech synthesis model, one or more of the synthesized versions of the input speech signal based on a set of features of the input speech signal and further based on one or more codebook entries from one or more codebooks; generating, for each of the one or more of the synthesized versions of the input speech signal, an associated error metric based on comparison of the one or more synthesized versions of the input speech signal to the input speech signal; and selecting a particular codebook entry based on the error metrics, wherein a bitstream output of the speech encoder includes an indication of the selected codebook entry.

25. The method of claim 24, wherein a synthesized version of the input speech signal is based on an output of a neural network speech synthesizer that receives a representation of the set of features and the particular codebook entry as inputs.

26. The method of claim 25, wherein the synthesized version of the input speech signal corresponds to the output of the neural network speech synthesizer combined with one or both of:the particular codebook entry; or an output of a filter that receives the particular codebook entry as an excitation signal.

27. The method of claim 24, wherein the generating of each of the one or more of the synthesized versions of the input speech signal and the generating of the associated error metrics are performed in an iterative loop over codebook entry indices.

28. The method of claim 22, further comprising sending a bitstream output that includes an indication of a codebook entry selection from the analysis-by-synthesis operation, wherein the codebook entry is related to a codebook coupled to the ML-based speech synthesis model.

29. A non-transitory computer-readable medium including instructions that, when executed by one or more processors that include a speech encoder, cause the one or more processors to: obtain an input speech signal; and perform an analysis-by-synthesis operation of the speech encoder that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal.

30. An apparatus comprising: means for obtaining an input speech signal; and means for performing an analysis-by-synthesis operation that includes generation, by a machine learning (ML)-based speech synthesis model, of synthesized versions of the input speech signal.

Citation Information

Patent Citations

  • Artificial intelligence based audio coding

    US11437050B2

  • Audio coding using machine learning based linear filters and non-linear neural sources

    WO2023064735A1