Content-based switchable audio codec

By using a content-based switchable decoder system that combines machine learning and waveform matching encoders, dynamically selects the encoder and uses hybrid techniques and switching hysteresis, the problem of high-quality, low-bit-rate encoding of various audio content in resource-constrained environments is solved, and the distortion introduced by codec switching is reduced.

CN122342004APending Publication Date: 2026-07-03QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480076838.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-12-02
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing technologies struggle to provide high-quality, low-bit-rate audio compression in resource-constrained use cases, especially for diverse audio content, and codec switching can introduce audio distortion.

Method used

A content-based switchable decoder system is adopted, which combines a machine learning audio encoder and a waveform matching audio encoder with an audio classifier and a controller to dynamically select the appropriate encoder to encode different types of audio segments, and uses mixing technology and switching hysteresis to reduce distortion.

Benefits of technology

It achieves high-quality, low-bitrate encoding of various audio content under resource constraints, while reducing audio distortion introduced by codec switching and improving the overall quality of audio reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122342004A_ABST
    Figure CN122342004A_ABST
Patent Text Reader

Abstract

An apparatus includes a machine learning audio encoder and a waveform matching audio encoder. The apparatus includes a controller configured to input segments of audio data into the machine learning audio encoder, the waveform matching audio encoder, or both, based on classifications associated with segments of audio data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of priority to jointly owned U.S. non-provisional patent application No. 18 / 538,006, filed on December 13, 2023, the entire contents of which are expressly incorporated herein by reference. Technical Field

[0003] This disclosure relates in general to content-based switchable audio codecs. Background Technology

[0004] Technological advancements have led to smaller and more powerful computing devices. For example, a wide variety of portable personal computing devices exist today, including small, lightweight, and easily portable cordless phones (such as mobile and smartphones, tablets, and laptops). These devices can transmit voice and data packets over wireless networks. Furthermore, many of these devices incorporate additional functionality, such as digital still cameras, digital camcorders, digital recorders, and audio file players. Moreover, these devices can process executable instructions, including software applications such as web browser applications that can be used to access the internet. Accordingly, these devices can include significant computing power.

[0005] Many common applications of such devices revolve around media (e.g., audio, video, games, extended reality, etc.), such as those involving the capture, transmission, and / or reproduction of media content. Taking audio data as an example, the digitization, storage, and transmission of audio data are challenging because high-fidelity sound reproduction is often desired (e.g., to improve user experience), but improving sound reproduction fidelity may require using more bits to represent audio content and / or increasing the computational complexity of processing audio data. Increased computational complexity requires more power, more processing resources, more memory, or all three. Increasing the number of bits used to represent audio content increases the bandwidth required to transmit audio data and / or the memory required to store it.

[0006] Encoding schemes are frequently used to process audio data to reduce the number of bits required to represent specific audio content. Many encoding techniques preserve sound reproduction fidelity while reducing the number of bits needed to represent audio content, but such techniques introduce additional computational complexity. Therefore, encoding audio data is challenging in resource-constrained use cases, such as on battery-powered mobile computing devices. Summary of the Invention

[0007] According to one embodiment of this disclosure, an apparatus includes a machine learning audio encoder and a waveform matching audio encoder. The apparatus includes a controller configured to input a segment of audio data into the machine learning audio encoder, the waveform matching audio encoder, or both, based on a classification associated with that segment.

[0008] According to another specific embodiment of this disclosure, a method includes obtaining, by one or more processors, an indication of the type of audio content associated with a segment of audio data. The method includes selectively feeding the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both, based on the indication.

[0009] According to another embodiment of this disclosure, a non-transitory computer-readable medium stores instructions that can be executed by one or more processors to cause the one or more processors to obtain an indication of the type of audio content associated with a segment of audio data. These instructions further cause the one or more processors to selectively transmit the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both, based on the indication.

[0010] According to another embodiment of this disclosure, an apparatus includes components for acquiring an indication of the type of audio content associated with a segment of audio data. The apparatus includes components for selectively transmitting the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both, based on the indication.

[0011] Other aspects, advantages, and features of this disclosure will become apparent upon reading the entire application, which comprises the following sections: description of the drawings, detailed description, and claims. Attached Figure Description

[0012] Figure 1 This is a block diagram illustrating specific exemplary aspects of a system including a content-switchable audio decoder, based on some examples of this disclosure.

[0013] Figure 2 Based on some examples of this disclosure and Figure 1 A diagram illustrating the illustrative aspects of the system's associated operations.

[0014] Figure 3 Based on some examples of this disclosure and Figure 1 A diagram illustrating the illustrative aspects of the system's associated operations.

[0015] Figure 4 Based on some examples of this disclosure and Figure 1 A diagram illustrating the illustrative aspects of the system's associated operations.

[0016] Figure 5Based on some examples of this disclosure and Figure 1 A table illustrating the system-related operations.

[0017] Figure 6 Based on some examples of this disclosure and Figure 1 A diagram illustrating the illustrative aspects of the system's associated operations.

[0018] Figure 7 Based on some examples of this disclosure and Figure 1 A diagram illustrating the illustrative aspects of the system's associated operations.

[0019] Figure 8 This is a diagram illustrating an example of an integrated circuit operable to switch an audio codec based on the content of one or more segments of audio data, according to some examples of this disclosure.

[0020] Figure 9 This is an illustration of a mobile device operable to switch audio codecs based on the content of one or more segments of audio data, according to some examples of this disclosure.

[0021] Figure 10 This is an illustration of a headset operable, according to some examples of this disclosure, to switch audio codecs based on the content of one or more segments of audio data.

[0022] Figure 11 This is an illustration of a wearable electronic device operable, according to some examples of this disclosure, to switch audio codecs based on the content of one or more segments of audio data.

[0023] Figure 12 This is a diagram illustrating a voice-controlled speaker system operable, according to some examples of this disclosure, to switch audio codecs based on the content of one or more segments of audio data.

[0024] Figure 13 This is an illustration of a camera operable, according to some examples of this disclosure, to switch audio codecs based on the content of one or more segments of audio data.

[0025] Figure 14 This is an illustration of a head-mounted device (such as a virtual reality, mixed reality, or augmented reality head-mounted device) operable to switch audio codecs based on the content of one or more segments of audio data, according to some examples of this disclosure.

[0026] Figure 15 This is a diagram illustrating a first example of a vehicle operable, according to some examples of this disclosure, to switch audio codecs based on the content of one or more segments of audio data.

[0027] Figure 16 This is an illustration of a mixed reality or augmented reality glasses device operable to switch audio codecs based on the content of one or more segments of audio data, according to some examples of this disclosure.

[0028] Figure 17 This is an illustration of an earphone operable, according to some examples of this disclosure, to switch audio codecs based on the content of one or more segments of audio data.

[0029] Figure 18 This is an illustration of a hearing aid device operable, according to some examples of this disclosure, to switch audio codecs based on the content of one or more segments of audio data.

[0030] Figure 19 This is a diagram illustrating a second example of a vehicle operable, based on some examples of this disclosure, to switch audio codecs based on the content of one or more segments of audio data.

[0031] Figure 20 This is an illustration of a specific implementation of a method for switching audio codecs based on the content of one or more segments of audio data, according to some examples of this disclosure. The method can be... Figure 1 The system execution.

[0032] Figure 21 This is a block diagram illustrating a particular exemplary example of a device operable to switch audio codecs based on the content of one or more segments of audio data, according to some examples of this disclosure. Detailed Implementation

[0033] Machine learning-based audio codecs can be trained to encode audio data representing specific types of audio content, such as speech, in a more efficient manner than traditional codecs (in terms of compression ratio, such as the number of bits used to represent a specific segment of audio data) without loss of quality. An example of such a machine learning-based audio codec is the Lyra codec. The Lyra codec uses a recurrent neural network to quantize audio data representing speech at a low bit rate and a generative neural network to decode the quantized audio data to generate an output representing speech. The Lyra codec achieves high compression ratios for audio representing speech and provides high-quality decoded speech output. However, the Lyra codec and other similar codecs sacrifice versatility to achieve this bit-efficiency and high-quality speech reproduction. For example, while the Lyra codec performs well when providing audio data representing speech, it cannot achieve the same performance when providing audio data representing other types of audio content, such as music.

[0034] Other more general-purpose machine learning-based audio codecs, such as the SoundStream codec, can provide high-quality audio reproduction, but at the cost of lower compression ratios compared to more specialized audio codecs that target specific audio types (e.g., Lyra, which targets speech data). Furthermore, such general-purpose machine learning-based audio codecs tend to be larger (e.g., in terms of model parameters, and correspondingly, in terms of memory footprint) and more complex (e.g., more resource-intensive to use), making their use challenging in resource-constrained use cases such as in-vehicle mobile devices.

[0035] Using machine learning codecs is challenging for audio streams containing various types of audio content (e.g., speech, noise, music, etc.) as well as audio streams with unknown audio content types beforehand. This is because the quality of the generated audio output cannot be guaranteed unless one relies on a general-purpose machine learning-based audio codec that is less efficient and has a higher memory footprint. Therefore, providing high-quality, low-bitrate audio compression under resource constraints is problematic.

[0036] Content-based switchable decoder systems, as described herein (also referred to herein as “content-based switchable decoder systems”), address the aforementioned problems associated with providing high-quality, low-bit-rate audio compression for resource-constrained use cases and various types of audio data. A content-based switchable decoder system includes multiple audio encoders and a controller that selectively feeds audio data segments to one or more audio encoders based on the content represented in each segment (e.g., the type of audio data). The multiple audio encoders may include, for example, machine learning audio encoders well-suited for encoding specific types of audio data, such as speech. In this example, a machine learning audio encoder may provide high compression ratios and high-quality audio reproduction for segments that include speech. The multiple audio encoders may include at least one audio encoder that is more general to provide high-quality audio reproduction for a variety of types of audio content, such as a waveform matching audio encoder. In this context, a waveform matching audio encoder refers to a decoder that attempts to represent a segment in a way that can reproduce the entire waveform of the segment of audio data (e.g., compared to a decoder that attempts to reproduce only the speech component of the segment).

[0037] The controller of a content-based switchable decoder system can provide input to a machine learning audio encoder that is well-suited to encoding the target audio type, including segments of audio data comprising a target audio type (e.g., speech, wind noise, noise, music, silence, etc.), and can provide input to a waveform matching audio encoder for other segments (e.g., segments comprising non-target audio types). The content-based switchable decoder system may include an audio classifier configured to provide the controller with an indication of whether each segment of audio data comprises the target audio or non-target audio. For example, the classifier may be a machine learning-based classifier configured to generate a classification output associated with the input segments of audio data. The classification output may be binary (e.g., a first value, such as 1, to indicate that the segment comprises the target audio type; and a second value, such as 0, to indicate that the segment does not comprise the target audio type). Alternatively, the classification output may indicate one of several categories associated with the segment (e.g., speech, wind noise, music, silence, etc.). The classifier may use machine learning techniques, procedural techniques (such as voice activity detection), or a combination thereof. For illustration, multiple classification techniques can be used, and voting or other selection mechanisms can be used to generate instructions for the controller based on various classification results from multiple classification techniques.

[0038] Content-based switchable decoder systems can therefore achieve high-quality, low-bit-rate representation of target audio data (e.g., speech) without sacrificing the quality of audio data segments including those of non-target audio types. Furthermore, the audio encoders used (e.g., target machine learning audio encoders and general waveform matching audio encoders) can have a smaller memory footprint and are less resource-intensive compared to general machine learning-based audio encoders. Therefore, content-based switchable decoder systems can be used in resource-constrained use cases.

[0039] One problem that can arise from switching between codecs is that such codec switching can introduce audio distortion, thereby reducing the overall quality of the reproduced audio output. The content-based switchable decoder system disclosed herein can use various switching techniques to mitigate the introduction of such distortion. For example, in some embodiments, when switching between decoders, the content-based switchable decoder system can provide one or more segments of audio data to both a machine learning audio encoder and a waveform matching audio encoder, and the output of each decoder can be transmitted to a decoder system. In such embodiments, the decoder system can use a machine learning audio decoder to decode the data from the machine learning audio encoder to generate a first decoded representation of the segment, and use a waveform matching decoder to decode the data from the waveform matching audio encoder to generate a second decoded representation of the segment. The decoder system can combine portions of the first and second decoded representations to gradually decrease from one decoder while gradually increasing from the other. To illustrate, when switching from a machine learning audio encoder to a waveform matching audio encoder, the decoder system can gradually attenuate (e.g., decrease) the first decoded representation while simultaneously gradually emphasizing (e.g., increase) the second decoded representation to generate output audio. Combining the decoded output data from two different decoders in this way will mix the audio in a way that reduces audio distortion caused by switching codecs.

[0040] As in the previous example, combining decoded output data from two different decoders increases the bit rate of data transmitted between the content-based switchable decoder system and the decoder system because two representations of at least one audio data segment are being transmitted to facilitate mixing. Additionally, providing a single segment to two different decoders uses additional resources (e.g., processor time and power) at the content-based switchable decoder system. In some implementations, these problems are addressed by providing each segment of audio data to only one audio encoder in the audio encoder of the content-based switchable decoder system. In such implementations, the decoder system uses extrapolation techniques to mix adjacent segments from different decoders. For example, mixing techniques (such as those used for frame error concealment) can be used to slow the transition between codecs to reduce audio distortion introduced by switching codecs. Such implementations do not increase the bit rate of data transmitted between the content-based switchable decoder system and the decoder system because mixing is performed using only one representation of each audio data segment.

[0041] In some implementations, the controller uses switching hysteresis to reduce audio distortion introduced by switching codecs. For example, when switching from a machine learning audio encoder to a waveform matching audio encoder, the controller can switch without delay. Conversely, when switching from a waveform matching audio encoder to a machine learning audio encoder, the controller can introduce a switching delay based on the content of the audio data segment. Waveform matching audio encoders are generally capable of encoding various types of audio content without significantly reducing fidelity; however, it is common for relatively short segments of non-target data to cause significant (e.g., audible) distortion in machine learning audio encoders. Additionally, switching between segments containing certain sounds to a machine learning audio encoder may result in more perceptible distortion compared to switching between segments representing other sounds or silence. For example, distortion can be introduced by switching to a machine learning audio encoder in the middle of a vowel sound. Therefore, the controller can delay the switch from a waveform matching audio encoder to a machine learning audio encoder until the vowel sound ends, until a period of silence is reached, or until a low-energy segment for decoding is received.

[0042] In some implementations, when switching from a first decoder to a second decoder, the controller populates the decoder state data of the second decoder based on data from the first decoder. For example, when switching from a machine learning-based decoder to a waveform-matched audio encoder, the controller may populate the activation signal memory of the waveform-matched audio encoder based on data from the machine learning decoder, providing a smoother sequence of tone pulses instead of initializing the activation signal memory of the waveform-matched audio encoder with default data (e.g., zero).

[0043] Specific aspects of this disclosure are described below with reference to the accompanying drawings. In this description, common features are designated by common reference numerals. As used herein, various terms are used only for the purpose of describing particular embodiments and are not intended to limit the scope of the embodiments. For example, the singular forms “a,” “an,” and “the” are intended to also include the plural forms unless the context clearly indicates otherwise. Furthermore, some features described herein are singular in some embodiments and plural in others. For example, Figure 1 It describes a system that includes one or more processors ( Figure 1 The device 102 (of which the “processor 190” is used) indicates that in some embodiments, device 102 includes a single processor 190, while in other embodiments, device 102 includes multiple processors 190. For ease of reference herein, such features are generally introduced as “one or more” features and are subsequently referred to in the singular or optional plural form (such as indicated by “multiple”), unless the aspect described relates to multiples of features.

[0044] In some figures, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference numerals are used for each feature, and these different instances are distinguished by adding letters to the reference numerals. Reference numerals are used without distinguishing letters when a feature is referenced herein as a group or a type of feature (e.g., when a specific feature among these features is not referenced). However, reference numerals are used with distinguishing letters when a specific feature among multiple features of the same type is mentioned herein. For example, see reference... Figure 1 The figure illustrates several segments, which are associated with reference numerals 114A and 114B. When referring to a specific segment among these segments (such as segment 114A), the distinguishing letter "A" is used. However, when referring to any segment among these segments or to these segments as a group, reference numeral 114 without the distinguishing letter is used.

[0045] As used herein, the term "comprising" may be used interchangeably with "including". Additionally, the term "wherein" may be used interchangeably with "where". As used herein, "exemplary" indicates an example, specific implementation, and / or aspect, and should not be construed as restrictive or indicating a preference or preferred implementation. As used herein, ordinal terms used to modify elements (such as structures, components, operations, etc.) (e.g., "first", "second", "third", etc.) do not themselves indicate any priority or order of that element relative to another element, but merely distinguish that element from another element with the same name (but using ordinal terms). As used herein, the term "set" refers to one or more specific elements within a set of specific elements, while the term "multiple" refers to multiple (e.g., two or more) specific elements.

[0046] As used herein, “coupling” can include “communicationally coupled,” “electrically coupled,” or “physically coupled,” and may also (or alternatively) include any combination thereof. Two devices (or components) may be coupled directly or indirectly (e.g., communicationally coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof). As an illustrative, non-limiting example, two electrically coupled devices (or components) may be included in the same device or in different devices and may be connected via electronics, one or more connectors, or inductive coupling. In some specific implementations, two communicationally coupled (such as electrical communication) devices (or components) may transmit and receive signals (e.g., digital or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein, “direct coupling” can include two devices coupled without intermediate components (e.g., communicationally coupled, electrically coupled, or physically coupled).

[0047] In this disclosure, terms such as “acquire,” “determine,” “calculate,” “estimate,” “shift,” and “adjust” are used to describe how one or more operations are performed. It should be noted that such terms should not be construed as restrictive, and similar operations can be performed using other techniques. Additionally, as mentioned herein, “acquire,” “generate,” “calculate,” “estimate,” “use,” “select,” “access,” and “determine” are used interchangeably. For example, “acquire,” “generate,” “calculate,” “estimate,” or “determine” a parameter (or signal) can refer to actively generating, estimating, calculating, or determining that parameter (or signal), or it can refer to using, selecting, or accessing a parameter (or signal) such as one already generated by another component or device.

[0048] As used herein, the term “machine learning” should be understood to have any of its usual and conventional meanings within the fields of computer science and data science. Such meanings include, for example, processes or techniques by which one or more computers can learn to perform certain operations or functions without being explicitly programmed to do so. As a typical example, machine learning can be used to enable one or more computers to analyze data to identify patterns in the data and generate results based on that analysis. For some types of machine learning, the generated results include data that indicates the underlying structure or patterns of the data itself. For example, such techniques include so-called “clustering” techniques, which identify clusters (e.g., groupings of data elements).

[0049] For some types of machine learning, the resulting output includes a data model (also known as a "machine learning model" or simply a "model"). Typically, a model is generated using a first dataset to facilitate analysis on a second dataset. For example, the first portion of a large dataset can be used to generate a model that can then be used to analyze the remaining portion of the large dataset. As another example, a set of historical data can be used to generate a model that can be used to analyze future data.

[0050] Because a model can be used to evaluate a dataset different from the data used to generate the model, it can be considered a type of software (e.g., instructions, parameters, or both) automatically generated by a computer during the machine learning process. Therefore, the model can be transferable (e.g., it can be generated at a first computer and subsequently moved to a second computer for further training, use, or both). Additionally, the model can be combined with one or more other models to perform a desired analysis. For example, first data can be provided as input to a first model to generate first model output data, and the first model output data (alone, with the first data, or with other data) can be provided as input to a second model to generate second model output data indicative of the results of the desired analysis. Depending on the analysis and data involved, different combinations of models can be used to generate such results. In some examples, multiple models can provide model outputs that are input to a single model. In some examples, a single model provides model outputs as input to multiple models.

[0051] Examples of machine learning models include, but are not limited to, perceptrons, neural networks, support vector machines, regression models, decision trees, Bayesian models, Boltzmann machines, adaptive neural fuzzy inference systems, and combinations, sets, and variations of these and other types of models. Variations of neural networks include, for example, but not limited to, prototype networks, autoencoders, transformers, self-attention networks, convolutional neural networks, deep neural networks, deep belief networks, etc. Variations of decision trees include, for example, but not limited to, random forests, boosted decision trees, etc.

[0052] Since machine learning models are generated by computers based on input data, they can be discussed within at least two distinct time windows: the creation / training phase and the runtime phase. During the creation / training phase, a model is created, trained, adapted, validated, or otherwise configured by a computer based on input data (often referred to as "training data" during the creation / training phase). It's important to note that a trained model corresponds to software that has been generated and / or refined during the creation / training phase to perform a specific operation (such as classification, prediction, encoding, or other data analysis or data synthesis operations). During the runtime phase (or "inference" phase), the model is used to analyze the input data to generate model outputs. The content of the model outputs depends on the type of model. For example, as a non-limiting example, a model can be trained to perform a classification task or a regression task. In some implementations, the model may be updated continuously, periodically, or occasionally, in which case training time and runtime may be interleaved, or one form of the model may be available for inference while a copy is updated, and subsequently, the updated copy may be deployed for inference.

[0053] In some implementations, machine learning techniques are used to train (or retrain) a previously generated model. In this context, "training" refers to adapting a model or its parameters to a specific dataset. Unless otherwise clearly understood from the specific context, the term "training" as used herein includes "retraining" or refining a model for a specific dataset. For example, training may include so-called "transfer learning." In transfer learning, a base model is trained using a general or typical dataset, and subsequently refined (e.g., retrained or further trained) the base model using a more specific dataset.

[0054] The dataset used during training is called the "training dataset" or simply "training data." The dataset can be labeled or unlabeled. "Labeled data" refers to data that has been assigned classification labels indicating the groups or categories associated with the data, and "unlabeled data" refers to unlabeled data. Typically, "supervised machine learning processes" use labeled data to train machine learning models, while "unsupervised machine learning processes" use unlabeled data; however, it should be understood that the labels associated with the data are simply another data element that can be used in any appropriate machine learning process. For example, many clustering operations can be performed using unlabeled data; however, such clustering operations can use labeled data by ignoring the labels assigned to the data or by treating the labels in the same way as other data elements.

[0055] Training a model on a training dataset typically involves modifying the model's parameters with the goal of making the model's output possess specific characteristics based on the data input to the model. To distinguish it from model generation operations, model training may be referred to as optimization or optimization training in this paper. In this context, "optimization" refers to improving a metric, not necessarily finding an ideal value for that metric (e.g., a global maximum or minimum). Examples of optimization trainers include, but are not limited to, backpropagation trainers, derivative-free optimizers (DFO), and extreme learning machines (ELM). As an example of training a model, during supervised training of a neural network, input data samples are associated with labels. When input data samples are fed to the model, the model generates output data, comparing that output data with the labels associated with the input data samples to generate error values. The model's parameters are modified to attempt to reduce (e.g., optimize) the error values. As another example of training a model, during unsupervised training of an autoencoder, data samples are fed as input to the autoencoder, and the autoencoder reduces the dimensionality of the data samples (a lossy operation) and attempts to reconstruct the data samples into output data. In this example, the output data is compared with the input data samples to generate the reconstruction loss, and the parameters of the autoencoder are modified to attempt to reduce (e.g., optimize) the reconstruction loss.

[0056] Figure 1 A block diagram of a system 100 illustrating aspects of a content-switchable codec is shown. System 100 includes a device 102 that includes a content-switchable decoder system 140. The content-switchable decoder system 140 includes two or more audio encoders of different types and a controller 142 configured to selectively route segments 114 of audio data 112 to one or more of the audio encoders. Specifically, the controller 142 selectively routes segments 114 of audio data 112 based on the audio content of each segment 114. This content-selective routing of segments 114 enables the content-switchable decoder system 140 to encode audio data 112 efficiently (in terms of both the computational resources used and the bit rate of the encoded data) in a manner that provides high-quality audio reproduction.

[0057] exist Figure 1 In this embodiment, device 102 is coupled to one or more audio sources 110. In some implementations, one or more audio sources 110 are integrated within device 102. For example, audio source 110 may include media files stored in memory 120 of the device. As another example, audio source 110 may include one or more microphones, which are integrated within or coupled to the device 102.

[0058] Audio data 112 includes multiple segments 114. Figure 1In the text, segment 114 includes segment 114A and segment 114B, which represent... Figure 1 The audio data 112 consists of a series of adjacent segments 114. Each segment 114 represents a portion of a time window of the audio data 112. Optionally, in some embodiments, adjacent segments in the series (e.g., segments 114A and 114B) partially overlap. In some embodiments, each segment 114 includes pulse-code-modulated (PCM) data.

[0059] exist Figure 1 In the content switchable decoder system 140, two or more audio encoders include at least a machine learning audio encoder 144 and a waveform matching audio encoder 146. The machine learning audio encoder 144 is a machine learning model (or a collection of machine learning models) configured and trained to generate low-bit-rate, high-quality representations of segments of audio content including a target type, such as speech. For example, the machine learning audio encoder 144 may include a neural homomorphic vocoder (NHV). In some embodiments, the NHV is based on a bimodal activation model of the human vocal cords, which enables the NHV to encode audio data representing speech at high fidelity and low bit rate. As an example, the NHV is configured to extract features from segment 114 representing audio data 112 and provide those features as input to a filter estimator. The filter estimator includes a neural network configured and trained to generate filter parameters for a noise filter (e.g., a first linear time-varying (LTV) filter) and a harmonic filter (e.g., a second LTV filter). The noise filter is configured to modify a random noise signal based on the noise filter parameters from the filter estimator to generate data representing silent speech components. Harmonic filters are configured to modify pulse trains representing the pitch of audio segments based on harmonic filter parameters from a filter estimator to generate data representing speech components. Filter parameters (e.g., noise filter parameters and harmonic filter parameters) and possibly other data are used to generate a representation of the segment. In other examples, the machine learning audio encoder 144 may include a low-complexity parametric neural network (LCPNet), a WaveNet decoder, a Lyra decoder, an EnCodec decoder, or another machine learning-based encoder specifically designed for audio content of the target content type.

[0060] Waveform matching audio encoder 146 is a procedural decoder that attempts to represent segment 114 of audio data 112 in a manner that allows the entire waveform of the segment to be reproduced regardless of the audio content represented by the waveform. For example, to limit the computational resources used by machine learning audio encoder 144, machine learning audio encoder 144 is optimized (e.g., configured and trained) to encode audio data including audio content of the target type (such as speech, music, etc.) at a low bit rate. In this example, machine learning audio encoder 144 may struggle to encode audio data that does not include audio content of the target type with the same fidelity and the same low bit rate. In contrast, waveform matching audio encoder 146 is a general-purpose decoder that can encode any audio content with approximately the same level of audio reproduction fidelity; however, to achieve this wide range of encoding, waveform matching audio encoder 146 has a higher bit rate than machine learning audio encoder 144 and can also use more computational resources to perform encoding. For example, a machine learning audio encoder 144 is configured to encode an input segment 114 to generate an output 154, the output including a first number of bits to represent the segment 114, and a waveform matching audio encoder 146 is configured to encode an input segment 114 to generate an output 156, the output including a second number of bits to represent the segment 114, wherein the first number is less than the second number.

[0061] Controller 142 is configured to input segment 114 to a machine learning audio encoder 144, a waveform matching audio encoder 146, or both, based on a classification associated with segment 114 of audio data 112. For example, content-switchable decoder system 140 includes an audio classifier 148 configured to generate an indicator 150 indicating a classification associated with a segment of segment 114, and controller 142 selects one or more audio encoders from audio encoders 144, 146 to process segment 114 based on indicator 150. For illustration, audio classifier 148 may be configured to generate indicator 150 indicating whether segment 114 represents a specific type of audio content (e.g., the target audio type of machine learning audio encoder 144). The target audio type may include, for example, speech, music, non-speech sounds, etc. Audio classifier 148 may include a machine learning model (e.g., a classification model such as a decision tree, neural network, support vector machine, etc.). Alternatively, audio classifier 148 may use non-machine learning techniques, such as speech activity detection.

[0062] In some cases, switching between the machine learning audio encoder 144 and the waveform matching audio encoder 146 can introduce distortion into the audio reproduced by the outputs 154, 156 of the content-switched decoder system 140. For example, when the audio data 112 includes speech, some speech may span more than one segment 114. In this example, switching the encoder during such speech sounds can cause the reproduced audio data 188 to include distortion from the switching that reduces the clarity of the speech and / or leads to a degraded user experience.

[0063] To limit the introduction of such distortion, controller 142 may optionally be configured to use one or more distortion mitigation techniques from a variety of distortion mitigation techniques. One example of a technique for limiting distortion introduction is to provide one or more segments 114 to both machine learning audio encoder 144 and waveform matching audio encoder 146. For example, in response to determining which audio encoder to which segment 114 of audio data 112 is provided to, controller 142 may provide at least one segment 114 of audio data 112 to both machine learning audio encoder 144 and waveform matching audio encoder 146. In this example, device 102, one or more remote devices 180, or both may use a mixing technique to combine a portion of the reproduced audio data 188 of the output 154 representing segment 114 based on machine learning audio encoder 144, and a portion of the reproduced audio data 188 of the output 156 representing segment 114 based on waveform matching audio encoder 146, to generate a mixed version of segment 114. Figure 3 and Figure 4 An example of this type of mixing is shown.

[0064] Another example of a technique for limiting distortion introduction is using encoder state data 160 of a first audio encoder (e.g., machine learning audio encoder 144 or waveform matching audio encoder 146) to initialize another audio encoder when switching from the first audio encoder to another. For example, when different audio encoders are selected to process two consecutive segments of audio data 112 (e.g., segment 114A and segment 114B), encoder state data 160 obtained from processing segment 114A (e.g., the first segment of the two consecutive segments) is used to process segment 114B (e.g., the second segment of the two consecutive segments). For example, in some embodiments, distortion can be reduced by populating the activation signal memory of waveform matching audio encoder 146 based on information from machine learning audio encoder 144. Such specific implementations allow waveform matching audio encoder 146 to begin generating a smoother evolution of the tonal pulse sequence, rather than starting from zero or some other default encoder state data 160.

[0065] Another example of a technique for limiting distortion introduction is to delay the switching between audio encoders based on the audio content of the segments. For illustration, in some implementations, controller 142 is configured to use a first delay when transitioning to inputting a segment to waveform-matching audio encoder 146, and is configured to use a second delay when transitioning to inputting a segment to machine learning audio encoder 144. In such implementations, the first delay differs from the second delay. For example, the first delay may be fixed, and the second delay may be variable and selected based on the content of segment 114. For illustration, when segment 114A includes speech (or other target audio data) and segment 114B (e.g., the segment immediately following segment 114A) includes non-speech (or other non-target audio data), controller 142 transmits segment 114A to machine learning audio encoder 144 and segment 114B to waveform-matching audio encoder 146 (e.g., with a zero-segment delay). In this example, a zero-segment delay is used because even encoding only a small number of non-speech signals of some types of non-speech signals using machine learning audio encoder 144 can result in significant audio distortion. Conversely, when switching in the opposite direction (e.g., from waveform matching audio encoder 146 to machine learning audio encoder 144), a delay based on the content of the audio data 112 can be used to avoid switching distortion. For example, controller 142 could delay the switch to machine learning audio encoder 144 until the end of a speech sound (e.g., a vowel sound) is detected or until a low-energy segment 114 (e.g., a segment indicating silence) is detected. Waveform matching audio encoder 146 can encode both target and non-target audio equally well, but at the cost of using more computational and communication resources. Therefore, delaying the transition from waveform matching audio encoder 146 to machine learning audio encoder 144 is less efficient than an immediate switch, but avoids introducing audio distortion.

[0066] Various combinations of the above techniques can be used together. For example, in some embodiments, controller 142 is configured to select one audio encoder from the audio encoders to process each corresponding segment 114 of the audio data. In such embodiments, a switching delay, the use of encoder state data from one encoder to initialize another encoder, or both, can be used to limit distortion. In other embodiments, controller 142 is configured to select two or more audio encoders from the audio encoders to encode a specific segment 114 of the audio data 112, in at least some cases. In some such embodiments, controller 142 may also use encoder state data 160 from one encoder to initialize another encoder to further limit distortion.

[0067] exist Figure 1In the illustrated example, device 102 includes a modem 170 coupled to processor 190 and configured to represent the output 154 of machine learning audio encoder 144, the output 156 of waveform matching audio encoder 146, or both, in bitstream 172. For example, bitstream 172 may be transmitted to remote device 180 via a communication channel. In this example, remote device 180 includes a decoder system 182 that includes decoders complementary to the audio encoders of device 102. For example, remote device 180 includes a machine learning audio decoder 184 configured to process data representing the output 154 of machine learning audio encoder 144 to generate audio data representing segment 114 input to machine learning audio encoder 144. Similarly, remote device 180 includes a waveform matching audio decoder 186 configured to process data representing the output 156 of waveform matching audio encoder 146 to generate audio data representing segment 114 input to waveform matching audio encoder 146. In some implementations, bitstream 172 includes information indicating which audio encoder of device 102 is used to decode each portion of bitstream 130 so that remote device 180 can select the corresponding audio decoder.

[0068] although Figure 1 The content-switchable decoder system 140 is illustrated as including two audio encoders 144, 146, but in some embodiments, the content-switchable decoder system 140 includes more than two audio encoders. In such embodiments, the controller 142 is configured to selectively transmit each segment of audio data 112 from segment 114 to any one or more audio encoders, depending on the content of each segment 114 and the specific distortion reduction technique employed.

[0069] In some specific implementations, device 102 corresponds to or is included in one of various types of devices. In an exemplary example, processor 190 is integrated into a headset device, as shown in reference [reference needed]. Figure 10 Further described. In other examples, processor 190 is integrated as described in reference . Figure 9 The described mobile phone or tablet computer device, as shown in the reference Figure 11 The wearable electronic devices described, as in the reference Figure 12 The described voice control speaker system, as in the reference Figure 13 The described camera equipment, or as referenced Figure 14 The virtual reality, mixed reality, or augmented reality head-mounted device described, as in the reference Figure 16 The described mixed reality or augmented reality glasses device, such as the reference Figure 17The described in-ear headphones, or as referenced Figure 18 In at least one of the described hearing aid devices. In another exemplary example, the processor 190 is integrated into a vehicle, such as a reference horoscope. Figure 15 and Figure 19 Further description.

[0070] Figures 2 to 4 This illustrates aspects of content-switched codec operation. Figures 2 to 4 Each figure in the diagram includes Figure 1 The content-switching decoder system 140 includes a machine learning (ML) audio encoder 144 and a waveform matching (WM) audio encoder 146. Additionally, Figures 2 to 4 Each figure in the diagram includes Figure 1 The decoder system 182 includes a machine learning audio decoder 184 and a waveform matching audio decoder 186.

[0071] exist Figures 2 to 4 In each of the diagrams, the decoder system 182 is configured to generate a decoded segment (DS) based on data representing the coded segment (ES) from the content switchable decoder system 140. For example, in Figure 2 In this configuration, the machine learning audio decoder 184 is configured to generate a decoded segment 232 based on data from the representation encoded segment 212 received from the machine learning audio encoder 144. Similarly, the waveform matching audio decoder 186 is configured to generate a decoded segment 230 based on data from the representation encoded segment 210 received from the waveform matching audio encoder 146. Figures 2 to 4 Each illustration in the diagram shows an encoded segment transmitted in bit stream 172 via channel 220 (such as one or more wired or wireless transmissions); however, in other embodiments, the encoded segment may be stored in memory and subsequently retrieved for decoding.

[0072] Figure 2 The operation of a content-switched codec is illustrated, wherein the content-switched decoder system 140 is configured to provide each segment 114 to a single audio encoder. For example, based on an instruction 150 from the audio classifier 148, Figure 1 The controller 142 causes segments 114 of the audio data 112 to be transmitted to either the machine learning audio encoder 144 or the waveform matching audio encoder 146. For illustration, when the machine learning audio encoder 144 is associated with a target content type, the indicator 150 may indicate for each segment 114 whether the segment 114 includes audio data representing the target content type. The indicator 150 has a first value when the segment 114 includes audio data representing the target content type, and a second value when the segment includes audio data that does not represent the target content type.

[0073] exist Figure 2 In this sequence, encoded segment 210 includes one or more segments 114 encoded by waveform matching audio encoder 146, encoded segment 212 includes one or more segments 114 encoded by machine learning audio encoder 144, encoded segment 214 includes one or more segments 114 encoded by waveform matching audio encoder 146, and encoded segment 216 includes one or more segments 114 encoded by machine learning audio encoder 144. Encoded segments 210 to 216 represent a time series, wherein encoded segment 210 precedes encoded segment 212, encoded segment 212 precedes encoded segment 214, and encoded segment 214 precedes encoded segment 216. Encoded segments 210 to 216 may each include representations of... Figure 1 The encoded segment 210 may include data representing one or more segments of segment 114 (such as segment 114A), or it may include data representing more than one segment of segment 114 (such as both segment 114A and segment 114B).

[0074] As described above, in some embodiments, the machine learning audio encoder 144 is configured to encode segments of audio data using a first number of bits, and the waveform matching audio encoder 146 is configured to encode segments of audio data using a second number of bits, greater than the first number of bits. Therefore, if encoded segment 210 and encoded segment 212 each represent the same number of segments, then encoded segment 210 includes more bits than encoded segment 212.

[0075] The audio output based on decoded segments 230 to 236 may include distortion due to switching codecs between segments. For example, switching from a machine learning audio codec to a waveform matching audio codec between decoded segments 232 and 234 may introduce distortion in the audio output. Figure 1 The content-switchable decoder system 140 can use any of a number of techniques to limit the introduction of such distortion. For example, the content-switchable decoder system 140 can delay the switching from waveform-matched audio encoder 146 to machine learning audio encoder 144 until the vowel sound ends or until a low-energy segment ends. For illustration, encoded segment 210 can represent the end of data at the end of a vowel sound, and encoded segment 212 can represent the start of the next sound or silence in audio data 112.

[0076] As another example, when switching from waveform matching audio encoder 146 to machine learning audio encoder 144, the decoder state data of machine learning audio encoder 144 can be initialized based on information from waveform matching audio encoder 146, or vice versa. For illustration, when machine learning audio encoder 144 begins generating coded segment 212, decoder state data based on the encoding of coded segment 210 can be used to initialize machine learning audio encoder 144. Additionally or alternatively, when waveform matching audio encoder 146 begins generating coded segment 214, encoder state data based on the encoding of coded segment 212 can be used to initialize waveform matching audio encoder 146.

[0077] Figure 3 The examples shown are similar Figure 2 The example shown differs in that Figure 3 In this process, the decoder system 182 performs a mixing operation to reduce distortion caused by switching audio codecs. Figure 2 In the illustrated example, each segment is provided as input to an audio encoder 144 or 146 to generate encoded segments 210 to 216. Each encoded segment 210 to 216 is decoded by a corresponding audio decoder 184 or 186 to generate decoded segments 230 to 236, and the audio output is generated by playing the decoded segments 230 to 236. Figure 2 ,exist Figure 3 In this process, each segment is provided as input to an audio encoder 144 or 146 to generate encoded segments 210 to 216, and each encoded segment 210 to 216 is decoded by a corresponding audio decoder 184 or 186 to generate decoded segments 230 to 236. However, with Figure 2 Conversely, in step 3, portions of decoded segments 330 through 336 on either side of the transition are mixed to generate the audio output. For example, portion 340 of decoded segment 330 may be mixed with portion 342 of decoded segment 332. An estimation process, such as interpolation or extrapolation for frame erasure hiding, can be used to estimate portions 340 and 342. In this example, the audio output is based on decoded segment 330 far prior to transition point 344; however, as transition point 344 approaches, a weighted mix of audio data from portion 340 and audio data from portion 342 is performed, where the weighting gradually decreases the influence of portion 340 and gradually increases the influence of portion 342. Similar mixing can be performed near other transition points, such as transition points 346 and 348. Mixing audio content near transition points 344 through 348 reduces distortion caused by switching audio codecs. In some implementations, Figure 3 The illustrated hybridization can be combined with other distortion reduction techniques (such as reference) Figure 2 The described technology is combined and implemented.

[0078] Figure 4 The operation of a content-switched codec is illustrated, wherein a content-switched decoder system 140 is configured to provide segments 114 to more than one audio encoder to reduce distortion due to transitions between audio encoders. For example, when transitioning between audio encoders 144, 146, a controller 142 of the content-switched decoder system 140 causes at least one segment 114 of audio data 112 to be transmitted to both the machine learning audio encoder 144 and the waveform matching audio encoder 146. Therefore, encoded segments 410 and 412 overlap at least in an overlap region 420. In some embodiments, encoded segments 410 and 412 completely overlap (e.g., the same audio data is encoded by different audio encoders in audio encoders 144, 146). Similarly, encoded segments 412 and 414 overlap at least in an overlap region 422, and encoded segments 414 and 416 overlap at least in an overlap region 424.

[0079] exist Figure 4 In this example, decoder system 182 smooths the content of overlapping regions 460 to 464 of adjacent decoded segments 430 to 436 to generate audio output. For example, a portion 442 of decoded segment 432 corresponding to overlapping region 460 can be used to smooth a portion 440 of decoded segment 430 corresponding to overlapping region 460. In this example, the audio output is based on portions 440 and 442 of decoded segments 430 and 432 associated with overlapping region 460. To illustrate, to generate audio output associated with overlapping region 460, weighted smoothing is performed on audio data from portion 440 and audio data from portion 442, wherein the weighting gradually decreases the influence of portion 440 and gradually increases the influence of portion 442.

[0080] Figure 5 This is an example that can be derived from Figure 1 The device 102 is used to transmit data in a packet structure for transmitting encoded segments 114 representing audio data 112, as shown in Table 500. Table 500 includes a header column 502 indicating information included in the packet header in various cases and a payload column 504 indicating the payload size in various cases.

[0081] exist Figure 5 In the illustrated example, the two-bit encoder ID field in the packet header is used to identify which audio encoder was used to encode the segment represented in the packet's payload. For example, in Figure 5 In the encoder ID field, a value of 00 indicates the first audio encoder (such as...). Figure 1A machine learning audio encoder 144 is used to encode segments represented by data in a grouped payload. As another example, the value 10 in the encoder ID field indicates a second audio encoder (such as...). Figure 1 The waveform-matching audio encoder 146 is used to encode segments represented by data in the grouped payload. As another example, the value 11 in the encoder ID field indicates that both the first and second audio encoders are used to encode segments represented by data in the grouped payload. When both audio encoders are used to encode segments in the grouped payload, the decoder can determine the position of each coded segment in the payload based on the number of bits used to encode each segment. For example, a coded segment from the first audio encoder may include a first number of bits, and a coded segment from the second audio encoder may include a second number of bits (as indicated in payload column 504). Therefore, the decoder can detect that a bit following the first number of bits of the first coded segment in the payload corresponds to the first bit of the second coded segment.

[0082] The specific values ​​listed in Table 500 are merely illustrative. In other embodiments, different values ​​are used to indicate which audio encoder is used to encode the segment. Furthermore, in some embodiments, the encoder ID field may include a different number of bits.

[0083] Figure 6 and Figure 7 An example of a switching scheme for transitioning between audio encoders is illustrated. Specifically, Figure 6 An example is given in a scenario with a fixed delay (e.g., Figure 6 An example of switching between two encoders (with zero latency). Conversely, Figure 7 An example is given in the case of variable delay (e.g., based on...). Figure 7 An example of switching between two encoders (the delay of the content of one or more segments in the code).

[0084] Figure 6 Including showing that will be by Figure 1 The content of Figure 600 shows the amplitude 602 and frequency 604 of the sound encoded by the decoder system 140. Below Figure 600, Figure 6 This illustrates the information associated with each segment of audio data representing sound. Figure 6 In this context, each segment is associated with an audio type 610 identifier and an encoder 620 identifier. The audio type 610 identifier of a segment indicates the category associated with that segment. For example, in... Figure 6 In this context, "TYPE1" can indicate that a segment is associated with audio data of the target type (e.g., speech), and "TYPE2" can indicate that a segment is associated with audio data of a non-target type (e.g., non-speech).

[0085] exist Figure 6 In the first time period 650, the sound comprises TYPE1 content (e.g., speech) represented by four segments of audio data. During the second time period 652, the sound comprises TYPE2 content (e.g., non-speech sound) represented by four segments of audio data. The transition between TYPE1 and TYPE2 content occurs at time 630. Because... Figure 6 An example of fixed zero latency is illustrated for the transition between encoders, so that the first encoder is used for four segments of TYPE1 content associated with the first time period 650, and the second encoder is used for four segments of TYPE2 content associated with the second time period 652.

[0086] Figure 7 Including showing that will be by Figure 1 The content of Figure 700 shows the amplitude 702 and frequency 704 of the sound encoded by the decoder system 140. Below Figure 700, Figure 7 This illustrates the information associated with each segment of audio data representing sound. Figure 7 In this context, each segment is associated with an audio type 710 identifier and an encoder 720 identifier. The audio type 710 identifier of a segment indicates the category associated with that segment. For example, in... Figure 7 In this context, "TYPE1" can indicate that a segment is associated with audio data of the target type (e.g., speech), and "TYPE2" can indicate that a segment is associated with audio data of a non-target type (e.g., non-speech).

[0087] exist Figure 7 In the first time period 750, the sound comprises TYPE2 content (e.g., non-speech sound) represented by four segments of audio data. During the second time period 752, the sound comprises TYPE1 content (e.g., speech) represented by two segments of audio data. During the third time period 754, the sound comprises TYPE1 content (e.g., speech) represented by three segments of audio data. Therefore, the transition between TYPE2 and TYPE1 content occurs at time 730.

[0088] exist Figure 7 In this circuit, the transition from the second encoder to the first encoder occurs at time 732. For example, during time period 752, the energy of the sound is relatively high (e.g., compared to immediately following time 732). During such high-energy periods, the transition from the second encoder to the first encoder is more likely to introduce distortion. Therefore, the transition is delayed until a low-energy period is detected.

[0089] Figure 8The device 102 is depicted as a specific implementation 800 of an integrated circuit 802 including one or more processors 190. The integrated circuit 802 also includes an audio input 804 (such as one or more bus interfaces) to enable receiving audio data 112 for processing. The integrated circuit 802 also includes a signal output 806, such as a bus interface, to enable transmitting output signals, such as encoded audio data 808. In this example, the encoded audio data 808 may correspond to or include... Figure 1 The output of the machine learning audio encoder 144 is 154. Figure 1 The waveform matching audio encoder 146 output 156, Figures 1 to 4 Bitstream 172 of any graph, including as Figure 5 The header and payload, or one or more packets, as shown in Table 500. The integrated circuit 802, including the content-switchable decoder system 140, enables the implementation of content-switchable audio decoding as a component of the system, such as... Figure 9 The mobile phone or tablet computer described, such as Figure 10 The described headphones, such as Figure 11 The described wearable electronic devices, such as Figure 12 The described voice control loudspeaker system, such as Figure 13 The camera described, such as Figure 14 The virtual reality, mixed reality, or augmented reality headsets depicted, such as references Figure 16 The described mixed reality or augmented reality glasses device, such as the reference Figure 17 The described in-ear headphones, as referenced Figure 18 The described hearing aid device or such Figure 15 or Figure 19 The vehicles depicted.

[0090] Figure 9A specific implementation 900 of a mobile device 902, in which device 102 includes a telephone or tablet computer (as an illustrative, non-limiting example), is depicted. Mobile device 902 includes one or more microphones 906, one or more speakers 908, and a display screen 904. Components including a processor 190 with a content-switchable decoder system 140 are integrated into mobile device 902 and are illustrated using dashed lines to indicate internal components of mobile device 902 that are not typically visible to the user. In a particular example, content-switchable decoder system 140 is operable to acquire audio data representing sound captured by microphone 906, determine the type of audio content associated with segments of audio data, and encode the segments using a machine learning audio encoder, a waveform matching audio encoder, or both, based on the type of audio content associated with the segments. Selectively encoding segments of audio data using a machine learning audio encoder, a waveform matching audio encoder, or both enables mobile device 902 to efficiently (in terms of computational resources and power) generate a representation of audio data suitable for high-quality sound reproduction.

[0091] Figure 10 A specific implementation 1000 of the device 102, including a headset device 1002, is depicted. The headset device 1002 includes one or more microphones 1006 and one or more speakers 1008. Components including a processor 190 with a content-switching decoder system 140 are integrated into the headset device 1002. In a particular example, the content-switching decoder system 140 is operable to acquire audio data representing sound captured by the microphones 1006, determine the type of audio content associated with segments of the audio data, and encode the segments using a machine learning audio encoder, a waveform matching audio encoder, or both, based on the type of audio content associated with the segments. Selectively encoding segments of audio data using a machine learning audio encoder, a waveform matching audio encoder, or both enables the headset device 1002 to efficiently (in terms of computational resources and power) generate a representation of audio data suitable for high-quality sound reproduction.

[0092] Figure 11A specific implementation 1100 of which device 102 includes wearable electronic device 1102 (exemplified as a "smartwatch") is depicted. Wearable electronic device 1102 includes a display screen 1104, one or more microphones 1106, and one or more speakers 1108. Components including a processor 190 with a content-switching decoder system 140 are integrated into wearable electronic device 1102. In a particular example, content-switching decoder system 140 is operable to acquire audio data representing sound captured by microphone 1106, determine the type of audio content associated with segments of audio data, and encode the segments using a machine learning audio encoder, waveform matching audio encoder, or both based on the type of audio content associated with the segments. Selectively encoding segments of audio data using a machine learning audio encoder, waveform matching audio encoder, or both enables wearable electronic device 1102 to efficiently (in terms of computational resources and power) generate a representation of audio data suitable for high-quality sound reproduction. In some embodiments, wearable electronic device 1102 is configured to generate notifications based on the content of one or more segments of the segments. For example, display screen 1104 may generate visual information based on the content of the segments. As another example, wearable electronic device 1102 may include a haptic device that provides haptic notifications (e.g., vibration) based on the content of a fragment.

[0093] Figure 12 This is a specific implementation 1200 of device 102, which includes a wireless speaker and a voice activation device 1202. The wireless speaker and voice activation device 1202 may have wireless network connectivity and is configured to perform auxiliary operations. The wireless speaker and voice activation device 1202 includes one or more microphones 1206 and one or more speakers 1208. Components including a processor 190 of a content-switchable decoder system 140 are integrated into the wireless speaker and voice activation device 1202. In a particular example, the content-switchable decoder system 140 is operable to acquire audio data representing sound captured by the microphones 1206, determine the type of audio content associated with segments of the audio data, and encode the segments using a machine learning audio encoder, a waveform matching audio encoder, or both, based on the type of audio content associated with the segments. Selectively encoding segments of audio data using a machine learning audio encoder, a waveform matching audio encoder, or both enables the wireless speaker and voice activation device 1202 to efficiently (in terms of computational resources and power) generate a representation of audio data suitable for high-quality sound reproduction.

[0094] Figure 13A specific implementation 1300 of a portable electronic device corresponding to a camera device 1302 is depicted in which device 132 includes a camera device 1302. Camera device 1302 includes one or more microphones 1306 and one or more speakers 1308. Components including a processor 190 with a content-switching decoder system 140 are integrated into camera device 1302. In a particular example, content-switching decoder system 140 is operable to acquire audio data representing sound captured by microphone 1306, determine the type of audio content associated with segments of audio data, and encode the segments using a machine learning audio encoder, a waveform matching audio encoder, or both, based on the type of audio content associated with the segments. Selectively encoding segments of audio data using a machine learning audio encoder, a waveform matching audio encoder, or both enables camera device 1302 to efficiently (in terms of computational resources and power) generate a representation of audio data suitable for high-quality sound reproduction.

[0095] Figure 14 A specific implementation 1400 of a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality head-mounted device 1402 is depicted. A visual interface device is positioned in front of the user's eyes to enable the display of augmented reality, mixed reality, or virtual reality images or scenes to the user when wearing the head-mounted device 1402. The head-mounted device 1402 also includes one or more microphones 1406 and one or more speakers 1408. Components including a processor 190 of a content-switching decoder system 140 are integrated into the head-mounted device 1402. In a particular example, the content-switching decoder system 140 is operable to acquire audio data representing sound captured by the microphones 1406, determine the type of audio content associated with segments of the audio data, and encode the segments using a machine learning audio encoder, a waveform matching audio encoder, or both based on the type of audio content associated with the segments. Selectively encoding segments of audio data using a machine learning audio encoder, a waveform matching audio encoder, or both enables the head-mounted device 1402 to efficiently (in terms of computational resources and power) generate a representation of audio data suitable for high-quality sound reproduction.

[0096] Figure 15A specific implementation 1500 is depicted in which device 102 corresponds to or is integrated within vehicle 1502 (exemplified as a manned or unmanned aerial device, such as a package delivery drone). Vehicle 1502 includes one or more microphones 1506 and one or more speakers 1508. Components including a processor 190 with a content-switching decoder system 140 are integrated into vehicle 1502. In a particular example, the content-switching decoder system 140 is operable to acquire audio data representing sound captured by microphone 1506, determine the type of audio content associated with segments of audio data, and encode the segments using a machine learning audio encoder, a waveform matching audio encoder, or both based on the type of audio content associated with the segments. Selectively encoding segments of audio data using a machine learning audio encoder, a waveform matching audio encoder, or both enables vehicle 1502 to efficiently (in terms of computational resources and power) generate a representation of audio data suitable for high-quality sound reproduction. For example, spoken instructions may be captured by microphone 1506 and used as a bitstream (e.g., Figure 1 The bit stream 172 is sent to a remote device for processing (e.g., to detect delivery instructions or determine whether spoken instructions are from an authorized user).

[0097] Figure 16 A specific implementation 1600 of a portable electronic device corresponding to augmented reality or mixed reality glasses 1602 is depicted in which device 102 includes such a device. Glasses 1602 include a holographic projection unit 1604 configured to project visual data onto the surface of a lens 1606, or to reflect the visual data from the surface of the lens 1606 onto the wearer's retina. Glasses 1602 also includes one or more microphones 1608 and one or more speakers 1610. Components including a processor 190 of a content-switchable decoder system 140 are integrated into glasses 1602. In a particular example, the content-switchable decoder system 140 is operable to acquire audio data representing sound captured by the microphones 1608, determine the type of audio content associated with a segment of audio data, and encode the segment using a machine learning audio encoder, a waveform matching audio encoder, or both, based on the type of audio content associated with the segment. Using a machine learning audio encoder, a waveform matching audio encoder, or both to selectively encode segments of audio data enables glasses 1602 to efficiently (in terms of computational resources and power) generate a representation of audio data suitable for high-quality sound reproduction. In a particular example, holographic projection unit 1604 is configured to display a notification indicating a detected audio event. For example, the notification may be based on the content of one or more segments of the data.

[0098] Figure 17A specific embodiment 1700 of a portable electronic device is depicted in which device 102 includes a pair of earbud-type headphones 1706, the pair of earbud-type headphones including a first earbud-type headphone 1702 and a second earbud-type headphone 1704. Although earbud-type headphones are described, it should be understood that this technology can be applied to other in-ear or over-ear audio devices.

[0099] The first earbud-type headphone 1702 includes: a first microphone 1720, such as a high signal-to-noise ratio microphone positioned to capture the speech of the wearer of the first earbud-type headphone 1702; an array of one or more other microphones configured to detect ambient sound and spatially distributed to support beamforming, exemplified as microphones 1722A, 1722B, and 1722C; an “internal” microphone 1724 located near the wearer’s ear canal (e.g., to assist in active noise cancellation); and a self-speech microphone 1726, such as a bone conduction microphone configured to convert sound vibrations from the wearer’s ear bones or skull into audio signals.

[0100] The second earbud 1704 can be configured in a manner substantially similar to that of the first earbud 1702. In some embodiments, the first earbud 1702 is also configured to receive one or more audio signals generated by one or more microphones of the second earbud 1704 (such as wireless transmission between earbuds 1702 and 1704, or wired transmission in embodiments where earbuds 1702 and 1704 are coupled via a transmission line).

[0101] In some embodiments, the earbuds 1702 and 1704 are configured to automatically switch between various operating modes, such as a pass-through mode in which ambient sounds are played via speaker 1730; a playback mode in which non-ambient sounds (e.g., streaming audio corresponding to telephone conversations, media playback, video games, etc.) are played back via speaker 1730; and an audio zoom mode or beamforming mode in which one or more ambient sounds are amplified and / or other ambient sounds are suppressed for playback at speaker 1730. In other embodiments, the earbuds 1702 and 1704 may support fewer modes, or may support one or more other modes in place of the described modes, or may support one or more other modes in addition to the described modes.

[0102] In exemplary examples, earbuds 1702 and 1704 can automatically switch from playback mode to pass-through mode in response to detecting the wearer's voice, and can automatically switch back to playback mode after the wearer stops speaking. In some examples, earbuds 1702 and 1704 can operate concurrently in two or more modes, such as performing audio zoom on a specific ambient sound (e.g., a dog barking) and playing an audio zoom sound superimposed on the playing sound while the wearer is listening to music (the volume can be reduced while playing the audio zoom sound). In this example, the wearer can be alerted to the ambient sound associated with the audio event without stopping music playback.

[0103] exist Figure 17 In this embodiment, components including the processor 190 of the content-switchable decoder system 140 are integrated into the earphones 1702 and 1704. In a particular example, the content-switchable decoder system 140 is operable to acquire audio data representing sound captured by one or more microphones 1720, 1722, 1724, and 1726, determine the type of audio content associated with segments of the audio data, and encode the segments using a machine learning audio encoder, a waveform matching audio encoder, or both, based on the type of audio content associated with the segments. The selective encoding of segments of audio data using a machine learning audio encoder, a waveform matching audio encoder, or both enables the earphones 1702 and 1704 to efficiently (in terms of computational resources and power) generate representations of audio data suitable for high-quality sound reproduction.

[0104] Figure 18 Examples of merging Figure 1 The hearing aid device 102 includes various aspects of hearing aid device 1800. Figure 18 In this embodiment, hearing aid device 1800 includes a housing 1802 that includes an earcup portion 1812 configured to be worn on a user's ear. A receiver 1810 is coupled to the housing 1802 and includes one or more speakers 1808. In some embodiments, one or more microphones 1806 are disposed on the housing 1802.

[0105] exist Figure 18In this device, components including the processor 190 of the content-switchable decoder system 140 are integrated into the hearing aid device 1800. In a particular example, the content-switchable decoder system 140 is operable to acquire audio data representing sound captured by the microphone 1806, determine the type of audio content associated with segments of the audio data, and encode the segments using a machine learning audio encoder, a waveform matching audio encoder, or both, based on the type of audio content associated with the segments. Selectively encoding segments of audio data using a machine learning audio encoder, a waveform matching audio encoder, or both enables the hearing aid device 1800 to efficiently (in terms of computational resources and power) generate a representation of audio data suitable for high-quality sound reproduction.

[0106] Figure 19 Another specific embodiment 1900 in which device 102 corresponds to or is integrated within a vehicle 1902 (illustrated as an automobile) is depicted. Vehicle 1902 includes a display screen 1920, one or more microphones 1906, and one or more speakers 1908. Components including a processor 190 with a content-switchable decoder system 140 are integrated into vehicle 1902. In a particular example, content-switchable decoder system 140 is operable to acquire audio data representing sound captured by microphone 1906, determine the type of audio content associated with segments of audio data, and encode the segments using a machine learning audio encoder, a waveform matching audio encoder, or both, based on the type of audio content associated with the segments. Selectively encoding segments of audio data using a machine learning audio encoder, a waveform matching audio encoder, or both enables vehicle 1902 to efficiently (in terms of computational resources and power) generate a representation of audio data suitable for high-quality sound reproduction.

[0107] refer to Figure 20 This illustrates a specific implementation of a method 2000 for switching audio codecs based on the content of one or more segments of audio data. In a particular aspect, one or more operations of method 2000 are performed by… Figure 1 The content can be executed by at least one of the following: decoder system 140, processor 190, device 102, system 100, or a combination thereof.

[0108] In some implementations, method 2000 includes, at block 2002, one or more processors obtaining an indication of the type of audio content associated with a segment of audio data. For example, Figure 1The controller 142 may receive indications (e.g., indicator 150) from the audio classifier 148. The indications may indicate the classification associated with a segment of audio data. The audio classifier 148 may include a machine learning-based classifier, a procedural classifier (e.g., a voice activity detector), or a combination thereof. The audio classifier 148 may generate indications based on whether the segment represents a specific type of audio content (such as speech, music, wind noise, etc.).

[0109] The method 2000 further includes, at block 2004, selectively feeding the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both, based on the instruction. For example, Figure 1 The controller 142 can cause segment 114 to be transmitted to a machine learning audio encoder 144, a waveform matching audio encoder 146, or both. In some embodiments, the controller is configured to transmit each segment of audio data as input to a single audio encoder. For example, each segment of audio data is transmitted to either a machine learning audio encoder or a waveform matching audio encoder. In other embodiments, the controller is configured to transmit at least some segments of audio data to both the machine learning audio encoder and the waveform matching audio encoder. For example, the controller can provide at least one segment of audio data to both the machine learning audio encoder and the waveform matching audio encoder based on a determination of which audio encoder to which segment of audio data is being provided.

[0110] In some implementations, method 2000 includes generating a bitstream representing the output of the machine learning audio encoder, the output of the waveform matching audio encoder, or both. For example, Figure 1 The modem 170 can generate a bitstream 172, which may include outputs 154, 156, or both representing a specific input segment 114 of audio data 112. In some embodiments, the machine learning audio encoder and the waveform matching audio encoder have different bit rates. For example, the machine learning audio encoder may be configured to encode the input segment using a first number of bits, and the waveform matching audio encoder may be configured to encode the input segment using a second number of bits, wherein the first number is less than the second number.

[0111] In some implementations, method 2000 includes: when a particular audio encoder is selected to process two consecutive segments of the audio data, processing a second segment of the two consecutive segments using encoder state data obtained from processing a first segment of the two consecutive segments, wherein the second segment follows the first segment. For example, when selecting... Figure 1When the waveform matching audio encoder 146 processes segment 114A and segment 114B immediately following segment 114A, encoder state data associated with the processing of segment 114A can be used to process segment 114B. As another example, when selecting... Figure 1 When the machine learning audio encoder 144 processes segment 114A and segment 114B immediately following segment 114A, encoder state data associated with the processing of segment 114A can be used to process segment 114B.

[0112] In some implementations, method 2000 includes: when different audio encoders are selected to process two consecutive segments of the audio data, processing a second segment of the two consecutive segments using encoder state data independent of the encoder state data used for processing the first segment of the two consecutive segments, wherein the second segment follows the first segment. For example, when selecting... Figure 1 When the waveform matching audio encoder 146 is selected to process segment 114A and the machine learning audio encoder 144 is selected to process segment 114B immediately following segment 114A, the encoder state data associated with the processing of segment 114A is not used to process segment 114B. For illustration, the encoder state data used to process segment 114B may include or correspond to default state data or state data representing the state of the machine learning audio encoder 144 before the waveform matching audio encoder 146 encodes segment 114A. As another example, when the waveform matching audio encoder 146 is selected... Figure 1 When the machine learning audio encoder 144 processes segment 114A and the waveform matching audio encoder 146 is selected to process segment 114B immediately following segment 114A, the encoder state data associated with the processing of segment 114A is not used to process segment 114B. For illustration, the encoder state data used to process segment 114B may include or correspond to default state data or state data representing the state of the waveform matching audio encoder 146 before the machine learning audio encoder 144 encodes segment 114A.

[0113] In some implementations, method 2000 includes: when different audio encoders are selected to process two consecutive segments of the audio data, processing a second segment of the two consecutive segments using encoder state data obtained from processing a first segment of the two consecutive segments, wherein the second segment follows the first segment. For example, when selecting... Figure 1 When the waveform matching audio encoder 146 is selected to process segment 114A and the machine learning audio encoder 144 is selected to process segment 114B immediately following segment 114A, the encoder state data associated with the processing of segment 114A can be provided to the machine learning-based audio encoder to initialize the processing of segment 114B. As another example, when the waveform matching audio encoder 146 is selected to process segment 114A, the encoder state data associated with the processing of segment 114A can be provided to the machine learning-based audio encoder to initialize the processing of segment 114B. Figure 1When the machine learning audio encoder 144 processes segment 114A and the waveform matching audio encoder 146 is selected to process segment 114B immediately following segment 114A, encoder state data associated with the processing of segment 114A can be provided to the waveform matching audio encoder to initialize the processing of segment 114B.

[0114] In some implementations, method 2000 includes applying a first delay when transitioning to inputting a segment into the waveform-matching audio encoder, and applying a second delay when transitioning to inputting the segment into the machine learning audio encoder, wherein the first delay is different from the second delay. In some such implementations, the first delay is fixed, and method 2000 further includes determining the second delay based on the content of one or more segments of the segment.

[0115] Figure 20 Method 2000 can be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 20 Method 2000 can be executed by a processor that executes instructions, such as references Figure 21 As described.

[0116] refer to Figure 21 This diagram depicts a specific, exemplary embodiment of the device, and generally designates the device as 2100. In various embodiments, device 2100 may have the same... Figure 21 The illustrated components may be more or fewer than the number of components. In an exemplary embodiment, device 2100 may correspond to device 102. In an exemplary embodiment, device 2100 may perform reference... Figures 1 to 20 One or more operations as described.

[0117] In a particular implementation, device 2100 includes a processor 2106 (e.g., a central processing unit (CPU)). Device 2100 may include one or more additional processors 2110 (e.g., one or more DSPs). In a particular aspect, Figure 1 The processor 190 corresponds to processor 2106, processor 2110, or a combination thereof. Processor 2110 may include a speech and music decoder-decoder (codec) 2108, which includes a voice decoder (“vocoder”) encoder 2136, a vocoder decoder 2138, a content switchable decoder system 140, or a combination thereof.

[0118] In this context, the term "processor" refers to an integrated circuit composed of logic units, interconnects, input / output blocks, clock management components, memory, and optional other dedicated hardware components, designed to execute instructions and perform various computational tasks. Examples of processors include, but are not limited to, central processing units (CPUs), digital signal processors (DSPs), neural processing units (NPUs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs), microcontrollers, quantum processors, coprocessors, vector processors, other similar circuits, and variations and combinations thereof. In some cases, processors may be integrated with other components, such as communication components, input / output components, etc., to form system-on-a-chip (SoC) devices or packaged electronics.

[0119] Starting with the CPU, a CPU typically includes one or more processor cores. Each of these cores comprises a complex interconnected network of transistors and other circuit components that define logic gates, memory elements, and so on. The cores are responsible for executing instructions to perform, for example, arithmetic and logical operations. Generally, a CPU includes an arithmetic logic unit (ALU) that handles mathematical operations and a control unit that generates signals to coordinate the operations of other CPU components (such as managing fetch-decode-execute loops).

[0120] CPUs and / or individual processor cores typically include local memory circuitry (such as registers and caches) to temporarily store data during operation. Registers consist of high-speed, small-size memory cells that are tightly connected to the CPU's logic units. Registers typically include transistors arranged as groups of flip-flops configured to store binary data. Caches include fast on-chip memory circuitry for storing frequently accessed data. Caches can be implemented, for example, using static random access memory (SRAM) circuitry.

[0121] CPU operations (e.g., arithmetic, logical, and flow control operations) are guided by software and firmware. At the lowest level, the CPU includes an Instruction Set Architecture (ISA), which specifies how hardware resources (e.g., registers, arithmetic units, etc.) are used to perform individual operations. Higher-level software and firmware are translated into various combinations of ISA operations to enable the CPU to perform specific higher-level operations. For example, an ISA generally specifies how the CPU's hardware components move and modify data to perform operations such as addition, multiplication, and subtraction, while higher-level software is translated into sets of such operations to accomplish larger tasks (e.g., adding two columns in a spreadsheet). Typically, the CPU operates on different levels of software, including the kernel, operating system, applications, etc., where each higher-level software is usually more closely associated with the abstraction of the ISA and is often more easily understood by human users.

[0122] GPUs, NPUs, DSPs, microcontrollers, coprocessors, FPGAs, ASICs, and vector processors include components similar to those described above for CPUs. The differences between these various types of processors typically involve the use of specialized interconnect schemes and ISAs to improve the processor's ability to perform specific types of operations. For example, the logic gates, local memory circuitry, and interconnects between them in a GPU are specifically designed to improve parallel processing, data sharing between processor cores, and vector operations, and the GPU's ISA defines the operations that utilize these structures. As another example, an ASIC is a highly specialized processor, comprising similar circuitry arranged and interconnected for specific tasks such as encryption or signal processing. As yet another example, an FPGA is a programmable device, comprising an array of configurable logic blocks (e.g., interconnects of transistors and memory elements) that can be configured (often supporting dynamic configuration) to perform customizable logic functions.

[0123] Device 2100 may include memory 2186 and codec 2134. Memory 2186 may include instructions 2156 that can be executed by one or more additional processors 2110 (or processor 2106) to implement the functionality described in Reference Content Switchable Decoder System 140 or both. Device 2100 may include modem 170 coupled to antenna 2152 via transceiver 2150.

[0124] Device 2100 may include a display 2128 coupled to display controller 2126. One or more speakers 2192 and microphones 2194 may be coupled to codec 2134. Codec 2134 may include digital-to-analog converter (DAC) 2102, analog-to-digital converter (ADC) 2104, or both. In a particular embodiment, codec 2134 may receive analog signals from microphone 2194, convert these analog signals to digital signals using ADC 2104, and provide these digital signals to speech and music codec 2108. Speech and music codec 2108 may process digital signals, and the digital signals may be further processed by content switchable decoder system 140. In a particular embodiment, speech and music codec 2108 may provide digital signals to codec 2134. Codec 2134 may use ADC 2102 to convert digital signals to analog signals and may provide the analog signals to speaker 2192.

[0125] In a particular embodiment, device 2100 may be included in a system-in-package (SiP) or a system-on-a-chip (SoC) 2122. In a particular embodiment, memory 2186, processor 2106, processor 2110, display controller 2126, codec 2134, and modem 170 are included in the SiP or SoC 2122. In a particular embodiment, input device 2130 and power supply 2144 are coupled to the SiP or SoC 2122. Furthermore, in a particular embodiment, such as... Figure 21 As illustrated, the display 2128, input device 2130, speaker 2192, microphone 2194, antenna 2152, and power supply 2144 are external to the system-in-package or system-on-chip device 2122. In a particular implementation, each of the display 2128, input device 2130, speaker 2192, microphone 2194, antenna 2152, and power supply 2144 may be coupled to a component of the system-in-package or system-on-chip device 2122, such as an interface or controller.

[0126] Device 2100 may include smart speakers, speaker bars, mobile communication devices, smartphones, cellular phones, laptops, computers, tablets, personal digital assistants, display devices, televisions, game consoles, music players, radios, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, vehicles, headsets, augmented reality headsets, mixed reality headsets, virtual reality headsets, aircraft, home automation systems, voice-activated devices, wireless speakers and voice-activated devices, portable electronic devices, automobiles, computing devices, communication devices, Internet of Things (IoT) devices, virtual reality (VR) devices, base stations, mobile devices, or any combination thereof.

[0127] In conjunction with the described specific embodiments, an apparatus includes components for acquiring an indication of the type of audio content associated with a segment of audio data. For example, the components for acquiring an indication of the type of audio content associated with a segment of audio data may include system 100, device 102, processor 190, content switchable decoder system 140, controller 142, audio classifier 148, integrated circuit 802, processor 2106, processor 2110, system-in-package or system-on-chip device 2122, device 2100, other circuitry configured to acquire an indication of the type of audio content associated with a segment of audio data, or combinations thereof.

[0128] The apparatus also includes components for selectively transmitting the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both based on the indication. For example, components for selectively transmitting the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both based on the indication may include system 100, device 102, processor 190, content switchable decoder system 140, controller 142, integrated circuit 802, processor 2106, processor 2110, system-in-package or system-on-chip device 2122, device 2100, or other circuitry or combinations thereof configured to selectively transmit the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both based on an indication of the type of audio content associated with the segment of audio data.

[0129] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as memory 2186) includes instructions (e.g., instructions 2156) that, when executed by one or more processors (e.g., one or more processors 2110 or 2106), cause the one or more processors to obtain an indication of the type of audio content associated with a segment of audio data. These instructions also cause the one or more processors to selectively pass the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both, based on the indication.

[0130] Specific aspects of this disclosure are described below in a collection of related embodiments:

[0131] According to Embodiment 1, an apparatus includes: a machine learning audio encoder; a waveform matching audio encoder; and a controller configured to input segments of audio data into the machine learning audio encoder, the waveform matching audio encoder, or both, based on a classification associated with segments of audio data.

[0132] Example 2 includes the device according to Example 1, the device further including an audio classifier configured to generate an indicator for the classification based on whether the segment represents a specific type of audio content, and configured to provide the indicator to the controller.

[0133] Example 3 includes the device according to Example 1 or Example 2, the device further including a modem coupled to the machine learning audio encoder and the waveform matching audio encoder and configured to represent the output of the machine learning audio encoder, the output of the waveform matching audio encoder, or both in a bit stream.

[0134] Example 4 includes a device according to any one of Examples 1 to 3, wherein the machine learning audio encoder is configured to encode an input segment using a first number of bits, wherein the waveform matching audio encoder is configured to encode the input segment using a second number of bits, and wherein the first number is less than the second number.

[0135] Example 5 includes a device according to any one of Examples 1 to 4, wherein the controller is configured to select the machine learning audio encoder to process a first set of segments representing speech, and to select the waveform matching audio encoder to process a second set of segments representing non-speech sounds.

[0136] Example 6 includes a device according to any one of Examples 1 to 5, wherein the controller is configured to select a single audio encoder to process each corresponding segment of the audio data.

[0137] Example 7 includes a device according to any one of Examples 1 to 5, wherein the controller is configured to provide at least one segment of the audio data to both the machine learning audio encoder and the waveform matching audio encoder in response to determining which segment of the audio data to provide to which audio encoder.

[0138] Example 8 includes the device according to any one of Examples 1 to 7 and further includes a modem coupled to the machine learning audio encoder and the waveform matching audio encoder and configured to represent in a bit stream the output of the machine learning audio encoder, the output of the waveform matching audio encoder, or both.

[0139] Example 9 includes a device according to any one of Examples 1 to 8, wherein, when a particular audio encoder is selected to process two consecutive segments of the audio data, encoder state data obtained from processing the first segment of the two consecutive segments is used to process the second segment of the two consecutive segments, wherein the second segment follows the first segment.

[0140] Example 10 includes a device according to any one of Examples 1 to 9, wherein, when different audio encoders are selected to process two consecutive segments of the audio data, a second segment of the two consecutive segments is processed using default encoder state data, wherein the second segment follows the first segment.

[0141] Example 11 includes a device according to any one of Examples 1 to 9, wherein, when different audio encoders are selected to process two consecutive segments of the audio data, encoder state data for processing the second segment of the two consecutive segments by the second audio encoder is based on the previous state of the second audio encoder, wherein the second segment follows the first segment.

[0142] Example 12 includes a device according to any one of Examples 1 to 9, wherein, when different audio encoders are selected to process two consecutive segments of the audio data, encoder state data for processing the second segment of the two consecutive segments is based on the processing of the first segment of the two consecutive segments, wherein the second segment follows the first segment.

[0143] Example 13 includes a device according to any one of Examples 1 to 12, wherein the controller is configured to use a first delay when transitioning to inputting a segment to the waveform matching audio encoder, and is configured to use a second delay when transitioning to inputting a segment to the machine learning audio encoder, wherein the first delay is different from the second delay.

[0144] Example 14 includes the device according to Example 13, wherein the first delay is fixed and the second delay is variable and selected based on the content of the segment.

[0145] Example 15 includes the device according to any one of Examples 1 to 14, wherein the controller is integrated into one or more processors.

[0146] Example 16 includes the device according to any one of Examples 1 to 14, wherein the controller is integrated into the processing circuitry.

[0147] Example 17 includes a device according to any one of Examples 1 to 16, wherein the machine learning audio encoder, the waveform matching audio encoder, or both are integrated into a processor.

[0148] Example 18 includes a device according to any one of Examples 1 to 17, wherein the controller, the machine learning audio encoder, and the waveform matching audio encoder are integrated into at least one of a mobile phone, a tablet computer device, a wearable electronic device, a camera device, a virtual reality headset, a mixed reality headset, or an augmented reality headset.

[0149] Example 19 includes a device according to any one of Examples 1 to 17, wherein the controller, the machine learning audio encoder, and the waveform matching audio encoder are integrated in a vehicle.

[0150] According to embodiment 20, a method includes: obtaining, by one or more processors, an indication of the type of audio content associated with a segment of audio data; and selectively transmitting, based on the indication, the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both.

[0151] Example 21 includes the method according to Example 20, the method further including using an audio classifier to generate the indication based on whether the segment represents a specific type of audio content.

[0152] Example 22 includes the method according to Example 20 or Example 21, the method further including generating a bitstream representing the output of the machine learning audio encoder, the output of the waveform matching audio encoder, or both.

[0153] Example 23 includes the method according to any one of Examples 20 to 22, wherein the machine learning audio encoder is configured to encode an input segment using a first number of bits, wherein the waveform matching audio encoder is configured to encode the input segment using a second number of bits, and wherein the first number is less than the second number.

[0154] Example 24 includes the method according to any one of Examples 20 to 23, wherein each segment of the audio data is transmitted as input to a single audio encoder.

[0155] Example 25 includes the method according to any one of Examples 20 to 23 and further includes: providing at least one segment of the audio data to both the machine learning audio encoder and the waveform matching audio encoder based on the determination of which audio encoder to which segment of the audio data is provided.

[0156] Example 26 includes the method according to any one of Examples 20 to 25 and further includes: when a particular audio encoder is selected to process two consecutive segments of the audio data, processing a second segment of the two consecutive segments using encoder state data obtained from processing a first segment of the two consecutive segments, wherein the second segment follows the first segment.

[0157] Example 27 includes the method according to any one of Examples 20 to 26 and further includes: when different audio encoders are selected to process two consecutive segments of the audio data, using default encoder state data to process a second segment of the two consecutive segments, wherein the second segment follows the first segment.

[0158] Example 28 includes the method according to any one of Examples 20 to 26 and further includes: when different audio encoders are selected to process two consecutive segments of the audio data, a second audio encoder processes a second segment of the two consecutive segments based on the previous state of the second audio encoder using encoder state data, wherein the second segment follows the first segment.

[0159] Example 29 includes the method according to any one of Examples 20 to 26 and further includes: when different audio encoders are selected to process two consecutive segments of the audio data, using encoder state data obtained from processing the first segment of the two consecutive segments to process the second segment of the two consecutive segments, wherein the second segment is after the first segment.

[0160] Example 30 includes the method according to any one of Examples 20 to 29 and further includes: applying a first delay when transitioning to inputting a segment to the waveform matching audio encoder, and applying a second delay when transitioning to inputting a segment to the machine learning audio encoder, wherein the first delay is different from the second delay.

[0161] Example 31 includes the method according to Example 30, wherein the first delay is fixed, and the method further includes determining the second delay based on the content of one or more segments of the segment.

[0162] According to embodiment 32, a non-transitory computer-readable medium storage instruction is provided, the instruction being executable by one or more processors to cause the one or more processors to: obtain an indication of the type of audio content associated with a segment of audio data; and selectively transmit the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both, based on the indication.

[0163] Example 33 includes a non-transitory computer-readable medium according to Example 32, wherein the instructions are executable to cause the one or more processors to generate the indication using an audio classifier based on whether the segment represents a specific type of audio content.

[0164] Example 34 includes a non-transitory computer-readable medium according to Example 32 or Example 33, wherein the instructions are executable to cause the one or more processors to generate a bitstream representing the output of the machine learning audio encoder, the output of the waveform matching audio encoder, or both.

[0165] Example 35 includes a non-transitory computer-readable medium according to any one of Examples 32 to 34, wherein the machine learning audio encoder is configured to encode an input segment using a first number of bits, wherein the waveform matching audio encoder is configured to encode the input segment using a second number of bits, and wherein the first number is less than the second number.

[0166] Example 36 includes a non-transitory computer-readable medium according to any one of Examples 32 to 35, wherein the instructions are executable to cause the one or more processors to transmit each segment of the audio data as input to a single audio encoder.

[0167] Example 37 includes a non-transitory computer-readable medium according to any one of Examples 32 to 35, wherein the instructions are executable to cause the one or more processors to: provide at least one segment of the audio data to both the machine learning audio encoder and the waveform matching audio encoder based on a determination of which audio encoder to which segment of the audio data is provided.

[0168] Example 38 includes a non-transitory computer-readable medium according to any one of Examples 32 to 37, wherein the instructions are executable to cause the one or more processors to: when a particular audio encoder is selected to process two consecutive segments of the audio data, process a second segment of the two consecutive segments using encoder state data obtained from processing a first segment of the two consecutive segments, wherein the second segment follows the first segment.

[0169] Example 39 includes a non-transitory computer-readable medium according to any one of Examples 32 to 38, wherein the instructions are executable to cause the one or more processors to: when different audio encoders are selected to process two consecutive segments of the audio data, use default encoder state data to process a second segment of the two consecutive segments, wherein the second segment follows the first segment.

[0170] Example 40 includes a non-transitory computer-readable medium according to any one of Examples 32 to 38, wherein the instructions are executable to cause the one or more processors to: when different audio encoders are selected to process two consecutive segments of the audio data, process a second segment of the two consecutive segments by a second audio encoder using encoder state data based on the previous state of the second audio encoder, wherein the second segment follows the first segment.

[0171] Example 41 includes a non-transitory computer-readable medium according to any one of Examples 32 to 38, wherein the instructions are executable to cause the one or more processors to: process a second segment of the two consecutive segments, wherein the second segment follows the first segment, using encoder state data obtained from processing a first segment of the two consecutive segments, when different audio encoders are selected to process two consecutive segments of the audio data.

[0172] Example 42 includes a non-transitory computer-readable medium according to any one of Examples 32 to 41, wherein the instructions are executable to cause the one or more processors to: apply a first delay when transitioning to inputting a segment to the waveform matching audio encoder, and apply a second delay when transitioning to inputting a segment to the machine learning audio encoder, wherein the first delay is different from the second delay.

[0173] Example 43 includes a non-transitory computer-readable medium according to Example 42, wherein the first delay is fixed, and wherein the instructions are executable to cause the one or more processors to determine the second delay based on the content of one or more segments of the segment.

[0174] According to embodiment 44, an apparatus includes: a component for acquiring an indication of the type of audio content associated with a segment of audio data; and a component for selectively transmitting the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both, based on the indication.

[0175] Example 45 includes the apparatus according to Example 44, the apparatus further including components for generating the indication using an audio classifier based on whether the segment represents a specific type of audio content.

[0176] Example 46 includes the apparatus according to Example 44 or Example 45, the apparatus further including components for generating a bitstream representing the output of the machine learning audio encoder, the output of the waveform matching audio encoder, or both.

[0177] Example 47 includes an apparatus according to any one of Examples 44 to 46, wherein the machine learning audio encoder is configured to encode an input segment using a first number of bits, wherein the waveform matching audio encoder is configured to encode the input segment using a second number of bits, and wherein the first number is less than the second number.

[0178] Example 48 includes an apparatus according to any one of Examples 44 to 47, wherein each segment of the audio data is transmitted as input to a single audio encoder.

[0179] Example 49 includes the apparatus according to any one of Examples 44 to 47 and further includes components for providing at least one segment of the audio data to both the machine learning audio encoder and the waveform matching audio encoder based on determining which segment of the audio data is provided to which audio encoder based on the transition.

[0180] Example 50 includes the apparatus according to any one of Examples 44 to 49 and further includes components for processing a second segment of the two consecutive segments, wherein, when a particular audio encoder is selected to process two consecutive segments of the audio data, encoder state data obtained from processing a first segment of the two consecutive segments is used, wherein the second segment follows the first segment.

[0181] Example 51 includes the apparatus according to any one of Examples 44 to 50 and further includes components for processing a second segment of the two consecutive segments of the audio data using default encoder state data when different audio encoders are selected to process two consecutive segments of the audio data, wherein the second segment follows the first segment.

[0182] Example 52 includes the apparatus according to any one of Examples 44 to 50 and further includes components for processing a second segment of two consecutive segments based on a previous state of the second audio encoder using encoder state data, wherein the second segment follows the first segment.

[0183] Example 53 includes the apparatus according to any one of Examples 44 to 50 and further includes components for processing a second segment of the two consecutive segments, wherein the second segment follows the first segment, when different audio encoders are selected to process two consecutive segments of the audio data.

[0184] Example 54 includes the apparatus according to any one of Examples 44 to 53 and further includes components for applying a first delay when transitioning to inputting a segment to the waveform matching audio encoder, and applying a second delay when transitioning to inputting a segment to the machine learning audio encoder, wherein the first delay is different from the second delay.

[0185] Example 55 includes the apparatus according to Example 54, wherein the first delay is fixed, and the apparatus further includes components for determining the second delay based on the content of one or more of the segments.

[0186] Those skilled in the art will also understand that the various exemplary logic blocks, configurations, modules, circuits, and algorithm steps described in connection with the specific embodiments disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various exemplary components, blocks, configurations, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, and such implementation decisions shall not be construed as departing from the scope of this disclosure.

[0187] The steps of the methods or algorithms described in conjunction with the specific embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, compressed optical disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integral with the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and storage medium may reside as discrete components in a computing device or a user terminal.

[0188] The prior description of the disclosed aspects is provided to enable those skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but should be granted the broadest scope that may be consistent with the principles and novel features as defined by the following claims.

Claims

1. An apparatus, the apparatus comprising: Machine learning audio encoder; Waveform matching audio encoder; and A controller configured to input segments of audio data into the machine learning audio encoder, the waveform matching audio encoder, or both, based on a classification associated with segments of audio data.

2. The device of claim 1, further comprising an audio classifier configured to generate an indicator for the classification based on whether the segment represents a specific type of audio content, and configured to provide the indicator to the controller.

3. The device of claim 1, further comprising a modem coupled to the machine learning audio encoder and the waveform matching audio encoder and configured to represent in a bit stream the output of the machine learning audio encoder, the output of the waveform matching audio encoder, or both.

4. The device of claim 1, wherein the machine learning audio encoder is configured to encode an input segment using a first number of bits, wherein the waveform matching audio encoder is configured to encode an input segment using a second number of bits, and wherein the first number is less than the second number.

5. The device of claim 1, wherein the controller is configured to select the machine learning audio encoder to process a first set of segments representing speech, and to select the waveform matching audio encoder to process a second set of segments representing non-speech sounds.

6. The device of claim 1, wherein the controller is configured to select a single audio encoder to process each corresponding segment of the audio data.

7. The apparatus of claim 1, wherein, When a specific audio encoder is selected to process two consecutive segments of the audio data, the second segment of the two consecutive segments is processed using encoder state data obtained from processing the first segment, wherein the second segment follows the first segment.

8. The apparatus of claim 1, wherein, When different audio encoders are selected to process two consecutive segments of the audio data, the second segment of the two consecutive segments, which is after the first segment, is processed using the default encoder state data.

9. The apparatus of claim 8, wherein, When different audio encoders are selected to process two consecutive segments of the audio data, the encoder state data used by the second audio encoder to process the second segment of the two consecutive segments is based on the previous state of the second audio encoder, wherein the second segment follows the first segment.

10. The apparatus of claim 1, wherein, When different audio encoders are selected to process two consecutive segments of the audio data, the encoder state data for processing the second segment of the two consecutive segments is based on the processing of the first segment of the two consecutive segments, wherein the second segment follows the first segment.

11. The device of claim 1, wherein the controller is configured to use a first delay when transitioning to inputting a segment to the waveform matching audio encoder, and is configured to use a second delay when transitioning to inputting a segment to the machine learning audio encoder, wherein the first delay is different from the second delay.

12. The device of claim 11, wherein the first delay is fixed and the second delay is variable and selected based on the content of the segment.

13. The device of claim 1, wherein the controller is configured to provide at least one segment of the audio data to both the machine learning audio encoder and the waveform matching audio encoder in response to determining which segment of the audio data to provide to which audio encoder.

14. The device of claim 13, further comprising a modem coupled to the machine learning audio encoder and the waveform matching audio encoder and configured to represent the output of the machine learning audio encoder and the output of the waveform matching audio encoder in a bit stream.

15. The device of claim 1, wherein the controller is integrated into one or more processors.

16. The device of claim 1, wherein the controller is integrated into the processing circuitry.

17. The device of claim 1, wherein the machine learning audio encoder, the waveform matching audio encoder, or both are integrated in the processor.

18. The device of claim 1, wherein the controller, the machine learning audio encoder, and the waveform matching audio encoder are integrated into at least one of a mobile phone, a tablet computer device, a wearable electronic device, a camera device, a virtual reality headset, a mixed reality headset, or an augmented reality headset.

19. The device of claim 1, wherein the controller, the machine learning audio encoder, and the waveform matching audio encoder are integrated into a vehicle.

20. A method, the method comprising: One or more processors obtain an indication of the type of audio content associated with a segment of audio data; as well as Based on the instructions, the segment is selectively fed as input to a machine learning audio encoder, a waveform matching audio encoder, or both.

21. The method of claim 20, further comprising generating a bitstream representing the output of the machine learning audio encoder, the output of the waveform matching audio encoder, or both.

22. The method of claim 20, further comprising: When a specific audio encoder is selected to process two consecutive segments of the audio data, the second segment of the two consecutive segments is processed using encoder state data obtained from processing the first segment, wherein the second segment follows the first segment.

23. The method of claim 20, further comprising: When different audio encoders are selected to process two consecutive segments of the audio data, the second segment of the two consecutive segments is processed using encoder state data independent of the first segment of the two consecutive segments, wherein the second segment follows the first segment.

24. The method according to claim 20, further comprising: When different audio encoders are selected to process two consecutive segments of the audio data, the second segment of the two consecutive segments is processed using encoder state data obtained from processing the first segment, wherein the second segment follows the first segment.

25. The method according to claim 20, further comprising: A first delay is applied when the transition is made to input a segment into the waveform matching audio encoder, and a second delay is applied when the transition is made to input a segment into the machine learning audio encoder, wherein the first delay is different from the second delay.

26. The method according to claim 20, further comprising: Based on the determination of which audio encoder to which segment of the audio data is provided by the transition, at least one segment of the audio data is provided to both the machine learning audio encoder and the waveform matching audio encoder.

27. A non-transitory computer-readable medium storing instructions, the instructions being executable by one or more processors to cause the one or more processors to: Obtain an indication of the type of audio content associated with a segment of audio data; and Based on the instructions, the segment is selectively fed as input to a machine learning audio encoder, a waveform matching audio encoder, or both.

28. The non-transitory computer-readable medium of claim 27, wherein the instructions are executable to cause the one or more processors to: apply a first delay when transitioning to inputting a segment to the waveform matching audio encoder, and apply a second delay when transitioning to inputting a segment to the machine learning audio encoder, wherein the first delay is different from the second delay.

29. The non-transitory computer-readable medium of claim 27, wherein the instructions are executable to cause the one or more processors to: provide at least one segment of the audio data to both the machine learning audio encoder and the waveform matching audio encoder, based on a determination of which audio encoder to which segment of the audio data is provided.

30. An apparatus comprising: A component used to obtain an indication of the type of audio content associated with a segment of audio data; and A component for selectively transmitting the segment as input to a machine learning audio encoder, a waveform matching audio encoder, or both, based on the indicated instruction.