Bundled Multirate Feedback Autoencoder

JP2025509113A5Pending Publication Date: 2026-01-15QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024550257
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-03-21
Filing Date
2023-01-23
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing data encoding techniques face challenges in efficiently transmitting audio data over wireless communication networks, where power constraints and resource congestion are prevalent. These techniques often compromise between data fidelity and efficiency, using either more bits for higher fidelity or fewer bits for efficiency, but not effectively balancing both.

Method used

The proposed solution involves a bundled multi-rate feedback autoencoder system that dynamically adjusts bitrates for different frames in a packet. This system uses a bottleneck layer with multiple bitrates, allowing for customized bit allocation. Reference frames are encoded with more bits, while predicted frames are encoded with fewer bits, optimizing bit usage and maintaining data fidelity.

Benefits of technology

The system achieves efficient data transmission by optimizing bit allocation across frames, reducing the overall bit count in packets, and improving header efficiency and transmission bandwidth. This approach balances data fidelity and efficiency, addressing the limitations of existing techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The method includes generating an input data state for each data sample in a time series of data samples of a portion of an audio data stream. The method also includes providing at least one input data state to a first bottleneck and providing at least one other input data state to a second bottleneck. The first bottleneck is associated with a first bit rate and the second bottleneck is associated with a second bit rate. The method further includes generating a first encoded frame based on a first output data state from the first bottleneck and generating a second encoded frame based on a second output data state from the second bottleneck. The first encoded frame and the second encoded frame are bundled into a packet.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority to commonly owned Greek Provisional Patent Application No. 20220100243, filed on March 21, 2022, the entire contents of which are expressly incorporated herein by reference.

[0002] The present disclosure generally relates to encoding and / or decoding data. [Background technology]

[0003] Advances in technology have resulted in smaller and more powerful computing devices. For example, there are now a variety of portable personal computing devices, including wireless telephones, such as mobile phones and smart phones, tablet and laptop computers, that are small, lightweight, and easily carried by users. These devices may communicate voice packets, data packets, or both over wired or wireless networks. In addition, many such devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Such devices may also process executable instructions, including software applications, such as web browser applications that may be used to access the Internet. Thus, these devices may contain significant computing power.

[0004] One common use of such wireless devices is communication (e.g., voice, video, and / or data communication). In wireless communication, a device with data to send generates a signal that represents the data as a set of bits. Often, the signal also includes other information, such as a packet header. Because wireless devices are often power constrained (e.g., battery powered) and because wireless communication resources (e.g., radio frequency channels) may be congested, it may be desirable to send certain data using as few bits as possible. However, many techniques for representing data using fewer bits are lossy. That is, encoding the data to be transmitted using fewer bits leads to a less faithful representation of the data. Thus, there may be a conflict between the goal of sending a higher fidelity representation of the data to be transmitted (e.g., using a larger number of bits) and the goal of efficiently sending the data (e.g., using fewer bits). Summary of the Invention

[0005] According to a particular aspect, a device includes a memory and one or more processors coupled to the memory. The one or more processors are operatively configured to generate a first input data state for a data sample in a time series of data samples of a portion of an audio data stream. The one or more processors are also operatively configured to provide the first input data state to a first bottleneck and to provide a second input data state to a second bottleneck, the second input data state being different from the first input data state. The first bottleneck is associated with a first bitrate, and the second bottleneck is associated with a second bitrate. According to one implementation, the first bottleneck and the second bottleneck may correspond to a common bottleneck operable to dynamically change the bitrate from the first bitrate to the second bitrate. The one or more processors are further operatively configured to generate a first encoded frame based on a first output data state from the first bottleneck and generate a second encoded frame based on a second output data state from the second bottleneck. The first encoded frame and the second encoded frame are bundled into a packet.

[0006] According to another particular aspect, the method includes generating a first input data state for a data sample in a time series of data samples of a portion of an audio data stream. The method also includes providing the first input data state to a first bottleneck and providing a second input data state to a second bottleneck, the second input data state being different from the first input data state. The first bottleneck is associated with a first bitrate and the second bottleneck is associated with a second bitrate. According to one implementation, the first bottleneck and the second bottleneck may correspond to a common bottleneck operable to dynamically change the bitrate from the first bitrate to the second bitrate. The method further includes generating a first encoded frame based on the first output data state from the first bottleneck and generating a second encoded frame based on the second output data state from the second bottleneck. The first encoded frame and the second encoded frame are bundled into a packet.

[0007] According to another particular aspect, the apparatus includes means for generating a first input data state for a data sample in a time series of data samples of a portion of an audio data stream. The apparatus also includes means for providing the first input data state to a first bottleneck and for providing a second input data state to a second bottleneck, the second input data state being different from the first input data state. The first bottleneck is associated with a first bitrate, and the second bottleneck is associated with a second bitrate. According to one implementation, the first bottleneck and the second bottleneck may correspond to a common bottleneck operable to dynamically change the bitrate from the first bitrate to the second bitrate. The apparatus further includes means for generating a first encoded frame based on a first output data state from the first bottleneck and for generating a second encoded frame based on a second output data state from the second bottleneck. The first encoded frame and the second encoded frame are bundled into a packet.

[0008] According to another particular aspect, a non-transitory computer-readable medium stores instructions executable by one or more processors to generate a first input data state for data samples in a time series of data samples of a portion of an audio data stream. Execution of the instructions also causes the one or more processors to provide the first input data state to a first bottleneck and to provide a second input data state to a second bottleneck, the second input data state being different from the first input data state. The first bottleneck is associated with a first bitrate and the second bottleneck is associated with a second bitrate. According to one implementation, the first bottleneck and the second bottleneck may correspond to a common bottleneck operable to dynamically change the bitrate from the first bitrate to the second bitrate. Execution of the instructions further causes the one or more processors to generate a first encoded frame based on a first output data state from the first bottleneck and to generate a second encoded frame based on a second output data state from the second bottleneck. The first encoded frame and the second encoded frame are bundled into a packet.

[0009] According to another particular aspect, a device includes a memory and one or more processors coupled to the memory. The one or more processors are operatively configured to receive a packet at a decoder network including a first encoded frame bundled with a second encoded frame. The first encoded frame includes a first output data state generated from a first bottleneck of a feedback autoencoder, and the second encoded frame includes a second output data state generated from a second bottleneck of the feedback autoencoder. The first bottleneck is associated with a first bitrate, and the second bottleneck is associated with a second bitrate. According to one implementation, the first bottleneck and the second bottleneck may correspond to a common bottleneck operable to dynamically change the bitrate from the first bitrate to the second bitrate. The one or more processors are also operatively configured to generate a reconstructed first data sample based on the first output data state. The reconstructed first data sample corresponds to a first data sample in a time sequence of data samples of a portion of the audio data stream. The one or more processors are further configured to be operable to generate a reconstructed second data sample based on the second output data state, the reconstructed second data sample corresponding to a second data sample in the time series of data samples.

[0010] According to another particular aspect, the method includes receiving a packet at a decoder network including a first encoded frame bundled with a second encoded frame. The first encoded frame includes a first output data state generated from a first bottleneck of a feedback autoencoder, and the second encoded frame includes a second output data state generated from a second bottleneck of the feedback autoencoder. The first bottleneck is associated with a first bitrate, and the second bottleneck is associated with a second bitrate. According to one implementation, the first bottleneck and the second bottleneck may correspond to a common bottleneck operable to dynamically change the bitrate from the first bitrate to the second bitrate. The method also includes generating a reconstructed first data sample based on the first output data state. The reconstructed first data sample corresponds to a first data sample in a time series of data samples of a portion of the audio data stream. The method further includes generating a reconstructed second data sample based on the second output data state. The reconstructed second data sample corresponds to a second data sample in the time series of data samples.

[0011] According to another particular aspect, the apparatus includes means for receiving a packet including a first encoded frame bundled with a second encoded frame. The first encoded frame includes a first output data state generated from a first bottleneck of a feedback autoencoder, and the second encoded frame includes a second output data state generated from a second bottleneck of the feedback autoencoder. The first bottleneck is associated with a first bit rate, and the second bottleneck is associated with a second bit rate. According to one implementation, the first bottleneck and the second bottleneck may correspond to a common bottleneck operable to dynamically change the bit rate from the first bit rate to the second bit rate. The apparatus also includes means for generating a reconstructed first data sample based on the first output data state. The reconstructed first data sample corresponds to a first data sample in a time series of data samples of a portion of the audio data stream. The apparatus further includes means for generating a reconstructed second data sample based on the second output data state. The reconstructed second data sample corresponds to a second data sample in the time series of data samples.

[0012] According to another particular aspect, a non-transitory computer-readable medium stores instructions executable by one or more processors to receive a packet at a decoder network including a first encoded frame bundled with a second encoded frame. The first encoded frame includes a first output data state generated from a first bottleneck of a feedback autoencoder, and the second encoded frame includes a second output data state generated from a second bottleneck of the feedback autoencoder. The first bottleneck is associated with a first bitrate, and the second bottleneck is associated with a second bitrate. According to one implementation, the first bottleneck and the second bottleneck may correspond to a common bottleneck operable to dynamically change the bitrate from the first bitrate to the second bitrate. Execution of the instructions also causes the one or more processors to generate a reconstructed first data sample based on the first output data state. The reconstructed first data sample corresponds to a first data sample in a time sequence of data samples of a portion of the audio data stream. Execution of the instructions further causes the one or more processors to generate a reconstructed second data sample based on the second output data state, the reconstructed second data sample corresponding to a second data sample in the time series of data samples. [Brief description of the drawings]

[0013] [Figure 1] FIG. 1 is a diagram of a particular illustrative embodiment of a system configured to encode data using multiple bit rates for different frames in a packet and jointly code them using a bidirectional gated recursive unit. [Diagram 2] FIG. 1 is a diagram of another particular illustrative embodiment of a system configured to encode data using multiple bit rates for different frames in a packet and jointly code them using a transformer-like attention mechanism. [Diagram 3]FIG. 3 is a diagram of a particular illustrative embodiment of a bottleneck architecture integrated into the system of FIG. 1 or the system of FIG. 2. [Figure 4] FIG. 2 is a diagram of a particular illustrative embodiment of a system operable to bundle frames generated at different bit rates into packets. [Diagram 5] FIG. 2 is a diagram of a particular illustrative embodiment of a system operable to decode bottleneck outputs generated at different bit rates. [Figure 6] FIG. 1 is a diagram of a particular illustrative embodiment of a system including two or more devices configured to communicate via transmission of encoded data. [Figure 7] 4 is a flow diagram of a particular embodiment of a method of operation of an encoding device. [Figure 8] 4 is a flow diagram of a particular embodiment of a method of operation of a decoding device; [Figure 9] 2 is a diagram of a particular embodiment of components of the encoding device of FIG. 1 in an integrated circuit. [Figure 10] 2 is a diagram of a particular embodiment of components of the decoding device of FIG. 1 in an integrated circuit. [Figure 11] FIG. 1 is a block diagram of a particular illustrative embodiment of a device operable to perform encoding, decoding, and both. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0014] An encoder may be used to encode data samples into frames that are transmitted to a receiving device. In some scenarios, the use of an encoder may be inefficient because frames are not bundled into packets. For example, a separate packet (including a large header section) may be used to send each encoded frame to the receiving device. Using a separate packet for each frame results in multiple bits being allocated to the header rather than the data portion of the frame, which is inefficient. In other scenarios where frames encoded using FRAE are bundled into packets, FRAE typically uses the same number of bits to encode each frame in a packet. Using the same number of bits to encode each frame in a packet also results in a relatively large amount of bits being allocated, which is inefficient and may result in an increase in transmission bandwidth.

[0015] Aspects disclosed herein enable a bundled multi-rate feedback autoencoder to selectively allocate different numbers of bits to frames and encode the bundled frames into packets. For example, aspects disclosed herein exploit redundant data in frames bundled into packets by allocating a relatively large amount of bits to a reference frame and a smaller amount of bits to other frames (e.g., "predicted frames") in the packet that may be reconstructed using data from the reference frame. To facilitate packetization of the frames, data associated with the reference frame may be encoded at a different bit rate than data associated with the predicted frame to allocate different numbers of bits to different frames to encode the frames at a similar frequency. For example, a bundled multi-rate feedback autoencoder may include a multi-rate bottleneck layer that encodes data at different bit rates. Illustratively, a first bottleneck layer may encode data of a first frame at a first bit rate, and a second bottleneck layer may encode data of a second frame at a second bit rate that is greater than the first bit rate. By encoding the data of the second frame at a higher bit rate, a larger number of bits can be allocated to the second frame while encoding each frame at a similar frequency (e.g., packet frequency). Thus, the techniques described herein allow frames encoded using bundled multi-rate feedback autoencoders to be bundled into packets, thereby improving header efficiency. In addition, the techniques described herein allow different amounts of bits to be allocated to frames in a packet, thereby improving transmission bandwidth.

[0016] Unless expressly limited by its context, the term "generate" is used herein to indicate any of its general meanings, such as calculating, or, in some cases, producing. Unless expressly limited by its context, the term "compute" is used herein to indicate any of its general meanings, such as calculating, determining a value, smoothing, and / or selecting from among multiple values. Unless expressly limited by its context, the term "obtain" is used herein to indicate any of its general meanings, such as calculating, deriving, receiving (e.g., from another component, block, or device), and / or reading (e.g., from a memory register or array of storage elements).

[0017] Unless expressly limited by the context, the term "produce" is used to indicate any of its general meanings, such as calculating, generating, and / or providing. Unless expressly limited by the context, the term "provide" is used to indicate any of its general meanings, such as calculating, generating, and / or producing. Unless expressly limited by the context, the term "coupled" is used to indicate a direct or indirect electrical or physical connection. If the connection is indirect, there may be other blocks or components between the structures that are "coupled." For example, a loudspeaker may be acoustically coupled to a nearby wall through an intervening medium (e.g., air) that allows for the propagation of waves (e.g., sound) from the loudspeaker to the wall (or vice versa).

[0018] The term "configuration" may be used in reference to a method, an apparatus, a device, a system, or any combination thereof, as indicated by its particular context. The term "comprising" is used in this specification and claims and does not exclude other elements or operations. The term "based on" (as in "A is based on B") is used to indicate any of its ordinary meanings, including (i) "based at least on" (e.g., "A is based at least on B"), and, if appropriate in the particular context, (ii) "equal to" (e.g., "A is equal to B"). In the case of (i) A is based on B, but including at least based on, this may include configurations in which A is coupled to B. Similarly, the term "responsive to" is used to indicate any of its ordinary meanings, including "responsive to at least". The term "at least one" is used to indicate any of its ordinary meanings, including "one or more". The term "at least two" is used to indicate any of its ordinary meanings, including "two or more".

[0019] The terms "apparatus" and "device" are used collectively and interchangeably unless otherwise indicated by specific context. Any disclosure of the operation of an apparatus having a particular feature is also expressly intended to disclose a method having similar features (and vice versa), and any disclosure of the operation of an apparatus with a particular configuration is also expressly intended to disclose a method with a similar configuration (and vice versa). The terms "method," "process," "procedure," and "technique" are used collectively and interchangeably, unless otherwise indicated by specific context. The terms "element" and "module" may be used to indicate a portion of a larger configuration. The term "packet" may correspond to a unit of data that includes a header portion and a payload portion. Any incorporation by reference of a portion of a document should also be understood to incorporate definitions of terms or variables referenced within that portion, if such definitions appear elsewhere in that document, as well as in any figures referenced in the incorporated portion.

[0020] As used herein, the term "communication device" refers to an electronic device that can be used for voice and / or data communication over a wireless communication network. Examples of communication devices include speaker bars, smart speakers, cellular phones, personal digital assistants (PDAs), handheld devices, headsets, wireless modems, laptop computers, personal computers, etc.

[0021] Various aspects will now be described with reference to the drawings. In the description, common features are designated by common reference numbers throughout the drawings. In some drawings, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference number is used for each, and the different instances are distinguished by the addition of a letter to the reference number. When features as a group or type are referred to herein (e.g., when no particular one of the features is referred to), the reference number is used without an identifying letter. However, when one particular feature of multiple features of the same type is referred to herein, the reference number is used with an identifying letter. For example, with reference to FIG. 1, multiple data samples are shown and associated with reference numbers 120A, 120B, 120C, etc. When referring to a particular one of these data samples, e.g., data sample 120A, the identifying letter "A" is used. However, when referring to any one of these data samples, or to these data samples as a group, the reference number 120 is used without an identifying letter.

[0022] 1 is a diagram of a particular illustrative embodiment of a system 100 configured to encode data using multiple bit rates for different frames in a packet and jointly code them using bidirectionally gated recursive units. For example, the system 100 may use a neural network architecture with a bottleneck operating at different bit rates to generate output data states 124 representing encoded data samples. The encoded data samples may be transmitted as frames bundled in packets to a receiving device, such as the receiving device 652 of FIG. 6, which may then decode the frames for playout. Thus, the system 100 may be integrated into a transmitting device, such as the transmitting device 602 of FIG. 6, configured to send one or more encoded data packets to the receiving device.

[0023] According to one implementation, the system 100 corresponds to a bundled multi-rate feedback autoencoder architecture. The system 100 includes an encoder portion 180A and a decoder portion 190A that provides feedback to the encoder portion 180A. The encoder portion 180A of the system 100 includes one or more front-end neural network preprocessing layers 102, a bidirectional gated recurrent unit (GRU) layer 105, and a bottleneck layer 107.

[0024] The bidirectional GRU layer 105 includes a bidirectional GRU network implemented across time instances 106A-106E. For example, the bidirectional GRU layer 105 includes a bidirectional GRU network implemented across time instances 106A, time instance 106B, time instance 106C, time instance 106D, and time instance 106E. Although the bidirectional GRU network of the bidirectional GRU layer 105 is shown as being implemented across five time instances 106A-106E, in other implementations, the bidirectional GRU network of the bidirectional GRU layer 105 may be implemented across fewer time instances 106 or across more time instances 106. As a non-limiting example, the bidirectional GRU network of the bidirectional GRU layer 105 may be implemented across four time instances 106. As another non-limiting example, the bidirectional GRU network of the bidirectional GRU layer 105 may be implemented across six time instances 106.

[0025] According to some implementations, the front-end neural network preprocessing layer(s) 102 may include a front-end GRU rather than a fully connected layer, and the number of outputs of the front-end neural network preprocessing layer(s) 102 may differ from the number of data samples 120 provided to the front-end neural network preprocessing layer(s) 102. For example, the front-end neural network preprocessing layer(s) 102 may summarize or upsample the input (e.g., data sample 120) based on the number of times a front-end GRU (in the front-end neural network preprocessing layer(s) 102) is tapped. In these scenarios, there may not be a one-to-one correspondence between the data samples 120A-120E and the time instances 106A-106E.

[0026] The bottleneck layer 107 includes multiple bottlenecks 108A-108E. For example, the bottleneck layer 107 includes bottleneck 108A, bottleneck 108B, bottleneck 108C, bottleneck 108D, and bottleneck 108E. Although five bottlenecks 108 are shown, in other implementations, the bottleneck layer 107 can include fewer (or additional) bottlenecks 108. As a non-limiting example, the bottleneck layer 107 can include four bottlenecks 108. As another non-limiting example, the bottleneck layer 107 can include six bottlenecks 108. The architecture of the bottlenecks 108 is described in more detail with respect to FIG. 3.

[0027] The decoder portion 190A of the system 100 includes a bidirectional GRU layer 109 and one or more back-end neural network post-processing layers 112. The bidirectional GRU layer 109 includes a bidirectional GRU network implemented across time instances 110A-110E. For example, the bidirectional GRU layer 109 includes a bidirectional GRU network implemented across time instances 110A, time instance 110B, time instance 110C, time instance 110D, and time instance 110E. Although the bidirectional GRU network of the bidirectional GRU layer 109 is shown as being implemented across five time instances 110A-110E, in other implementations the bidirectional GRU network of the bidirectional GRU layer 109 may be implemented across fewer or more time instances 110. As a non-limiting example, the bidirectional GRU network of the bidirectional GRU layer 109 may be implemented across four time instances 110. As another non-limiting example, the bidirectional GRU network of the bidirectional GRU layer 109 may be implemented at six time instances 110. According to some implementations, the back-end neural network post-processing layer(s) 112 may include back-end GRU(s) rather than a fully connected layer, and the number of outputs of the back-end neural network post-processing layer(s) 112 may differ from the inputs provided to the back-end neural network post-processing layer 112. In these scenarios, there may not be a one-to-one correspondence between the reconstructed data samples 126A-126E and the time instances 110A-110E.

[0028] A data stream including time-ordered data may be provided to the encoder portion 180A of the system 100. For example, the data stream may include a time sequence of data samples 120, with each data sample 120 representing a time-windowed portion of the data. Although described as a "data sample," in other implementations, each data sample 120 may correspond to a "frame" of the data stream. As shown in FIG. 1, the data samples 120 include data sample 120A, data sample 120B, data sample 120C, data sample 120D, and data sample 120E. Data sample 120A includes data (e.g., extracted features) generated at an earlier time instance than data included in data sample 120B, and data sample 120B includes data generated at an earlier time instance than data included in data sample 120C. According to some implementations, adjacent data samples 120 may include overlapping data (e.g., temporal redundancy). For example, a portion of the data in data sample 120A may also be included in data sample 120B. In some examples, the data included in data sample 120 includes media data, such as voice data, audio data, video data, gaming data, augmented reality data, other media data, or combinations thereof. As described below and further shown in FIG. 4, each data sample 120A-120E may be encoded into a corresponding frame, such as frames 420A-420E, and packetized for transmission to a receiving device.

[0029] To reduce the amount of bits used to encode the data samples 120, the system 100 may select data samples (e.g., reference frame data samples) to be encoded into a reference frame (e.g., frame 420A) and may select data samples (e.g., predicted frame data samples) to be encoded into a predicted frame (e.g., frames 420B-420E). For ease of illustration and explanation, the reference frames and data samples associated with the reference frames are shaded gray. The temporal redundancy discussed above is exploited during encoding and decoding of the data samples 120 by using the reference frames to predict other frames.

[0030] To illustrate, the system 100 may designate M bits for encoding the reference frame data sample and N bits for encoding each of the predicted frame data samples, where M is significantly greater than N (e.g., M>>N). For illustrative purposes, the data sample 120A is designated as a reference frame data sample, indicated by gray shading, and the data samples 120B-120E are designated as predicted frame data samples. Thus, in the example of FIG. 1, the total number of bits for encoding the data samples 120A-120E is equal to 4×N+M. Because M is significantly greater than N, the total number of bits for encoding the data samples 120A-120E is significantly less than encoding each data sample 120A-120E as a reference frame data sample (e.g., 4×N+M<<5×M). Thus, the total number of bits for encoding data samples 120A-120E using system 100 is significantly less than the 5×M bits that may be used by a conventional feedback recurrent autoencoder (FRAE) architecture.

[0031] However, given the increased number of bits used to encode the reference frame data sample 120A, if each of the data samples 120A-120E are encoded at the same frequency (e.g., packet frequency), the reference frame data sample 120A must be encoded at an increased bit rate (compared to the bit rate of the predicted frame data samples 120B-120E). Thus, as described below, the bottleneck 108A associated with the reference frame data sample 120A can operate at a higher bit rate than the bottlenecks 108B-108E associated with the predicted frame data samples 120B-120E, such that each data sample 120 is encoded into a frame at the same frequency (e.g., packet frequency) to facilitate bundling the corresponding frames 420A-420E into a packet 430, as shown in FIG. 4. It should be understood that the names in FIG. 1 are for illustrative purposes only and different names may be used. As a non-limiting example, two or more data samples may be designated as reference frame data samples, or different data samples may be designated as reference frame data samples.

[0032] The data samples 120A-120E are provided to one or more front-end neural network pre-processing layers 102. In some implementations, the one or more front-end neural network pre-processing layers 102 include one or more fully-connected layers. As described herein, a "fully-connected layer" is a feed-forward neural network that includes multiple input nodes and generates one or more outputs based on a weighting function and a mapping function. According to some implementations, a fully-connected layer may include multiple node levels (e.g., input level nodes, intermediate level nodes, and output level nodes) with unique weighting and mapping patterns. For ease of explanation, a fully-connected layer is described as receiving one or more inputs and generating one or more outputs based on neural network operations. However, it should be understood that the architecture of each fully-connected layer described herein can be unique and can have unique weighting and mapping patterns to generate simple or complex neural networks. One or more front-end neural network pre-processing layers 102 may receive the data samples 120A-120E and generate corresponding neural network encoded data that is provided to the bidirectional GRU layer 105.

[0033] The bidirectional GRU layer 105 is configured to generate, at each time instance 106A-106E, an input data state 122A-122E, respectively, which is provided to a corresponding bottleneck 108A-108E in the bottleneck layer 107. As non-limiting examples, the bidirectional GRU layer 105 is configured to generate input data state 122A associated with data sample 120A at time instance 106A, the bidirectional GRU layer 105 is configured to generate input data state 122B associated with data sample 120B at time instance 106B, the bidirectional GRU layer 105 is configured to generate input data state 122C associated with data sample 120C at time instance 106C, the bidirectional GRU layer 105 is configured to generate input data state 122D associated with data sample 120D at time instance 106D, and the bidirectional GRU layer 105 is configured to generate input data state 122E associated with data sample 120E at time instance 106E.

[0034] The bidirectional GRU layer 105 may also generate each of the input data states 122A-122E based on data associated with previous and future time steps. For example, each of the bidirectional GRU layer 105 may access input data states 122 from adjacent time instances 106 to generate each of the input data states 122A-122E. Thus, the bidirectional GRU layer 105 may generate the input data states in a manner that takes into account data states associated with other data samples 120.

[0035] Additionally, the input data states 122A-122E may be generated based on feedback 150 (e.g., output data state 124) from previously decoded data samples. To illustrate, the decoder portion 190A of the system 100 may provide feedback 150 to the bidirectional GRU layer 105. The feedback 150 may include data states (e.g., output data states) of previously decoded packets. As a result, the bidirectional GRU layer 105 may encode data associated with the data samples 120A-120E (e.g., outputs of one or more front-end neural network pre-processing layers 102) in a manner that takes into account previously encoded / decoded data samples. Although not shown in FIG. 1, the previously decoded frames may correspond to decoded frames in packet 634 of FIG. 6 (e.g., packets generated prior to packet 430 associated with encoded data sample 120).

[0036] In Figure 1, input data states 122A-122E may correspond to encoder hidden states generated by bidirectional GRU layer 105. According to the example of Figure 1, input data state 122A is provided to bottleneck 108A, input data state 122B is provided to bottleneck 108B, input data state 122C is provided to bottleneck 108C, input data state 122D is provided to bottleneck 108D, and input data state 122E is provided to bottleneck 108E.

[0037] The bottlenecks 108B-108E are associated with a first bit rate, and the bottleneck 108A is associated with a second bit rate that is greater than the first bit rate. That is, the bit rate of the bottleneck 108A associated with the data sample 120A is higher than the bit rates of the bottlenecks 108B-108E associated with the other data samples 120B-120E because more bits are used to encode the data sample 120A into a reference frame (e.g., frame 420A) than to encode the data samples 120B-120E into a predicted data frame (e.g., frame 420B-420E). Although the bottlenecks 108B-108E are shown as four separate bottlenecks, they may be a single bottleneck that encodes the input data states 122B-122E at different time instances. According to one implementation, to reduce the bitrates of bottlenecks 108B-108E compared to the bitrate of bottleneck 108A, fewer bits (e.g., units) are assigned to the latent codes generated in bottlenecks 108B-108E compared to the number of bits (e.g., units) assigned to the latent codes generated in bottlenecks 108A, as further described with respect to Figure 3. Additionally or alternatively, to reduce the bitrates of bottlenecks 108B-108E compared to the bitrate of bottleneck 108A, the codebooks associated with bottlenecks 108B-108E may have a smaller size (or allocate fewer bits) than the codebooks associated with bottleneck 108A, as further described with respect to Figure 3.

[0038] Bottleneck 108A is configured to generate output data state 124A based on input data state 122A. According to one implementation, output data state 124A may correspond to a latent code (e.g., a post-quantization latent code) generated based on input data state 122A using a codebook, as further described with respect to FIG. 3. Similarly, bottleneck 108B is configured to generate output data state 124B based on input data state 122B, bottleneck 108C is configured to generate output data state 124C based on input data state 122C, bottleneck 108D is configured to generate output data state 124D based on input data state 122D, and bottleneck 108E is configured to generate output data state 124E based on input data state 122E. Generation of output data state 124 is described in more detail with respect to FIG. 3. As will be described in more detail with respect to FIG. 4, the output data states 124A-124E may be included in corresponding encoded data frames 420A-420E that are bundled into packets 430 and transmitted to a receiving device.

[0039] The decoder section 190A of the system 100 is configured to reconstruct the data samples 120A-120E based on the output data states 124A-124E. By way of example, the output data states 124 are provided to the bidirectional GRU layer 109. The bidirectional GRU layer 109 may use feedback 150 (e.g., output data states from a previous packet) to initialize the bidirectional GRU layer 109. The bidirectional GRU layer 109 may perform decoding operations on the output data states 124 based on the feedback 150 to generate outputs that are provided to one or more back-end neural network post-processing layers 112. The operation of the bidirectional GRU layer 109 is described in more detail with respect to FIG. 5. The one or more back-end neural network post-processing layers 112 are configured to reconstruct the data samples 120D-120A to generate reconstructed data samples 126D-126A, respectively.

[0040] It should be appreciated that the system 100 of FIG. 1 can preserve the temporal structure of the data samples 120 during encoding. For example, instead of concatenating features of the data samples 120 to generate a single input data state that removes the temporal structure, the bidirectional GRU layer 105 generates input data states 122A-122E for each data sample 120A-120E such that the temporal structure is preserved. Thus, the system 100 reduces the complexity of reconstructing a frame by preserving the temporal structure in the code that might otherwise be lost (or implicitly learned). Preservation of the temporal structure results in improved, more efficient reconstruction of the data samples at the decoder portion 190A of the system 100 (or at a decoder of a receiving device). In addition, using the bidirectional GRU layer 105 rather than a conventional GRU layer allows for exploiting temporal redundancy during encoding. For example, as described above, the bidirectional GRU layer 105 can access data states from previous and future time steps to exploit temporal redundancy in the generation of the input data states 122A-122E.

[0041] The system 100 further allows customized bit allocation for encoding of different data samples 120. For example, a larger number of bits are allocated to the reference frame data sample (e.g., data sample 120A) than to the predicted frame data sample (e.g., data samples 120B-120E). Thus, if the data samples 120A-120E are encoded into five respective frames (e.g., frames 420A-420E in FIG. 4) bundled into a packet (e.g., packet 430 in FIG. 4), the boundary frame of the packet 430 (e.g., frame 420A associated with data sample 120A) may be encoded with additional bits to be used as a reference frame. As a result, the remaining frames 420B-420E may be encoded with a relatively small number of bits, which reduces the amount of bits used to encode the frames bundled into the packet 430. It should be understood that the system 100 has the flexibility to allocate additional bits to any data sample 120 and increase the bit rate of the corresponding bottleneck 108 in order to encode the data sample 120 at the packet frequency. In some scenarios, the system 100 may allocate zero bits for encoding to a particular data sample 120. In these scenarios, nothing is transmitted for a particular data sample 120, and the decoder will predict the corresponding frame based on the encoded data in adjacent frames.

[0042] According to some implementations, system complexity is reduced by sharing the front-end neural network pre-processing layer 102 and the back-end neural network post-processing layer 112 across time steps. For example, the front-end neural network pre-processing layer 102 and the back-end neural network post-processing layer 112 can perform pre-processing and post-processing operations in parallel. As a result, the system 100 can have reduced memory usage and reduced complexity. In addition, by sharing the front-end neural network pre-processing layer 102 and the back-end neural network post-processing layer 112 across time steps, the system 100 can adapt to network conditions on the fly by changing the bit rate with reduced weight loads. For example, the bit rate of the bottleneck 108 can change in response to changes in network conditions, but the front-end neural network pre-processing layer 102, the bidirectional GRU layer 105, and the back-end neural network post-processing layer 112 can remain unchanged.

[0043] 2 is a diagram of another particular illustrative embodiment of a system 200 configured to encode data using multiple bit rates for different frames in a packet and jointly code them using a transformer-like attention mechanism. Similar to the system 100 of FIG. 1, the system 200 may use a neural network architecture to generate output data states 124 representing encoded audio data samples.

[0044] System 200 of Figure 2 has a substantially similar architecture and operates in a substantially similar manner as system 100 of Figure 1. However, system 200 of Figure 2 replaces bidirectional GRU layers 105 and 109 with encoder-side attention mechanism 205 and decoder-side attention mechanism 209, respectively. For example, encoder portion 180B of system 200 includes encoder-side attention mechanism 205 rather than bidirectional GRU layer 105, and decoder portion 190B of system 200 includes decoder-side attention mechanism 209 rather than bidirectional GRU layer 109.

[0045] The encoder-side attention mechanism 205 receives the output of the front-end neural network pre-processing layer 102 and can receive feedback 150 from a decoded frame of a previous packet. According to one implementation, the encoder-side attention mechanism 205 includes a transformer. Instead of accessing adjacent tokens, the encoder-side attention mechanism 205 can directly access each token (e.g., each input data state 122A-122E). For example, in FIG. 2, the encoder-side attention mechanism 205 can directly access each input data state 122A-122E and the feedback 150. As a result, the encoder-side attention mechanism 205 can generate the input data state 122 in a manner that takes into account data associated with other data samples 120A, 120B, 120D, 120E and in a manner that takes into account previously encoded / decoded data samples.

[0046] The decoder-side attention mechanism 209 can receive the output data states 124A-124E from the bottlenecks 108A-108E and can receive feedback 150 from a decoded frame of a previous packet. The decoder-side attention mechanism 209 can perform decoding operations on the output data states 124A-124E based on the feedback 150 to generate an output that is provided to the back-end neural network post-processing layer 112 for processing. According to one implementation, the decoder-side attention mechanism 209 includes a transformer.

[0047] Figure 3 is a diagram of a particular example embodiment of a bottleneck architecture integrated into the system of Figure 1 or the system of Figure 2. For example, a non-limiting example of bottleneck 108A is shown in Figure 3, and a non-limiting example of bottleneck 108B is shown in Figure 3. As discussed above, bottleneck 108A is associated with a higher bitrate than bottleneck 108B.

[0048] The bottleneck 108A includes a fully connected layer 302A, a quantizer 304A, one or more codebooks 306A, and a fully connected layer 308A. The input data state 122A is provided to the fully connected layer 302A. The fully connected layer 302A is configured to generate a pre-quantization potential 350A based on the input data state 122A. The pre-quantization potential 350A may correspond to an encoding indicating an array of floating-point values. The pre-quantization potential 350A is provided to the quantizer 304A. The quantizer 304A is configured to map each floating-point value of the pre-quantization potential 350A to a representative value of one or more codebooks 306A to generate a post-quantization potential 352A. According to one implementation, the post-quantization potential 352A may correspond to an output data state 124A of the bottleneck 108A. According to another implementation, the quantized latents 352A can be provided to a fully connected layer 308A, which can generate the output data state 124A based on the quantized latents 352A.

[0049] The bottleneck 108B includes a fully connected layer 302B, a quantizer 304B, one or more codebooks 306B, and a fully connected layer 308B. The input data state 122B is provided to the fully connected layer 302B. The fully connected layer 302B is configured to generate a pre-quantized potential 350B based on the input data state 122B. The pre-quantized potential 350B may correspond to an encoding indicating an array of floating-point values. The pre-quantized potential 350B is provided to the quantizer 304B. The quantizer 304B is configured to map each floating-point value of the pre-quantized potential 350B to a representative value of one or more codebooks 306B to generate a quantized potential 352B. According to one implementation, the quantized potential 352B may correspond to an output data state 124B of the bottleneck 108B. According to another implementation, the quantized latents 352B can be provided to a fully connected layer 308B, which can generate the output data state 124B based on the quantized latents 352B.

[0050] As described above with respect to FIG. 1, the bottleneck 108B is associated with a first bit rate, and the bottleneck 108A is associated with a second bit rate that is greater than the first bit rate. That is, the bit rate of the bottleneck 108A associated with the data sample 120A is higher than the bit rates of the bottlenecks 108 associated with the other data samples because more bits are used to encode the input data state 122A (e.g., the input data state associated with the data sample 120A) into a reference frame than to encode the input data state 122B (e.g., the input data state associated with the data sample 120b) into a predicted data frame. According to one implementation, to reduce the bit rate of the bottleneck 108B compared to the bit rate of the bottleneck 108A, fewer bits (e.g., units) are assigned to the latent code generated at the bottleneck 108B compared to the number of units assigned to the latent code generated at the bottleneck 108A. For example, the quantized latent 352B may have fewer units (e.g., bits) than the quantized latent 352A. Additionally or alternatively, to reduce the bitrate of bottleneck 108B compared to the bitrate of bottleneck 108A, one or more codebooks 306B associated with bottleneck 108B may have a smaller size (or may be allocated fewer bits) than one or more codebooks 306A associated with bottleneck 108A. In some scenarios, to reduce the bitrate of bottleneck 108B compared to the bitrate of bottleneck 108A, fewer quantization stages may be used in bottleneck 108B compared to the number of quantization stages used in bottleneck 108B.

[0051] 4 is a diagram of a particular illustrative embodiment of a system 400 operable to bundle frames generated at different bit rates into packets. System 400 includes a frame generator 402 and a packet generator 404. The components of system 400 may be integrated into a transmitting device, including system 100 of FIG. 1 or system 200 of FIG. 2.

[0052] The frame generator 402 is configured to receive the output data states 124A-124E from each bottleneck 108A-108E and generate corresponding frames 420A-420E. By way of example, the frame generator 402 can generate the frame 420A based on the output data state 124A. For example, the output data state 124A may correspond to an encoded version of the data samples 120A included in the frame 420A. The frame generator 402 can also generate the frame 420B based on the output data state 124B. For example, the output data state 124B may correspond to an encoded version of the data samples 120B included in the frame 420B. The frame generator 402 can also generate the frame 420C based on the output data state 124C. For example, the output data state 124C may correspond to an encoded version of the data samples 120C included in the frame 420C. The frame generator 402 may also generate a frame 420D based on the output data state 124D. For example, the output data state 124D may correspond to an encoded version of the data samples 120D included in the frame 420D. The frame generator 402 may also generate a frame 420e based on the output data state 124E. For example, the output data state 124E may correspond to an encoded version of the data samples 120E included in the frame 420E.

[0053] As discussed above, frame 420A may correspond to a reference frame that includes more bits than frames 420B-420E. Because of the additional bits associated with reference frame 420A, the bottleneck 108A associated with generating the corresponding output data state 124A operates at a higher bit rate than the bottlenecks 108B-108E associated with generating the other output data states 124B-124E.

[0054] The packet generator 404 is configured to receive each frame 420A-420E from the frame generator 402. The packet generator 404 may be further configured to bundle the frames 420A-420E into a packet 430 to be transmitted to a receiving device. For example, the packet generator 404 may operate as a frame bundler that bundles (e.g., combines) the frames 420A-420E into a single packet 430. Because the frames 420A-420E are jointly coded and temporal redundancy is exploited during the coding process, the frames 420A-420E are smaller in size than frames generated without exploiting temporal redundancy, and the number of bits that make up the packet 430 is reduced.

[0055] 5 is a diagram of a particular illustrative embodiment of a system 500 operable to decode bottleneck outputs generated at different bit rates. According to one implementation, the system 500 may correspond to the decoder portion 190A of the system 100. According to another implementation, the system 500 may correspond to a stand-alone decoder in a receiving device that receives the packets 430.

[0056] System 500 includes a bidirectional GRU layer 109 and one or more back-end neural network post-processing layers 112. At each time instance 110, bidirectional GRU layer 109 is configured to generate a left hidden data state and a right hidden data state to facilitate decoding of data samples 120. Thus, in the example of Figure 5, bidirectional GRU layer 109 is shown as generating a left data state 520 and a right data state 522 at different time instances 110.

[0057] The bidirectional GRU layer 109 may be initialized based on data states from a previous frame or packet. For example, at time instance 110, the bidirectional GRU layer 109 may receive a left data state 530 from the previous frame and a right data state 532 from the previous frame. According to one implementation, the data states 530, 532 from the previous frame correspond to the feedback 150 provided to the GRU layer 109. The bidirectional GRU layer 109 may use the data states 530, 532 from the previous frame that were used in a scheme that takes into account previously encoded / decoded data samples.

[0058] The bidirectional GRU layer 109 is configured to generate left data state 520E based on left data state 530 from the previous frame and output data state 124E. According to some implementations, the bidirectional GRU layer 109 can generate left data state 520E based on left data state 520D. Based on left data state 520E, output data state 124D, left data state 520C, or a combination thereof, the bidirectional GRU layer 109 is configured to generate left data state 520D. Based on left data state 520D, output data state 124C, left data state 520B, or a combination thereof, the bidirectional GRU layer 109 is configured to generate left data state 520C. Based on left data state 520C, output data state 124B, left data state 520A, or a combination thereof, the bidirectional GRU layer 109 is configured to generate left data state 520B. Based on left data state 520B, output data state 124A, or both, bidirectional GRU layer 109 is configured to generate left data state 520A.

[0059] Based on the right data state 532 from the previous frame and the output data state 124E, the bidirectional GRU layer 109 is configured to generate the right data state 522A. According to some implementations, the bidirectional GRU layer 109 can generate the right data state 522A based on the right data state 522B. Based on the right data state 522A, the output data state 124B, the right data state 522C, or a combination thereof, the bidirectional GRU layer 109 is configured to generate the right data state 522B. Based on the right data state 522B, the output data state 124C, the right data state 522D, or a combination thereof, the bidirectional GRU layer 109 is configured to generate the right data state 522C. Based on the right data state 522C, the output data state 124D, the right data state 522E, or a combination thereof, the bidirectional GRU layer 109 is configured to generate the right data state 522D. Based on right data state 522D, output data state 124E, or both, bidirectional GRU layer 109 is configured to generate right data state 522E.

[0060] The left data state 520A and the right data state 522A can be provided to one or more back-end neural network post-processing layers 112, which are configured to generate reconstructed data samples 126A based on the data states 520A, 522A. Similarly, the left data state 520B and the right data state 522B can be provided to one or more back-end neural network post-processing layers 112, which are configured to generate reconstructed data samples 126B based on the data states 520B, 522B. Similarly, the left data state 520C and the right data state 522C can be provided to one or more back-end neural network post-processing layers 112, which are configured to generate reconstructed data samples 126C based on the data states 520C, 522C.

[0061] Similarly, the left data state 520D and the right data state 522D can be provided to one or more back-end neural network post-processing layers 112, which are configured to generate reconstructed data samples 126D based on the data states 520D, 522D. Similarly, the left data state 520E and the right data state 522E can be provided to one or more back-end neural network post-processing layers 112, which are configured to generate reconstructed data samples 126E based on the data states 520E, 522E.

[0062] 5, it should be appreciated that the bidirectional GRU layer 109 can jointly decode a frame 420 (e.g., an output data state 124 in a frame 420) of a packet 430 based on knowledge of a “future” frame. For example, the bidirectional GRU layer 109 can access and use data states 520, 522 associated with a future frame or time step to generate the respective data states 520, 522.

[0063] FIG. 6 is a diagram of a particular illustrative embodiment of a system 600 including two or more devices configured to communicate via transmission of encoded data. The example of FIG. 6 shows a first device 602 configured to encode and transmit data and a second device 652 configured to receive, decode, and use the data. For ease of reference herein, the first device 602 is also referred to herein as an encoding device and / or a transmitting device, and the second device 652 is also referred to herein as a decoding device and / or a receiving device. Although the system 600 shows one transmitting device 602, the system 600 can include two or more transmitting devices 602. For example, a two-way communication system may include two devices (e.g., mobile phones), each of which may transmit data to the other device and receive data from the other device. That is, each device may operate as both a transmitting device 602 and a receiving device 652. In another example, a single receiving device 652 may receive data from two or more transmitting devices 602. Additionally or alternatively, the system 600 may include two or more receiving devices 652. For example, a single transmitting device 602 may transmit (e.g., multicast or broadcast) data to multiple receiving devices 652. Thus, the one-to-one pairing of transmitting devices 602 and receiving devices 652 shown in FIG. 6 is merely illustrative of one configuration and is not intended to be limiting.

[0064] In the example of FIG. 6, the transmitting device 602 includes multiple components configured to obtain data from a data stream 604 and process the data to generate data packets (e.g., data packet 634 and data packet 430) that are transmitted over a transmission medium 632. In FIG. 6, the components of the transmitting device 602 include a feature extractor 606, a subsystem 610, a frame generator 402, a packet generator 404, a modem 628, and a transmitter 630. The subsystem 610 may correspond to the system 100 of FIG. 1, the system 200 of FIG. 2, or both. In other examples, the transmitting device 602 may include more, fewer, or different components. To illustrate, in some examples, the transmitting device 602 includes one or more data generating devices configured to generate the data stream 604. Examples of such data generating devices include, but are not limited to, for example, a microphone, a camera, a game engine, a media processor (e.g., a computer-generated image engine), an augmented reality engine, a sensor, or other devices and / or instructions configured to output the data stream 604. To further illustrate, in some examples, the transmitting device 602 includes a transceiver in place of (or in the location of) the transmitter 630.

[0065] 1 includes data arranged in a time sequence. For example, the data stream 604 can include a sequence of data frames, with each data frame representing a time-windowed portion of the data. In some examples, the data includes media data, such as voice data, audio data, video data, gaming data, augmented reality data, other media data, or a combination thereof.

[0066] The feature extractor 606 is configured to generate data samples 120 based on the data stream 604. The data samples 120 include data representing a portion of the data stream 604 (e.g., a single frame of data, multiple frames of data, or a segment or subset of a frame of data). The feature extraction technique(s) used by the feature extractor 606 may include, for example, data aggregation, interpolation, compression, windowing, domain transformation, sampling, smoothing, statistical analysis, and the like. By way of example, if the data stream 604 includes voice data or other audio data, the feature extractor 606 may be configured to determine time-domain or frequency-domain spectral information describing a time-windowed portion of the data stream 604. In this example, the data samples 120 may include spectral information. As one non-limiting example, the data samples 120 may include data describing a cepstrum of the voice data of the data stream 604, data describing a pitch associated with the voice data, other data indicative of characteristics of the voice data, or a combination thereof. As another illustrative example, if data stream 604 includes video data, gaming data, or both, feature extractor 606 may be configured to determine pixel information associated with image frames of data stream 604. In the same or other examples, data sample 120 may include other information, such as metadata associated with data stream 604, compressed data (e.g., key frame identifiers), or other information used by subsystem 610 to encode data sample 120.

[0067] The subsystem 610 includes an encoder portion 180 of a bundled multi-rate feedback autoencoder. The encoder portion 180 may correspond to the encoder portion 180A of the system 100, the encoder portion 180B of the system 200, or both. In some implementations, the encoder portion 180 may include a bottleneck layer 107. In other implementations, the bottleneck layer 107 may be coupled to the encoder portion 180. Similar to the description with respect to FIGS. 1 and 2, the encoder portion 180 may generate input data states 122A-122E (not shown in FIG. 6) that are provided to the bottleneck layer 107. The bottleneck layer 107 may generate output data states 124A-124E based on the input data states 122A-122E. As described above, the bottleneck layer 107 may include multiple bottlenecks 108 associated with different bit rates to facilitate encoding frames having different bit sizes at similar frequencies. For ease of illustration, Figure 6 shows output data states 124A and 124B. However, it should be understood that the bottleneck layer 107 shown in Figure 6 can generate other output data states 124C-124E, as described above. The output data states 124A, 124B are provided to the frame generator 402. The subsystem 610 may also include a decoder portion 190 of a bundled multi-rate feedback autoencoder. The decoder portion 190 may operate in a manner substantially similar to the system 500 of Figure 5.

[0068] Frame generator 402 is configured to generate frame 420A (e.g., a reference frame) and frame 420B (e.g., a predicted frame). It should be understood that frame generator 402 may generate additional frames (e.g., frames 420C-420E) as described with respect to FIG. 4. Frames 420A, 420B are provided to packet generator 404. Packet generator 404 is configured to generate packets 430 based on frames 420A, 420B (and other frames 420C-420E not shown in FIG. 6).

[0069] The modem 628 is configured to modulate baseband according to a particular communication protocol to generate signals representing the packet 430 and the previous packet 634. The transmitter 630 is configured to send the signals representing the packets 430, 634 over a transmission medium 632. The transmission medium 632 may include a wireline medium, an optical medium, or a wireless medium. By way of example, the transmitter 630 may include or correspond to a wireless transmitter configured to send signals via free-space propagation of electromagnetic waves.

[0070] According to one implementation, the bit rate of the bottleneck 108 in the bottleneck layer 107 may be dynamically changed based on network conditions associated with the transmission medium 632. As a non-limiting example, if the network is congested such that packets are lost or delayed more frequently, the bit rate may be increased to allocate additional bits to the frames. As another example, if the network has a relatively large bandwidth such that packets are rarely lost or delayed, the bit rate may be decreased.

[0071] 6, the receiving device 652 is configured to receive packets 430, 634 from the transmitting device 602. As discussed above, the transmission medium 632 may be lossy. For example, one or more of the packets 430, 634 may be delayed in transmission or may not be received at the receiving device 652. The receiving device 652 includes multiple components configured to process the received packets 430, 634 and generate an output based on the received packets 430, 634.

[0072] 6, components of the receiving device 652 include a receiver 654, a modem 656, a depacketizer 658, one or more buffers 660, a decoder controller 665, one or more decoder networks 670, a renderer 678, and a user interface device 680. In other examples, the receiving device 652 may include more, fewer, or different components. By way of example, in some embodiments, the receiving device 652 includes two or more user interface devices 680, such as one or more displays, one or more speakers, one or more haptic output devices, etc. By way of further example, in some embodiments, the receiving device 652 includes a transceiver in place of (or in which) the receiver 654 is located.

[0073] The receiver 654 is configured to receive signals representing the packets 430, 634 and provide the signals (after initial signal processing, such as amplification, filtering, etc.) to the modem 656. As mentioned above, the receiving device 652 may not receive all of the packets 430, 634 sent by the transmitting device 602. Additionally or alternatively, the packets 430, 634 may be received in an order different than the order in which they were transmitted by the transmitting device 602.

[0074] The modem 656 is configured to demodulate the signal to generate bits representative of the received packets 430, 634 and provide the bits representative of the received data packets to a depacketizer 658. The depacketizer 658 is configured to extract one or more data frames 420 from the payload of each received packet 430, 634 and store the frames 420 in a buffer(s) 660. For example, in FIG. 6, the buffer(s) 660 includes a jitter buffer(s) 662 configured to store the data frames 420. The buffer(s) 660 store the data frames 420 to allow reordering of the data frames 420 to allow time for delayed data frames to arrive.

[0075] 6, the decoder controller 665 retrieves data from the buffer(s) 660 to generate the output data state 124 for the decoder network(s) 670. In some implementations, the decoder controller 665 also performs buffer management operations, such as managing the depth of the jitter buffer(s) 662, the depth of the playout buffer(s) 674, or both. If the decoder network(s) 670 includes multiple decoders, the decoder controller 665 may also determine which decoder to use at a particular time.

[0076] To decode a particular data sample, the decoder controller 665 extracts the output data state 124 from the frame 420 and provides the output data state 124 to a decoder 672 of the decoder network 670. The decoder 672 may include components of the system 500 and may operate in a substantially similar manner. For example, the decoder 672 may generate a reconstructed data sample 126 based on the output state 124 in a manner similar to that described with respect to FIG.

[0077] The reconstructed data samples 126 may be stored in a buffer(s) 660 (e.g., in one or more playout buffers 674). During playback, a renderer 678 retrieves the reconstructed data samples 126 from the buffer(s) 660 and processes the reconstructed data samples 126 to generate output signals, such as audio signals, video signals, game update signals, etc. The renderer 678 provides signals to a user interface device 680 to generate user-perceptible outputs based on the reconstructed data samples 126. For example, the user-perceptible outputs may include one or more of a sound, an image, or a vibration. In some implementations, the renderer 678 includes or corresponds to a game engine that generates user-perceptible outputs in response to modifying a game state based on the reconstructed data samples 126.

[0078] Fig. 7 is a flow diagram of a particular example of an encoding device operation method 700. In various implementations, the method 700 may be performed by one or more of the system 100 of Fig. 1, the system 200 of Fig. 2, the bottlenecks 108A, 108B of Fig. 3, the system 400 of Fig. 4, or the transmitting device 602 of Fig. 6.

[0079] In the example of FIG. 7, the method 700 includes generating a first input data state for a data sample in a time series of data samples of a portion of an audio data stream at block 702. For example, referring to FIG. 1, the bidirectional GRU layer 105 may generate an input data state 122B and an input data state 122A. In this example, the input data states 122A, 122B of each data sample 120A, 120B respectively correspond to encoder hidden states generated in the bidirectional GRU layer 105 of the bundled multi-rate feedback autoencoder. As another example, referring to FIG. 2, the encoder-side attention mechanism 205 may generate an input data state 122B and an input data state 122A.

[0080] The method 700 also includes, at block 704, providing a first input data state to a first bottleneck and providing a second input data state to a second bottleneck, the second input data state being different from the first input data state. The first bottleneck is associated with a first bitrate and the second bottleneck is associated with a second bitrate. According to some implementations, the second bitrate is different from the first bitrate. For example, with reference to FIGS. 1 and 2, the input data state 122B is provided to the bottleneck 108B and the input data state 122A is provided to the bottleneck 108A. The bottleneck 108B is associated with a first bitrate and the bottleneck 108A is associated with a second bitrate, the second bitrate being different from the first bitrate. For example, the first bitrate can be less than the second bitrate such that the bottleneck 108A can encode more data bits than the second bottleneck during a similar period of time. The bottlenecks 108A, 108B are integrated into the bottleneck layer 107 of a bundled multi-rate feedback autoencoder.

[0081] According to one implementation, the method 700 includes allocating fewer units to a latent code generated at a first bottleneck than a latent code generated at a second bottleneck to reduce the bitrate of the first bottleneck. For example, referring to FIG. 3, fewer units may be allocated to a latent code generated at bottleneck 108B compared to the number of units allocated to the latent code generated at bottleneck 108A to reduce the bitrate of bottleneck 108B compared to the bitrate of bottleneck 108A. For example, quantized latent 352B may have fewer units than quantized latent 352A.

[0082] According to one implementation of the method 700, the first codebook associated with the first bottleneck has a smaller size than the second codebook associated with the second bottleneck. For example, referring to FIG. 3, in order to reduce the bit rate of the bottleneck 108B compared to the bit rate of the bottleneck 108A, the codebook 306B associated with the bottleneck 108B may have a smaller size than the codebook 306A associated with the bottleneck 108A.

[0083] According to one implementation, the method 700 includes dynamically varying the first bit rate and the second bit rate based on network conditions. For example, if the network is congested, resulting in more frequent lost or delayed packets, the second bit rate may be increased to allocate additional bits to the reference frame 402A. As another example, if the network has a relatively large bandwidth, resulting in fewer lost or delayed packets, the second bit rate may be decreased.

[0084] The method 700 also includes, at block 706, generating a first encoded frame based on a first output data state from the first bottleneck and generating a second encoded frame based on a second output data state from the second bottleneck. The first encoded frame and the second encoded frame are bundled into a packet. For example, referring to FIG. 3, the bottleneck 108B generates the output data state 124B and the bottleneck 108A generates the output data state 124A. Referring to FIG. 4, the frame generator 402 generates the encoded frame 420B based on the output data state 124B and generates the encoded frame 420A based on the output data state 124A. The encoded frames 420A, 420B are bundled into a packet 430.

[0085] The method 700 of FIG. 7 allows for customized bit allocation for encoding different data samples 120. For example, a greater number of bits may be assigned to a reference frame data sample (e.g., data sample 120A) than to a predicted frame data sample (e.g., data samples 120B-120E). Thus, if data samples 120A-120E are encoded into five respective frames (e.g., frames 420A-420E of FIG. 4) bundled into a packet (e.g., packet 430 of FIG. 4), a boundary frame of the packet 430 (e.g., frame 420A associated with data sample 120A) may be encoded with additional bits to be used as a reference frame. As a result, the remaining frames 420B-420E may be encoded with a relatively small number of bits, which reduces the amount of bits used to encode the frames bundled into the packet 430. In some scenarios, the system 100 may allocate zero bits to a particular data sample 120 for encoding. In these scenarios, nothing is transmitted for a particular data sample 120, and the decoder will predict the corresponding frame based on the coded data in adjacent frames.

[0086] The method 700 of Figure 7 may be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a digital signal processor (DSP), a graphics processing unit (GPU), a controller, another hardware device, a firmware device, or any combination thereof. As an example, the method 700 of Figure 7 may be performed by a processor executing instructions, such as those described with reference to the processor(s) 1110 of Figure 11.

[0087] 8 is a flow diagram of a particular example of a method of operation 800 of a decoding device. In various implementations, the method 800 may be performed by one or more of the system 100 of FIG. 1, the system 200 of FIG. 2, the system 500 of FIG. 5, the transmitting device 602 of FIG. 6, or the receiving device 652 of FIG. 6.

[0088] In the example of FIG. 8, the method 800 includes receiving a packet at a decoder network, at block 802, including a first encoded frame bundled with a second encoded frame. The first encoded frame includes a first output data state generated from a first bottleneck of a feedback autoencoder, and the second encoded frame includes a second output data state generated from a second bottleneck of the feedback autoencoder. The first bottleneck is associated with a first bit rate, and the second bottleneck is associated with a second bit rate. According to some implementations, the second bit rate is different from the first bit rate. For example, the receiving device 652 receives a packet 430 including an encoded frame 420B bundled with an encoded frame 420A. The encoded frame 420B includes an output data state 124B generated from a bottleneck 108B, and the encoded frame 420A includes an output data state 124A generated from a bottleneck 108A. Bottleneck 108B is associated with a first bitrate and bottleneck 108A is associated with a second bitrate that is different from (eg, higher than) the first bitrate.

[0089] The method 800 also includes generating a reconstructed first data sample based on the first output data state at block 804. The reconstructed data sample corresponds to a first data sample in the time series of data samples of the portion of the audio data stream. For example, referring to FIG. 5, the GRU layer 109 and the back-end neural network post-processing layer 112 may generate a reconstructed data sample 126B based on the output data state 124B. The reconstructed data sample 126B corresponds to data sample 120B in the time series of data samples 120 of the audio data stream 604.

[0090] The method 800 also includes generating a reconstructed second data sample based on the second output data state at block 806. The reconstructed second data sample corresponds to a second data sample in the time series of data samples. For example, referring to FIG. 5, the GRU layer 109 and the back-end neural network post-processing layer 112 may generate a reconstructed data sample 126A based on the output data state 124A. The reconstructed data sample 126A corresponds to the data sample 120A in the time series of data samples 120 of the audio data stream 604.

[0091] The method 800 of Figure 8 may be implemented by an FPGA device, an ASIC, a processing unit such as a CPU, DSP, GPU, a controller, another hardware device, a firmware device, or any combination thereof. As an example, the method 800 of Figure 8 may be performed by a processor executing instructions, such as described with reference to the processor(s) 1110 of Figure 11.

[0092] 9 shows an implementation 900 in which device 902 includes one or more processors 910 that include components of transmitting device 602 of FIG. 6. Device 902 also includes an input interface 904 (e.g., one or more buses or wireless interfaces) configured to receive input data, such as data stream 604, and an output interface 906 (e.g., one or more buses or wireless interfaces) configured to output data 914, such as packets 430. Device 902 may correspond to a system-on-chip or other modular device that may be integrated into other systems to provide data encoding, such as in a mobile phone, another communication device, an entertainment system, or a vehicle, as an illustrative and non-limiting example. According to some implementations, the device 902 may be integrated into a server, a mobile communications device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an automobile such as a car, or any combination thereof.

[0093] In the illustrated implementation 900, the device 902 includes a memory 920 (e.g., one or more memory devices) that includes instructions 922 and one or more codebooks 306. The device 902 also includes one or more processors 910 coupled to the memory 920 and configured to execute the instructions 922 from the memory 920. In this implementation 900, the feature extractor 606, the subsystem 610, the encoder portion 180 of the bundled multi-rate feedback autoencoder, the frame generator 402, and the packet generator 404 may correspond to or be implemented via the instructions 922. For example, when the instructions 922 are executed by the processor(s) 910, the processor(s) 910 may generate input data states 122A-122E for each data sample 120A-122E in the time series of data samples 120 of the portion of the audio data stream. The processor(s) 910 may also provide at least one input data state 122B to the first bottleneck 108B and at least one other input data state 122A to the second bottleneck 108A. The first bottleneck 108B may be associated with a first bitrate and the second bottleneck 108A may be associated with a second bitrate different from the first bitrate. The processor(s) 910 may also generate a first encoded frame 420B based on a first output data state 124B from the first bottleneck 108B and generate a second encoded frame 420A based on a second output data state 124A from the second bottleneck 108A. The first encoded frame 420B and the second encoded frame 420A are bundled into a packet 430.

[0094] Figure 10 shows an implementation 1000 in which a device 1002 includes one or more processors 1010 that include components of the receiving device 652 of Figure 6. The device 1002 also includes an input interface 1004 (e.g., one or more buses or wireless interfaces) configured to receive input data 1012, such as packets 430 from the receiver 654 of Figure 6, and an output interface 1006 (e.g., one or more buses or wireless interfaces) configured to provide output 1014 based on the input data 1012, such as signals provided to the user interface device 680 of Figure 6. The device 1002 may correspond to a system-on-chip or other modular device that may be integrated into other systems to provide data decoding, such as in a mobile phone, another communication device, an entertainment system, or a vehicle, as illustrative and not limiting examples. According to some implementations, the device 1002 may be integrated into a server, a mobile communications device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a DVD player, a tuner, a camera, a navigation device, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an automobile such as a car, or any combination thereof.

[0095] In the illustrated implementation 1000, the device 1002 includes a memory 1020 (e.g., one or more memory devices) that includes instructions 1022 and one or more buffers 660. The device 1002 also includes one or more processors 1010 coupled to the memory 1020 and configured to execute the instructions 1022 from the memory 1020. In this implementation 1000, the depacketizer 658, the decoder controller 665, the decoder network(s) 670, the decoder(s) 672, and / or the renderer 678 may correspond to or be implemented via the instructions 1022. For example, when the instructions 1022 are executed by the processor(s) 1010, the processor(s) 1010 may receive a packet 430 that includes a first encoded frame 420B bundled with a second encoded frame 420A. The first encoded frame 420B may include a first output data state 124B generated from a first bottleneck 108B of the bundled multi-rate feedback autoencoder, and the second encoded frame 420A may include a second output data state 124A generated from a second bottleneck 108A of the bundled multi-rate feedback autoencoder. The first bottleneck may be associated with a first bit rate, and the second bottleneck may be associated with a second bit rate different from the first bit rate. The processor(s) 1010 may further generate a reconstructed first data sample 126B based on the first output data state 124B. The reconstructed first data sample 126B may correspond to a first data sample 120B in the time series of data samples 120 of the portion of the audio data stream 604. The processor(s) 1010 may further generate a reconstructed second data sample 126A based on the second output data state 124A. The reconstructed second data sample 126A may correspond to the second data sample 120A in the time series of data samples 120.

[0096] 11, a block diagram of a particular example implementation of a device is shown, generally designated 1100. In various implementations, the device 1100 may have more or fewer components than shown in FIG. 11. In an example implementation, the device 1100 may correspond to the transmitting device 602 of FIG. 6, the receiving device 652 of FIG. 6, or both. In an example implementation, the device 1100 may perform one or more operations described with reference to FIGS. 1-10.

[0097] In certain implementations, the device 1100 includes a processor 1106 (e.g., a CPU). The device 1100 may include one or more additional processors 1110 (e.g., one or more DSPs, one or more GPUs, or a combination thereof). The processor(s) 1110 may include a speech and music coder-decoder (CODEC) 1108. The speech and music codec 1108 may include a voice coder ("vocoder") encoder 1136, a vocoder decoder 1138, or both. In certain aspects, the vocoder encoder 1136 includes an encoder portion 180 of a bundled multi-rate feedback autoencoder. In certain aspects, the vocoder decoder 1138 includes a decoder portion of a bundled multi-rate feedback autoencoder.

[0098] The device 1100 includes a memory 1186 and a CODEC 1134. The memory 1186 may include instructions 1156 executable by one or more additional processors 1110 (or processor 1106) to implement functions described with reference to the transmitting device 602 of FIG. 6, the receiving device 652 of FIG. 6, or both. The device 1100 may include a modem 1140 coupled to an antenna 1190 via a transceiver 1150.

[0099] The device 1100 may include a display 1128 coupled to a display controller 1126. A speaker 1196 and a microphone 1194 may be coupled to a CODEC 1134. The CODEC 1134 may include a digital-to-analog converter (DAC) 1102 and an analog-to-digital converter (ADC) 1104. In certain implementations, the CODEC 1134 may receive an analog signal from the microphone 1194, convert the analog signal to a digital signal using the analog-to-digital converter 1104, and provide the digital signal to the speech and music CODEC 1108 (e.g., as data stream 604 of FIG. 6). The speech and music CODEC 1108 may process the digital signal. In certain implementations, the speech and music CODEC 1108 may provide a digital signal (e.g., an output from a renderer 678 of FIG. 6) to the CODEC 1134. The CODEC 1134 may convert the digital signal to an analog signal using a digital-to-analog converter 1102 and may provide the analog signal to a speaker 1196 .

[0100] In particular implementations, the device 1100 can be included in a system-in-package or system-on-chip device 1122 that corresponds to the transmitting device 602 of Figure 6, the system 100 of Figure 1, the system 200 of Figure 2, the system 400 of Figure 4, the device 902 of Figure 9, or any combination thereof. Additionally or alternatively, the system-in-package or system-on-chip device 1122 corresponds to the receiving device 652 of Figure 1, the system 100 of Figure 1, the system 200 of Figure 2, the system 500 of Figure 5, the device 1002 of Figure 10, or any combination thereof.

[0101] In certain implementations, the memory 1186, the processor 1106, the processors 1110, the display controller 1126, the CODEC 1134, and the modem 1140 are included in the system-in-package or system-on-chip device 1122. In certain implementations, the input device 1130 and the power supply 1144 are coupled to the system-in-package or system-on-chip device 1122. Furthermore, in certain implementations, the display 1128, the input device 1130, the speaker 1196, the microphone 1194, the antenna 1190, and the power supply 1144 are external to the system-in-package or system-on-chip device 1122, as shown in FIG. 11. In certain implementations, each of the display 1128, the input device 1130, the speaker 1196, the microphone 1194, the antenna 1190, and the power supply 1144 may be coupled to a component of the system-in-package or system-on-chip device 1122, such as an interface or a controller. In some implementations, the device 1100 includes additional memory that is external to the system-in-package or system-on-chip device 1122 and coupled to the system-in-package or system-on-chip device 1122 via an interface or controller.

[0102] The device 1100 may include a smart speaker (e.g., the processor 1106 may execute instructions 1156 to run a voice-controlled digital assistant application), a speaker bar, a mobile communications device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a DVD player, a tuner, a camera, a navigation device, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, a vehicle, or any combination thereof.

[0103] In accordance with the described implementation, the apparatus includes means for generating a first input data state for a data sample in a time series of data samples of a portion of the audio data stream. For example, the means for generating the input data state includes the front-end neural network preprocessing layer 102, the bidirectional GRU layer 105, the encoder-side attention mechanism 205, the encoder portion 180 of the bundled multi-rate feedback autoencoder, the transmitting device 602, the device 902, the processor(s) 910, the processor 1106, the processor(s) 1110, the speech and music codec 1108, the vocoder decoder 1138, one or more other circuits or components configured to generate the input data state, or any combination thereof.

[0104] The apparatus also includes means for providing a first input data state to the first bottleneck and a second input data state to the second bottleneck, the second input data state being different from the first input data state. The first bottleneck is associated with a first bitrate, and the second bottleneck is associated with a second bitrate being different from the first bitrate. For example, the means for providing includes the front-end neural network preprocessing layer 102, the bidirectional GRU layer 105, the encoder-side attention mechanism 205, the encoder portion 180 of the bundled multi-rate feedback autoencoder, the transmitting device 602, the device 902, the processor(s) 910, the processor 1106, the processor(s) 1110, the speech and music codec 1108, the vocoder decoder 1138, one or more other circuits or components configured to provide the input data state to the bottlenecks, or any combination thereof.

[0105] The apparatus further includes means for generating a first encoded frame based on a first output data condition from the first bottleneck and for generating a second encoded frame based on a second output data condition from the second bottleneck. The first encoded frame and the second encoded frame are bundled into a packet. For example, the generating means includes the frame generator 402, the transmitting device 602, the device 902, the processor(s) 910, the processor 1106, the processor(s) 1110, the speech and music codec 1108, the vocoder decoder 1138, one or more other circuits or components configured to generate the first and second encoded frames, or any combination thereof.

[0106] In accordance with the described implementation, the apparatus includes means for receiving a packet including a first encoded frame bundled with a second encoded frame. The first encoded frame includes a first output data state generated from a first bottleneck of a feedback autoencoder, and the second encoded frame includes a second output data state generated from a second bottleneck of the feedback autoencoder. The first bottleneck is associated with a first bit rate, and the second bottleneck is associated with a second bit rate that is different from the first bit rate. For example, the means for receiving the packet includes the receiver 654, the modem 656, the depacketizer 658, the input interface 1004, the processor(s) 1010, the antenna 1190, the transceiver 1150, the modem 1140, the processor(s) 1110, the processor 1106, one or more other circuits or components configured to receive the packet, or any combination thereof.

[0107] The apparatus also includes a means for generating a reconstructed first data sample based on the first output data state. The reconstructed first data sample corresponds to a first data sample in a time sequence of data samples of a portion of the audio data stream. For example, the means for generating the reconstructed first data sample includes the bidirectional GRU layer 109, the back-end neural network post-processing layer 112, the decoder controller 665, the decoder network(s) 670, the decoder 672, the processor(s) 1010, the processor(s) 1110, the processor 1106, one or more other circuits or components configured to generate the reconstructed first data sample, or any combination thereof.

[0108] The apparatus further includes a means for generating a reconstructed second data sample based on the second output data state. The reconstructed second data sample corresponds to a second data sample in the time series of data samples. For example, the means for generating the reconstructed second data sample includes the bidirectional GRU layer 109, the back-end neural network post-processing layer 112, the decoder controller 665, the decoder network(s) 670, the decoder 672, the processor(s) 1010, the processor(s) 1110, the processor 1106, one or more other circuits or components configured to generate the reconstructed second data sample, or any combination thereof.

[0109] In some implementations, the non-transitory computer-readable medium includes instructions that, when executed by one or more processors of the device, cause the one or more processors to generate an input data state (e.g., input data state 122) for each data sample (e.g., data sample 120) in a time series of data samples of a portion of an audio data stream (e.g., data stream 604). Execution of the instructions also causes the one or more processors to provide at least one input data state to a first bottleneck (e.g., bottleneck 108B) and provide at least one other input data state to a second bottleneck (e.g., bottleneck 108A). The first bottleneck is associated with a first bitrate and the second bottleneck is associated with a second bitrate that is different from the first bitrate. Execution of the instructions further causes the one or more processors to generate a first encoded frame (e.g., frame 420B) based on a first output data state (e.g., output data state 124B) from the first bottleneck and generate a second encoded frame (e.g., frame 420A) based on a second output data state (e.g., output data state 124A) from the second bottleneck. The first encoded frame and the second encoded frame are bundled into a packet (e.g., packet 430).

[0110] In some implementations, the non-transitory computer-readable medium includes instructions that, when executed by one or more processors of a device, cause the one or more processors to receive at a decoder network a packet (e.g., packet 430) including a first encoded frame (e.g., frame 420B) bundled with a second encoded frame (e.g., frame 420A). The first encoded frame includes a first output data state (e.g., output data state 124B) generated from a first bottleneck (e.g., bottleneck 108B) of the feedback autoencoder, and the second encoded frame includes a second output data state (e.g., output data state 124A) generated from a second bottleneck (e.g., bottleneck 108A) of the feedback autoencoder. The first bottleneck is associated with a first bitrate, and the second bottleneck is associated with a second bitrate that is different from the first bitrate. Execution of the instructions also causes the one or more processors to generate a reconstructed first data sample (e.g., reconstructed data sample 126B) based on the first output data state. The reconstructed first data sample corresponds to a first data sample (e.g., data sample 120B) in a time series of data samples (e.g., data sample 120) of a portion of the audio data stream (e.g., data stream 604). Execution of the instructions further causes the one or more processors to generate a reconstructed second data sample (e.g., reconstructed data sample 126A) based on the second output data state. The reconstructed second data sample corresponds to a second data sample (e.g., data sample 120A) in the time series of data samples.

[0111] Certain aspects of the present disclosure are described below in a set of interrelated examples.

[0112] According to Example 1, a device includes a memory and one or more processors coupled to the memory, where the one or more processors are configured to operate to generate a first input data state for data samples in a time series of data samples of a portion of an audio data stream, provide the first input data state to a first bottleneck associated with a first bit rate, provide a second input data state, different from the first input data state, to a second bottleneck associated with a second bit rate, generate a first encoded frame based on a first output data state from the first bottleneck, and generate a second encoded frame based on a second output data state from the second bottleneck, where the first encoded frame and the second encoded frame are bundled into a packet.

[0113] Example 2 includes the device of example 1, wherein the first bottleneck and the second bottleneck are integrated into a bottleneck layer of the feedback autoencoder.

[0114] Example 3 includes the device of example 1 or 2, wherein the first input data state and the second input data state correspond to first encoder hidden states and second encoder hidden states generated in a bidirectional gated recurrent unit (GRU) layer of a feedback autoencoder.

[0115] Example 4 includes the device of any of Examples 1 to 3, where the first bit rate is different from the second bit rate.

[0116] Example 5 includes a device described in any of Examples 1 to 4, wherein the one or more processors are configured to be operable to allocate fewer bits to the latent code generated at the first bottleneck than the latent code generated at the second bottleneck.

[0117] Example 6 includes the device of any of Examples 1 to 5, wherein a first codebook associated with the first bottleneck has a smaller size than a second codebook associated with the second bottleneck.

[0118] Example 7 includes a device described in any of Examples 1 to 6, wherein the packet includes a predicted frame and a reference frame, and an input data state associated with the predicted frame is provided to a first bottleneck, and an input data state associated with the reference frame is provided to a second bottleneck.

[0119] Example 8 includes the device according to any one of examples 1 to 7, wherein a bit size of the prediction frame is smaller than a bit size of the reference frame.

[0120] Example 9 includes the device of any of Examples 1-8, wherein the input data state for each frame of a packet is generated using an attention mechanism.

[0121] Example 10 includes the device of any of Examples 1-9, wherein the attention mechanism includes a transformer.

[0122] Example 11 includes the device of any of Examples 1 to 10, wherein the one or more processors are configured to be operable to dynamically change the first bit rate and the second bit rate based on network conditions.

[0123] Example 12 includes a method including generating a first input data state for data samples in a time series of data samples of a portion of an audio data stream; providing the first input data state to a first bottleneck associated with a first bit rate and providing a second input data state, different from the first input data state, to a second bottleneck associated with a second bit rate; generating a first encoded frame based on a first output data state from the first bottleneck and generating a second encoded frame based on a second output data state from the second bottleneck, wherein the first encoded frame and the second encoded frame are bundled into a packet.

[0124] Example 13 includes the method of example 12, in which the first bottleneck and the second bottleneck are integrated into a bottleneck layer of the feedback autoencoder.

[0125] Example 14 includes the method of example 12 or 13, wherein the first input data state and the second input data state correspond to a first encoder hidden state and a second encoder hidden state generated in a bidirectional gated recurrent unit (GRU) layer of a feedback autoencoder.

[0126] Example 15 includes the method of any of examples 12-14, wherein the first bit rate is different from the second bit rate.

[0127] Example 16 includes the method of any of examples 12-15, further including allocating fewer bits to the latent code generated at the first bottleneck than the latent code generated at the second bottleneck.

[0128] Example 17 includes the method of any of examples 12-16, wherein the first codebook associated with the first bottleneck has a smaller size than the second codebook associated with the second bottleneck.

[0129] Example 18 includes a method described in any of Examples 12 to 17, wherein the packet includes a predicted frame and a reference frame, and an input data state associated with the predicted frame is provided to a first bottleneck, and an input data state associated with the reference frame is provided to a second bottleneck.

[0130] Example 19 includes the method according to any one of Examples 12 to 18, wherein a bit size of the prediction frame is smaller than a bit size of the reference frame.

[0131] Example 20 includes the method according to any of Examples 12 to 19, wherein the input data state of each frame of a packet is generated using an attention mechanism.

[0132] Example 21 includes the method of any of examples 12-20, wherein the attention mechanism includes a transformer.

[0133] Example 22 includes the method of any of examples 12-21, further including dynamically varying the first bit rate and the second bit rate based on network conditions.

[0134] Example 23 includes a non-transitory computer-readable medium storing instructions, the instructions executable by one or more processors to generate a first input data state for data samples in a time series of data samples of a portion of an audio data stream, provide the first input data state to a first bottleneck associated with a first bit rate, provide a second input data state, different from the first input data state, to a second bottleneck associated with a second bit rate, generate a first encoded frame based on a first output data state from the first bottleneck, and generate a second encoded frame based on a second output data state from the second bottleneck, wherein the first encoded frame and the second encoded frame are bundled into a packet.

[0135] Example 24 includes the non-transitory computer-readable medium of example 23, wherein the first bottleneck and the second bottleneck are integrated into a bottleneck layer of the feedback autoencoder.

[0136] Example 25 includes the non-transitory computer-readable medium of example 23 or 24, wherein the first input data state and the second input data state correspond to a first encoder hidden state and a second encoder hidden state generated in a bidirectional gated recurrent unit (GRU) layer of a feedback autoencoder.

[0137] Example 26 includes the non-transitory computer-readable medium of any of Examples 23-25, wherein the first bit rate is different from the second bit rate.

[0138] Example 27 includes the non-transitory computer-readable medium of any of Examples 23-26, wherein the instructions, when executed, further cause the one or more processors to allocate fewer bits to the latent code generated at the first bottleneck than the latent code generated at the second bottleneck.

[0139] Example 28 includes the non-transitory computer-readable medium of any of Examples 23-27, wherein a first codebook associated with a first bottleneck has a smaller size than a second codebook associated with a second bottleneck.

[0140] Example 29 includes the non-transitory computer-readable medium of any of Examples 23 to 28, wherein the packet includes a predicted frame and a reference frame, and an input data state associated with the predicted frame is provided to a first bottleneck, and an input data state associated with the reference frame is provided to a second bottleneck.

[0141] Example 30 includes the non-transitory computer-readable medium of any of Examples 23-29, wherein a bit size of the predicted frame is smaller than a bit size of the reference frame.

[0142] Example 31 includes the non-transitory computer-readable medium of any of Examples 23-30, wherein the input data state for each frame of a packet is generated using an attention mechanism.

[0143] Example 32 includes the non-transitory computer-readable medium of any of examples 23-31, wherein the attention mechanism includes a transformer.

[0144] Example 33 includes the non-transitory computer-readable medium of any of Examples 23-32, wherein the instructions, when executed, cause one or more processors to dynamically change the first bit rate and the second bit rate based on network conditions.

[0145] Example 34 includes an apparatus including means for generating a first input data state for data samples in a time series of data samples of a portion of an audio data stream; means for providing the first input data state to a first bottleneck associated with a first bit rate and for providing a second input data state, different from the first input data state, to a second bottleneck associated with a second bit rate; and means for generating a first encoded frame based on a first output data state from the first bottleneck and generating a second encoded frame based on a second output data state from the second bottleneck, wherein the first encoded frame and the second encoded frame are bundled into a packet.

[0146] Example 35 includes the apparatus of example 34, wherein the first bottleneck and the second bottleneck are integrated into a bottleneck layer of the feedback autoencoder.

[0147] Example 36 includes the apparatus of example 34 or 35, wherein the first input data state and the second input data state correspond to a first encoder hidden state and a second encoder hidden state generated in a bidirectional gated recurrent unit (GRU) layer of a feedback autoencoder.

[0148] Example 37 includes the apparatus of any of Examples 34 to 36, wherein the first bit rate is different from the second bit rate.

[0149] Example 38 includes the apparatus of any of examples 34-37, wherein fewer bits are assigned to the latent code generated at the first bottleneck than the latent code generated at the second bottleneck.

[0150] Example 39 includes the apparatus of any of examples 34-38, wherein a first codebook associated with the first bottleneck has a smaller size than a second codebook associated with the second bottleneck.

[0151] Example 40 includes the apparatus of any of Examples 34 to 39, wherein the packet includes a predicted frame and a reference frame, and an input data state associated with the predicted frame is provided to a first bottleneck, and an input data state associated with the reference frame is provided to a second bottleneck.

[0152] Example 41 includes the apparatus according to any one of examples 34 to 40, wherein a bit size of the prediction frame is smaller than a bit size of the reference frame.

[0153] Example 42 includes the apparatus of any of Examples 34-41, wherein the input data state for each frame of a packet is generated using an attention mechanism.

[0154] Example 43 includes the device of any of Examples 34 to 42, wherein the attention mechanism includes a transformer.

[0155] Example 44 includes the apparatus of any of Examples 34 to 43, further including means for dynamically varying the first bit rate and the second bit rate based on network conditions.

[0156] Example 45 includes a device including a memory and one or more processors coupled to the memory, wherein the one or more processors are configured to execute instructions from the memory to receive a packet at a decoder network including a first encoded frame bundled with a second encoded frame, the first encoded frame including a first output data state generated from a first bottleneck of a feedback autoencoder associated with a first bit rate, and the second encoded frame including a second output data state generated from a second bottleneck of the feedback autoencoder associated with a second bit rate, generate a reconstructed first data sample corresponding to a first data sample in a time series of data samples of a portion of the audio data stream based on the first output data state, and generate a second reconstructed data sample corresponding to a second data sample in the time series of data samples based on the second output data state.

[0157] Example 46 includes the device of example 45, wherein the first output data state is different from the second output data state.

[0158] Example 47 includes the device of example 45 or 46, wherein the first output data state and the second output data state are received at a bidirectional gated recurrent unit (GRU) layer of the decoder network.

[0159] Example 48 includes the device of any of Examples 45 to 47, wherein the first bit rate is smaller than the second bit rate.

[0160] Example 49 includes the device of any of examples 45-48, wherein the first output data state and the second output data state are received by an attention mechanism.

[0161] Example 50 includes the device of any of Examples 45 to 49, wherein the attention mechanism includes a transformer.

[0162] Example 51 includes a method including: receiving a packet at a decoder network, the packet including a first encoded frame bundled with a second encoded frame, the first encoded frame including a first output data state generated from a first bottleneck of a feedback autoencoder associated with a first bit rate, and the second encoded frame including a second output data state generated from a second bottleneck of the feedback autoencoder associated with a second bit rate; generating a reconstructed first data sample corresponding to a first data sample in a time series of data samples of a portion of an audio data stream based on the first output data state; and generating a reconstructed second data sample corresponding to a second data sample in the time series of data samples based on the second output data state.

[0163] Example 52 includes the method of example 51, wherein the first output data state is different from the second output data state.

[0164] Example 53 includes the method of example 51 or 52, wherein the first output data state and the second output data state are received at a bidirectional gated recurrent unit (GRU) layer of the decoder network.

[0165] Example 54 includes the method of any of examples 51 to 53, wherein the first bit rate is smaller than the second bit rate.

[0166] Example 55 includes the method of any of examples 51-54, wherein the first output data state and the second output data state are received by an attention mechanism.

[0167] Example 56 includes the method of any of examples 51 to 55, wherein the attention mechanism includes a transformer.

[0168] Example 57 includes a non-transitory computer-readable medium including instructions executable by one or more processors to receive a packet at a decoder network including a first encoded frame bundled with a second encoded frame, the first encoded frame including a first output data state generated from a first bottleneck of a feedback autoencoder associated with a first bit rate, and the second encoded frame including a second output data state generated from a second bottleneck of the feedback autoencoder associated with a second bit rate, generate a reconstructed first data sample corresponding to a first data sample in a time series of data samples of a portion of the audio data stream based on the first output data state, and generate a second reconstructed data sample corresponding to a second data sample in the time series of data samples based on the second output data state.

[0169] Example 58 includes the non-transitory computer-readable medium of example 56, wherein the first output data state is different from the second output data state.

[0170] Example 59 includes the non-transitory computer-readable medium of example 56 or 57, wherein the first output data state and the second output data state are received at a bidirectional gated recursive unit (GRU) layer of the decoder network.

[0171] Example 60 includes the non-transitory computer-readable medium of any of Examples 57-59, wherein the first bit rate is less than the second bit rate.

[0172] Example 61 includes the non-transitory computer-readable medium of any of Examples 57-60, wherein the first output data state and the second output data state are received by an attention mechanism.

[0173] Example 62 includes the non-transitory computer-readable medium of any of Examples 57-61, wherein the attention mechanism includes a transformer.

[0174] Example 63 includes an apparatus including means for receiving a packet at a decoder network, the packet including a first encoded frame bundled with a second encoded frame, the first encoded frame including a first output data state generated from a first bottleneck of a feedback autoencoder associated with a first bit rate, and the second encoded frame including a second output data state generated from a second bottleneck of the feedback autoencoder associated with a second bit rate; means for generating a reconstructed first data sample corresponding to a first data sample in a time series of data samples of a portion of an audio data stream based on the first output data state; and means for generating a reconstructed second data sample corresponding to a second data sample in the time series of data samples based on the second output data state.

[0175] Example 64 includes the apparatus of example 63, wherein the first output data state is different from the second output data state.

[0176] Example 65 includes the apparatus of example 63 or 64, wherein the first output data state and the second output data state are received at a bidirectional gated recurrent unit (GRU) layer of the decoder network.

[0177] Example 66 includes the apparatus of any of Examples 63 to 65, wherein the first bit rate is smaller than the second bit rate.

[0178] Example 67 includes the apparatus of any of Examples 63-66, wherein the first output data state and the second output data state are received by an attention mechanism.

[0179] Example 68 includes the apparatus of any of examples 63 to 67, wherein the attention mechanism includes a transformer.

[0180] Those skilled in the art will further appreciate that the various exemplary logical blocks, configurations, modules, circuits, and algorithmic steps described with respect to the implementations disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or a combination of both. Various exemplary components, blocks, configurations, modules, circuits, and steps have been described above generally with respect to their functionality. Whether such functionality is implemented as hardware or as processor-executable instructions depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0181] The steps of a method or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a computing device or user terminal In the alternative, the processor and the storage medium may reside as discrete components in a computing device or user terminal.

[0182] The foregoing description of the disclosed implementations is provided to enable any person skilled in the art to make or use the disclosed implementations. Various modifications to these implementations will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other implementations without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the implementations shown herein, but is to be accorded the widest possible scope consistent with the principles and novel features defined by the following claims.

Claims

1. Memory and one or more processors coupled to the memory; 10. A device comprising: generating a first input data state for a data sample in a time series of data samples of a portion of the audio data stream; providing the first input data state to a first bottleneck and a second input data state different from the first input data state to a second bottleneck, the first bottleneck being associated with a first bitrate and the second bottleneck being associated with a second bitrate; generating a first coded frame based on a first output data condition from the first bottleneck and a second coded frame based on a second output data condition from the second bottleneck, the first coded frame and the second coded frame being bundled into a packet; configured to operate device.

2. The device of claim 1 , wherein the first bottleneck and the second bottleneck are integrated into a bottleneck layer of a feedback autoencoder.

3. 3. The device of claim 2, wherein the first input data state and the second input data state correspond to first and second encoder hidden states generated in a bidirectional gated recurrent unit (GRU) layer of the feedback autoencoder.

4. The device of claim 1 , wherein the first bit rate is different from the second bit rate.

5. 5. The device of claim 4, wherein the one or more processors are configured to operate to allocate fewer bits to latent codes generated at the first bottleneck than to latent codes generated at the second bottleneck.

6. The device of claim 4 , wherein a first codebook associated with the first bottleneck has a smaller size than a second codebook associated with the second bottleneck.

7. 5. The device of claim 4, wherein the packet includes a predicted frame and a reference frame, a first input data state associated with the predicted frame being provided to the first bottleneck, and a second input data state associated with the reference frame being provided to the second bottleneck.

8. The device of claim 7 , wherein a bit size of the predicted frame is smaller than a bit size of the reference frame.

9. The device of claim 1 , wherein the input data state for each frame of the packet is generated using an attention mechanism.

10. The device of claim 9 , wherein the attention mechanism includes a transformer.

11. The device of claim 1 , wherein the one or more processors are configured to be operable to dynamically change the first bit rate and the second bit rate based on network conditions.

12. A device-implemented method comprising: generating a first input data state for a data sample in a time series of data samples of a portion of the audio data stream; providing the first input data state to a first bottleneck and a second input data state different from the first input data state to a second bottleneck, the first bottleneck being associated with a first bitrate and the second bottleneck being associated with a second bitrate; generating a first coded frame based on a first output data condition from the first bottleneck and a second coded frame based on a second output data condition from the second bottleneck, the first coded frame and the second coded frame being bundled into a packet; A method comprising:

13. Memory and one or more processors coupled to the memory; 10. A device comprising: receiving, in a decoder network, a packet including a first encoded frame bundled with a second encoded frame, the first encoded frame including a first output data state generated from a first bottleneck of a feedback autoencoder, the second encoded frame including a second output data state generated from a second bottleneck of the feedback autoencoder, the first bottleneck being associated with a first bitrate and the second bottleneck being associated with a second bitrate; generating a reconstructed first data sample based on the first output data state, the reconstructed first data sample corresponding to a first data sample in a time sequence of data samples of a portion of an audio data stream; generating a reconstructed second data sample based on the second output data state, the reconstructed second data sample corresponding to a second data sample in the time series of data samples; configured to operate device.

14. A device-implemented method comprising: receiving, in a decoder network, a packet including a first encoded frame bundled with a second encoded frame, the first encoded frame including a first output data state generated from a first bottleneck of a feedback autoencoder, the second encoded frame including a second output data state generated from a second bottleneck of the feedback autoencoder, the first bottleneck being associated with a first bitrate and the second bottleneck being associated with a second bitrate; generating a reconstructed first data sample based on the first output data state, the reconstructed first data sample corresponding to a first data sample in a time sequence of data samples of a portion of an audio data stream; generating a reconstructed second data sample based on the second output data state, the reconstructed second data sample corresponding to a second data sample in the time series of data samples; A method comprising:

15. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method described in claim 12 or the method described in claim 14.