Efficient signalling of sub-band prediction parameters

IL328785A0Pending Publication Date: 2026-07-01DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
IL · IL
Patent Type
Applications
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2024-12-19
Publication Date
2026-07-01

AI Technical Summary

Technical Problem

Existing audio coding techniques face challenges in efficiently coding audio signals in the complex low delay filterbank (CLDFB) domain, leading to increased delay and computational complexity due to signal domain conversions.

Method used

The proposed solution involves a method for coding audio signals by obtaining a representation of the input audio signal in sub-bands, dividing a current frame into a subset of sub-bands, determining prediction parameters and active/passive flags, and encoding these parameters and flags in a bitstream for efficient transmission and decoding.

Benefits of technology

This approach enables efficient sub-band audio coding, reducing additional delay and computational complexity by avoiding signal domain conversions, while maintaining effective prediction and coding gain, especially for tonal signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000051_0000
    Figure 00000051_0000
  • Figure 00000052_0000
    Figure 00000052_0000
  • Figure 00000053_0000
    Figure 00000053_0000
Patent Text Reader

Abstract

Systems and methods for sub-band audio coding are disclosed. One example provides a method for coding input audio signals in a bitstream from a first device to a second device. The method includes, with the first device, obtaining a representation of the input audio signal in sub- bands, dividing a current frame into a current subset of the sub-bands that excludes other sub-bands, and determining a prediction parameter for the current subset of sub-bands. The method also includes, with the first device, determining an active / passive flag for at least one of the sub-bands and encoding the bitstream with the current subset of the sub-bands, a number of subsets, the prediction parameter for the current subset of sub-bands, and the active / passive flag for the other subsets of sub-bands. The method may include, with the second device, decoding the bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

EFFICIENT SIGNALING OF SUB-BAND PREDICTION PARAMETERSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from U.S. Provisional Application Ser. No. 63 / 613,331, filed on 21 December 2023, U.S. Provisional Application Ser. No. 63 / 561,231, filed on 4 March 2024, and U.S. Provisional Application Ser. No. 63 / 718,182, filed on 8 November 2024, each of which is incorporated by reference herein in its entirety.TECHNICAL FIELD

[0002] This application relates generally to audio processing and more specifically to sub-band audio coding.BACKGROUND

[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted as prior art by inclusion in this section.

[0004] Linear Predictive Coding (LPC) is an audio coding technique that can be used for speech coding in a broad band time domain. On the other hand, transform-based audio codecs like MPEG AAC benefit from the high frequency resolution and associated coding gain especially for tonal signals. For some applications, it may be beneficial to encode the audio signal in a filterbank domain with relatively low frequency resolution such that there might not be significant coding gain for very tonal signals.

[0005] In one example, a split-rendering topology for an immersive voice and audio services (IVAS) codec may encode the audio signal in the filterbank domain. For this example, the IVAS decoder may encode the audio signal in the complex low delay filterbank (CLDFB) domain with a frequency resolution of about 400 Hz for 48 kHz audio sample rate. The encoded signal can then be transmitted (e.g., wirelessly) to a mobile device, which then performs rendering in the same CLDFB domain, resulting in the lowest possible delay.

[0006] One aim of examples described herein is to efficiently code audio signals directly in the CLDFB domain such that any additional delay and computational complexity is avoided whichwould otherwise come from signal domain conversions. This type of coding is labelled CLDFB codec in the following text. Since the CLDFB domain is complex valued, the signal may be oversampled by a factor of 2, which may make the compression task challenging.

[0007] It is with respect to these and other considerations that the disclosure made herein is presented.BRIEF SUMMARY OF THE DISCLOSURE

[0008] Techniques for coding, signaling, and decoding audio signals in the CLDFB domain are described herein. In some examples, these techniques may be applied to split rendering for an IV AS codec.

[0009] Briefly stated, systems and methods for sub-band audio coding are disclosed. One example provides a method for coding input audio signals in a bitstream from a first device to a second device. The example method includes, with the first device, obtaining a representation of the input audio signal in sub-bands, dividing a current frame into a current subset of the sub-bands that excludes other sub-bands, and determining a prediction parameter for the current subset of subbands. The example method may also include, with the first device, determining an active / passive flag for at least one of the sub-bands and encoding the bitstream with the current subset of the subbands, a number of subsets, the prediction parameter for the current subset of sub-bands, and the active / passive flag for the other subsets of sub-bands. The example method may include, with the second device, decoding the bitstream.

[0010] In some described embodiments, audio signals are coded for transmission or signaling between a first device and a second device by way of an efficient sub-band parameter signaling and processing scheme. The first device is configured to process a current frame into sub-bands that are divided into subsets, where a current subset is signaled (or transmitted) from the first device to the second device along with active / passive flags. The second device is configured to receive the prediction parameters for the current subset along with the active / passive flags. The received prediction parameters for the current subset can be used to set or reset the prediction parameters for the current subset, while previously sent prediction parameters can be used for active subsets that are not in the current subset.

[0011] In some embodiments, the active / passive flags may indicate whether prediction parameters are to be applied for a current audio channel of the input audio signal. In some further embodiments, the active / passive flag may indicate the active sub-bands that are enabled for prediction coding.

[0012] In some embodiments, a method to code an input audio signal in a bitstream from a first device to a second device is described, the method comprising: obtaining, with the first device, a representation of the input audio signal in sub-bands; dividing a current frame into a current subset of the sub-bands that excludes other sub-bands; determining, with the first device, prediction parameters for the current subset of the sub-bands; determining, with the first device, an active / passive flag for at least one of the sub-bands; and encoding the bitstream, with the first device, the current subset of sub-bands, the number of subsets, the prediction parameters for the current subset of sub-bands, and the active / passive flag for the other subsets of sub-bands; wherein the second device is configured to decode the current subset of sub-bands, the number of subsets, the prediction parameters, and the active / passive flag, and apply previously sent parameters for active subsets that are not for the current subset to continue the prediction for sub-bands in the active subsets. In some further embodiments, the representation of the input audio signal in sub-bands is obtained by converting the input audio signal in a time domain into sub-bands of a banded domain. In some further embodiments, a combination of the current subset of sub-bands with the other subbands includes all of the sub-bands.

[0013] In some additional embodiments, converting the input audio signal in the time domain into sub-bands of the banded domain includes processing the input audio signal in the time domain to a current frame in a CLDFB domain that includes a plurality of sub-bands. In other embodiments converting the input audio signal in the time domain into sub-bands of the banded domain includes either applying a modified discrete Fourier transform (MDFT) to the input audio signal in the time domain, or applying a modified discrete cosine transform (MDCT) to the input audio signal in the time domain.

[0014] In some additional embodiments, a method to decode audio signals from a bitstream from a first device to a second device is described, the method comprising: decoding from the bitstream, with the second device, a current subset of sub-bands, the prediction parameters associated to the current subset of sub-bands, and an active / passive flag for the current subset of sub-bands; andapplying previously sent parameters for active subsets that are not for the current subset to continue the prediction for sub-bands in the active subsets.

[0015] In still other embodiments, a method to decode audio signals from a bitstream from a first device to a second device, the method comprising: decoding from the bitstream for a current frame, with the second device, a current subset of sub-bands, prediction parameters, and an active / passive flag; identifying, with the second device, active sub-bands that are different from the current subset of sub-bands; determining, with the second device, a prediction gain for the identified active subbands that are different from the current subset of sub-bands based on one or more previously decoded prediction parameters; and concealing the active sub-bands that are different from the current subset of sub-bands when: one or more previous frames are lost; and the determined prediction gain exceeds a prediction gain threshold.

[0016] In yet other embodiments, methods to code an input audio signal in a bitstream from a first device to a second device are described. Some methods for the first device may comprise: obtaining audio data that represents the input audio signal arranged in sub-bands, dividing a current frame into a current subset of the sub-bands that excludes other sub-bands, determining: a prediction parameter for the current subset of the sub-bands, a prediction enable flag the current subset of the sub-bands, and an active / passive flag for at least one of the sub-bands, arranging the audio data for the current subset of sub-bands to be contained within one block, and encoding the bitstream with: the arranged audio data for the current subset of sub-bands, the number of subsets, the prediction parameter for the current subset of sub-bands, and the active / passive flag for the other subsets of sub-bands, wherein the encoded bitstream enables the second device to decode the current subset of sub-bands, the number of subsets, the prediction parameter, and the active / passive flag, and apply previously sent parameters for active subsets that are not for the current subset to continue the prediction for sub-bands in the active subsets.

[0017] In yet still other embodiments, methods are described to decode an input audio signal received by a first device and encoded into a bitstream, wherein the bitstream is transmitted from the first device to a second device. Some methods for the second device may comprise: receiving the bitstream, wherein the bitstream includes encoded audio data that represents the input audio signal in sub-bands, decoding from the bitstream a current subset of sub-bands, a prediction parameter, a prediction enable flag, and an active / passive flag, wherein the prediction enable flag enables theselection of one or more code books for the decoding, evaluating the active subsets to determine if the active subsets correspond to the current subset of sub-bands, and applying previously sent parameters when the active subsets that are not for the current subset of sub-bands to continue the prediction for sub-bands in the active subsets.

[0018] In some further embodiments, a decoder status is set based on frame loss. For example, methods may further include evaluating a bitstream to identify a frame loss, designating each of the subsets of sub-bands with a decoding unresolved status after the frame loss is identified, and setting a decoder status to decoding failed. When a valid frame is identified after the frame loss, methods may include designating the current subset of sub-bands and inactive (without prediction) subsets of sub-bands with a decoding resolved status after the valid frame is identified and setting the decoder status to decoding not failed. In some embodiments, methods may include concealing the subsets of sub-bands when the decoder status corresponds to decoding failed and enabling output of the subsets of sub-bands with a status of resolved when the decoder status corresponds to decoding not failed.

[0019] Various aspects of the present disclosure provide for processing of audio signals, and effect improvements in at least the technical fields of audio processing, audio encoding, audio decoding, virtual reality, and the like.

[0020] The embodiments described herein may be generally described as techniques, where the term “technique” may refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested by the context as applied herein.

[0021] Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associate drawings. This Summary is provided to introduce a selection of techniques in a simplified form, and not intended to identify key or essential features of the claimed subject matter, which are defined by the appended claims.DESCRIPTION OF THE DRAWINGS

[0022] These and other more detailed and specific features of various embodiments are more fully disclosed in the following description, reference being had to the accompanying drawings, in which:

[0023] FIG. 1 illustrates a block diagram of an IVAS coder / decoder framework for encoding and decoding IVAS bitstreams in which various aspects of the present disclosure can be practiced.

[0024] FIG. 2 illustrates a block diagram of another IVAS encoder / decoder framework for encoding and decoding IVAS bitstreams, which employs sub-band encoding and decoding in which various aspects of the present disclosure can be practiced.

[0025] FIG. 3 illustrates a block diagram of an encoding framework that is arranged according to one or more embodiments that utilize sub-band processing according to some aspects of the present disclosure.

[0026] FIG. 4 illustrates a diagram showing parameter prediction in a complex low-delay filter bank (CLDFB) topology that is arranged according to one or more embodiments that utilize sub-band according to some aspects of the present disclosure.

[0027] FIG. 5 illustrates a diagram of complex valued first order prediction parameters for a particular sub-band plotted over time.

[0028] FIG. 6 illustrates a table of a bitstream processing syntax that may be utilized in various embodiments that utilize sub-band processing according to some aspects of the present disclosure.

[0029] FIG. 7 illustrates a graph illustrating the results of a MUSHRA listening test for five listeners.

[0030] FIGS. 8A, 8B, 8C, and 8D illustrate example pseudocodes.

[0031] FIG. 9 illustrates a block diagram of various example methods for coding input audio signals in a bitstream from a first device to a second device according to some aspects of the present disclosure.

[0032] FIG. 10 illustrates a block diagram of various example methods for decoding audio signals from a bitstream from a first device to a second device according to some aspects of the present disclosure.

[0033] FIG. 11 illustrates a block diagram of various example methods for decoding audio signals from a bitstream from a first device to a second device according to some aspects of the present disclosure.

[0034] FIG. 12 illustrates a block diagram of various example methods for coding an input audio signal in a bitstream from a first device to a second device according to some aspects of the present disclosure.

[0035] FIG. 13 illustrates a block diagram of various example methods for decoding an input audio signal received by a first device and encoded into a bitstream according to some aspects of the present disclosure.

[0036] FIG. 14A illustrates a schematic block diagram of an example device architecture that may be used to implement various aspects of the present disclosure.

[0037] FIG. 14B illustrates a schematic block diagram of an example CPU implemented in the device architecture of FIG. 14A that may be used to implement various aspects of the present disclosure.DETAILED DESCRIPTION

[0038] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of various described embodiments with reference to the accompanying drawings. The illustrative embodiments in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes made, without departing from the spirit or scope of the present disclosure. In light of the present disclosure, it will be apparent to one of ordinary skill in the art that the various described features and implementations may be practiced without many of these specific details. In some instances, well- known methods, procedures, components, and circuits, have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features are described hereafter that can each be used independently of one another or with any combination of other features. Thus, the features may be arranged, substituted, combined, separated, or designed into other configurations, which is contemplated in light of the present disclosure.

[0039] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless the context clearly indicates otherwise. Such terms are to be read as having an inclusive meaning. For example, “A and B” may mean at least the following: “both A and B”, “at least both A and B”. As another example, “A or B” may mean at least the following: “at least A”, “at least B”, “both A and B”, “atleast both A and B”. As another example, “A and / or B” may mean at least the following: “A and B”, “A or B”. When an exclusive-or is intended, such will be specifically noted (e.g., “either A or B”, “at most one of A and B”). The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting, or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0040] Various Acronyms that may appear throughout this disclosure and in the associated claims and / or drawings are listed below. Other commonly used acronyms and terms of art may be excluded from this list in the interest of brevity. Thus, a short list of acronyms is provided below as an easy reference for the reader.IVAS - Immersive Voice and Audio Services LPC - Linear Predictive Coding CLDFB - Complex Low Delay Filter Bank SBA - Scene Based Audio SD - Spatial Metadata BS - BitstreamEVS - Enhanced Voice Services HOA - Higher Order Ambisonics FOA - First Order Ambisonics MAS A - Metadata Assisted Spatial Audio MDFT - Modified Discrete Fourier Transform MDCT - Modified Discrete Cosine Transform

[0041] FIG. 1 illustrates a block diagram of an immersive voice and audio services (IVAS) coder / decoder (“codec”) framework 100 for encoding and decoding IVAS bitstreams, according to one or more embodiments. IVAS is expected to support a range of audio service capabilities, including but not limited to mono to stereo upmixing and fully immersive audio encoding, decoding,and rendering. IVAS is also intended to be supported by a wide range of devices, endpoints, and network nodes, including but not limited to: mobile and smart phones, electronic tablets, personal computers, conference phones, conference rooms, virtual reality (VR) and augmented reality (AR) devices, home theatre devices, and other suitable devices.

[0042] The example IVAS codec 100 includes an IVAS encoder 101 and an IVAS decoder 104. The IVAS encoder 101 may be considered a first device that is located upstream from a second device (e.g., a mobile device) such as the IVAS decoder 104. Thus, the first device may also be referred to as an upstream device or an upstream encoder device, while the second device may be referred to as a downstream device or a downstream decoder device.

[0043] The IVAS encoder 101 includes a spatial encoder 102 and a core audio encoder 103. The input of the spatial encoder 102 corresponds to a first path 110. The spatial encoder 102 receives input audio (e.g., input audio content) via the first path 110. The spatial encoder 102 processes and encodes the received input audio. In some implementations, the spatial encoder 102 implements SPAR and Dir AC for analyzing / downmixing N_dmx spatial audio channels, as described in further detail below. The outputs of the spatial encoder 102 correspond to a second path 111 and a third path 112. The spatial encoder 102 is coupled to the core audio encoder 103 via the second path 111. The spatial encoder 102 is coupled to the IVAS decoder 104 via the third path 112. The output of the spatial encoder 102 includes a spatial metadata (MD) bitstream (BS) and N_dmx channels of spatial downmix. The N_dmx channels of spatial downmix are provided by the spatial encoder 102 to the core audio encoder 103 via the second path 111. The spatial MD BS is provided by the spatial encoder 102 to the IVAS decoder 104 via the third path 112. The spatial MD is quantized and entropy coded. In some implementations, quantization can include fine, moderate, coarse, and extra coarse quantization strategies and entropy coding can include Huffman or Arithmetic coding. The framework permits not more than three levels of quantization at a given operating mode; however, with decreasing bitrates, the three levels become increasingly coarser overall, to meet bitrate requirements.

[0044] The input of the core audio encoder 103 corresponds to the second path 111. The output of the core audio encoder 103 corresponds to a fourth path 113. The core audio encoder 103 (e.g., based on mono Enhanced Voice Services (EVS) encoding unit) encodes N_dmx channels N_dmx = 1-16 channels) of the spatial downmix into an audio bitstream, which is combined (via the fourthpath 113) with the spatial MD bitstream into an IVAS encoded bitstream transmitted to IVAS decoder 104 via the third path 112.

[0045] The IVAS decoder 104 includes a core audio decoder 105 (e.g., an EVS decoder) and a spatial decoder / renderer 106 (e.g., SPAR / DirAC). The input of the core audio decoder 105 corresponds to a fifth path 114. The core audio decoder 105 receives the audio bitstream via the fifth path 114. The core audio decoder 105 is configured to decode the audio bitstream extracted from the IVAS bitstream to recover the N_dmx audio channels. The output of the core audio decoder 105 corresponds to a sixth path 115. The core audio decoder 105 is coupled to the spatial decoder / renderer 106 via the sixth path 115. The core audio decoder 105 is configured to provide the N_dmx audio channels (e.g., the decoded spatial downmix) to the spatial decoder / renderer 106 via the sixth path 115.

[0046] The inputs of the spatial decoder / renderer 106 correspond to the third path 112 and the sixth path 115. The spatial decoder / renderer 106 receives the spatial MD bitstream from the IVAS encoder 101 via the third path 112 and receives the decoded spatial downmix from the core audio decoder 105 via the sixth path 115. The spatial decoder / renderer 106 decodes the spatial MD bitstream extracted from the IVAS bitstream to recover the spatial MD and synthesizes (e.g., renders) output audio channels using the spatial MD and a spatial upmix for playback on various audio systems with different speaker configurations and capabilities. The output of the spatial decoder / renderer 106 corresponds to a seventh path 116. The spatial decoder / renderer 106 provides the output audio (e.g., decoded audio) via the seventh path 116.

[0047] FIG. 2 illustrates a block diagram of another IVAS encoder / decoder (“codec”) framework 200 for encoding and decoding IVAS bitstreams, which employs sub-band encoding and decoding, according to one or more embodiments. The framework 200 includes the IVAS encoder 101, a first device 202, and a second device 208. The first device 202 may be, for example, a decoding device and / or a rendering device with a high amount of processing power, such as a server or personal desktop computer. The second device 208 may be, for example, a decoding device and / or a rendering device with less processing power than the first device 202, such as a mobile device or a virtual reality headset.

[0048] The input of the IVAS encoder 101 corresponds to a first path 201. The IVAS encoder 101 receives input audio via the first path 201. The first path 201 may correspond, for example, to thefirst path 110 of FIG. 1. The input audio may include, for example, First Order Ambisonics (FOA), Higher Order Ambisonics (HOA), Metadata Assisted Spatial Audio (MASA), multichannel, stereo, and / or mono audio content. The IVAS encoder 101 may operate substantially the same as that described with respect to FIG. 1 to generate an IVAS bitstream. The IVAS encoder 101 is configured to provide an IVAS bitstream to the first device 202 via a second path 203. The second path 203 may correspond, for example, to the third path 112 of FIG. 1.

[0049] The first device 202 includes the IVAS decoder 104, an IVAS Tenderer and metadata encoder 204, and a sub-band encoder 206. The input of the IVAS decoder 104 corresponds to the second path 203. The IVAS decoder 104 is coupled to the IVAS encoder 101 via the second path 203. The output of the IVAS decoder 104 corresponds to a third path 205. The IVAS decoder 104 is coupled to the IVAS Tenderer and metadata encoder 204 via the third path 205. The third path 205 may corresponds, for example, to the seventh path 116 of FIG. 1. The IVAS decoder 104 may operate substantially the same as that described with respect to FIG. 1 to generate decoded audio. The IVAS decoder 104 decodes the audio in the decoder’s native sub-band domain. The IVAS decoder 104 is configured to provide the decoded audio to the IVAS Tenderer and metadata encoder 204 via the third path 205.

[0050] The inputs of the IVAS Tenderer and metadata encoder 204 correspond to the third path 205 and a fourth path 207. The IVAS Tenderer and metadata encoder 204 is coupled to the IVAS decoder 104 via the third path 205. The IVAS Tenderer and metadata encoder 204 receives the decoded audio from the IVAS decoder 104 via the third path 205. The IVAS Tenderer and metadata encoder 204 is coupled to a tracking module 214 in the second device 208 via the fourth path 207. The IVAS Tenderer and metadata encoder 204 receives pose information from the tracking module 214 via the fourth path 207, as described below in more detail. The outputs of the IVAS Tenderer and metadata encoder 204 correspond to a fifth path 209 and a sixth path 211. The IVAS Tenderer and metadata encoder 204 is coupled to the sub-band encoder 206 via the fifth path 209. The IVAS Tenderer and metadata encoder 204 is coupled to a pose correction Tenderer 212 in the second device 208 via the sixth path 211.

[0051] The IVAS Tenderer and metadata encoder 204 processes the sub-band domain decoded audio received from the IVAS decoder 104 using a head pose indicated by the pose information from the tracking module 214. In instances where the first device 202 is coupled to the second device 208 viaa wireless connection, the head pose may be received with a time delay. Based on the head pose, the IVAS Tenderer and metadata encoder 204 produces one or multiple binaural representations of the decoded audio. The IVAS Tenderer and metadata encoder 204 also produces pose correction metadata. The IVAS Tenderer and metadata encoder 204 provides the pose correction metadata to the pose correction Tenderer 212 via the sixth path 211. The IVAS Tenderer and metadata encoder 204 provides rendered audio, which includes the one or multiple binaural representations of the decoded audio and the corresponding sub-bands, to the sub-band encoder 206 via the fifth path 209.

[0052] The inputs to the sub-band encoder 206 correspond to the fifth path 209 and a seventh path 213. The sub-band encoder 206 is coupled to the IVAS Tenderer and metadata encoder 204 via the fifth path 209. The sub-band encoder 206 receives rendered audio from the IVAS Tenderer and metadata encoder 204 via the fifth path 209. The sub-band encoder 206 receives frame length information via the seventh path 213. The frame length information may include the default frame length (in milliseconds) and / or a short frame length (in milliseconds). In some instances, such as for an IVAS codec, the default frame length is 20 milliseconds. The short frame length may be, for example, 5 milliseconds, 10 milliseconds, or the like. The default frame length and / or the short frame length may be defined by a user input to the sub-band encoder 206. In other instances, the default frame length and / or the short frame length are default values stored in a memory associated with the sub-band encoder 206 (for example, memory 1421 of FIG. 14B).

[0053] The sub-band encoder 206 is configured to encode the rendered binaural audio received from the sub-band encoder 206 to generate a sub-band coder bitstream. In some instances, the sub-band encoder 206 uses linear predictive coding (LPC) when encoding the rendered audio. The sub-band encoder 206 also encodes prediction parameters metadata using subset processing and signaling methods, as described in more detail below. The output of the sub-band encoder 206 corresponds to an eighth path 215. The sub-band encoder 206 is coupled to a sub-band decoder 210 via the eighth path 215. The sub-band encoder 206 provides the sub-band coder bitstream, which includes the prediction parameters metadata, to the sub-band decoder 210 via the eighth path 215.

[0054] The second device 208 includes the sub-band decoder 210, the pose correction Tenderer 212, the tracking module 214, and an output device 216. The input of the sub-band decoder 210 corresponds to the eighth path 215. The sub-band decoder 210 is coupled to the sub-band encoder 206 via the eighth path 215. The sub-band decoder 210 receives the sub-band coder bitstream fromthe sub-band encoder 206 via the eighth path 215. The sub-band decoder 210 is configured to decode the received sub-band coder bitstream. The output of the sub-band decoder 210 corresponds to a ninth path 217. The sub-band decoder 210 is coupled to the pose correction Tenderer 212 via the ninth path 217. The sub-band decoder 210 provides the decoded bitstream to the pose correction Tenderer 212 via the ninth path 217. Further operations of the sub-band encoder 206 and the second device 208 are described with respect to FIG. 3.

[0055] The inputs of the pose correction Tenderer 212 corresponds to the sixth path 211, the ninth path 217, and a tenth path 219. The pose correction Tenderer 212 is coupled to the IVAS Tenderer and metadata encoder 204 via the sixth path 211. The pose correction Tenderer 212 receives pose correction metadata from the IVAS Tenderer and metadata encoder 204 via the sixth path 211. The pose correction Tenderer 212 is coupled to the sub-band decoder 210 via the ninth path 217. The pose correction Tenderer 212 receives the decoded bitstream from the sub-band decoder 210 via the ninth path 217. The pose correction Tenderer 212 is coupled to the tracking module 214 via the tenth path 219. The pose correction Tenderer 212 receives pose information from the tracking module 214 via the tenth path 219.

[0056] The pose correction Tenderer 212 is configured to render the received decoded bitstream based on the instant head pose (indicated in the pose information provided by the tracking module 214) and the pose correction metadata provided by the IVAS Tenderer and metadata encoder 204, thereby generating rendered audio. The output of the pose correction Tenderer 212 corresponds to an eleventh path 221. The pose correction Tenderer 212 is coupled to the output device 216 via the eleventh path 221. The pose correction Tenderer 212 is configured to provide the rendered audio to the output device 216 via the eleventh path 221.

[0057] The input of the output device 216 corresponds to the eleventh path 221. The output device 216 is coupled to the pose correction Tenderer 212 via the eleventh path 221. The output device 216 receives rendered audio from the pose correction Tenderer 212 via the eleventh path 221. The output device 216 is configured to output the rendered audio to an environment. The output device 216 may be, for example, a speaker or a plurality of speakers. The environment may be, for example, a listening environment in which a listener of the rendered audio is situated, such as movie theater, a home theater, a conference room, a physical space for use of a virtual reality headset, or the like.

[0058] Since intermediate sub-band to time domain conversions are avoided using the framework 200, the framework 200 is computationally efficient and offers a shortened possible delay if the IV AS decoded audio is present in the sub-band domain. In addition, the rendering operation in the second device 208 is computationally efficient, as rendering in the second device 208 may not be necessary for examples where pose correction is not needed. In such instances, where rendering in the second device 208 may not be necessary, rendering may instead be performed by the first device 202. Additionally, the tracked head pose correction may be applied instantaneously in the second device 208 (e.g., without any significant delay that may occur in the transmission of the head pose from the second device 208 to the IVAS Tenderer and metadata encoder 204).

[0059] To improve coding efficiency for tonal signals, LPC may be applied to sub-bands in the filterbank domain, which utilizes transmission of prediction parameters and creates side data rate overhead. In some configurations of such a system, codecs may be operated with a short frame length, which may lead to increased prediction metadata rates, and which in turn may reduce or eliminate the coding gain from prediction. As one example, a default frame length (also referred to as a long frame length) as described herein may refer to a frame length of approximately 20 ms with 960 samples at a 48 kHz sampling rate. An example default frame length may be signaled at a rate of approximately 10.5 kbit per second with 210 bits per frame. A short frame length as described herein may refer to a frame length that is shorter than the default frame length. An example short frame length may be approximately 5 ms long with 240 samples. An example short frame length may be signaled at a rate of approximately 10.5 kbit per second with 52.5 bits per short frame. Accordingly, the 52.5 bits of four short frames are equal to the 210 bits per frame of the default frame length.

[0060] One solution identified for this above-described problem is to send prediction parameters for all sub-bands only at a certain frame interval. However, this may result in a high bitrate peak for the first frame (i-frame) in the interval, and recovery after packet loss would be problematic if not starting with an i-frame.

[0061] Examples described herein provide a prediction metadata signaling and processing scheme that keeps the prediction metadata rate approximately constant for a reduced frame length (e.g., the short frame length), which enables significant coding gain using prediction and limits the impact from frame losses.

[0062] FIG. 3 illustrates a block diagram of an encoding framework 300 that is arranged according to one or more embodiments that utilize sub-band processing as described herein. The framework 300 includes a sub-band processor 302, an audio buffer 304, a predictor computation module 306, a prediction and quantization coding module 308, a prediction parameters buffer 310, a quantization coding module 312, and a multiplexer 314. In some examples, operations of the encoding framework 300 may be performed by the sub-band encoder 206 of FIG. 2.

[0063] The input of the sub-band processor 302 corresponds to a first path 301. The sub-band processor 302 receives input audio signals via the first path 301. The sub-band processor 302 processes a frame of the input audio signal and generates N-sub-bands for the audio in the filterbank domain. For example, the input audio signal from the time domain may be transformed into a banded domain using a modified discrete Fourier transform (MDFT) process that creates a complex modulated multi-channel filter bank. In some other examples, the input audio signal is transformed into a banded domain including time domain-based filters or using a modified discrete cosine transform (MDCT) process, without requiring complex values. In other examples, the input audio signal may already be obtained in the banded domain. In such an example, the sub-band processor 302 may be omitted from the encoding framework 300. The output of the sub-band processor 302 corresponds to second paths 303. The sub-band processor 302 is coupled to the audio buffer 304 via the second paths 303. The sub-band processor 302 provides the audio sub-bands to the audio buffer 304 via the second paths 303, where the audio sub-bands include N sub-bands. The audio sub-bands may be provided in packets that combine to the default frame length (e.g., four 5 ms short frames that combine to form a 20 ms total frame length). In other instances, the audio sub-bands may be provided in 5 ms short frames.

[0064] The input of the audio buffer 304 corresponds to the second paths 303. The audio buffer 304 is coupled to the sub-band processor 302 via the second paths 303. The audio buffer 304 receives the audio sub-bands from the sub-band processor 302 via the second paths 303. The audio buffer 304 is configured to buffer (e.g., store) the audio sub-bands for the frame. The audio buffer 304 may store, for example, 20 ms of audio frames. In such an example, when the audio sub-bands are provided in 5 ms short frames, four short frames of audio data may be stored in the audio buffer 304. The output of the audio buffer 304 corresponds to third paths 305. The audio buffer 304 is coupled to the predictor computation module 306 via the third paths 305. The audio buffer 304 is configured to provide the buffered audio sub-bands to the predictor computation module 306 via the third paths305, where the buffered audio sub-bands include M sub-bands and where M may be different than N or may be the same as N.

[0065] The input of the predictor computation module 306 corresponds to the third paths 305. The predictor computation module 306 is coupled to the audio buffer 304 via the third paths 305. The predictor computation module 306 receives the buffered audio sub-bands from the audio buffer 304 via the third paths 305. The predictor computation module 306 processes the received buffered audio sub-bands to determine (e.g., estimate or calculate) prediction coefficients (e.g., prediction parameters) for the sub-bands. For example, a subset of the audio sub-bands may be evaluated by the predictor computation module 306 using LPC techniques for each of the subsets, thereby generating prediction parameters for each of the subsets of the bands. The predictor computation module 306 may also quantize the prediction parameters for the current subset of sub-bands. In some instances, the predictor computation module 306 estimates a prediction gain and a decision on whether prediction shall be used for each audio channel (e.g., the decision may result in setting a prediction enabled flag as one of the prediction parameters).

[0066] The predictor computation module 306 may also set the active / passive flags for the current subset of sub-bands. For example, the predictor computation module 306 may flag the current subset as active for any sub-band in the current subset where the prediction gain exceeds a threshold (for example, the prediction gain is greater than 4). In instances where any sub-band in the current subset does not have a prediction gain that exceeds the threshold, the current subset is flagged as passive (e.g., inactive). In some instances, the predictor computation module 306 may also set the active / passive flag for subsets that are not the current subset.

[0067] In some instances, the predictor computation module 306 determines an identifier for the current subset (e.g., a current subset ID). For example, the predictor computation module 306 may assign, for each subset of sub-bands, a subset identifier by incrementing the subset identifier for each subsequent short frame. As one example, a first identifier is 0 in a first short frame. In a second short frame, the identifier is incremented to 1. In a third short frame, the identifier is incremented to 2. In a fourth short frame, the identifier is incremented to 3. Once a maximum value is reached (for example, 3), the identifier is set back to the minimum value (for example, 0) in the next subsequent short frame. The blocks of the bitstream may be arranged according to the assigned subset IDs. For example, the blocks of the bitstream may be arranged beginning with the currentsubset ID and then decrementing the subset IDs, while the predictor computation module 306 concurrently increments the current subset ID each short frame.

[0068] The outputs of the predictor computation module 306 corresponds to fourth paths 307 and a fifth path 309. The predictor computation module 306 is coupled to the prediction and quantization coding module 308 via the fourth paths 307. The predictor computation module 306 is coupled to the prediction parameters buffer 310 and the quantization coding module 312 via the fifth path 309. The predictor computation module 306 is configured to provide the current prediction parameters for the current subset of sub-bands to the prediction and quantization coding module 308, the prediction parameters buffer 310 and the quantization coding module 312 via the fourth paths 307 and the fifth path 309, respectively.

[0069] The inputs of the prediction and quantization coding module 308 correspond to the fourth paths 307 and a sixth path 311. The prediction and quantization coding module 308 is coupled to the predictor computation module 306 via the fourth paths 307. The prediction and quantization coding module 308 receives the current prediction parameters from the predictor computation module 306 via the fourth paths 307. The prediction and quantization coding module 308 is coupled to the prediction parameters buffer 310 via the sixth path 311. The prediction and quantization coding module 308 receives previous prediction parameters from the prediction parameters buffer 310 via the sixth path 311.

[0070] The prediction and quantization coding module 308 is configured to apply the prediction parameters to audio data if the prediction enabled flag is enabled (e.g., high) for the corresponding audio data. The prediction and quantization coding module 308 is also configured to quantize and code residual audio data. The output of the prediction and quantization coding module 308 corresponds to a seventh path 313. The prediction and quantization coding module 308 is coupled to the multiplexer 314 via the seventh path 313. The prediction and quantization coding module 308 is configured to provide the current coded prediction parameters to the multiplexer 314 via the seventh path 313.

[0071] The inputs of the prediction parameters buffer 310 correspond to the fifth path 309 and an eighth path 315. The prediction parameters buffer 310 is coupled to the predictor computation module 306 via the fifth path 309. The prediction parameters buffer 310 receives the current prediction parameters from the predictor computation module 306 via the fifth path 309. Theprediction parameters buffer 310 receives subset metadata via the eighth path 315. The prediction parameters buffer 310 is configured to store (e.g., buffer) the received current prediction parameters. As the predictor computation module 306 only determines prediction parameters for a current subset of sub-bands of interest, the prediction parameters from previous subsets of sub-bands (e.g., other subsets) may be retrieved from the prediction parameters buffer 310, where they are stored. The subset metadata may indicate a current subset of sub-bands (via, for example, a subset identifier) that are being analyzed by the predictor computation module 306. Accordingly, the prediction parameters buffer 310 may refer to the subset metadata to identify the other sub-bands and previous prediction parameters associated with the other sub-bands to provide to the prediction and quantization coding module 308. The subset metadata may be provided, for example, by the subband encoder 206.

[0072] The output of the prediction parameters buffer 310 corresponds to the sixth path 311. The prediction parameters buffer 310 is coupled to the prediction and quantization coding module 308 via the sixth path 311. In a subsequent short frame, the prediction parameters buffer 310 is configured to provide previous prediction parameters (which relate to sub-bands that were previously analyzed by the predictor computation module 306) to the prediction and quantization coding module 308 via the sixth path 311.

[0073] The input of the quantization coding module 312 corresponds to the fifth path 309. The quantization coding module 312 is coupled to the predictor computation module 306 via the fifth path 309. The quantization coding module 312 receives the current prediction parameters from the predictor computation module 306 via the fifth path 309. The quantization coding module 312 is configured to quantize the current prediction parameter to generate a current coded prediction parameter. The output of the quantization coding module 312 corresponds to a ninth path 317. The quantization coding module 312 is configured to provide the current coded prediction parameter to the multiplexer 314 via the ninth path 317.

[0074] The inputs of the multiplexer 314 correspond to the seventh path 313, the eighth path 315, and the ninth path 317. The multiplexer 314 is coupled to the prediction and quantization coding module 308 via the seventh path 313. The multiplexer 314 receives the coded prediction parameters from the prediction and quantization coding module 308 via the seventh path 313. The multiplexer 314 is coupled to the quantization coding module 312 via the ninth path 317. The multiplexer 314receives the current coded prediction parameter from the quantization coding module 312 via the ninth path 317. The multiplexer 314 receives subset metadata via the eighth path 315.

[0075] The multiplexer 314 is configured to selectively combine the prediction parameters for the current subset and the other sub-bands for signaling (e.g., transmission) in the bitstream to a downstream device (e.g., the second device 208). The output of the multiplexer 314 corresponds to a tenth path 319. The multiplexer 314 is configured to provide the bitstream to a downstream device via the tenth path 319.

[0076] The current subset may be signaled or transmitted from the upstream encoder device to the downstream decoder device, including transmission of active / passive flags for other subsets. The active / passive flags may accordingly be included in the subset metadata received via the eighth path 315. Prediction may be disabled by the encoder in the sub-bands of a non-current subset by setting the corresponding active flag to 0. For sub-bands that belong to the active subsets that are not included in the current subset, previously transmitted prediction parameters (for example, prediction parameters stored in the prediction parameters buffer 310) are applied similarly by the encoder and the decoder.

[0077] The number of possible subsets may be determined based on the ratio of the default frame length to the current frame length. In other instances, the number of possible subsets is transmitted in the bitstream. The audio quality impact compared to the operation at the default frame length is small since the sub-band prediction parameters for tonal signals typically vary slowly over time.

[0078] In some examples, a default frame length is used, prediction parameters are transmitted in a current frame with a shorter frame length than the default frame length, and the prediction parameters that are transmitted over the shorter frame length are only for the sub-bands belonging to a certain subset of the sub-bands. For example, with a default frame length of 20 ms, the total number of sub-bands may be divided into four subsets. The prediction parameters are transmitted in a frame having a frame length that is one-fourth the length of the default frame length (e.g., 5 ms). Over time, the multiple shorter frames of 5 ms length are transmitted from the upstream device to the downstream device so that the prediction parameters for all of the sub-bands are eventually received at the downstream device.

[0079] FIG. 4 illustrates a diagram 400 showing parameter prediction in a complex low-delay filter bank (CLDFB) topology that is arranged according to one or more embodiments that utilize subband processing as described herein.

[0080] As illustrated in FIG. 4, audio frequency sub-bands 402 are input to the prediction encoder, where each of the sub-bands 402 is represented by a row and the time dimension is represented by columns. In the example shown, one audio frame covers four time samples (columns) of audio subband data. Sub-bands 402 belonging to the current subset of sub-bands are marked in bolded solid diagonal lines. Sub-bands belonging to a certain subset are arranged in an interleaved way as indicated by the different patterns.

[0081] Prediction parameters may be updated by the upstream device (e.g., by a sub-band encoder) only for the current subset, where the indicator for the current subset (for example, the active / passive flag) is transmitted in the bitstream from the upstream device (e.g., the first device 202) to the downstream device (e.g., the second device 208). In an example implementation, the selection of the current subset by the upstream device can be determined by a rotational selection from among all of the possible subsets (four in this example). However, this is merely one example, and other selection methods are contemplated such as random selection, least recently / frequently selected, or a simple queue such as a linear queue (e.g., FIFO, LIFO, etc.), a rotational queue, or any other reasonable topology. In general, any subset may be signaled as the current subset. The number of subsets may also be transmitted from the upstream device to the downstream device, giving full control over the transmission of prediction parameters.

[0082] Prediction parameters for the current subset may be updated by the upstream device based on the buffered sub-bands, where the buffer length corresponds for the illustrated example has a 20ms duration (the prediction analysis window length). If the estimated prediction gains are sufficiently high (considering that prediction coefficients for the current subset are transmitted only every four frames in this example), then the current subset active flag is set to true, prediction parameters are transmitted from the upstream device to the downstream device, and prediction is applied by the downstream device (except for the first sample in every sub-band). Otherwise, the current subset active flag is set to false and prediction is not applied by the downstream device. Sub-bands belonging to active, non-current subsets, are processed by the downstream device using previously computed prediction coefficients and using the prediction state from the previous frame. In theexample shown, prediction is applied to all sub-bands but only the prediction coefficients for the sub-bands marked in black have been updated.

[0083] FIG. 5 illustrates a diagram of complex valued first order prediction parameters for a particular sub-band plotted over time. In the example of FIG. 5, the complex valued first order prediction parameters are plotted over time for a tonal signal (e.g., a pitch pipe), for a complex low- delay filter bank (CLDFB) topology that is arranged according to one or more embodiments that utilize sub-band processing as described herein.

[0084] As shown in FIG. 5, the prediction parameters for one sub-band may vary slowly over time. The optimal prediction coefficient(s) in the evaluated sub-band are computed every frame (for example, every 20 ms when the prediction analysis window length is 20 ms, or every 5 ms when the prediction analysis window length is 5 ms) and quantized using magnitude and phase representation. As shown, the optimal quantized prediction parameter remains substantially constant over an extended period of time. The period of time may even exceed the default frame length of 20 ms. Accordingly, for some signals, it may be sufficient to transmit prediction coefficients less frequently than the short frame length (or, in some cases, less frequently than the default frame length). It may be assumed that the prediction parameter for a corresponding sub-band remains substantially constant for a certain amount of time. In a situation where the frame length is shorter than the prediction analysis window length of 20 ms, it may be assumed that the optimal prediction parameter remains unchanged at least for the duration of the analysis window and the side data rate may be reduced by transmitting prediction parameters less frequently with limited or no loss of coding gain. Accordingly, in some instances, prediction may be applied only when signals have a tonal character that does not change between short frame lengths.

[0085] FIG. 6 illustrates a table 600 of a bitstream processing syntax that may be utilized in various embodiments that utilize sub-band processing as described herein. In the example table 600, two bits are implemented for the number of subsets available (e.g., 1 through 4). However, in other examples, this number of bits may be changed to accommodate any number of subsets as may be required for a specific implementation (e.g., 3 bits for 8 subsets, 4 bits for 16 subsets, etc.). Additionally, in the example of table 600, the total number of bands is 60, and in other examples the number of bands per subset may need to be adjusted to accommodate a different number of bands.

[0086] FIG. 7 illustrates a graph illustrating the results of a MUSHRA listening test for five listeners. For the listening test, the stereo audio items at a bitrate of 384 kb / s were compared under three test conditions: (i) 20 ms default frame length (default operation, prediction parameters transmitted for all bands every frame), (ii) 5 ms short frame length old (no prediction applied, no prediction parameters transmitted due to high side rate and no prediction gain), and (iii) 5 ms frame length new (with prediction applied using signaling and processing schemes described herein). The results illustrated in FIG. 7 illustrates that the condition for a 5 ms short frame length that leverages the sub-band signaling and processing scheme described herein yield a significantly better result than without the sub-band signaling and processing scheme.

[0087] The above-described methods, systems and devices may be applied for lost data recovery. For example, frames are transmitted in the bitstream from an upstream device to a downstream device over time. The prediction parameters (e.g., prediction coefficients) for each current frame are decoded from the bitstream and stored in a buffer at the downstream device. When a packet or frame is lost in the transmission, the previously acquired and stored prediction coefficients may be evaluated to determine if the prediction parameters are reliable. The prediction parameters are then used to generate the lost information.

[0088] An example process flow for lost frame recovery by the downstream device may be described as follows. The bitstream for a current frame is decoded by the downstream device, which includes a current subset of sub-bands, a prediction parameter(s), and an active / passive flag. Active sub-bands that are different from the current set of sub-bands are identified based on the decoded current frame. A prediction gain is determined or computed for the identified active sub-bands that are different from the current subset of sub-bands based on one or more previously decoded prediction parameters. The prediction gain is evaluated to see if the prior prediction parameter(s) is / are reliable. If the prediction gain is sufficient (e.g., exceeds a threshold), then when one or more of the previous packets or frames are lost, the active sub-bands that are different from the current subset of sub-bands are concealed.

[0089] Concealment of the lost packet / frames may be accomplished by generating replacement audio for the concealed active sub-bands. The replacement audio is based on the one or more associated prediction parameters received in prior frames that were not lost. One or more decoded sub-band residual signals may also be used in reconstructing the lost information.

[0090] Another example process flow for lost frame recovery is described below. Before decoding the first good (or valid) frame after a frame loss, all subsets of sub-bands may be given an “unresolved” status in persistent decoder memory. When reading a new good (or valid) frame after the frame loss, the current subset of sub-bands and passive subsets of sub-bands may be given a status of “resolved”, while the “unresolved” status of other subsets remains unchanged. Subsets of sub-bands that have “unresolved” status may be concealed, while subsets of sub-bands with “resolved” status may be used for output.

[0091] In various examples of frame loss, prediction information for subsets of sub-bands may be required to be able to decode the actual audio data. In such examples, the audio data associated with a subset with an “unresolved” status may not be able to be decoded. In some instances, it may not be possible to further parse the bitstream at all. However, by employing the herein described lost frame recovery techniques, the status information “unresolved” may be used to detect such a situation and stop parsing the bitstream when the first “unresolved” is reached. Therefore, in cases where a dependency exists between prediction metadata and the audio data in the bitstream, the bitstream may be arranged differently to achieve an improved frame loss recovery topology.

[0092] An example bitstream arrangement for improved frame loss recovery may be described as follows. In general, all audio data (including all audio channels) corresponding to a subset of subbands may be written in one block. Furthermore, since the current subset of sub-bands does not have a status of “unresolved”, the data block corresponding to the current subset may be written as a first block of data. Therefore, the first block of data may always be successfully read and decoded correctly. If the next subset of sub-bands to be read has a status of “unresolved”, then the decoding process is stopped. In an example process, the current subset as specified by the encoder loops through the number of subsets of sub-bands in an ascending manner. Thus, the rest of the data blocks may be written by looping through the number of subsets of sub-bands in a descending manner. This described approach ensures that with every new good frame after a frame loss, the number of correctly decodable subsets of sub-bands increases at least by 1. An example representation of the blocks in the bitstream for two audio channels using N as the number of subsets and C as the current subset is provided below:• Channel 1 o C: N: NumTotalBands• Channel! o C: N: NumTotalBands....• Channel 1 o Modulo(C-l, N): N: NumTotalBands• Channel 2 o Modulo(C-l, N): N: NumTotalBands o ...• Channel 1 o Modulo(C-N+l, N): NumSubSets: NumTotalBands• Channel 2 o Modulo(C-N+l, N): NumSubSets: NumTotalBands

[0093] Similar to the way the decoder tracs the “unresolved” status of prediction metadata, the decoder may also track subsets of sub-bands that have been decoded successfully after frame loss. This information may be used to decide which sub-bands should remain concealed after a frame loss. One way to provide concealment of the not-yet decoded sub-bands is to fill the corresponding audio samples with zeros.

[0094] As a particular example of the described frame recovery process that may include steps for encoding bitstreams, decoding bitstreams, as well as detection and recovery logic, refer to Pseudocodes shown in FIGS. 8A-8D.

[0095] FIG. 9 illustrates a block diagram of various example methods 900 for coding input audio signals in a bitstream from a first device to a second device, which may be performed by the framework 200 of FIG. 2, the encoding framework 300 of FIG. 3, or a combination thereof. The methods 900 may be performed by a one or more processors, which may be configured to perform methods 900 via machine-executable instructions. The methods 900 may be broken into various blocks or partitions, such as blocks 902, 904, 906, 908, and 910. The various process blocks illustrated in FIG. 9 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 902.

[0096] At block 902, “Obtaining A Representation Of The Input Audio Signal In Sub-bands,” an example method 900 may include obtaining, with the first device 202, a representation of the input audio signal in sub-bands. For example, with reference to FIG. 3, the sub-band processor 302receiving an input audio signal via the first path 301 and processing the received input audio signal to obtain a plurality of audio sub-bands. Processing may proceed from block 902 to block 904.

[0097] At block 904, “Dividing A Current Frame Into A Current Subset Of The Sub-bands That Excludes Other Sub-bands,” an example method 900 may include dividing a current frame into a current subset of the sub-bands that excludes other sub-bands. For example, with reference to FIG. 3, the predictor computation module 306 may divide the plurality of audio sub-bands into a current subset of the sub-bands that excludes other sub-bands. Processing may proceed from block 904 to block 906.

[0098] At block 906, “Determining A Prediction Parameter For The Current Subset Of Sub-bands,” an example method 900 may include determining, with the first device 202, a prediction parameter for the current subset of the sub-bands. For example, with reference to FIG. 3, the predictor computation module 306 may determine (e.g., calculate or estimate) a prediction parameter for the current subset of the sub-bands. Processing may proceed from block 906 to block 908.

[0099] At block 908, “Determining An Active / Passive Flag For At Least One Of The Sub-bands,” an example method 900 may include determining, with the first device 202, an active / passive flag for at least one of the sub-bands. For example, with reference to FIG. 2, the sub-band encoder 206 may generate an active / passive flag indicating which of the sub-bands are currently active and correspond to the determined prediction parameters. As another example, with reference to FIG. 3, the predictor computation module 306 may generate an active / passive flag indicating which of the sub-bands are currently active and correspond to the determined prediction parameters. Processing may proceed from block 908 to block 910.

[0100] At block 910, “Encoding The Bitstream, The Current Subset Of Sub-bands, The Number Of Sub-bands, The Prediction Parameter For The Current Set Of Sub-bands, And The Active / Passive Flag For The Other Set Of Sub-bands,” an example method 900 may include encoding the bitstream, with the first device 202, the current subset of the sub-bands, a number of subsets, the prediction parameter for the current subset of sub-bands, and the active / passive flag for the other subsets of sub-bands. For example, with reference to FIG. 2, the sub-band encoder 206 encodes the bitstream including the current subset of the sub-bands, the number of subsets, the prediction parameter for the current subset of sub-bands, and the active / passive flag for the othersubsets of sub-bands. In another example, with reference to FIG. 3, the multiplexer 314 encodes the bitstream.

[0101] FIG. 10 illustrates a block diagram of various example methods 1000 for decoding audio signals from a bitstream from a first device to a second device, which may be performed by the framework 200 of FIG. 2, the encoding framework 300 of FIG. 3, or a combination thereof. The methods 1000 may be performed by one or more processors, which may be configured to perform methods 1000 via machine-executable instructions. The methods 1000 may be broken into various blocks or partitions, such as blocks 1002 and 1004. The various process blocks illustrated in FIG.10 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 1002.

[0102] At block 1002, “Decoding From The Bitstream A Current Subset Of Sub-bands, The Prediction Parameter, And An Active / Passive Flag,” an example method 1000 may include decoding from the bitstream, with the second device 208, a current subset of sub-bands, a prediction parameter, and an active / passive flag. Processing may proceed from block 1002 to block 1004.

[0103] At block 1004, “Applying Previously Sent Parameters For Active Subsets That Are Not For The Current Subset To Continue The Prediction For Sub-bands In The Active Subsets,” an example method 1000 may include applying previously sent parameters for active subsets that are not for the current subset to continue the prediction for sub-bands in the active sets. In this manner, previously- received prediction parameters are used for rendering non-current subsets of sub-bands.

[0100] FIG. 11 illustrates a block diagram of various example methods 1100 for decoding audio signals from a bitstream from a first device to a second device, which may be performed by the framework 200 of FIG. 2, the encoding framework 300 of FIG. 3, or a combination thereof. The methods 1100 may be performed by one or more processors, which may be configured to perform methods 1100 via machine-executable instructions. The methods 1100 may be broken into various blocks or partitions, such as blocks 1102, 1104, 1106, and 1108. The various process blocks illustrated in FIG. 11 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may bedescribed as processes, methods, steps, blocks, operations, or functions, may commence at block 1102.

[0101] At block 1102, “Decoding From The Bitstream For A Current Frame A Current Subset Of Sub-bands, A Prediction Parameter, And An Active / Passive Flag,” an example method 1100 may include decoding from the bitstream for a current frame, with the second device 208, a current subset of sub-bands, a prediction parameter, and an active / passive flag. Processing may proceed from block 1102 to block 1104.

[0102] At block 1104, “Identifying Active Sub-bands That Are Different From The Current Subset Of Sub-bands,” an example method 1100 may include identifying, with the second device 208, active sub-bands that are different from the current subset of sub-bands. In some instances, the second device 208 identifies the active sub-bands that are different from the current subset of subbands based on the active / passive flag. Processing may proceed from block 1104 to block 1106.

[0103] At block 1106, “Determining A Prediction Gain For The Identified Active Sub-bands That Are Different From The Current Subset Of Sub-bands Based On One Or More Previously Decoded Prediction Parameters,” an example method 1100 may include determining, with the second device 208, a prediction gain for the identified active sub-bands that are different from the current subset of sub-bands based on one or more previously decoded prediction parameters. Processing may proceed from block 1106 to block 1108.

[0104] At block 1108, “Concealing The Active Sub-bands That Are Different From The Current Subset Of sub-bands,” an example method 1100 may include concealing the active sub-bands that are different from the current subset of sub-bands. The active sub-bands may be concealed when, for example, one or more previous frames are lost and / or the determined prediction gain exceeds a prediction gain threshold.

[0105] FIG. 12 illustrates a block diagram of various example methods 1200 for coding an input audio signal in a bitstream from a first device to a second device, which may be performed by the framework 200 of FIG. 2, the encoding framework 300 of FIG. 3, or a combination thereof. The methods 1200 may be performed by one or more processors, which may be configured to perform methods 1200 via machine-executable instructions. The methods 1200 may be broken into various blocks or partitions, such as blocks 1202, 1204, 1206, 1208, and 1210. The various process blocksillustrated in FIG. 12 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 1202.

[0106] At block 1202, “Obtaining Audio Data That Represents The Input Audio Signal Arranged In Sub-bands,” an example method 1200 may include obtaining audio data that represents the input audio signal arranged in sub-bands. Processing may proceed from block 1202 to block 1204.

[0107] At block 1204, “Dividing A Current Frame Into A Current Subset Of The Sub-bands That Excludes Other Sub-bands,” an example method 1200 may include dividing a current frame into a current subset of the sub-bands that excludes other sub-bands. Processing may proceed from block 1204 to block 1206.

[0108] At block 1206, “Determining A Prediction Parameter For The Current Subset Of Sub-bands, A Prediction Enable Flag For The Current Subset Of The Sub-bands, And An Active / Passive Flag For At Least One Of The Sub-bands, an example method 1200 may include determining a prediction parameter for the current subset of the sub-bands, a prediction enable flag for the current subset of the sub-bands, and an active / passive flag for at least one of the sub-bands. Processing may proceed from block 1206 to block 1208.

[0109] At block 1208, “Arranging The Audio Data For The Current Subset Of Sub-bands To Be Contained Within One Block,” an example method 1200 may include arranging the audio data for the current subset of sub-bands to be contained within one block. Processing may proceed from block 1208 to block 1210.

[0110] At block 1210, “Encoding The Bitstream,” an example method 1200 may include encoding the bitstream. The bitstream may be encoded with, for example, the arranged audio data for the current subset of sub-bands, a number of subsets, the prediction parameter for the current subset of sub-bands, the active / passive flag for the other subsets of sub-bands, or a combination thereof. In some instances, the encoded bitstream enables the second device to decode the current subset of subbands, the number of subsets, the prediction parameter, and the active / passive flag, and applypreviously sent parameters for active subsets that are not for the current subset to continue the prediction for sub-bands in the active subsets.

[0111] FIG. 13 illustrates a block diagram of various example methods 1300 for decoding an input audio signal received by a first device and encoded into a bitstream, which may be performed by the framework 200 of FIG. 2, the encoding framework 300 of FIG. 3, or a combination thereof. The methods 1300 may be performed by one or more processors, which may be configured to perform methods 1300 via machine-executable instructions. The methods 1300 may be broken into various blocks or partitions, such as blocks 1302, 1304, 1306, and 1308. The various process blocks illustrated in FIG. 13 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed, added, combined, or modified without departing from the spirit of the present disclosure. For some examples, processing of the various blocks, which may be described as processes, methods, steps, blocks, operations, or functions, may commence at block 1302.

[0112] At block 1302, “Receiving the Bitstream Including Encoded Audio Data That Represents The Input Audio Signal In Sub-bands,” an example method 1300 may include receiving a bitstream, wherein the bitstream includes encoded audio data that represents the input audio signal in subbands. Processing may proceed from block 1302 to block 1304.

[0113] At block 1304, “Decoding From The Bitstream A Current Subset Of Sub-bands, A Prediction Parameter, A Prediction Enable Flag, And An Active / Passive Flag,” an example method 1300 may include decoding from the bitstream a current subset of sub-bands, a prediction parameter, a prediction enable flag, and an active / passive flag. The prediction enable flag may enable selection of one or more code books for the decoding. Processing may proceed from block 1304 to block 1306.

[0114] At block 1306, “Evaluating The Active Subsets To Determine If The Active Subsets Correspond To The Current Subset Of Sub-bands,” an example method 1300 may include evaluating active subsets to determine if the active subsets correspond to the current subset of subbands. Processing may proceed from block 1306 to block 1308.

[0115] At block 1308, “Applying Previously Sent Parameters When The Active Subsets Are Not For The Current Subset Of Sub-bands To Continue The Prediction For Sub-bands In The ActiveSubsets,” an example method 1300 may include applying previously sent parameters when the active subsets are not for the current subset of sub-bands to continue the prediction for sub-bands in the active subsets.

[0116] FIG. 14A illustrates a schematic block diagram of an example device architecture 1400 (e.g., an apparatus 1400) that may be used to implement various aspects of the present disclosure.Architecture 1400 includes but is not limited to servers and client devices, systems, and methods as described in reference to FIGS. 1-13. As shown, the architecture 1400 includes central processing unit (CPU) 1401 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 1402 or a program loaded from, for example, storage unit 1408 to random access memory (RAM) 1403. The CPU 1401 may be, for example, an electronic processor 1401. In RAM 1403, the data required when CPU 1401 performs the various processes is also stored, as required. CPU 1401, ROM 1402, and RAM 1403 are connected to one another via bus 1404. Input / output interface 1405 is also connected to bus 1404.

[0117] The following components are connected to RO interface 1405: input unit 1406, that may include a keyboard, a mouse, or the like; output unit 1407 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 1408 including a hard disk, or another suitable storage device; and communication unit 1409 including a network interface card such as a network card (e.g., wired or wireless).

[0118] In some implementations, input unit 1406 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0119] In some implementations, output unit 1407 include systems with various number of speakers. Output unit 1407 (depending on the capabilities of the hose device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0120] In some embodiments, communication unit 1409 is configured to communicate with other devices (e.g., via a network). Drive 1410 is also connected to RO interface 1405, as required. Removable medium 1411, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 1410, so that a computer program read therefrom is installed into storage unit 1408, as required. A person skilled in the art wouldunderstand that although apparatus 1400 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.

[0121] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 509, and / or installed from the removable medium 1411, as shown in FIG. 14A.

[0122] FIG. 14B illustrates a schematic block diagram of an example CPU 1401 implemented in the device architecture 1400 of FIG. 14A that may be used to implement various aspects of the present disclosure. The CPU 1401 includes an electronic processor 1420 and a memory 1421. The electronic processor 1420 is electrically and / or communicatively connected to the memory 1421 for bidirectional communication. The memory 1421 stores a sub-band coding software 1422. In some examples, memory 1421 may be located internal to the electronic processor 1420, such as for an internal cache memory or some other internally located ROM, RAM, or flash memory. In other examples, memory 1421 may be located external to the electronic processor 1420, such as in a ROM 1402, a RAM 1403, flash memory or a removable medium 1411, or another non-transitory computer readable medium that is contemplated for device architecture 1400. In some instances, the electronic processor 1420 may implement the sub-band coding software 1422 stored in the memory 1421 to perform, among other things, any of the methods 900 of FIG. 9, the methods 1000 of FIG. 10, the methods 1100 of FIG. 11, the methods 1200 of FIG. 12, and / or the methods 1300 of FIG. 13.

[0123] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units and modules discussed above can be executed by control circuitry (e.g., CPU 1401 in combination with other components of FIG. 14A), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While variousaspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as nonlimiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0124] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.

[0125] In the context of the disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non- transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine -readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a randomaccess memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD- ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0126] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alonesoftware package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.

[0127] A person skilled in the art realizes that the present invention by no means is limited to the embodiments described above. On the contrary, many modifications and variations are possible and considered within the scope of the appended claims. Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims, and which may represent systems, methods, and devices, all arranged in accordance with aspects of the present disclosure.

[0128] EEE1. A method to code an input audio signals in a bitstream from a first device to a second device, the method comprising: obtaining, with the first device, a representation of the input audio signal in sub-bands; dividing a current frame into a current subset of the sub-bands that excludes other sub-bands; determining, with the first device, a prediction parameter for the current subset of the sub-bands; determining, with the first device, an active / passive flag for at least one of the subbands; and encoding the bitstream, with the first device, the current subset of the sub-bands, a number of subsets, the prediction parameter for the current subset of sub-bands, and the active / passive flag for the other subsets of sub-bands, wherein the second device is configured to decode the current subset of sub-bands, the number of subsets, the prediction parameter, and the active / passive flag, and apply previously sent parameters for active subsets that are not for the current subset to continue the prediction for sub-bands in the active subsets.

[0129] EEE2. The method of EEE1, wherein the second device is configured to store the prediction parameter for the current subset of sub-bands for subsequent use.

[0130] EEE3. The method of any one of EEE1 to EEE2, wherein a combination of the current subset of sub-bands with the other sub-bands include all of the sub-bands.

[0131] EEE4. The method of any one of EEE1 to EEE3, wherein obtaining the representation of the input audio signal in sub-bands comprises converting the input audio signal in a time domain into sub-bands of a banded domain.

[0132] EEE5. The method of EEE4, wherein converting the input audio signal in the time domain into sub-band of the banded domain comprises processing the input audio signal in the time domain to a current frame in a complex filter bank domain (CLDFB) that includes a plurality of sub-bands.

[0133] EEE6. The method of EEE4, wherein converting the input audio signal in the time domain into sub-bands of the banded domain comprises either applying a modified discrete fourier transform (MDFT) to the input audio signal in the time domain, or applying a modified discrete cosine transform (MDCT) to the input audio signal in the time domain.

[0134] EEE7. The method of any one of EEE1 to EEE6, wherein determining the active / passive flag for at least one of the sub-bands comprises: setting the active / passive flag that corresponds to the current sub-band to indicate that prediction parameters are to be applied for a current audio channel of the input audio signal.

[0135] EEE8. The method of any one of EEE1 to EEE7, wherein determining the active / passive flag for at least one of the sub-bands comprises: setting a respective active / passive flag for each corresponding audio channel associated with the input audio signal.

[0136] EEE9. The method of EEE8, wherein setting the respective active / passive flag for each corresponding audio channel comprises: setting a bit mask to indicate the active sub-bands from the subset are enabled for prediction coding.

[0137] EEE10. The method of any one of EEE1 to EEE9, further comprising: comparing a current frame length for the input audio signal to a default frame length; and applying linear predictive coding to determine the prediction parameters when the current frame length is less than the default frame length.

[0138] EEE11. The method of EEE10, wherein the number of subsets corresponds to an integer value based on a ratio of the default frame length to the current frame length.

[0139] EEE12. The method of EEE11, wherein the number of subsets is 4 when the default frame length is 20ms and the current frame length is 5ms.

[0140] EEE13. The method of any one of EEE1 to EEE 12, wherein the bitstream is an IVAS encoded bitstream.

[0141] EEE14. A method to decode audio signals from a bitstream from a first device to a second device, the method comprising: decoding from the bitstream, with the second device, a current subset of sub-bands, a prediction parameter, and an active / passive flag; and applying previously sentparameters for active subsets that are not for the current subset to continue the prediction for subbands in the active subsets.

[0142] EEE15. The method of EEE 14, further comprising: storing the prediction parameter for the current subset of sub-bands that is decoded from the bitstream for subsequent use in prediction.

[0143] EEE16. The method of any one of EEE 14 to EEE 15, wherein an input audio signal is divided into sub-bands that include the current subset of sub-bands and other sub-bands.

[0144] EEE17. The method of EEE16, wherein a combination of the current subset of sub-bands with the other sub-bands include all of the sub-bands.

[0145] EEE18. The method of any one of EEE 14 to EEE 17, wherein input audio signal sub-bands correspond to a conversion of the input audio signal in a time domain into sub-bands of a banded domain.

[0146] EEE19. The method of EEE18, wherein the banded domain is a complex filter bank domain (CLDFB) that includes a plurality of sub-bands.

[0147] EEE20. The method of EEE 19, wherein the conversion to the complex filter bank domain corresponds to either a modified discrete Fourier transform (MDFT) or a modified discrete cosine transform (MDCT).

[0148] EEE21. The method of any one of EEE 14 to EEE20, wherein applying previously sent parameters for active subsets comprises: evaluating the active / passive flag; and retrieving previously sent prediction parameters when the active / passive flag corresponds to active.

[0149] EEE22. The method of any one of EEE14 to EEE21, wherein decoding from the bitstream further comprises decoding a number of subsets; and wherein applying previously sent parameters for the active subsets is at least partially based on the decoded number of subsets.

[0150] EEE23. The method of any one of EEE 14 to EEE22, wherein the bitstream is an IVAS encoded bitstream.

[0151] EEE24. A method to decode audio signals from a bitstream from a first device to a second device, the method comprising: decoding from the bitstream for a current frame, with the second device, a current subset of sub-bands, a prediction parameter, and an active / passive flag; identifying,with the second device, active sub-bands that are different from the current subset of sub-bands; determining, with the second device, a prediction gain for the identified active sub-bands that are different from the current subset of sub-bands based on one or more previously decoded prediction parameters; and concealing the active sub-bands that are different from the current subset of subbands when: one or more previous frames are lost; and the determined prediction gain exceeds a prediction gain threshold.

[0152] EEE25. The method of EEE24, further comprising: generating replacement audio for the identified active sub-bands based on: one or more prediction parameters received from previous frames that are not lost; and one or more decoded sub-band residual signals.

[0153] EEE26. A method to code an input audio signal in a bitstream from a first device to a second device, the method for the first device comprising: obtaining audio data that represents the input audio signal arranged in sub-bands; dividing a current frame into a current subset of the sub-bands that excludes other sub-bands; determining: a prediction parameter for the current subset of the subbands; a prediction enable flag the current subset of the sub-bands; and an active / passive flag for at least one of the sub-bands; arranging the audio data for the current subset of sub-bands to be contained within one block; and encoding the bitstream with: the arranged audio data for the current subset of sub-bands, a number of subsets, the prediction parameter for the current subset of subbands, and the active / passive flag for the other subsets of sub-bands; wherein the encoded bitstream enables the second device to decode the current subset of sub-bands, the number of subsets, the prediction parameter, and the active / passive flag, and apply previously sent parameters for active subsets that are not for the current subset to continue the prediction for sub-bands in the active subsets.

[0154] EEE27. The method of EEE26, wherein the audio data is arranged such that a first block of the encoded bitstream has a dependency between the audio data and the prediction enable flag for the current subset of sub-bands.

[0155] EEE28. The method of any one of EEE26 to EEE27, wherein arranging the audio data further comprises arranging the audio data for all audio channels corresponding to a subset of subbands into one block of the bitstream.

[0156] EEE29. The method of any one of EEE26 to EEE28, further comprising: for each subset of sub-bands, assigning the blocks of the audio data a subset identifier (subset ID); and encoding the subset identifier in the bitstream.

[0157] EEE30. The method of EEE29, further comprising: incrementing the subset identifier assigned to the blocks of the audio data for each subsequent frame.

[0158] EEE31. The method of EEE30, wherein the subset identifier has a maximum value determined by the number of subsets - 1, and wherein the subset identifier is set to a minimum value of 0 when the incrementing exceeds the maximum value.

[0159] EEE32. The method of EEE31, further comprising: decrementing the subset identifier assigned to the blocks of the audio data when the blocks of the audio data are associated with a second subset that is different from the current subset.

[0160] EEE33. The method of EEE32, wherein encoding further comprises writing the block of the audio data to the bitstream.

[0161] EEE34. The method of EEE33, wherein encoding further comprises reading the blocks of the audio data from the bitstream.

[0162] EEE35. The method of any one of EEE32 to EEE34, wherein the subset identifier is set to the maximum value when the decrementing is below the minimum value.

[0163] EEE36. A method to decode an input audio signal received by a first device and encoded into a bitstream, wherein the bitstream is transmitted from the first device to a second device, the method for the second device comprising: receiving the bitstream, wherein the bitstream includes encoded audio data that represents the input audio signal in sub-bands; decoding from the bitstream a current subset of sub-bands, a prediction parameter, a prediction enable flag, and an active / passive flag, wherein the prediction enable flag enables a selection of one or more code books for the decoding; evaluating active subsets to determine if the active subsets correspond to the current subset of sub-bands; and applying previously sent parameters when the active subsets that are not for the current subset of sub-bands to continue the prediction for sub-bands in the active subsets.

[0164] EEE37. The method of EEE36, wherein the decoded audio data is arranged such that a first block of the encoded bitstream has a dependency between the audio data and the prediction parameter for the current subset of sub-bands.

[0165] EEE38. The method of any one of EEE36 to EEE37, further comprising arranging the audio data for all audio channels corresponding to a subset of sub-bands into one block of the bitstream.

[0166] EEE39. The method of any one of EEE36 to EEE38, further comprising: evaluating the bitstream to identify a frame loss; designating each of the subsets of sub-bands with a decoding unresolved status after the frame loss is identified; and setting a decoder status to decoding failed.

[0167] EEE40. The method of EEE39, further comprising: evaluating the bitstream to identify a valid frame after the frame loss; designating the current subset of sub-bands with a decoding resolved status after the valid frame is identified; and setting the decoder status to decoding not failed.

[0168] EEE41. The method of EEE40, further comprising: designating each passive subset with a decoding resolved status after the valid frame is identified.

[0169] EEE42. The method of any one of EEE39 to EEE41, further comprising: concealing the subsets of sub-bands when the decoder status corresponds to decoding failed.

[0170] EEE43. The method of any one of EEE39 to EEE42, further comprising: enabling output of the subsets of sub-bands with a status of resolved when the decoder status corresponds to decoding not failed.

[0171] EEE44. The method of any one of EEE36 to EEE43, further comprising: continuing to decode the bitstream until a subset with a decoding unresolved status is identified; and setting a decoder status to decoding failed for all subsets when frame loss is identified.

[0172] EEE45. The method of any of EEE36 to EEE43, further comprising: continuing to decode the bitstream until a subset with a decoding unresolved status is identified; and when the subset with an unresolved status is identified, setting a decoder status to decoding failed for the current subset and each subset subsequent to the current subset.

[0173] EEE46. The method of any one of EEE36 to EEE39, further comprising: continuing to decode the bitstream until a subset with a decoding unresolved status is identified; marking validly decoded subsets as pass; and when the subset with an unresolved status is identified: setting a decoder status to decoder stopped; mark this subset and all subsequent subsets with decoding failed; and applying a concealment process to the not decoded data.

[0174] EEE47. The method of EEE46, wherein applying the concealment process comprises setting the not decoded data to a value of zero.

[0175] EEE48. The method of any one of EEE46 to EEE47, further comprising: continuing to conceal the subsets of sub-bands in undecoded subsets while the decoder status corresponds to decoding failed.

[0176] EEE49. The method of any one of EEE46 to EEE48, further comprising: setting the decoder status to decoding not failed when a valid frame is identified.

[0177] EEE50. The method of any one of EEE46 to EEE49, further comprising: enabling output of the subsets of sub-bands when the decoder status corresponds to decoding not failed.

[0178] EEE51. An apparatus comprising: an electronic processor configured to perform operations including the method of any one of EEE1 to EEE50.

[0179] EEE52. A non-transitory computer-readable storage medium recording a program of instructions that is executable by a device to perform the method of any one of EEE1 to EEE50.

[0180] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be replaced, amended, or omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.

[0181] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would beapparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.

[0182] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary in made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

[0183] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.

Claims

CLAIMSWhat is claimed is:

1. A method (900) to code an input audio signals in a bitstream from a first device (202) to a second device (208), the method comprising: obtaining (902), with the first device (202), a representation of the input audio signal in subbands; dividing (904) a current frame into a current subset of the sub-bands that excludes other subbands; determining (906), with the first device (202), a prediction parameter for the current subset of the sub-bands; determining (908), with the first device (202), an active / passive flag for at least one of the sub-bands; and encoding (910) the bitstream, with the first device (202), the current subset of the sub-bands, a number of subsets, the prediction parameter for the current subset of sub-bands, and the active / passive flag for the other subsets of sub-bands; wherein the second device (208) is configured to decode the current subset of sub-bands, the number of subsets, the prediction parameter, and the active / passive flag, and apply previously sent parameters for active subsets that are not for the current subset to continue the prediction for subbands in the active subsets.

2. The method (900) of claim 1, wherein the second device (208) is configured to store the prediction parameter for the current subset of sub-bands for subsequent use.

3. The method (900) of claim 1 or claim 2, wherein a combination of the current subset of subbands with the other sub-bands include all of the sub-bands.

4. The method (900) of any one of claims 1 to 3, wherein obtaining (902) the representation of the input audio signal in sub-bands comprises converting the input audio signal in a time domain into sub-bands of a banded domain.

5. The method (900) of claim 4, wherein converting the input audio signal in the time domain into sub-band of the banded domain comprises processing the input audio signal in the time domain to a current frame in a complex filter bank domain (CLDFB) that includes a plurality of sub-bands.

6. The method (900) of claim 4, wherein converting the input audio signal in the time domain into sub-bands of the banded domain comprises either applying a modified discrete Fourier transform (MDFT) to the input audio signal in the time domain, or applying a modified discrete cosine transform (MDCT) to the input audio signal in the time domain.

7. The method (900) of any one of claims 1 to 6, wherein determining (908) the active / passive flag for at least one of the sub-bands comprises: setting the active / passive flag that corresponds to the current sub-band to indicate that prediction parameters are to be applied for a current audio channel of the input audio signal.

8. The method (900) of any one of claims 1 to 7, wherein determining (908) the active / passive flag for at least one of the sub-bands comprises: setting a respective active / passive flag for each corresponding audio channel associated with the input audio signal.

9. The method (900) of claim 8, wherein setting the respective active / passive flag for each corresponding audio channel comprises: setting a bit mask to indicate the active sub-bands from the subset are enabled for prediction coding.

10. The method (900) of any one of claims 1 to 9, further comprising: comparing a current frame length for the input audio signal to a default frame length; and applying linear predictive coding to determine the prediction parameters when the current frame length is less than the default frame length.

11. The method (900) of claim 10, the number of subsets corresponds to an integer value based on a ratio of the default frame length to the current frame length.

12. The method (900) of claim 11, the number of subsets is 4 when the default frame length is 20ms and the current frame length is 5ms.

13. The method (900) of any one of claims 1 to 12, wherein the bitstream is an IVAS encoded bitstream.

14. A method (1000) to decode audio signals from a bitstream from a first device (202) to a second device (208), the method (1000) comprising: decoding (1002) from the bitstream, with the second device (208), a current subset of subbands, a prediction parameter, and an active / passive flag; and applying (1004) previously sent parameters for active subsets that are not for the current subset to continue the prediction for sub-bands in the active subsets.

15. The method (1000) of claim 14, further comprising: storing the prediction parameter for the current subset of sub-bands that is decoded from the bitstream for subsequent use in prediction.

16. The method (1000) of any one of claims 14 to 15, wherein an input audio signal is divided into sub-bands that include the current subset of sub-bands and other sub-bands.

17. The method (1000) of claim 16, wherein a combination of the current subset of sub-bands with the other sub-bands include all of the sub-bands.

18. The method (1000) of any one of claims 14 to 17, wherein input audio signal sub-bands correspond to a conversion of the input audio signal in a time domain into sub-bands of a banded domain.

19. The method (1000) of claim 18, wherein the banded domain is a complex filter bank domain (CLDFB) that includes a plurality of sub-bands.

20. The method (1000) of claim 19, wherein the conversion to the complex filter bank domain corresponds to either a modified discrete Fourier transform (MDFT) or a modified discrete cosine transform (MDCT).

21. The method (1000) of any one of claims 14 to 20, wherein applying previously sent parameters for active subsets comprises: evaluating the active / passive flag; and retrieving previously sent prediction parameters when the active / passive flag corresponds to active.

22. The method (1000) of any one of claims 14 to 21, wherein decoding from the bitstream further comprises decoding a number of subsets; and wherein applying previously sent parameters for the active subsets is at least partially based on the decoded number of subsets.

23. The method (1000) of any one of claims 14 to 22, wherein the bitstream is an IVAS encoded bitstream.

24. A method (1100) to decode audio signals from a bitstream from a first device (202) to a second device (208), the method (1100) comprising: decoding (1102) from the bitstream for a current frame, with the second device (208), a current subset of sub-bands, a prediction parameter, and an active / passive flag; identifying (1104), with the second device (208), active sub-bands that are different from the current subset of sub-bands; determining (1106), with the second device (208), a prediction gain for the identified active sub-bands that are different from the current subset of sub-bands based on one or more previously decoded prediction parameters; and concealing (1108) the active sub-bands that are different from the current subset of subbands when: one or more previous frames are lost; and the determined prediction gain exceeds a prediction gain threshold.

25. The method (1100) of claim 24, where further comprising: generating replacement audio for the identified active sub-bands based on: one or more prediction parameters received from previous frames that are not lost; and one or more decoded sub-band residual signals.

26. A method (1200) to code an input audio signal in a bitstream from a first device (202) to a second device (208), the method (1200) for the first device (202) comprising: obtaining (1202) audio data that represents the input audio signal arranged in sub-bands; dividing (1204) a current frame into a current subset of the sub-bands that excludes other sub-bands;determining (1206): a prediction parameter for the current subset of the sub-bands; a prediction enable flag the current subset of the sub-bands; and an active / passive flag for at least one of the sub-bands; arranging (1208) the audio data for the current subset of sub-bands to be contained within one block; and encoding (1210) the bitstream with: the arranged audio data for the current subset of sub-bands, a number of subsets, the prediction parameter for the current subset of sub-bands, and and the active / passive flag for the other subsets of sub-bands; wherein the encoded bitstream enables the second device (208) to decode the current subset of sub-bands, the number of subsets, the prediction parameter, and the active / passive flag, and apply previously sent parameters for active subsets that are not for the current subset to continue the prediction for sub-bands in the active subsets.

27. The method (1200) of claim 26, wherein the audio data is arranged such that a first block of the encoded bitstream has a dependency between the audio data and the prediction enable flag for the current subset of sub-bands.

28. The method (1200) of any one of claims 26 to 27, wherein arranging (1208) the audio data further comprises arranging the audio data for all audio channels corresponding to a subset of subbands into one block of the bitstream.

29. The method (1200) of any one of claims 26 to 28, further comprising: for each subset of sub-bands, assigning the blocks of the audio data a subset identifier (subset ID); and encoding the subset identifier in the bitstream.

30. The method (1200) of claim 29, further comprising:incrementing the subset identifier assigned to the blocks of the audio data for each subsequent frame.

31. The method (1200) of claim 30, wherein the subset identifier has a maximum value determined by the number of subsets - 1, and wherein the subset identifier is set to a minimum value of 0 when the incrementing exceeds the maximum value.

32. The method (1200) of claim 31, further comprising: decrementing the subset identifier assigned to the blocks of the audio data when the blocks of the audio data are associated with a second subset that is different from the current subset.

33. The method (1200) of claim 32, wherein encoding (1210) further comprises writing the blocks of the audio data to the bitstream.

34. The method (1200) of claim 33, wherein encoding (1210) further comprises reading the blocks of the audio data from the bitstream.

35. The method (1200) of any one of claims 32 to 34, wherein the subset identifier is set to the maximum value when the decrementing is below the minimum value.

36. A method (1300) to decode an input audio signal received by a first device (202) and encoded into a bitstream, wherein the bitstream is transmitted from the first device (202) to a second device (208), the method (1300) for the second device (208) comprising: receiving (1302) the bitstream, wherein the bitstream includes encoded audio data that represents the input audio signal in sub-bands; decoding (1304) from the bitstream a current subset of sub-bands, a prediction parameter, a prediction enable flag, and an active / passive flag, wherein the prediction enable flag enables a selection of one or more code books for the decoding; evaluating (1306) active subsets to determine if the active subsets correspond to the current subset of sub-bands; and applying (1308) previously sent parameters when the active subsets that are not for the current subset of sub-bands to continue the prediction for sub-bands in the active subsets.

37. The method (1300) of claim 36, wherein the decoded audio data is arranged such that a first block of the encoded bitstream has a dependency between the audio data and the prediction parameter for the current subset of sub-bands.

38. The method (1300) of any one of claims 36 to 37, further comprising arranging the audio data for all audio channels corresponding to a subset of sub-bands into one block of the bitstream.

39. The method (1300) of any one of claims 36 to 38, further comprising: evaluating the bitstream to identify a frame loss; designating each of the subsets of sub-bands with a decoding unresolved status after the frame loss is identified; and setting a decoder status to decoding failed.

40. The method (1300) of claim 39, further comprising: evaluating the bitstream to identify a valid frame after the frame loss; designating the current subset of sub-bands with a decoding resolved status after the valid frame is identified; and setting the decoder status to decoding not failed.

41. The method (1300) of claim 40, further comprising: designating each passive subset with a decoding resolved status after the valid frame is identified.

42. The method (1300) of any one of claims 39 to 41, further comprising: concealing the subsets of sub-bands when the decoder status corresponds to decoding failed.

43. The method (1300) of any one of claims 39 to 42, further comprising: enabling output of the subsets of sub-bands with a status of resolved when the decoder status corresponds to decoding not failed.

44. The method (1300) of any one of claims 36 to 43, further comprising: continuing to decode the bitstream until a subset with a decoding unresolved status is identified; and setting a decoder status to decoding failed for all subsets when frame loss is identified.

45. The method (1300) of any one of claims 36 to 43, further comprising: continuing to decode the bitstream until a subset with a decoding unresolved status is identified; and when the subset with an unresolved status is identified, setting a decoder status to decoding failed for the current subset and each subset subsequent to the current subset.

46. The method (1300) of any one of claims 36 to 39, further comprising: continuing to decode the bitstream until a subset with a decoding unresolved status is identified; marking validly decoded subsets as pass; and when the subset with an unresolved status is identified: setting a decoder status to decoder stopped; mark this subset and all subsequent subsets with decoding failed; and applying a concealment process to the not decoded data.

47. The method (1300) of claim 46, wherein applying the concealment process comprises setting the not decoded data to a value of zero.

48. The method (1300) of any one of claims 46 to 47, further comprising: continuing to conceal the subsets of sub-bands in undecoded subsets while the decoder status corresponds to decoding failed.

49. The method (1300) of any one of claims 46 to 48, further comprising: setting the decoder status to decoding not failed when a valid frame is identified.

50. The method (1300) of any one of claims 46 to 49, further comprising:enabling output of the subsets of sub-bands when the decoder status corresponds to decoding not failed.

51. An apparatus (1400) comprising: an electronic processor (1420) configured to perform operations including the method of any one of claims 1-50.

52. A non-transitory computer-readable storage medium (1421) recording a program of instructions (1422) that is executable by a device (1400) to perform the method of any one of claims 1-50.