Bit rate distribution in immersive voice and audio services
Patent Information
- Application Number
- JP2025114818
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-10-16
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-15
AI Technical Summary
Existing voice and audio encoder/decoder (codec) standards, such as IVAS, face challenges in efficiently allocating bitrates for immersive voice and audio services across various devices and network nodes, including mobile devices, tablets, and VR/AR equipment, due to diverse acoustic interfaces and requirements.
A method for encoding IVAS bitstreams involves downmixing input audio signals, determining bitrates and quantization levels, and combining downmix channels with spatial metadata to optimize bitrate allocation, using a bitrate allocation control table and quantization strategies tailored to specific audio formats like FoA and stereo, ensuring efficient encoding and decoding.
The method optimizes bitrate allocation between mono codecs and spatial metadata, reducing spatial metadata overhead and minimizing bit waste, resulting in an optimized IVAS bitstream for immersive audio services.
Smart Images

Figure 2025157315000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 927,772, filed October 30, 2019, and U.S. Provisional Patent Application No. 63 / 092,830, filed October 16, 2020, which are incorporated herein by reference.
[0002] Technical Field TECHNICAL FIELD This disclosure relates generally to encoding and decoding audio bitstreams. [Background technology]
[0003] Voice and audio encoder / decoder ("codec") standards development has focused in recent years on the development of codecs for immersive voice and audio services (IVAS). IVAS is expected to support a range of audio service functions, including, but not limited to, mono-to-stereo upmixing and fully immersive audio encoding, decoding, and rendering. IVAS is intended to be supported by a wide range of devices, endpoints, and network nodes, including, but not limited to, mobile and smartphones, electronic tablets, personal computers, conference phones, conference rooms, virtual reality (VR) and augmented reality (AR) devices, home theater equipment, and other suitable devices. These devices, endpoints, and network nodes may have a variety of acoustic interfaces for sound capture and rendering. Summary of the Invention [Problem to be solved by the invention]
[0004] An implementation for bitrate allocation in immersive voice and audio services is disclosed. [Means for solving the problem]
[0005] In one embodiment, a method for encoding an Immersive Speech and Audio Services (IVAS) bitstream includes the steps of: receiving, using one or more processors, an input audio signal; downmixing, using the one or more processors, the input audio signal into one or more downmix channels and spatial metadata associated with one or more channels of the input audio signal; reading, using the one or more processors, a set of one or more bitrates for the downmix channels and a set of quantization levels for the spatial metadata from a bitrate allocation control table; determining, using the one or more processors, a combination of one or more bitrates for the downmix channels; quantizing and encoding the spatial metadata using the metadata quantization levels using the one or more processors; generating a downmix bitstream for the one or more downmix channels using the combination of the one or more processors and one or more bitrates; combining the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into the IVAS bitstream using the one or more processors; and streaming or storing the IVAS bitstream for playback on an IVAS-enabled device.
[0006] In one embodiment, the input audio signal is a four-channel first-order Ambisonic (FoA) audio signal, a three-channel planar FoA signal, or a two-channel stereo audio signal.
[0007] In one embodiment, the one or more bit rates are the bit rates of one or more channels of a mono audio coder / decoder (codec) bit rate.
[0008] In one embodiment, the mono audio codec is an enhanced voice services (EVS) codec, and the downmix bitstream is an EVS bitstream.
[0009] In an embodiment, using the one or more processors to obtain one or more bitrates for the downmix channels and spatial metadata using a bitrate allocation control table further includes: identifying a row in the bitrate allocation control table using a table index including the format of the input audio signal, the bandwidth of the input audio signal, the allowed spatial coding tools, the transition mode, and the mono downmix backward compatible mode; and extracting a target bitrate, a bitrate ratio, a minimum bitrate, and a bitrate deviation step from the identified row in the bitrate allocation control table. wherein the bitrate ratio indicates a ratio at which the total bitrate is allocated among downmix audio signal channels, the minimum bitrate is a value below which the total bitrate is not allowed to fall, and the bitrate deviation increment is a target bitrate reduction increment when a first priority for the downmix signal is equal to or greater than or lower than a second priority of the spatial metadata; and determining the one or more bitrates for the downmix channels and the spatial metadata based on the target bitrate, the bitrate ratio, the minimum bitrate, and the bitrate deviation increment.
[0010] In an embodiment, when quantizing the spatial metadata for the one or more channels of the input audio signal using a set of quantization levels, quantization is performed in a quantization loop that applies an increasingly coarser quantization strategy based on the difference between a target metadata bitrate and an actual metadata bitrate.
[0011] In one embodiment, the quantization is determined based on characteristics extracted from the input audio signal and channel-banded covariance values, according to a mono codec priority and a spatial metadata priority.
[0012] In an embodiment, the input audio signal is a stereo signal and the downmix signal comprises a mid signal from the stereo signal, a representation of a residual and the spatial metadata.
[0013] In one embodiment, the spatial metadata includes prediction coefficients (PR), cross-prediction coefficients (C), and decorrelation (P) coefficients for the spatial reconstructor (SPAR) format, and prediction coefficients (PR) and decorrelation coefficients (P) for the complex advanced coupling (CACPL) format.
[0014] In one embodiment, a method for encoding an Immersive Speech and Audio Services (IVAS) bitstream includes the steps of: receiving, using one or more processors, an input audio signal; extracting, using the one or more processors, characteristics of the input audio signal; calculating, using the one or more processors, spatial metadata for channels of the input audio signal; reading, using the one or more processors, a set of one or more bitrates for downmix channels and a set of quantization levels for spatial metadata from a bitrate allocation control table; determining, using the one or more processors, a combination of the one or more bitrates for downmix channels; and calculating, using the one or more processors, spatial metadata for channels of the input audio signal. determining, using a bitrate allocation process, metadata quantization levels from the set of metadata quantization levels; using the one or more processors, quantizing and encoding spatial metadata using the metadata quantization levels; using the combination of the one or more processors and one or more bitrates, generating a downmix bitstream for the one or more downmix channels using the one or more bitrates; using the one or more processors, combining the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into the IVAS bitstream; and streaming or storing the IVAS bitstream for playback on an IVAS-enabled device.
[0015] In one embodiment, the characteristics of the input audio signal include one or more of bandwidth, speech / music classification data, and voice activity detection (VAD) data.
[0016] In one embodiment, the number of downmix channels to be encoded in the IVAS bitstream is selected based on a residual level indicator in spatial metadata.
[0017] In an embodiment, the method for encoding an Immersive Speech and Audio Services (IVAS) bitstream further includes the steps of: receiving, using one or more processors, a First Order Ambisonic (FoA) input audio signal; extracting, using one or more processors and an IVAS bitrate, characteristics of the FoA input audio signal, one of the characteristics being a bandwidth of the FoA input audio signal; generating, using the one or more processors, spatial metadata for the FoA input audio signal using the FoA signal characteristics; selecting, using the one or more processors, a number of residual channels to transmit based on a residual level indicator and a decorrelation coefficient in the spatial metadata; obtaining, using the one or more processors, a bitrate allocation control table index based on the IVAS bitrate, bandwidth, and number of downmix channels; and extracting, using the one or more processors, a bitrate allocation control table index based on the IVAS bitrate, bandwidth, and number of downmix channels. reading a spatial reconstructor (SPAR) configuration from a row of the bitrate allocation control table pointed to by a row index; using the one or more processors, determining a target metadata bitrate from the IVAS bitrates, the sum of the target EVS bitrates, and an IVAS header length; using the one or more processors, determining a maximum metadata bitrate from the IVAS bitrates, the sum of a minimum EVS bitrate, and an IVAS header length; using the one or more processors and a quantization loop, quantizing the spatial metadata in a non-temporal differential manner according to a first quantization strategy; using the one or more processors, entropy coding the quantized spatial metadata; using the one or more processors, calculating a first actual metadata bitrate; and using the one or more processors, determining whether the first actual metadata bitrate is less than or equal to a target metadata bitrate;and terminating the quantization loop in response to the first actual metadata bitrate being less than or equal to the target metadata bitrate;
[0018] In one embodiment, the method further includes: using the one or more processors, determining a first total actual EVS bitrate by adding a first amount of bits to a total EVS target bitrate, the first amount of bits being equal to the difference between the metadata target bitrate and the first actual metadata bitrate; using the one or more processors, generating an EVS bitstream using the first total actual EVS bitrate; using the one or more processors, generating an IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy coded spatial metadata; and determining a first total actual EVS bitrate by adding a first amount of bits to the total EVS target bitrate, the first actual metadata bitrate being equal to the difference between the metadata target bitrate and the first actual metadata bitrate. is greater than the target metadata bitrate: using the one or more processors, quantizing spatial metadata in a temporal differential manner according to the first quantization strategy; using the one or more processors, entropy coding the quantized spatial metadata; calculating a second actual metadata bitrate using the one or more processors; determining, using the one or more processors, whether the second actual metadata bitrate is less than or equal to the target metadata bitrate; and terminating the quantization loop in response to the second actual metadata bitrate being less than or equal to the target metadata bitrate.
[0019] In one embodiment, the method further includes: determining, using the one or more processors, a second total actual EVS bitrate by adding a second amount of bits to the total EVS target bitrate, the second amount of bits being equal to the difference between the metadata target bitrate and the second actual metadata bitrate; generating, using the one or more processors, an EVS bitstream using the second total actual EVS bitrate; generating, using the one or more processors, the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy coded spatial metadata; and in response to the second actual metadata bitrate being greater than the target metadata bitrate: quantizing, using the one or more processors, the spatial metadata in a non-temporal differential manner according to the first quantization strategy; encoding the quantized spatial metadata using the one or more processors and a base2 coder; calculating, using the one or more processors, a third actual metadata bitrate; and in response to the third actual metadata bitrate being less than or equal to the target metadata bitrate, terminating the quantization loop.
[0020] In one embodiment, the method further includes: using the one or more processors, determining a third total actual EVS bitrate by adding a third amount of bits to a total EVS target bitrate, the third amount of bits being equal to the difference between the metadata target bitrate and the third actual metadata bitrate; using the one or more processors, generating an EVS bitstream using the third total actual EVS bitrate; using the one or more processors, generating the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy coded spatial metadata; and determining that the third actual metadata bitrate is greater than the target metadata bitrate. in response: using the one or more processors, setting a fourth actual metadata bitrate to the minimum of the first, second, and third actual metadata bitrates; determining, using the one or more processors, whether the fourth actual metadata bitrate is less than or equal to the maximum metadata bitrate; in response to the fourth actual metadata bitrate being less than or equal to the maximum metadata bitrate: determining, using the one or more processors, whether the fourth actual metadata bitrate is less than or equal to the target metadata bitrate; and in response to the fourth actual metadata bitrate being less than or equal to the target metadata bitrate, terminating the quantization loop.
[0021] In one embodiment, the method further includes: using the one or more processors, determining a fourth total actual EVS bitrate by adding a fourth amount of bits to the total target EVS bitrate, the fourth amount of bits equal to the difference between the metadata target bitrate and the fourth actual metadata bitrate; using the one or more processors, generating an EVS bitstream using the fourth total actual EVS bitrate; using the one or more processors, generating the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy coded spatial metadata; and terminating the quantization loop in response to the fourth actual metadata bitrate being greater than the target metadata bitrate and less than or equal to the maximum metadata bitrate.
[0022] In one embodiment, the method further includes: using the one or more processors, determining a fifth total actual EVS bitrate by subtracting from the total target EVS bitrate an amount of bits equal to the difference between the fourth actual metadata bitrate and the target metadata bitrate; using the one or more processors, generating an EVS bitstream using the fifth actual EVS bitrate; using the one or more processors, generating the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy coded spatial metadata; and in response to the fourth actual metadata bitrate being greater than the maximum metadata bitrate: changing the first quantization strategy to a second quantization strategy and re-entering the quantization loop using the second quantization strategy, the second quantization strategy being coarser than the first quantization strategy. In one embodiment, a third quantization strategy guaranteed to provide an actual MD bitrate less than the maximum MD bitrate can be used.
[0023] In one embodiment, the SPAR configuration is defined by a downmix string, an active W flag, a complex spatial metadata flag, a spatial metadata quantization strategy, minimum, maximum, and target bit rates for one or more instances of an Enhanced Voice Services (EVS) mono coder / decoder (codec), and a time-domain decorrelator ducking flag.
[0024] In one embodiment, the actual total number of EVS bits equals the number of IVAS bits minus the number of header bits minus the actual metadata bitrate; if the actual total number of EVS bits is less than the total number of EVS target bits, bits are removed from the EVS channels in the order Z, X, Y, W; the maximum number of bits that can be removed from any channel is the number of EVS target bits for that channel minus the minimum number of EVS bits for that channel; if the actual number of EVS bits exceeds the number of EVS target bits, all additional bits are assigned to the downmix channels in the order W, Y, X, Z; and the maximum number of additional bits that can be added to any channel is the maximum number of EVS bits minus the number of EVS target bits.
[0025] In one embodiment, a method for decoding an Immersive Speech and Audio Services (IVAS) bitstream includes: receiving an IVAS bitstream using one or more processors; obtaining an IVAS bitrate from a bitlength of the IVAS bitstream using one or more processors; obtaining a bitrate allocation control table index from the IVAS bitstream using the one or more processors; parsing a metadata quantization strategy from a header of the IVAS bitstream using the one or more processors; parsing and dequantizing the quantized spatial metadata bits based on the metadata quantization strategy using the one or more processors; and decoding the IVAS bitstream using the one or more processors. setting an actual number of Enhanced Voice Services (EVS) bits equal to the remaining bit length of the bit stream; using the one or more processors and the bit rate allocation control table index, reading a table entry in the bit rate allocation control table containing an EVS target, and an EVS minimum bit rate and a maximum EVS bit rate for one or more EVS instances; using the one or more processors, obtaining an actual EVS bit rate for each downmix channel; using the one or more processors, decoding each EVS channel using the actual EVS bit rate for that channel; and using the one or more processors, upmixing the EVS channels to a primary Ambisonic (FoA) channel.
[0026] In one embodiment, a system includes: one or more processors; and a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operations of any one of the methods described above.
[0027] In one embodiment, a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of any one of the methods described above.
[0028] Other implementations disclosed herein are directed to systems, devices, and computer-readable media. Details of the disclosed implementations are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
[0029] Specific implementations disclosed herein provide one or more of the following advantages: IVAS codec bitrate is allocated between the mono codec and spatial metadata (MD) and between multiple instances of the mono codec. For a given audio frame, the IVAS codec determines the spatial audio coding mode (parametric or residual coding). The IVAS bitstream is optimized to reduce spatial MD, reduce mono codec overhead, and minimize bit waste to zero. [Brief explanation of the drawings]
[0030] The figures show a particular arrangement or order of schematic elements, such as those representing devices, units, instruction blocks, and data elements, for ease of explanation. However, those skilled in the art should understand that the particular order or arrangement of the schematic elements in the figures is not intended to imply that a particular order or sequence of processing, or separation of processes, is required. Furthermore, the inclusion of a schematic element in a figure is not intended to imply that such element is required in all embodiments, or that features represented by such element cannot be included in or combined with other elements in some implementations.
[0031] Furthermore, when connecting elements such as solid or dashed lines or arrows are used in the drawings to indicate a connection, relationship, or association between two or more other schematic elements, the absence of such connecting elements is not intended to imply that the connection, relationship, or association cannot exist. In other words, some connections, relationships, or associations between elements may not be shown in the drawings so as not to obscure the present disclosure. Furthermore, for ease of explanation, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, when a connecting element represents communication of signals, data, or instructions, those skilled in the art will understand that such element represents one or more signal paths necessary to affect the communication.
[0032] [Figure 1] 1 illustrates a use case for an IVAS codec, according to an embodiment.
[0033] [Figure 2] FIG. 1 is a block diagram of a system for encoding and decoding an IVAS bitstream, according to an embodiment.
[0034] [Figure 3] 1 is a block diagram of a spatial reconstructor (SPAR) first-order Ambisonics (FoA) coder / decoder ("codec") for encoding and decoding an IVAS bitstream in FoA format, according to an embodiment.
[0035] [Figure 4A] FIG. 1 is a block diagram of an IVAS signal chain for FoA and stereo input signals according to an embodiment.
[0036] [Figure 4B] FIG. 10 is a block diagram of an alternative IVAS signal chain for FoA and stereo input signals, according to an embodiment.
[0037] [Figure 5A]FIG. 10 is a flow diagram of a bitrate allocation process for stereo, planar FoA and FoA input signals according to an embodiment.
[0038] [Figure 5B] FIG. 1 is a flow diagram of a bitrate allocation process for a spatial reconstructor (SPAR) FoA input signal, according to an embodiment. [Figure 5C] FIG. 1 is a flow diagram of a bitrate allocation process for a spatial reconstructor (SPAR) FoA input signal, according to an embodiment.
[0039] [Figure 6] FIG. 10 is a flow diagram of a bitrate allocation process for stereo, planar FoA and FoA input signals according to an embodiment.
[0040] [Figure 7] FIG. 1 is a flow diagram of a bitrate allocation process for a SPAR FoA input signal, according to an embodiment.
[0041] [Figure 8] FIG. 1 is a block diagram of an exemplary device architecture, according to an embodiment.
[0042] The use of the same reference symbols in the various drawings indicates like elements. DETAILED DESCRIPTION OF THE INVENTION
[0043] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of various described embodiments. It will be apparent to those skilled in the art that various described embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features are described below that can each be used independently of each other or in any combination with other features.
[0044] name As used herein, the term "comprises" and variations thereof should be read as open-ended terms meaning "including, but not limited to." The term "or" should be read as "and / or" unless the context clearly dictates otherwise. The term "based on" should be read as "based at least in part on." The terms "one exemplary implementation" and "an exemplary implementation" should be read as "at least one exemplary implementation." The term "another implementation" should be read as "at least one other implementation." The terms "determined," "determining," or "determining" should be read as obtaining, receiving, calculating, computing, estimating, predicting, or deriving. Moreover, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0045] Examples of IVAS use cases FIG. 1 illustrates a use case 100 for an IVAS codec 100 according to one or more implementations. In some implementations, various devices communicate through a call server 102 configured to receive audio signals from, for example, a public switched telephone network (PSTN) or public land mobile network (PLMN) device, indicated by PSTN / other PLMN 104. Use case 100 supports legacy devices 106 that capture and render audio in mono only, including, but not limited to, devices that support Enhanced Voice Services (EVS), Multi-Rate Wideband (AMR-WB), and Adaptive Multi-Rate Narrowband (AMR-NB). Use case 100 also supports user equipment (UE) 108, 114 that capture and render stereo audio signals, or UE 110 that captures a mono signal and binaurally renders it into a multi-channel signal. Use case 100 also supports immersive and stereo signals that are captured and rendered by videoconferencing room systems 116, 118, respectively. Use case 100 also supports stereo capture and immersive rendering of stereo audio signals for a home theater system 120, a computer 112 for mono capture and immersive rendering, and immersive rendering of audio signals for virtual reality (VR) equipment 122, and immersive content ingestion 124.
[0046] Exemplary IVAS Encode / Decode System 2 is a block diagram of a system 200 for encoding and decoding an IVAS bitstream, according to one or more implementations. For encoding, the IVAS encoder includes a spatial decomposition and downmix unit 202 that receives audio data 201, including, but not limited to, mono, stereo, binaural, spatial audio (e.g., multi-channel spatial audio objects), FoA, Higher-Order Ambisonics (HoA), and any other audio data. In some implementations, the spatial decomposition and downmix unit 202 implements Complex Advanced Combining (CACPL) for decomposing / downmixing stereo / FoA audio signals and / or SPAR for decomposing / downmixing FoA audio signals. In other implementations, the spatial decomposition and downmix unit 202 implements other formats.
[0047] The output of the spatial decomposition and downmix unit 202 includes spatial metadata and 1-N downmix channels of audio, where N is the number of input channels. The spatial metadata is input to a quantization and entropy coding unit 203, which quantizes and entropy codes the spatial data. In some implementations, the quantization can include several levels of increasingly coarse quantization, such as fine, medium, coarse, and very coarse quantization strategies, and the entropy coding can include Huffman coding or arithmetic coding. The enhanced voice services (EVS) encoding unit 206 encodes the 1-N channels of audio into one or more EVS bitstreams.
[0048] In some implementations, the EVS encoding unit 206 complies with 3GPP TS 26.445 and provides a wide range of features, such as improved quality and coding efficiency for narrowband (EVS-NB) and wideband (EVS-WB) speech services, improved quality with ultra-wideband (EVS-SWB) speech, improved quality for mixed content and music in conversational applications, robustness to packet loss and delay jitter, and backward compatibility with the AMR-WB codec. In some implementations, the EVS encoding unit 206 includes a preprocessing and mode selection unit that selects between a speech coder for encoding the speech signal and a perceptual coder for encoding the audio signal at a specified bitrate based on a mode / bitrate control 207. In some implementations, the speech encoder is an improved variant of algebraic code-excited linear prediction (ACELP) extended with specialized linear prediction (LP)-based modes for different speech classes. In some implementations, the audio encoder is a low-latency / low-bitrate, efficient modified discrete cosine transform (MDCT) encoder designed to perform seamless and reliable switching between the speech and audio encoders.
[0049] In some implementations, the IVAS decoder includes a quantization and entropy decoding unit 204 configured to recover the spatial metadata and an EVS decoder(s) 208 configured to recover 1-N channel audio signals. The recovered spatial metadata and audio signals are input to a spatial synthesis / rendering unit 209, which synthesizes / renders the audio signals using the spatial metadata for playback on various audio systems 210.
[0050] Example IVAS / SPAR codec FIG. 3 is a block diagram of a FoA codec 300 for encoding and decoding FoA in the SPAR format, according to some implementations. The FoA codec 300 includes a SPAR FoA encoder 301, an EVS encoder 305, a SPAR FoA decoder 306, and an EVS decoder 307. The SPAR FoA encoder 301 converts the FoA input signal into a set of downmix channels and parameters used to regenerate the input signal in the SPAR FoA decoder 306. The downmix signal can vary from one channel to four channels, and the parameters include prediction coefficients (PR), cross-prediction coefficients (C), and decorrelation coefficients (P). Note that SPAR is a process used to reconstruct an audio signal from a downmix version of the audio signal using the PR, C, and P parameters. This will be described in more detail below.
[0051] Note that the example implementation shown in Figure 3 illustrates a nominal two-channel downmix, where a W (passive prediction) or W' (active prediction) channel is sent to the decoder 306 along with a single predicted channel Y'. In some implementations, W may be an active channel. An active W channel allows some mixing of the X, Y, and Z channels into the W channel, as follows: W'=W+f*pr y *Y+f*pr z *Z+f*pr x *X where f is a constant (e.g., 0.5) that allows some of the X, Y, and Z channels to be mixed into the W channel, and pr y , pr x , pr z are the prediction (PR) coefficients. For passive W, f=0, and there is no mixing of the X, Y, and Z channels into the W channel.
[0052] The cross-prediction coefficients (C) allow some of the parametric channels to be reconstructed from the residual channel when at least one channel is transmitted as a residual and at least one is transmitted parametrically, i.e., for two- and three-channel downmixes. For two-channel downmixes (as described in more detail below), the C coefficients allow some of the X and Z channels to be reconstructed from Y', and the remaining channels are reconstructed by a decorrelated version of the W channel, as described in more detail below. For three-channel downmixes, Y' and X' are used to reconstruct only Z.
[0053] In some implementations, the SPAR FoA encoder 301 includes a passive / active predictor unit 302, a remix unit 303, and an extraction / downmix selection unit 304. The passive / active predictor receives FoA channels in a 4-channel B-format (W, Y, Z, X) and computes downmix channels (representations of W, Y', Z', X').
[0054] The extraction / downmix selection unit 304 extracts the SPAR FoA metadata from the metadata payload section of the IVAS bitstream, as described in more detail below. The passive / active predictor unit 302 and remix unit 303 use the SPAR FoA metadata to generate remixed FoA channels (W or W' and A'), which are input to the EVS encoder 305 and encoded into an EVS bitstream. The EVS bitstream is encapsulated into the IVAS bitstream sent to the decoder 306. In this example, the Ambisonic B format channels are arranged in the AmbiX convention. However, other conventions, such as the Furse-Malham (FuMa) convention (W, X, Y, Z), can also be used.
[0055] Referring to the SPAR FoA decoder 306, the EVS bitstream is decoded by the EVS decoder 307, resulting in N_dmx (e.g., N_dmx=2) downmix channels. In some implementations, the SPAR FoA decoder 306 performs the inverse of the operations performed by the SPAR encoder 301. For example, in the example of FIG. 3, remixed FoA channels (representing W', A', B', and C') are recovered from the two downmix channels using SPAR FoA spatial metadata. The remixed SPAR FoA channels are input to the inverse mixer 311 to recover the SPAR FoA downmix channels (representing W', Y', Z', and X'). The predicted SPAR FoA channels are then input to the inverse predictor 312, which recovers the original unmixed SPAR FoA channels (W, Y, Z, and X). Note that in this two-channel example, decorrelator blocks 309A (dec1) and 309B (dec2) are used to generate a decorrelated version of the W channel using a time-domain or frequency-domain decorrelator. The downmix and decorrelated channels are used in combination with SPAR FoA metadata to fully or parametrically reconstruct the X and Z channels. C block 308 refers to the multiplication of the residual channel by a 2×1 C coefficient matrix, and the resulting two inter-predicted signals are summed to form the parametrically reconstructed channel, as shown in FIG. 3. P1 block 310A and P2 block 310B refer to the multiplication of the decorrelator outputs by columns of a 2×2 P coefficient matrix, and the resulting four outputs are summed to form the parametrically reconstructed channel, as shown in FIG. 3.
[0056] In some implementations, depending on the number of downmix channels, one of the FoA inputs is sent untouched to the SPAR FoA decoder 306 (the W channel), and one to three of the other channels (Y, Z, X) are sent to the SPAR FoA decoder 306 as residuals or fully parametrically. The PR coefficients remain the same regardless of the number of downmix channels, N, and are used to minimize the predictable energy in the residual downmix channels. The C coefficients are used to further aid in regenerating fully parameterized channels from the residuals. Thus, for the 1-channel and 4-channel downmixes where there are no residual or parameterized channels to predict from, the C coefficients are not required. The P coefficients are used to fill in the remaining energy not accounted for by the PR and C coefficients. The number of P coefficients depends on the number of downmix channels, N, in each band. In some implementations, the SPAR PR coefficients (passive W only) are calculated as follows:
[0057] Step 1. Predict all side signals (Y, Z, X) from the main W signal using equation [1].
number
number
[0058] Step 2. Remix W and the predicted (Y', Z', X') signals from acoustically most important to least important, where "remix" means to rearrange or combine signals in some way.
number
[0059] One implementation of remix is to reorder the input signal into W, Y', X', Z' under the assumption that audio cues from left and right are acoustically more important than front and back, which in turn are acoustically more important than up and down cues.
[0060] Step 3. Calculate the covariance of the 4-channel post-prediction and remix downmix as shown in equations [4] and [5].
number
number
[0061] For the example of a WABC downmix with channels 1 to 4, d and u represent the following channels as shown in Table I: [Table 1]
[0062] Of primary interest for the SPAR FoA metadata calculation are the R_dd, R_ud, and R_uu quantities. From the R_dd, R_ud, and R_uu quantities, the codec 300 determines whether it is possible to cross-predict the remaining part of the fully parametric channel from the residual channel sent to the decoder. In some implementations, the required extra C coefficients are given as:
number
[0063] Thus, the C parameter has the shape (1×2) for a 3-channel downmix and (2×1) for a 2-channel downmix.
[0064] Step 4. Calculate the residual energy of the parameterized channels that have to be reconstructed by the decorrelators 309A, 309B. The residual energy of the upmix channel Res_uu is the difference between the actual energy R_uu (post-predicted) and the regenerated cross-predicted energy Reg_uu.
number
number
number
[0065] Example IVAS signal chain (FoA or stereo input) FIG. 4A is a block diagram of an IVAS signal chain 400 for FoA and stereo input audio signals, according to one embodiment. In this exemplary configuration, the audio input to the signal chain 400 can be a four-channel FoA audio signal or a two-channel stereo audio signal. A downmix unit 401 generates a downmix audio channel (dmx_ch) and a spatial MD. The downmix channel is input to a bitrate allocation unit 402, which is configured to quantize the spatial MD and provide a mono codec bitrate for the downmix audio channel using a BR allocation control table and an IVAS bitrate, as described in more detail below. The output of the BR allocation unit 402 is input to an EVS unit 403, which encodes the downmix audio channel into an EVS bitstream. The EVS bitstream and the quantized and coded spatial MD are input to an IVAS bitstream packer 405 to form an IVAS bitstream that is sent to an IVAS decoder and / or stored for further processing or playback on one or more IVAS devices.
[0066] For a stereo input signal, the downmix unit 401 is configured to generate a mid signal (M') and a residual (Re) representation from the stereo signal and spatial MD. The spatial MD includes PR, C, and P coefficients for SPAR, and PR and P coefficients for CACPL, as described in detail below. The M' signal, Re, spatial MD, and BR allocation control table are input to a BR (bitrate) allocation unit 402, which is configured to quantize the spatial metadata and provide a mono codec bitrate for the downmix channel using the signal characteristics of the M' signal and the BR allocation control table. The M' signal, Re, and mono codec BR are input to an EVS unit 403, which encodes the M' signal and Re into an EVS bitstream. The EVS bitstream and the quantized and coded spatial MD are input to an IVAS bitstream packer 405 to form an IVAS bitstream, which is transmitted to an IVAS decoder and / or stored for subsequent processing or playback on one or more IVAS devices.
[0067] For an FoA input signal, the downmix unit 401 is configured to generate one to four FoA downmix channels W', Y', X', and Z' and spatial MD. The spatial MD includes PR, C, and P coefficients for SPAR and PR and P coefficients for CACPL, as described in detail below. The one to four FoA downmix channels (W', Y', X', Z') are input to a BR allocation unit 402, which is configured to quantize the spatial MD and provide a mono codec bitrate for the FoA downmix channels using the signal characteristics of the FoA downmix channels and a BR allocation control table. The FoA downmix channels are input to an EVS unit 403, which encodes the FoA downmix channels into an EVS bitstream. The EVS bitstream and the quantized and encoded spatial MD are input to an IVAS bitstream packer 405 to form an IVAS bitstream that is transmitted to an IVAS decoder and / or stored for subsequent processing or playback on one or more IVAS devices. The IVAS decoder can perform the inverse of the operations performed by the IVAS encoder to reconstruct the input audio signal for playback on an IVAS device.
[0068] 4B is a block diagram of an alternative IVAS signal chain 405 for FoA and stereo input audio signals, according to one embodiment. In this exemplary configuration, the audio input to the signal chain 405 can be a four-channel FoA audio signal or a two-channel stereo audio signal. In this embodiment, the pre-processor 406 extracts signal characteristics from the input audio signal, such as bandwidth (BW), speech / music classification data, and voice activity detection (VAD) data.
[0069] The spatial MD unit 407 generates spatial MD from the input audio signal using the extracted signal characteristics. The input audio signal, signal characteristics, and spatial MD are input to the BR allocation unit 408, which is configured to quantize the spatial MD and provide a mono codec bitrate for the downmix audio channel using a BR allocation control table and an IVAS bitrate, as described in more detail below.
[0070] The input audio signal, the quantized spatial MD, and the number of downmix channels (d_dmx) output by the BR allocation unit 408 are input to the downmix unit 409, which generates the downmix channels. For example, for an FoA signal, the downmix channels may include W′ and N_dmx−1 residuals (Re).
[0071] The EVS bitrate and downmix channels output by the BR allocation unit 408 are input to an EVS unit 410, which encodes the downmix channels into an EVS bitstream. The EVS bitstream and the quantized and encoded spatial MD are input to an IVAS bitstream packer 411 to form an IVAS bitstream, which is transmitted to an IVAS decoder and / or stored for further processing or playback on one or more IVAS devices. The IVAS decoder can perform the inverse of the operations performed by the IVAS encoder to reconstruct the input audio signal for playback on an IVAS device.
[0072] Exemplary Bitrate Allocation Control Strategies In one embodiment, the IVAS bitrate allocation control strategy includes two components. The first component is a BR allocation control table that provides the initial conditions for the BR allocation control process. The index into the BR allocation control table is determined by codec configuration parameters. The codec configuration parameters include the IVAS bitrate, the input format such as stereo, FoA, planar FoA or any other format, the audio bandwidth (BW), the spatial coding mode (or the number of residual channels N), and the number of audio streams. re ), which can include the priority and spatial MD for mono codecs. For stereo coding, N re =0 corresponds to full parametric (FP) mode, and N re =1 corresponds to mid-residual (MR) mode. In one embodiment, the BR allocation control table index points to target, minimum, and maximum mono codec bitrates for each of the downmix channels, and multiple quantization strategies (e.g., fine, medium-coarse, coarse) for encoding spatial MD. In another embodiment, the BR allocation control table index points to overall target and minimum bitrates for all mono codec instances, the ratio in which the available bitrate should be divided among all downmix channels, and multiple quantization strategies for encoding spatial MD. The second component of the IVAS bitrate allocation control strategy is the process of determining the spatial metadata quantization level and bitrate and the bitrate for each downmix channel using the BR allocation control table output and input audio signal characteristics, as described with reference to Figures 5A and 5B.
[0073] Bitrate Allocation Process - An Overview The main processing components of the bitrate allocation process disclosed herein include: Audio Bandwidth (BW) detection (narrowband (NB), wideband (WB), super wideband (SWB), fullband (FB), etc.). In this step, the BW of the mid or W signal is detected and the metadata is quantized accordingly. EVS then treats the IVAS BW as an upper limit and encodes the downmix channels accordingly. Input audio signal characteristics extraction (e.g. speech or music). Spatial coding mode (e.g., Full Parametric (FP), Mid Residual (MR)) or number of residual channels selection N_re, where for stereo coding, FP mode is selected when N_re=0 and MR mode is selected when N_re=1. · decisionTarget bitrate for mono codec and spatial MD priority, minimum and maximum bitrate for each downmix channel, or ratio in which the total mono codec bitrate is divided between the downmix channels.
[0074] Audio BW Detection This component detects the BW of the mid or W signal. In an embodiment, the IVAS codec uses the EVS BW detector described in EVS TS 26.445.
[0075] Extraction of input signal characteristics This component classifies each frame of the input audio signal as either speech or music. In one embodiment, the IVAS codec uses the EVS speech / music classifier as described in EVS TS 26.445.
[0076] Mono Codec vs. Spatial MD Prioritization This component determines the priority of the mono codec over spatial MD based on downmix signal characteristics. Examples of downmix signal characteristics include speech or music as determined by speech / music classifier data, and mid-side (MS) banded covariance estimates for stereo and WY, WX, and WZ banded covariance estimates for FoA. The speech / music classifier data can be used to give higher priority to the mono codec if the input audio signal is music, and the covariance estimates can be used to give higher priority to spatial MD if the input audio signal is hard panned left or right.
[0077] In one embodiment, a priority determination is calculated for each frame of the input audio signal. For a given IVAS bitrate, mid- or W-signal BW, and input configuration, the bitrate allocation starts from the target or desired bitrate for the downmix channel (e.g., the mono codec bitrate is determined based on subjective or objective evaluation) and the finest quantization strategy for metadata present in the BR Allocation Control Table. If the initial conditions do not fit within the given IVAS bitrate budget, the mono codec bitrate and / or spatial MD quantization level are iteratively reduced in the quantization loop based on their respective priorities until both fit within the IVAS bitrate budget.
[0078] Bitrate distribution between downmix channels Full Parametric and Mid Residual In FP mode, only the M' or W' channel is coded by the mono codec, and an additional parameter is coded in spatial MD to indicate the level of the residual channels or the level of decorrelation to be added by the decoder. For bit rates where both FP and MR are feasible, the IVAS BR allocation process dynamically selects the number of residual channels to be coded by the mono codec and transmitted / streamed to the decoder per frame based on spatial MD. If the level of any residual channel is higher than a threshold, that residual channel is coded by the mono codec; otherwise, the process runs in FP mode. A transition frame process is performed to reset the codec state buffer when the number of residual channels to be coded by the mono codec changes.
[0079] MR Downmix Bitrate Allocation Listening evaluations were performed using various input signals and bitrate distributions between the mid and residual channels. Based on focused listening tests, the most effective mid-to-residual bitrate ratio was 3:2. However, other ratios can be used based on application requirements. In one embodiment, the bitrate distribution uses a fixed ratio, which is further adjusted in the adjustment phase. During the iterative process of selecting the quantization strategy and BR for the downmix channels, the BR for each downmix channel is modified according to a given ratio.
[0080] In one embodiment, instead of maintaining a fixed ratio between downmix channel bit rates, the target bit rate and minimum and maximum bit rates for each downmix channel are listed separately in the BR allocation control table. These bit rates are selected based on careful subjective and objective evaluation. During the iterative process of selecting a quantization strategy and BR for the downmix channels, bits are added to or subtracted from the downmix channels based on the priorities of all downmix channels. The priorities of the downmix channels can be fixed or dynamic from frame to frame. In one embodiment, the priorities of the downmix channels are fixed.
[0081] Bitrate Allocation Process - Process Flow 5A is a flow diagram of a bitrate allocation process 500 for stereo and FoA input signals, according to one embodiment. The inputs to process 500 are IVAS bitrate, constants (e.g., bitrate allocation control table, IVAS bitrate), downmix channels, spatial MD, input format (e.g., stereo, FoA, planar FoA), and mandatory command line parameters (e.g., maximum bandwidth, encoding mode, mono downmix EVS backward compatibility mode). The output of process 500 is the EVS bitrate, metadata quantization level, and encoded metadata bits for each downmix channel. The following steps are performed as part of process 500:
[0082] Downmix audio feature extraction In step 501, the following signal characteristics are extracted from the input audio signal: bandwidth (e.g., narrowband, wideband, ultra-wideband, fullband), speech / music classification data, and voice activity detection (VAD) data. The bandwidth (BW) is the minimum of the actual bandwidth of the input audio signal and the command-line maximum bandwidth specified by the user. In one embodiment, the downmix audio signal can be in pulse code modulation (PCM) format.
[0083] Determining table indexes In step 502, process 500 extracts an IVAS bitrate allocation control table index from the IVAS bitrate allocation control table using the IVAS bitrate. In step 503, process 500 determines an input format table index based on the signal parameters extracted in step 501 (i.e., BW and speech / music classification), the input audio signal format, the IVAS bitrate allocation control table index extracted in step 502, and the EVS mono downmix backward compatibility mode. In step 504, process 500 selects a spatial coding mode (i.e., FP or MR) or the number of residual channels (i.e., N_re = 0 to 3) based on the bitrate allocation control table index, the transitional audio coding mode, and spatial MD. In step 505, process 500 determines the final accurate table index based on the above six parameters. In one embodiment, the selection of the spatial audio coding mode in step 504 is based on the residual channel level indicator in the spatial MD. The spatial audio coding mode indicates either an MR coding mode in which the representation of the mid or W channel (M' or W') in the downmixed audio signal is accompanied by one or more residual channels, or an FP coding mode in which only the representation of the mid or W channel (M' or W') is present in the downmixed audio signal. In an embodiment, if the spatial audio coding mode in the previous frame includes residual channel coding, but the current frame requires only M' or W' channel coding, the transitional audio coding mode is set to 1. Otherwise, the transitional audio coding mode is set to 0. If the number of residual channels differs between the current and previous frames, the transitional audio coding mode is set to 1.
[0084] Mono Codec and Spatial MD Priority Calculation In step 506, process 500 determines a mono codec / spatial MD priority based on the input audio signal characteristics extracted in step 1 and the mid-side or WY, WX, WZ channel banded covariance estimates. In one embodiment, there are four possible priority outcomes: mono codec high priority and spatial MD low priority, mono codec low priority and spatial MD high priority, mono codec high priority and spatial MD high priority, and mono codec low priority and spatial MD low priority.
[0085] Extract mono codec bitrate related variables from the table In step 507, the following parameters are read from the table entry pointed to by the final table index calculated in step 505: mono codec (EVS) target bitrate, bitrate ratio, EVS minimum bitrate, and EVS bitrate deviation step. The actual mono codec (EVS) bitrate may be higher or lower than the mono codec (EVS) target bitrate specified in the BR allocation control table, depending on the mono codec / spatial MD priority determined in step 506 and the spatial MD bitrates with various quantization levels. The bitrate ratio indicates the ratio by which the total EVS bitrate must be distributed among the input audio signal channels. The EVS minimum bitrate is the value below which the total EVS bitrate is not allowed to fall. The EVS bitrate deviation step is the EVS target bitrate reduction step when the EVS priority is equal to or greater than the spatial MD priority or lower than the spatial MD priority.
[0086] Calculating the best EVS bitrate and metadata quantization level based on input parameters In step 508, an optimal EVS bitrate and metadata quantization strategy are calculated based on the input parameters obtained in steps 501-503, according to the following substeps: A high bitrate and a coarse quantization strategy for the downmix channels may lead to spatial problems, while a fine quantization strategy and a low downmix audio channel bitrate may lead to mono codec encoding artifacts. As used herein, "optimal" refers to the most balanced allocation of the IVAS bitrate between the EVS bitrate and the metadata quantization level while utilizing all available bits within the IVAS bitrate budget, or at least significantly reducing bit waste.
[0087] Step 508.1: Quantize the metadata at the finest quantization level and check condition 508.a (below). If condition 508.a is true, execute step 508.b (below). Otherwise, proceed to either step 508.2 or 508.3 or 508.4 based on the priority calculated in step 503.
[0088] Step 508.2: If EVS priority is high and spatial MD priority is low, decrease spatial MD quantization level and check condition 508.a. If condition 508.a is true, execute step 508.b. Otherwise, decrease EVS target bitrate based on step 507 (EVS bitrate deviation increment) and check condition 508.a. If condition 508a is true, execute step 508.b, otherwise repeat step 508.2.
[0089] Step 508.3: If the EVS priority is low and the spatial MD priority is high, decrease the EVS target bitrate based on step 507 (EVS bitrate deviation increment) and check condition 508.a. If condition 508.a is true, execute step 508.b. Otherwise, decrease the spatial MD quantization level and check condition 508.a. If condition 508.a is true, execute step 508.b. Otherwise, repeat step 508.3.
[0090] Step 508.4: If the EVS priority is equal to the spatial MD priority, decrease the EVS target bitrate based on step 507 (EVS bitrate deviation increment) and check condition 508.a. If condition 508.a is true, execute step 508.b. Otherwise, decrease the quantization level of the spatial metadata and check condition 508.a. If condition 508.a is true, execute step 508.b. Otherwise, repeat step 508.4.
[0091] Condition 508.a above checks whether the sum of the metadata bitrate, EVS target bitrate, and overhead bits is less than or equal to the IVAS bitrate.
[0092] Step 508.b above calculates the EVS bitrate to be equal to the IVAS bitrate minus the metadata bitrate minus the overhead bits. The EVS bitrate is then allocated among the downmix audio channels according to the bitrate ratios mentioned in step 507.
[0093] If the minimum EVS target bitrate and the coarsest quantization level do not fit within the IVAS bitrate budget, the bitrate allocation process 500 runs at a lower bandwidth.
[0094] In one embodiment, the table index and metadata quantization level information are included in the overhead bits of the IVAS bitstream sent to the IVAS decoder. The IVAS decoder reads the table index and metadata quantization level from the overhead bits in the IVAS bitstream and decodes the spatial MD. This leaves the IVAS decoder processing only the EVS bits in the IVS bitstream. The EVS bits are divided among the input audio signal channels according to the ratio indicated by the table index (step 508.b). Each EVS decoder instance is then invoked with the corresponding bits, which leads to the reconstruction of the downmix audio channels.
[0095] Example IVAS Bitrate Allocation Control Table Below is an exemplary IVAS bitrate allocation control table. The following parameters shown in the table have the values shown below:
[0096] Input Format: Stereo-1, Planar FoA-2, FoA-3
[0097] BW:NB-0, WB-1, SWB-2, FB-3
[0098] Accepted spatial coding tools: FP-1, MR-2
[0099] Transition mode: 1 → Transition from MR to FP, 0 → Other
[0100] Mono downmix backward compatibility mode: 1 → Mid channel is compatible with 3GPP® EVS, 0 → Other [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5]
[0101] Also shown in Figure 5A is an IVAS bitstream. In one embodiment, the IVAS bitstream includes a fixed-length common IVAS header (CH) 509 and a variable-length common tool header (CTH) 510. In one embodiment, the bit length of the CTH section is calculated based on the number of entries corresponding to a given IVAS bitrate in the IVAS bitrate allocation control table. A relative table index (offset from the first index for that IVAS bitrate in the table) is stored in the CTH section. When operating in mono downmix backward compatibility mode, the CTH 510 is followed by an EVS payload 511, which is followed by a spatial MD payload 513. When operating in IVAS mode, the CTH 510 is followed by a spatial MD payload 512, which is followed by an EVS payload 514. In other embodiments, the order may be different.
[0102] Example Process An exemplary process for bitrate allocation can be performed by an IVAS codec, encoding / decoding, or a system including one or more processors executing instructions stored on a non-transitory computer-readable storage medium.
[0103] In one embodiment, a system for encoding audio receives an audio input and metadata, and determines one or more indices for a bitrate allocation control table based on the audio input, the metadata, and parameters of an IVAS codec used in encoding the audio input, where the parameters include an IVAS bitrate, an input format, and a mono backward compatible mode, and the one or more indices include a spatial audio coding mode and a bandwidth of the audio input.
[0104] The system performs a lookup in a bitrate allocation control table based on the IVAS bitrate, the input format, the spatial audio coding mode, and the one or more indexes, which identifies an entry in the bitrate allocation control table, the entry including a representation of an EVS target bitrate, a bitrate ratio, an EVS minimum bitrate, and an EVS bitrate deviation step.
[0105] The system provides the identified entries to a bitrate calculation process programmed to determine an audio input bitrate (e.g., a downmix channel), a metadata bitrate, and a metadata quantization level. The system provides the downmix channel bitrate and at least one of the metadata bitrate or the metadata quantization level to a downstream IVAS device.
[0106] In some implementations, the system can extract characteristics from the audio input, including an indication of whether the audio input is speech or music and a bandwidth of the audio input. Based on the characteristics, the system determines a priority between a bit rate for the downmix channels and a bit rate for the metadata. The system provides the priority to a bit rate calculation process.
[0107] In some implementations, the system extracts one or more parameters including a residual (side channel prediction error) level from the spatial MD, determines a spatial audio coding mode indicating the need for one or more residual channels in the IVAS bitstream based on the parameters, and provides the spatial audio coding mode to the bitrate calculation process.
[0108] In some implementations, the bitrate allocation control table index is stored in the Common Tool Header (CTH) of the IVAS bitstream.
[0109] A system for decoding audio is configured to receive an IVAS bitstream. The system determines an IVAS bitrate and a bitrate allocation control table index based on the IVAS bitstream. The system performs a lookup in a bitrate allocation control table based on the table index and extracts the input format, the spatial coding mode, the mono backward compatible mode, and the one or more indices, an EVS target bitrate, and a bitrate ratio. The system extracts and decodes downmix audio bits and spatial MD bits for each downmix channel. The system provides the extracted downmix signal bits and spatial MD bits to a downstream IVAS device. The downstream IVAS device may be an audio processing device or a storage device.
[0110] SPAR FoA Bitrate Allocation Process In one embodiment, the bitrate allocation process described above for a stereo input signal can also be modified and applied to SPAR FoA bitrate allocation using the SPAR FoA bitrate allocation control table shown below. Definitions of the terms contained in the table are provided below to assist the reader, followed by the SPAR FoA bitrate allocation control table. Metadata target bits (MDtar) = IVAS_bits-header_bits-evs_target_bits (EVStar) Metadata maximum bits (MDmax) = IVAS_bits - header_bits - evs_minimum_bits (EVSmin) · Metadata target bits should always be less than "MDmax". [Table 3]
[0111] Some example calculations of maximum MD bitrate (real coefficients) are shown in the table below. [Table 4]
[0112] Exemplary Metadata Quantization Loop: In one embodiment, the metadata quantization loop is implemented as described below: The metadata quantization loop includes two thresholds, MDtar and MDmax (defined above).
[0113] Step 1: For each frame of the input audio signal, the MD parameters are quantized in a non-time differential manner and coded using an arithmetic coder. The actual metadata bitrate (MDact) is calculated based on the MD coded bits. If MDact is less than MDtar, this step is considered successful, the process terminates the quantization loop, and the MDact bits are integrated into the IVAS bitstream. Any excess available bits (MDtar - MDAT) are fed to the mono codec (EVS) encoder to increase the intrinsic bitrate of the downmix audio channels. The additional bitrate allows more information to be encoded by the mono codec, resulting in a relatively lossless decoded audio output.
[0114] Step 2: If step 1 fails, a subset of the MD parameter values in the frame is quantized and then subtracted from the quantized MD parameter values in the previous frame, and the resulting quantized parameter values are coded using an arithmetic coder (i.e., temporal differential coding). MDact is calculated based on the MD coded bits. If MDact is less than MDtar, this step is considered successful, the process terminates the quantization loop, and the MDact bits are integrated into the IVAS bitstream. Any excess available bits (MDtar - MDAT) are fed to the mono codec (EVS) encoder to increase the intrinsic bitrate of the downmix audio channel.
[0115] Step 3: If step 2 fails, the bit rate of the quantized MD parameters (MDact) is calculated without entropy.
[0116] Step 4: The MDact bitrate value calculated in steps 1-3 is compared with MDmax. If the minimum of the MDact bitrates calculated in steps 1, 2, and 3 is within MDmax, this step is considered successful, the process ends the quantization loop, and the MD bitstream with the minimum MDact is integrated into the IVAS bitstream. If MDact is higher than MDtar, bits (MDact - MDtar) are removed from the mono codec (EVS) encoder.
[0117] Step 5: If step 4 fails, the parameters are quantized more coarsely and the above steps are repeated as the first fallback strategy (Fallback 1).
[0118] Step 6: If step 5 fails, the parameters are quantized with a quantization scheme that is guaranteed to fall within MDmax as a second fallback strategy (Fallback 2).
[0119] After all the above iterations, it is guaranteed that the metadata bitrate is within MDmax and the encoder produces the actual metadata bits or MDact.
[0120] Downmix Channels / EVS Bitrate Allocation (EVSbd): In one embodiment, EVS actual bits (EVSact) = IVAS_bits - header_bits - MDact. If "EVSact" is less than "EVStar", bits are removed from the EVS channels in the following order (Z, X, Y, W). The maximum number of bits that can be taken from any channel is EVStar(ch) minus EVSmin(ch). If "EVSact" is greater than "EVStar", all additional bits are allocated to downmix channels in the order W, Y, X, Z. The maximum number of additional bits that can be added to any channel is EVSmax(ch) - EVStar(ch).
[0121] SPAR decoder unpacking In one embodiment, the SPAR decoder unpacks the IVAS bitstream as follows: 1. Get the IVAS bit rate from the bit length and get the table index from the tool header (CTH) in the IVAS bitstream. 2. Parse the header / metadata bits in the IVAS bitstream 3. Parse the metadata bits and dequantize. 4. Set EVSact = remaining bit length. 5. Read the table entries related to the EVS target, min and max bitrates and repeat the "EVSbd" step in the decoder to get the actual EVS bitrate for each channel. 6. Decode the EVS channels and upmix them to the FoA channels
[0122] BR allocation process for SPAR FoA input audio signals 5B and 5C are flow diagrams of a bitrate allocation process 515 for a SPAR FoA input signal, according to one embodiment. The process 515 begins by preprocessing 517 the FoA input (W, Y, Z, X) 516 to extract signal characteristics using the IVAS bitrate, such as BW, speech / music classification data, and VAD data. The process 515 continues by generating a spatial MD (e.g., PR, C, P coefficients) 518, selecting (520) the number of residual channels to send to the IVAS decoder based on a residual level indicator in the spatial MD, and deriving (521) a BR allocation control table index based on the IVAS bitrate, BW, and number of downmix channels (N_dmx). In some embodiments, the P coefficient in the spatial MD can serve as the residual level indicator. The BR allocation control table index is sent to the IVAS bitpacker (see FIGS. 4A and 4B) to be stored and / or included in the IVAS bitstream sent to the IVAS decoder.
[0123] Process 515 continues by reading 521 the SPAR configuration from the row in the BR Allocation Control Table pointed to by the table index. As shown in Table II above, the SPAR configuration is defined by one or more features, including but not limited to: downmix string (remix), active W flag, composite spatial MD flag, spatial MD quantization strategy, EVS min / target / max bitrates, and time-domain decorrelator ducking flag.
[0124] The process 515 continues by determining MDmax, MDtar bitrates from the IVAS bitrate, EVSmin, and EVStar bitrate values (522), quantizing the spatial MD in a non-temporal differential manner using a quantization strategy, encoding the quantized spatial MD with an entropy coder (e.g., an arithmetic coder), and entering a quantization loop (523), which includes calculating MDact, as previously described. In one embodiment, the first iteration of the quantization loop uses a fine quantization strategy.
[0125] Process 515 continues by checking (524) whether MDact is less than or equal to MDtar. If MDact is less than or equal to MDtar, the MD bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, and (MDtar - MDact) bits are added to the EVStar bitrate in the following order W, Y, X, Z (532), generating N_dmx EVS bitstreams (channels). The EVS bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, as previously described. If MDact is not less than or equal to MDtar, process 515 quantizes the spatial MD using a temporal differential scheme with a fine quantization strategy, encodes the quantized spatial MD using an entropy coder, and again calculates MDact (525). If MDact is less than or equal to MDtar, the MD bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, and (MDtar - MDact) bits are added to the EVStar bitrate in the following order W, Y, X, Z (532), generating N_dmx EVS bitstreams (channels). The EVS bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, as described above. If MDact is greater than MDtar, the spatial MD is non-temporally quantized using a fine quantization strategy and entropy, and binary coded to calculate a new value for MDact (527). Note that the maximum number of bits that can be added to any EVS instance is equal to EVSmax - EVStar.
[0126] Process 515 again determines whether MDact is less than or equal to MDtar (528). If MDact is less than or equal to MDtar, the MD bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, and (MDtar - MDact) bits are added to the EVStar bitrate in the following order W, Y, X, Z (532), generating N_dmx EVS bitstreams (channels). The EVS bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, as previously described. If MDact is greater than MDtar, process 515 sets MDact as the minimum of the three MDact bitrates calculated in (523), (525), and (527) and compares MDact to MDmax (529). If MDact is greater than MDmax (530), the quantization loop (steps 523-530) is repeated using a coarser quantization strategy, as previously described.
[0127] If MDact is less than or equal to MDmax, the MD bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, and process 515 again determines whether MDact is less than or equal to MDtar (531). If MDact is less than or equal to MDtar, (MDtar - MDact) bits are added to the EVStar bitrate in the following order: W, Y, X, Z (532), generating N_dmx EVS bitstreams (channels), and the EVS bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, as previously described. If MDact is greater than MDtar, (MDtar - MDact) bits are subtracted from the EVStar bitrate in the following order: Z, X, Y, W (532), generating N_dmx EVS bitstreams (channels), and the EVS bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, as previously described. Note that the maximum number of bits that can be subtracted from any EVS instance is equal to EVStar - EVSmin.
[0128] Example Process 6 is a flow diagram of an IVAS encoding process 600 according to one embodiment. Process 600 can be implemented using the device architecture as described with reference to FIG.
[0129] The process 600 includes receiving an input audio signal (601), downmixing the input audio signal into one or more downmix channels and spatial metadata associated with the one or more downmix channels (602), reading a set of one or more bitrates for the downmix channels and a set of quantization levels for the spatial metadata from a bitrate allocation control table (603), determining one or more bitrate combinations for the downmix channels (604), determining metadata quantization levels from a set of metadata quantization levels using a bitrate allocation process (605), quantizing and encoding the spatial metadata using the metadata quantization levels (606), generating a downmix bitstream for the one or more downmix channels using the one or more bitrate combinations (607), combining the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into an IVAS bitstream (608), and streaming or storing the IVAS bitstream for playback on an IVAS-enabled device (609).
[0130] 7 is a flow diagram of an alternative IVAS encoding process 700, according to one embodiment. Process 700 can be implemented using an apparatus architecture such as that described with reference to FIG.
[0131] The process 700 includes the steps of receiving an input audio signal (701), extracting characteristics of the input audio signal (702), calculating spatial metadata for channels of the input audio signal (703), reading a set of one or more bit rates for downmix channels and a set of quantization levels for spatial metadata from a bit rate allocation control table (704), determining a combination of the one or more bit rates for the downmix channels (705), and determining a metadata quantization level from the set of metadata quantization levels using a bit rate allocation process. The method includes a step of generating a downmix bitstream for the one or more downmix channels using one or more bitrate combinations (706), a step of quantizing and encoding the spatial metadata using metadata quantization levels (707), a step of generating a downmix bitstream for the one or more downmix channels using one or more bitrate combinations (708), a step of combining the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into an IVAS bitstream (709), and a step of streaming or storing the IVAS bitstream for playback on an IVAS-enabled device (710).
[0132] Exemplary System Architecture 8 shows a block diagram of an exemplary system 800 suitable for implementing exemplary embodiments of the present disclosure. System 800 includes one or more server computers or any client devices, including, but not limited to, any of the devices shown in FIG. 1 , such as call server 102, legacy devices 106, user devices 108, 114, conference room systems 116, 118, home theater systems, VR gear 122, and immersive content consumers 124. System 800 includes any consumer device, including, but not limited to, smartphones, tablet computers, wearable computers, vehicle computers, gaming consoles, surround systems, and kiosks.
[0133] As shown, system 800 includes a central processing unit (CPU) 801 that can execute various processes according to programs stored in, for example, read-only memory (ROM) 802 or loaded, for example, from a storage unit 808 into random access memory (RAM) 803. RAM 803 also stores data needed by CPU 801 to execute the various processes, as needed. CPU 801, ROM 802, and RAM 803 are connected to one another via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.
[0134] The following components are connected to the I / O interface 805: an input unit 806, which may include a keyboard, a mouse, etc.; an output unit 807, which may include a display, such as a liquid crystal display (LCD), and one or more speakers; a storage unit 808, which may include a hard disk or another suitable storage device; and a communication unit 809, which may include a network interface card, such as a network card (e.g., wired or wireless).
[0135] In some implementations, the input unit 806 includes one or more microphones in different locations (depending on the host device) that enable the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0136] In some implementations, the output unit 807 includes a system with various numbers of speakers. As shown in Figure 1, the output unit 807 can (depending on the capabilities of the host device) render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats). The communication unit 809 is configured to communicate with other devices (e.g., via a network). If necessary, a drive 810 is also connected to the I / O interface 805. A removable medium 811, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or other suitable removable medium, is mounted on the drive 810, and if necessary, a computer program read from the removable medium is installed in the storage unit 808. Those skilled in the art will understand that although the system 800 is described as including the above-mentioned components, in actual applications, some of these components may be added, removed, and / or substituted, and all such modifications or variations are within the scope of the present disclosure.
[0137] According to exemplary embodiments of the present disclosure, the above-described processes may be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing a method. In such embodiments, the computer program may be downloaded from a network via a communication unit 809, mounted, and / or installed from a removable medium 811, as shown in FIG. 8.
[0138] In general, various exemplary embodiments of the present disclosure may be implemented in hardware or special purpose circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the units described above may be executed by a control circuit (e.g., a CPU in combination with other components of FIG. 8 ), which may then perform the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device (e.g., control circuitry). While various aspects of the exemplary embodiments of the present disclosure have been illustrated and described as block diagrams, flowcharts, or using some other pictorial representations, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special purpose circuitry or logic, general purpose hardware or controller, or other computing device, or some combination thereof.
[0139] Additionally, the various blocks illustrated in the flowcharts may be viewed as method steps and / or as operations resulting from computer program code operations and / or as multiple coupled logic circuit elements configured to perform the associated functions. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to perform the above-described method.
[0140] In the context of the present disclosure, a machine-readable medium may be any tangible medium that includes an instruction execution system, apparatus, or device, or that can store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may be non-transitory and may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. Further examples of machine-readable storage media include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0141] Computer program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. The computer program code may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing device having control circuitry, such that the program code, when executed by the processor of the computer or other programmable data processing device, causes the computer to perform the functions / acts specified in the flowcharts and / or block diagrams. The program code can run entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.
[0142] While this document contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in certain combinations and may even initially be claimed as such, one or more features from a claimed combination may, in some cases, be separated from that combination, and the claimed combination may be directed to a subcombination or variation of the subcombination. The logic flow depicted in the figures does not require the particular order shown, or sequential order, to achieve desired results. Additionally, other steps may be provided or steps may be removed from the described flow, and other components may be added to or removed from the described system. Accordingly, other implementations are within the scope of the claims.
[0143] Several aspects will be described. [Aspect 1] 1. A method for encoding an Immersive Speech and Audio Services (IVAS) bitstream, the method comprising: receiving, using one or more processors, an input audio signal; using the one or more processors, downmixing the input audio signal into one or more downmix channels and spatial metadata associated with one or more channels of the input audio signal; using the one or more processors to read a set of one or more bit rates for the downmix channels and a set of quantization levels for the spatial metadata from a bitrate allocation control table; determining, using the one or more processors, the one or more bitrate combinations for the downmix channels; using the one or more processors to determine a metadata quantization level from the set of metadata quantization levels using a bitrate allocation process; quantizing and encoding the spatial metadata using the one or more processors using the metadata quantization levels; generating a downmix bitstream for the one or more downmix channels using the combination of the one or more processors and one or more bitrates; combining, using the one or more processors, the downmix bitstream, the quantized and coded spatial metadata, and the set of quantization levels into the IVAS bitstream; streaming or storing the IVAS bitstream for playback on an IVAS-enabled device. method. [Aspect 2] 2. The method of aspect 1, wherein the input audio signal is a four-channel first-order Ambisonic (FoA) audio signal, a three-channel planar FoA signal, or a two-channel stereo audio signal. Aspect 3 3. The method of claim 1 or 2, wherein the one or more bit rates are bit rates of one or more instances of a mono audio coder / decoder (codec) bit rate. Aspect 4 3. The method of claim 1, wherein the mono audio codec is an Enhanced Voice Services (EVS) codec and the downmix bitstream is an EVS bitstream. Aspect 5 The step of obtaining, using the one or more processors, one or more bitrates for the downmix channels and the spatial metadata using a bitrate allocation control table further comprises: identifying a row in the bitrate allocation control table using a table index that includes the format of the input audio signal, the bandwidth of the input audio signal, the allowed spatial coding tools, the transition mode, and the mono downmix backward compatible mode; extracting a target bitrate, a bitrate ratio, a minimum bitrate and a bitrate deviation increment from the identified row of the bitrate allocation control table, wherein the bitrate ratio indicates a ratio at which the total bitrate is allocated among the downmix audio signal channels, the minimum bitrate is a value below which the total bitrate is not allowed to fall, and the bitrate deviation increment is a target bitrate reduction increment when a first priority for the downmix signal is equal to or greater than or lower than a second priority of the spatial metadata; determining the one or more bit rates for the downmix channel and the spatial metadata based on the target bit rate, the bit rate ratio, the minimum bit rate, and the bit rate deviation step; 3. The method of embodiment 1 or 2. Aspect 6 3. The method of claim 1 or 2, wherein quantizing the spatial metadata for the one or more channels of the input audio signal using a set of quantization levels is performed in a quantization loop that applies an increasingly coarser quantization strategy based on a difference between a target metadata bitrate and an actual metadata bitrate. Aspect 7 3. The method of claim 1 or 2, wherein the quantization is determined based on characteristics extracted from the input audio signal and channel banded covariance values and according to a mono codec priority and a spatial metadata priority. Aspect 8 3. The method of claim 1 or 2, wherein the input audio signal is a stereo signal, and the downmix signal includes a mid signal from the stereo signal, a representation of a residual, and the spatial metadata. Aspect 9 3. The method of claim 1 or 2, wherein the spatial metadata includes prediction coefficients (PR), cross-prediction coefficients (C), and decorrelation (P) coefficients for a spatial reconstructor (SPAR) format, and prediction coefficients (PR) and decorrelation coefficients (P) for a complex advanced combining (CACPL) format. Aspect 10 1. A method for encoding an Immersive Speech and Audio Services (IVAS) bitstream, the method comprising: receiving, using one or more processors, an input audio signal; extracting characteristics of the input audio signal using the one or more processors; calculating, using the one or more processors, spatial metadata for channels of the input audio signal; using the one or more processors, reading from a bitrate allocation control table one or more sets of bitrates for the downmix channels and a set of quantization levels for the spatial metadata; determining, using the one or more processors, the one or more bitrate combinations for the downmix channels; using the one or more processors to determine a metadata quantization level from the set of metadata quantization levels using a bitrate allocation process; quantizing and encoding the spatial metadata using the one or more processors using the metadata quantization levels; generating, using the combination of the one or more processors and one or more bit rates, a downmix bitstream for the one or more downmix channels using the one or more bit rates; combining, using the one or more processors, the downmix bitstream, the quantized and coded spatial metadata, and the set of quantization levels into the IVAS bitstream; streaming or storing the IVAS bitstream for playback on an IVAS-enabled device. method. Aspect 11 11. The method of claim 10, wherein the characteristics of the input audio signal include one or more of bandwidth, speech / music classification data, and voice activity detection (VAD) data. Aspect 12 12. The method of claim 10 or 11, wherein the input audio signal is a four-channel first-order Ambisonic (FoA) audio signal, a three-channel planar FoA signal, or a two-channel stereo audio signal. Aspect 13 12. The method of claim 10 or 11, wherein the one or more bit rates are bit rates of one or more instances of a mono audio coder / decoder (codec) bit rate. Aspect 14 14. The method of claim 13, wherein the mono audio codec is an Enhanced Voice Services (EVS) codec and the downmix bitstream is an EVS bitstream. Aspect 15 The step of obtaining, using the one or more processors, the set of one or more bit rates for the downmix channels and quantization levels for spatial metadata using a bit rate allocation control table further comprises: identifying a row in the bitrate allocation control table using a table index that includes the format of the input audio signal, the bandwidth of the input audio signal, the allowed spatial coding tools, the transition mode, and the mono downmix backward compatible mode; extracting a target bit rate, a bit rate ratio, a minimum bit rate and a bit rate deviation increment from the identified row of the bit rate allocation control table, wherein the bit rate ratio indicates a ratio at which the total bit rate is allocated among the input audio signal channels, the minimum bit rate is a value below which the total bit rate is not allowed to fall, and the bit rate deviation increment is a target bit rate reduction increment when a first priority for the downmix signal is equal to or greater than or lower than a second priority of the spatial metadata; determining the one or more bit rates for the downmix channel and the spatial metadata based on the target bit rate, the bit rate ratio, the minimum bit rate, and the bit rate deviation step; 12. The method of embodiment 10 or 11. Aspect 16 12. The method of claim 10 or 11, wherein when quantizing the spatial metadata for the one or more channels of the input audio signal using a set of quantization levels, the quantization is performed in a quantization loop that applies an increasingly coarser quantization strategy based on a difference between a target metadata bit rate and an actual metadata bit rate. Aspect 17 12. The method of claim 10 or 11, wherein the quantization is determined based on characteristics extracted from the input audio signal and channel banded covariance values and according to a mono codec priority and a spatial metadata priority. Aspect 18 12. The method of claim 10 or 11, wherein the input audio signal is a stereo signal, and the downmix signal includes a mid signal from the stereo signal, a representation of a residual, and the spatial metadata. Aspect 19 12. The method of claim 10 or 11, wherein the spatial metadata includes prediction coefficients (PR), cross-prediction coefficients (C), and decorrelation (P) coefficients for a spatial reconstructor (SPAR) format, and prediction coefficients (PR) and decorrelation coefficients (P) for a complex advanced combination (CACPL) format. Aspect 20 12. The method of claim 10 or 11, wherein the number of downmix channels to be encoded in the IVAS bitstream is selected based on a residual level indicator in the spatial metadata. Aspect 21 1. A method of encoding an Immersive Voice and Audio Services (IVAS) bitstream, comprising: receiving, using one or more processors, a first-order Ambisonic (FoA) input audio signal; extracting characteristics of the FoA input audio signal using the one or more processors and an IVAS bitrate, one of the characteristics being a bandwidth of the FoA input audio signal; and generating spatial metadata about the FoA input audio signal using the FoA signal characteristics using the one or more processors; using the one or more processors to select a number of residual channels to transmit based on the residual level indicator and decorrelation coefficients in the spatial metadata; using the one or more processors, obtaining a bitrate allocation control table index based on the IVAS bitrate, bandwidth, and number of downmix channels; using the one or more processors, reading a spatial reconstructor (SPAR) configuration from a row in the bitrate allocation control table pointed to by the bitrate allocation control table index; using the one or more processors to determine a target metadata bitrate from the IVAS bitrate, the sum of the target EVS bitrate, and a length of the IVAS header; using the one or more processors, determining a sum of a maximum metadata bitrate and a minimum EVS bitrate from the IVAS bitrate, and a length of the IVAS header; quantizing the spatial metadata in a non-temporal differential manner according to a first quantization strategy using the one or more processors and quantization loop; entropy encoding the quantized spatial metadata using the one or more processors; calculating a first actual metadata bitrate using the one or more processors; and using the one or more processors to determine whether the first actual metadata bitrate is less than or equal to a target metadata bitrate; and terminating the quantization loop in response to the first actual metadata bitrate being less than or equal to the target metadata bitrate. method. Aspect 22 moreover: using the one or more processors to determine a first total actual EVS bitrate by adding a first amount of bits to the total EVS target bitrate, the first amount of bits being equal to the difference between the metadata target bitrate and the first actual metadata bitrate; generating an EVS bitstream using the one or more processors using the first total actual EVS bitrate; generating, using the one or more processors, an IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized entropy coded spatial metadata; In response to the first actual metadata bitrate being greater than the target metadata bitrate: quantizing the spatial metadata in a temporally differential manner according to the first quantization strategy using the one or more processors; entropy encoding the quantized spatial metadata using the one or more processors; calculating a second actual metadata bitrate using said one or more processors; using the one or more processors, determining whether the second actual metadata bitrate is less than or equal to the target metadata bitrate; and terminating the quantization loop in response to the second actual metadata bitrate being less than or equal to the target metadata bitrate. 22. The method of embodiment 21. Aspect 23 moreover: using the one or more processors to determine a second total actual EVS bitrate by adding a second amount of bits to the total EVS target bitrate equal to the difference between the metadata target bitrate and the second actual metadata bitrate; generating an EVS bitstream using the one or more processors using the second total actual EVS bitrate; generating, using the one or more processors, the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized entropy coded spatial metadata; In response to the second actual metadata bitrate being greater than the target metadata bitrate: quantizing the spatial metadata in a non-temporal differential manner according to the first quantization strategy using the one or more processors; encoding the quantized spatial metadata using the one or more processors and a binary encoder; calculating, using said one or more processors, a third actual metadata bitrate; and responsive to the third actual metadata bitrate being less than or equal to the target metadata bitrate; terminating the quantization loop. 23. The method of embodiment 22. Aspect 24 moreover: using the one or more processors to determine a third total actual EVS bitrate by adding a third amount of bits to the total EVS target bitrate equal to the difference between the metadata target bitrate and the third actual metadata bitrate; generating, using the one or more processors, an EVS bitstream using the third total actual EVS bitrate; generating, using the one or more processors, the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized entropy coded spatial metadata; In response to the third actual metadata bitrate being greater than the target metadata bitrate: using the one or more processors, setting a fourth actual metadata bitrate to the minimum of the first, second, and third actual metadata bitrates; using the one or more processors, determining whether the fourth actual metadata bitrate is less than or equal to the maximum metadata bitrate; 4. where the actual metadata bitrate is less than or equal to the maximum metadata bitrate: using the one or more processors to determine whether the fourth actual metadata bitrate is less than or equal to the target metadata bitrate; and terminating the quantization loop in response to the fourth actual metadata bitrate being less than or equal to the target metadata bitrate. 24. The method of embodiment 23. Aspect 25 moreover: using the one or more processors to determine a fourth total actual EVS bitrate by adding a fourth amount of bits to the total EVS target bitrate equal to the difference between the metadata target bitrate and the fourth actual metadata bitrate; generating, using the one or more processors, an EVS bitstream using the fourth total actual EVS bitrate; generating, using the one or more processors, the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized entropy coded spatial metadata; and terminating the quantization loop in response to the fourth actual metadata bitrate being greater than the target metadata bitrate and less than or equal to the maximum target metadata bitrate. 25. The method of embodiment 24. Aspect 26 moreover: using the one or more processors to determine a fifth total actual EVS bitrate by subtracting from the total EVS target bitrate an amount of bits equal to the difference between the fourth actual metadata bitrate and the target metadata bitrate; generating, using the one or more processors, an EVS bitstream using the fifth actual EVS bitrate; generating, using the one or more processors, the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized entropy coded spatial metadata; and in response to the fourth actual metadata bitrate being greater than the maximum target metadata bitrate: changing the first quantization strategy to a second quantization strategy and re-entering the quantization loop using the second quantization strategy, the second quantization strategy being coarser than the first quantization strategy. The method according to embodiment 25. Aspect 27 27. The method of any one of aspects 21 to 26, wherein the SPAR configuration is defined by a downmix string, an active W flag, a complex spatial metadata flag, a spatial metadata quantization strategy, minimum, maximum, and target bit rates for one or more instances of an Enhanced Voice Services (EVS) mono coder / decoder (codec), and a time-domain decorrelator ducking flag. Aspect 28 27. The method of any one of aspects 21 to 26, wherein the actual total number of EVS bits equals the number of IVAS bits minus the number of header bits minus the actual metadata bitrate; if the actual total number of EVS bits is less than the total number of EVS target bits, bits are removed from the EVS channels in the order Z, X, Y, W; the maximum number of bits that can be removed from any channel is the number of EVS target bits for that channel minus the minimum number of EVS bits for that channel; and if the actual total number of EVS bits exceeds the number of EVS target bits, all additional bits are allocated to downmix channels in the order W, Y, X, Z; and the maximum number of additional bits that can be added to any channel is the maximum number of EVS bits minus the number of EVS target bits. Aspect 29 How to decode an Immersive Voice and Audio Services (IVAS) bitstream: receiving, using one or more processors, an IVAS bitstream; using one or more processors to obtain an IVAS bitrate from the bit length of the IVAS bitstream; using the one or more processors to obtain a bitrate allocation control table index from the IVAS bitstream; parsing a metadata quantization strategy from a header of the IVAS bitstream using the one or more processors; parsing and dequantizing the quantized spatial metadata bits based on the metadata quantization strategy using the one or more processors; using the one or more processors to set an actual number of Enhanced Voice Services (EVS) bits equal to a remaining bit length of the IVAS bitstream; using the one or more processors and the bitrate allocation control table index, reading a table entry in the bitrate allocation control table that includes an EVS target, an EVS minimum bitrate, and a maximum EVS bitrate for one or more EVS instances; obtaining an actual EVS bitrate for each downmix channel using the one or more processors; using the one or more processors, decoding each EVS channel using the actual EVS bit rate for that channel; and upmixing the EVS channels into a primary Ambisonic (FoA) channel using the one or more processors. method. Aspect 30 one or more processors; a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operations of the method of any one of aspects 1 to 29; system. Aspect 31 A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations of the method of any one of aspects 1 to 29.
Claims
1. 1. A method for encoding an immersive speech and audio service (IVAS) bitstream, the method comprising: receiving, using one or more processors, an input audio signal; extracting characteristics of the input audio signal using the one or more processors; calculating, using the one or more processors, spatial metadata for channels of the input audio signal; using the one or more processors to obtain a set of one or more bitrates for the downmix channels and a set of metadata quantization levels for the spatial metadata from a bitrate allocation control table; determining, using the one or more processors, the one or more bitrate combinations for the downmix channels; using the one or more processors to determine a metadata quantization level from the set of metadata quantization levels using a bitrate allocation process; quantizing and encoding the spatial metadata using the one or more processors using the metadata quantization levels; generating a downmix bitstream for the one or more downmix channels using the one or more bitrates using the combination of the one or more processors and one or more bitrates; and combining, using the one or more processors, the downmix bitstream, the quantized and coded spatial metadata, and metadata quantization level information into the IVAS bitstream. method.
2. The method of claim 1 , wherein the characteristics of the input audio signal include one or more of bandwidth, speech / music classification data, and voice activity detection (VAD) data.
3. 3. The method of claim 1, wherein the input audio signal is a four-channel first-order Ambisonics (FoA) audio signal, a three-channel planar FoA signal, or a two-channel stereo audio signal.
4. 3. The method of claim 1, wherein the one or more bit rates are bit rates of one or more instances of a bit rate of a mono audio coder / decoder (codec).
5. The method of claim 1 , wherein the mono audio codec is an Enhanced Voice Services (EVS) codec and the downmix bitstream is an EVS bitstream.
6. The step of using the one or more processors to obtain the set of one or more bitrates for the downmix channels and the set of metadata quantization levels for spatial metadata using a bitrate allocation control table further includes: identifying a row in the bitrate allocation control table using a table index that includes one or more of the format of the input audio signal, the bandwidth of the input audio signal, the allowed spatial coding tools, a transition mode, and a mono downmix backward compatible mode; extracting from the identified row of the bitrate allocation control table one or more of a target bitrate, a bitrate ratio, a minimum bitrate and a bitrate deviation increment, wherein the bitrate ratio indicates the ratio at which the total bitrate is allocated among the input audio signal channels, the minimum bitrate is a value below which the total bitrate is not allowed to fall, and the bitrate deviation increment is a target bitrate reduction increment when a first priority for the downmix signal is greater than or equal to or lower than a second priority of the spatial metadata, determining the combination of the one or more bitrates for the downmix channels and the metadata quantization level for the spatial metadata based on the target bitrate, the bitrate ratio, the minimum bitrate, and the bitrate deviation step; 3. The method according to claim 1 or 2.
7. 3. The method of claim 1, wherein quantizing and encoding the spatial metadata for the one or more channels of the input audio signal using the set of metadata quantization levels is performed in a quantization loop that applies an increasingly coarser quantization strategy based on a difference between a target metadata bitrate and an actual metadata bitrate.
8. The method of claim 1 or 2, wherein the quantization is determined based on characteristics extracted from the input audio signal and channel banding covariance values, according to a mono codec priority and a spatial metadata priority.
9. The method of claim 1 or 2, wherein the input audio signal is a stereo signal and the downmix signal comprises a mid signal from the stereo signal, a representation of a residual and the spatial metadata.
10. 3. The method of claim 1, wherein the spatial metadata includes prediction coefficients (PR), cross-prediction coefficients (C), and decorrelation coefficients (P) for a spatial reconstructor (SPAR) format, and prediction coefficients (PR) or decorrelation coefficients (P) for a complex advanced combination (CACPL) format.
11. The method of claim 1 or 2, wherein the number of downmix channels to be coded into the IVAS bitstream is selected based on a residual level indicator in the spatial metadata.
12. one or more processors; a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operations of the method of any one of claims 1 to 11; system.
13. A non-transitory computer readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of the method of any one of claims 1 to 11.
14. The step of using the one or more processors to obtain the set of one or more bitrates and the set of metadata quantization levels for spatial metadata for the downmix channels using a bitrate allocation control table further comprises: identifying a row in the bitrate allocation control table using a table index including one or more of the format of the input audio signal, the bandwidth of the input audio signal, and an IVAS bitrate; extracting, for each of the downmix channels from the identified row of the bitrate allocation control table, one or more of a target bitrate, a minimum bitrate, and a maximum bitrate, wherein the minimum bitrate and the maximum bitrate define a bitrate range for the bitrate of that downmix channel, and the target bitrate is a preferred bitrate for that downmix channel; calculating a total downmix bitrate by subtracting a metadata bitrate and an IVAS header bitrate from the IVAS bitrate; determining the combination of the one or more bitrates for the downmix channels based on one or more of the target bitrate, the minimum bitrate, the maximum bitrate, the total downmix bitrate and priorities assigned to the downmix channels; The method of claim 1 , comprising:
15. 2. The method of claim 1 , wherein the bitrate allocation process reduces at least one of the target bitrates or at least one of the metadata quantization levels of the spatial metadata based, at least in part, on a bitstream budget for the IVAS bitstream.
16. A computer program product for causing one or more processors to perform the method of claim 1.