Bitrate Allocation in Submerged Audio and Audio Services
The method optimizes bitrate allocation in IVAS codecs by downmixing and quantizing spatial metadata, addressing inefficiencies in immersive audio services across diverse devices, enhancing encoding and decoding efficiency.
Patent Information
- Application Number
- JP2022524623
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-16
- Filing Date
- 2020-10-28
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2040-10-28
AI Technical Summary
Existing voice and audio encoder/decoder (codec) standards, such as IVAS, face challenges in optimizing bitrate allocation for immersive audio services across various devices, leading to inefficiencies in encoding and decoding processes.
A method for encoding immersive audio services (IVAS) bitstreams involves downmixing input audio signals, determining bitrates and quantization levels using a bitrate allocation control table, and combining downmix channels with quantized spatial metadata to generate an optimized IVAS bitstream for playback on compatible devices.
The solution ensures efficient distribution of bitrates between mono codecs and spatial metadata, reducing spatial metadata and mono codec overhead, and minimizing bit waste, thereby optimizing the IVAS codec performance.
Smart Images

Figure 0007712050000018 
Figure 0007712050000019 
Figure 0007712050000020
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 62 / 927,772, filed Oct. 30, 2019, and U.S. Provisional Patent Application No. 63 / 092,830, filed Oct. 16, 2020, which are hereby incorporated by reference herein.
[0002] Technical Field The present disclosure generally relates to the encoding and decoding of audio bitstreams.
Background Art
[0003] Voice and audio encoder / decoder ( "codec") standard development has focused in recent years on the development of codecs for immersive voice and audio services (IVAS). IVAS is expected to support a range of audio service functions including, but not limited to, upmixing from monaural to stereo and fully immersive audio encoding, decoding, and rendering. IVAS is intended to be supported by a wide range of devices, endpoints, and network nodes including, but not limited to, cellular phones and smartphones, electronic tablets, personal computers, conference phones, conference rooms, virtual reality (VR) and augmented reality (AR) devices, home theater devices, and other suitable devices. These devices, endpoints, and network nodes can have various acoustic interfaces for sound capture and rendering.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Implementations for bitrate allocation in immersive voice and audio services are disclosed.
Means for Solving the Problems
[0005] In one embodiment, a method of encoding an immersive audio and audio service (IVAS) bitstream, the method comprising: receiving an input audio signal using one or more processors; using the one or more processors to downmix the input audio signal into one or more downmix channels and spatial metadata associated with one or more channels of the input audio signal; using the one or more processors to read from a bitrate allocation control table a set of one or more bitrates for the downmix channels and a set of quantization levels for the spatial metadata; using the one or more processors to determine a combination of the one or more bitrates for the downmix channels; using the one or more processors to determine a metadata quantization level from the set of metadata quantization levels using a bitrate allocation process; using the one or more processors to quantize and encode the spatial metadata using the metadata quantization level; using the one or more processors and the combination of the one or more bitrates to generate a downmix bitstream for the one or more downmix channels; using the one or more processors to combine the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into the IVAS bitstream; and streaming or storing the IVAS bitstream for playback on an IVAS-compatible device.
[0006] In one embodiment, the input audio signal is a 4-channel first-order ambisonic (FoA) audio signal, a 3-channel planar FoA signal, or a 2-channel stereo audio signal.
[0007] In one embodiment, the one or more bitrates are the bitrates of one or more channels of a mono audio coder / decoder (codec).
[0008] In one embodiment, the mono audio codec is an enhanced voice services (EVS) codec and the downmix bitstream is an EVS bitstream.
[0009] In one embodiment, the step of obtaining one or more bitrates for the downmix channels and the spatial metadata using the bitrate allocation control table using the one or more processors further comprises: identifying a row in the bitrate allocation control table using a table index including the format of the input audio signal, the bandwidth of the input audio signal, allowed spatial coding tools, a transition mode, and a mono downmix backward compatibility mode; extracting a target bitrate, a bitrate ratio, a minimum bitrate, and a bitrate deviation quantization from the identified row in the bitrate allocation control table, wherein the bitrate ratio indicates a ratio at which the total bitrate is distributed among the downmix audio signal channels, the minimum bitrate is a value below which the total bitrate is not allowed to fall, and the bitrate deviation quantization is a target bitrate reduction quantization when a first priority for the downmix signal is greater than or less than a second priority for the spatial metadata; and determining the one or more bitrates for the downmix channels and the spatial metadata based on the target bitrate, the bitrate ratio, the minimum bitrate, and the bitrate deviation quantization.
[0010] In one embodiment, when quantizing the spatial metadata for the one or more channels of the input audio signal using a set of quantization levels, quantization is performed in a quantization loop that applies a quantization strategy that gradually coarsens based on the difference between a target metadata bitrate and an actual metadata bitrate.
[0011] In one embodiment, the quantization is determined according to a mono codec priority and a spatial metadata priority based on characteristics extracted from the input audio signal and channel - banded covariance values.
[0012] In one embodiment, the input audio signal is a stereo signal, and the downmix signal includes a mid signal from the stereo signal, a residual representation, and the spatial metadata.
[0013] In one embodiment, the spatial metadata includes prediction coefficients (PR), cross - prediction coefficients (C), and decorrelation (P) coefficients for a spatial reconstructor (SPAR) format, and prediction coefficients ( PR ) and decorrelation coefficients ( P ) for a complex advanced coupling (CACPL) format.
[0014] In one embodiment, a method for encoding an immersive audio and audio service (IVAS) bitstream, the method comprising: receiving an input audio signal using one or more processors; extracting characteristics of the input audio signal using the one or more processors; calculating spatial metadata for channels of the input audio signal using the one or more processors; reading, from a bitrate allocation control table, a set of one or more bitrates for downmix channels and a set of quantization levels for spatial metadata using the one or more processors; determining a combination of the one or more bitrates for the downmix channels using the one or more processors; determining a metadata quantization level from the set of metadata quantization levels using a bitrate allocation process with the one or more processors; quantizing and encoding the spatial metadata using the metadata quantization level with the one or more processors; generating a downmix bitstream for the one or more downmix channels using the one or more bitrates with the one or more processors and the combination of the one or more bitrates; combining the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into the IVAS bitstream using the one or more processors; and streaming or storing the IVAS bitstream for playback on an IVAS-compatible device.
[0015] In one embodiment, the characteristics of the input audio signal include one or more of bandwidth, speech / music classification data, and voice activity detection (VAD) data.
[0016] In one embodiment, the number of downmix channels encoded in the IVAS bitstream is selected based on a residual level indicator in the spatial metadata.
[0017] In one embodiment, a method for encoding an immersive audio and audio service (IVAS) bitstream further comprises: receiving, using one or more processors, a first-order Ambisonics (FoA) input audio signal; extracting, using one or more processors and an IVAS bitrate, characteristics of the FoA input audio signal, wherein one of the characteristics is a bandwidth of the FoA input audio signal; generating, using the one or more processors and using the FoA signal characteristics, spatial metadata for the FoA input audio signal; selecting, using the one or more processors, a number of residual channels to transmit based on a residual level indicator and a decorrelation coefficient in the spatial metadata; obtaining, using the one or more processors, a bitrate allocation control table index based on the IVAS bitrate, the bandwidth, and a number of downmix channels; reading, using the one or more processors, a spatial reconstructor (SPAR) configuration from a row of the bitrate allocation control table pointed to by the bitrate allocation control table index; determining, using the one or more processors, a target metadata bitrate from the IVAS bitrate, a sum of the target EVS bitrate, and a length of an IVAS header; determining, using the one or more processors, a maximum metadata bitrate from the IVAS bitrate, a sum of a minimum EVS bitrate, and the length of the IVAS header; quantizing, using the one or more processors and a quantization loop, the spatial metadata in a non-time-differential manner according to a first quantization strategy; entropy encoding, using the one or more processors, the quantized spatial metadata; calculating, using the one or more processors, a first actual metadata bitrate; and determining, using the one or more processors, whether the first actual metadata bitrate is less than or equal to a target metadata bitrate;ending the quantization loop in response to the first actual metadata bitrate being less than or equal to the target metadata bitrate;
[0018] In certain embodiments, the method further comprises: using the one or more processors to determine a first total actual EVS bitrate by adding a first bit amount equal to a difference between the metadata target bitrate and the first actual metadata bitrate to a total EVS target bitrate; using the one or more processors to generate an EVS bitstream using the first total actual EVS bitrate; using the one or more processors to generate an IVAS bitstream comprising the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy-coded spatial metadata; in response to the first actual metadata bitrate being greater than the target metadata bitrate: using the one or more processors to quantize spatial metadata in a temporal differential manner according to the first quantization strategy; using the one or more processors to entropy-code the quantized spatial metadata; using the one or more processors to calculate a second actual metadata bitrate; using the one or more processors to determine whether the second actual metadata bitrate is less than or equal to the target metadata bitrate; and ending the quantization loop in response to the second actual metadata bitrate being less than or equal to the target metadata bitrate.
[0019] In one embodiment, the method further comprises: using the one or more processors to determine a second total actual EVS bitrate by adding a second bit amount equal to a difference between the metadata target bitrate and the second actual metadata bitrate to the total EVS target bitrate; using the one or more processors to generate an EVS bitstream using the second total actual EVS bitrate; using the one or more processors to generate the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy-encoded spatial metadata; in response to the second actual metadata bitrate being greater than the target metadata bitrate: using the one or more processors to quantize the spatial metadata in a non-temporal difference manner according to the first quantization strategy; using the one or more processors and a base2 coder to encode the quantized spatial metadata; using the one or more processors to calculate a third actual metadata bitrate; and in response to the third actual metadata bitrate being less than or equal to the target metadata bitrate, ending the quantization loop.
[0020] In one embodiment, the method further comprises: using the one or more processors to determine a third total actual EVS bitrate by adding a third bit amount equal to the difference between the metadata target bitrate and the third actual metadata bitrate to the total EVS target bitrate; using the one or more processors to generate an EVS bitstream using the third total actual EVS bitrate; using the one or more processors to generate the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy-coded spatial metadata; in response to the third actual metadata bitrate being greater than the target metadata bitrate: using the one or more processors to set a fourth actual metadata bitrate to the minimum value of the first, second, and third actual metadata bitrates; using the one or more processors to determine whether the fourth actual metadata bitrate is less than or equal to the maximum metadata bitrate; in response to the fourth actual metadata bitrate being less than or equal to the maximum metadata bitrate: using the one or more processors to determine whether the fourth actual metadata bitrate is less than or equal to the target metadata bitrate; and in response to the fourth actual metadata bitrate being less than or equal to the target metadata bitrate, ending the quantization loop.
[0021] In one embodiment, the method further comprises: using the one or more processors to determine a fourth total actual EVS bitrate by adding a fourth bit amount equal to a difference between the metadata target bitrate and the fourth actual metadata bitrate to the total target EVS bitrate; using the one or more processors to generate an EVS bitstream using the fourth total actual EVS bitrate; using the one or more processors to generate the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy-coded spatial metadata; and ending the quantization loop in response to the fourth actual metadata bitrate being greater than the target metadata bitrate and less than or equal to the maximum metadata bitrate.
[0022] In one embodiment, the method further comprises: using the one or more processors to determine a fifth total actual EVS bitrate by subtracting from the total target EVS bitrate an amount of bits equal to the difference between the fourth actual metadata bitrate and the target metadata bitrate; using the one or more processors to generate an EVS bitstream using the fifth actual EVS bitrate; using the one or more processors to generate the IVAS bitstream comprising the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy-coded spatial metadata; and in response to the fourth actual metadata bitrate being greater than the maximum metadata bitrate: changing the first quantization strategy to a second quantization strategy and entering the quantization loop again using the second quantization strategy, the second quantization strategy being coarser than the first quantization strategy. In one embodiment, a third quantization strategy can be used that ensures an actual MD bitrate less than the maximum MD bitrate.
[0023] In one embodiment, the SPAR configuration is defined by a downmix string, an active W flag, a complex spatial metadata flag, a spatial metadata quantization strategy, minimum, maximum, and target bitrates for one or more instances of an extended voice service (EVS) mono coder / decoder (codec), and a time domain decorrelator ducking flag.
[0024] In one embodiment, the actual total number of EVS bits is equal to the number of IVAS bits minus the number of header bits minus the actual metadata bit rate. When the actual total number of EVS bits is less than the total number of EVS target bits, bits are removed from the EVS channel in the order of Z, X, Y, W. The maximum number of bits that can be removed from any channel is the number of EVS target bits for that channel minus the minimum number of EVS bits for that channel. When the actual number of EVS bits is greater than the number of EVS target bits, all additional bits are assigned to the downmix channels in the order of W, Y, X, Z. The maximum number of additional bits that can be added to any channel is the maximum number of EVS bits minus the number of EVS target bits.
[0025] In one embodiment, a method for decoding an immersive voice and audio service (IVAS) bitstream includes: receiving the IVAS bitstream using one or more processors; obtaining an IVAS bitrate from the bit length of the IVAS bitstream using one or more processors; obtaining a bitrate allocation control table index from the IVAS bitstream using one or more processors; parsing a metadata quantization strategy from the header of the IVAS bitstream using one or more processors; parsing and dequantizing the quantized spatial metadata bits based on the metadata quantization strategy using one or more processors; setting an actual number of extended voice service (EVS) bits equal to the remaining bit length of the IVAS bitstream using one or more processors; reading a table entry of the bitrate allocation control table including an EVS target, and an EVS minimum bitrate and a maximum EVS bitrate for one or more EVS instances using the one or more processors and the bitrate allocation control table index; obtaining an actual EVS bitrate for each downmix channel using one or more processors; decoding each EVS channel using the actual EVS bitrate for that channel using one or more processors; and upmixing the EVS channels to first-order ambisonic (FoA) channels using one or more processors.
[0026] In one embodiment, a system includes: one or more processors; and a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any one of the operations of the above method.
[0027] In one embodiment, when executed by one or more processors, a non-transitory computer-readable medium storing instructions that cause the one or more processors to perform any one of the operations of the methods described above.
[0028] Other implementations disclosed herein are directed to systems, apparatuses, and computer-readable media. Details of the disclosed implementations are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the specification, drawings, and claims.
[0029] The specific implementations disclosed herein provide one or more of the following advantages. The IVAS codec bitrate is distributed between the mono codec and the spatial metadata (MD), and between multiple instances of the mono codec. For a given audio frame, the IVAS codec determines the spatial audio coding mode (parametric or residual coding). The IVAS bitstream is optimized to reduce spatial MD, reduce mono codec overhead, and minimize bit waste to zero.
Brief Description of the Drawings
[0030] In the drawings, for ease of explanation, a particular arrangement or order of schematic elements, such as those representing devices, units, instruction blocks, and data elements, is shown. However, it should be understood by those skilled in the art that a particular order or arrangement of schematic elements in the drawings is not intended to imply a particular order or sequence of processing, or that separation of processes is required. Further, the inclusion of a particular schematic element in a drawing is not intended to imply that such an element is required in all embodiments, or that features represented by such an element cannot be included in or combined with other elements in some implementations.
[0031] Furthermore, in the drawings, when connection elements such as solid lines, dashed lines, or arrows are used to indicate a connection, relationship, or association between two or more other schematic elements, the absence of such a connection element is not intended to imply that a connection, relationship, or association cannot exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the present disclosure. Further, for ease of explanation, a single connection element is used to represent multiple connections, relationships, or associations between elements. For example, when a connection element represents the communication of signals, data, or instructions, one of ordinary skill in the art should understand that such an element represents one or more signal paths necessary to affect the communication.
[0032]
Figure 1
[0033]
Figure 2
[0034]
Figure 3
[0035]
Figure 4A
[0036]
Figure 4B
[0037]
Figure 5A
[0038]
Figure 5B
Figure 5C
[0039]
Figure 6
[0040]
Figure 7
[0041]
Figure 8
[0042] The same reference symbols used in the various drawings denote similar elements.
DETAILED DESCRIPTION OF THE INVENTION
[0043] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. It will be apparent to those skilled in the art that the various described embodiments may be practiced without these specific details. On the other hand, well-known methods, procedures, components, and circuits are not described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features that can be used either independently of each other or in any combination with other features are described below.
[0044] Title As used in this specification, the term "comprising" and variations thereof should be read as an open-ended term meaning "including, but not limited to." The term "or" should be read as "and / or" unless the context clearly indicates otherwise. The term "based on" should be read as "at least in part based on." The terms "one exemplary implementation" and "an exemplary implementation" should be read as "at least one exemplary implementation." The term "another implementation" should be read as "at least one other implementation." The terms "determined," "determining," or "determines" should be read as obtaining, receiving, calculating, computing, estimating, predicting, or deriving. Further, in the following description and claims, unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0045] Examples of IVAS Usage Scenarios Figure 1 shows a use case 100 for an IVAS codec 100 according to one or more implementations. In some implementations, various devices communicate through a call server 102 configured to receive audio signals from, for example, a public switched telephone network (PSTN) or a public land mobile network device (PLMN) indicated by PSTN / other PLMN 104. The use case 100 supports legacy devices 106 that render and capture audio in mono only, including but not limited to devices that support extended voice service (EVS), multi-rate wideband (AMR-WB), and adaptive multi-rate narrowband (AMR-NB). The use case 100 also supports user equipment (UE) 108, 114 that capture and render stereo audio signals, or UE 110 that captures a mono signal and performs binaural rendering to a multi-channel signal. The use case 100 also supports immersive and stereo signals captured and rendered by video conference room systems 116, 118, respectively. The use case 100 also supports stereo capture and immersive rendering of stereo audio signals for a home theater system 120, a computer 112 for mono capture and immersive rendering, immersive rendering of audio signals for a virtual reality (VR) headset 122, and immersive content consumption 124.
[0046] Exemplary IVAS encoding / decoding system Figure 2 is a block diagram of a system 200 for encoding and decoding an IVAS bitstream according to one or more implementations. For encoding, the IVAS encoder includes a spatial decomposition and downmix unit 202 that receives audio data 201 including, but not limited to, mono signals, stereo signals, binaural signals, spatial audio signals (e.g., multi-channel spatial audio objects), FoA, higher-order ambisonics (HoA), and any other audio data. In some implementations, the spatial decomposition and downmix unit 202 implements a complex advanced coupling (CACPL) for decomposing / downmixing stereo / FoA audio signals and / or a SPAR for decomposing / downmixing FoA audio signals. In other implementations, the spatial decomposition and downmix unit 202 implements other formats.
[0047] The output of the spatial decomposition and downmix unit 202 includes spatial metadata and one to N downmix channels of audio, where N is the number of input channels. The spatial metadata is input to a quantization entropy coding unit 203 that quantizes and entropy-codes the spatial data. In some implementations, the quantization can include several levels of increasingly coarse quantization, such as fine, medium, coarse, and very coarse quantization strategies, and the entropy coding can include Huffman coding or arithmetic coding. An extended voice service (EVS) encoding unit 206 encodes the one to N channels of audio into one or more EVS bitstreams.
[0048] In some implementations, the EVS encoding unit 206 complies with 3GPP (registered trademark) TS 26.445 and provides a wide range of features, such as improved quality and coding efficiency for narrowband (EVS-NB) and wideband (EVS-WB) speech services, improved quality using ultra-wideband (EVS-SWB) speech, improved quality for mixed content and music in conversational applications, robustness against packet loss and delay jitter, and backward compatibility with the AMR-WB codec. In some implementations, the EVS encoding unit 206 includes a preprocessing and mode selection unit that selects, based on mode / bitrate control 207, between a speech encoder for encoding speech signals at a specified bitrate and a perceptual encoder for encoding audio signals. In some implementations, the speech encoder is an improved variant of algebraic code-excited linear prediction (ACELP) extended with specialized linear prediction (LP)-based modes for different speech classes. In some implementations, the audio encoder is a modified discrete cosine transform (MDCT) encoder with improved efficiency at low latency / low bitrate and is designed to perform seamless and reliable switching between the speech encoder and the audio encoder.
[0049] In some implementations, the IVAS decoder includes a quantization and entropy decoding unit 204 configured to recover spatial metadata and one or more EVS decoders 208 configured to recover 1-N channel audio signals. The recovered spatial metadata and audio signals are input to a spatial synthesis / rendering unit 209. The unit synthesizes / renders the audio signals using the spatial metadata for playback on various audio systems 210.
[0050] Exemplary IVAS / SPAR codec Figure 3 is a block diagram of a FoA codec 300 for encoding and decoding FoA in the SPAR format according to some implementations. The FoA codec 300 includes a SPAR FoA encoder 301, an EVS encoder 305, a SPAR FoA decoder 306, and an EVS decoder 307. The SPAR FoA encoder 301 converts a FoA input signal into a set of downmix channels and parameters used to regenerate the input signal in the SPAR FoA decoder 306. The downmix signal can vary from 1 channel to 4 channels, and the parameters include prediction coefficients (PR), cross-prediction coefficients (C), and decorrelation coefficients (P). Note that SPAR is a process used to reconstruct an audio signal from a downmixed version of the audio signal using the PR, C, and P parameters. This will be described in more detail later.
[0051] Note that the exemplary implementation shown in Figure 3 shows a nominal 2-channel downmix, where the W (passive prediction) or W' (active prediction) channel is sent to the decoder 306 along with a single predicted channel Y'. In some implementations, W can be an active channel. The active W channel allows some mixing of the X, Y, Z channels into the W channel as follows: W' = W + f * pr y * Y + f * pr z * Z + f * pr x * X where f is a constant (e.g., 0.5) that allows some mixing of the X, Y, Z channels into the W channel, and pr y 、pr x 、pr z are prediction (PR) coefficients. For passive W, f = 0 and there is no mixing of the X, Y, Z channels into the W channel.
[0052] The cross prediction coefficient (C) allows for the reconstruction of some of the parametric channels from the residual channels when at least one channel is sent as residual and at least one is sent parametrically, i.e., for 2 and 3 channel downmixes. For the 2 channel downmix (described in more detail below), the C coefficient allows for some of the X and Z channels to be reconstructed from Y', and the remaining channels are reconstructed by a decorrelated version of the W channel, as described in more detail below. For the 3 channel downmix, Y' and X' are used to reconstruct only Z.
[0053] In some implementations, the SPAR FoA encoder 301 includes a passive / active predictor unit 302, a remix unit 303, and an extract / downmix selection unit 304. The passive / active predictor receives the FoA channels in 4 channel B format (W, Y, Z, X) and calculates the downmix channels (representation of W, Y', Z', X').
[0054] The extract / downmix selection unit 304 extracts the SPAR FoA metadata from the metadata payload section of the IVAS bitstream, as described in more detail below. The passive / active predictor unit 302 and the remix unit 303 use the SPAR FoA metadata to generate the remixed FoA channels (W or W' and A'), which are input to the EVS encoder 305 and encoded into an EVS bitstream. The EVS bitstream is encapsulated into the IVAS bitstream that is sent to the decoder 306. In this example, the channels of the ambisonic B format are arranged according to the AmbiX convention. However, other conventions such as the Furse-Malham (FuMa) convention (W, X, Y, Z) can also be used.
[0055] Referring to the SPAR FoA decoder 306, the EVS bitstream is decoded by the EVS decoder 307, resulting in N_dmx (e.g., N_dmx = 2) downmix channels. In some implementations, the SPAR FoA decoder 306 performs the inverse of the operations performed by the SPAR encoder 301. For example, in the example of FIG. 3, the remixed FoA channels (represented by W', A', B', C') are recovered from two downmix channels using the SPAR FoA spatial metadata. The remixed SPAR FoA channels are input to an inverse mixer 311 to recover the SPAR FoA downmix channels (represented by W', Y', Z', X'). The predicted SPAR FoA channels are then input to an inverse predictor 312 to recover the original unmixed SPAR FoA channels (W, Y, Z, X). Note that in this two-channel example, decorrelator blocks 309A (dec1) and 309B (dec2) are used to generate a decorrelated version of the W channel using time-domain or frequency-domain decorrelators. The downmix channels and the decorrelated channels are used in combination with the SPAR FoA metadata to fully or parametrically reconstruct the X channel and the Z channel. Block C 308 refers to the multiplication of the residual channel by a 2×1 C coefficient matrix, and the two resulting cross-prediction signals are added to be parametrically reconstructed channels as shown in FIG. 3. Blocks P1 310A and P2 310B refer to multiplying the columns of a 2×2 P coefficient matrix by the decorrelator output, and the four resulting outputs are added to be parametrically reconstructed channels as shown in FIG. 3.
[0056] In some implementations, depending on the number of downmix channels, one of the FoA inputs is sent to the SPAR FoA decoder 306 in a hands-off state (W channel), and one to three of the other channels (Y, Z, X) are sent to the SPAR FoA decoder 306 as residuals or completely parametrically. The PR coefficients remain the same regardless of the number N of downmix channels and are used to minimize the predictable energy in the residual downmix channels. The C coefficients are used to further assist in regenerating the completely parameterized channels from the residuals. Thus, in the case of the one-channel and four-channel downmixes where there are no residual channels or parameterized channels on which to base the prediction, the C coefficients are not required. The P coefficients are used to fill in the remaining energy not accounted for by the PR and C coefficients. The number of P coefficients depends on the number N of downmix channels in each band. In some implementations, the SPAR PR coefficients (passive W only) are calculated as follows.
[0057] Step 1. Predict all side signals (Y, Z, X) from the main W signal using Equation [1].
Equation
Equation
[0058] Step 2. Remix the predicted (Y', Z', X') signals with W in order from the most acoustically important to the least important. Here, "remix" means changing the order of the signals or changing the combination based on some methodology.
Number
[0059] One implementation of the remix is to rearrange the input signals into W, Y', X', Z' under the assumption that the audio cues from left and right are more acoustically important than those from front and back, and the cues from front and back are more acoustically important than those from top and bottom.
[0060] Step 3. Calculate the covariance of the 4-channel post-prediction and the downmix of the remix as shown in equations [4] and [5].
Number
Number
[0061] For an example of a WABC downmix with 1 to 4 channels, d and u represent the following channels shown in Table I:
Table 1
[0062] Of primary interest for the calculation of SPAR FoA metadata are the R_dd, R_ud, and R_uu quantities. From the R_dd, R_ud, and R_uu quantities, codec 300 determines whether it is possible to predict the remainder of the parametric channel from the residual channel sent to the decoder with perfect tolerance. In some implementations, the extra C coefficients required are given as follows:
Number
[0063] Thus, the C parameters have the form (1×2) for 3-channel downmix and (2×1) for 2-channel downmix.
[0064] Step 4. Calculate the remaining energy of the parameterized channels that must be reconstructed by decorrelators 309A, 309B. The residual energy of the upmix channel Res_uu is the difference between the actual energy R_uu (post-prediction) and the regenerated cross-prediction energy Reg_uu.
Number
Number
Number
[0065] Exemplary IVAS signal chain (FoA or stereo input) Figure 4A is a block diagram of an IVAS signal chain 400 for FoA and stereo input audio signals according to an embodiment. In this exemplary configuration, the audio input to the signal chain 400 can be a 4-channel FoA audio signal or a 2-channel stereo audio signal. The downmix unit 401 generates a downmix audio channel (dmx_ch) and a spatial MD. The downmix channel is input to the bitrate allocation unit 402. The bitrate allocation unit 402 is configured to quantize the spatial MD and provide a mono codec bitrate for the downmix audio channel using a BR allocation control table and an IVAS bitrate, as will be described in detail below. The output of the BR allocation unit 402 is input to the EVS unit 403 that encodes the downmix audio channel into an EVS bitstream. The EVS bitstream and the quantized and encoded spatial MD are input to the IVAS bitstream packer 405 to form an IVAS bitstream, which is transmitted to an IVAS decoder and / or stored for subsequent processing or playback on one or more IVAS devices.
[0066] For the stereo input signal, the downmix unit 401 is configured to generate the representation of the mid signal (M'), the residual (Re) from the stereo signal and the spatial MD. The spatial MD includes the PR, C and P coefficients for SPAR, and the PR and P coefficients for CACPL, details of which will be described later. The M' signal, Re, spatial MD and BR distribution control table are input to the BR (bit rate) distribution unit 402. The BR distribution unit is configured to quantize the spatial metadata and provide the mono codec bit rate for the downmix channel using the signal characteristics of the M' signal and the BR distribution control table. The M' signal, Re and mono codec BR are input to the EVS unit 403, and the EVS unit 403 encodes the M' signal and Re into an EVS bit stream. The EVS bit stream and the quantized and encoded spatial MD are input to the IVAS bit stream packer 405 to form an IVAS bit stream, which is transmitted to the IVAS decoder and / or stored for subsequent processing or playback on one or more IVAS devices.
[0067] For the FoA input signal, the downmix unit 401 is configured to generate one to four FoA downmix channels W', Y', X' and Z' and the spatial MD. The spatial MD includes the PR, C and P coefficients for SPAR, and the PR and P coefficients for CACPL, the details of which will be described later. The one to four FoA downmix channels (W', Y', X', Z') are input to the BR distribution unit 402. The BR distribution unit quantizes the spatial MD and is configured to provide a mono-coder bitrate for the FoA downmix channels using the signal characteristics of the FoA downmix channels and a BR distribution control table. The FoA downmix channels are input to the EVS unit 403. The EVS unit 403 encodes the FoA downmix channels into an EVS bitstream. The EVS bitstream and the quantized and encoded spatial MD are input to the IVAS bitstream packer 405 to form an IVAS bitstream, which is transmitted to the IVAS decoder and / or stored for subsequent processing or playback on one or more IVAS devices. The IVAS decoder can perform the reverse of the operations performed by the IVAS encoder to reconstruct an input audio signal for playback on the IVAS device.
[0068] Figure 4B is a block diagram of an alternative IVAS signal chain 405 for FoA and stereo input audio signals according to an embodiment. In this exemplary configuration, the audio input to the signal chain 405 can be a 4-channel FoA audio signal or a 2-channel stereo audio signal. In this embodiment, the preprocessor 406 extracts signal characteristics from the input audio signal such as the bandwidth (BW), speech / music classification data, voice activity detection (VAD) data, etc.
[0069] The spatial MD unit 407 generates spatial MD from the input audio signal using the extracted signal characteristics. The input audio signal, the signal characteristics, and the spatial MD are input to the BR distribution unit 408. The BR distribution unit quantizes the spatial MD and is configured to provide a mono codec bit rate for the downmix audio channels using a BR distribution control table and the IVAS bit rate, which will be described in detail below.
[0070] The input audio signal, the quantized spatial MD, and the number of downmix channels (d_dmx) output by the BR distribution unit 408 are input to the downmix unit 409, which generates the downmix channels. For example, for the FoA signal, the downmix channels can include W' and N_dmx - 1 residuals (Re).
[0071] The EVS bit rate and the downmix channels output by the BR distribution unit 408 are input to the EVS unit 410, which encodes the downmix channels into an EVS bit stream. The EVS bit stream and the quantized and encoded spatial MD are input to the IVAS bit stream packer 411 to form an IVAS bit stream, which is transmitted to the IVAS decoder and / or stored for subsequent processing or playback on one or more IVAS devices. The IVAS decoder can perform the reverse of the operations performed by the IVAS encoder to reconstruct the input audio signal for playback on the IVAS device.
[0072] Exemplary bit rate distribution control strategies In one embodiment, the IVAS bitrate allocation control strategy includes two components. The first component is a BR allocation control table that provides the initial conditions for the BR allocation control process. The index to the BR allocation control table is determined by codec configuration parameters. The codec configuration parameters may include the IVAS bitrate, stereo, FoA, planar FoA or any other input format such as format, audio bandwidth (BW), spatial coding mode (or number of residual channels N re ), mono-codec priority, and spatial MD. For stereo coding, N re = 0 corresponds to the full parametric (FP) mode, and N re = 1 corresponds to the midresidual (MR) mode. In one embodiment, the index to the BR allocation control table points to the target, minimum, and maximum mono-codec bitrates for each of the downmix channels, and multiple quantization strategies (e.g., fine, moderately coarse, coarse) for encoding the spatial MD. In another embodiment, the index to the BR allocation control table points to the total target and minimum bitrates for all mono-codec instances, the ratio by which the available bitrate needs to be divided among all downmix channels, and multiple quantization strategies for encoding the spatial MD. The second component of the IVAS bitrate allocation control strategy is a process that uses the BR allocation control table output and the input audio signal characteristics to determine the spatial metadata quantization level and bitrate, and the bitrate for each downmix channel, as described with reference to FIGS. 5A and 5B.
[0073] Bitrate Allocation Process - Overview The main processing components of the bitrate allocation process disclosed herein include the following: · Audio bandwidth (BW) detection (narrowband (NB), wideband (WB), super-wideband (SWB), fullband (FB), etc.). In this step, the BW of the mid or W signal is detected, and the metadata is quantized accordingly. Then, EVS treats the IVAS BW as an upper limit and encodes the downmix channel accordingly. · Input audio signal characteristic extraction (e.g., speech or music). · Spatial coding mode (e.g., full parametric (FP), midresidual (MR)) or the number of residual channels selected N_re. Here, for stereo coding, the FP mode is selected when N_re = 0, and the MR mode is selected when N_re = 1. · Mono codec and spatial MD priority decisionTarget [decision target] bitrate, the minimum and maximum bitrates for each downmix channel, or the ratio by which the total mono codec bitrate is divided among the downmix channels.
[0074] Audio BW Detection This component detects the BW of the mid or W signal. In an embodiment, the IVAS codec uses the EVS BW detector described in EVS TS 26.445.
[0075] Extraction of Input Signal Characteristics This component classifies each frame of the input audio signal as speech or music. In one embodiment, the IVAS codec uses the EVS speech / music classifier as described in EVS TS 26.445.
[0076] Mono Codec vs. Spatial MD Priority Decision This component determines the priority of the mono codec for spatial MD based on the downmix signal characteristics. Examples of downmix signal characteristics include speech or music determined by the speech / music classifier data, and mid-side (M-S) banded covariance estimates for stereo, and W-Y, W-X, W-Z banded covariance estimates for FoA. The speech / music classifier data can be used to give a higher priority to the mono codec when the input audio signal is music, and the covariance estimates can be used to give a higher priority to spatial MD when the input audio signal is hard panned left or right.
[0077] In one embodiment, the priority determination is calculated for each frame of the input audio signal. For a given IVAS bitrate, intermediate or W signal BW and input configuration, the bitrate allocation starts from the target or desired bitrate for the downmix channels (e.g., the mono codec bitrate is determined based on subjective or objective evaluation) and the finest quantization strategy for the metadata that exist in the BR allocation control table. If the initial conditions do not fit within the given IVAS bitrate budget, the mono codec bitrate or the quantization level of spatial MD or both are sequentially and iteratively reduced in the quantization loop based on their respective priorities until both fit within the IVAS bitrate budget.
[0078] Bitrate Allocation between Downmix Channels Full Parametric and Mid Residual In the FP mode, only the M' or W' channel is encoded by the mono - codec, additional parameters are encoded in the spatial MD, indicating the level of the residual channel or the level of decorrelation to be added by the decoder. For bitrates achievable with both FP and MR, the IVAS BR allocation process dynamically selects, on a per - frame basis, based on the spatial MD, the number of residual channels to be encoded by the mono - codec and transmitted / streamed to the decoder. If the level of any residual channel is higher than a threshold, that residual channel is encoded by the mono - codec; otherwise, the process is executed in the FP mode. Transition frame processing is executed to reset the codec state buffer when the number of residual channels to be encoded by the mono - codec changes.
[0079] MR Downmix Bitrate Allocation Listening evaluations were performed using various input signals and bitrate allocations between the mid - channel and the residual channels. Based on a focused listening test, the most effective mid - to - residual bitrate ratio is 3:2. However, other ratios can be used based on the requirements of the application. In one embodiment, the bitrate allocation uses a fixed ratio, which is further adjusted during an adjustment phase. During the sequential iterative process of selecting a quantization strategy and BR for the downmix channels, the BR for each downmix channel is modified according to a given ratio.
[0080] In one embodiment, instead of maintaining a fixed ratio between the downmix channel bitrates, the target bitrate as well as the minimum and maximum bitrates for each downmix channel are listed separately in a BR allocation control table. These bitrates are selected based on careful subjective and objective evaluations. During the sequential iterative process of selecting the quantization strategy and BR for the downmix channels, bits are added to or removed from the downmix channels based on the priority of all the downmix channels. The priority of the downmix channels can be fixed or dynamic for each frame. In one embodiment, the priority of the downmix channels is fixed.
[0081] Bitrate Allocation Process - Process Flow FIG. 5A is a flowchart of a bitrate allocation process 500 for stereo and FoA input signals according to an embodiment. The inputs to process 500 are the IVAS bitrate, constants (e.g., bitrate allocation control table, IVAS bitrate), downmix channels, spatial MD, input format (e.g., stereo, FoA, planar FoA), and forced command line parameters (e.g., maximum bandwidth, encoding mode, mono downmix EVS backward compatibility mode). The outputs of process 500 are the EVS bitrates for each downmix channel, the metadata quantization level, and the encoded metadata bits. The following steps are executed as part of process 500.
[0082] Extraction of Downmix Audio Features In step 501, the following signal characteristics are extracted from the input audio signal: bandwidth (e.g., narrowband, wideband, ultra-wideband, full-band) and speech / music classification data, voice activity detection (VAD) data. The bandwidth (BW) is the minimum of the actual bandwidth of the input audio signal and the maximum bandwidth of the command line specified by the user. In one embodiment, the downmix audio signal can be in pulse code modulation (PCM) format.
[0083] Determination of table index In step 502, process 500 extracts an IVAS bitrate distribution control table index from the IVAS bitrate distribution control table using the IVAS bitrate. In step 503, process 500 determines an input format table index based on the signal parameters extracted in step 501 (i.e., BW and speech / music classification), the input audio signal format, the IVAS bitrate distribution control table index extracted in step 502, and the EVS mono-downmix backward compatibility mode. In step 504, process 500 selects a spatial encoding mode (i.e., FP or MR) or the number of residual channels (i.e., N_re = 0 to 3) based on the bitrate distribution control table index, the transition audio encoding mode, and the spatial MD. In step 505, process 500 determines the final accurate table index based on the above six parameters. In one embodiment, the selection of the spatial audio encoding mode in step 504 is based on the residual channel level indicator in the spatial MD. The spatial audio encoding mode indicates either an MR encoding mode in which the representation of the mid or W channel (M' or W') in the downmixed audio signal is accompanied by one or more residual channels, or an FP encoding mode in which only the representation of the mid or W channel (M' or W') exists in the downmixed audio signal. In one embodiment, if the spatial audio encoding mode in the previous frame includes residual channel encoding, but the current frame requires only M' or W' channel encoding, the transition audio encoding mode is set to 1. Otherwise, the transition audio encoding mode is set to 0. If the number of residual channels differs between the current frame and the previous frame, the transition audio encoding mode is set to 1.
[0084] Calculation of Mono-Coder and Spatial MD Priority In step 506, process 500 determines the mono codec / space MD priority based on the input audio signal characteristics extracted in step 1 and the mid-side or W-Y, W-X, W-Z channel banded covariance estimates. In certain embodiments, there are four possible priority results: high mono codec priority and low space MD priority, low mono codec priority and high space MD priority, high mono codec priority and high space MD priority, low mono codec priority and low space MD priority.
[0085] Extract the mono codec bitrate related variable from the table In step 507, the following parameters are read from the table entry pointed to by the final table index calculated in step 505: mono codec (EVS) target bitrate, bitrate ratio, EVS minimum bitrate, and EVS bitrate deviation step. The actual mono codec (EVS) bitrate may be higher or lower than the mono codec (EVS) target bitrate specified in the BR allocation control table, depending on the mono codec / space MD priority determined in step 506 and the space MD bitrates with various quantization levels. The bitrate ratio indicates the ratio by which the total EVS bitrate must be distributed among the input audio signal channels. The EVS minimum bitrate is the value below which the total EVS bitrate is not allowed to fall. The EVS bitrate deviation step is the EVS target bitrate reduction step when the EVS priority is greater than or equal to, or lower than, the priority of space MD.
[0086] Calculation of the best EVS bitrate and metadata quantization level based on the input parameters In step 508, an optimal EVS bitrate and metadata quantization strategy are calculated according to the following sub-steps based on the input parameters obtained in steps 501 to 503. A high bitrate and a coarse quantization strategy for the downmix channel may lead to spatial problems, while a fine quantization strategy and a low downmix audio channel bitrate may lead to mono codec encoding artifacts. As used herein, "optimal" is the most balanced distribution of the IVAS bitrate between the EVS bitrate and the metadata quantization level that utilizes all available bits within the IVAS bitrate budget or at least significantly reduces bit waste.
[0087] Step 508.1: Quantize the metadata at the finest quantization level and check condition 508.a (shown below). If condition 508.a is true, execute step 508.b (shown below). Otherwise, proceed to either step 508.2 or 508.3 or 508.4 based on the priority calculated in step 503.
[0088] Step 508.2: If the EVS priority is high and the spatial MD priority is low, lower the quantization level of the spatial MD and check condition 508.a. If condition 508.a is true, execute step 508.b. Otherwise, reduce the EVS target bitrate based on step 507 (EVS bitrate deviation slope) and check condition 508a. If condition 508a is true, execute step 508.b, otherwise repeat step 508.2.
[0089] Step 508.3: When the EVS priority is low and the spatial MD priority is high, reduce the EVS target bitrate based on Step 507 (EVS bitrate deviation slope), and check Condition 508.a. If Condition 508.a is true, execute Step 508.b. Otherwise, lower the quantization level of the spatial MD and check Condition 508.a. If Condition 508.a is true, execute Step 508.b. Otherwise, repeat Step 508.3.
[0090] Step 508.4: When the EVS priority is equal to the spatial MD priority, reduce the EVS target bitrate based on Step 507 (EVS bitrate deviation slope), and check Condition 508.a. If Condition 508.a is true, execute Step 508.b. Otherwise, lower the quantization level of the spatial metadata and check Condition 508.a. If Condition 508.a is true, execute Step 508.b. Otherwise, repeat Step 5.4.
[0091] The above Condition 508.a checks whether the sum of the metadata bitrate, the EVS target bitrate, and the overhead bits is less than or equal to the IVAS bitrate.
[0092] The above Step 508.b calculates the EVS bitrate to be equal to the IVAS bitrate minus the metadata bitrate minus the overhead bits. Then, according to the bitrate ratio described in Step 507, the EVS bitrate is distributed among the downmix audio channels.
[0093] If the minimum EVS target bitrate and the coarsest quantization level do not fit within the IVAS bitrate budget, the bitrate allocation process 500 is executed at a lower bandwidth.
[0094] In one embodiment, the table index and metadata quantization level information are included in the overhead bits of the IVAS bitstream that is sent to the IVAS decoder. The IVAS decoder reads the table index and metadata quantization level from the overhead bits in the IVAS bitstream and decodes the spatial MD. As a result, all that the IVAS decoder processes are the EVS bits in the IVS bitstream. The EVS bits are split among the input audio signal channels according to the ratio indicated by the table index (step 508.b). Next, each EVS decoder instance is invoked with the corresponding bits, which leads to the reconstruction of the downmix audio channels.
[0095] Exemplary IVAS Bitrate Allocation Control Table The following is an exemplary IVAS bitrate allocation control table. The following parameters shown in the table have the values shown below:
[0096] Input Format: Stereo-1, Planar FoA-2, FoA-3
[0097] BW: NB-0, WB-1, SWB-2, FB-3
[0098] Allowed Spatial Encoding Tools: FP-1, MR-2
[0099] Transition Mode: 1 → Transition from MR to FP, 0 → Others
[0100] Mono Downmix Backward Compatibility Mode: 1 → When the mid channel is compatible with 3GPP (registered trademark) EVS, 0 → Others [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5]
[0101] Figure 5A also shows the IVAS bitstream. In one embodiment, the IVAS bitstream includes a fixed-length common IVAS header (CH) 509 and a variable-length common tool header (CTH) 510. In one embodiment, the bit length of the CTH section is calculated based on the number of entries corresponding to a given IVAS bit rate in the IVAS bit rate allocation control table. The relative table index (the offset from the first index for that IVAS bit rate within the table) is stored in the CTH section. When operating in the mono-downmix backward compatibility mode, the CTH 510 is followed by an EVS payload 511, which is followed by a spatial MD payload 513. When operating in the IVAS mode, the CTH 510 is followed by a spatial MD payload 512, which is followed by an EVS payload 514. In other embodiments, the order may be different.
[0102] Exemplary Process An exemplary process for bit rate allocation can be executed by a system that includes an IVAS codec, an encoder / decoder, or one or more processors that execute instructions stored on a non-transitory computer-readable storage medium.
[0103] In one embodiment, a system for encoding audio receives an audio input and metadata. The system determines one or more indices of a bitrate allocation control table based on the audio input, the metadata, and parameters of an IVAS codec used when encoding the audio input. The parameters include an IVAS bitrate, an input format, and a mono backward compatibility mode, and the one or more indices include a spatial audio coding mode and a bandwidth of the audio input.
[0104] The system performs a lookup in the bitrate allocation control table based on the IVAS bitrate, the input format, the spatial audio coding mode, and the one or more indices. The lookup identifies an entry in the bitrate allocation control table, and the entry includes representations of an EVS target bitrate, a bitrate ratio, an EVS minimum bitrate, and an EVS bitrate deviation quantization.
[0105] The system provides the identified entry to a bitrate calculation process programmed to determine a bitrate of the audio input (e.g., a downmix channel), a bitrate of the metadata, and a quantization level of the metadata. The system provides the bitrate of the downmix channel and at least one of the bitrate of the metadata or the quantization level of the metadata to a downstream IVAS device.
[0106] In some implementations, the system can extract characteristics from the audio input, and the characteristics include an indicator of whether the audio input is speech or music and a bandwidth of the audio input. The system determines a priority between the bitrate of the downmix channel and the bitrate of the metadata based on the characteristics. The system provides the priority to the bitrate calculation process.
[0107] In some implementations, the system extracts one or more parameters from the spatial MD, including the residual (side channel prediction error) level. The system determines a spatial audio coding mode indicating the need for one or more residual channels in the IVAS bitstream based on the parameters. The system provides the spatial audio coding mode to the bitrate calculation process.
[0108] In some implementations, the bitrate allocation control table index is stored in the common tool header (CTH) of the IVAS bitstream.
[0109] A system for decoding audio is configured to receive an IVAS bitstream. The system determines the IVAS bitrate and the bitrate allocation control table index based on the IVAS bitstream. The system performs a lookup in the bitrate allocation control table based on the table index and extracts the input format, the spatial coding mode, the mono backward compatibility mode, and the one or more indexes, the EVS target bitrate, and the bitrate ratio. The system extracts and decodes the downmix audio bits and the spatial MD bits for each downmix channel. The system provides the extracted downmix signal bits and spatial MD bits to a downstream IVAS device. The downstream IVAS device can be an audio processing device or a storage device.
[0110] SPAR FoA Bitrate Allocation Process In one embodiment, the bitrate allocation process described above for a stereo input signal can also be modified and applied to SPAR FoA bitrate allocation using the SPAR FoA bitrate allocation control table shown below. Definitions of terms included in the table are provided below to assist the reader. The SPAR FoA bitrate allocation control table follows. · Metadata target bits (MDtar) = IVAS_bits - header_bits - evs_target_bits (EVStar) · Metadata maximum bits (MDmax) = IVAS_bits - header_bits - evs_minimum_bits (EVSmin) · The metadata target bits should always be less than "MDmax".
Table 3
[0111] Some exemplary calculations of the maximum MD bit rate (actual coefficient) are shown in the following table.
Table 4
[0112] Exemplary metadata quantization loop: In one embodiment, the metadata quantization loop is implemented as described below. The metadata quantization loop includes two thresholds, MDtar and MDmax (defined above).
[0113] Step 1: For each frame of the input audio signal, the MD parameter is quantized in a non - time - differential manner and encoded by an arithmetic encoder. The actual metadata bit rate (MDact) is calculated based on the MD encoded bits. If MDact is less than MDtar, this step is considered a pass, the process ends the quantization loop, and the MDact bits are integrated into the IVAS bit stream. If there are surplus available bits (MDtar - MDAT), they are fed to the mono - codec (EVS) encoder to increase the bit rate of the essence of the down - mixed audio channels. The additional bit rate allows more information to be encoded by the mono - codec, and the decoded audio output is relatively less lossy.
[0114] Step 2: If step 1 fails, a subset of the MD parameter values within the frame is quantized and then subtracted from the quantized MD parameter values in the previous frame, and the quantized parameter values of the difference are encoded by an arithmetic coder (i.e., temporal differential encoding). MDact is calculated based on the MD encoded bits. If MDact is less than MDtar, this step is considered a pass, the process ends the quantization loop, and the MDact bits are integrated into the IVAS bitstream. If there are remaining available bits (MDtar - MDAT), they are supplied to the mono codec (EVS) encoder to increase the bitrate of the essence of the downmixed audio channels.
[0115] Step 3: If step 2 fails, the bitrate (MDact) of the quantized MD parameters is calculated without entropy.
[0116] Step 4: The MDact bitrate values calculated in steps 1 to 3 are compared with MDmax. If the minimum value among the MDact bitrates calculated in step 1, step 2, and step 3 is within MDmax, this step is considered a pass, the process ends the quantization loop, and the MD bitstream of the minimum MDact is integrated into the IVAS bitstream. If MDact is higher than MDtar, bits (MDact - MDtar) are removed from the mono codec (EVS) encoder.
[0117] Step 5: If step 4 fails, the parameters are quantized more coarsely, and the above steps are repeated as the first fallback strategy (fallback 1).
[0118] Step 6: If step 5 fails, the parameters are quantized with a quantization scheme that guarantees to fit within MDmax as the second fallback strategy (fallback 2).
[0119] After all of the above sequential iterations, it is guaranteed that the metadata bitrate stays within MDmax and the encoder generates the actual metadata bits or MDact.
[0120] Downmix channel / EVS bitrate allocation (EVSbd): In one embodiment, the actual bits of EVS (EVSact) = IVAS_bits - header_bits - MDact. If "EVSact" is less than "EVStar", bits are removed from the EVS channel in the following order (Z, X, Y, W). The maximum number of bits that can be taken from any channel is EVStar(ch) minus EVSmin(ch). If "EVSact" is greater than "EVStar", all additional bits are assigned to the downmix channels in the order of W, Y, X, Z. The maximum number of additional bits that can be added to any channel is EVSmax(ch) - EVStar(ch).
[0121] SPAR decoder unpacking In one embodiment, the SPAR decoder unpacks the IVAS bitstream as follows: 1. Obtain the IVAS bitrate from the bit length and obtain the table index from the tool header (CTH) in the IVAS bitstream. 2. Parse the header / metadata bits in the IVAS bitstream 3. Parse the metadata bits and dequantize them. 4. Set EVSact = the remaining bit length. 5. Read the table entries related to the EVS target, minimum, and maximum bitrates and repeat the "EVSbd" step in the decoder to obtain the actual EVS bitrate for each channel. 6. Decode the EVS channel and upmix to the FoA channel
[0122] BR allocation process for the SPAR FoA input audio signal Figures 5B and 5C are flow diagrams of a bitrate allocation process 515 for an SPAR FoA input signal according to an embodiment. Process 515 begins by preprocessing 517 the FoA inputs (W, Y, Z, X) 516 to extract signal characteristics using the IVAS bitrate, such as BW, speech / music classification data, VAD data. Process 515 generates spatial MDs (e.g., PR, C, P coefficients) 518 and selects (520) the number of residual channels to send to the IVAS decoder based on the residual level indicators within the spatial MDs, and continues by obtaining (521) a BR allocation control table index based on the IVAS bitrate, BW, and downmix channel number (N_dmx). In some embodiments, the P coefficients in the spatial MD can serve as the residual level indicators. The BR allocation control table index is sent to an IVAS bit packer (see FIGS. 4A, 4B) to be included in the IVAS bitstream that is stored and / or sent to the IVAS decoder.
[0123] Process 515 continues by reading (521) the SPAR configuration from the rows within the BR allocation control table pointed to by the table index. As shown in Table II above, the SPAR configuration is defined by one or more characteristics including, but not limited to, a downmix string (remix), an active W flag, a composite spatial MD flag, a spatial MD quantization strategy, EVS minimum / target / maximum bitrates, and a time domain decorrelator ducking flag.
[0124] As described above, process 515 continues by entering a quantization loop that determines (522) the MDmax and MDtar bitrates from the IVAS bitrate, EVSmin, and EVStar bitrate values, quantizes the spatial MD in a non-temporal difference manner using a quantization strategy, encodes the quantized spatial MD with an entropy encoder (e.g., an arithmetic coder), and calculates MDact (523). In one embodiment, the first iteration step of the quantization loop uses a fine quantization strategy.
[0125] Process 515 continues by checking (524) whether MDact is less than or equal to MDtar. If MDact is less than or equal to MDtar, the MD bits are sent to the IVAS bit packer to be included in the IVAS bit stream, and (MDtar - MDact) bits are added to the EVStar bitrate in the following order: W, Y, X, Z (532), N_dmx EVS bit streams (channels) are generated, and the EVS bits are sent to the IVAS bit packer to be included in the IVAS bit stream as described above. If MDact is not less than or equal to MDtar, process 515 quantizes the spatial MD in a temporal difference manner using a fine quantization strategy, encodes the quantized spatial MD with an entropy encoder, and calculates MDact again (525). If MDact is less than or equal to MDtar, the MD bits are sent to the IVAS bit packer to be included in the IVAS bit stream, (MDtar - MDact) bits are added to the EVStar bitrate in the following order: W, Y, X, Z (532), N_dmx EVS bit streams (channels) are generated, and the EVS bits are sent to the IVAS bit packer to be included in the IVAS bit stream as described above. If MDact is greater than MDtar, the spatial MD is quantized in a non-temporal difference manner using a fine quantization strategy and entropy, binary encoded, and a new value for MDact is calculated (527). Note that the maximum number of bits that can be added to any EVS instance is equal to EVSmax - EVStar.
[0126] Process 515 determines again whether MDact is less than or equal to MDtar (528). If MDact is less than or equal to MDtar, the MD bits are sent to the IVAS bit packer to be included in the IVAS bit stream, and (MDtar - MDact) bits are added to the EVStar bit rate in the following order W, Y, X, Z (532). N_dmx EVS bit streams (channels) are generated, and the EVS bits are sent to the IVAS bit packer to be included in the IVAS bit stream as described above. If MDact is greater than MDtar, process 515 sets MDact to the minimum value of the three MDact bit rates calculated in (523), (525), and (527), and compares MDact with MDmax (529). If MDact is greater than MDmax (530), the quantization loop (steps 523 - 530) is repeated using a coarser quantization strategy as described above.
[0127] If MDact is less than or equal to MDmax, the MD bits are sent to the IVAS bit packer, included in the IVAS bit stream, and process 515 determines again whether MDact is less than or equal to MDtar (531). If MDact is less than or equal to MDtar, (MDtar - MDact) bits are added to the EVStar bit rate in the following order W, Y, X, Z (532). N_dmx EVS bit streams (channels) are generated, and the EVS bits are sent to the IVAS bit packer to be included in the IVAS bit stream as described above. If MDact is greater than MDtar, (MDtar - MDact) bits are subtracted from the EVStar bit rate in the following order Z, X, Y, W (532). N_dmx EVS bit streams (channels) are generated, and the EVS bits are sent to the IVAS bit packer to be included in the IVAS bit stream as described above. Note that the maximum number of bits that can be subtracted from any EVS instance is equal to EVStar - EVSmin.
[0128] Exemplary Process Figure 6 is a flowchart of an IVAS encoding process 600 according to an embodiment. The process 600 can be implemented using the apparatus architecture as described with reference to Figure 8.
[0129] The process 600 receives (601) an input audio signal, downmixes (602) the input audio signal into one or more downmix channels and spatial metadata associated with the one or more downmix channels, reads (603) from a bitrate allocation control table a set of one or more bitrates for the downmix channels and a set of quantization levels for the spatial metadata, determines (604) a combination of one or more bitrates for the downmix channels, determines (605) a metadata quantization level from the set of metadata quantization levels using a bitrate allocation process, quantizes and encodes (606) the spatial metadata using the metadata quantization level, generates (607) a downmix bitstream for the one or more downmix channels using the combination of one or more bitrates, combines (608) the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into an IVAS bitstream, and streams or stores (609) the IVAS bitstream for playback on an IVAS-compatible device.
[0130] Figure 7 is a flowchart of an alternative IVAS encoding process 700 according to an embodiment. The process 700 can be implemented using the apparatus architecture as described with reference to Figure 8.
[0131] Process 700 includes steps of receiving an input audio signal (701), extracting characteristics of the input audio signal (702), calculating spatial metadata for channels of the input audio signal (703), reading from a bitrate allocation control table a set of one or more bitrates for downmix channels and a set of quantization levels for the spatial metadata (704), determining a combination of the one or more bitrates for the downmix channels (705), determining a metadata quantization level from the set of metadata quantization levels using a bitrate allocation process (706), quantizing and encoding the spatial metadata using the metadata quantization level (707), generating a downmix bitstream for the one or more downmix channels using the combination of the one or more bitrates (708), combining the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into an IVAS bitstream (709), and streaming or storing the IVAS bitstream for playback on an IVAS-compatible device (710).
[0132] Exemplary system architecture FIG. 8 shows a block diagram of an exemplary system 800 suitable for implementing an exemplary embodiment of the present disclosure. System 800 includes one or more server computers or any client devices, including, but not limited to, call server 102, legacy device 106, user devices 108, 114, conference room systems 116, 118, home theater systems, VR gear 122, and immersive content ingestor 124, such as any of the devices shown in FIG. 1. System 800 includes any consumer device including, but not limited to, smartphones, tablet computers, wearable computers, vehicle computers, gaming consoles, surround systems, kiosks.
[0133] As shown in the figure, system 800 includes a central processing unit (CPU) 801 that can execute various processes according to, for example, a program stored in a read-only memory (ROM) 802 or a program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 also stores data required for the CPU 801 to execute various processes as needed. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0134] The following components are connected to the I / O interface 805: an input unit 806 that may include a keyboard, a mouse, etc.; an output unit 807 that may include a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 808 that includes a hard disk or another suitable storage device; a communication unit 809 that includes a network interface card such as a network card (e.g., wired or wireless).
[0135] In some implementations, the input unit 806 includes one or more microphones at different positions (depending on the host device) that enable the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0136] In some implementations, the output unit 807 includes a system with various numbers of speakers. As shown in FIG. 1, the output unit 807 can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats) (depending on the capabilities of the host device). The communication unit 809 is configured to communicate with other devices (e.g., via a network). Optionally, the drive 810 is also connected to the I / O interface 805. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive, or other suitable removable medium, is mounted on the drive 810, and a computer program read therefrom is installed in the storage unit 808 as needed. Those skilled in the art will understand that although the system 800 is described as including the above-described components, in actual applications, some of these components can be added, removed, and / or replaced, and all of these modifications or changes are within the scope of the present disclosure.
[0137] According to an exemplary embodiment of the present disclosure, the above-described process can be implemented as a computer software program or on a computer-readable storage medium. For example, an embodiment of the present disclosure includes a computer program product including a computer program embodied in a tangible manner on a machine-readable medium, the computer program including program code for performing the method. In such an embodiment, the computer program may be downloaded from a network via the communication unit 809, mounted, and / or installed from the removable medium 811, as shown in FIG. 8.
[0138] In general, the various exemplary embodiments of the present disclosure can be implemented in hardware or special-purpose circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the units described above can be executed by a control circuit (e.g., a CPU combined with other components in FIG. 8), and thus the control circuit can execute the actions described in the present disclosure. Some aspects can be implemented in hardware, and other aspects can be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing device (e.g., control circuitry). Although the various aspects of the exemplary embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, the blocks, devices, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special-purpose circuitry or logic, general-purpose hardware or controllers, or other computing devices, or any combination thereof.
[0139] Furthermore, the various blocks shown in the flowchart can be regarded as method steps and / or acts resulting from the operation of computer program code and / or as a plurality of interconnected logic circuit elements configured to perform related functions. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to execute the above method.
[0140] In the context of the present disclosure, a machine-readable medium may be any tangible medium that can store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may be non-transitory and may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Further specific examples of the machine-readable storage medium include electrical connections having one or more wires, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0141] The computer program code for carrying out the methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus having a control circuit, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowchart and / or block diagram to be executed. The program codes may be executed entirely on the computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer, or entirely on the remote computer or server, or distributed across one or more remote computers and / or servers.
[0142] This document contains details of many individual implementations, which should not be construed as limitations on the scope of what can be claimed, but rather as descriptions and interpretations of features that may be specific to particular embodiments. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, features may be described above as acting in certain combinations and may initially be claimed as such, but one or more features from the claimed combination may, in some cases, be excised from that combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination. The logical flow shown in the figures does not require a particular order or sequential order to achieve the desired result. In addition, other steps may be provided, or steps may be removed from the flow described, other components may be added to the system described, or components described may be removed from the system described. Thus, other implementations are within the scope of the claims. Some aspects will be described. [Aspect 1] A method for encoding an immersive audio and audio service (IVAS) bitstream, the method comprising: receiving an input audio signal using one or more processors; downmixing the input audio signal to one or more downmix channels and spatial metadata associated with one or more channels of the input audio signal using the one or more processors; reading, using the one or more processors, from a bitrate allocation control table, a set of one or more bitrates for the downmix channels and a set of quantization levels for the spatial metadata; Using the one or more processors, determining one or more combinations of bitrates for the downmix channels; Using the one or more processors, determining a metadata quantization level from the set of metadata quantization levels using a bitrate allocation process; Using the one or more processors, quantizing and encoding the spatial metadata using the metadata quantization level; Using the one or more processors and the combination of one or more bitrates, generating a downmix bitstream for the one or more downmix channels; Using the one or more processors, combining the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into the IVAS bitstream; Streaming or storing the IVAS bitstream for playback on an IVAS-compatible device, Method. [Aspect 2] The method according to aspect 1, wherein the input audio signal is a 4-channel first-order ambisonic (FoA) audio signal, a 3-channel planar FoA signal, or a 2-channel stereo audio signal. [Aspect 3] The method according to aspect 1 or 2, wherein the one or more bitrates are the bitrates of one or more instances of a mono audio codec / decoder (codec). [Aspect 4] The method according to aspect 1 or 2, wherein the mono audio codec is an Enhanced Voice Service (EVS) codec and the downmix bitstream is an EVS bitstream. [Aspect 5] Using the one or more processors, obtaining one or more bitrates for the downmix channels and the spatial metadata using a bitrate allocation control table further comprises: Identifying a row in the bitrate allocation control table using a table index including the format of the input audio signal, the bandwidth of the input audio signal, allowed spatial encoding tools, a transition mode, and a mono-downmix backward compatibility mode; Extracting a target bitrate, a bitrate ratio, a minimum bitrate, and a bitrate deviation quantization step from the identified row of the bitrate allocation control table, wherein the bitrate ratio indicates a ratio by which the total bitrate is distributed among the downmix audio signal channels, the minimum bitrate is a value below which the total bitrate is not allowed to fall, and the bitrate deviation quantization step is a target bitrate reduction quantization step when a first priority for the downmix signal is greater than or less than a second priority for the spatial metadata; Determining the one or more bitrates for the downmix channels and the spatial metadata based on the target bitrate, the bitrate ratio, the minimum bitrate, and the bitrate deviation quantization step. The method according to aspect 1 or 2. [Aspect 6] The method according to aspect 1 or 2, wherein quantization is performed in a quantization loop that applies a quantization strategy of gradually coarsening based on a difference between a target metadata bitrate and an actual metadata bitrate when quantizing the spatial metadata for the one or more channels of the input audio signal using a set of quantization levels. [Aspect 7] The method according to aspect 1 or 2, wherein the quantization is determined according to a mono-codec priority and a spatial metadata priority based on characteristics extracted from the input audio signal and channel-bandwidth covariance values. 〔Aspect 8〕 The method according to aspect 1 or 2, wherein the input audio signal is a stereo signal, and the downmix signal includes a mid signal from the stereo signal, a residual representation, and the spatial metadata. 〔Aspect 9〕 The spatial metadata includes prediction coefficients (PR), cross-prediction coefficients (C), and decorrelation (P) coefficients for the spatial reconstruction (SPAR) format, and prediction coefficients ( PR ) and decorrelation coefficients ( P ) for the complex advanced coupling (CACPL) format, in the method according to aspect 1 or 2. 〔Aspect 10〕 A method for encoding an immersive audio and audio service (IVAS) bitstream, the method comprising: receiving an input audio signal using one or more processors; extracting characteristics of the input audio signal using the one or more processors; calculating spatial metadata for channels of the input audio signal using the one or more processors; reading, from a bitrate allocation control table using the one or more processors, a set of one or more bitrates for the downmix channels and a set of quantization levels for the spatial metadata; determining a combination of the one or more bitrates for the downmix channels using the one or more processors; determining a metadata quantization level from the set of metadata quantization levels using a bitrate allocation process with the one or more processors; quantizing and encoding the spatial metadata using the metadata quantization level with the one or more processors; Using the one or more processors and the combination of the one or more bitrates, generating a downmix bitstream for the one or more downmix channels using the one or more bitrates; Using the one or more processors, combining the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into the IVAS bitstream; Streaming or storing the IVAS bitstream for playback on an IVAS-compatible device, Method. [Aspect 11] The method according to aspect 10, wherein the characteristics of the input audio signal include one or more of bandwidth, speech / music classification data, and voice activity detection (VAD) data. [Aspect 12] The method according to aspect 10 or 11, wherein the input audio signal is a 4-channel first-order ambisonic (FoA) audio signal, a 3-channel planar FoA signal, or a 2-channel stereo audio signal. [Aspect 13] The method according to aspect 10 or 11, wherein the one or more bitrates are the bitrates of one or more instances of a mono audio codec / decoder (codec). [Aspect 14] The method according to aspect 10 or 11, wherein the mono audio codec is an enhanced voice service (EVS) codec and the downmix bitstream is an EVS bitstream. [Aspect 15] The step of obtaining, using the one or more processors and using a bitrate allocation control table, the one or more bitrates for the downmix channels and the set of quantization levels for the spatial metadata further comprises: Identifying a row in the bitrate distribution control table using a table index including the format of the input audio signal, the bandwidth of the input audio signal, allowed spatial encoding tools, a transition mode, and a mono-downmix backward compatibility mode; Extracting a target bitrate, a bitrate ratio, a minimum bitrate, and a bitrate deviation quantization step from the identified row of the bitrate distribution control table, wherein the bitrate ratio indicates a ratio at which the total bitrate is distributed among the input audio signal channels, the minimum bitrate is a value below which the total bitrate is not allowed to fall, and the bitrate deviation quantization step is a target bitrate reduction quantization step when a first priority for the downmix signal is equal to or higher than a second priority for the spatial metadata, or lower than it; Determining the one or more bitrates for the downmix channel and the spatial metadata based on the target bitrate, the bitrate ratio, the minimum bitrate, and the bitrate deviation quantization step. The method according to aspect 10 or 11. [Aspect 16] The method according to aspect 10 or 11, wherein quantization is performed in a quantization loop that applies a quantization strategy that gradually coarsens based on a difference between a target metadata bitrate and an actual metadata bitrate when quantizing the spatial metadata for the one or more channels of the input audio signal using a set of quantization levels. [Aspect 17] The method according to aspect 10 or 11, wherein the quantization is determined according to a mono-coder priority and a spatial metadata priority based on characteristics extracted from the input audio signal and a channel banding co-variance value. [Aspect 18] The input audio signal is a stereo signal, and the downmix signal includes a mid signal from the stereo signal, a residual representation, and the spatial metadata, according to the method of aspect 10 or 11. [Aspect 19] The spatial metadata includes prediction coefficients (PR), cross-prediction coefficients (C), and decorrelation (P) coefficients for the spatial reconstructor (SPAR) format, and prediction coefficients ( PR ) and decorrelation coefficients ( P ) for the complex advanced coupling (CACPL) format, according to the method of aspect 10 or 11. [Aspect 20] The number of downmix channels encoded in the IVAS bitstream is selected based on the residual level indicator in the spatial metadata, according to the method of aspect 10 or 11. [Aspect 21] A method for encoding an immersive audio and audio service (IVAS) bitstream, comprising: receiving a first-order ambisonic (FoA) input audio signal using one or more processors; extracting characteristics of the FoA input audio signal using the one or more processors and an IVAS bitrate, wherein one of the characteristics is the bandwidth of the FoA input audio signal; generating spatial metadata for the FoA input audio signal using the FoA signal characteristics using the one or more processors; selecting the number of residual channels to transmit based on a residual level indicator and a decorrelation coefficient in the spatial metadata using the one or more processors; obtaining a bitrate allocation control table index based on the IVAS bitrate, the bandwidth, and the number of downmix channels using the one or more processors; Using the one or more processors, reading a spatial reconfigurer (SPAR) configuration from a row in the bitrate allocation control table pointed to by the bitrate allocation control table index; Using the one or more processors, determining a sum of a target metadata bitrate from the IVAS bitrate, the target EVS bitrate, and a length of the IVAS header; Using the one or more processors, determining a sum of a maximum metadata bitrate from the IVAS bitrate, a minimum EVS bitrate, and a length of the IVAS header; Using the one or more processors and a quantization loop, quantizing the spatial metadata in a non-temporal differential manner according to a first quantization strategy; Using the one or more processors, entropy encoding the quantized spatial metadata; Using the one or more processors, calculating a first actual metadata bitrate; Using the one or more processors, determining whether the first actual metadata bitrate is less than or equal to a target metadata bitrate; ending the quantization loop in response to the first actual metadata bitrate being less than or equal to the target metadata bitrate, Method. [Aspect 22] Further: Using the one or more processors, determining a first total actual EVS bitrate by adding a first bit amount equal to a difference between the metadata target bitrate and the first actual metadata bitrate to a total EVS target bitrate; Using the one or more processors, generating an EVS bitstream using the first total actual EVS bitrate; Using the one or more processors, generating an IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy-coded spatial metadata; In response to the first actual metadata bitrate being greater than the target metadata bitrate: Using the one or more processors, quantizing the spatial metadata in a temporal difference manner according to the first quantization strategy; Using the one or more processors, entropy-coding the quantized spatial metadata; Using the one or more processors, calculating a second actual metadata bitrate; Using the one or more processors, determining whether the second actual metadata bitrate is less than or equal to the target metadata bitrate; In response to the second actual metadata bitrate being less than or equal to the target metadata bitrate, ending the quantization loop, including the steps of The method according to aspect 21. [Aspect 23] Furthermore: Using the one or more processors, determining a second total actual EVS bitrate by adding a second bit amount equal to the difference between the metadata target bitrate and the second actual metadata bitrate to the total EVS target bitrate; Using the one or more processors, generating an EVS bitstream using the second total actual EVS bitrate; Using the one or more processors, generating the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy-coded spatial metadata; In response to the second actual metadata bitrate being greater than the bitrate of the target metadata: Using the one or more processors, quantize the spatial metadata in a non-temporal difference manner according to the first quantization strategy; Using the one or more processors and a binary encoder, encode the quantized spatial metadata; Using the one or more processors, calculating a third actual metadata bitrate; In response to the third actual metadata bitrate being less than or equal to the target metadata bitrate, ending the quantization loop, including the step of The method according to aspect 22. [Aspect 24] Furthermore: Using the one or more processors, determining a third total actual EVS bitrate by adding a third bit amount equal to the difference between the metadata target bitrate and the third actual metadata bitrate to the total EVS target bitrate; Using the one or more processors, generating an EVS bitstream using the third total actual EVS bitrate; Using the one or more processors, generating the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy encoded spatial metadata; In response to the third actual metadata bitrate being greater than the target metadata bitrate: Using the one or more processors, setting a fourth actual metadata bitrate to the minimum value of the first, second, and third actual metadata bitrates; Determining, using the one or more processors, whether the fourth actual metadata bitrate is less than or equal to the maximum metadata bitrate; In response to the fourth actual metadata bitrate being less than or equal to the maximum metadata bitrate: Determining, using the one or more processors, whether the fourth actual metadata bitrate is less than or equal to the target metadata bitrate; In response to the fourth actual metadata bitrate being less than or equal to the target metadata bitrate, ending the quantization loop, including the method according to aspect 23. The method according to aspect 23. 〔Aspect 25〕 Furthermore: Determining a fourth total actual EVS bitrate by adding, using the one or more processors, a fourth bit amount equal to the difference between the metadata target bitrate and the fourth actual metadata bitrate to the total EVS target bitrate; Generating an EVS bitstream using, using the one or more processors, the fourth total actual EVS bitrate; Generating the IVAS bitstream including, using the one or more processors, the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy-coded spatial metadata; Ending the quantization loop in response to the fourth actual metadata bitrate being greater than the target metadata bitrate and less than or equal to the maximum target metadata bitrate, including the method according to aspect 24. The method according to aspect 24. 〔Aspect 26〕 Furthermore: Using the one or more processors, determining a fifth total actual EVS bitrate by subtracting from the total EVS target bitrate an amount of bits equal to the difference between the fourth actual metadata bitrate and the target metadata bitrate; Using the one or more processors, generating an EVS bitstream using the fifth actual EVS bitrate; Using the one or more processors, generating the IVAS bitstream including the EVS bitstream, the bitrate allocation control table index, and the quantized and entropy-coded spatial metadata; In response to the fourth actual metadata bitrate being greater than the maximum target metadata bitrate: changing the first quantization strategy to a second quantization strategy and entering the quantization loop again using the second quantization strategy, the second quantization strategy being coarser than the first quantization strategy; The method according to aspect 25. [Aspect 27] The SPAR configuration is defined by a downmix string, an active W flag, a complex spatial metadata flag, a spatial metadata quantization strategy, minimum, maximum, and target bitrates for one or more instances of an extended voice service (EVS) mono coder / decoder (codec), and a time domain decorrelator dithering flag, the method according to any one of aspects 21 to 26. [Aspect 28] The actual total number of EVS bits is equal to subtracting the number of header bits from the number of IVAS bits and subtracting the actual metadata bit rate. When the actual total number of EVS bits is less than the total number of EVS target bits, bits are removed from the EVS channel in the order of Z, X, Y, W, and the maximum number of bits that can be removed from any channel is the result of subtracting the minimum number of EVS bits for that channel from the number of EVS target bits for that channel. When the actual total number of EVS bits is greater than the number of EVS target bits, all additional bits are assigned to the downmix channels in the order of W, Y, X, Z, and the maximum number of additional bits that can be added to any channel is the result of subtracting the number of EVS target bits from the maximum number of EVS bits. The method according to any one of aspects 21 to 26. 〔Aspect 29〕 A method for decoding an immersive audio and audio service (IVAS) bitstream is: Receiving the IVAS bitstream using one or more processors; Obtaining the IVAS bit rate from the bit length of the IVAS bitstream using one or more processors; Obtaining a bit rate allocation control table index from the IVAS bitstream using the one or more processors; Parsing a metadata quantization strategy from the header of the IVAS bitstream using the one or more processors; Parsing and dequantizing the quantized spatial metadata bits based on the metadata quantization strategy using the one or more processors; Setting the actual number of extended voice service (EVS) bits equal to the remaining bit length of the IVAS bitstream using the one or more processors; Using the one or more processors and the bitrate allocation control table index, reading a table entry of the bitrate allocation control table including an EVS target, and an EVS minimum bitrate and a maximum EVS bitrate for one or more EVS instances; Using the one or more processors to obtain an actual EVS bitrate for each downmix channel; Using the one or more processors to decode each EVS channel using the actual EVS bitrate for that channel; Using the one or more processors to upmix the EVS channel to a first-order ambisonic (FoA) channel, including: Method. 〔Aspect 30〕 One or more processors; A non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operations of the method according to any one of Aspects 1 to 29; System. 〔Aspect 31〕 A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of the method according to any one of Aspects 1 to 29.
Claims
**Claim 1** A method for encoding immersive audio and an Audio Service (IVAS) bitstream, the method comprising: Receiving an input audio signal using one or more processors; Using the one or more processors to downmix the input audio signal into one or more downmix channels and generate spatial metadata associated with one or more channels of the input audio signal; Using the one or more processors to obtain from a bitrate allocation control table a set of one or more bitrates for the downmix channels and a set of metadata quantization levels for the spatial metadata; Using the one or more processors to determine a combination of the one or more bitrates for the downmix channels; Using the one or more processors to determine a metadata quantization level from the set of metadata quantization levels; Using the one or more processors to quantize and encode the spatial metadata using the metadata quantization level; Using the one or more processors and the combination of the one or more bitrates to generate a downmix bitstream for the one or more downmix channels; Using the one or more processors to combine the downmix bitstream, the quantized and encoded spatial metadata, and information on the metadata quantization level into the IVAS bitstream, Method. **Claim 2** The method of claim 1, wherein the one or more bitrates are the bitrates of one or more instances of a mono audio codec / decoder (codec). **Claim 3** The method of claim 2, wherein the mono audio codec is an Enhanced Voice Service (EVS) codec and the downmix bitstream is an EVS bitstream. **Claim 4** The quantization is determined according to the mono codec priority and the spatial metadata priority based on the characteristics extracted from the input audio signal and the channel banded covariance value, according to the method of claim 1.
5. The input audio signal is a stereo signal, and the downmix channel includes a mid signal and a residual representation from the stereo signal, according to the method of claim 1.
6. One or more processors; A non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operations of the method according to any one of claims 1 to 5. System.
7. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Adaptive Bit Allocation in Multichannel Speech Coding
JP2008529056A
A method for parametric multichannel encoding
JP2016509260A
Method and device for allocating a bit-budget between sub-frames in a CELP codec
WO2019056107A1
Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to dirac based spatial audio coding
WO2019068638A1