Bit rate distribution in immersive voice and audio services
By using a bit rate distribution control table and quantization strategy, the problem of inefficient bit rate allocation in immersive speech and audio service codecs is solved, efficient audio signal processing is achieved on various devices, and bit loss is reduced.
Patent Information
- Application Number
- CN202511154550.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-16
- Filing Date
- 2020-10-28
- Publication Date
- 2025-09-26
AI Technical Summary
Existing immersive speech and audio service codecs suffer from inefficient bit rate allocation and high bit loss, making it difficult to achieve efficient audio signal encoding and decoding on various devices and network nodes.
Using a bit rate distribution control table and quantization strategy, the processor receives the input audio signal, downmixes it into multiple channels and associates it with spatial metadata, determines the bit rate and quantization level, quantizes and encodes the spatial metadata, generates a downmix bitstream, and finally forms an IVAS bitstream to adapt to the playback requirements of different devices.
The bit rate allocation of the IVAS codec has been optimized, and the addition of spatial metadata and mono codec has been reduced, thus reducing bit loss and enabling efficient audio signal processing on various devices.
Smart Images

Figure CN120708627A_ABST
Abstract
Description
[0001] Information about divisional applications
[0002] This application is a divisional application. The parent application is an invention patent application filed on October 28, 2020, with application number 202080075350.8 and titled “Bit Rate Distribution in Immersive Voice and Audio Services.”
[0003] Cross-reference to related applications
[0004] This application claims priority to U.S. Provisional Patent Application No. 62 / 927,772, filed on October 30, 2019, and U.S. Provisional Patent Application No. 63 / 092,830, filed on October 16, 2020, which are incorporated herein by reference. Technical Field
[0005] The present disclosure generally relates to audio bitstream encoding and decoding. Background Art
[0006] Speech and audio coder / decoder ("codec") standards development has recently focused on developing codecs for immersive voice and audio services (IVAS). IVAS is expected to support a range of audio service capabilities, including (but not limited to) mono to stereo upmixing and fully immersive audio encoding, decoding, and rendering. IVAS is expected to be supported by a wide range of devices, endpoints, and network nodes, including (but not limited to): mobile phones and smartphones, electronic tablets, personal computers, conference phones, conference rooms, virtual reality (VR) and augmented reality (AR) devices, home theater devices, and other suitable devices. These devices, endpoints, and network nodes may have various acoustic interfaces for sound capture and rendering. Summary of the Invention
[0007] Implementations of bit rate distribution in immersive speech and audio services are disclosed.
[0008] In an embodiment, a method of encoding an immersive voice and audio service (IVAS) bitstream, the method comprising: receiving an input audio signal using one or more processors; downmixing the input audio signal into one or more downmix channels and spatial metadata associated with the one or more channels of the input audio signal using the one or more processors; reading a set of one or more bit rates for the downmix channels and a set of quantization levels for the spatial metadata from a bit rate distribution control table using the one or more processors; determining a combination of the one or more bit rates for the downmix channels using the one or more processors; and The one or more processors determine a metadata quantization level from the set of metadata quantization levels using a bitrate profile process; quantize and encode the spatial metadata using the metadata quantization level using the one or more processors; generate a downmix bitstream for the one or more downmix channels using the one or more processors and the combination of one or more bitrates; combine the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into the IVAS bitstream using the one or more processors; and stream or store the IVAS bitstream for playback on an IVAS-capable device.
[0009] In an embodiment, the input audio signal is a four-channel first-order Ambisonics (FoA) audio signal, a three-channel planar FoA signal, or a two-channel stereo audio signal.
[0010] In an embodiment, the one or more bit rates are bit rates of one or more channels of a mono audio coder / decoder (codec) bit rate.
[0011] In an embodiment, the mono audio codec is an Enhanced Voice Services (EVS) codec and the downmix bitstream is an EVS bitstream.
[0012] In an embodiment, obtaining the one or more bit rates of the downmix channels and the spatial metadata using a bit rate distribution control table using the one or more processors further comprises: identifying a row in the bit rate distribution control table using a table index, the row comprising a format of the input audio signal, a bandwidth of the input audio signal, allowed spatial coding tools, a transition mode, and a mono downmix backward compatibility mode; and extracting a target bit rate, a bit rate ratio, a minimum bit rate, and a bit rate deviation step from the identified row of the bit rate distribution control table, wherein the bit rate ratio indicates a ratio of a total bit rate distributed among the downmix audio signal channels, the minimum bit rate is a value below which the total bit rate is not allowed to be achieved, and the bit rate deviation step is a target bit rate reduction step when a first priority of the downmix signal is higher than, equal to, or lower than a second priority of the spatial metadata; and determining the one or more bit rates of the downmix channels and the spatial metadata based on the target bit rate, the bit rate ratio, the minimum bit rate, and the bit rate deviation step.
[0013] In an embodiment, quantizing the spatial metadata of the one or more channels of the input audio signal using a set of quantization levels is performed in a quantization loop that applies increasingly coarser quantization strategies based on a difference between a target metadata bitrate and an actual metadata bitrate.
[0014] In an embodiment, the quantization is determined based on properties extracted from the input audio signal and channel band covariance values according to a mono codec priority and a spatial metadata priority.
[0015] In an embodiment, the input audio signal is a stereo signal and the downmix signal comprises an intermediate signal, a residual from the stereo signal and a representation of the spatial metadata.
[0016] In an embodiment, the spatial metadata includes prediction coefficients (PR), cross-prediction coefficients (C), and decorrelation coefficients (P) for a spatial reconstructor (SPAR) format and prediction coefficients (P) and decorrelation coefficients (PR) for a complex advanced coupled (CACPL) format.
[0017] In an embodiment, a method for encoding an immersive voice and audio service (IVAS) bitstream comprises: receiving an input audio signal using one or more processors; extracting properties of the input audio signal using the one or more processors; computing spatial metadata of channels of the input audio signal using the one or more processors; reading a set of one or more bit rates for the downmix channels and a set of quantization levels for the spatial metadata from a bit rate distribution control table using the one or more processors; determining a combination of the one or more bit rates for the downmix channels using the one or more processors; and computing spatial metadata of the channels of the input audio signal using the one or more processors. The invention also provides a method for transmitting the spatial metadata of the one or more downmix channels to an IVAS device, wherein the one or more processors are configured to determine a metadata quantization level from the set of metadata quantization levels using a bitrate profile process; quantizing and encoding the spatial metadata using the metadata quantization level using the one or more processors; generating a downmix bitstream for the one or more downmix channels using the one or more bitrates using the combination of the one or more processors and the one or more bitrates; combining the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into the IVAS bitstream using the one or more processors; and streaming or storing the IVAS bitstream for playback on an IVAS-capable device.
[0018] In an embodiment, the properties of the input audio signal include one or more of bandwidth, speech / music classification data, and voice activity detection (VAD) data.
[0019] In an embodiment, the number of downmix channels to be encoded into the IVAS bitstream is selected based on a residual level indicator in the spatial metadata.
[0020] In an embodiment, a method of encoding an immersive voice and audio service (IVAS) bitstream further comprises: receiving, using one or more processors, a first-order ambisonic (FoA) input audio signal; extracting, using the one or more processors and an IVAS bitrate, properties of the FoA input audio signal, wherein one of the properties is a bandwidth of the FoA input audio signal; generating, using the one or more processors, spatial metadata for the FoA input audio signal using the FoA signal properties; selecting, using the one or more processors, a number of residual channels to send based on a residual level indicator and a decorrelation coefficient in the spatial metadata; obtaining, using the one or more processors, a bitrate profile control table index based on the IVAS bitrate, the bandwidth, and the number of downmix channels; and reading, using the one or more processors, a bitrate profile control table index from a row of the bitrate profile control table pointed to by the bitrate profile control table index. The method includes: assuming a spatial reconstructor (SPAR) configuration; determining, using the one or more processors, a target metadata bit rate from the sum of the IVAS bit rate, the target EVS bit rate, and the length of the IVAS header; determining, using the one or more processors, a maximum metadata bit rate from the sum of the IVAS bit rate, a minimum EVS bit rate, and the length of the IVAS header; quantizing, using the one or more processors and a quantization loop, the spatial metadata in a non-temporal difference manner according to a first quantization strategy; entropy encoding the quantized spatial metadata, using the one or more processors; calculating, using the one or more processors, a first actual metadata bit rate; determining, using the one or more processors, whether the first actual metadata bit rate is less than or equal to a target metadata bit rate; and exiting the quantization loop in response to the first actual metadata bit rate being less than or equal to the target metadata bit rate.
[0021] In an embodiment, the method further includes: determining, using the one or more processors, a first total actual EVS bit rate by adding a first amount of bits equal to the difference between the metadata target bit rate and the first actual metadata bit rate to the total EVS target bit rate; generating, using the one or more processors, an EVS bit stream using the first total actual EVS bit rate; generating, using the one or more processors, an IVAS bit stream comprising the EVS bit stream, the bit rate distribution control table index, and the quantized and entropy encoded spatial metadata; based on the first actual metadata bit rate being greater than the target metadata bit rate: quantizing, using the one or more processors, the spatial metadata in a temporal differential manner according to the first quantization strategy; entropy encoding, using the one or more processors, the quantized spatial metadata; calculating, using the one or more processors, a second actual metadata bit rate; determining, using the one or more processors, whether the second actual metadata bit rate is less than or equal to the target metadata bit rate; and exiting the quantization loop based on the second actual metadata bit rate being less than or equal to the target metadata bit rate.
[0022] In an embodiment, the method further includes: determining, using the one or more processors, a second total actual EVS bit rate by adding a second amount of bits equal to the difference between the metadata target bit rate and the second actual metadata bit rate to the total EVS target bit rate; generating, using the one or more processors, an EVS bit stream using the second total actual EVS bit rate; generating, using the one or more processors, the IVAS bit stream comprising the EVS bit stream, the bit rate distribution control table index, and the quantized and entropy encoded spatial metadata; based on the second actual metadata bit rate being greater than the target metadata bit rate: quantizing, using the one or more processors, the spatial metadata in a non-temporal difference manner according to the first quantization strategy; encoding the quantized spatial metadata using the one or more processors and a base2 encoder; calculating, using the one or more processors, a third actual metadata bit rate; and exiting the quantization loop based on the third actual metadata bit rate being less than or equal to the target metadata bit rate.
[0023] In an embodiment, the method further includes: determining, using the one or more processors, a third total actual EVS bitrate by adding a third amount of bits equal to the difference between the metadata target bitrate and the third actual metadata bitrate to the total EVS target bitrate; generating, using the one or more processors, an EVS bitstream using the third total actual EVS bitrate; generating, using the one or more processors, the IVAS bitstream comprising the EVS bitstream, the bitrate distribution control table index, and the quantized and entropy coded spatial metadata; based on the third actual metadata bitrate being greater than the target metadata bitrate: setting, using the one or more processors, a fourth actual metadata bitrate to the minimum of the first, second, and third actual metadata bitrates; determining, using the one or more processors, whether the fourth actual metadata bitrate is less than or equal to a maximum metadata bitrate; based on the fourth actual metadata bitrate being less than or equal to the maximum metadata bitrate: determining, using the one or more processors, whether the fourth actual metadata bitrate is less than or equal to the target metadata bitrate; and based on the fourth actual metadata bitrate being less than or equal to the target metadata bitrate, exiting the quantization loop.
[0024] In an embodiment, the method further includes: determining, using the one or more processors, a fourth total actual EVS bit rate by adding a fourth amount of bits equal to the difference between the metadata target bit rate and the fourth actual metadata bit rate to the total target EVS bit rate; generating, using the one or more processors, an EVS bit stream using the fourth total actual EVS bit rate; generating, using the one or more processors, the IVAS bit stream including the EVS bit stream, the bit rate distribution control table index, and the quantized and entropy coded spatial metadata; and exiting the quantization loop based on the fourth actual metadata bit rate being greater than the target metadata bit rate and less than or equal to the maximum metadata bit rate.
[0025] In an embodiment, the method further comprises: determining, using the one or more processors, a fifth total actual EVS bitrate by subtracting a number of bits equal to the difference between the fourth actual metadata bitrate and the target metadata bitrate from the total target EVS bitrate; generating, using the one or more processors, an EVS bitstream using the fifth actual EVS bitrate; generating, using the one or more processors, the IVAS bitstream comprising the EVS bitstream, the bitrate profile control table index, and the quantized and entropy-encoded spatial metadata; and, based on the fourth actual metadata bitrate being greater than the maximum metadata bitrate: changing the first quantization strategy to a second quantization strategy and re-entering the quantization loop using the second quantization strategy, wherein the second quantization strategy is coarser than the first quantization strategy. In an embodiment, a third quantization strategy may be used that ensures an actual MD bitrate less than the maximum MD bitrate.
[0026] In an embodiment, the SPAR configuration is defined by a downmix string, an active W flag, a composite spatial metadata flag, a spatial metadata quantization strategy, minimum, maximum, and target bit rates for one or more instances of an Enhanced Voice Services (EVS) mono encoder / decoder (codec), and a time-domain decorrelator volume reduction flag.
[0027] In an embodiment, the actual total number of EVS bits is equal to the number of IVAS bits minus the number of header bits minus the actual metadata bit rate, and wherein if the number of total actual EVS bits is less than the total number of EVS target bits, then bits are taken from the EVS channels in the following order: Z, X, Y, and W, and wherein the maximum number of bits that can be taken from any channel is the number of EVS target bits for that channel minus the minimum number of EVS bits for that channel, and wherein if the number of actual EVS bits is greater than the number of EVS target bits, then all extra bits are assigned to the downmix channels in the following order: W, Y, X, and Z, and the maximum number of extra bits that can be added to any channel is the maximum number of EVS bits minus the number of EVS target bits.
[0028] In an embodiment, a method for decoding an immersive voice and audio service (IVAS) bitstream includes: receiving an IVAS bitstream using one or more processors; obtaining an IVAS bitrate from a bit length of the IVAS bitstream using one or more processors; obtaining a bitrate profile control table index from the IVAS bitstream using the one or more processors; parsing a metadata quantization policy from a header of the IVAS bitstream using the one or more processors; parsing and dequantizing the quantized spatial metadata bits based on the metadata quantization policy using the one or more processors; and converting, using the one or more processors, an enhanced voice service (EVAS) bitstream into an IVAS bitstream. The method includes setting an actual number of IVAS (EVS) bits to be equal to the remaining bit length of the IVAS bitstream; reading a table entry of the bitrate distribution control table containing an EVS target and minimum EVS bitrates and a maximum EVS bitrate for one or more EVS instances using the one or more processors and the bitrate distribution control table index; obtaining an actual EVS bitrate for each downmix channel using the one or more processors; and decoding each EVS channel using the actual EVS bitrate for the channel using the one or more processors; and upmixing the EVS channels to first-order ambisonic (FoA) channels using the one or more processors.
[0029] In an embodiment, a system includes: one or more processors; and a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operations of any of the methods described above.
[0030] In an embodiment, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of any of the methods described above.
[0031] Other embodiments disclosed herein relate to a system, apparatus, and computer-readable medium. Details of the disclosed embodiments are set forth in the accompanying drawings and description below. Other features, objects, and advantages are apparent from the description, drawings, and claims.
[0032] Certain embodiments disclosed herein provide one or more of the following advantages. The IVAS codec bit rate is distributed between the mono codec and spatial metadata (MD) and between multiple instances of the mono codec. For a given audio frame, the IVAS codec determines the spatial audio coding mode (parametric or residual coding). The IVAS bitstream is optimized to reduce spatial MD, reduce mono codec overhead, and minimize bit loss to zero. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In the drawings, for ease of description, a specific arrangement or order of schematic elements (e.g., elements representing devices, units, instruction blocks, and data elements) is shown. However, those skilled in the art will understand that the specific order or arrangement of the schematic elements in the drawings is not intended to imply that a specific order or sequence of processing or separation of processes is required. Furthermore, in some embodiments, the inclusion of a schematic element in a drawing is not intended to imply that the element is required in all embodiments or that the features represented by the element may not be included in or combined with other elements.
[0034] In addition, in a diagram in which a connecting element (e.g., a solid or dashed line or arrow) is used to illustrate a connection, relationship, or association between or among two or more other schematic elements, the lack of any such connecting element does not intend to imply that a connection, relationship, or association may not exist. In other words, some connections, relationships, or associations between elements are not shown in the diagram in order to avoid making the present disclosure unclear. In addition, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, in the case where a connecting element represents the communication of a signal, data, or instruction, it will be understood by those skilled in the art that such elements may represent one or more signal paths for communication as needed.
[0035] Figure 1 The use of the IVAS codec according to the embodiment is described.
[0036] Figure 2 is a block diagram of a system for encoding and decoding an IVAS bitstream, according to an embodiment.
[0037] Figure 3 is a block diagram of a Spatial Reconstructor (SPAR) First Order Ambisonics (FoA) encoder / decoder ("codec") for encoding and decoding an IVAS bitstream in FoA format, according to an embodiment.
[0038] Figure 4A is a block diagram of an IVAS signal chain for FoA and stereo input signals, according to an embodiment.
[0039] Figure 4B is a block diagram of an alternative IVAS signal chain for FoA and stereo input signals, according to an embodiment.
[0040] Figure 5A is a flow chart of a bitrate profiling process for stereo, planar FoA, and FoA input signals, according to an embodiment.
[0041] Figure 5B and 5C is a flow chart of a bit rate profiling process for a spatial reconstructor (SPAR) FoA input signal, according to an embodiment.
[0042] Figure 6 is a flow chart of a bitrate profiling process for stereo, planar FoA, and FoA input signals, according to an embodiment.
[0043] Figure 7 is a flow chart of a bit rate profiling process for a SPAR FoA input signal according to an embodiment.
[0044] Figure 8 is a block diagram of an example device architecture according to an embodiment.
[0045] Like reference numbers used in the various drawings indicate like elements. DETAILED DESCRIPTION
[0046] In the following detailed description, numerous specific details are set forth to provide a thorough explanation of the various described embodiments. One of ordinary skill in the art will appreciate that the various described embodiments can be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail to avoid unnecessarily obscuring aspects of the embodiments. Several features are described below, each of which can be used independently of one another or in any combination with other features.
[0047] Nomenclature
[0048] As used herein, the term "including" and its variations should be considered open-ended terms meaning "including, but not limited to". The term "or" should be considered "and / or" unless the context clearly indicates otherwise. The term "based on" should be considered "based at least in part on". The terms "one example implementation" and "example implementation" should be considered "at least one example implementation". The term "another implementation" should be considered "at least one other implementation". The terms "determined", "determined" or "in determining" should be considered to obtain, receive, calculate, compute, estimate, predict or derive. In addition, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0049] IVAS Use Case Examples
[0050] Figure 1A use case 100 of the IVAS codec 100 is illustrated, according to one or more implementations. In some implementations, various devices communicate through a call server 102 configured to receive audio signals from, for example, a public switched telephone network (PSTN) or public land mobile network device (PLMN), illustrated by PSTN / OTHER PLMN 104. Use case 100 supports legacy devices 106 that only render and capture audio in mono, including but not limited to devices that support Enhanced Voice Service (EVS), Multi-Rate Wideband (AMR-WB), and Adaptive Multi-Rate Narrowband (AMR-NB). Use case 100 also supports user equipment (UE) 108, 114 that captures and renders stereo audio signals, or UE 110 that captures mono signals and renders them binaurally as multi-channel signals. Use case 100 also supports immersive and stereo signals captured and rendered by video conferencing room systems 116, 118, respectively. Use case 100 also supports stereo capture and immersive presentation of stereo audio signals for a home theater system 120 , and mono capture and immersive presentation of audio signals for a virtual reality (VR) device 122 and immersive content ingestion 124 by computer 112 .
[0051] Example IVAS encoding / decoding system
[0052] Figure 2 1 is a block diagram of a system 200 for encoding and decoding an IVAS bitstream, according to one or more implementations. For encoding, the IVAS encoder includes a spatial analysis and downmixing unit 202 that receives audio data 201 (including but not limited to: a mono signal, a stereo signal, a binaural signal, a spatial audio signal (e.g., a multi-channel spatial audio object), FoA, higher-order ambisonics (HoA), and any other audio data). In some implementations, the spatial analysis and downmixing unit 202 implements Complex Advanced Coupling (CACPL) for analyzing / downmixing stereo / FoA audio signals and / or SPAR for analyzing / downmixing FoA audio signals. In other implementations, the spatial analysis and downmixing unit 202 implements other formats.
[0053] The output of the spatial analysis and downmix unit 202 includes spatial metadata and 1 to N downmix channels of audio, where N is the number of input channels. The spatial metadata is input to the quantization and entropy coding unit 203, which quantizes and entropy codes the spatial data. In some implementations, quantization may include several levels of increasingly coarse quantization, such as, for example, fine, medium, coarse, and extra coarse quantization strategies, and entropy coding may include Huffman or arithmetic coding. The enhanced voice service (EVS) coding unit 206 encodes 1 to N channels of audio into one or more EVS bitstreams.
[0054] In some embodiments, the EVS coding unit 206 complies with 3GPP TS 26.445 and provides a wide range of functionality, such as enhanced quality and coding efficiency for narrowband (EVS-NB) and enhanced quality and coding efficiency for wideband (EVS-WB) voice services, enhanced quality of voice using ultra-wideband (EVS-SWB), enhanced quality of mixed content and music in conversational applications, robustness against packet loss and delay jitter, and backward compatibility with the AMR-WB codec. In some embodiments, the EVS coding unit 206 includes a pre-processing and mode selection unit that selects between a speech encoder for encoding speech signals and a perceptual encoder for encoding audio signals at a specified bit rate based on a mode / bitrate control 207. In some embodiments, the speech encoder is an improved variant of Algebraic Coded Excitation Linear Prediction (ACELP) extended with dedicated linear prediction (LP)-based modes for different speech classes. In some implementations, the audio encoder is a modified discrete cosine transform (MDCT) encoder with increased efficiency at low delay / low bit rate and is designed to perform seamless and reliable switching between speech and audio encoders.
[0055] In some implementations, the IVAS decoder includes a quantization and entropy decoding unit 204 configured to recover spatial metadata and an EVS decoder 208 configured to recover 1 to N channel audio signals. The recovered spatial metadata and audio signals are input to a spatial synthesis / rendering unit 209 that uses the spatial metadata to synthesize / render the audio signal for playback on various audio systems 210.
[0056] Example IVAS / SPAR Codec
[0057] Figure 3 is a block diagram of a FoA codec 300 for encoding and decoding FoA in SPAR format, according to some implementations. The FoA codec 300 includes a SPAR FoA encoder 301, an EVS encoder 305, a SPAR FoA decoder 306, and an EVS decoder 307. The SPAR FoA encoder 301 converts a FoA input signal into a set of downmix channels and parameters used to regenerate the input signal at the SPAR FoA decoder 306. The downmix signal can vary between one and four channels, and the parameters include prediction coefficients (PR), cross-prediction coefficients (C), and decorrelation coefficients (P). It should be noted that SPAR is a process for reconstructing an audio signal from a downmix version of the audio signal using the PR, C, and P parameters, as described in further detail below.
[0058] It should be noted that Figure 3The example implementation shown in depicts a nominal 2-channel downmix, where either W (passive prediction) or W' (active prediction) channels are sent to the decoder 306 along with a single prediction channel, Y'. In some implementations, W may be an active channel. An active W channel allows for a certain mix of the X, Y, and Z channels into the W channel as follows:
[0059] W'=W+f*pr y *Y+f*pr z *Z+f*pr x *X,
[0060] where f is a constant (e.g., 0.5) that allows some of the X, Y, and Z channels to be mixed into the W channel, and pr y 、pr x and pr z is the prediction (PR) coefficient. In passive W, f = 0, so there is no mixing of the X, Y, Z channels into the W channel.
[0061] In the case where at least one channel is sent as a residual and at least one is sent parametrically, i.e., for 2- and 3-channel downmixes, the cross-prediction coefficients (C) allow some portions of the parameterized channels to be reconstructed from the residual channels. For a two-channel downmix (described in further detail below), the C coefficients allow some of the X and Z channels to be reconstructed from Y', with the remaining channels being reconstructed from a decorrelated version of the W channel, as described in further detail below. In the case of a 3-channel downmix, Y' and X' are used to reconstruct Z alone.
[0062] In some implementations, the SPAR FoA encoder 301 includes a passive / active predictor unit 302, a remix unit 303, and an extraction / downmix selection unit 304. The passive / active predictor receives the FoA channels in a 4-channel B-format (W, Y, Z, X) and computes the downmix channels (representations of W, Y', Z', X').
[0063] The extraction / downmix selection unit 304 extracts SPAR FoA metadata from the metadata payload section of the IVAS bitstream, as described in more detail below. The passive / active predictor unit 302 and the remix unit 303 use the SPAR FoA metadata to generate remixed FoA channels (W or W' and A'), which are input to the EVS encoder 305 for encoding into an EVS bitstream, which is encapsulated in the IVAS bitstream and sent to the decoder 306. It should be noted that in this example, the Ambisonics B-format channels are arranged in the AmbiX convention. However, other conventions, such as the Furse-Malham (FuMa) convention (W, X, Y, Z), may also be used.
[0064] Referring to the SPAR FoA decoder 306, the EVS bitstream is decoded by the EVS decoder 307, thereby generating N_dmx (e.g., N_dmx=2) downmix channels. In some embodiments, the SPAR FoA decoder 306 performs the inverse of the operations performed by the SPAR encoder 301. For example, in Figure 3 In the example of , SPAR FoA spatial metadata is used to recover the remixed FoA channels (representations of W', A', B', C') from the two downmix channels. The remixed SPAR FoA channels are input to an inverse mixer 311 to recover the SPARFoA downmix channels (representations of W', Y', Z', X'). The predicted SPAR FoA channels are then input to an inverse predictor 312 to recover the original unmixed SPAR FoA channels (W, Y, Z, X). It should be noted that in this two-channel example, decorrelator blocks 309A (dec1) and 309B (dec2) are used to generate a decorrelated version of the W channel using either a time-domain or frequency-domain decorrelator. The downmix channels and decorrelated channels are used in combination with the SPAR FoA metadata to fully or parametrically reconstruct the X and Z channels. The C block 308 refers to the multiplication of the residual channel by a 2×1 C coefficient matrix, thereby generating two cross-prediction signals that are summed to form the parameterized reconstructed channel, such as Figure 3 The P1 block 310A and the P2 block 310B refer to the multiplication of the decorrelator outputs by the columns of the 2×2P coefficient matrix, thereby producing four outputs that are summed into the parameterized reconstructed channels, as shown in FIG. Figure 3 In display.
[0065] In some implementations, one of the FoA inputs is sent in its entirety to the SPARFoA decoder 306 (the W channel), and one to three of the other channels (Y, Z, and X) are sent as residuals or fully parametrically to the SPAR FoA decoder 306, depending on the number of downmix channels. The PR coefficients (which remain the same regardless of the number of downmix channels N) are used to minimize the predictable energy in the residual downmix channels. The C coefficients are used to further assist in regenerating the fully parametric channels from the residuals. Thus, the C coefficients are not needed in the one- and four-channel downmix cases, where there are no residual or parametric channels for prediction. The P coefficients are used to fill in the remaining energy not accounted for by the PR and C coefficients. The number of P coefficients depends on the number of downmix channels N in each band. In some implementations, the SPAR PR coefficients (passive W only) are calculated as follows.
[0066] Step 1. Predict all side signals (Y, Z, X) from the main W signal using equation [1].
[0067]
[0068] As an example, the prediction parameters of the predicted channel Y' are calculated using equation [2].
[0069]
[0070] where R AB = cov(A, B) is the element of the input covariance matrix corresponding to signals A and B and can be operated on for each frequency band. Similarly, the Z' and X' residual channels have corresponding prediction parameters prz and prx. PR is the prediction coefficient [pr Y ,pr Z ,pr X ] T Vector.
[0071] Step 2. Remix the W and predicted (Y', Z', X') signals from most acoustically correlated to least acoustically correlated, where "remix" means reordering or recombining the signals based on a certain methodology,
[0072]
[0073] One implementation of remixing is to reorder the input signals to W, Y', X', Z' based on the assumption that audio cues from left and right are more acoustically correlated than front-back, and front-back cues are more acoustically correlated than up-down cues.
[0074] Step 3. Calculate the covariance of the 4-channel predicted and remixed downmix as shown in equations [4] and [5].
[0075] R pr =[remix]PR.R.PR H [remix] [4]
[0076]
[0077] Where d represents the residual channels (ie, the 2nd to N_dmxth channels), and u represents the parameterized channels that need to be completely regenerated (ie, the (N_dmx+1)th to the 4th channels).
[0078] For the example of WABC downmix using 1 to 4 channels, d and u represent the following channels shown in Table 1:
[0079] Table Id and u channel representation
[0080] N D channel U channel 1 -- A′, B′, C′ 2 A′ B′, C′ 3 A′, B′ C′ 4 A′, B′, C′ --
[0081] The main focus of the calculation of the SPAR FoA metadata is the R_dd, R_ud, and R_uu quantities. From the R_dd, R_ud, and R_uu quantities, the codec 300 determines whether any remaining portion of the fully parameterized channel can be cross-predicted from the residual channel sent to the decoder. In some implementations, the required additional C coefficient is given by:
[0082] C=R ud (R dd +Imax(∈,tr(R dd )*0.005)) -1 [6]
[0083] Therefore, the C parameters have the shape (1×2) for a 3-channel downmix and (2×1) for a 2-channel downmix.
[0084] Step 4. Calculate the residual energy in the parameterized channels that must be reconstructed by the decorrelators 309A, 309B. The residual energy in the upmix channel Res_uu is the difference between the actual energy R_uu (after prediction) and the regenerated cross-prediction energy Reg_uu.
[0085] Reg uu =CR dd C H , [7]
[0086] Res uu =R uu -Reg uu [8]
[0087]
[0088] In an embodiment, in the normalized Res uu The matrix square root is taken after the matrix has had its off-diagonal elements set to zero. P is also a covariance matrix and is therefore Hermitian symmetric, and therefore only parameters from the upper or lower triangle need to be sent to the decoder 306. The diagonal entries are real numbers, while the off-diagonal elements can be complex numbers. In an embodiment, the P coefficients can be further separated into diagonal and off-diagonal elements P_d and P_o.
[0089] Example IVAS Signal Chain (FoA or Stereo Input)
[0090] Figure 4Ais a block diagram of an IVAS signal chain 400 for FoA and stereo input audio signals, according to an embodiment. In this example configuration, the audio input to signal chain 400 can be a 4-channel FoA audio signal or a 2-channel stereo audio signal. A downmix unit 401 generates downmix audio channels (dmx_ch) and spatial MDs. The downmix channels are input to a bitrate (BR) distribution unit 402, which is configured to quantize the spatial MDs using a BR distribution control table and the IVAS bitrate and provide a mono codec bitrate for the downmix audio channels, as described in detail below. The output of the BR distribution unit 402 is input to an EVS unit 403, which encodes the downmix audio channels into an EVS bitstream. The EVS bitstream and the quantized and encoded spatial MDs are input to an IVAS bitstream wrapper 404 to form an IVAS bitstream, which is transmitted to an IVAS decoder and / or stored for subsequent processing or playback on one or more IVAS devices.
[0091] For a stereo input signal, a downmix unit 401 is configured to generate an intermediate signal (M') and a representation of the residual (Re) from the stereo signal and the spatial MD. The spatial MD includes the PR, C, and P coefficients of SPAR and the PR and P coefficients of CACPL, as described more fully below. The M' signal, Re, the spatial MD, and the BR distribution control table are input to a BR (bitrate) distribution unit 402. The BR distribution unit 402 is configured to quantize the spatial metadata using the signal characteristics of the M' signal and the BR distribution control table and provide a mono codec bit rate for the downmixed channel. The M' signal, Re, and the mono codec BR are input to an EVS unit 403, which encodes the M' signal and Re into an EVS bitstream. The EVS bitstream and the quantized and encoded spatial MD are input to an IVAS bitstream wrapper 404 to form an IVAS bitstream, which is transmitted to an IVAS decoder and / or stored for subsequent processing or playback on one or more IVAS devices.
[0092] For a FoA input signal, the downmix unit 401 is configured to generate one to four FoA downmix channels, W', Y', X', and Z', and a spatial MD. The spatial MD includes the PR, C, and P coefficients of SPAR and the PR and P coefficients of CACPL, as described more fully below. The one to four FoA downmix channels (W', Y', X', and Z') are input to the BR distribution unit 402, which is configured to quantize the spatial MD using the signal characteristics of the FoA downmix channels and a BR distribution control table and provide a mono codec bit rate for the FoA downmix channels. The FoA downmix channels are input to the EVS unit 403, which encodes the FoA downmix channels into an EVS bitstream. The EVS bitstream and the quantized and coded spatial MD are input to the IVAS bitstream wrapper 404 to form an IVAS bitstream, which is transmitted to the IVAS decoder and / or stored for subsequent processing or playback on one or more IVAS devices. The IVAS decoder can perform the inverse of the operations performed by the IVAS encoder to reconstruct the input audio signal for playback on the IVAS device.
[0093] Figure 4B FIG4 is a block diagram of an alternative IVAS signal chain 405 for FoA and stereo input audio signals, according to an embodiment. In this example configuration, the audio input to the signal chain 405 can be a 4-channel FoA audio signal or a 2-channel stereo audio signal. In this embodiment, a pre-processor 406 extracts signal properties from the input audio signal, such as bandwidth (BW), speech / music classification data, voice activity detection (VAD) data, etc.
[0094] The spatial MD unit 407 generates a spatial MD from the input audio signal using the extracted signal properties. The input audio signal, signal properties, and spatial MD are input to the BR distribution unit 408, which is configured to quantize the spatial MD using the BR distribution control table and IVAS bit rate described in detail below and provide a mono codec bit rate for the downmixed audio channel.
[0095] The input audio signal, the quantized spatial MD and a number of downmix channels (d_dmx) output by the BR distribution unit 408 are input to the downmix unit 409, which generates downmix channels. For example, for FoA signals, the downmix channels may include W' and N_dmx-1 residuals (Re).
[0096] The EVS bitrate and downmix channels output by the BR distribution unit 408 are input to the EVS unit 410, which encodes the downmix channels into an EVS bitstream. The EVS bitstream and the quantized, coded spatial MD are input to the IVAS bitstream wrapper 411 to form an IVAS bitstream, which is transmitted to the IVAS decoder and / or stored for subsequent processing or playback on one or more IVAS devices. The IVAS decoder can perform the inverse of the operations performed by the IVAS encoder to reconstruct the input audio signal for playback on the IVAS device.
[0097] Instance bit rate distribution control strategy
[0098] In an embodiment, the IVAS bitrate profile control strategy includes two components. The first component is a BR profile control table that provides the initial conditions for the BR profile control process. The index to the BR profile control table is determined by the codec configuration parameters. The codec configuration parameters may include IVAS bitrate, input format (e.g., stereo, FoA, planar FoA, or any other format), audio bandwidth (BW), spatial coding mode (or number of residual channels N), and the number of channels N. re ), priority and spatial MD of mono codec. For stereo coding, N re = 0 corresponds to the full parameter (FP) mode and N re =1 corresponds to medium residual (MR) mode. In an embodiment, the BR distribution control table index points to the target, minimum and maximum mono codec bitrates for each downmix channel and multiple quantization strategies (e.g., fine, medium coarse, coarse) to encode the spatial MD. In another embodiment, the BR distribution control table index points to the overall target and minimum bitrate for all mono codec instances, the ratio in which the available bitrate needs to be divided among all downmix channels and multiple quantization strategies to encode the spatial MD. The second component of the IVAS bitrate distribution control strategy is the process of using the BR distribution control table output and input audio signal properties to determine the spatial metadata quantization level and bitrate and the bitrate for each downmix channel, as referenced. Figure 5A and 5B describe.
[0099] Bit Rate Distribution Process - Overview
[0100] The main processing components of the bitrate profile process disclosed herein include:
[0101] Audio bandwidth (BW) detection (e.g., narrowband (NB), wideband (WB), ultra-wideband (SWB), fullband (FB)). In this step, the BW of the middle or W signal is detected and the metadata is quantized accordingly. EVS then considers the IVAS BW as an upper limit and encodes the downmix channels accordingly.
[0102] Input audio signal property extraction (e.g., speech or music)
[0103] Spatial coding mode (e.g., full parametric (FP), medium residual (MR)) or number of residual channels selected N_re, where for stereo coding, when N_re=0, FP mode is selected, and when N_re=1, MR mode is selected
[0104] Mono codec and spatial MD priority decisions target bitrate, minimum and maximum bitrate per downmix channel, or the ratio of the total mono codec bitrate to be divided among the downmix channels
[0105] Audio BW detection
[0106] This component detects the BW of the intermediate or W signal.In an embodiment, the IVAS codec uses the EVS BW detector described in EVS TS 26.445.
[0107] Input signal property extraction
[0108] This component classifies each frame of the input audio signal as speech or music.In an embodiment, the IVAS codec uses the EVS speech / music classifier as described in EVS TS 26.445.
[0109] Mono codec vs. spatial MD priority decision
[0110] This component determines the priority of the mono codec over spatial MD based on the downmix signal properties. Examples of downmix signal properties include speech or music, as determined by the speech / music classifier data, and mid-side (MS) band covariance estimates for stereo, and WY, WX, and WZ band covariance estimates for FoA. If the input audio signal is music, the speech / music classifier data can be used to give higher priority to the mono codec, and when the input audio signal is hard-panned to the left or right, the covariance estimates can be used to give more priority to spatial MD.
[0111] In an embodiment, a priority decision is calculated for each frame of the input audio signal. For a given IVAS bitrate, intermediate or W signal BW, and input configuration, the bitrate profile starts with the target or desired bitrate for the downmix channel in the finest quantization strategy present in the BR profile control table and metadata (e.g., the mono codec bitrate is determined based on a subjective or objective assessment). If the initial conditions do not meet the given IVAS bitrate budget, the mono codec bitrate or quantization level of the spatial MD, or both, is iteratively reduced in the quantization loop based on their respective priorities until both meet the IVAS bitrate budget.
[0112] Bit rate distribution between downmix channels
[0113] Full parameter alignment residuals
[0114] In FP mode, only the M' or W' channel is encoded by a mono codec, and additional parameters are encoded in spatial MD, indicating the level or decorrelation of the residual channel to be added by the decoder. For bitrates where both FP and MR are feasible, the IVAS BR distribution process dynamically selects, on a frame-by-frame basis, the number of residual channels to be encoded by the mono codec and transmitted / streamed to the decoder based on spatial MD. If the level of any residual channel is above a threshold, that residual channel is encoded by the mono codec; otherwise, the process operates in FP mode. When the number of residual channels to be encoded by the mono codec changes, a transition frame process is performed to reset the codec state buffer.
[0115] MR downmix bitrate distribution
[0116] Listening evaluations were performed using various input signals and bitrate distributions between the center and residual channels. Based on concentrated listening tests, the most effective center-to-residual bitrate ratio was 3:2. However, other ratios can be used based on the application's requirements. In one embodiment, the bitrate distribution uses a fixed ratio that is further tuned during a tuning phase. During the iterative process of selecting a quantization strategy and BR for the downmix channels, the BR for each downmix channel is modified according to the given ratio.
[0117] In one embodiment, instead of maintaining a fixed ratio between downmix channel bit rates, the target bit rate, as well as the minimum and maximum bit rates, for each downmix channel are individually listed in the BR distribution control table. These bit rates are selected based on careful subjective and objective evaluations. During the iterative process of selecting a quantization strategy and BR for the downmix channels, bits are added to or taken from the downmix channels based on the priority of all downmix channels. The priority of the downmix channels can be fixed or dynamic on a frame-by-frame basis. In one embodiment, the priority of the downmix channels is fixed.
[0118] Bit Rate Distribution Process - Process Flow
[0119] Figure 5AFlowchart of a bitrate profile process 500 for stereo and FoA input signals, according to an embodiment. Inputs to process 500 are IVAS bitrate, constants (e.g., bitrate profile control table, IVAS bitrate), downmix channels, spatial MD, input format (e.g., stereo, FoA, planar FoA), and mandatory command line parameters (e.g., maximum bandwidth, encoding mode, mono downmix EVS backward compatibility mode). Outputs of process 500 are the EVS bitrate, metadata quantization level, and encoded metadata bits for each downmix channel. The following steps are performed as part of process 500.
[0120] Downmix audio feature extraction
[0121] In step 501, the following signal properties are extracted from the input audio signal: bandwidth (e.g., narrowband, wideband, ultra-wideband, fullband), voice / music classification data, and voice activity detection (VAD) data. The bandwidth (BW) is the minimum of the actual bandwidth of the input audio signal and a user-specified maximum bandwidth on the command line. In one embodiment, the downmixed audio signal may be in pulse code modulation (PCM) format.
[0122] Determine table index
[0123] In step 502, process 500 extracts an IVAS bitrate profile control table index from the IVAS bitrate profile control table using the IVAS bitrate. In step 503, process 500 determines the input format table index based on the signal parameters extracted in step 501 (i.e., BW and speech / music classification), the input audio signal format, the IVAS bitrate profile control table index extracted in step 502, and the EVS mono downmix backward compatibility mode. In step 504, process 500 selects a spatial coding mode (i.e., FP or MR) or the number of residual channels (i.e., N_re = 0 to 3) based on the bitrate profile control table index, the transition audio coding mode, and the spatial MD. In step 505, process 500 determines the final extraction table index based on the six parameters described above. In an embodiment, the selection of the spatial audio coding mode in step 504 is based on the residual channel level indicator in the spatial MD. The spatial audio coding mode indicates either MR coding mode (wherein representation of the middle or W channel (M' or W') is accompanied by one or more residual channels in the downmix audio signal) or FP coding mode (wherein only representation of the middle or W channel (M' or W') is present in the downmix audio signal). In an embodiment, if the spatial audio coding mode in the previous frame includes residual channel coding and the current frame requires only M' or W' channel coding, the transition audio coding mode is set to 1. Otherwise, the transition audio coding mode is set to 0. If the number of residual channels to be coded differs between the current frame and the previous frame, the transition audio coding mode is set to 1.
[0124] Calculate mono codec and spatial MD priority
[0125] In step 506, the process 500 determines the mono codec / spatial MD priority based on the input audio signal properties extracted in step 1 and the covariance estimates of the mid-side or WY, WX, WZ channel bands. In an embodiment, there are four possible priority outcomes: mono codec high priority and spatial MD low priority, mono codec low priority and spatial MD high priority, mono codec high priority and spatial MD high priority, and mono codec low priority and spatial MD low priority.
[0126] Extract mono codec bitrate-related variables from the table
[0127] In step 507, the following parameters are read from the table entry pointed to by the final table index calculated in step 505: mono codec (EVS) target bit rate, bit rate ratio, EVS minimum bit rate, and EVS bit rate deviation step size. Depending on the mono codec / spatial MD priority determined in step 506 and the spatial MD bit rates at various quantization levels, the actual mono codec (EVS) bit rate may be higher or lower than the mono codec (EVS) target bit rate specified in the BR distribution control table. The bit rate ratio indicates the ratio by which the total EVS bit rate must be distributed among the input audio signal channels. The EVS minimum bit rate is the value below which the total EVS bit rate is not permitted. When the EVS priority is higher than, equal to, or lower than the spatial MD priority, the EVS bit rate deviation step size is the step size by which the EVS target bit rate is reduced.
[0128] Calculates the optimal EVS bit rate and metadata quantization level based on input parameters
[0129] In step 508, an optimal EVS bit rate and metadata quantization strategy are calculated based on the input parameters obtained in steps 501 to 503 according to the following sub-steps. A high bit rate and a coarse quantization strategy for the downmix channels may lead to spatial issues, while a fine quantization strategy and a low downmix audio channel bit rate may lead to mono codec encoding artifacts. As used herein, "optimal" means the IVAS bit rate that is the most balanced distribution between the EVS bit rate and the metadata quantization level while utilizing all available bits in the IVAS bit rate budget or at least significantly reducing bit loss.
[0130] Step 508.1: Quantize the metadata using the finest quantization level and check condition 508.a (shown below). If condition 508.a is true, proceed to step 508.b (shown below). Otherwise, based on the priority calculated in step 503, continue to step 508.2 or 508.3 or 508.4.
[0131] Step 508.2: If the EVS priority is high and the spatial MD priority is low, then reduce the spatial MD quantization level and check condition 508.a. If condition 508.a is true, proceed to step 508.b. Otherwise, reduce the EVS target bit rate based on step 507 (EVS bit rate deviation step size) and check condition 508a. If condition 508a is true, proceed to step 508.b; otherwise, repeat step 508.2.
[0132] Step 508.3: If the EVS priority is low and the spatial MD priority is high, then based on step 507 (EVS bitrate deviation step size), reduce the EVS target bitrate and check condition 508.a. If condition 508.a is true, proceed to step 508.b. Otherwise, reduce the spatial MD quantization level and check condition 508.a. If condition 508.a is true, proceed to step 508.b. Otherwise, repeat step 508.3.
[0133] Step 508.4: If the EVS priority is equal to the spatial MD priority, then based on step 507 (EVS bitrate deviation step size), reduce the EVS target bitrate and check condition 508.a. If condition 508.a is true, proceed to step 508.b. Otherwise, reduce the quantization level of the spatial metadata and check condition 508.a. If condition 508.a is true, proceed to step 508.b, otherwise repeat step 5.4.
[0134] The condition 508.a mentioned above checks whether the sum of the metadata bit rate, the EVS target bit rate and the overhead bits is less than or equal to the IVAS bit rate.
[0135] Step 508.b mentioned above calculates the EVS bit rate as the IVAS bit rate minus the metadata bit rate minus the overhead bits. Next, the EVS bit rate is distributed among the downmix audio channels according to the bit rate ratios mentioned in step 507.
[0136] If the minimum EVS target bit rate and the coarsest quantization level do not meet the IVAS bit rate budget, then the bit rate profile process 500 is performed using a lower bandwidth.
[0137] In one embodiment, the table index and metadata quantization level information are included in the overhead bits of the IVAS bitstream sent to the IVAS decoder. The IVAS decoder reads the table index and metadata quantization level from the overhead bits in the IVAS bitstream and decodes the spatial MD. This leaves only the EVS bits in the IVAS bitstream for the IVAS decoder to process. The EVS bits are divided among the input audio signal channels according to the ratio indicated by the table index (step 508.b). Each EVS decoder instance is then called using the corresponding bit, resulting in the reconstruction of the downmix audio channels.
[0138] Example IVAS Bit Rate Distribution Control Table
[0139] The following is an example IVAS bit rate distribution control table. The following parameters shown in the table have the values indicated below:
[0140] Input formats: Stereo–1, Planar FoA–2, FoA-3
[0141] BW:NB–0, WB–1, SWB–2, FB-3
[0142] Spatial coding tools with permission: FP-1, MR-2
[0143] Transition mode: 1→MR to FP transition, 0→other
[0144] Mono downmix backward compatibility mode: 1 → if the center channel is 3GPP EVS compliant, 0 → otherwise
[0145] Table I - Example IVAS Bit Rate Distribution Table
[0146]
[0147]
[0148]
[0149]
[0150]
[0151]
[0152] exist Figure 5AAlso shown is the IVAS bitstream. In one embodiment, the IVAS bitstream includes a fixed-length common IVAS header (CH) 509 and a variable-length common tool header (CTH) 510. In one embodiment, the bit length of the CTH segment is calculated based on the number of entries corresponding to a given IVAS bit rate in the IVAS bit rate profile control table. The relative table index (offset from the first index of the IVAS bit rate in the table) is stored in the CTH segment. If operating in mono downmix backward compatibility mode, CTH 510 is followed by EVS payload 511, which is followed by spatial MD payload 513. If operating in IVAS mode, CTH 510 is followed by spatial MD payload 512, which is followed by EVS payload 514. In other embodiments, the order may be different.
[0153] Example Process
[0154] The example process of bitrate profiling may be performed by an IVAS codec or encoding / decoding system including one or more processors executing instructions stored on a non-transitory computer-readable storage medium.
[0155] In one embodiment, a system for encoding audio receives an audio input and metadata. The system determines, based on the audio input, the metadata, and parameters of the IVAS codec used when encoding the audio input, one or more indices into a bitrate profile control table, parameters including IVAS bitrate, input format, and mono backward compatibility mode, and one or more indices including a spatial audio coding mode and a bandwidth of the audio input.
[0156] The system performs a lookup in the bitrate distribution control table based on the IVAS bitrate, input format, spatial audio coding mode, and one or more indices. The lookup table identifies entries in the bitrate distribution control table, each of which includes a representation of an EVS target bitrate, a bitrate ratio, an EVS minimum bitrate, and an EVS bitrate deviation step size.
[0157] The system provides the identified items to a bit rate calculation process, which is programmed to determine the bit rate of the audio input (e.g., the downmix channels), the bit rate of the metadata, and the quantization level of the metadata. The system provides the bit rate of the downmix channels and at least one of the bit rate of the metadata or the quantization level of the metadata to a downstream IVAS device.
[0158] In some implementations, the system can extract properties from the audio input, including an indicator of whether the audio input is speech or music and the bandwidth of the audio input. Based on the properties, the system determines the priority between the bit rate of the downmix channels and the bit rate of the metadata. The system provides the priority to the bit rate calculation process.
[0159] In some embodiments, the system extracts one or more parameters from the spatial MD, including the level of the residual (side channel prediction error). Based on the parameters, the system determines a spatial audio coding mode indicating the need for one or more residual channels in the IVAS bitstream. The system provides the spatial audio coding mode to the bitrate calculation process.
[0160] In some implementations, the bitrate profile control table index is stored in the Common Tool Header (CTH) of the IVAS bitstream.
[0161] The system for decoding audio is configured to receive an IVAS bitstream. Based on the IVAS bitstream, the system determines the IVAS bitrate and a bitrate distribution control table index. Based on the table index, the system performs a lookup in the bitrate distribution control table and extracts the input format, spatial coding mode, mono backward compatibility mode, and one or more indices, the EVS target bitrate, and the bitrate ratio. The system extracts and decodes the downmix audio bits and spatial MD bits for each downmix channel. The system provides the extracted downmix signal bits and spatial MD bits to a downstream IVAS device. The downstream IVAS device can be an audio processing device or a storage device.
[0162] SPAR FoA bit rate distribution process
[0163] In an embodiment, the bit rate distribution process described above for a stereo input signal may also be modified and applied to SPAR FoA bit rate distribution using the SPAR FoA bit rate distribution control table shown below. Definitions of the terms contained in the table are provided below to assist the reader, followed by the SPAR FoA bit rate distribution control table.
[0164] Metadata target bits (MDtar) = IVAS_bits - header_bits - evs_target_bits (EVStar)
[0165] Metadata maximum bits (MDmax) = IVAS_bits - header_bits - evs_minimum_bits (EVSmin)
[0166] The metadata target bits should always be less than "MDmax".
[0167] Table II - Example SPAR FoA Bit Rate Distribution Control Table
[0168]
[0169] Some example operations for maximum MD bit rates (real coefficients) are shown in the table below.
[0170]
[0171] Instance metadata quantization loop:
[0172] In an embodiment, the metadata quantization loop is implemented as described below.The metadata quantization loop includes two thresholds (defined above): MDtar and MDmax.
[0173] Step 1: For each frame of the input audio signal, the MD parameters are quantized in a non-time-difference manner and encoded using an arithmetic encoder. The actual metadata bit rate (MDact) is calculated based on the MD encoding bits. If MDact is lower than MDtar, this step is considered a pass and the process exits the quantization loop and integrates the MDact bits into the IVAS bitstream. Any additional available bits (MDtar - MDact) are fed to the mono codec (EVS) encoder to increase the bit rate of the base data for the downmixed audio channels. The higher bit rate allows more information to be encoded by the mono codec, and the loss of the decoded audio output will be relatively small.
[0174] Step 2: If step 1 fails, a subset of the MD parameter values in the frame is quantized and then subtracted from the quantized MD parameter values in the previous frame and the difference quantization parameter values are encoded using an arithmetic encoder (i.e., temporal difference coding). MDact is calculated based on the MD encoding bits. If MDact is lower than MDtar, this step is considered a pass and the process exits the quantization loop and integrates the MDact bits into the IVAS bitstream. Any additional available bits (MDtar - MDact) are supplied to the mono codec (EVS) encoder to increase the bit rate of the base data for the downmixed audio channel.
[0175] Step 3: If step 2 fails, then the bit rate of the quantized MD parameters is calculated without using entropy (MDact).
[0176] Step 4: Compare the MDact bitrate values calculated in steps 1 to 3 with MDmax. If the minimum of the MDact bitrates calculated in steps 1, 2, and 3 is within MDmax, then this step is considered a pass and the process leaves the quantization loop and integrates the MD bitstream with the minimum MDact into the IVAS bitstream. If MDact is higher than MDtar, then bits (MDact - MDtar) are obtained from the mono codec (EVS) encoder.
[0177] Step 5: If step 4 fails, quantize the parameters more coarsely and repeat the above steps as the first fallback strategy (Fallback 1).
[0178] Step 6: If step 5 fails, then use the quantization scheme that guarantees compliance with MDmax to quantize the parameters as the second fallback strategy (fallback 2).
[0179] After all the iterations mentioned above, it is guaranteed that the metadata bit rate will comply with MDmax and the encoder will produce actual metadata bits or MDact.
[0180] Downmix channel / EVS bit rate distribution (EVSbd):
[0181] In an embodiment, EVS actual bits (EVSact) = IVAS_bits - header_bits - MDact. If "EVSact" is less than "EVStar", bits are taken from the EVS channels in the following order (Z, X, Y, W). The maximum number of bits that can be taken from any channel is EVStar(ch) minus EVSmin(ch). If "EVSact" is greater than "EVStar", all extra bits are assigned to the downmix channels in the following order: W, Y, X, and Z. The maximum number of extra bits that can be added to any channel is EVSmax(ch) - EVStar(ch).
[0182] SPAR decoder unpacking
[0183] In an embodiment, the SPAR decoder unpacks the IVAS bitstream as follows:
[0184] 1. Get the IVAS bit rate from the bit length and the table index from the tool header (CTH) in the IVAS bitstream
[0185] 2. Dissecting the header / metadata bits in the IVAS bitstream
[0186] 3. Parse and dequantize metadata bits.
[0187] 4. Set "EVSact" = remaining bit length
[0188] 5. Read the table entry related to EVS target, minimum and maximum bitrates and repeat the "EVSbd" step at the decoder to get the actual EVS bitrate for each channel
[0189] 6. Decode EVS channels and upmix to FoA channels
[0190] BR distribution process of SPAR FoA input audio signal
[0191] Figure 5B and 5Cis a flow chart of a bitrate profile process 515 for a SPAR FoA input signal, according to an embodiment. The process 515 begins by pre-processing 517 the FoA input (W, Y, Z, X) 516 to extract signal properties (e.g., BW, speech / music classification data, VAD data, etc.) using the IVAS bitrate. The process 515 continues by generating a spatial MD (e.g., PR, C, P coefficients) 518 and based on the residual level indicator in the spatial MD, selecting a number of residual channels to send to the IVAS decoder (520) and obtaining a BR profile control table index (521) based on the IVAS bitrate, BW, and number of downmix channels (N_dmx). In some embodiments, the P coefficient in the spatial MD may be used as a residual level indicator. The BR profile control table index is sent to the IVAS bit packer (see Figure 4A 、 4B ) for inclusion in an IVAS bitstream that may be stored and / or sent to an IVAS decoder.
[0192] The process 515 continues by reading the SPAR configuration from the row in the BR distribution control table pointed to by the table index (521). As shown in Table II above, the SPAR configuration is defined by one or more features including (but not limited to) the following: downmix string (remix), active W flag, complex spatial MD flag, spatial MD quantization strategy, EVS minimum / target / maximum bit rate, and time domain decorrelator volume reduction flag.
[0193] The process 515 continues by determining MDmax, MDtar bitrates from the IVAS bitrate, EVSmin, and EVStar bitrate values (522), as previously described above, and enters a quantization loop that includes: quantizing the spatial MD in a non-temporal difference manner using a quantization strategy; encoding the quantized spatial MD using an entropy encoder (e.g., an arithmetic encoder); and computing MDact (523). In an embodiment, the first iteration of the quantization loop uses a fine quantization strategy.
[0194] The process 515 continues by checking whether MDact is less than or equal to MDtar (524). If MDact is not less than or equal to MDtar, the MD bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream and (MDtar-MDact) bits are added to the EVStar bitrate (532) in the following order: W, Y, X, Z, N_dmx EVS bitstreams (channels) are generated and the EVS bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream as previously described. If MDact is less than or equal to MDtar, the process 515 quantizes the spatial MD in a time-difference manner using a fine quantization strategy, encodes the quantized spatial MD using an entropy encoder, and computes MDact again (525). If MDact is less than or equal to MDtar, the MD bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream and (MDtar - MDact) bits are added to the EVStar bitrate in the following order (532): W, Y, X, Z, N_dmx EVS bitstreams (channels) are generated and the EVS bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream as previously described. If MDact is greater than MDtar, the spatial MD is quantized in a non-temporal difference manner using a fine quantization strategy and entropy and base2 encoded, and the new value of MDact is calculated (527). It should be noted that the maximum number of bits that can be added to any EVS instance is equal to EVSmax - EVStar.
[0195] Process 515 again determines whether MDact is less than or equal to MDtar (528). If MDact is less than or equal to MDtar, the MD bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream and (MDtar-MDact) bits are added to the EVStar bitrate (532) in the following order: W, Y, X, Z, N_dmx EVS bitstreams (channels) are generated and the EVS bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, as previously described. If MDact is greater than MDtar, process 515 sets MDact to the minimum of the three MDact bitrates operated on in (523), (525), (527) and compares MDact to MDmax (529). If MDact is greater than MDmax (530), the quantization loop (steps 523 to 530) is repeated using a coarse quantization strategy, as previously described above.
[0196] If MDact is less than or equal to MDmax, then the MD bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, and the process 515 again determines whether MDact is less than or equal to MDtar (531). If MDact is less than or equal to MDtar, then (MDtar-MDact) bits are added to the EVStar bitrate (532) in the following order: W, Y, X, Z, N_dmx EVS bitstream (channels) are generated and the EVS bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, as previously described. If MDact is greater than MDtar, then (MDtar-MDact) bits are subtracted from the EVStar bitrate (532) in the following order: Z, X, Y, W, N_dmx EVS bitstream (channels) are generated and the EVS bits are sent to the IVAS bit packer for inclusion in the IVAS bitstream, as previously described. It should be noted that the maximum number of bits that can be subtracted from any EVS instance is equal to EVStar-EVSmin.
[0197] Example Process
[0198] Figure 6 is a flow chart of an IVAS encoding process 600 according to an embodiment. The process 600 may be used as described in reference Figure 8 The described device architecture is implemented.
[0199] Process 600 includes receiving an input audio signal (601); downmixing the input audio signal into one or more downmix channels and spatial metadata associated with the one or more channels of the input audio signal (602); reading a set of one or more bitrates for the downmix channels and a set of quantization levels for the spatial metadata from a bitrate profile control table (603); determining a combination of one or more bitrates for the downmix channels (604); determining a metadata quantization level from the set of metadata quantization levels using a bitrate profile process (605); quantizing and encoding the spatial metadata using the metadata quantization level (606); generating a downmix bitstream for the one or more downmix channels using the combination of the one or more bitrates (607); combining the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into an IVAS bitstream (608); and streaming or storing the IVAS bitstream for playback on an IVAS-capable device (609).
[0200] Figure 7 is a flow chart of an alternative IVAS encoding process 700 according to an embodiment. Process 700 may be used as described in reference Figure 8 The described device architecture is implemented.
[0201] Process 700 includes receiving an input audio signal (701); extracting properties of the input audio signal (702); computing spatial metadata for channels of the input audio signal (703); reading a set of one or more bit rates for downmix channels and a set of quantization levels for the spatial metadata from a bit rate profile control table (704); determining a combination of one or more bit rates for the downmix channels (705); determining a metadata quantization level from the set of metadata quantization levels using a bit rate profile process (706); quantizing and encoding the spatial metadata using the metadata quantization level (707); generating a downmix bitstream for one or more downmix channels using the one or more bit rates using the combination of the one or more bit rates (708); combining the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels into an IVAS bitstream (709); and streaming or storing the IVAS bitstream for playback on an IVAS-capable device (710).
[0202] Example system architecture
[0203] Figure 8 A block diagram of an example system 800 suitable for implementing example embodiments of the present disclosure is shown. The system 800 includes one or more server computers or any client devices, including but not limited to Figure 1 , such as call server 102, legacy device 106, user equipment 108, 114, conference room systems 116, 118, home theater system, VR device 122, and immersive content ingestion 124. System 800 includes any consumer device, including but not limited to: smartphones, tablet computers, wearable computers, vehicle computers, gaming consoles, perimeter systems, and kiosks.
[0204] As shown, the system 800 includes a central processing unit (CPU) 801 capable of executing various processes according to a program stored in, for example, a read-only memory (ROM) 802 or a program loaded from, for example, a storage unit 808 to a random access memory (RAM) 803. Data required when the CPU 801 executes various processes is also stored in the RAM 803 as needed. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0205] The following components are connected to the I / O interface 805: an input unit 806, which may include a keyboard, a mouse, or the like; an output unit 807, which may include a display (such as a liquid crystal display (LCD)) and one or more speakers; a storage unit 808, which includes a hard disk or another suitable storage device; and a communication unit 809, which includes a network interface card, such as a network card (e.g., wired or wireless).
[0206] In some implementations, input unit 806 includes one or more microphones in different locations (depending on the host device) that enable capture of audio signals in various formats (eg, mono, stereo, spatial, immersive, and other suitable formats).
[0207] In some embodiments, the output unit 807 includes a system with various numbers of speakers. Figure 1 As illustrated in FIG, output unit 807 (depending on the capabilities of the host device) can present the audio signal in various formats (eg, mono, stereo, immersive, binaural, and other suitable formats).
[0208] The communication unit 809 is configured to communicate with other devices (e.g., via a network). The drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811 (e.g., a magnetic disk, an optical disk, a magneto-optical disk, a flash drive, or another suitable removable medium) is installed on the drive 810 so that a computer program read therefrom is installed in the storage unit 808 as needed. Those skilled in the art will understand that although the system 800 is described as including the above components, in actual applications, some of these components may be added, removed, and / or replaced, and all such modifications or changes are within the scope of the present disclosure.
[0209] According to example embodiments of the present disclosure, the processes described above may be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product comprising a computer program embodied on a machine-readable medium, the computer program comprising program code for executing the method. In such embodiments, the computer program may be downloaded and installed from a network via the communication unit 809 and / or installed from a removable medium 811, such as Figure 8 In display.
[0210] In general, the various example embodiments of the present disclosure may be implemented as hardware or dedicated circuits (e.g., control circuitry), software, logic, or any combination thereof. For example, the units discussed above may be implemented by control circuitry (e.g., Figure 8 The control circuitry may be executed by a CPU (combined with other components) and, therefore, the control circuitry may perform the actions described in the present disclosure. Some aspects may be implemented as hardware, while other aspects may be implemented as firmware or software that may be executed by a controller, microprocessor, or other computing device (e.g., control circuitry). Although various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flow charts, or using some other illustration, it should be understood that, as non-limiting examples, the blocks, devices, systems, techniques, or methods described herein may be implemented as hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0211] In addition, the various blocks shown in the flowcharts may be viewed as method steps and / or as operations resulting from the operation of computer program code and / or as a plurality of coupled logic circuit elements constructed to perform the associated functions. For example, embodiments of the present disclosure include a computer program product comprising a computer program embodied on a machine-readable medium, the computer program containing program code configured to perform the method as described above.
[0212] In the context of this disclosure, a machine-readable medium may be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0213] The computer program code for implementing the disclosed method can be written in any combination of one or more process design languages. These computer program codes can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment with a control circuit system so that the program code causes the function / operation specified in the flow chart and / or block diagram to be implemented when the processor of the computer or other programmable data processing equipment is executed. The program code can be executed entirely on a computer, partially on a computer, as an independent packaged software, partially on a computer and partially on a remote computer, or entirely on a remote computer or server or distributed on one or more remote computers and / or servers.
[0214] Although this document contains many specific implementation details, these should not be understood as limitations on the scope of what can be claimed, but rather as descriptions of features unique to particular embodiments. Specific features described in this specification in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable subcombination in multiple embodiments. In addition, although features are described above as acting in a specific combination and even initially claimed as such, in some cases, one or more features from the claimed combination may be removed from the combination and the claimed combination may be a variation on a subcombination or subcombination. The logical flow depicted in the figure does not require the specific order or sequential order shown to achieve the desired result. In addition, other steps may be provided, or steps may be eliminated from the process, and other components may be added to or removed from the system. Therefore, other embodiments are within the scope of the appended claims.
Claims
1. A method for encoding an Immersive Voice and Audio Service (IVAS) bitstream, the method comprising: using one or more processors to receive an input audio signal; extracting properties of the input audio signal using the one or more processors; calculating, using the one or more processors, spatial metadata for channels of the input audio signal; obtaining, using the one or more processors, from a bitrate profile control table a set of one or more bitrates for downmix channels and a set of quantization levels for the spatial metadata; determining, using the one or more processors, a combination of the one or more bit rates for the downmix channels; determining, using the one or more processors, a metadata quantization level from the set of metadata quantization levels; quantizing and encoding the spatial metadata using the metadata quantization level using the one or more processors; generating, using the combination of the one or more processors and the one or more bit rates, a downmix bitstream for the one or more downmix channels using the one or more bit rates; and Using the one or more processors, the downmix bitstream, the quantized and encoded spatial metadata, and the set of quantization levels are combined into the IVAS bitstream.
2. A system comprising: one or more processors; and A non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operations of method claim 1.
3. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of method claim 1.