Method and system for coding metadata in audio streams and for flexible intra- and inter-object bitrate adaptation

The framework optimizes bitrate allocation and metadata coding for object-based audio formats, enhancing efficiency and robustness in immersive audio experiences by managing bitrates across multiple audio objects.

JP7739255B2Active Publication Date: 2025-09-16VOICEAGE CORPORATION
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022500960
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-08
Filing Date
2020-07-07
Publication Date
2025-09-16
Estimated Expiration
2040-07-07

AI Technical Summary

Technical Problem

Existing audio coding technologies struggle to efficiently manage bitrate allocation and metadata coding for object-based audio formats, leading to inefficiencies and reduced robustness in immersive audio experiences.

Method used

A framework for coding object-based audio signals that includes metadata processing and bitrate adaptation, utilizing intra-object and inter-object bitrate allocation strategies to optimize bit distribution among multiple audio objects, ensuring efficient use of bitrates and robustness in noisy channels.

Benefits of technology

The framework enhances bitrate efficiency and robustness in decoding object-based audio, providing flexible and adaptive bitrate management for immersive audio experiences, minimizing errors and maintaining audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007739255000011
    Figure 0007739255000011
  • Figure 0007739255000012
    Figure 0007739255000012
  • Figure 0007739255000013
    Figure 0007739255000013
Patent Text Reader

Abstract

A system and method codes an object-based audio signal that includes audio objects according to an audio stream with associated metadata. In the system and method, an audio stream processor analyzes the audio stream. A metadata processor responds to information about the audio stream from the analysis by the audio stream processor to code the metadata. The metadata processor uses logic to control the bit budget of the coding of the metadata. An encoder codes the audio stream.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to audio coding, and more particularly to techniques for digitally coding object-based audio, such as human voice, music, or general audio. In particular, this disclosure relates to systems and methods for coding and decoding object-based audio signals that include audio objects according to an audio stream with associated metadata.

[0002] In this disclosure and the accompanying claims:

[0003] (a) The term "object-based audio" is intended to represent a complex audio auditory scene as a collection of individual elements, also known as audio objects. Also, as indicated herein above, "object-based audio" may include, for example, human voice, music, or general audio sounds.

[0004] (b) The term "audio object" is intended to refer to an audio stream that has associated metadata. For example, in this disclosure, an "audio object" is referred to as an independent audio stream with metadata (ISm).

[0005] (c) The term "audio stream" is intended to represent an audio waveform, such as human voice, music, or general audio, in a bitstream, which may consist of one channel (mono), although two channels (stereo) may also be considered. "Mono" is short for "monophonic" and "stereo" is short for "stereophonic."

[0006] (d) The term "metadata" is intended to refer to a set of information describing an audio stream and artistic intent used to convey original or coded audio objects to a playback system. Typically, the metadata describes spatial characteristics of each individual audio object, such as position, orientation, volume, width, etc. In the context of this disclosure, two sets of metadata are considered: Input Metadata: Unquantized metadata representation used as input for the codec. This disclosure is not limited to input metadata in any particular format; and - Coded Metadata: The quantized and coded metadata that forms part of the bitstream sent from the encoder to the decoder.

[0007] (e) The term "audio format" is intended to refer to a technique for achieving an immersive audio experience.

[0008] (f) The term "playback system" is intended to refer, for example, but not limited to, to an element in a decoder that, on the playback side, can use the transmitted metadata and artistic intent to render audio objects in a 3D (three-dimensional) audio space around the listener. Rendering can be performed for a target loudspeaker layout (e.g., 5.1 surround) or headphones, while metadata can be dynamically modified, for example, in response to head-tracking device feedback. Other types of rendering can be envisioned. [Background technology]

[0009] In recent years, audio generation, recording, representation, coding, transmission, and playback have been moving toward enhanced, interactive, and immersive experiences for listeners. An immersive experience can be described, for example, as a state of deep involvement and participation in an audio scene, with sounds coming from all directions. In immersive audio (also called 3D audio), a sound image is reproduced in all three dimensions around the listener, taking into account a wide range of sound characteristics, such as timbre, directionality, reverberation, clarity, and accuracy of (auditory) spaciousness. Immersive audio is generated for a given playback system, i.e., a loudspeaker configuration, an integrated playback system (sound bar), or headphones. The interactivity of an audio playback system can then include, for example, the ability to adjust audio levels, change audio position, or select different languages ​​for playback.

[0010] There are three basic approaches (hereafter also referred to as audio formats) to achieving an immersive audio experience.

[0011] The first approach is channel-based audio, where multiple spaced microphones are used to capture sound from different directions, while one microphone corresponds to one audio channel in a particular loudspeaker layout. Each recorded channel is fed to a loudspeaker in a particular position. Examples of channel-based audio include, for example, stereo, 5.1 surround, 5.1.4, etc.

[0012] The second approach is scene-based audio, which represents a desired sound field in a localized space as a function of time using a combination of dimensional components. While the signal representing scene-based audio is independent of the position of the sound source, the sound field must be transformed into a selected loudspeaker layout in a rendering playback system. An example of scene-based audio is Ambisonics.

[0013] A third and final approach to immersive audio is object-based audio, which represents an auditory scene as a set of individual audio elements (e.g., singer, drums, guitar) along with information about their positions within the audio scene so that they can be rendered in their intended positions in the playback system. This gives object-based audio great flexibility and interactivity, as each object is kept separate and can be manipulated individually.

[0014] Each of the above audio formats has its own advantages and disadvantages. Therefore, it is common that not only one particular format is used in an audio system, but they may also be combined into a complex audio system to create an immersive auditory scene. An example could be a system that combines scene-based or channel-based audio with object-based audio, for example, combining Ambisonics with several separate audio objects. [Prior art documents] [Patent documents]

[0015] [Patent Document 1] PCT patent application PCT / CA2018 / 51175 [Non-patent literature]

[0016] [Non-Patent Document 1] 3GPP Specification TS 26.445: "Codec for Enhanced Voice Services (EVS). Detailed Algorithmic Description", v.12.0.0, September 2014 Summary of the Invention [Problem to be solved by the invention]

[0017] This disclosure presents, in the following description, a framework for encoding and decoding object-based audio. Such a framework can be a stand-alone system for coding object-based audio formats, or it can form part of a composite immersive codec that may also include coding of other audio formats and / or combinations thereof. [Means for solving the problem]

[0018] According to a first aspect, the present disclosure provides a system for coding an object-based audio signal including audio objects in response to an audio stream having associated metadata, the system comprising: an audio stream processor for analyzing the audio stream; a metadata processor responsive to information about the audio stream from the analysis by the audio stream processor for coding the metadata, the metadata processor using logic for controlling a bit budget for coding the metadata; and an encoder for coding the audio stream.

[0019] The present disclosure also provides a method for coding an object-based audio signal including audio objects according to an audio stream having associated metadata, the method including the steps of analyzing the audio stream, coding the metadata using (a) information about the audio stream from the analysis of the audio stream and (b) logic for controlling a bit budget for coding the metadata, and encoding the audio stream.

[0020] According to a third aspect, there is provided an encoder device for coding complex audio auditory scenes including scene-based audio signals, multi-channel signals and object-based audio signals, the encoder device comprising the system defined above for coding object-based audio signals.

[0021] The present disclosure further provides an encoding method for coding complex audio auditory scenes, including scene-based audio signals, multi-channel signals, and object-based audio signals, including the above-mentioned method for coding object-based audio signals.

[0022] The above and other objects, advantages, and features of systems and methods for coding object-based audio signals and systems and methods for decoding object-based audio signals will become more apparent upon reading the following non-limiting description of exemplary embodiments thereof, given by way of example only with reference to the accompanying drawings. [Brief explanation of the drawings]

[0023] [Figure 1] 1 is a schematic block diagram illustrating simultaneously a system for coding an object-based audio signal and a corresponding method for coding an object-based audio signal; [Figure 2] 1A-1C illustrate different scenarios for coding of a bitstream of one metadata parameter. [Figure 3a] 1 is a graph showing the values ​​of absolute coding flags flagabs for metadata parameters of three audio objects when no inter-object metadata coding logic is used, where the arrows indicate frames where the value of some absolute coding flags is equal to 1. [Figure 3b]10 is a graph showing the values ​​of absolute coding flags flagabs for metadata parameters of three audio objects when using inter-object metadata coding logic; [Figure 4] 1 is a graph showing an example of bitrate adaptation for three core-encoders; [Figure 5] 10 is a graph showing an example of bitrate adaptation based on ISm (Independent Audio Streams with Metadata) importance logic. [Figure 6] 8 is a schematic diagram illustrating the structure of a bitstream transmitted from the coding system of FIG. 1 to the decoding system of FIG. 7. [Figure 7] 1 is a schematic block diagram illustrating simultaneously a system for decoding audio objects according to an audio stream with associated metadata and a corresponding method for decoding the audio objects; [Figure 8] FIG. 1 is a simplified block diagram of an exemplary configuration of hardware components for implementing a system and method for coding an object-based audio signal and a system and method for decoding an object-based audio signal. DETAILED DESCRIPTION OF THE INVENTION

[0024] This disclosure provides example mechanisms for coding metadata. This disclosure also provides mechanisms for flexible intra-object and inter-object bitrate adaptation, i.e., mechanisms that distribute the available bitrate as efficiently as possible. This disclosure also considers that the bitrate is fixed (constant). However, it is within the scope of this disclosure to also consider adaptive bitrates, for example, (a) in adaptive bitrate-based codecs or (b) as a result of coding a combination of audio formats otherwise coded at a fixed total bitrate.

[0025] In this disclosure, there is no description of how the audio streams are actually coded in the so-called "core encoder." In general, the core encoder for coding one audio stream can be any mono codec that uses adaptive bit rate coding. An example is a codec based on the EVS codec described in Reference [1], which uses a variable bit budget that is flexibly and efficiently distributed among the modules of the core encoder as described in Reference [2]. The entire contents of References [1] and [2] are incorporated herein by reference.

[0026] 1. A framework for coding audio objects As a non-limiting example, this disclosure considers a framework that supports simultaneous coding of several audio objects (e.g., up to 16 audio objects) while considering a fixed, constant ISm total bitrate, called ism_total_rate, for coding audio objects including audio streams with associated metadata. Note that, for example, in the case of non-diegetic content, metadata is not necessarily transmitted for at least some of the audio objects. Extra-world audio in movies, television programs, and other videos is audio that cannot be heard by the characters. A soundtrack is an example of extra-world audio because only the viewer hears the music.

[0027] When coding a combination of audio formats in the framework, for example the combination of an Ambisonics audio format with two audio objects, a constant total codec bitrate called codec_total_brate represents the sum of the bitrate of the Ambisonics audio format (i.e. the bitrate for encoding the Ambisonics audio format) and the total bitrate of ISm, ism_total_brate (i.e. the sum of the bitrates for coding the audio objects, i.e. the audio streams with associated metadata).

[0028] This disclosure considers a basic, non-limiting example of input metadata consisting of two parameters stored per audio frame for each object: azimuth and elevation. In this example, the azimuth range [-180°, 180°] and the elevation range [-90°, 90°] are considered. However, it is within the scope of this disclosure to consider only one or more than two metadata parameters.

[0029] 2. Object-Based Coding FIG. 1 is a schematic block diagram illustrating simultaneously a system 100 for coding an object-based audio signal, including several processing blocks, and a corresponding method 150 for coding an object-based audio signal.

[0030] 2.1 Input Buffering 1, a method 150 for coding an object-based audio signal includes an input buffering operation 151. To perform the input buffering operation 151, the system 100 for coding an object-based audio signal includes an input buffer 101.

[0031] The input buffer 101 buffers N input audio objects 102, i.e., N audio streams with their respective N associated pieces of metadata. The N input audio objects 102, comprising the N audio streams and the N pieces of metadata associated with each of these N audio streams, are buffered for one frame, for example a frame of length 20 ms. As is well known in the field of audio signal processing, audio signals are sampled at a given sampling frequency and processed in successive blocks of these samples, called "frames", each of which is divided into a number of "subframes".

[0032] 2.2 Audio stream analysis and front pre-processing 1, the method 150 for coding an object-based audio signal includes an operation of analyzing and pre-processing N audio streams 153. To perform operation 153, the system 100 for coding an object-based audio signal includes audio stream processors 103 for analyzing and pre-processing, e.g., in parallel, N buffered audio streams transmitted from an input buffer 101 to the audio stream processors 103 via N transport channels 104, respectively.

[0033] The analysis and forward pre-processing operations 153 performed by the audio stream processor 103 may include, for example, at least one of the following sub-operations: time-domain transient detection, spectral analysis, long-term prediction analysis, pitch tracking and voicing analysis, voice activity detection / sound activity detection (VAD / SAD), bandwidth detection, noise estimation, and signal classification (which, in a non-limiting embodiment, may include: (a) selection of a core encoder, for example, from an ACELP core encoder, a TCX core encoder, an HQ core encoder, etc.; (b) classification of signal type, such as inactive core encoder type, unvoiced core encoder type, voiced core encoder type, generic core encoder type, transition core encoder type, and audio core encoder type; (c) human voice / music classification, etc.). Information obtained from the analysis and forward pre-processing operations 153 is provided to the configuration and decision processor 106 via line 121. Examples of the above sub-operations are described in reference [1] in connection with the EVS codec and therefore will not be further described in this disclosure.

[0034] 2.3 Metadata Analysis, Quantization, and Coding 1 includes a metadata analysis, quantization, and coding operation 155. To perform operation 155, the system 100 for coding an object-based audio signal includes a metadata processor 105.

[0035] 2.3.1 Metadata Analysis Signal classification information 120 from the audio stream processor 103 (e.g., the VAD or localVAD flag used in the EVS codec (see reference [1])) is provided to the metadata processor 105. The metadata processor 105 includes an analyzer (not shown) of the metadata for each of the N audio objects to determine whether the current frame is inactive (e.g., VAD = 0) or active (e.g., VAC ≠ 0) for this particular audio object. In an inactive frame, no metadata is coded by the metadata processor 105 in association with that object. In an active frame, metadata is quantized and coded for this audio object using a variable bit rate. Further details on metadata quantization and coding are given in sections 2.3.2 and 2.3.3 below.

[0036] 2.3.2 Metadata Quantization The metadata processor 105 of FIG. 1, in the described non-limiting exemplary embodiment, quantizes and codes the metadata of N audio objects in turn in a loop, while certain dependencies may be used between the quantization of the audio objects and the metadata parameters of these audio objects.

[0037] As indicated herein above, in this disclosure, two metadata parameters are considered: azimuth angle and elevation angle (contained in the N input metadata). By way of non-limiting example, the metadata processor 105 includes a quantizer (not shown) for the indexes of the following metadata parameters using the following exemplary resolutions to reduce the number of bits used: - Azimuth parameter: 12-bit azimuth parameter from the input metadata file, index B az The bit index (for example, B az= 7). Given the minimum and maximum azimuth limits (-180° and +180°), (B az = 7)-bit uniform scalar quantizer has a quantization step of 2.835°. - Elevation parameter: The index of the 12-bit elevation parameter from the input metadata file is B. el The bit index (for example, B el = 6). Given the minimum and maximum elevation limits (-90° and +90°), (B el = 6)-bit uniform scalar quantizer has a quantization step of 2.857°.

[0038] The total metadata bit budget for coding the N metadata items and the total number of quantization bits for quantizing the metadata parameter indices (i.e., the granularity and therefore the resolution of the quantization indexes) may be made dependent on the bit rates codec_total_brate, ism_total_brate, and / or element_brate (the last resulting from the sum of the metadata bit budgets and / or core encoder bit budgets associated with one audio object).

[0039] The azimuth and elevation parameters may be represented as one parameter, for example, by a point on a sphere. In such cases, it is within the scope of this disclosure to implement different metadata that includes two or more parameters.

[0040] 2.3.3 Coding Metadata Once quantized, both the azimuth index and the elevation index can be coded by a metadata encoder (not shown) of the metadata processor 105 using either absolute coding or differential coding. As is known, absolute coding means that the current value of the parameter is coded. Differential coding means that the difference between the current value and the previous value of the parameter is coded. Since the indices of the azimuth and elevation parameters usually evolve smoothly (i.e., the change in the azimuth or elevation position can be considered continuous and smooth), differential coding is used by default. However, absolute coding may be used, for example, in the following cases: - The difference between the current and previous values ​​of the parameter index is too large, which results in a greater or equal number of bits for using differential coding compared to using absolute coding (this can occur exceptionally). - No metadata was coded and transmitted in the previous frame. - Too many consecutive frames using differential coding, to control decoding in noisy channels (Bad Frame Indicator, BFI = 1). For example, the metadata encoder codes the index of the metadata parameter using absolute coding if the number of consecutive frames coded using differential coding exceeds the maximum number of consecutive frames coded using differential coding. The maximum number of consecutive frames is set to β. In a non-limiting illustrative example, β = 10 frames.

[0041] The metadata encoder uses a 1-bit absolute coding flag to distinguish between absolute and differential coding. abs Generate.

[0042] In the case of absolute coding, the coding flag abs is set to 1, and then B is coded using absolute coding. az Bit (or B el followed by the index of the B az and B el and refer to the above-mentioned indexes of the azimuth and elevation parameters to be coded, respectively.

[0043] In the case of differential coding, a 1-bit coding flag abs is set to 0, then B in the current and previous frames is equal to 0. az The bit index (or B el A 1-bit zero-coding flag signaling the difference Δ between the zero If the difference Δ is not equal to 0, the metadata encoder may use a 1-bit code flag flag followed by a difference index, the number of bits of which is adaptive, e.g., in the form of a unary code indicating the value of the difference Δ. sign The coding continues by generating

[0044] FIG. 2 illustrates different scenarios for coding one metadata parameter in the bitstream.

[0045] With reference to Figure 2, it is noted that not all metadata parameters are always transmitted in every frame. Some may only be transmitted every y frames, and some may not be transmitted at all, for example, when they do not evolve, or when they are not important, or when the available bit budget is low.

[0046] - In the case of absolute coding (first line of Figure 2), the absolute coding flag abs and B az The bit index (or B el The bit index) is transmitted.

[0047] - B in the current and previous frames az The bit index (or B el In the case of differential coding where the difference Δ between the two (the index of the bit) is equal to 0 (line 2 of Figure 2), the absolute coding flag abs =0 and zero coding flag zero =1 is sent.

[0048] - B in the current and previous frames az The bit index (or B el In the case of differential coding where there is a positive difference Δ between the bit indices (line 3 of Figure 2), the absolute coding flag abs =0, zero coding flag zero =0, sign flag sign = 0, and the difference index (1 to (B az -3) bit index (or 1 to (B el -3) The index of the bit)) is transmitted, and

[0049] - B in the current and previous frames az The bit index (or B el In the case of differential coding where there is a negative difference Δ between the two (the index of the bit) (last line of Figure 2), the absolute coding flag abs =0, zero coding flag zero =0, sign flag sign =1, and the difference index (1 to (B az -3) bit index (or 1 to (B el -3) The index of the bit)) is transmitted.

[0050] 2.3.3.1 Coding logic for metadata within objects The logic used to set up absolute or differential coding may be further extended by metadata coding logic within the object. In particular, to limit the amount of variation in the metadata coding bit budget between frames, and thus prevent the core encoder 109 from being left with too little bit budget, the metadata encoder limits absolute coding in a given frame to one, or generally to as few metadata parameters as possible.

[0051] In a non-limiting example of coding azimuth and elevation metadata parameters, the metadata encoder uses logic to avoid absolute coding of an elevation index in a given frame if the azimuth index has already been coded using absolute coding in the same frame. In other words, the azimuth and elevation parameters of an audio object are (substantially) never both coded using absolute coding in the same frame. As a result, the absolute coding flag for the azimuth parameter is set to abs.azi If is equal to 1, the absolute coding flag for the elevation parameter abs.ele is not transmitted in the bitstream of the audio object.

[0052] It is also within the scope of this disclosure to make the coding logic of the metadata in the object dependent on the bit rate. For example, if the bit rate is high enough, then the absolute coding flag for the elevation angle parameter is set to 0. abs.ele and the absolute coding flag for the azimuth parameter abs.azi Both may be transmitted in the same frame.

[0053] 2.3.3.2 Inter-object metadata coding logic The metadata encoder may apply similar logic to the coding of metadata of different audio objects. The implemented inter-object metadata coding logic minimizes the number of metadata parameters of different audio objects that are coded using absolute coding in the current frame. This is achieved by the metadata encoder mainly by controlling a frame counter of the metadata parameters that are coded using absolute coding, which is chosen for robustness purposes and is represented by the parameter β. As a non-limiting example, a scenario in which the metadata parameters of the audio objects evolve slowly and smoothly is considered. To control decoding in a noisy channel, where the index is coded using absolute coding every β frames, the B of the azimuth angle of audio object #1 is az The bit index is coded using absolute coding in frame M, and the elevation angle B of audio object #1 is el The bit index is coded using absolute coding in frame M+1, and the azimuth angle B of audio object #2 is az The bit index is coded using absolute coding in frame M+2, and the elevation angle B of object #2 is el The bit index is coded using absolute coding in frame M+3, and so on.

[0054] Figure 3a shows the absolute coding flags for the metadata parameters of three audio objects when no inter-object metadata coding logic is used. abs 3b is a graph showing the values ​​of the absolute coding flag , for the metadata parameters of the three audio objects when using the inter-object metadata coding logic. abs 3a is a graph showing the values ​​of .times. ...

[0055] More specifically, FIG. 3a shows the absolute coding flags for two metadata parameters of an audio object (azimuth and elevation in this particular example) when no inter-object metadata coding is used. abs 3a and 3b show the same values, but with the coding logic of metadata between objects implemented. The graphs in 3a and 3b correspond (from top to bottom) to: - the audio stream of audio object #1, - Audio object #2's audio stream, - Audio object #3's audio stream, - Absolute coding flag for the azimuth parameter of audio object #1 abs,azi , - Absolute coding flag for the elevation parameter of audio object #1 abs,ele , - Absolute coding flag for the azimuth parameter of audio object #2 abs,azi , - Absolute coding flag for the elevation parameter of audio object #2 abs,ele , - Absolute coding flag for the azimuth parameter of audio object #3 abs,azi , and - Absolute coding flag for the elevation parameter of audio object #3 abs,ele .

[0056] When the inter-object metadata coding logic is not used, several flags in the same frame abs It can be seen from Fig. 3a that there are cases where , has a value equal to 1 (see arrow). In contrast, Fig. 3b shows that when the inter-object metadata coding logic is used, there is only one absolute flag, flag absindicates that only one may have a value equal to 1.

[0057] The coding logic for inter-object metadata may also be made bitrate dependent, in which case, for example, if the bitrate is large enough, it may be possible to have more than one absolute flag in a given frame even when the coding logic for inter-object metadata is used. abs may have a value equal to 1.

[0058] A technical advantage of the inter-object metadata coding logic and intra-object metadata coding is that it limits the range of variation in the bit budget of inter-frame metadata coding. Another technical advantage is that it increases the robustness of the codec in noisy channels: when a frame is lost, only a limited number of metadata parameters from an audio object coded using absolute coding are lost. Therefore, any error propagated from a lost frame only affects a small number of metadata parameters of the entire audio object, and therefore does not affect the entire audio scene (or several different channels).

[0059] The overall technical advantage of analyzing, quantizing and coding the metadata separately from the audio stream is that it is specifically adapted to the metadata, as described above, and allows for more efficient processing in terms of metadata coding bit rate, metadata coding bit budget variations, robustness in noisy channels, and propagation of errors due to lost frames.

[0060] The quantized and coded metadata 112 from the metadata processor 105 is fed to a multiplexer 110 for insertion into an output bitstream 111 that is sent to a distant decoder 700 (FIG. 7).

[0061] Once the metadata of the N audio objects has been analyzed, quantized and encoded, information 107 from the metadata processor 105 about the bit budget for coding the metadata for each audio object is provided to a configuration and decision processor 106 (bit budget allocator) which will be described in more detail in the following section 2.4. Once the configuration and bit rate distribution among the audio streams is completed in processor 106 (bit budget allocator), coding continues with further pre-processing 158 which will be described below. Finally, the N audio streams are encoded using an encoder comprising N variable bit rate core encoders 109, e.g., mono core encoders.

[0062] 2.4 Configuring and Determining Bit Rates Per Channel 1 for coding an object-based audio signal includes an operation 156 of configuring and determining a bit rate per transport channel 104. To perform operation 156, the system 100 for coding an object-based audio signal includes a configuration and determination processor 106 that forms a bit budget allocator.

[0063] The configuration and decision processor 106 (hereinafter bit budget allocator 106) uses a bit rate adaptation algorithm to allocate the available bit budget for core-encoding the N audio streams over the N transport channels 104.

[0064] The bit rate adaptation algorithm of the configuration and decision operation 156 includes the following sub-operations 1-6 performed by the bit budget allocator 106:

[0065] 1. Total bit budget of ISm per frame (bits) ismis calculated from the total bitrate of ISm_total_brate (or the total bitrate of the codec, codec_total_brate, if only audio objects are coded), for example using the following relation:

[0066]

number

[0067] The denominator 50 corresponds to the number of frames per second, assuming a frame length of 20ms. If the frame size is different from 20ms, the value 50 will be different.

[0068] 2. The element bitrate element_rate defined above for N audio objects (resulting from the sum of the bit budget of the metadata associated with one audio object and the bit budget of the core encoder) is assumed to be constant during a session for the total bit rate of a given codec and approximately the same for the N audio objects. A "session" is defined as, for example, a telephone call or offline compression of an audio file. The corresponding element bit budget bits element For example, for the audio stream objects n = 0, ..., N-1, the following relation

[0069]

number

[0070] is calculated using the formula:

[0071]

number

[0072] denotes the largest integer less than or equal to x. The total bit budget of ISm available is bitsism For example, to use all of the bit budget bits of the last audio object element, element But finally, the following relation

[0073]

number

[0074] where "mod" indicates a modulo operation. Finally, the bit budget of the elements of the N audio objects is element For example, the value element_brate for audio object n = 0, ..., N-1 is calculated using the following relation: element_brate[n] = bits element [n]*50 where the number 50 corresponds to the number of frames per second, assuming 20 ms long frames as described above.

[0075] 3. Bit budget for metadata per frame for N audio objects (bits) meta But the following relation

[0076]

number

[0077] The resulting value is summed using bits meta_all However, the bit budget for ISm common signaling is Ism_signalling and the codec's side bit-budget bits side = bits meta_all + bits ISm_signalling results.

[0078] 4. Codec side bit budget bits per frame side is divided evenly among the N audio objects, and the core encoder bit budget bits for each of the N audio streams is CoreCoder For example, the following relation

[0079]

number

[0080] while the core encoder bit budget for the final audio stream is calculated using, for example, the following relation to finally use all the available core encoding bit budget:

[0081]

number

[0082] Then, the corresponding total bitrate, i.e., the bitrate for coding one audio stream in the core encoder, may be adjusted using, for example, the following relation: total_brate[n] = bits CoreCoder [n]*50 where the number 50 corresponds to the number of frames per second, again assuming frames of 20 ms length.

[0083] 5. The total bit rate, total_brate, in inactive frames (or frames with very low energy or otherwise no meaningful content) may be reduced and set to a constant value in the associated audio stream. The bit budget thus saved is then redistributed evenly among the audio streams with active content within the frame. Such redistribution of the bit budget is further described in Section 2.4.1 below.

[0084] 6. The total bitrate, total_brate, among audio streams (with active content) within an active frame is further adjusted between these audio streams based on the importance classification of ISm. Such bitrate adjustment is further described in Section 2.4.2 below.

[0085] When the audio streams are all in inactive segments (or have no meaningful content), the last two sub-operations 5 and 6 above may be omitted. Thus, the bitrate adaptation algorithm described in sections 2.4.1 and 2.4.2 below is used when at least one audio stream has active content.

[0086] 2.4.1 Bitrate Adaptation Based on Signal Activity In inactive frames (VAD = 0), the total bit rate total_rate is reduced and the saved bit budget is redistributed, for example evenly, among the audio streams in active frames (VAD ≠ 0). As a prerequisite, waveform coding of audio streams in frames classified as inactive is not required and audio objects may be muted. The logic used in every frame can be expressed by the following sub-operations 1 to 3:

[0087] 1. For a particular frame, set a smaller core encoder bit budget for any audio stream n that has inactive content; bits CoreCoder '[n] = B VAD0 ∀n where VAD=0 In the formula, B VAD0 is a lower constant core encoder bit budget set in inactive frames, e.g., B VAD0 = 140 (equivalent to 7 kbps for a 20 ms frame) or B VAD0 = 49 (corresponding to 2.45 kbps for a 20 ms frame).

[0088] 2. The saved bit budget is then calculated using, for example, the following relation:

[0089]

number

[0090] It is calculated using

[0091] 3. Finally, the saved bit budget is calculated, for example, between the core encoder bit budget of an audio stream with active content in a given frame, using the following relationship:

[0092]

number

[0093] are evenly redistributed using the formula, where N VAD1 is the number of audio streams with active content, and the core encoder bit budget for the first audio stream with active content is, for example,

[0094]

number

[0095] Finally, the total bitrate of the corresponding core encoder, total_brate, is obtained for each audio stream n = 0, ..., N-1 as follows: total_brate'[n] = bits CoreCoder '[n]*50

[0096] Figure 4 is a graph showing an example of bitrate adaptation for three core encoders. In particular, in Figure 4, the first row shows the core encoder total bitrate total_brate for audio stream #1, the second row shows the core encoder total bitrate total_brate for audio stream #2, the third row shows the core encoder total bitrate total_brate for audio stream #3, the fourth row is audio stream #1, the fifth row is audio stream #2, and the sixth row is audio stream #3.

[0097] In the example of Figure 4, the adaptation of the total bitrate of the three core encoders, total_brate, is based on the VAD activity (active / inactive frames). As can be seen from Figure 4, in most cases, the variable side bit budget (bits) side There are small fluctuations in the core encoder total bitrate total_brate as a result of VAD activity, and there are infrequent large changes in the core encoder total bitrate total_brate as a result of VAD activity.

[0098] For example, referring to Figure 4, case A) corresponds to a frame in which the VAD activity of audio stream #1 changes from 1 (active) to 0 (inactive). According to this logic, the smallest core encoder total bitrate, total_brate, is allocated to audio object #1, while the core encoder total bitrates for active audio objects #2 and #3, total_brate, are increased. Case B) corresponds to a frame in which the VAD activity of audio stream #3 changes from 1 (active) to 0 (inactive), while the VAD activity of audio stream #1 remains 0. According to this logic, the smallest core encoder total bitrate, total_brate, is allocated to audio streams #1 and #3, while the core encoder total bitrate for active audio stream #2, total_brate, is further increased.

[0099] The above logic in section 2.4.1 can be made dependent on the total bit rate ism_total_rate. For example, the bit budget B in the above sub-operation 1 VAD0 may be set higher for higher total bitrates ism_total_brate and lower for lower total bitrates ism_total_brate.

[0100] 2.4.2 Bitrate Adaptation Based on ISm Importance The logic described in the previous section 2.4.1 results in approximately the same core encoder bitrate for all audio streams with active content (VAD = 1) in a given frame, but it may be beneficial to introduce inter-object core encoder bitrate adaptation based on ISm importance classification (or more broadly, an indicator of how important the coding of a particular audio object in the current frame is for a given (satisfactory) quality of the decoded synthesis).

[0101] The classification of the importance of ISm may be based on several parameters and / or combinations of parameters, such as the core encoder type (coder_type), FEC (Forward Error Correction), speech signal class, human voice / music classification decision, and / or SNR (Signal-to-Noise Ratio) estimates (snr_celp, snr_tcx) from the open-loop ACELP / TCX (Algebraic Code Excited Linear Prediction / Transform Coding Excitation) core decision module described in [1]. Other parameters may be used to determine the classification of the importance of ISm.

[0102] In a non-limiting example, a simple classification of ISm importance based on core encoder type as defined in Reference [1] is implemented. To that end, the bit budget allocator 106 of Figure 1 includes a classifier (not shown) for assessing the importance of a particular ISm stream. As a result, four different ISm importance classes are defined: ISm is defined. - No Metadata Class ISM_NO_META: Frames without metadata coding, e.g. inactive frames with VAD=0 - Low importance class ISM_LOW_IMP: frames with coder_type = UNVOICED or INACTIVE - Medium importance class ISM_MEDIUM_IMP: frames with coder_type = VOICED - High importance class ISM_HIGH_IMP: frames with coder_type = GENERIC

[0103] The ISm importance classes are then used by bit budget allocator 106 in the bit rate adaptation algorithm (see sub-operation 6 in section 2.4 above) to allocate larger bit budgets to audio streams with higher ISm importance and lower bit budgets to audio streams with lower ISm importance. Thus, for every audio stream n, n = 0, ..., N-1, the following bit rate adaptation algorithm is used by bit budget allocator 106: 1. class ISm = For frames classified as ISM_NO_META, a constant low bitrate B VAD0 is assigned. 2. class ISm = ISM_LOW_IMP, the total bitrate total_brate is, for example, total_brate new [n] = max(α low *total_brate[n], B low ) where the constant α low is set to a value less than 1.0, for example 0.6. And the constant B low represents the minimum bitrate threshold supported by the codec for a particular configuration, which may depend, for example, on the codec's internal sampling rate, the bandwidth of the audio being coded, etc. (See reference [1] for further details on these values). 3. class ISm = ISM_MEDIUM_IMP, the total bitrate of the core encoder is, for example, total_brate new [n] = max(α med *total_brate[n], B low ) where the constant αmed is less than 1.0, but α low , for example set to a value greater than 0.8. 4. class ISm = ISM_HIGH_IMP, bitrate adaptation is not used. 5. Finally, calculate the saved bit budget (old total bitrate (total_brate) vs. new total bitrate (total_brate new ) is redistributed evenly among the audio streams with active content within the frame. The same bit budget redistribution logic as described in sub-actions 2 and 3 of Section 2.4.1 may be used.

[0104] Figure 5 is a graph showing an example of bitrate adaptation based on ISm importance logic. From top to bottom, the graph in Figure 5 synchronously shows: - the active speech segment of the audio stream for audio object #1, - the active speech segment of the audio stream for audio object #2, - total_bitrate, the total bitrate of the audio stream for audio object #1 without using the bitrate adaptation algorithm; - total_bitrate, the total bitrate of the audio stream for audio object #2 without using the bitrate adaptation algorithm, - total_brate, the total bitrate of the audio stream for audio object #1 when the bitrate adaptation algorithm is used, and - total_brate, the total bitrate of the audio stream for audio object #2 when the bitrate adaptation algorithm is used.

[0105] 5, with two audio objects (N=2) and a fixed total bit rate ism_total_brate equal to 48 kbps, the core encoder total bit rate total_brate for active frames of audio object #1 varies between 23.45 kbps and 23.65 kbps when the bit rate adaptation algorithm is not used, and between 19.15 kbps and 28.05 kbps when the bit rate adaptation algorithm is used. Similarly, the core encoder total bit rate total_brate for active frames of audio object #2 varies between 23.40 kbps and 23.65 kbps when the bit rate adaptation algorithm is not used, and between 19.10 kbps and 28.05 kbps when the bit rate adaptation algorithm is used. This results in a better and more efficient distribution of the available bit budget between the audio streams.

[0106] 2.5 Preprocessing 1, the method 150 for coding an object-based audio signal includes an operation 158 of pre-processing N audio streams conveyed from the configuration and decision processor 106 (bit budget allocator) over the N transport channels 104. To perform operation 158, the system 100 for coding an object-based audio signal includes a pre-processor 108.

[0107] Once the configuration and bitrate distribution among the N audio streams has been completed by the configuration and decision processor 106 (bit budget allocator), the preprocessor 108 performs sequential further preprocessing 158 on each of the N audio streams. Such preprocessing 158 may include, for example, further signal classification, selection of further core encoders (e.g., selection from an ACELP core, a TCX core, and an HQ core), selection of different internal sampling frequencies F adapted to the bitrates used for the core encoders, and so on. s Examples of such pre-processing can be found, for example, in reference [1] in connection with the EVS codec and therefore will not be described further in this disclosure.

[0108] 2.6 Core Encoding 1, the method 150 for coding an object-based audio signal includes a core encoding operation 159. To perform operation 159, the system 100 for coding an object-based audio signal includes, for example, the above-mentioned encoder of N audio streams, including N core encoders 109 for coding the N audio streams conveyed from the pre-processor 108 via the N transport channels 104, respectively.

[0109] In particular, the N audio streams are encoded using N variable bit rate core encoders 109, e.g., mono core encoders. The bit rate used by each of the N core encoders is the bit rate selected by the configuration and decision processor 106 (bit budget allocator) for the corresponding audio stream. For example, the core encoder described in Reference [1] may be used as the core encoder 109.

[0110] 3.0 Bitstream Structure 1, the method 150 for coding an object-based audio signal includes an operation of multiplexing 160. To perform operation 160, the system 100 for coding an object-based audio signal includes a multiplexer 110.

[0111] Figure 6 is a schematic diagram illustrating the structure, in terms of frames, of the bitstream 111 generated by multiplexer 110 and transmitted from coding system 100 of Figure 1 to decoding system 700 of Figure 7. Whether or not metadata is present and transmitted, the structure of bitstream 111 may be assembled as shown in Figure 6.

[0112] Referring to Figure 6, the multiplexer 110 writes the indices of the N audio streams from the beginning of the bitstream 111, while the indices of the ISm common signaling 113 from the configuration and decision processor 106 (bit budget allocator) and the metadata 112 from the metadata processor 105 are written from the end of the bitstream 111.

[0113] 3.1 ISm Common Signaling The multiplexer writes ISm common signaling 113 from the end of the bitstream 111. The ISm common signaling is generated by the configuration and decision processor 106 (bit budget allocator) and contains a variable number of bits representing:

[0114] (a) Number of Audio Objects N: The signaling regarding the number N of coded audio objects present in the bitstream 111 is, for example, in the form of a unary code with a stop bit (e.g., for N = 3 audio objects, the first three bits of the ISm common signaling are "110").

[0115] (b) Metadata existence flag meta : flag metais present when the signal activity based bitrate adaptation described in section 2.4.1 is used and metadata for that particular audio object is present in the bitstream 111 (flag meta = 1) or does not exist (flag meta = 0), or (c) ISm importance class: This signaling is present when ISM importance-based bitrate adaptation as described in 2.4.2 is used, and contains one bit per audio object to indicate whether the ISm importance class defined in 2.4.2 is used. ISm Contains 2 bits per audio object to indicate (ISM_NO_META, ISM_LOW_IMP, ISM_MEDIUM_IMP, ISM_HIGH_IMP).

[0116] (d) ISm VAD flag VAD : ISm VAD flag is flag meta = 0 or class ISm = ISM_NO_META and distinguishes between the following two cases: 1) There is no input metadata or the metadata is not coded, and therefore the audio stream needs to be coded by the active coding mode (flag VAD = 1), and 2) Input metadata is present and transmitted, so the audio stream can be coded in the inactive coding mode (flag VAD = 0).

[0117] 3.2 Encoded Metadata Payload The multiplexer 110 is supplied with coded metadata 112 from the metadata processor 105 and determines whether metadata is coded in the current frame (flag meta = 1 or class ISm≠ ISM_NO_META) Write the metadata payload starting from the end of the bitstream for the audio object. The metadata bit budget for each audio object is not constant, but rather adaptive between objects and frames. Different metadata format scenarios are shown in Figure 2.

[0118] If metadata is not present or not transmitted for at least some of the N audio objects, then for these audio objects the metadata flag is set to 0, i.e., flag meta = 0 or class ISm = ISM_NO_META, then no metadata indexes are sent in association with those audio objects, i.e. bits meta [n] = 0.

[0119] 3.3 Audio Stream Payload The multiplexer 110 receives the N audio streams 114 coded by the N core encoders 109 via the N transport channels 104 and writes the payloads of the audio streams sequentially for the N audio streams in chronological order from the beginning of the bitstream 111 (see Figure 6). The bit budget of each of the N audio streams is fluctuating as a result of the bitrate adaptation algorithm described in Section 2.4.

[0120] 4.0 Decoding Audio Objects FIG. 7 is a schematic block diagram illustrating together a system 700 for decoding audio objects according to an audio stream with associated metadata and a corresponding method 750 for decoding the audio objects.

[0121] 4.1 Demultiplexing 7, a method 750 for decoding audio objects according to an audio stream with associated metadata includes a demultiplexing operation 755. To perform operation 755, the system 700 for decoding audio objects according to an audio stream with associated metadata includes a demultiplexer 705.

[0122] The demultiplexer receives a bitstream 701 transmitted from the coding system 100 of Figure 1 to the decoding system 700 of Figure 7. In particular, the bitstream 701 of Figure 7 corresponds to the bitstream 111 of Figure 1.

[0123] The demultiplexer 110 extracts from the bitstream 701 (a) the N coded audio streams 114, (b) the coded metadata 112 for the N audio objects, and (c) the ISm common signaling 113 read from the end of the received bitstream 701.

[0124] 4.2 Metadata Decoding and Dequantization 7, a method 750 for decoding audio objects according to an audio stream having associated metadata includes a metadata decoding and dequantization operation 756. To perform operation 756, the system 700 for decoding audio objects according to an audio stream having associated metadata includes a metadata decoding and dequantization processor 706.

[0125] The metadata decoding and inverse quantization processor 706 is provided with output settings 709 for decoding and inverse quantizing the coded metadata 112 for transmitted audio objects, the ISm common signaling 113, and the metadata for audio streams / objects with active content. The output settings 709 are command line parameters for the number M of decoded audio objects / transport channels and / or audio formats, which can be equal to or different from the number N of coded audio objects / transport channels. The metadata decoding and inverse quantization processor 706 generates decoded metadata 704 for the M audio objects / transport channels and provides information on the respective bit budgets for the M decoded metadata on line 708. Obviously, the decoding and inverse quantization performed by the processor 706 is the inverse of the quantization and coding performed by the metadata processor 105 of FIG. 1.

[0126] 4.3 Bitrate Configuration and Decisions 7, a method 750 for decoding audio objects according to an audio stream with associated metadata includes an operation of configuring and determining a bit rate per channel 757. To perform operation 757, the system 700 for decoding audio objects according to an audio stream with associated metadata includes a configuration and determination processor 707 (bit budget allocator).

[0127] The bit budget allocator 707 receives (a) information about the bit budget for each of the M decoded metadata on line 708 and (b) the ISm importance class from common signaling 113. ISmand determines the core decoder bit rate total_rate[n] for each audio stream. The bit budget allocator 707 determines the core decoder bit rate using the same procedure as the bit budget allocator 106 in Figure 1 (see Section 2.4).

[0128] 4.4 Core-decoding 7, a method 750 for decoding audio objects according to audio streams with associated metadata includes a core decoding operation 760. To perform the operation 760, a system 700 for decoding audio objects according to audio streams with associated metadata includes decoders for the N audio streams 114, including N core decoders 710, e.g., N variable bit rate core decoders.

[0129] The N audio streams 114 from the demultiplexer 705 are decoded, e.g., in N variable bit rate core decoders 710, in turn, at their respective core decoder bit rates as determined by the bit budget allocator 707. If the number M of decoded audio objects requested by the output settings 709 is less than the number of transport channels, i.e., M < N, then a fewer number of core decoders are used. Similarly, in such cases, it is possible that not all metadata payloads are decoded.

[0130] Depending on the N audio streams 114 from the demultiplexer 705, the core decoder bit rate determined by the bit budget allocator 707, and the output settings 709, the core decoder 710 generates M decoded audio streams 703 on each of the M transport channels.

[0131] Rendering 5.0 audio channels In a render audio channels operation 761, an audio object renderer 711 converts the M decoded metadata 704 and the M decoded audio streams 703 into several output audio channels 702, taking into account output settings 712 that indicate the number and content of the output audio channels to be generated. Again, the number of output audio channels 702 may be equal to or different from the number M.

[0132] The renderer 711 may be designed in a variety of different configurations to obtain the desired output audio channels, and as such, the renderer will not be further described in this disclosure.

[0133] 6.0 Source Code According to a non-limiting exemplary embodiment, the system and method for coding an object-based audio signal disclosed in the above description may be implemented by the following source code (expressed in C code) provided below as additional disclosure:

[0134] void ism_metadata_enc( const long ism_total_brate, / * i : ISm total bitrate * / const short n_ISms, / * i : number of objects * / ISM_METADATA_HANDLE hIsmMeta[], / * i / o: Handle to ISM metadata * / ENC_HANDLE hSCE[], / * i / o: Handle to element encoder * / BSTR_ENC_HANDLE hBstr, / * i / o: Bitstream handle * / short nb_bits_metadata[], / * o : number of bits of metadata * / short localVAD[] ) { short i, ch, nb_bits_start, diff; short idx_azimuth, idx_azimuth_abs, flag_abs_azimuth[MAX_NUM_OBJECTS], nbits_diff_azimuth; short idx_elevation, idx_elevation_abs, flag_abs_elevation[MAX_NUM_OBJECTS], nbits_diff_elevation; float valQ; ISM_METADATA_HANDLE hIsmMetaData; long element_brate[MAX_NUM_OBJECTS], total_brate[MAX_NUM_OBJECTS]; short ism_metadata_flag_global; short ism_imp[MAX_NUM_OBJECTS]; / * Initialize * / ism_metadata_flag_global = 0; set_s( nb_bits_metadata, 0, n_ISms ); set_s( flag_abs_azimuth, 0, n_ISms ); set_s( flag_abs_elevation, 0, n_ISms ); / *----------------------------------------------------------------* * Set metadata presence / importance flags *----------------------------------------------------------------* / for( ch = 0; ch < n_ISms; ch++ ) { if( hIsmMeta[ch]->ism_metadata_flag ) { hIsmMeta[ch]->ism_metadata_flag = localVAD[ch]; } else { hIsmMeta[ch]->ism_metadata_flag = 0; } if ( hSCE[ch]->hCoreCoder[0]->tcxonly ) { / * Metadata is transmitted in every frame at the highest bitrate (using only the TCX core) * / hIsmMeta[ch]->ism_metadata_flag = 1; } } rate_ism_importance( n_ISms, hIsmMeta, hSCE, ism_imp ); / *----------------------------------------------------------------* * Write ISm common signaling *----------------------------------------------------------------* / / * Write some objects - unary encoding * / for( ch = 1; ch < n_ISms; ch++ ) { push_indice( hBstr, IND_ISM_NUM_OBJECTS, 1, 1 ); } push_indice( hBstr, IND_ISM_NUM_OBJECTS, 0, 1 ); / * Write ISm metadata flags (one per object) * / for( ch = 0; ch < n_ISms; ch++ ) { push_indice( hBstr, IND_ISM_METADATA_FLAG, ism_imp[ch], ISM_METADATA_FLAG_BITS ); ism_metadata_flag_global |= hIsmMeta[ch]->ism_metadata_flag; } / * Write VAD flags * / for( ch = 0; ch < n_ISms; ch++ ) { if( hIsmMeta[ch]->ism_metadata_flag == 0 ) { push_indice( hBstr, IND_ISM_VAD_FLAG, localVAD[ch], VAD_FLAG_BITS ); } } if( ism_metadata_flag_global ) { / *----------------------------------------------------------------* * Metadata quantization and coding, looping over all objects *----------------------------------------------------------------* / for( ch = 0; ch < n_ISms; ch++ ) { hIsmMetaData = hIsmMeta[ch]; nb_bits_start = hBstr->nb_bits_tot; if( hIsmMeta[ch]->ism_metadata_flag ) { / *----------------------------------------------------------------* * Azimuth angle quantization and coding *----------------------------------------------------------------* / / * Azimuth angle quantization * / idx_azimuth_abs = usquant( hIsmMetaData->azimuth, &valQ, ISM_AZIMUTH_MIN, ISM_AZIMUTH_DELTA, (1 << ISM_AZIMUTH_NBITS) ); idx_azimuth = idx_azimuth_abs; nbits_diff_azimuth = 0; flag_abs_azimuth[ch] = 0; / * default to differential coding * / if( hIsmMetaData->azimuth_diff_cnt == ISM_FEC_MAX / * Differentially encode at most ISM_FEC_MAX consecutive frames (to control decoding in FEC) * / || hIsmMetaData->last_ism_metadata_flag == 0 / * If the last frame did not code metadata, do not use delta coding * / ) { flag_abs_azimuth[ch] = 1; } / * Try differential coding * / if( flag_abs_azimuth[ch] == 0 ) { diff = idx_azimuth_abs - hIsmMetaData->last_azimuth_idx; if( diff == 0 ) { idx_azimuth = 0; nbits_diff_azimuth = 1; } else if( ABSVAL( diff ) < ISM_MAX_AZIMUTH_DIFF_IDX ) / * When diff bit >= abs bit, abs takes precedence * / { idx_azimuth = 1 << 1; nbits_diff_azimuth = 1; if( diff < 0 ) { idx_azimuth += 1; / * negative sign * / diff *= -1; } else { idx_azimuth += 0; / * positive sign * / } idx_azimuth = idx_azimuth << diff; nbits_diff_azimuth++; / * Unary encoding of "diff" * / idx_azimuth += ((1< <diff) - 1); nbits_diff_azimuth += diff; if( nbits_diff_azimuth < ISM_AZIMUTH_NBITS - 1 ) { / * Add stop bits - only for codewords shorter than ISM_AZIMUTH_NBITS * / idx_azimuth = idx_azimuth << 1; nbits_diff_azimuth++; } } else { flag_abs_azimuth[ch] = 1; } } / * Update the counter * / if( flag_abs_azimuth[ch] == 0 ) { hIsmMetaData->azimuth_diff_cnt++; hIsmMetaData->elevation_diff_cnt = min( hIsmMetaData->elevation_diff_cnt, ISM_FEC_MAX ); } else { hIsmMetaData->azimuth_diff_cnt = 0; } / * Write the azimuth angle * / push_indice( hBstr, IND_ISM_AZIMUTH_DIFF_FLAG, flag_abs_azimuth[ch], 1 ); if( flag_abs_azimuth[ch] ) { push_indice( hBstr, IND_ISM_AZIMUTH, idx_azimuth, ISM_AZIMUTH_NBITS ); } else { push_indice( hBstr, IND_ISM_AZIMUTH, idx_azimuth, nbits_diff_azimuth ); } / *----------------------------------------------------------------* * Elevation angle quantization and coding *----------------------------------------------------------------* / / * Elevation angle quantization * / idx_elevation_abs = usquant( hIsmMetaData->elevation, &valQ, ISM_ELEVATION_MIN, ISM_ELEVATION_DELTA, (1 << ISM_ELEVATION_NBITS) ); idx_elevation = idx_elevation_abs; nbits_diff_elevation = 0; flag_abs_elevation[ch] = 0; / * differential coding by default * / if( hIsmMetaData->elevation_diff_cnt == ISM_FEC_MAX / * Differential encoding is performed on at most ISM_FEC_MAX consecutive frames (to control decoding in FEC) * / || hIsmMetaData->last_ism_metadata_flag == 0 / * If the last frame did not code metadata, do not use delta coding * / ) { flag_abs_elevation[ch] = 1; } / * NOTE: Elevation is coded only from the second frame onwards (it has no meaning in init_frame) * / if( hSCE[0]->hCoreCoder[0]->ini_frame == 0 ) { flag_abs_elevation[ch] = 1; hIsmMetaData->last_elevation_idx = idx_elevation_abs; } diff = idx_elevation_abs - hIsmMetaData->last_elevation_idx; / * Avoid absolute coding of elevation angles if absolute coding is already used for azimuth angles * / if( flag_abs_azimuth[ch] == 1 ) { flag_abs_elevation[ch] = 0; if( diff >= 0 ) { diff = min( diff, ISM_MAX_ELEVATION_DIFF_IDX ); } else { diff = -1 * min( -diff, ISM_MAX_ELEVATION_DIFF_IDX ); } } / * Try differential coding * / if( flag_abs_elevation[ch] == 0 ) { if( diff == 0 ) { idx_elevation = 0; nbits_diff_elevation = 1; } else if( ABSVAL( diff ) < ISM_MAX_ELEVATION_DIFF_IDX ) / * When diff bit >= abs bit, abs takes precedence * / { idx_elevation = 1 << 1; nbits_diff_elevation = 1; if( diff < 0 ) { idx_elevation += 1; / * negative sign * / diff *= -1; } else { idx_elevation += 0; / * positive sign * / } idx_elevation = idx_elevation << diff; nbits_diff_elevation++; / * Unary encoding of "diff" * / idx_elevation += ((1 << diff) - 1); nbits_diff_elevation += diff; if( nbits_diff_elevation < ISM_ELEVATION_NBITS - 1 ) { / * Add stop bits * / idx_elevation = idx_elevation << 1; nbits_diff_elevation++; } } else { flag_abs_elevation[ch] = 1; } } / * Update the counter * / if( flag_abs_elevation[ch] == 0 ) { hIsmMetaData->elevation_diff_cnt++; hIsmMetaData->elevation_diff_cnt = min( hIsmMetaData->elevation_diff_cnt, ISM_FEC_MAX ); } else { hIsmMetaData->elevation_diff_cnt = 0; } / * Write the elevation angle * / if( flag_abs_azimuth[ch] == 0 ) / * If "flag_abs_azimuth == 1", do not write "flag_abs_elevation" * / / * VE: TBV for VAD 0->1 * / { push_indice( hBstr, IND_ISM_ELEVATION_DIFF_FLAG, flag_abs_elevation[ch], 1 ); } if( flag_abs_elevation[ch] ) { push_indice( hBstr, IND_ISM_ELEVATION, idx_elevation, ISM_ELEVATION_NBITS ); } else { push_indice( hBstr, IND_ISM_ELEVATION, idx_elevation, nbits_diff_elevation ); } / *----------------------------------------------------------------* * Update *----------------------------------------------------------------* / hIsmMetaData->last_azimuth_idx = idx_azimuth_abs; hIsmMetaData->last_elevation_idx = idx_elevation_abs; / * Store the number of bits of metadata written * / nb_bits_metadata[ch] = hBstr->nb_bits_tot - nb_bits_start; } } / *----------------------------------------------------------------* *Inter-object logic to minimize the use of several absolute coded indices in the same frame *----------------------------------------------------------------* / i = 0; while( i == 0 || i < n_ISms / INTER_OBJECT_PARAM_CHECK ) { short num, abs_num, abs_first, abs_next, pos_zero; short abs_matrice[INTER_OBJECT_PARAM_CHECK * 2]; num = min( INTER_OBJECT_PARAM_CHECK, n_ISms - i * INTER_OBJECT_PARAM_CHECK ); i++; set_s( abs_matrice, 0, INTER_OBJECT_PARAM_CHECK * ISM_NUM_PARAM ); for( ch = 0; ch < num; ch++ ) { if( flag_abs_azimuth[ch] == 1 ) { abs_matrice[ch*ISM_NUM_PARAM] = 1; } if( flag_abs_elevation[ch] == 1 ) { abs_matrice[ch*ISM_NUM_PARAM + 1] = 1; } } abs_num = sum_s( abs_matrice, INTER_OBJECT_PARAM_CHECK * ISM_NUM_PARAM ); abs_first = 0; while( abs_num > 1 ) { / * Find the first "1" entry * / while( abs_matrice[abs_first] == ​​0 ) { abs_first++; } / * Find the next "1" entry * / abs_next = abs_first + 1; while( abs_matrice[abs_next] == ​​0 ) { abs_next++; } / * Find the position of "0" * / pos_zero = 0; while( abs_matrice[pos_zero] == 1 ) { pos_zero++; } ch = abs_next / ISM_NUM_PARAM; if( abs_next % ISM_NUM_PARAM == 0 ) { hIsmMeta[ch]->azimuth_diff_cnt = abs_num - 1; } if( abs_next % ISM_NUM_PARAM == 1 ) { hIsmMeta[ch]->elevation_diff_cnt = abs_num - 1; / *hIsmMeta[ch]->elevation_diff_cnt = min( hIsmMeta[ch]->elevation_diff_cnt, ISM_FEC_MAX );* / } abs_first++; abs_num--; } } } / *----------------------------------------------------------------* * Configuring and determining bitrate per channel *----------------------------------------------------------------* / ism_config( ism_total_brate, n_ISms, hIsmMeta, localVAD, ism_imp, element_brate, total_brate, nb_bits_metadata ); for( ch = 0; ch < n_ISms; ch++ ) { hIsmMeta[ch]->last_ism_metadata_flag = hIsmMeta[ch]->ism_metadata_flag; hSCE[ch]->hCoreCoder[0]->low_rate_mode = 0; if ( hIsmMeta[ch]->ism_metadata_flag == 0 && localVAD[ch][0] == 0 && ism_metadata_flag_global ) { hSCE[ch]->hCoreCoder[0]->low_rate_mode = 1; } hSCE[ch]->element_brate = element_brate[ch]; hSCE[ch]->hCoreCoder[0]->total_brate = total_brate[ch]; / * Write metadata only in the active frame * / if( hSCE[0]->hCoreCoder[0]->core_brate > SID_2k40 ) { reset_indices_enc( hSCE[ch]->hMetaData, MAX_BITS_METADATA ); } } return; } void rate_ism_importance( const short n_ISms, / * i : number of objects * / ISM_METADATA_HANDLE hIsmMeta[], / * i / o: Handle to ISM metadata * / ENC_HANDLE hSCE[], / * i / o: Handle to element encoder * / short ism_imp[] / * o : ISM importance flag * / ) { short ch, ctype; for( ch = 0; ch < n_ISms; ch++ ) { ctype = hSCE[ch]->hCoreCoder[0]->coder_type_raw; if( hIsmMeta[ch]->ism_metadata_flag == 0 ) { ism_imp[ch] = ISM_NO_META; } else if( ctype == INACTIVE || ctype == UNVOICED ) { ism_imp[ch] = ISM_LOW_IMP; } else if( ctype == VOICED ) { ism_imp[ch] = ISM_MEDIUM_IMP; } else / * GENERIC * / { ism_imp[ch] = ISM_HIGH_IMP; } } return; } void ism_config( const long ism_total_brate, / * i : ISm total bitrate * / const short n_ISms, / * i : number of objects * / ISM_METADATA_HANDLE hIsmMeta[], / * i / o: Handle to ISM metadata * / short localVAD[], const short ism_imp[], / * i : ISM importance flag * / long element_brate[], / * o : element bitrate per object * / long total_brate[], / * o : total bitrate per object * / short nb_bits_metadata[] / * i / o: number of bits of metadata * / ) { short ch; short bits_element[MAX_NUM_OBJECTS], bits_CoreCoder[MAX_NUM_OBJECTS]; short bits_ism, bits_side; long tmpL; short ism_metadata_flag_global; / * Initialization * / ism_metadata_flag_global = 0; bits_side = 0; if( hIsmMeta != NULL ) { for( ch = 0; ch < n_ISms; ch++ ) { ism_metadata_flag_global |= hIsmMeta[ch]->ism_metadata_flag;} } / * Decision about bitrate per channel - constant (at one ism_total_brate) for the duration of the session * / bits_ism = ism_total_brate / FRMS_PER_SECOND; set_s( bits_element, bits_ism / n_ISms, n_ISms ); bits_element[n_ISms - 1] += bits_ism % n_ISms; bitbudget_to_brate( bits_element, element_brate, n_ISms ); / * Count bits in ISm common signaling * / if( hIsmMeta != NULL ) { nb_bits_metadata[0] += n_ISms * ISM_METADATA_FLAG_BITS + n_ISms; for( ch = 0; ch < n_ISms; ch++ ) { if( hIsmMeta[ch]->ism_metadata_flag == 0 ) { nb_bits_metadata[0] += ISM_METADATA_VAD_FLAG_BITS; } } } / * Divide the metadata bit budget evenly among the channels * / if( nb_bits_metadata != NULL ) { bits_side = sum_s( nb_bits_metadata, n_ISms ); set_s( nb_bits_metadata, bits_side / n_ISms, n_ISms ); nb_bits_metadata[n_ISms - 1] += bits_side % n_ISms; v_sub_s( bits_element, nb_bits_metadata, bits_CoreCoder, n_ISms ); bitbudget_to_brate( bits_CoreCoder, total_brate, n_ISms ); mvs2s( nb_bits_metadata, nb_bits_metadata, n_ISms ); } / * Allocate less of the CoreCoder bit budget to inactive streams (at least one stream must be active) * / if( ism_metadata_flag_global ) { long diff; short n_higher, flag_higher[MAX_NUM_OBJECTS]; set_s( flag_higher, 1, MAX_NUM_OBJECTS ); diff = 0; for( ch = 0; ch < n_ISms; ch++ ) { if( hIsmMeta[ch]->ism_metadata_flag == 0 && localVAD[ch] == 0 ) { diff += bits_CoreCoder[ch] - BITS_ISM_INACTIVE; bits_CoreCoder[ch] = BITS_ISM_INACTIVE; flag_higher[ch] = 0; } } n_higher = sum_s( flag_higher, n_ISms ); if( diff > 0 && n_higher > 0 ) { tmpL = diff / n_higher; for( ch = 0; ch < n_ISms; ch++ ) { if( flag_higher[ch] ) { bits_CoreCoder[ch] += tmpL; } } tmpL = diff % n_higher; ch = 0; while( flag_higher[ch] == 0 ) { ch++; } bits_CoreCoder[ch] += tmpL; } bitbudget_to_brate( bits_CoreCoder, total_brate, n_ISms ); diff = 0; for( ch = 0; ch < n_ISms; ch++ ) { long limit; limit = MIN_BRATE_SWB_BWE / FRMS_PER_SECOND; if( element_brate[ch] < MIN_BRATE_SWB_STEREO ) { limit = MIN_BRATE_WB_BWE / FRMS_PER_SECOND; } else if( element_brate[ch] >= SCE_CORE_16k_LOW_LIMIT ) { / *限度(limit) = SCE_CORE_16k_LOW_LIMIT;* / limit = (ACELP_16k_LOW_LIMIT + SWB_TBE_1k6) / FRMS_PER_SECOND; } if( ism_imp[ch] == ISM_NO_META && localVAD[ch] == 0 ) { tmpL = BITS_ISM_INACTIVE; } else if( ism_imp[ch] == ISM_LOW_IMP ) { tmpL = BETA_ISM_LOW_IMP * bits_CoreCoder[ch]; tmpL = max( limit, bits_CoreCoder[ch] - tmpL ); } else if( ism_imp[ch] == ISM_MEDIUM_IMP ) { tmpL = BETA_ISM_MEDIUM_IMP * bits_CoreCoder[ch]; tmpL = max( limit, bits_CoreCoder[ch] - tmpL ); } else / * ism_imp[ch] == ISM_HIGH_IMP * / { tmpL = bits_CoreCoder[ch]; } diff += bits_CoreCoder[ch] - tmpL; bits_CoreCoder[ch] = tmpL; } if( diff > 0 && n_higher > 0 ) { tmpL = diff / n_higher; for( ch = 0; ch < n_ISms; ch++ ) { if( flag_higher[ch] ) { bits_CoreCoder[ch] += tmpL; } } tmpL = diff % n_higher; ch = 0; while( flag_higher[ch] == 0 ) { ch++; } bits_CoreCoder[ch] += tmpL; } / * Validate for max bitrate @ 12.8kHz core * / diff = 0; for ( ch = 0; ch < n_ISms; ch++ ) { limit_high = STEREO_512k / FRMS_PER_SECOND; if ( element_brate[ch] < SCE_CORE_16k_LOW_LIMIT ) / * Reproduce the function set_ACELP_flag() -> It is not intended to switch the internal sampling rate of ACELP within an object * / { limit_high = ACELP_12k8_HIGH_LIMIT / FRMS_PER_SECOND; } tmpL = min( bits_CoreCoder[ch], limit_high ); diff += bits_CoreCoder[ch] - tmpL; bits_CoreCoder[ch] = tmpL; } if ( diff > 0 ) { ch = 0; for ( ch = 0; ch < n_ISms; ch++ ) { if ( flag_higher[ch] == 0 ) { if ( diff > limit_high ) { diff += bits_CoreCoder[ch] - limit_high; bits_CoreCoder[ch] = limit_high; } else { bits_CoreCoder[ch] += diff; break; } } } } bitbudget_to_brate( bits_CoreCoder, total_brate, n_ISms ); } return; }

[0135] 7.0 Hardware Implementation FIG. 8 is a simplified block diagram of an exemplary configuration of hardware components forming the coding and decoding systems and methods described above.

[0136] Each of the coding and decoding systems may be implemented as part of a mobile terminal, as part of a portable media player, or in any similar device. Each of the coding and decoding systems (identified as 1200 in FIG. 8) includes an input 1202, an output 1204, a processor 1206, and a memory 1208.

[0137] Input 1202 is configured to receive an input signal, e.g., N audio objects 102 (N audio streams with corresponding N metadata) of Figure 1 or bitstream 701 of Figure 7, in digital or analog format. Output 1204 is configured to provide an output signal, e.g., bitstream 111 of Figure 1, or M decoded audio channels 703 and M decoded metadata 704 of Figure 7. Input 1202 and output 1204 may be implemented in a common module, e.g., a serial input / output device.

[0138] The processor 1206 is operatively connected to the input 1202, the output 1204, and the memory 1208. The processor 1206 is implemented as one or more processors for supporting the functions of the various processors and other modules in Figures 1 and 7 to execute code instructions.

[0139] The memory 1208 may include non-transitory memory for storing code instructions executable by the processor 1206, in particular processor-readable memory containing non-transitory instructions that, when executed, cause the processor to perform the operations and processors / modules of the coding and decoding systems and methods as described in this disclosure. The memory 1208 may also include random access memory or buffers for storing intermediate processed data from various functions performed by the processor 1206.

[0140] Those skilled in the art will recognize that the description of the coding and decoding systems and methods is illustrative only and is not intended to be limiting in any way. Other embodiments will readily occur to such skilled artisans having the benefit of this disclosure. Furthermore, the disclosed coding and decoding systems and methods may be customized to provide valuable solutions to existing needs and problems of encoding and decoding audio.

[0141]

[0013] For purposes of clarity, not all routine features of an implementation of the coding and decoding systems and methods are shown and described. Of course, it will be understood that in developing any such actual implementation of the coding and decoding systems and methods, numerous implementation-specific decisions may need to be made to achieve the developer's particular objectives, such as compliance with application-, system-, network-, and business-related constraints, and that these particular objectives will vary from implementation to implementation and from developer to developer. Moreover, it will be understood that the development effort may be complex and time-consuming, but will nevertheless be a routine undertaking of engineering for those of ordinary skill in the art of speech processing having the benefit of this disclosure.

[0142] In accordance with this disclosure, the processors / modules, processing operations, and / or data structures described herein may be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. In addition, those skilled in the art will recognize that less general-purpose devices, such as hardwired devices, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., may also be used. When a method comprising a series of operations and sub-operations is performed by a processor, computer, or machine, and the operations and sub-operations may be stored as a series of non-transitory code instructions readable by a processor, computer, or machine, the operations and sub-operations may be stored on a tangible and / or non-transitory medium.

[0143] The coding and decoding systems and methods described herein may use software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.

[0144] In the coding and decoding systems and methods described herein, various operations and sub-operations may be performed in various orders, and some of the operations and sub-operations may be optional.

[0145] Although the present disclosure has been described above through non-limiting exemplary embodiments of the present disclosure, these embodiments may be modified at will within the scope of the appended claims without departing from the spirit and essence of the present disclosure.

[0146] 8.0 References The following references are referenced in this disclosure, the entire contents of which are incorporated herein by reference: [1] 3GPP Specification TS 26.445: "Codec for Enhanced Voice Services (EVS). Detailed Algorithmic Description", v.12.0.0, September 2014 [2] V. Eksler, “Method and Device for Allocating a Bit-budget Between Sub-frames in a CELP Codec”, PCT Patent Application PCT / CA2018 / 51175

[0147] 9.0 Further Embodiments The following embodiments (embodiments 1 to 83) are part of the present disclosure related to the present invention.

[0148] Embodiment 1. A system for coding an object-based audio signal including audio objects according to an audio stream having associated metadata, comprising: an audio stream processor for analyzing the audio stream; a metadata processor responsive to information about the audio stream from analysis by the audio stream processor for encoding metadata of the input audio stream.

[0149] Embodiment 2. The system of embodiment 1, wherein the metadata processor outputs information about the bit budget of the metadata of the audio objects, and wherein the system further includes a bit budget allocator responsive to the information about the bit budget of the metadata of the audio objects from the metadata processor for allocating bit rates to the audio stream.

[0150] Embodiment 3. The system of embodiment 1 or 2, including an encoder for an audio stream including coded metadata.

[0151] Embodiment 4. The system of any one of embodiments 1 to 3, wherein the encoder includes several core-coders that use the bit rate allocated to the audio stream by the bit budget allocator.

[0152] Embodiment 5. The system of any one of embodiments 1 to 4, wherein the object-based audio signal includes at least one of human voice, music, and general audio sounds.

[0153] Embodiment 6. The system of any one of embodiments 1 to 5, wherein the object-based audio signal represents or encodes a complex audio auditory scene as a collection of individual elements, said audio objects.

[0154] Embodiment 7. The system of any one of embodiments 1 to 6, wherein each audio object includes an audio stream with associated metadata.

[0155] Embodiment 8. The system of any one of embodiments 1 to 7, wherein the audio stream is a separate stream having metadata.

[0156] Embodiment 9. The system of any one of embodiments 1 to 8, wherein the audio stream represents an audio waveform and typically includes one or two channels.

[0157] Embodiment 10. The system of any one of embodiments 1 to 9, wherein the metadata is a set of information describing the audio stream and artistic intent used to convey the original or coded audio object to a final playback system.

[0158] Embodiment 11. The system of any one of embodiments 1 to 10, wherein the metadata typically describes spatial characteristics of each audio object.

[0159] Embodiment 12. The system of any one of embodiments 1 to 11, wherein the spatial characteristics include one or more of the position, orientation, volume, and width of the audio object.

[0160] Embodiment 13. The system of any one of embodiments 1 to 12, wherein each audio object includes a set of metadata, referred to as input metadata, defined as an unquantized metadata representation used as input to the codec.

[0161] Embodiment 14. The system of any one of embodiments 1 to 13, wherein each audio object includes a set of metadata, called coded metadata, defined as quantized coded metadata that is part of the bitstream transmitted from the encoder to the decoder.

[0162] Embodiment 15. The system of any one of embodiments 1 to 14, wherein the playback system is configured to render audio objects in a 3D audio space around the listener at the playback side using the transmitted metadata and artistic intent.

[0163] Embodiment 16. The system of any one of embodiments 1 to 15, wherein the playback system includes a head tracking device for dynamically modifying metadata during rendering of audio objects.

[0164] Embodiment 17. The system of any one of embodiments 1 to 16, comprising a framework for simultaneous coding of several audio objects.

[0165] Embodiment 18. The system of any one of embodiments 1 to 17, wherein the simultaneous coding of several audio objects uses a fixed, constant overall bit rate for encoding the audio objects.

[0166] Embodiment 19. The system of any one of embodiments 1 to 18, including a transmitter for transmitting some or all of the audio objects.

[0167] Embodiment 20. The system of any one of embodiments 1 to 19, wherein when coding a combination of audio formats in the framework, the constant overall bit rate represents the sum of the bit rates of the formats.

[0168] Embodiment 21. The system of any one of embodiments 1 to 20, wherein the metadata includes two parameters including an azimuth angle and an elevation angle.

[0169] Embodiment 22. The system of any one of embodiments 1 to 21, wherein azimuth and elevation parameters are stored for each audio frame for each audio object.

[0170] Embodiment 23. The system of any one of embodiments 1 to 22, including an input buffer for buffering at least one input audio stream and input metadata associated with the audio stream.

[0171] Embodiment 24. The system of any one of embodiments 1 to 23, wherein the input buffer buffers each audio stream for one frame.

[0172] Embodiment 25. The system of any one of embodiments 1 to 24, wherein the audio stream processor analyzes and processes the audio stream.

[0173] Embodiment 26. The system of any one of embodiments 1 to 25, wherein the audio stream processor includes at least one of the following elements: a time-domain transient detector, a spectrum analyzer, a long-term prediction analyzer, a pitch tracker and voicing analyzer, a voice / sound activity detector, a bandwidth detector, a noise estimator, and a signal classifier.

[0174] Embodiment 27. The system of any one of embodiments 1 to 26, wherein the signal classifier performs at least one of coder type selection, signal classification, and human voice / music classification.

[0175] Embodiment 28. The system of any one of embodiments 1 to 27, wherein the metadata processor analyzes, quantizes, and encodes metadata of the audio stream.

[0176] Embodiment 29. The system of any one of embodiments 1 to 28, wherein in an inactive frame, metadata is not encoded by the metadata processor and is not transmitted by the system in the bitstream of the corresponding audio object.

[0177] Embodiment 30. The system of any one of embodiments 1 to 29, wherein in the active frame, the metadata is encoded by the metadata processor for the corresponding object using a variable bit rate.

[0178] Embodiment 31. The system of any one of claims 1 to 30, wherein the bit budget allocator sums up the bit budgets of the metadata of the audio objects and adds the sum of the bit budgets to the bit budget of the signaling to allocate bit rate to the audio stream.

[0179] Embodiment 32. The system of any one of embodiments 1 to 31, including a preprocessor for further processing the audio streams once the configuration and bitrate distribution between the audio streams has been performed.

[0180] Embodiment 33. The system of any one of embodiments 1 to 32, wherein the preprocessor performs at least one of further classification of the audio stream, selection of a core encoder, and resampling.

[0181] Embodiment 34. The system of any one of embodiments 1 to 33, wherein the encoder encodes the audio streams in sequence.

[0182] Embodiment 35. The system of any one of embodiments 1 to 34, wherein the encoder sequentially encodes the audio stream using several variable bit rate core coders.

[0183] Embodiment 36. The system of any one of embodiments 1 to 35, wherein the metadata processor encodes the metadata sequentially in a loop using dependencies between the quantization of the audio objects and the metadata parameters of the audio objects.

[0184] Embodiment 37. The system of any one of embodiments 1 to 36, wherein the metadata processor quantizes an index of the metadata parameter using a quantization step to encode the metadata parameter.

[0185] Embodiment 38. The system of any one of embodiments 1 to 37, wherein the metadata processor quantizes the azimuth angle index using a quantization step to encode the azimuth angle parameter, and quantizes the elevation angle index using a quantization step to encode the elevation angle parameter.

[0186] Embodiment 39. The system of any one of embodiments 1 to 38, wherein the total metadata bit budget and number of quantization bits depend on the total bit rate of the codecs associated with one audio object, the total bit rate of the metadata, or the sum of the metadata bit budget and the core coder bit budget.

[0187] Embodiment 40. The system of any one of embodiments 1 to 39, wherein the azimuth parameter and the elevation parameter are expressed as one parameter.

[0188] Embodiment 41. The system of any one of embodiments 1 to 40, wherein the metadata processor encodes the index of the metadata parameter either absolutely or differentially.

[0189] Embodiment 42. The system of any one of embodiments 1 to 41, wherein the metadata processor encodes the index of the metadata parameter using absolute coding when there is a difference between the index of the current parameter and the index of the previous parameter that results in the number of bits required for differential coding being equal to or greater than the number of bits required for absolute coding.

[0190] Embodiment 43. The system of any one of embodiments 1 to 42, wherein the metadata processor encodes the index of the metadata parameter using absolute coding when no metadata was present in the previous frame.

[0191] Embodiment 44. The system of any one of embodiments 1 to 43, wherein the metadata processor encodes the index of the metadata parameter using absolute coding when the number of consecutive frames using differential coding is greater than the maximum number of consecutive frames coded using differential coding.

[0192] Embodiment 45. The system of any one of embodiments 1 to 44, wherein when the metadata processor encodes an index of a metadata parameter using absolute coding, it writes an absolute coding flag following the absolutely coded index of the metadata parameter, which distinguishes between absolute coding and differential coding.

[0193] Embodiment 46. The system of any one of embodiments 1 to 45, wherein the metadata processor, when encoding the index of the metadata parameter using differential coding, sets an absolute coding flag to 0 and writes a zero coding flag following the absolute coding flag, signaling whether the difference between the index of the current frame and the index of the previous frame is 0.

[0194] Embodiment 47. The system of any one of embodiments 1 to 46, wherein if the difference between the index of the current frame and the index of the previous frame is not equal to 0, the metadata processor continues coding by writing a sign flag followed by an adaptive-bits difference index.

[0195] Embodiment 48. The system of any one of embodiments 1 to 47, wherein the metadata processor uses in-object metadata coding logic to limit the range of variation in the metadata bit budget between frames and prevent the bit budget remaining for core coding from becoming too small.

[0196] Embodiment 49. The system of any one of embodiments 1 to 48, wherein the metadata processor limits the use of absolute coding in a given frame to only one metadata parameter, or to as few metadata parameters as possible, according to the coding logic of the metadata in the object.

[0197] Embodiment 50. The system of any one of embodiments 1 to 49, wherein the metadata processor avoids absolute coding of an index of a metadata parameter if an index of one metadata parameter's coding logic has already been coded using absolute coding in the same frame according to the coding logic of the metadata in the object.

[0198] Embodiment 51. The system of any one of embodiments 1 to 50, wherein the coding logic of the metadata in the object is bit rate dependent.

[0199] Embodiment 52. The system of any one of embodiments 1 to 51, wherein the metadata processor uses inter-object metadata coding logic that is used between coding of metadata of different objects to minimize the number of absolutely coded metadata parameters of different audio objects in the current frame.

[0200] Embodiment 53. The system of any one of embodiments 1 to 52, wherein the metadata processor controls frame counters for absolutely coded metadata parameters using inter-object metadata coding logic.

[0201] Embodiment 54. The system of any one of embodiments 1 to 53, wherein the metadata processor uses inter-object metadata coding logic to, when the metadata parameters of the audio objects evolve slowly and smoothly, (a) code an index of a first metadata parameter of a first audio object using absolute coding at frame M, (b) code an index of a second metadata parameter of the first audio object using absolute coding at frame M+1, (c) code an index of a first metadata parameter of a second audio object using absolute coding at frame M+2, and (d) code an index of a second metadata parameter of a second audio object using absolute coding at frame M+3.

[0202] Embodiment 55. The system of any one of embodiments 1 to 54, wherein the coding logic for the inter-object metadata is bit rate dependent.

[0203] Embodiment 56. The system of any one of embodiments 1 to 55, wherein the bit budget allocator uses a bit rate adaptive algorithm to allocate the bit budget for encoding the audio stream.

[0204] Embodiment 57. The system of any one of embodiments 1 to 56, wherein the bit budget allocator uses a bit rate adaptation algorithm to derive the total bit budget for the metadata from the total bit rate of the metadata or the total bit rate of the codec.

[0205] Embodiment 58. The system of any one of embodiments 1 to 57, wherein the bit budget allocator uses a bit rate adaptation algorithm to calculate the bit budget of an element by dividing the total bit budget of the metadata by the number of audio streams.

[0206] Embodiment 59. The system of any one of embodiments 1 to 58, wherein the bit budget allocator uses a bit rate adaptation algorithm to adjust the bit budget of the elements of the final audio stream to use all of the available metadata bit budget.

[0207] Embodiment 60. The system of any one of embodiments 1 to 59, wherein the bit budget allocator uses a bit rate adaptation algorithm to sum the bit budgets of the metadata of all audio objects and add the sum to the bit budget of the metadata common signaling to produce a side bit budget for the core coder.

[0208] Embodiment 61. The system of any one of embodiments 1 to 60, wherein the bit budget allocator uses a bit rate adaptation algorithm to (a) divide the core coder's side bit budget evenly among the audio objects, and (b) calculate the core coder's bit budget for each audio stream using the divided core coder's side bit budget and the element bit budgets.

[0209] Embodiment 62. The system of any one of embodiments 1 to 61, wherein the bit budget allocator uses a bit rate adaptation algorithm to adjust the core coder bit budget of the last audio stream to use all of the available core coder bit budget.

[0210] Embodiment 63. The system of any one of embodiments 1 to 62, wherein the bit budget allocator uses a bit rate adaptation algorithm to calculate a bit rate for encoding one audio stream in the core coder using the bit budget of the core coder.

[0211] Embodiment 64. The system of any one of embodiments 1 to 63, wherein the bit budget allocator uses a bit rate adaptation algorithm in inactive frames or frames with low energy to reduce the bit rate for encoding one audio stream in the core coder, set it to a constant value, and redistribute the saved bit budget among the audio streams in active frames.

[0212] Embodiment 65. The system of any one of embodiments 1 to 64, wherein the bit budget allocator uses a bit rate adaptation algorithm in active frames to adjust the bit rate for encoding one audio stream in the core coder based on metadata importance classification.

[0213] Embodiment 66. The system of any one of embodiments 1 to 65, wherein the bit budget allocator reduces the bit rate for encoding one audio stream in the core coder in inactive frames (VAD = 0) and redistributes the bit budget saved by the bit rate reduction among the audio streams in frames classified as active.

[0214] Embodiment 67. The system of any one of embodiments 1 to 66, wherein the bit budget allocator (a) sets a lower constant core coder bit budget for any audio stream with inactive content in a frame, (b) calculates a saved bit budget as the difference between the lower constant core coder bit budget and the core coder bit budget, and (c) reallocates the saved bit budget among the core coder bit budgets of the audio streams of the active frame.

[0215] Embodiment 68. The system of any one of embodiments 1 to 67, wherein the lower constant bit budget depends on the total bit rate of the metadata.

[0216] Embodiment 69. The system of any one of embodiments 1 to 68, wherein the bit budget allocator calculates a bit rate for encoding one audio stream in the core coder using a lower constant core coder bit budget.

[0217] Embodiment 70. The system of any one of embodiments 1 to 69, wherein the bit budget allocator uses core coder bit rate adaptation between objects based on metadata importance classification.

[0218] Embodiment 71. The system of any one of embodiments 1 to 70, wherein the importance of the metadata is based on an indicator of how important coding of a particular audio object in the current frame is for obtaining a satisfactory quality of the decoded synthesis.

[0219] Embodiment 72. The system of any one of embodiments 1 to 71, wherein the bit budget allocator classifies the importance of the metadata based on at least one of the following parameters: coder type (coder_type), FEC signal classification (class), human voice / music classification decision, and SNR estimates (snr_celp, snr_tcx) from the open-loop ACELP / TCX core decision module.

[0220] Embodiment 73. The system of any one of embodiments 1 to 72, wherein the bit budget allocator classifies the importance of the metadata based on the coder type (coder_type).

[0221] Embodiment 74. The bit budget allocator considers four different metadata importance classes: ISm ), i.e., - No Metadata Class ISM_NO_META: Frames without metadata coding, e.g. inactive frames with VAD = 0 - Low importance class ISM_LOW_IMP: frames with coder_type = UNVOICED or INACTIVE - Medium importance class ISM_MEDIUM_IMP: frames with coder_type = VOICED - High importance class ISM_HIGH_IMP: frames with coder_type = GENERIC 74. The system of any one of embodiments 1 to 73, wherein

[0222] Embodiment 75. The system of any one of embodiments 1 to 74, wherein the bit budget allocator uses metadata importance classes in a bit rate adaptation algorithm to allocate more bit budget to audio streams with higher importance and less bit budget to audio streams with lower importance.

[0223] Embodiment 76. The bit budget allocator implements the following logic in a frame: 1. class ISm = ISM_NO_META frames: Lower constant core coder bit rate is allocated. 2. class ISm = ISM_LOW_IMP frame: The bitrate (total_brate) for encoding one audio stream in the core coder is [number]total_brate new [n] = max(α low *total_brate[n], B low ) where the constant α low is set to a value less than 1.0, and the constant B low is the minimum bitrate threshold supported by the core coder. 3. class ISm = ISM_MEDIUM_IMP frame: The bitrate (total_brate) for encoding one audio stream in the core coder is [number]total_brate new [n] = max(α med *total_brate[n], B low ) where the constant α med is less than 1.0, but the value α low is set to a value greater than 4. class ISm Frames with ISM_HIGH_IMP = ISM_HIGH_IMP: Bitrate adaptation is not used 76. The system of any one of embodiments 1 to 75, wherein

[0224] Embodiment 77. The system of any one of embodiments 1 to 76, wherein the bit budget allocator redistributes the saved bit budget, expressed as the sum of the difference between the previous bit rate total_brate and the new bit rate total_brate, among the audio streams of frames classified as active.

[0225] Embodiment 78. A system for decoding an audio object according to an audio stream having associated metadata, comprising: a metadata processor for decoding metadata of an audio stream having active content; a bit budget allocator responsive to the bit budgets of the decoded metadata and each of the audio objects for determining a bit rate for a core coder of the audio stream; a decoder for the audio stream using the bit rate of the core coder determined in the bit budget allocator.

[0226] Embodiment 79. The system of embodiment 78, wherein the metadata processor is responsive to metadata common signaling read from the end of the received bitstream.

[0227] Embodiment 80. The system of embodiment 78 or 79, wherein the decoder includes a core decoder for decoding the audio stream.

[0228] Embodiment 81. The system of any one of embodiments 78 to 80, wherein the core decoder includes a variable bit rate core decoder for sequentially decoding the audio streams at the bit rate of their respective core coders.

[0229] Embodiment 82. The system of any one of embodiments 78 to 81, wherein the number of audio objects to be decoded is less than the number of core decoders.

[0230] Embodiment 83. The system of any one of embodiments 78 to 83, including a renderer of audio objects responsive to the decoded audio stream and the decoded metadata.

[0231] Any of embodiments 2 to 77 that further describe elements of embodiments 78 to 83 may be implemented in any of these embodiments 78 to 83. As an example, the bit rate of the core coder per audio stream in the decoding system is determined using the same procedure as in the coding system.

[0232] The present invention also relates to a coding method and a decoding method. In this respect, system embodiments 1 to 83 may be drafted as method embodiments in which elements of the system embodiments are replaced by operations performed by such elements. [Explanation of symbols]

[0233] 100 systems 101 Input Buffer 102 Input Audio Objects 103 Audio Stream Processor 104 Transport Channels 105 Metadata Processor 106 Configuration and Decision Processor 107 Information 108 Preprocessor 109 Core Encoder 110 Multiplexer 111 Output Bitstream 112 Quantized and Encoded Metadata 113 ISm Common Signaling 114 N audio streams 120 Signal classification information 121 line 150 methods 151 Input buffering behavior 153 Analysis and forward preprocessing operations 155 Metadata Analysis, Quantization, and Coding Operations 156 Composition and Decision-Making Behavior 158 Further Preprocessing; Preprocessing Actions 159 Core Coding in Action 160 Multiplexing Operation 700 decoder; decoding system 701 Bitstream 702 output audio channels 703 decoded audio stream 704 Decoded Metadata 705 Demultiplexer 706 Metadata Decoding and Inverse Quantization Processor 707 Configuration and Decision Processor 708 lines 709 Output Settings 710 Core Decoder 711 Renderer 712 Output Settings 750 method 755 Demultiplexing Operation 756 Metadata Decoding and Dequantization Operations 757 Configuration and decision behavior for bit rates per channel 760 Core Decryption Operation 761 Audio Channel Rendering Behavior 1200 Coding and Decoding System 1202 Input 1204 Output 1206 processor 1208 memory

Claims

1. 1. A system for coding an object-based audio signal including audio objects according to an audio stream having associated metadata, comprising: the metadata for each audio object includes azimuth and elevation parameters, and the audio objects are processed for successive frames; an audio stream processor for analyzing the audio stream to extract information about the audio stream; a metadata processor responsive to the information about the audio stream from the analysis by the audio stream processor for coding the metadata, the metadata processor using metadata coding logic to control the use of absolute coding of the metadata parameters on a frame, audio object, and metadata parameter basis to determine the metadata parameters to be coded using absolute coding in order to limit (a) the range of variation in metadata coding bit budget between frames and (b) the number of metadata parameters lost from the audio objects coded using absolute coding when a frame is lost; an encoder for coding the audio stream; Including, the system.

2. the metadata processor: The system of claim 1 , wherein the coding logic of the metadata within the object is used as the coding logic of the metadata to limit absolute coding at a given frame for each audio object to one metadata parameter.

3. the metadata processor:

2. The system of claim 1, wherein the coding logic of metadata within an object is used as the coding logic of metadata to avoid absolute coding of one of the azimuth angle parameter and the elevation angle parameter of an audio object in the same frame if the other of the azimuth angle parameter and the elevation angle parameter of the same audio object has already been coded using absolute coding.

4. the metadata processor:

2. The system of claim 1, wherein if the bit rate is sufficiently large, a coding logic for metadata in the bit rate dependent object is used as the coding logic for metadata to enable absolute coding of multiple metadata parameters in the same frame.

5. the metadata processor:

2. The system of claim 1, wherein inter-object metadata coding logic is applied as metadata coding logic to coding of metadata of different audio objects in a current frame, in order to minimize the number of metadata parameters of different audio objects that are coded using absolute coding.

6. the metadata processor: The system of claim 5 , wherein the inter-object metadata coding logic is used to control frame counters for metadata parameters coded using absolute coding.

7. the metadata processor:

7. The system of claim 5 or 6, wherein the inter-object metadata coding logic is used to code metadata parameters of one audio object per frame.

8. the metadata processor: Using said inter-object metadata coding logic, (a) coding a first one of the azimuth and elevation parameters of a first audio object using absolute coding at frame M; (b) coding a second one of the azimuth and elevation parameters of the first audio object using absolute coding at frame M+1; (c) coding the first of the azimuth and elevation parameters of a second audio object using absolute coding at frame M+2; 8. The system of claim 5, wherein (d) coding the second of the azimuth and elevation parameters of the second audio object using absolute coding in frame M+3.

9. 9. The system of claim 5, wherein the inter-object metadata coding logic depends on the bitrate to allow absolute coding of multiple metadata parameters of the audio object in the same frame if the bitrate is sufficiently large.

10. - the audio stream processor analyzing the audio stream to detect voice activity; - the metadata processor includes an analyzer of the metadata of the audio objects that uses voice activity detection from the audio stream processor to determine whether the current frame is inactive or active for each audio object; - during an inactive frame, the metadata processor does not code metadata associated with the audio object; 10. The system of claim 1, wherein in an active frame, the metadata processor codes the metadata relating to the audio object.

11. A system described in any one of claims 1 to 10, wherein the metadata processor includes a quantizer for an index of an azimuth angle parameter using an azimuth angle quantization step and an index of an elevation angle parameter using an elevation angle quantization step to quantize the azimuth angle parameter and the elevation angle parameter.

12. 12. The system of claim 11, wherein a total metadata bit budget for coding the metadata and a total number of quantization bits for quantizing the metadata parameter indices depend on a total codec bit rate, a total metadata bit rate, or a sum of a metadata bit budget and a core encoder bit budget associated with one audio object.

13. The metadata processor according to claim 1, wherein the metadata processor represents the azimuth parameter and the elevation parameter as a single parameter; 13. The system of claim 1, wherein the metadata processor includes a quantizer for the index of the one parameter.

14. the metadata processor includes a metadata parameter index quantizer using a quantization step to quantize the metadata parameters of the audio objects; 14. The system of claim 1, wherein the metadata processor includes a metadata encoder for coding indices of the quantized metadata parameters using either absolute coding or differential coding.

15. 15. The system of claim 14, wherein if metadata was not present in the previous frame, the metadata encoder codes the index of the quantized metadata parameter using absolute coding.

16. 16. The system of claim 14 or 15, wherein the metadata encoder codes the index of the quantized metadata parameter using absolute coding when the number of consecutive frames using differential coding is greater than the maximum number of consecutive frames coded using differential coding.

17. 17. The system of claim 14, wherein the metadata encoder, when coding an index of the quantized metadata parameter using absolute coding, generates an absolute coding flag followed by an index of the quantized metadata parameter coded using absolute coding, the absolute coding flag distinguishing between absolute coding and differential coding.

18. 18. The system of claim 17, wherein when the metadata encoder encodes the quantized metadata parameter index using differential coding, it sets the absolute coding flag to 0 and generates a zero coding flag following the absolute coding flag, the zero coding flag signaling a difference between the quantized metadata parameter index of a current frame and the quantized metadata parameter index of a previous frame, the difference being equal to 0.

19. 20. The system of claim 18, wherein if the difference between the quantized metadata parameter index of the current frame and the quantized metadata parameter index of the previous frame is not equal to 0, the metadata encoder generates a sign flag indicating a positive or negative sign of the difference followed by a difference index indicating the value of the difference.

20. the metadata processor outputs information about a bit budget for coding the metadata of the audio object; 20. The system of claim 1, further comprising a bit budget allocator responsive to the information about the bit budget for the coding of the metadata of the audio objects from the metadata processor for allocating a bit rate for coding of the audio stream.

21. 21. The system of claim 20, wherein the bit budget allocator sums the bit budgets for the coding of the metadata of the audio objects and adds the sum to a bit budget for signaling to perform bit rate distribution among the audio streams.

22. 1. A method for coding an object-based audio signal comprising audio objects according to an audio stream with associated metadata, comprising: the metadata for each audio object includes azimuth and elevation parameters, and the audio objects are processed for successive frames; analyzing the audio stream to extract information about the audio stream; (i) the information about the audio stream from the analysis of the audio stream; and (ii) metadata coding logic for controlling the use of absolute coding of the metadata parameters on a frame, audio object, and metadata parameter basis to determine the metadata parameters to be coded using absolute coding in order to limit (a) the range of variation in the bit budget of metadata coding between frames and (b) the number of metadata parameters lost from the audio object coded using absolute coding when a frame is lost; and coding the metadata using encoding the audio stream; A method comprising:

23. 23. The method of claim 22, wherein the metadata coding logic is intra-object metadata coding logic for limiting absolute coding at a given frame for each audio object to one metadata parameter.

24. 23. The method of claim 22, wherein the metadata coding logic is intra-object metadata coding logic to avoid absolute coding of one of the azimuth angle parameter and the elevation angle parameter of an audio object in the same frame if the other of the azimuth angle parameter and the elevation angle parameter of the same audio object has already been coded using absolute coding.

25. 23. The method of claim 22, wherein the metadata coding logic is a bitrate-dependent intra-object metadata coding logic to allow absolute coding of multiple metadata parameters in the same frame if the bitrate is sufficiently large.

26. 23. The method of claim 22, wherein using the metadata coding logic comprises using inter-object metadata coding logic for coding metadata of different audio objects to minimize a number of metadata parameters of different audio objects that are coded using absolute coding in the current frame.

27. 27. The method of claim 26, wherein using the inter-object metadata coding logic includes controlling a frame counter for metadata parameters coded using absolute coding.

28. 28. The method of claim 26 or 27, wherein using the inter-object metadata coding logic comprises coding metadata parameters of one audio object per frame.

29. using said inter-object metadata coding logic, (a) coding a first one of the azimuth and elevation parameters of a first audio object using absolute coding at frame M; (b) coding a second one of the azimuth and elevation parameters of the first audio object using absolute coding at frame M+1; (c) coding the first of the azimuth and elevation parameters of a second audio object using absolute coding at frame M+2; 29. The method of claim 26, comprising: (d) coding the second of the azimuth and elevation parameters of the second audio object using absolute coding in frame M+3.

30. 30. The method of any one of claims 26 to 29, wherein the inter-object metadata coding logic depends on the bitrate to allow absolute coding of multiple metadata parameters of the audio object in the same frame if the bitrate is sufficiently large.

31. - analyzing the audio stream to detect voice activity; - analyzing the metadata of each audio object using the voice activity detection to determine whether the current frame is inactive or active for that audio object; - not encoding metadata associated with said audio object in inactive frames; - encoding, in an active frame, said metadata relating to said audio object; 31. The method of any one of claims 22 to 30, comprising:

32. A method described in any one of claims 22 to 31, comprising quantizing an index of an azimuth angle parameter using an azimuth angle quantization step and quantizing an index of an elevation angle parameter using an elevation angle quantization step to quantize the azimuth angle parameter and the elevation angle parameter.

33. 33. The method of claim 32, wherein a total metadata bit budget for coding the metadata and a total number of quantization bits for quantizing the metadata parameter indices depend on a total codec bit rate, a total metadata bit rate, or a sum of a metadata bit budget and a core encoder bit budget associated with one audio object.

34. The method of claim 33, further comprising expressing the azimuth angle parameter and the elevation angle parameter as a single parameter; quantizing the index of the one parameter; 34. The method of any one of claims 22 to 33, comprising:

35. quantizing the indices of the metadata parameters using a quantization step to quantize the metadata parameters of the audio objects; coding the index of the quantized metadata parameter using either absolute coding or differential coding; 32. The method of any one of claims 22 to 31, comprising:

36. 36. The method of claim 35, wherein coding the index of the quantized metadata parameter comprises using absolute coding if metadata was not present in the previous frame.

37. 37. The method of claim 35 or 36, wherein the step of coding an index of the quantized metadata parameter comprises using absolute coding when the number of consecutive frames using differential coding is greater than the maximum number of consecutive frames coded using differential coding.

38. 38. The method of claim 35, wherein the step of coding the index of the quantized metadata parameter using absolute coding comprises generating an absolute coding flag followed by the index of the quantized metadata parameter coded using absolute coding, the absolute coding flag distinguishing between absolute coding and differential coding.

39. 39. The method of claim 38, wherein coding the quantized metadata parameter index using differential coding comprises: setting the absolute coding flag to 0; and generating a zero coding flag following the absolute coding flag, the zero coding flag signaling a difference between the quantized metadata parameter index of a current frame and the quantized metadata parameter index of a previous frame, the difference being equal to 0.

40. 40. The method of claim 39, wherein coding the quantized metadata parameter index using differential coding comprises generating a sign flag indicating a positive or negative sign of the difference followed by a difference index indicating the value of the difference if the difference between the quantized metadata parameter index of the current frame and the quantized metadata parameter index of the previous frame is not equal to 0.

41. the step of coding the metadata includes outputting information about a bit budget for coding the metadata of the audio object; 41. The method of claim 22, wherein the method comprises allocating a bit budget responsive to information about the bit budget for the coding of the metadata of the audio objects for allocating a bit rate for the coding of the audio stream.

42. 42. The method of claim 41 , wherein the allocating bit budget comprises summing the bit budgets for the coding of the metadata of the audio objects and adding the sum to a bit budget for signaling to perform bit rate distribution among the audio streams.

Citation Information

Patent Citations

  • Acoustic signal auxiliary information conversion transmission apparatus and program

    JP2019003185A

  • PCT/CA2018/51175

  • Exploiting metadata redundancy in immersive audio metadata

    US20170013387A1

  • Encoding device and method, decoding device and method, and program

    WO2014192602A1

  • Audio encoding device and audio decoding device

    WO2015056383A1