Method and system for coding metadata in an audio stream and for efficient bitrate allocation to coding of the audio stream

The system addresses the challenges of encoding and decoding object-based audio signals by using a metadata processor and bit-budget allocator to efficiently manage metadata and bitrate, resulting in high-quality immersive audio experiences.

JP7699095B2Active Publication Date: 2025-06-26VOICEAGE CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022500962
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-08
Filing Date
2020-07-07
Publication Date
2025-06-26
Estimated Expiration
2040-07-07

AI Technical Summary

Technical Problem

Current audio coding technologies face challenges in efficiently encoding and decoding object-based audio signals, which require precise metadata handling and bitrate allocation to maintain immersive audio experiences.

Method used

A system and method for coding and decoding object-based audio signals that include a metadata processor for coding metadata, a bit-budget allocator for determining the bitrate for the audio stream, and a decoder that uses the allocated bitrates to decode the audio objects.

Benefits of technology

This approach enables efficient encoding and decoding of object-based audio signals, ensuring high-quality immersive audio experiences by effectively managing metadata and bitrate allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699095000011
    Figure 0007699095000011
  • Figure 0007699095000012
    Figure 0007699095000012
  • Figure 0007699095000013
    Figure 0007699095000013
Patent Text Reader

Abstract

A system and method for coding an object-based audio signal including audio objects according to an audio stream having associated metadata, wherein a metadata processor codes the metadata and generates information about a bit budget for coding the metadata of the audio objects, an encoder codes the audio stream, and a bit budget allocator is responsive to the information about the bit budget for coding the metadata of the audio objects from the metadata processor to allocate a bit rate for coding the audio stream by the encoder.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to audio coding, and more particularly to techniques for digitally coding object-based audio such as, for example, human voice, music, or general audio voice. In particular, the present disclosure relates to systems and methods for coding and decoding object-based audio signals including audio objects according to an audio stream having associated metadata.

[0002] In the present disclosure and the appended claims,

[0003] (a) The term "object-based audio" is intended to represent a complex auditory scene of audio as a collection of individual elements also known as audio objects. Also, as shown above in this specification, "object-based audio" may include, for example, human voice, music, or general audio voice.

[0004] (b) The term "audio object" is intended to refer to an audio stream having associated metadata. For example, in the present disclosure, an "audio object" is referred to as an independent audio stream with metadata (ISm).

[0005] (c) The term "audio stream" is intended to represent an audio waveform such as human voice, music, or general audio voice within a bitstream, and may consist of one channel (mono), but two channels (stereo) may also be considered. "Mono" is an abbreviation for "monophonic", and "stereo" is an abbreviation for "stereophonic".

[0006] (d) The term "metadata" is intended to represent a set of information that describes an audio stream and artistic intent and is used to convey an original or coded audio object to a playback system. Typically, metadata describes the spatial characteristics of each individual audio object, such as position, orientation, volume, width, etc. In the context of the present disclosure, two sets of metadata are considered. - Input metadata: An unquantized metadata representation used as input to a codec. The present disclosure is not limited to a specific format of input metadata. And - Coded metadata: Quantized and coded metadata that forms part of a bitstream transmitted from an encoder to a decoder.

[0007] (e) The term "audio format" is intended to refer to a technique for realizing an immersive audio experience.

[0008] (f) The term "playback system" is intended to refer to an element within a decoder that, on the playback side, can use the transmitted metadata and artistic intent to render audio objects within a 3D (three-dimensional) audio space around the listener. Rendering can be performed for a target loudspeaker layout (e.g., 5.1 surround) or headphones, while the metadata can be dynamically modified, for example, in response to feedback from a head tracking device. Other types of rendering may be envisioned. BACKGROUND ART

[0009] In recent years, the generation, recording, representation, coding, transmission, and playback of audio have been moving towards enhanced, interactive, immersive experiences for the listener. An immersive experience can be described as being deeply involved and participating in an audio scene, for example, with voices heard from all directions. In immersive audio (also called 3D audio), the sound image is reproduced in all three dimensions around the listener, taking into account a wide range of sound characteristics such as timbre, directivity, reverberation, transparency, and the accuracy of (auditory) spread. Immersive audio is generated for a given playback system, i.e., a loudspeaker configuration, an all-in-one playback system (soundbar), or headphones. At that time, the interactivity of an audio playback system can include, for example, the ability to adjust the level of the voice, change the position of the voice, or select a different language for playback.

[0010] There are three basic techniques (hereinafter also referred to as audio formats) for realizing an immersive audio experience.

[0011] The first technique is channel-based audio where multiple spaced microphones are used to capture voices from different directions, while one microphone corresponds to one audio channel in a specific loudspeaker layout. Each recorded channel is fed to a loudspeaker at a specific position. Examples of channel-based audio include, for example, stereo, 5.1 surround, 5.1.4, etc.

[0012] The second technique is scene-based audio where a desired sound field in a localized space is represented as a function of time by a combination of dimensional components. The signal representing scene-based audio is independent of the position of the sound source, while the sound field must be converted to the loudspeaker layout selected in the rendering playback system. An example of scene-based audio is ambisonics.

[0013] The last, third immersive audio approach is object-based audio that represents an auditory scene as a set of those audio elements, for example, with information about the positions of those audio elements within the audio scene, such that individual audio elements (e.g., singer, drums, guitar) can be rendered at their intended positions within the playback system. This gives object-based audio high flexibility and interactivity as each object can be kept separately and manipulated individually.

[0014] Each of the above-described audio formats has its respective advantages and disadvantages. Thus, it is common that not just one particular format is used in an audio system, but they may be combined in a complex audio system to generate an immersive auditory scene. An example is a system that combines scene-based or channel-based audio with object-based audio, for example, combining ambisonics with several separate audio objects.

Prior Art Documents

Patent Documents

[0015]

Patent Document 1

Non-Patent Documents

[0016]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0017] The present disclosure presents a framework for encoding and decoding object-based audio in the following description. Such a framework can be an independent system for coding object-based audio formats or can form part of a complex immersive codec that may include coding of other audio formats and / or combinations thereof. **Means for Solving the Problems**

[0018] According to a first aspect, the present disclosure provides a system for coding an object-based audio signal including audio objects in response to an audio stream having associated metadata, the system including a metadata processor for coding the metadata, the metadata processor generating information about a bit-budget for coding the metadata of the audio objects. An encoder codes the audio stream, and a bit-budget allocator responds to the information about the bit-budget for coding the metadata of the audio objects from the metadata processor to allocate a bitrate for coding the audio stream by the encoder.

[0019] The present disclosure also provides a method for coding an object-based audio signal including audio objects in response to an audio stream having associated metadata, the method including coding the metadata, generating information about a bit-budget for coding the metadata of the audio objects, encoding the audio stream, and allocating a bitrate for coding the audio stream in response to the information about the bit-budget for coding the metadata of the audio objects.

[0020] According to a third aspect, there is provided a system for decoding an audio object according to an audio stream having associated metadata, the system comprising: a metadata processor for decoding the metadata of the audio object and providing information for each bit budget of the metadata of the audio object; a bit budget allocator responsive to the bit budget of the metadata of the audio object for determining the bit rate of a core-decoder of the audio stream; and a decoder of the audio stream that uses the bit rate of the core-decoder determined by the bit budget allocator.

[0021] The present disclosure further provides a method for decoding an audio object according to an audio stream having associated metadata, the method comprising: decoding the metadata of the audio object and providing information for each bit budget of the metadata of the audio object; determining the bit rate of a core-decoder of the audio stream using the bit budget of the metadata of the audio object; and decoding the audio stream using the determined bit rate of the core-decoder.

[0022] The above and other objects, advantages, and features of the systems and methods for coding an object-based audio signal and the systems and methods for decoding an object-based audio signal will become more apparent from the following non-limiting description of exemplary embodiments of those systems and methods given by way of example with reference to the accompanying drawings.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

DETAILED DESCRIPTION OF THE INVENTION

[0024] The present disclosure provides an example of a mechanism for coding metadata. The present disclosure also provides a mechanism for bitrate adaptation within and between flexible objects, i.e., a mechanism for distributing the available bitrate as efficiently as possible. In the present disclosure, it is further considered that the bitrate is fixed (constant). However, for example, it is also within the scope of the present disclosure to consider an adaptive bitrate as a result of coding a combination of (a) in an adaptive bitrate-based codec or (b) other audio formats coded in some other way at a fixed total bitrate.

[0025] In the present disclosure, there is no description of how an audio stream is actually coded in a so-called "core encoder". Generally, a core encoder for coding one audio stream can be any mono codec using adaptive bitrate coding. An example is a codec based on the EVS codec described in reference [1], which uses a varying bit budget that is flexibly and efficiently distributed among the modules of the core encoder, as described in reference [2] for example. The entire contents of references [1] and [2] are incorporated herein by reference.

[0026] 1. Framework for Coding Audio Objects As a non-limiting example, the present disclosure considers a framework that supports the simultaneous coding of several audio objects (e.g., up to 16 audio objects) while taking into account a fixed total bitrate of ISm called ism_total_brate for coding an audio object that includes an audio stream with associated metadata. Note that for non-diegetic content, for example, metadata is not necessarily transmitted for at least a part of the audio object. Non-diegetic audio in movies, television shows, and other videos is audio that the characters cannot hear. A soundtrack is an example of non-diegetic audio since only the viewer listens to the music.

[0027] When coding a combination of audio formats in the framework, for example, for a combination of an ambisonics audio format with two audio objects, a fixed total codec bitrate called codec_total_brate represents the sum of the bitrate of the ambisonics audio format (i.e., the bitrate for encoding the ambisonics audio format) and the total bitrate ism_total_brate of the ISm (i.e., the sum of the bitrates for coding the audio objects, i.e., the audio streams with associated metadata).

[0028] The present disclosure considers a basic and non-limiting example of input metadata consisting of two parameters, namely azimuth and elevation, that are stored for each object per audio frame. In this example, an azimuth range of [-180°, 180°] and an elevation range of [-90°, 90°] are considered. However, it is within the scope of the present disclosure to consider only one or more than two metadata parameters.

[0029] 2. Object-based coding FIG. 1 is a schematic block diagram showing simultaneously a system 100 including several processing blocks for coding an object-based audio signal and a corresponding method 150 for coding an object-based audio signal.

[0030] 2.1 Input buffering Referring to FIG. 1, a method 150 for coding an object-based audio signal includes an operation 151 of input buffering. To perform the operation 151 of input buffering, a system 100 for coding an object-based audio signal includes an input buffer 101.

[0031] The input buffer 101 buffers N input audio objects 102, i.e., N audio streams each having associated respective N metadata. The N audio streams and the N metadata associated with each of these N audio streams in the N input audio objects 102 are buffered for a frame, e.g., a frame of 20 ms length. As is well known in the field of audio signal processing, an audio signal is sampled at a given sampling frequency and processed for each consecutive block of these samples called a "frame" which is divided into several "sub-frames".

[0032] 2.2 Analysis and front pre-processing of audio streams Continuing to refer to FIG. 1, a method 150 for coding an object-based audio signal includes an operation 153 of analyzing and pre-processing N audio streams. To perform operation 153, a system 100 for coding an object-based audio signal includes an audio stream processor 103 that analyzes and pre-processes, for example, in parallel, the N buffered audio streams transmitted from an input buffer 101 to the audio stream processor 103 via N transport channels 104, respectively.

[0033] The analysis and pre-processing operations 153 performed by the audio stream processor 103 may include, for example, the following sub-operations, namely, transient detection in the time domain, spectral analysis, long-term prediction analysis, pitch tracking and voicing analysis, voice activity detection / sound activity detection (VAD / SAD), bandwidth detection, noise estimation, and signal classification (which may include, in non-limiting embodiments, (a) selection of a core encoder, for example, from an ACELP core encoder, a TCX core encoder, an HQ core encoder, etc., (b) classification of signal types such as inactive core encoder types, unvoiced core encoder types, voiced core encoder types, general-purpose core encoder types, transition core encoder types, and audio core encoder types, (c) classification of human voice / music, etc.). The information obtained from the analysis and pre-processing operations 153 is supplied to the configuration and decision processor 106 via line 121. Examples of the above sub-operations are described in reference [1] in relation to the EVS codec and are therefore not further described in the present disclosure.

[0034] 2.3 Analysis, Quantization, and Coding of Metadata The method 150 of Figure 1 for coding an object-based audio signal includes operations 155 of metadata analysis, quantization, and coding. To perform operation 155, a system 100 for coding an object-based audio signal includes a metadata processor 105.

[0035] 2.3.1 Metadata Analysis Signal classification information 120 (e.g., the VAD or localVAD flag used in the EVS codec (see reference [1])) from the audio stream processor 103 is supplied to the metadata processor 105. The metadata processor 105 includes an analyzer (not shown) for each of the N audio objects to determine whether the current frame is inactive (e.g., VAD = 0) or active (e.g., VAC≠0) with respect to this particular audio object. In an inactive frame, the metadata is not coded by the metadata processor 105 in relation to that object. In an active frame, the metadata is quantized and coded for this audio object using a variable bit rate. Further details regarding the quantization and coding of the metadata are given in Sections 2.3.2 and 2.3.3 below.

[0036] 2.3.2 Metadata Quantization In the non-limiting exemplary embodiment being described, the metadata processor 105 of Figure 1 quantizes and codes the metadata of the N audio objects in sequence in a loop, while specific dependencies can be used between the quantization of the audio objects and their metadata parameters.

[0037] As shown above in this specification, in the present disclosure, two metadata parameters, the azimuth angle and the elevation angle (included in N input metadata), are considered. As a non-limiting example, the metadata processor 105 includes a quantizer (not shown) of the indices of the following exemplary metadata parameters that uses the following exemplary resolutions to reduce the number of bits used. - Azimuth parameter: The index of the 12-bit azimuth parameter from the input metadata file is quantized to B az bits of index (e.g., B az = 7). Given the minimum azimuth limit and the maximum azimuth limit (-180° and +180°), the quantization step of the (B az = 7)-bit uniform scalar quantizer is 2.835°. - Elevation parameter: The index of the 12-bit elevation parameter from the input metadata file is quantized to B el bits of index (e.g., B el = 6). Given the minimum elevation limit and the maximum elevation limit (-90° and +90°), the quantization step of the (B el = 6)-bit uniform scalar quantizer is 2.857°.

[0038] The total metadata bit budget for coding N metadata and the total number of quantization bits for quantizing the indices of the metadata parameters (i.e., the granularity and thus the resolution of the quantization index) may be made dependent on the bitrates codec_total_brate, ism_total_brate, and / or element_brate (the last one resulting from the sum of the metadata bit budget and / or the core encoder bit budget related to one audio object).

[0039] The azimuth and elevation parameters can be represented as one parameter by a point on a sphere, for example. In such a case, implementing different metadata including two or more parameters is within the scope of the present disclosure.

[0040] 2.3.3 Coding of Metadata Both the azimuth index and the elevation index, when quantized, can be coded by an encoder (not shown) of the metadata processor 105 using either absolute coding or differential coding. As is known, absolute coding means that the current value of a parameter is coded. Differential coding means that the difference between the current value and the previous value of a parameter is coded. Since the indices of the azimuth parameter and the elevation parameter usually evolve smoothly (i.e., the change in the azimuth or elevation position can be considered continuous and smooth), differential coding is used by default. However, for example, absolute coding may be used in the following cases. - The difference between the current value and the previous value of the parameter index is too large, resulting in more or equal number of bits for using differential coding compared to using absolute coding (which may occur exceptionally). - In the previous frame, the metadata was not coded and not transmitted. - There are too many consecutive frames using differential coding. To control decoding in a noisy channel (Bad Frame Indicator, BFI = 1). For example, if the number of consecutive frames coded using the difference exceeds the maximum number of consecutive frames coded using differential coding, the metadata encoder codes the index of the metadata parameter using absolute coding. The maximum number of consecutive frames is set to β. In an example for illustrative purposes, β = 10 frames.

[0041] The metadata encoder generates a 1-bit absolute coding flag flag to distinguish between absolute coding and differential coding. abs is generated.

[0042] In the case of absolute coding, the coding flag flag abs is set to 1, and then followed by the index of the B az bits (or B el bits) coded using absolute coding, where B az and B el each refer to the above-mentioned indices of the azimuth and elevation parameters to be coded.

[0043] In the case of differential coding, the 1-bit coding flag flag abs is set to 0, and then followed by a 1-bit zero coding flag flag az that signals the difference Δ between the indices of the B el bits (or the indices of the B zero bits) in the current frame and the previous frame, which is equal to 0. If the difference Δ is not equal to 0, the metadata encoder continues the coding by generating, for example, a 1-bit sign flag flag sign followed by a differentially adaptive bit-index that indicates the value of the difference Δ in the form of a unary code.

[0044] Figure 2 is a diagram showing different scenarios of coding the bit stream of one metadata parameter.

[0045] Referring to Figure 2, it is noted that not all metadata parameters are always transmitted in every frame. Some may be transmitted only every y frames, and some may not be transmitted at all, for example, when they do not change, they are not important, or when the available bit budget is low. Referring to Figure 2, for example,

[0046] - In the case of absolute coding (the first line of FIG. 2), the absolute coding flag flag abs and B az bit index (or B el bit index) is transmitted.

[0047] - In the case of differential coding where the difference Δ between the B az bit index (or B el bit index) in the current frame and the previous frame is equal to 0 (the second line of FIG. 2), the absolute coding flag flag abs = 0 and the zero coding flag flag zero = 1 are transmitted.

[0048] - In the case of differential coding where there is a positive difference Δ between the B az bit index (or B el bit index) in the current frame and the previous frame (the third line of FIG. 2), the absolute coding flag flag abs = 0, the zero coding flag flag zero = 0, the sign flag flag sign = 0, and the difference index (from 1 to (B az - 3) bit index (or from 1 to (B el - 3) bit index)) are transmitted. And

[0049] - In the case of differential coding where there is a negative difference Δ between the B az bit index (or B el bit index) in the current frame and the previous frame (the last line of FIG. 2), the absolute coding flag flag abs = 0, the zero coding flag flag zero = 0, the sign flag flag sign = 1, and the difference index (from 1 to (B az - 3) bit index (or from 1 to (B el - 3) bit index)) are transmitted.

[0050] 2.3.3.1 Coding Logic for Metadata within an Object The logic used to set absolute coding or differential coding may be further extended by the coding logic for metadata within an object. In particular, to limit the width of the variation of the bit budget for coding metadata between frames and thus prevent the bit budget left for the core encoder 109 from becoming too small, the metadata encoder restricts the absolute coding in a given frame to one or generally the fewest possible number of metadata parameters.

[0051] In a non - limiting example of coding azimuth and elevation metadata parameters, the metadata encoder uses logic to avoid absolute coding of the elevation index in a given frame if the azimuth index has already been coded using absolute coding in the same frame. In other words, (substantially) it is not the case that the azimuth and elevation parameters of one audio object are both coded using absolute coding in the same frame. As a result, when the absolute coding flag flag abs.azi for the azimuth parameter is equal to 1, the absolute coding flag flag abs.ele for the elevation parameter is not transmitted within the bitstream of the audio object.

[0052] It is also within the scope of the present disclosure to make the coding logic for metadata within an object dependent on the bitrate. For example, when the bitrate is high enough, both the absolute coding flag flag abs.ele for the elevation parameter and the absolute coding flag flag abs.azi for the azimuth parameter can be transmitted in the same frame.

[0053] 2.3.3.2 Coding Logic for Metadata between Objects The metadata encoder may apply similar logic to the coding of metadata for different audio objects. The coding logic of the metadata among the objects to be implemented minimizes the number of metadata parameters of different audio objects coded using absolute coding in the current frame. This is mainly achieved by the metadata encoder controlling the frame counter of the metadata parameters coded using absolute coding, which is selected mainly for robustness and is represented by the parameter β. As a non-limiting example, a scenario where the metadata parameters of an audio object develop slowly and smoothly is considered. To control decoding in a noisy channel where the index is coded using absolute coding every β frames, the azimuth B az bits of index of audio object #1 are coded using absolute coding in frame M, and the elevation B el bits of index of audio object #1 are coded using absolute coding in frame M + 1, the azimuth B az bits of index of audio object #2 are coded using absolute coding in frame M + 2, and the elevation B el bits of index of object #2 are coded using absolute coding in frame M + 3, and so on.

[0054] Figure 3a is a graph showing the values of the absolute coding flag for the metadata parameters of three audio objects when the coding logic of the metadata among the objects is not used, and Figure 3b is a graph showing the values of the absolute coding flag for the metadata parameters of three audio objects when the coding logic of the metadata among the objects is used. In Figure 3a, the arrows indicate the frames in which the values of some of the absolute coding flags are equal to 1. abs Figure 3b is a graph showing the values of the absolute coding flag for the metadata parameters of three audio objects when the coding logic of the metadata among the objects is used. In Figure 3a, the arrows indicate the frames in which the values of some of the absolute coding flags are equal to 1. abs Figure 3b is a graph showing the values of the absolute coding flag for the metadata parameters of three audio objects when the coding logic of the metadata among the objects is used. In Figure 3a, the arrows indicate the frames in which the values of some of the absolute coding flags are equal to 1.

[0055] More specifically, FIG. 3a shows the absolute coding flag flag for two metadata parameters (azimuth and elevation in this particular example) of an audio object when not using the coding of metadata between objects. abs FIG. 3b shows the same values, but with the coding logic of metadata between objects implemented. The graphs in FIGS. 3a and 3b correspond (from top to bottom) to the following. - The audio stream of audio object #1, - The audio stream of audio object #2, - The audio stream of audio object #3, - The absolute coding flag flag for the azimuth parameter of audio object #1 abs,azi 、 - The absolute coding flag flag for the elevation parameter of audio object #1 abs,ele 、 - The absolute coding flag flag for the azimuth parameter of audio object #2 abs,azi 、 - The absolute coding flag flag for the elevation parameter of audio object #2 abs,ele 、 - The absolute coding flag flag for the azimuth parameter of audio object #3 abs,azi 、and - The absolute coding flag flag for the elevation parameter of audio object #3 abs,ele 。

[0056] As can be seen from FIG. 3a, when the coding logic of metadata between objects is not used, some flags abs may have values equal to 1 in the same frame (see the arrow). In contrast, FIG. 3b shows that when the coding logic of metadata between objects is used, there is one absolute flag flag in a given frame absindicates that only may have a value equal to 1.

[0057] Also, the coding logic of the metadata between objects may be made dependent on the bit rate. In this case, for example, if the bit rate is large enough, even when the coding logic of the metadata between objects is used, two or more absolute flags flag abs in a given frame may have a value equal to 1.

[0058] The technical advantages of the coding logic of the metadata between objects and the coding of the metadata within an object are to limit the range of variation of the bit budget of the coding of the metadata between frames. Another technical advantage is to enhance the robustness of the codec in a noisy channel. When a frame is lost, only a limited number of metadata parameters from the audio object coded using absolute coding are lost. Therefore, any error propagated from the lost frame affects only a few metadata parameters of the entire audio object and thus does not affect the entire audio scene (or several different channels).

[0059] The overall technical advantage of analyzing, quantizing, and coding the metadata separately from the audio stream is, as described above, to enable more efficient processing in terms of being specifically adapted to the metadata, the bit rate of the coding of the metadata, the variation of the bit budget of the coding of the metadata, the robustness in a noisy channel, and the propagation of errors caused by lost frames.

[0060] The quantized and coded metadata 112 from the metadata processor 105 is supplied to the multiplexer 110 for insertion into the output bit stream 111 transmitted to the remote decoder 700 (FIG. 7).

[0061] When the metadata of N audio objects is analyzed, quantized, and encoded, information 107 from the metadata processor 105 about the bit budget for the coding of the metadata for each audio object is supplied from the metadata processor 105 to a configuration and decision processor 106 (bit budget allocator) with a configuration and decision process to be described in more detail in the following section 2.4. When the configuration and bit rate distribution among the audio streams are completed in the processor 106 (bit budget allocator), the coding continues with further preprocessing 158 to be described later. Finally, the N audio streams are encoded using an encoder including N variable bit rate core encoders 109 such as a mono core encoder.

[0062] 2.4 Configuration and Decision of Bit Rate per Channel The method 150 of FIG. 1 for coding an object-based audio signal includes an operation 156 of configuration and decision for the bit rate for each transport channel 104. To perform the operation 156, the system 100 for coding an object-based audio signal includes a configuration and decision processor 106 that forms a bit budget allocator.

[0063] The configuration and decision processor 106 (hereinafter, bit budget allocator 106) uses a bit rate adaptation algorithm to allocate the available bit budget for core-encoding N audio streams in N transport channels 104.

[0064] The bit rate adaptation algorithm of the configuration and decision operation 156 includes the following sub-operations 1 to 6 performed by the bit budget allocator 106.

[0065] 1. Total bit budget bits of ISm per frame ismis calculated from the total bitrate of the ISm, ism_total_brate (or, if only audio objects are coded, the total bitrate of the codec, codec_total_brate), for example, using the following relationship.

Number

[0066] 2. The bitrate of the element, element_brate, defined for N audio objects (obtained as the sum of the bit budget of the metadata related to one audio object and the bit budget of the core encoder) is assumed to be constant during the session and approximately the same for the N audio objects in the total bitrate of a given codec. The "session" is defined, for example, as a telephone call or the offline compression of an audio file. The bit budget of the corresponding element, bits element for the audio stream objects n = 0, ..., N-1, for example, using the following relationship

Number

Number

Number

[0067] 3. The bit-budget bits of the metadata for each frame of the N audio objects meta are totaled using the following relational expression

Number

[0068] 4. The side bit-budget bits of the codec for each frame side are evenly divided among the N audio objects, and the bit-budget bits of the core encoder for each of the N audio streams CoreCoder are, for example, set according to the following relational expression

Number

[0069] 5. The total bitrate total_brate in non-active frames (or frames with very low energy or otherwise meaningless content) may be reduced and set to a constant value in the relevant audio stream. And the bit budget thus saved is evenly redistributed among the audio streams with active content within the frame. Such redistribution of the bit budget is further explained in Section 2.4.1 below.

[0070] 6. The total bitrate total_brate in the audio streams (with active content) within an active frame is further adjusted among these audio streams based on the importance classification of ISm. Such adjustment of the bitrate is further explained in Section 2.4.2 below.

[0071] When the audio streams are all within a segment that is inactive (or has no meaningful content), the last two sub-operations 5 and 6 described above may be omitted. Therefore, the bitrate adaptation algorithms described in Sections 2.4.1 and 2.4.2 below are used when at least one audio stream has active content.

[0072] 2.4.1 Bitrate adaptation based on signal activity In inactive frames (VAD = 0), the total bitrate total_brate is reduced, and the saved bit budget is redistributed, for example, evenly among the audio streams of the active frames (VAD ≠ 0). On the premise that the coding of the waveform of the audio stream in the frames classified as inactive is not required, the audio object may be muted. The logic used in all frames can be represented by the following sub-operations 1 to 3.

[0073] 1. For a specific frame, set the bit budget of the smaller core encoder for any audio stream n having inactive content, bits CoreCoder '[n] = B VAD0 ∀n where VAD = 0 where B VAD0 is the lower fixed bit budget of the core encoder set in the inactive frame, for example, B VAD0 = 140 (corresponding to 7 kbps for a 20 ms frame) or B VAD0 = 49 (corresponding to 2.45 kbps for a 20 ms frame).

[0074] 2. Next, the saved bit budget is calculated using, for example, the following relational expression

Equation

[0075] 3. Finally, the saved bit budget is evenly redistributed among, for example, the bit budgets of the core encoders of the audio streams having active content within a given frame using the following relationship,

Number

Number

[0076] Figure 4 is a graph showing an example of bitrate adaptation for three core encoders. In particular, in Figure 4, the first row shows the total bitrate total_brate of the core encoder for audio stream #1, the second row shows the total bitrate total_brate of the core encoder for audio stream #2, the third row shows the total bitrate total_brate of the core encoder for audio stream #3, the fourth row is audio stream #1, the fifth row is audio stream #2, and the sixth row is audio stream #3.

[0077] In the example of Figure 4, the adaptation of the total bitrate total_brate of the three core encoders is based on VAD activity (active / inactive frames). As can be seen from Figure 4, in most cases, the varying side bit budget bits sideAs a result, there are minor fluctuations in the total bitrate total_brate of the core encoder. And as a result of VAD activity, there are rare significant changes in the total bitrate total_brate of the core encoder.

[0078] For example, referring to FIG. 4, case A) corresponds to a frame in which the VAD activity of audio stream #1 changes from 1 (active) to 0 (inactive). According to this logic, the minimum total bitrate total_brate of the core encoder is allocated to audio object #1, while the total bitrate total_brate of the core encoder for active audio objects #2 and #3 is increased. Case B) corresponds to a frame in which the VAD activity of audio stream #3 changes from 1 (active) to 0 (inactive), while the VAD activity of audio stream #1 remains 0. According to the logic, the minimum total bitrate total_brate of the core encoder is allocated to audio streams #1 and #3, while the total bitrate total_brate of the core encoder for active audio stream #2 is further increased.

[0079] The above-mentioned logic in section 2.4.1 can be made dependent on the total bitrate ism_total_brate. For example, the bit budget B in the above-mentioned lower operation 1 VAD0 can be set higher for a higher total bitrate ism_total_brate and lower for a lower total bitrate ism_total_brate.

[0080] 2.4.2 Adaptation of Bitrate Based on the Importance of ISm The logic described in the previous section 2.4.1 yields approximately the same core encoder bitrate in any audio stream that has active content (VAD = 1) within a given frame. However, it may be beneficial to introduce an adaptation of the core encoder bitrate between objects based on the classification of the importance of the ISm (or more generally, an indicator showing how important the coding of a particular audio object in the current frame is to obtain a given (satisfactory) quality of decoded synthesis).

[0081] The classification of the importance of the ISm can be obtained based on several parameters and / or combinations of parameters, such as the core encoder type (coder_type), FEC (Forward Error Correction), the classification of the audio signal (class), the determination of the classification of human voice / music, and / or the SNR (Signal-to-Noise Ratio) estimated values (snr_celp, snr_tcx) from the open-loop ACELP / TCX (Algebraic Code-Excited Linear Prediction / Transform Coded Excitation) core decision module described in reference [1]. Other parameters may potentially be used to determine the classification of the importance of the ISm.

[0082] In a non-limiting example, a simple classification of the importance of the ISm based on the core encoder type defined in reference [1] is implemented. For that purpose, the bit budget allocator 106 in FIG. 1 includes a classifier (not shown) for evaluating the importance of a particular ISm stream. As a result, four different ISm importance classes class ISm are defined. - No metadata class ISM_NO_META: A frame without metadata coding, e.g., an inactive frame with VAD = 0 - Low importance class ISM_LOW_IMP: A frame where coder_type = UNVOICED or INACTIVE - Medium importance class ISM_MEDIUM_IMP: frames where coder_type = VOICED - High importance class ISM_HIGH_IMP: frames where coder_type = GENERIC

[0083] At that time, the ISm importance class is used by the bit budget allocator 106 in a bit rate adaptation algorithm (see sub-operation 6 in section 2.4 above) to allocate a larger bit budget to the audio stream with a higher ISm importance and a lower bit budget to the audio stream with a lower ISm importance. Therefore, for every audio stream n, n = 0, ..., N-1, the following bit rate adaptation algorithm is used by the bit budget allocator 106. 1. class ISm = For frames classified as ISM_NO_META, a certain low bit rate B VAD0 is allocated. 2. class ISm = For frames classified as ISM_LOW_IMP, the total bit rate total_brate is reduced, for example, total_brate new [n] = max(α low *total_brate[n], B low ) where the constant α low is set to a value less than 1.0, for example, 0.6. And the constant B low represents the threshold of the minimum bit rate supported by the codec for a specific configuration, and this minimum bit rate threshold may depend on, for example, the internal sampling rate of the codec, the bandwidth of the audio being coded, etc. (see reference [1] for more details on these values). 3. class ISm= In frames classified as ISM_MEDIUM_IMP, the total bitrate total_brate of the core encoder is, for example, total_brate new [n] = max(α med *total_brate[n], B low ) is lowered as, for example, med where the constant α low is less than 1.0 but is set to a value greater than, for example, 0.8. 4. class ISm = In frames classified as ISM_HIGH_IMP, bitrate adaptation is not used. 5. Finally, the saved bit budget (the sum of the differences between the old total bitrate (total_brate) and the new total bitrate (total_brate new )) is evenly redistributed among the audio streams having active content within the frame. The same bit budget redistribution logic as described in sub-operations 2 and 3 of section 2.4.1 may be used.

[0084] Figure 5 is a graph showing an example of bitrate adaptation based on ISm importance logic. From top to bottom, the graph of Figure 5 synchronously shows the following. - The active speech segments of the audio stream for audio object #1, - The active speech segments of the audio stream for audio object #2, - The total bitrate total_brate of the audio stream for audio object #1 when not using the bitrate adaptation algorithm, - The total bitrate total_brate of the audio stream for audio object #2 when not using the bitrate adaptation algorithm, - Total bitrate total_brate of the audio stream related to audio object #1 when a bitrate adaptation algorithm is used, and - Total bitrate total_brate of the audio stream related to audio object #2 when a bitrate adaptation algorithm is used.

[0085] In the non-limiting example of FIG. 5, when using two audio objects (N = 2) and a fixed total bitrate ism_total_brate equal to 48 kbps, the total bitrate total_brate of the core encoder in the active frames of audio object #1 varies between 23.45 kbps and 23.65 kbps when no bitrate adaptation algorithm is used, while it varies between 19.15 kbps and 28.05 kbps when a bitrate adaptation algorithm is used. Similarly, the total bitrate total_brate of the core encoder in the active frames of audio object #2 varies between 23.40 kbps and 23.65 kbps when not using a bitrate adaptation algorithm, and varies between 19.10 kbps and 28.05 kbps when using a bitrate adaptation algorithm. Thereby, a better and more efficient distribution of the available bit budget between the audio streams is obtained.

[0086] 2.5 Pretreatment Referring to FIG. 1, a method 150 for coding an object-based audio signal includes an operation 158 of preprocessing N audio streams carried via N transport channels 104 from a configuration and determination processor 106 (bit budget allocator). To execute operation 158, a system 100 for coding an object-based audio signal includes a preprocessor 108.

[0087] Once the configuration and bitrate distribution among the N audio streams are completed by the configuration and decision processor 106 (bit budget allocator), the pre-processor 108 performs sequential further pre-processing 158 for each of the N audio streams. Such pre-processing 158 may include, for example, further signal classification, selection of a further core encoder (e.g., selection from an ACELP core, a TCX core, and an HQ core), different internal sampling frequencies F s at adapted bitrates for the core encoder, and other resampling, etc. Examples of such pre-processing can be found, for example, in reference [1] in relation to the EVS codec and are thus not further described in this disclosure.

[0088] 2.6 Core Encoding Referring to FIG. 1, a method 150 for coding an object-based audio signal includes an operation 159 of core encoding. To perform operation 159, a system 100 for coding an object-based audio signal includes, for example, the above-described encoders for N audio streams for respectively coding the N audio streams carried from the pre-processor 108 via N transport channels 104, including N core encoders 109.

[0089] In particular, the N audio streams are encoded using N variable bitrate core encoders 109, e.g., mono core encoders. The bitrate used by each of the N core encoders is the bitrate selected by the configuration and decision processor 106 (bit budget allocator) for the corresponding audio stream. For example, the core encoders described in reference [1] may be used as the core encoders 109.

[0090] 3.0 Bitstream Structure Referring to FIG. 1, a method 150 for coding an object-based audio signal includes a multiplexing operation 160. To perform the operation 160, a system 100 for coding an object-based audio signal includes a multiplexer 110.

[0091] FIG. 6 is a schematic diagram showing the structure of a bitstream 111 generated by the multiplexer 110 and transmitted from the coding system 100 of FIG. 1 to the decoding system 700 of FIG. 7 with respect to a frame. Whether or not there is metadata, the structure of the bitstream 111 may be assembled as shown in FIG. 6.

[0092] Referring to FIG. 6, while the multiplexer 110 writes the indexes of N audio streams from the beginning of the bitstream 111, the indexes of the ISm common signaling 113 from the configuration and determination processor 106 (bit budget allocator) and the metadata 112 from the metadata processor 105 are written from the end of the bitstream 111.

[0093] 3.1 ISm Common Signaling The multiplexer writes the ISm common signaling 113 from the end of the bitstream 111. The ISm common signaling is generated by the configuration and determination processor 106 (bit budget allocator) and includes bits of a variable representing the following.

[0094] (a) Number N of audio objects: Signaling regarding the number N of coded audio objects present in the bitstream 111 is, for example, in the form of a unary code having a stop bit (for example, for N = 3 audio objects, the first 3 bits of the ISm common signaling are "110").

[0095] (b) Metadata presence flag flag meta : Flag flag metaExists when bitrate adaptation based on signal activity as described in section 2.4.1 is used and the metadata for that particular audio object is present in bitstream 111 (flag meta = 1) or not present (flag meta = 0), and includes 1 bit per audio object to indicate which, or (c) ISm importance class: This signaling exists when bitrate adaptation based on the importance of ISM as described in section 2.4.2 is used and includes 2 bits per audio object to indicate the ISm importance class class ISm (ISM_NO_META, ISM_LOW_IMP, ISM_MEDIUM_IMP, ISM_HIGH_IMP) defined in section 2.4.2.

[0096] (d) ISm VAD flag flag VAD : The ISm VAD flag is sent when flag meta = 0 or class ISm = ISM_NO_META and distinguishes between the following two cases. 1) No input metadata exists or the metadata is not coded, so the audio stream needs to be coded in active coding mode (flag VAD = 1), and 2) Input metadata exists, is sent, so the audio stream can be coded in non - active coding mode (flag VAD = 0).

[0097] 3.2 Payload of Coded Metadata The multiplexer 110 is supplied with coded metadata 112 from the metadata processor 105 and metadata is coded in the current frame (flag meta = 1 or class ISmWrite the metadata payload in order from the end of the bitstream for the audio object (≠ISM_NO_META). The bit budget for the metadata for each audio object is not constant; rather, it is adaptive between objects and between frames. Scenarios of different metadata formats are shown in Figure 2.

[0098] If metadata does not exist or is not transmitted for at least some of the N audio objects, for these audio objects, the metadata flag is set to 0, that is, flag meta = 0, or class ISm = ISM_NO_META. At that time, the metadata index is not transmitted in relation to those audio objects, that is, bits meta [n] = 0.

[0099] 3.3 Payload of the Audio Stream The multiplexer 110 receives N audio streams 114 encoded by N core encoders 109 via N transport channels 104, and writes the payload of the audio stream in order in time series for the N audio streams starting from the beginning of the bitstream 111 (see Figure 6). The bit budget of each of the N audio streams varies as a result of the bitrate adaptation algorithm described in Section 2.4.

[0100] 4.0 Decoding of Audio Objects Figure 7 is a schematic block diagram showing simultaneously a system 700 for decoding audio objects according to an audio stream having associated metadata and a corresponding method 750 for decoding audio objects.

[0101] 4.1 Demultiplexing Referring to FIG. 7, a method 750 for decoding an audio object according to an audio stream having associated metadata includes an operation 755 of multiplex separation. To perform operation 755, a system 700 for decoding an audio object according to an audio stream having associated metadata includes a demultiplexer 705.

[0102] The demultiplexer receives a bitstream 701 transmitted from the coding system 100 of FIG. 1 to the decoding system 700 of FIG. 7. In particular, the bitstream 701 of FIG. 7 corresponds to the bitstream 111 of FIG. 1.

[0103] The demultiplexer 110 extracts from the bitstream 701: (a) N coded audio streams 114, (b) coded metadata 112 regarding N audio objects, and (c) ISm common signaling 113 read from the end of the received bitstream 701.

[0104] 4.2 Decoding and Inverse Quantization of Metadata Referring to FIG. 7, a method 750 for decoding an audio object according to an audio stream having associated metadata includes an operation 756 of decoding and inverse quantization of metadata. To perform operation 756, a system 700 for decoding an audio object according to an audio stream having associated metadata includes a metadata decoding and inverse quantization processor 706.

[0105] The metadata decoding and inverse quantization processor 706 is supplied with output settings 709 for decoding and inverse quantizing the coded metadata 112 regarding the transmitted audio object, the ISm common signaling 113, and the metadata regarding the audio stream / object having the active content. The output settings 709 are command line parameters regarding the number M of decoded audio objects / transport channels and / or audio formats, which can be equal to or different from the number N of coded audio objects / transport channels. The metadata decoding and inverse quantization processor 706 generates decoded metadata 704 regarding M audio objects / transport channels and supplies information regarding each bit budget for the M decoded metadata on line 708. Obviously, the decoding and inverse quantization executed by the processor 706 is the inverse of the quantization and coding executed by the metadata processor 105 in FIG. 1.

[0106] 4.3 Configuration and determination regarding bit rate Referring to FIG. 7, a method 750 for decoding an audio object according to an audio stream having related metadata includes an operation 757 of configuration and determination regarding the bit rate per channel. To execute the operation 757, a system 700 for decoding an audio object according to an audio stream having related metadata includes a configuration and determination processor 707 (bit budget allocator).

[0107] The bit budget allocator 707 includes (a) information regarding each bit budget for the M decoded metadata on line 708 and (b) the ISm importance class class from the common signaling 113 ISmReceives the above, and determines the bit rate total_brate[n] of the core decoder for each audio stream. The bit budget allocator 707 determines the bit rate of the core decoder using the same procedure as the bit budget allocator 106 in FIG. 1 (see section 2.4).

[0108] 4.4 Core Decoding Referring to FIG. 7, a method 750 for decoding an audio object according to an audio stream having associated metadata includes an operation 760 of core decoding. To perform operation 760, a system 700 for decoding an audio object according to an audio stream having associated metadata includes N core decoders 710, for example, decoders for N audio streams 114 of N variable bit rate core decoders.

[0109] The N audio streams 114 from the demultiplexer 705 are decoded, for example, in N variable bit rate core decoders 710, in order at their respective core decoder bit rates determined by the bit budget allocator 707. If the number M of decoded audio objects required by the output setting 709 is less than the number of transport channels, i.e., M < N, then a smaller number of core decoders are used. Similarly, in such a case, it is possible that not all metadata payloads are decoded.

[0110] Depending on the N audio streams 114 from the demultiplexer 705, the core decoder bit rates determined by the bit budget allocator 707, and the output setting 709, the core decoder 710 generates M decoded audio streams 703 on their respective M transport channels.

[0111] 5.0 Rendering of Audio Channels In the operation 761 of rendering the audio channel, the renderer 711 of the audio object converts M decoded metadata 704 and M decoded audio streams 703 into several output audio channels 702 in consideration of the output setting 712 indicating the number and content of the output audio channels to be generated. Again, the number of the output audio channels 702 may be equal to or different from the number M.

[0112] The renderer 711 may be designed in various different structures to obtain the desired output audio channels. Therefore, the renderer will not be further described in this disclosure.

[0113] 6.0 Source Code According to a non-limiting exemplary embodiment, the system and method for coding the object-based audio signal disclosed in the above description may be implemented by the following source code (represented in C code) given below as additional disclosure.

[0114] void ism_metadata_enc( const long ism_total_brate, / * i : Total bitrate of ISm * / const short n_ISms, / * i : Number of objects * / ISM_METADATA_HANDLE hIsmMeta[], / * i / o: Handle of ISM metadata * / ENC_HANDLE hSCE[], / * i / o: Handle of element encoder * / BSTR_ENC_HANDLE hBstr, / * i / o: Handle of bitstream * / short nb_bits_metadata[], / * o : Number of bits of metadata * / short localVAD[] ) { short i, ch, nb_bits_start, diff; short idx_azimuth, idx_azimuth_abs, flag_abs_azimuth[MAX_NUM_OBJECTS], nbits_diff_azimuth; short idx_elevation, idx_elevation_abs, flag_abs_elevation[MAX_NUM_OBJECTS], nbits_diff_elevation; float valQ; ISM_METADATA_HANDLE hIsmMetaData; long element_brate[MAX_NUM_OBJECTS], total_brate[MAX_NUM_OBJECTS]; short ism_metadata_flag_global; short ism_imp[MAX_NUM_OBJECTS]; / * Initialization * / ism_metadata_flag_global = 0; set_s( nb_bits_metadata, 0, n_ISms ); set_s( flag_abs_azimuth, 0, n_ISms ); set_s( flag_abs_elevation, 0, n_ISms ); / *----------------------------------------------------------------* * Set metadata presence / importance flags *----------------------------------------------------------------* / for( ch = 0; ch < n_ISms; ch++ ) { if( hIsmMeta[ch]->ism_metadata_flag ) { hIsmMeta[ch]->ism_metadata_flag = localVAD[ch]; } else { hIsmMeta[ch]->ism_metadata_flag = 0; } if ( hSCE[ch]->hCoreCoder[0]->tcxonly ) { / * Metadata is sent in every frame at the highest bitrate (using only the TCX core) * / hIsmMeta[ch]->ism_metadata_flag = 1; } } rate_ism_importance( n_ISms, hIsmMeta, hSCE, ism_imp ); / *----------------------------------------------------------------* * Write ISm common signaling *----------------------------------------------------------------* / / * Write several objects - single forward encoding * / for( ch = 1; ch < n_ISms; ch++ ) { push_indice( hBstr, IND_ISM_NUM_OBJECTS, 1, 1 ); } push_indice( hBstr, IND_ISM_NUM_OBJECTS, 0, 1 ); / * Write the ISm metadata flag (one per object) * / for( ch = 0; ch < n_ISms; ch++ ) { push_indice( hBstr, IND_ISM_METADATA_FLAG, ism_imp[ch], ISM_METADATA_FLAG_BITS ); ism_metadata_flag_global |= hIsmMeta[ch]->ism_metadata_flag; } / * Write the VAD flag * / for( ch = 0; ch < n_ISms; ch++ ) { if( hIsmMeta[ch]->ism_metadata_flag == 0 ) { push_indice( hBstr, IND_ISM_VAD_FLAG, localVAD[ch], VAD_FLAG_BITS ); } } if( ism_metadata_flag_global ) { / *----------------------------------------------------------------* * Quantize and code the metadata. Loop over all objects *----------------------------------------------------------------* / for( ch = 0; ch < n_ISms; ch++ ) { hIsmMetaData = hIsmMeta[ch]; nb_bits_start = hBstr->nb_bits_tot; if( hIsmMeta[ch]->ism_metadata_flag ) { / *----------------------------------------------------------------* * Quantization and coding of azimuth *----------------------------------------------------------------* / / * Azimuth quantization * / idx_azimuth_abs = usquant( hIsmMetaData->azimuth, &valQ, ISM_AZIMUTH_MIN, ISM_AZIMUTH_DELTA, (1 << ISM_AZIMUTH_NBITS) ); idx_azimuth = idx_azimuth_abs; nbits_diff_azimuth = 0; flag_abs_azimuth[ch] = 0; / * Default to differential coding * / if( hIsmMetaData->azimuth_diff_cnt == ISM_FEC_MAX / * Perform differential coding in up to ISM_FEC_MAX consecutive frames to control decoding in FEC * / || hIsmMetaData->last_ism_metadata_flag == 0 / * If the last frame did not code metadata, do not use differential coding * / ) { flag_abs_azimuth[ch] = 1; } / * Try differential coding * / if( flag_abs_azimuth[ch] == 0 ) { diff = idx_azimuth_abs - hIsmMetaData->last_azimuth_idx; if( diff == 0 ) { idx_azimuth = 0; nbits_diff_azimuth = 1; } else if( ABSVAL( diff ) < ISM_MAX_AZIMUTH_DIFF_IDX ) / * When diff bit >= abs bit, abs is preferred * / { idx_azimuth = 1 << 1; nbits_diff_azimuth = 1; if( diff < 0 ) { idx_azimuth += 1; / * negative sign * / diff *= -1; } else { idx_azimuth += 0; / * positive sign * / } idx_azimuth = idx_azimuth << diff; nbits_diff_azimuth++; / * Unicode encoding of "diff" * / idx_azimuth += ((1< <diff) - 1); nbits_diff_azimuth += diff; if( nbits_diff_azimuth < ISM_AZIMUTH_NBITS - 1 ) { / * Add stop bit - only for codewords shorter than ISM_AZIMUTH_NBITS * / idx_azimuth = idx_azimuth << 1; nbits_diff_azimuth++; } } else { flag_abs_azimuth[ch] = 1; } } / * Update counter * / if( flag_abs_azimuth[ch] == 0 ) { hIsmMetaData->azimuth_diff_cnt++; hIsmMetaData->elevation_diff_cnt = min( hIsmMetaData->elevation_diff_cnt, ISM_FEC_MAX ); } else { hIsmMetaData->azimuth_diff_cnt = 0; } / * Write azimuth * / push_indice( hBstr, IND_ISM_AZIMUTH_DIFF_FLAG, flag_abs_azimuth[ch], 1 ); if( flag_abs_azimuth[ch] ) { push_indice( hBstr, IND_ISM_AZIMUTH, idx_azimuth, ISM_AZIMUTH_NBITS ); } else { push_indice( hBstr, IND_ISM_AZIMUTH, idx_azimuth, nbits_diff_azimuth ); } / *----------------------------------------------------------------* * Quantization and encoding of elevation angle *----------------------------------------------------------------* / / * Quantization of elevation angle * / idx_elevation_abs = usquant( hIsmMetaData->elevation, &valQ, ISM_ELEVATION_MIN, ISM_ELEVATION_DELTA, (1 << ISM_ELEVATION_NBITS) ); idx_elevation = idx_elevation_abs; nbits_diff_elevation = 0; flag_abs_elevation[ch] = 0; / * Default differential coding * / if( hIsmMetaData->elevation_diff_cnt == ISM_FEC_MAX / * Perform differential coding in up to ISM_FEC_MAX consecutive frames to control decoding in FEC * / || hIsmMetaData->last_ism_metadata_flag == 0 / * If the last frame did not code metadata, do not use differential coding * / ) { flag_abs_elevation[ch] = 1; } / * Note: Elevation is only coded after the second frame (it has no meaning in init_frame) * / if( hSCE[0]->hCoreCoder[0]->ini_frame == 0 ) { flag_abs_elevation[ch] = 1; hIsmMetaData->last_elevation_idx = idx_elevation_abs; } diff = idx_elevation_abs - hIsmMetaData->last_elevation_idx; / * Avoid absolute coding of elevation if absolute coding has already been used for azimuth * / if( flag_abs_azimuth[ch] == 1 ) { flag_abs_elevation[ch] = 0; if( diff >= 0 ) { diff = min( diff, ISM_MAX_ELEVATION_DIFF_IDX ); } else { diff = -1 * min( -diff, ISM_MAX_ELEVATION_DIFF_IDX ); } } / * Try differential coding * / if( flag_abs_elevation[ch] == 0 ) { if( diff == 0 ) { idx_elevation = 0; nbits_diff_elevation = 1; } else if(ABSVAL(diff) < ISM_MAX_ELEVATION_DIFF_IDX) / * When diff bit >= abs bit, give priority to abs * / { idx_elevation = 1 << 1; nbits_diff_elevation = 1; if(diff < 0) { idx_elevation += 1; / * Negative sign * / diff *= -1; } else { idx_elevation += 0; / * Positive sign * / } idx_elevation = idx_elevation << diff; nbits_diff_elevation++; / * Unary encoding of "diff" * / idx_elevation += ((1 << diff) - 1); nbits_diff_elevation += diff; if(nbits_diff_elevation < ISM_ELEVATION_NBITS - 1) { / * Add a stop bit * / idx_elevation = idx_elevation << 1; nbits_diff_elevation++; } } else { flag_abs_elevation[ch] = 1; } } / * Update the counter * / if( flag_abs_elevation[ch] == 0 ) { hIsmMetaData->elevation_diff_cnt++; hIsmMetaData->elevation_diff_cnt = min( hIsmMetaData->elevation_diff_cnt, ISM_FEC_MAX ); } else { hIsmMetaData->elevation_diff_cnt = 0; } / * Write the elevation * / if( flag_abs_azimuth[ch] == 0 ) / * If "flag_abs_azimuth == 1", do not write "flag_abs_elevation" * / / * VE: Regarding VAD 0->1, TBV * / { push_indice( hBstr, IND_ISM_ELEVATION_DIFF_FLAG, flag_abs_elevation[ch], 1 ); } if( flag_abs_elevation[ch] ) { push_indice( hBstr, IND_ISM_ELEVATION, idx_elevation, ISM_ELEVATION_NBITS ); } else { push_indice( hBstr, IND_ISM_ELEVATION, idx_elevation, nbits_diff_elevation ); } / *----------------------------------------------------------------* * Update *----------------------------------------------------------------* / hIsmMetaData->last_azimuth_idx = idx_azimuth_abs; hIsmMetaData->last_elevation_idx = idx_elevation_abs; / * Save the number of bits of the written metadata * / nb_bits_metadata[ch] = hBstr->nb_bits_tot - nb_bits_start; } } / *----------------------------------------------------------------* * Inter-object logic to minimize the use of several absolutely-coded indexes in the same frame *----------------------------------------------------------------* / i = 0; while( i == 0 || i < n_ISms / INTER_OBJECT_PARAM_CHECK ) { short num, abs_num, abs_first, abs_next, pos_zero; short abs_matrice[INTER_OBJECT_PARAM_CHECK * 2]; num = min( INTER_OBJECT_PARAM_CHECK, n_ISms - i * INTER_OBJECT_PARAM_CHECK ); i++; set_s( abs_matrice, 0, INTER_OBJECT_PARAM_CHECK * ISM_NUM_PARAM ); for( ch = 0; ch < num; ch++ ) { if( flag_abs_azimuth[ch] == 1 ) { abs_matrice[ch*ISM_NUM_PARAM] = 1; } if( flag_abs_elevation[ch] == 1 ) { abs_matrice[ch*ISM_NUM_PARAM + 1] = 1; } } abs_num = sum_s( abs_matrice, INTER_OBJECT_PARAM_CHECK * ISM_NUM_PARAM ); abs_first = 0; while( abs_num > 1 ) { / * Find the first "1" entry * / while( abs_matrice[abs_first] == 0 ) { abs_first++; } / * Find the next "1" entry * / abs_next = abs_first + 1; while( abs_matrice[abs_next] == 0 ) { abs_next++; } / * Find the position of '0' * / pos_zero = 0; while( abs_matrice[pos_zero] == 1 ) { pos_zero++; } ch = abs_next / ISM_NUM_PARAM; if( abs_next % ISM_NUM_PARAM == 0 ) { hIsmMeta[ch]->azimuth_diff_cnt = abs_num - 1; } if( abs_next % ISM_NUM_PARAM == 1 ) { hIsmMeta[ch]->elevation_diff_cnt = abs_num - 1; / *hIsmMeta[ch]->elevation_diff_cnt = min( hIsmMeta[ch]->elevation_diff_cnt, ISM_FEC_MAX );* / } abs_first++; abs_num--; } } } / *----------------------------------------------------------------* * Configuration and determination of bit rate for each channel *----------------------------------------------------------------* / ism_config( ism_total_brate, n_ISms, hIsmMeta, localVAD, ism_imp, element_brate, total_brate, nb_bits_metadata ); for( ch = 0; ch < n_ISms; ch++ ) { hIsmMeta[ch]->last_ism_metadata_flag = hIsmMeta[ch]->ism_metadata_flag; hSCE[ch]->hCoreCoder[0]->low_rate_mode = 0; if ( hIsmMeta[ch]->ism_metadata_flag == 0 && localVAD[ch][0] == 0 && ism_metadata_flag_global ) { hSCE[ch]->hCoreCoder[0]->low_rate_mode = 1; } hSCE[ch]->element_brate = element_brate[ch]; hSCE[ch]->hCoreCoder[0]->total_brate = total_brate[ch]; / * Write metadata only for active frames * / if( hSCE[0]->hCoreCoder[0]->core_brate > SID_2k40 ) { reset_indices_enc( hSCE[ch]->hMetaData, MAX_BITS_METADATA ); } } return; } void rate_ism_importance( const short n_ISms, / * i : Number of objects * / ISM_METADATA_HANDLE hIsmMeta[], / * i / o: Handle of ISM metadata * / ENC_HANDLE hSCE[], / * i / o: Handle of element encoder * / short ism_imp[] / * o : ISM importance flag * / ) { short ch, ctype; for( ch = 0; ch < n_ISms; ch++ ) { ctype = hSCE[ch]->hCoreCoder[0]->coder_type_raw; if( hIsmMeta[ch]->ism_metadata_flag == 0 ) { ism_imp[ch] = ISM_NO_META; } else if( ctype == INACTIVE || ctype == UNVOICED ) { ism_imp[ch] = ISM_LOW_IMP; } else if( ctype == VOICED ) { ism_imp[ch] = ISM_MEDIUM_IMP; } else / * GENERIC * / { ism_imp[ch] = ISM_HIGH_IMP; } } return; } void ism_config( const long ism_total_brate, / * i : Total bit rate of ISm * / const short n_ISms, / * i : Number of objects * / ISM_METADATA_HANDLE hIsmMeta[], / * i / o: Handle of ISM metadata * / short localVAD[], const short ism_imp[], / * i : ISM importance flag * / long element_brate[], / * o : Element bit rate per object * / long total_brate[], / * o : Total bit rate per object * / short nb_bits_metadata[] / * i / o: Number of bits of metadata * / ) { short ch; short bits_element[MAX_NUM_OBJECTS], bits_CoreCoder[MAX_NUM_OBJECTS]; short bits_ism, bits_side; long tmpL; short ism_metadata_flag_global; / * Initialization * / ism_metadata_flag_global = 0; bits_side = 0; if( hIsmMeta != NULL ) { for( ch = 0; ch < n_ISms; ch++ ) { ism_metadata_flag_global |= hIsmMeta[ch]->ism_metadata_flag;} } / * Judgment on bitrate for each channel - Constant during session (with one ism_total_brate) * / bits_ism = ism_total_brate / FRMS_PER_SECOND; set_s( bits_element, bits_ism / n_ISms, n_ISms ); bits_element[n_ISms - 1] += bits_ism % n_ISms; bitbudget_to_brate( bits_element, element_brate, n_ISms ); / * Count the bits of ISm common signaling * / if( hIsmMeta != NULL ) { nb_bits_metadata[0] += n_ISms * ISM_METADATA_FLAG_BITS + n_ISms; for( ch = 0; ch < n_ISms; ch++ ) { if( hIsmMeta[ch]->ism_metadata_flag == 0 ) { nb_bits_metadata[0] += ISM_METADATA_VAD_FLAG_BITS; } } } / * Divide the bit budget of metadata equally among channels * / if( nb_bits_metadata != NULL ) { bits_side = sum_s( nb_bits_metadata, n_ISms ); set_s(nb_bits_metadata, bits_side / n_ISms, n_ISms); nb_bits_metadata[n_ISms - 1] += bits_side % n_ISms; v_sub_s(bits_element, nb_bits_metadata, bits_CoreCoder, n_ISms); bitbudget_to_brate(bits_CoreCoder, total_brate, n_ISms); mvs2s(nb_bits_metadata, nb_bits_metadata, n_ISms); } / * Allocate less CoreCoder bit budget for non - active streams (at least one stream must be active) * / if(ism_metadata_flag_global) { long diff; short n_higher, flag_higher[MAX_NUM_OBJECTS]; set_s(flag_higher, 1, MAX_NUM_OBJECTS); diff = 0; for(ch = 0; ch < n_ISms; ch++) { if(hIsmMeta[ch]->ism_metadata_flag == 0 && localVAD[ch] == 0) { diff += bits_CoreCoder[ch] - BITS_ISM_INACTIVE; bits_CoreCoder[ch] = BITS_ISM_INACTIVE; flag_higher[ch] = 0;} } n_higher = sum_s( flag_higher, n_ISms ); if( diff > 0 && n_higher > 0 ) { tmpL = diff / n_higher; for( ch = 0; ch < n_ISms; ch++ ) { if( flag_higher[ch] ) { bits_CoreCoder[ch] += tmpL; } } tmpL = diff % n_higher; ch = 0; while( flag_higher[ch] == 0 ) { ch++; } bits_CoreCoder[ch] += tmpL; } bitbudget_to_brate( bits_CoreCoder, total_brate, n_ISms ); diff = 0; for( ch = 0; ch < n_ISms; ch++ ) { long limit; limit = MIN_BRATE_SWB_BWE / FRMS_PER_SECOND; if( element_brate[ch] < MIN_BRATE_SWB_STEREO ) { limit = MIN_BRATE_WB_BWE / FRMS_PER_SECOND; } else if( element_brate[ch] >= SCE_CORE_16k_LOW_LIMIT ) { / *限度(limit) = SCE_CORE_16k_LOW_LIMIT;* / limit = (ACELP_16k_LOW_LIMIT + SWB_TBE_1k6) / FRMS_PER_SECOND; } if( ism_imp[ch] == ISM_NO_META && localVAD[ch] == 0 ) { tmpL = BITS_ISM_INACTIVE; } else if( ism_imp[ch] == ISM_LOW_IMP ) { tmpL = BETA_ISM_LOW_IMP * bits_CoreCoder[ch]; tmpL = max( limit, bits_CoreCoder[ch] - tmpL ); } else if( ism_imp[ch] == ISM_MEDIUM_IMP ) { tmpL = BETA_ISM_MEDIUM_IMP * bits_CoreCoder[ch]; tmpL = max( limit, bits_CoreCoder[ch] - tmpL ); } else / * ism_imp[ch] == ISM_HIGH_IMP * / { tmpL = bits_CoreCoder[ch]; } diff += bits_CoreCoder[ch] - tmpL; bits_CoreCoder[ch] = tmpL; } if (diff > 0 && n_higher > 0) { tmpL = diff / n_higher; for (ch = 0; ch < n_ISms; ch++) { if (flag_higher[ch]) { bits_CoreCoder[ch] += tmpL; } } tmpL = diff % n_higher; ch = 0; while (flag_higher[ch] == 0) { ch++; } bits_CoreCoder[ch] += tmpL; } / * Verify for the maximum bitrate @12.8kHz core * / diff = 0; for (ch = 0; ch < n_ISms; ch++) { limit_high = STEREO_512k / FRMS_PER_SECOND; if (element_brate[ch] < SCE_CORE_16k_LOW_LIMIT) / * To reproduce the function set_ACELP_flag() -> It is not intended to switch the internal sampling rate of ACELP within the object * / { limit_high = ACELP_12k8_HIGH_LIMIT / FRMS_PER_SECOND; } tmpL = min(bits_CoreCoder[ch], limit_high); diff += bits_CoreCoder[ch] - tmpL; bits_CoreCoder[ch] = tmpL; } if ( diff > 0 ) { ch = 0; for ( ch = 0; ch < n_ISms; ch++ ) { if ( flag_higher[ch] == 0 ) { if ( diff > limit_high ) { diff += bits_CoreCoder[ch] - limit_high; bits_CoreCoder[ch] = limit_high; } else { bits_CoreCoder[ch] += diff; break; } } } } bitbudget_to_brate( bits_CoreCoder, total_brate, n_ISms ); } return; }

[0115] 7.0 Hardware Implementation Figure 8 is a simplified block diagram of an exemplary configuration of hardware components forming the coding and decoding system and method described above.

[0116] Each of the coding and decoding systems may be implemented as part of a mobile terminal, as part of a portable media player, or in any similar device. (Identified as 1200 in FIG. 8) Each of the coding and decoding systems includes an input 1202, an output 1204, a processor 1206, and a memory 1208.

[0117] Input 1202 is configured to receive an input signal, e.g., the N audio objects 102 of FIG. 1 (N audio streams with corresponding N metadata) or the bitstream 701 of FIG. 7, in digital or analog form. Output 1204 is configured to supply an output signal, e.g., the bitstream 111 of FIG. 1, or the M decoded audio channels 703 and M decoded metadata 704 of FIG. 7. Input 1202 and output 1204 may be implemented in a common module, e.g., a serial input / output device.

[0118] Processor 1206 is operably connected to input 1202, output 1204, and memory 1208. Processor 1206 is realized as one or more processors for executing code instructions to assist with the functions of the various processors and other modules of FIGS. 1 and 7.

[0119] Memory 1208 includes a non-transitory memory for storing code instructions executable by processor 1206, in particular, a processor-readable memory including non-transitory instructions that, when executed, cause the processor to implement the operation of the coding and decoding systems and methods and the processors / modules as described in the present disclosure. Memory 1208 may also include a random access memory or buffer for storing intermediate processing data from the various functions executed by processor 1206.

[0120] Those skilled in the art will recognize that the description of the coding and decoding systems and methods is merely exemplary and not at all intended to be limiting. Other embodiments will be readily apparent to such skilled artisans who have the benefit of this disclosure. Further, the disclosed coding and decoding systems and methods may be customized to provide valuable solutions to existing needs and problems in encoding and decoding audio.

[0121] For clarity, not all well - defined features of the implementation of the coding and decoding systems and methods are shown and described. Of course, in developing any such actual implementation of the coding and decoding systems and methods, numerous implementation - specific decisions must be made to achieve the developer's particular goals, such as conforming to application, system, network, and business - related constraints, and it will be understood that these particular goals will vary from implementation to implementation and from developer to developer. Further, although the development effort can be complex and time - consuming, it will be understood that it is routine work of engineering for those of ordinary skill in the field of audio processing who have the benefit of this disclosure.

[0122] According to the present disclosure, the processors / modules, processing operations, and / or data structures described herein may be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. Additionally, those skilled in the art will recognize that devices of a less general-purpose nature, such as wired devices, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., may also be used. Where a method comprising a series of operations and sub-operations is performed by a processor, computer, or machine, and those operations and sub-operations may be stored as a series of non-transitory code instructions readable by the processor, computer, or machine, those operations and sub-operations may be stored on a tangible and / or non-transitory medium.

[0123] The coding and decoding systems and methods described herein may use software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.

[0124] In the coding and decoding systems and methods described herein, the various operations and sub-operations may be executed in various orders, and some of the operations and sub-operations may be optional.

[0125] The present disclosure has been described above through non-limiting exemplary embodiments of the present disclosure, but these embodiments may be optionally modified within the scope of the appended claims without departing from the spirit and essence of the present disclosure.

[0126] 8.0 References The following references are cited in the present disclosure, and all of the contents of those references are incorporated herein by reference. [1] 3GPP Specification TS 26.445: "Codec for Enhanced Voice Services (EVS). Detailed Algorithmic Description", v.12.0.0, September 2014 [2] V. Eksler, "Method and Device for Allocating a Bit-budget Between Sub-frames in a CELP Codec", PCT Patent Application PCT / CA2018 / 51175

[0127] 9.0 Further Embodiments The following embodiments (Embodiments 1 to 83) are part of the present disclosure related to the present invention.

[0128] Embodiment 1. A system for coding an object-based audio signal including an audio object according to an audio stream having associated metadata, an audio stream processor for analyzing the audio stream, a metadata processor responsive to information about the audio stream from an analysis by the audio stream processor for encoding metadata of the input audio stream.

[0129] Embodiment 2. The system of Embodiment 1, further including a bit-budget allocator responsive to information about the bit-budget of the metadata of the audio object from the metadata processor, wherein the metadata processor outputs information about the bit-budget of the metadata of the audio object and the system allocates a bitrate to the audio stream.

[0130] Embodiment 3. The system of Embodiment 1 or 2 including an encoder for the audio stream including the coded metadata.

[0131] Embodiment 4. A system according to any one of Embodiments 1 to 3, wherein the encoder includes several Core-Coders that use the bitrate assigned to the audio stream by a bit budget allocator.

[0132] Embodiment 5. A system according to any one of Embodiments 1 to 4, wherein the object-based audio signal includes at least one of a human voice, music, and general audio speech.

[0133] Embodiment 6. A system according to any one of Embodiments 1 to 5, wherein the object-based audio signal represents or encodes an auditory scene of complex audio as a set of individual elements, the audio objects.

[0134] Embodiment 7. A system according to any one of Embodiments 1 to 6, wherein each audio object includes an audio stream having associated metadata.

[0135] Embodiment 8. A system according to any one of Embodiments 1 to 7, wherein the audio stream is an independent stream having metadata.

[0136] Embodiment 9. A system according to any one of Embodiments 1 to 8, wherein the audio stream represents an audio waveform and typically includes one or two channels.

[0137] Embodiment 10. A system according to any one of Embodiments 1 to 9, wherein the metadata is a set of information that describes the audio stream and the artistic intent for communicating the original or coded audio object to a final playback system.

[0138] Embodiment 11. A system according to any one of Embodiments 1 to 10, wherein the metadata typically describes the spatial characteristics of each audio object.

[0139] Embodiment 12. A system according to any one of Embodiments 1 to 11, wherein the spatial characteristics include one or more of the position, orientation, volume, and width of the audio object.

[0140] Embodiment 13. A system according to any one of Embodiments 1 to 12, wherein each audio object includes a set of metadata called input metadata, which is defined as an unquantized metadata representation used as an input to the codec.

[0141] Embodiment 14. A system according to any one of Embodiments 1 to 13, wherein each audio object includes a set of metadata called coded metadata, which is defined as quantized and coded metadata that is part of the bitstream transmitted from the encoder to the decoder.

[0142] Embodiment 15. A system according to any one of Embodiments 1 to 14, wherein the playback system is assembled on the playback side to render audio objects in the 3D audio space around the listener using the transmitted metadata and artistic intent.

[0143] Embodiment 16. A system according to any one of Embodiments 1 to 15, wherein the playback system includes a head tracking device for dynamically modifying the metadata during the rendering of the audio object.

[0144] Embodiment 17. A system according to any one of Embodiments 1 to 16, including a framework for the simultaneous coding of several audio objects.

[0145] Embodiment 18. A system according to any one of Embodiments 1 to 17, wherein the simultaneous coding of several audio objects uses a fixed overall bitrate determined for encoding the audio objects.

[0146] Embodiment 19. A system according to any one of Embodiments 1 to 18, including a transmitter for transmitting some or all of the audio objects.

[0147] Embodiment 20. A system according to any one of Embodiments 1 to 19, where when coding a combination of audio formats in a framework, a certain overall bitrate represents the sum of the bitrates of the formats.

[0148] Embodiment 21. A system according to any one of Embodiments 1 to 20, where the metadata includes two parameters including azimuth and elevation angles.

[0149] Embodiment 22. A system according to any one of Embodiments 1 to 21, where the azimuth parameter and the elevation parameter are stored for each audio frame for each audio object.

[0150] Embodiment 23. A system according to any one of Embodiments 1 to 22, including an input buffer for buffering at least one input audio stream and input metadata associated with the audio stream.

[0151] Embodiment 24. A system according to any one of Embodiments 1 to 23, where the input buffer buffers each audio stream for one frame.

[0152] Embodiment 25. A system according to any one of Embodiments 1 to 24, where an audio stream processor analyzes and processes the audio stream.

[0153] Embodiment 26. A system according to any one of Embodiments 1 to 25, where the audio stream processor includes at least one of the following elements: a transient detector in the time domain, a spectrum analyzer, a long-term prediction analyzer, a pitch tracker and a voice analyzer, a voice / sound activity detector, a bandwidth detector, a noise estimator, and a signal classifier.

[0154] Embodiment 27. The system according to any one of Embodiments 1 to 26, wherein the signal classifier performs at least one of coder type selection, signal classification, and human voice / music classification.

[0155] Embodiment 28. The system according to any one of Embodiments 1 to 27, wherein the metadata processor analyzes, quantizes, and encodes the metadata of the audio stream.

[0156] Embodiment 29. The system according to any one of Embodiments 1 to 28, wherein in an inactive frame, the metadata is not encoded by the metadata processor and is not transmitted by the system in the bitstream of the corresponding audio object.

[0157] Embodiment 30. The system according to any one of Embodiments 1 to 29, wherein in an active frame, the metadata is encoded by the metadata processor for the corresponding object using a variable bit rate.

[0158] Embodiment 31. The system according to any one of Claims 1 to 30, wherein the bit budget allocator sums the bit budgets of the metadata of the audio objects and adds the sum of the bit budgets to the bit budget of the signaling to allocate a bit rate to the audio stream.

[0159] Embodiment 32. The system according to any one of Embodiments 1 to 31, further comprising a preprocessor for further processing the audio stream when the configuration and bit rate distribution between the audio streams are performed.

[0160] Embodiment 33. The system according to any one of Embodiments 1 to 32, wherein the preprocessor performs at least one of further classification of the audio stream, selection of a core encoder, and resampling.

[0161] Embodiment 34. A system according to any one of Embodiments 1 to 33, wherein an encoder sequentially encodes an audio stream.

[0162] Embodiment 35. A system according to any one of Embodiments 1 to 34, wherein an encoder sequentially encodes an audio stream using several variable bit rate core encoders.

[0163] Embodiment 36. A system according to any one of Embodiments 1 to 35, wherein a metadata processor sequentially encodes metadata in a loop using a dependency between quantization of an audio object and metadata parameters of the audio object.

[0164] Embodiment 37. A system according to any one of Embodiments 1 to 36, wherein a metadata processor quantizes an index of metadata parameters using a quantization step in order to encode the metadata parameters.

[0165] Embodiment 38. A system according to any one of Embodiments 1 to 37, wherein a metadata processor quantizes an index of azimuth using a quantization step in order to encode an azimuth parameter, and quantizes an index of elevation using a quantization step in order to encode an elevation parameter.

[0166] Embodiment 39. A system according to any one of Embodiments 1 to 38, wherein the total bit budget and quantization bit number of metadata depend on the total bit rate of a codec related to one audio object, the total bit rate of metadata, or the sum of the bit budget of metadata and the bit budget of a core encoder.

[0167] Embodiment 40. A system according to any one of Embodiments 1 to 39, wherein an azimuth parameter and an elevation parameter are represented as one parameter.

[0168] Embodiment 41. A system according to any one of Embodiments 1 to 40, wherein the metadata processor encodes the index of the metadata parameter either absolutely or differentially.

[0169] Embodiment 42. A system according to any one of Embodiments 1 to 41, wherein the metadata processor uses absolute coding to encode the index of the metadata parameter when there is a difference between the index of the current parameter and the index of the previous parameter that results in the number of bits required for differential coding being greater than or equal to the number of bits required for absolute coding.

[0170] Embodiment 43. A system according to any one of Embodiments 1 to 42, wherein the metadata processor uses absolute coding to encode the index of the metadata parameter when there is no metadata in the previous frame.

[0171] Embodiment 44. A system according to any one of Embodiments 1 to 43, wherein the metadata processor uses absolute coding to encode the index of the metadata parameter when the number of consecutive frames using differential coding is greater than the maximum number of consecutive frames coded using differential coding.

[0172] Embodiment 45. A system according to any one of Embodiments 1 to 44, wherein when the metadata processor encodes the index of the metadata parameter using absolute coding, it writes an absolute coding flag that distinguishes between absolute coding and differential coding following the absolutely coded index of the metadata parameter.

[0173] Embodiment 46. When the metadata processor encodes the index of the metadata parameter using differential coding, it sets the absolute coding flag to 0, and following the absolute coding flag, writes a zero coding flag that signals whether the difference between the index of the current frame and the index of the previous frame is 0. A system according to any one of Embodiments 1 to 45.

[0174] Embodiment 47. When the difference between the index of the current frame and the index of the previous frame is not equal to 0, the metadata processor continues coding by writing a sign flag and an adaptive-bits difference index that follows. A system according to any one of Embodiments 1 to 46.

[0175] Embodiment 48. The metadata processor uses the coding logic of the metadata within the object to limit the range of variation of the bit budget of the metadata between frames and prevent the bit budget left for core coding from becoming too small. A system according to any one of Embodiments 1 to 47.

[0176] Embodiment 49. The metadata processor limits the use of absolute coding in a given frame to only one metadata parameter or the smallest possible number of metadata parameters according to the coding logic of the metadata within the object. A system according to any one of Embodiments 1 to 48.

[0177] Embodiment 50. When the index of the coding logic of one metadata has already been coded using absolute coding within the same frame according to the coding logic of the metadata within the object, the metadata processor avoids absolute coding of the index of another metadata parameter. A system according to any one of Embodiments 1 to 49.

[0178] Embodiment 51. A system according to any one of Embodiments 1 to 50, wherein the coding logic of the metadata in the object depends on the bit rate.

[0179] Embodiment 52. A system according to any one of Embodiments 1 to 51, wherein the metadata processor uses the coding logic of the metadata between different objects to minimize the number of metadata parameters to be absolutely coded for different audio objects in the current frame.

[0180] Embodiment 53. A system according to any one of Embodiments 1 to 52, wherein the metadata processor uses the coding logic of the metadata between different objects to control the frame counter of the metadata parameters to be absolutely coded.

[0181] Embodiment 54. A system according to any one of Embodiments 1 to 53, wherein the metadata processor uses the coding logic of the metadata between different objects, and when the metadata parameters of the audio object develop slowly and smoothly, (a) codes the index of the first metadata parameter of the first audio object using absolute coding in frame M, (b) codes the index of the second metadata parameter of the first audio object using absolute coding in frame M+1, (c) codes the index of the first metadata parameter of the second audio object using absolute coding in frame M+2, and (d) codes the index of the second metadata parameter of the second audio object using absolute coding in frame M+3.

[0182] Embodiment 55. A system according to any one of Embodiments 1 to 54, wherein the coding logic of the metadata between different objects depends on the bit rate.

[0183] Embodiment 56. A system according to any one of Embodiments 1 to 55, wherein a bit budget allocator uses a bit rate adaptation algorithm to allocate a bit budget for encoding an audio stream.

[0184] Embodiment 57. A system according to any one of Embodiments 1 to 56, wherein a bit budget allocator uses a bit rate adaptation algorithm to obtain a total bit budget of metadata from a total bit rate of metadata or a total bit rate of a codec.

[0185] Embodiment 58. A system according to any one of Embodiments 1 to 57, wherein a bit budget allocator uses a bit rate adaptation algorithm to calculate an element bit budget by dividing the total bit budget of metadata by the number of audio streams.

[0186] Embodiment 59. A system according to any one of Embodiments 1 to 58, wherein a bit budget allocator uses a bit rate adaptation algorithm to adjust the element bit budget of the last audio stream in order to use up all the bit budgets of the available metadata.

[0187] Embodiment 60. A system according to any one of Embodiments 1 to 59, wherein a bit budget allocator uses a bit rate adaptation algorithm to sum up the bit budgets of metadata of all audio objects, add the sum to the bit budget of metadata common signaling, and generate a side bit budget of a core coder.

[0188] Embodiment 61. A system according to any one of Embodiments 1 to 60, wherein a bit budget allocator uses a bit rate adaptation algorithm to (a) evenly divide the side bit budget of a core coder among audio objects, and (b) calculate the bit budget of the core coder for each audio stream using the divided side bit budget of the core coder and the element bit budget.

[0189] Embodiment 62. A system according to any one of Embodiments 1 to 61, wherein the bit budget allocator uses a bit rate adaptation algorithm to adjust the bit budget of the core coder of the last audio stream in order to use up all the bit budgets of the available core coders.

[0190] Embodiment 63. A system according to any one of Embodiments 1 to 62, wherein the bit budget allocator uses a bit rate adaptation algorithm to calculate the bit rate for encoding one audio stream in the core coder using the bit budget of the core coder.

[0191] Embodiment 64. A system according to any one of Embodiments 1 to 63, wherein the bit budget allocator uses a bit rate adaptation algorithm in non-active frames or frames with low energy to lower the bit rate for encoding one audio stream in the core coder, set it to a constant value, and redistribute the saved bit budget among the audio streams of the active frames.

[0192] Embodiment 65. A system according to any one of Embodiments 1 to 64, wherein the bit budget allocator uses a bit rate adaptation algorithm in active frames to adjust the bit rate for encoding one audio stream in the core coder based on the importance classification of the metadata.

[0193] Embodiment 66. A system according to any one of Embodiments 1 to 65, wherein the bit budget allocator lowers the bit rate for encoding one audio stream in the core coder in non-active frames (VAD = 0), and redistributes the bit budget saved by the lowering of the bit rate among the audio streams of the frames classified as active.

[0194] Embodiment 67. A system according to any one of Embodiments 1 to 66, wherein the bit budget allocator sets, in a frame, (a) a lower constant core coder bit budget for any audio stream having non-active content, (b) calculates the saved bit budget as the difference between the lower constant core coder bit budget and the core coder bit budget, and (c) redistributes the saved bit budget among the core coder bit budgets of the audio streams of the active frames.

[0195] Embodiment 68. A system according to any one of Embodiments 1 to 67, wherein the lower constant bit budget depends on the total bit rate of the metadata.

[0196] Embodiment 69. A system according to any one of Embodiments 1 to 68, wherein the bit budget allocator calculates the bit rate for encoding one audio stream in the core coder using the lower constant core coder bit budget.

[0197] Embodiment 70. A system according to any one of Embodiments 1 to 69, wherein the bit budget allocator uses adaptation of the core coder bit rate among objects based on classification of the importance of the metadata.

[0198] Embodiment 71. A system according to any one of Embodiments 1 to 70, wherein the importance of the metadata is based on an indicator indicating how important the coding of a particular audio object in the current frame is for obtaining a satisfactory quality of the decoded synthesis.

[0199] Embodiment 72. A system according to any one of Embodiments 1 to 71, wherein the bit budget allocator classifies the importance of metadata based on at least one of the following parameters: coder type (coder_type), FEC signal classification (class), determination of human voice / music classification, and SNR estimation values (snr_celp, snr_tcx) from an open-loop ACELP / TCX core determination module.

[0200] Embodiment 73. A system according to any one of Embodiments 1 to 72, wherein the bit budget allocator classifies the importance of metadata based on the coder type (coder_type).

[0201] Embodiment 74. A system according to any one of Embodiments 1 to 73, wherein the bit budget allocator defines the following four different metadata importance classes (class ISm ), namely: - No-metadata class ISM_NO_META: A frame without metadata coding, for example, an inactive frame with VAD = 0 - Low-importance class ISM_LOW_IMP: A frame with coder_type = UNVOICED or INACTIVE - Medium-importance class ISM_MEDIUM_IMP: A frame with coder_type = VOICED - High-importance class ISM_HIGH_IMP: A frame with coder_type = GENERIC

[0202] Embodiment 75. A system according to any one of Embodiments 1 to 74, wherein the bit budget allocator uses the metadata importance class in a bitrate adaptation algorithm to allocate more bit budgets to audio streams with higher importance and fewer bit budgets to audio streams with lower importance.

[0203] Embodiment 76. The bit budget allocator performs the following logic in the frame, that is, 1. class ISm = FRAME of ISM_NO_META: A lower fixed core coder bit rate is allocated. 2. class ISm = FRAME of ISM_LOW_IMP: The bit rate (total_brate) for encoding one audio stream in the core coder is [number]total_brate new [n] = max(α low *total_brate[n], B low ) lowered as shown, where the constant α low is set to a value less than 1.0, and the constant B low is the threshold of the minimum bit rate supported by the core coder. 3. class ISm = FRAME of ISM_MEDIUM_IMP: The bit rate (total_brate) for encoding one audio stream in the core coder is, [number]total_brate new [n] = max(α med *total_brate[n], B low ) lowered as shown, where the constant α med is less than 1.0 but is set to a value greater than the value α low . 4. class ISm = FRAME of ISM_HIGH_IMP: Any one of the systems of Embodiments 1 to 75 that does not use bit rate adaptation is used.

[0204] Embodiment 77. A system according to any one of Embodiments 1 to 76, wherein a bit budget allocator redistributes a saved bit budget represented as the sum of the difference between the previous bit rate total_brate and the new bit rate total_brate among the audio streams of frames classified as active.

[0205] Embodiment 78. A system for decoding an audio object according to an audio stream having associated metadata, a metadata processor for decoding the metadata of an audio stream having active content, a bit budget allocator responsive to the decoded metadata and the respective bit budgets of the audio objects for determining the bit rate of the core coder of the audio stream, and a decoder of the audio stream that uses the bit rate of the core coder determined by the bit budget allocator.

[0206] Embodiment 79. The system of Embodiment 78, wherein the metadata processor responds to metadata common signaling read from the end of the received bit stream.

[0207] Embodiment 80. The system of Embodiment 78 or 79, wherein the decoder includes a core decoder for decoding the audio stream.

[0208] Embodiment 81. A system according to any one of Embodiments 78 to 80, wherein the core decoder includes a variable bit rate core decoder for sequentially decoding the audio stream at the bit rate of each respective core coder.

[0209] Embodiment 82. A system according to any one of Embodiments 78 to 81, wherein the number of audio objects to be decoded is less than the number of core decoders.

[0210] Embodiment 83. A system according to any one of Embodiments 78 to 83, including a renderer of an audio object that responds to a decoded audio stream and decoded metadata.

[0211] Any of Embodiments 2 to 77, which further describe the elements of Embodiments 78 to 83, may be implemented in any of these Embodiments 78 to 83. By way of example, the bitrate of the core coder for each audio stream in the decoding system is determined using the same procedure as the coding system.

[0212] The present invention also relates to a coding method and a decoding method. In this regard, Embodiments 1 to 83 of the system may be drafted as method embodiments in which the elements of the system embodiment are replaced by the operations performed by such elements.

Description of Signs

[0213] 100 System 101 Input buffer 102 Input audio object 103 Audio stream processor 104 Transport channel 105 Metadata processor 106 Configuration and determination processor 107 Information 108 Preprocessor 109 Core encoder 110 Multiplexer 111 Output bitstream 112 Quantized and encoded metadata 113 ISm common signaling 114 N audio streams 120 Signal classification information 121 Line 150 Method 151 Operation of buffering the input 153 Operations of analysis and prefrontal processing Operations of analyzing, quantizing, and coding metadata Operations of configuring and determining 158 Further preprocessing; Operations of preprocessing Operations of core coding Operations of multiplexing 700 Decoder; Decoding system 701 Bitstream 702 Output audio channel 703 Decoded audio stream 704 Decoded metadata 705 Demultiplexer 706 Metadata decoding and inverse quantization processor 707 Configuration and determination processor 708 Line 709 Output setting 710 Core decoder 711 Renderer 712 Output setting 750 Method Operations of demultiplexing Operations of decoding and inverse quantizing metadata Operations of configuring and determining for per-channel bitrate Operations of core decoding Operations of rendering audio channels 1200 Coding and decoding system 1202 Input 1204 Output 1206 Processor 1208 Memory

Claims

1. A system for coding an object-based audio signal including an audio object in accordance with an audio stream having associated metadata, a metadata processor, coding the metadata separately from and prior to coding of the audio stream using a process adapted to the metadata, after the metadata is coded, generating information about a bit budget to be used by the metadata processor for coding the metadata of the audio object; a bit budget allocator that allocates a bit rate for coding the audio stream in response to the information about the bit budget to be used by the metadata processor for coding the metadata of the audio object; and an encoder for coding the audio stream using the bit rate allocated by the bit budget allocator for coding the audio stream. A system comprising the above components.

2. The system according to claim 1, further comprising an audio stream processor for analyzing the audio stream and providing information about the audio stream to the metadata processor and the bit budget allocator.

3. The system according to claim 1 or 2, wherein the bit budget allocator uses a bit rate adaptation algorithm to allocate available bit budgets for coding the audio stream.

4. The system according to claim 3, wherein the bit budget allocator uses the bit rate adaptation algorithm to calculate a total bit budget for the audio stream and metadata (ISm) or a total bit rate of a codec for coding the audio stream and the associated metadata from a total bit rate of the audio stream and metadata (ISm) or a total bit rate of the codec for coding the audio stream and the associated metadata.

5. The system according to claim 4, wherein the bit budget allocator uses the bit rate adaptation algorithm to calculate the bit budget of an element by dividing the total bit budget of the ISm by the number of audio streams.

6. The system according to claim 5, wherein the bit budget allocator uses the bit rate adaptation algorithm to adjust the bit budget of the element of the last audio object in order to use all of the total bit budget of the ISm.

7. The system according to claim 5 or 6, wherein the bit budget allocator uses the bit rate adaptation algorithm to sum the bit budget for the coding of the metadata of the audio object, add the sum to the bit budget of the ISm common signaling, and produce a codec side bit budget.

8. The system according to claim 7, wherein the bit budget allocator uses the bit rate adaptation algorithm to (a) evenly divide the codec side bit budget among the audio objects, and (b) calculate the coding bit budget for each audio stream using the divided codec side bit budget and the bit budget of the element.

9. The system according to claim 8, wherein the bit budget allocator uses the bit rate adaptation algorithm to adjust the coding bit budget of the last audio stream in order to use all of the available coding bit budget.

10. The system according to claim 8 or 9, wherein the bit budget allocator uses the bit rate adaptation algorithm to calculate the bit rate for coding one of the audio streams using the coding bit budget for the audio stream.

11. The system according to any one of claims 3 to 10, wherein the bit budget allocator uses the bit rate adaptation algorithm with an audio stream having non-active content or meaningless content to reduce the value of the bit rate for coding one of the audio streams, and redistributes the saved bit budget among the audio streams having active content.

12. The system according to any one of claims 3 to 11, wherein the bit budget allocator uses the bit rate adaptation algorithm with an audio stream having active content to adjust the bit rate for coding one of the audio streams based on the classification of the importance of the audio stream and metadata (ISm).

13. The system according to claim 11, wherein the bit budget allocator uses the bit rate adaptation algorithm with an audio stream having non-active content or meaningless content to reduce the bit budget for coding the audio stream and set it to a constant value.

14. The system according to claim 11 or 13, wherein the bit budget allocator calculates the saved bit budget as the difference between the reduced value of the bit budget for coding the audio stream and the non-reduced value of the bit budget for coding the audio stream.

15. The system according to claim 13 or 14, wherein the bit budget allocator calculates the bit rate for coding the audio stream using the reduced value of the bit budget.

16. The system according to claim 12, wherein the bit budget allocator classifies the importance of the ISm based on an indicator indicating how important the coding of the audio object is to obtain a given quality of the decoded synthesis.

17. The system according to claim 12 or 16, wherein the bit budget allocator classifies the importance of the ISm based on at least one of the following parameters: encoder type of the audio stream, FEC (Forward Error Correction), classification of the audio signal, classification of human voice / music, and SNR (Signal-to-Noise Ratio) estimate value.

18. The system according to any one of claims 12, 16, and 17, wherein the bit budget allocator uses the classification of the importance of the ISm in the bit rate adaptation algorithm to increase the bit budget for coding the audio stream with a higher importance of the ISm and decrease the bit budget for coding the audio stream with a lower importance of the ISm.

19. A method for coding an object-based audio signal including an audio object according to an audio stream having associated metadata, coding the metadata before and separately from coding the audio stream using processing adapted to the metadata; generating information about a bit budget used for coding the metadata of the audio object after the metadata is coded; allocating a bit rate for coding the audio stream according to the information about the bit budget used for coding the metadata of the audio object; coding the audio stream using the bit rate allocated for coding the audio stream; and including the steps.

20. The method according to claim 19, including analyzing the audio stream and providing information about the audio stream for coding the metadata and for allocating the bit rate for coding the audio stream.

21. The method according to claim 19 or 20, wherein the allocation of the bitrate for the coding of the audio stream includes using a bitrate adaptation algorithm for allocating an available bit budget for coding the audio stream.

22. The method according to claim 21, wherein the allocation of the bitrate for the coding of the audio stream using the bitrate adaptation algorithm includes calculating the total bit budget of the audio stream and metadata (ISm) or the total bitrate of the codec for coding the audio stream and the associated metadata, and subtracting the total bit budget of the ISm therefrom.

23. The method according to claim 22, wherein the allocation of the bitrate for the coding of the audio stream using the bitrate adaptation algorithm includes calculating the bit budget of an element by dividing the total bit budget of the ISm by the number of audio streams.

24. The method according to claim 23, wherein the allocation of the bitrate for the coding of the audio stream using the bitrate adaptation algorithm includes adjusting the bit budget of the element of the last audio object to use up the total bit budget of the ISm.

25. The method according to claim 23 or 24, wherein the allocation of the bitrate for the coding of the audio stream using the bitrate adaptation algorithm includes summing the bit budget for coding the metadata of the audio object, adding the sum to the bit budget of the ISm common signaling, and generating a side bit budget of the codec.

26. The allocation of the bitrate for the coding of the audio stream using the bitrate adaptation algorithm includes (a) evenly dividing the side bit budget of the codec among the audio objects, and (b) calculating the coding bit budget for each audio stream using the divided side bit budget of the codec and the bit budget of the elements. The method according to claim 25.

27. The method according to claim 26, wherein the allocation of the bitrate for the coding of the audio stream using the bitrate adaptation algorithm includes adjusting the coding bit budget of the last audio stream to use all available coding bit budgets.

28. The method according to claim 26 or 27, wherein the allocation of the bitrate for the coding of the audio stream using the bitrate adaptation algorithm includes calculating the bitrate for coding one of the audio streams using the coding bit budget of the audio stream.

29. The method according to any one of claims 21 to 28, wherein the allocation of the bitrate for the coding of the audio stream using the bitrate adaptation algorithm with an audio stream having non-active content or no meaningful content includes reducing the value of the bitrate for coding one of the audio streams and redistributing the saved bit budget among the audio streams having active content.

30. The method according to any one of claims 21 to 29, wherein the allocation of the bitrate for the coding of the audio stream using the bitrate adaptation algorithm with an audio stream having active content includes adjusting the bitrate for coding one of the audio streams based on the classification of the importance of the audio stream and metadata (ISm).

31. The method according to claim 29, wherein the allocation of the bitrate for the coding of the audio stream uses the bitrate adaptation algorithm with an audio stream having non-active content or having no meaningful content, and includes reducing and setting to a constant value a bit budget for coding the audio stream.

32. The method according to claim 29 or 31, wherein the allocation of the bitrate for the coding of the audio stream includes calculating the saved bit budget as the difference between the reduced value of the bit budget for coding the audio stream and the non-reduced value of the bit budget for coding the audio stream.

33. The method according to claim 31 or 32, wherein the allocation of the bitrate for the coding of the audio stream includes calculating a bitrate for coding the audio stream using the reduced value of the bit budget.

34. The method according to claim 30, wherein the allocation of the bitrate for the coding of the audio stream includes classifying the importance of the ISm based on an indicator indicating how important the coding of the audio object is to obtain a given quality of the decoded synthesis.

35. The method according to claim 30 or 34, wherein the allocation of the bitrate for the coding of the audio stream includes classifying the importance of the ISm based on at least one of the following parameters: encoder type of the audio stream, FEC (Forward Error Correction), classification of the audio signal, classification of human voice / music, and SNR (Signal-to-Noise Ratio) estimated value.

36. The method according to any one of claims 30, 34, and 35, wherein the allocation of the bitrate for the coding of the audio stream includes using the classification of the importance of the ISm in the bitrate adaptation algorithm to increase the bit budget for the coding of the audio stream having a higher importance of the ISm and decrease the bit budget for the coding of the audio stream having a lower importance of the ISm.

Citation Information

Patent Citations

  • Post-encoding bitrate reduction for multiple object audio

    JP2017507365A

  • PCT/CA2018/51175

  • Audio encoding device and audio decoding device

    WO2015056383A1

  • Method and device for allocating a bit-budget between sub-frames in a CELP codec

    WO2019056107A1

  • Encoding device and method, decoding device and method, and program

    WO2019069710A1