Method and system for metadata in coded audio stream and for allocation of effective bitrate for coding of audio stream

By combining metadata processing and a bit budget allocator, efficient object-based encoding and decoding of audio signals is achieved, solving the problem of insufficient flexible bitrate adaptation in existing technologies and improving the interactivity and metadata processing capabilities of the immersive audio experience.

CN114072874BActive Publication Date: 2025-10-17VOICEAGE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080050126.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-08
Filing Date
2020-07-07
Publication Date
2025-10-17
Estimated Expiration
2040-07-07

AI Technical Summary

Technical Problem

Existing audio codec technologies find it difficult to effectively implement flexible bitrate adaptation and efficient encoding and decoding of object-based audio signals, resulting in insufficient interactivity and flexibility in the immersive audio experience.

Method used

A metadata processor is used to encode and decode the metadata of audio objects, and a bit budget allocator is used to allocate bit rates for audio streams to achieve bit rate adaptation within and between objects. Encoding and decoding are combined with the core encoder to ensure efficient metadata processing and flexible allocation of audio streams.

Benefits of technology

It achieves efficient encoding and decoding of object-based audio signals, improves the interactivity and flexibility of immersive audio experience, enhances the ability to process metadata, and reduces the risk of error propagation in noisy channels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114072874B_ABST
    Figure CN114072874B_ABST
Patent Text Reader

Abstract

A system and method encode object-based audio signals including audio objects in response to an audio stream having associated metadata. In the system and method, a metadata processor encodes the metadata and generates information regarding a bit budget for encoding metadata of the audio objects. An encoder encodes the audio stream and a bit budget allocator allocates a bit rate for encoding the audio stream by the encoder in response to the information from the metadata processor regarding the bit budget for encoding metadata of the audio objects.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to sound coding, and more specifically, to techniques for digitally coding object-based audio, e.g. speech, music or general audio sounds. In particular, the present disclosure relates to systems and methods for coding, and to systems and methods for decoding object-based audio signals comprising audio objects in response to audio streams with associated metadata.

[0002] In the present disclosure and the appended claims:

[0003] (a) The term "object-based audio" is intended to represent a complex audio auditory scene as a collection of individual elements, also called audio objects. Moreover, as mentioned above, "object-based audio" can comprise e.g. speech, music or general audio sounds.

[0004] The term "audio object" is intended to designate an audio stream with associated metadata. For example, in the present disclosure, an "audio object" is referred to as an independent audio stream with metadata (ISm).

[0005] The term "audio stream" is intended to represent an audio waveform, e.g. speech, music or general audio sounds, in a bitstream, and can consist of one channel (mono), although two channels (stereo) are also considered. "Mono" is an abbreviation for "monophonic" and "stereo" is an abbreviation for "stereophonic".

[0006] The term "metadata" is intended to represent a set of information describing the audio stream and the artistic intension for translating the original or coded audio objects to a reproduction system. The metadata typically describes spatial properties of each individual audio object, such as position, direction, volume, width, etc. In the context of the present disclosure, two sets of metadata are considered:

[0007] - input metadata: unquantized metadata representation used as input for the codec; the present disclosure is not limited to a specific format of input metadata; and

[0008] - coded metadata: quantized and coded metadata constituting a part of the bitstream sent from the encoder to the decoder.

[0009] (e) The term "audio format" is intended to designate a method to achieve an immersive audio experience.

[0010] (f) The term "rendering system" is intended to designate elements in the decoder that can use the transmitted metadata and artistic intent on the rendering side, for example, but not limited to, rendering audio objects in a 3D (three-dimensional) audio space around the listener. The rendering can be performed to a target loudspeaker layout (e.g., 5.1 surround sound) or headphones, while the metadata can be dynamically modified, for example, in response to feedback from a head-tracking device. Other types of rendering can be considered. BACKGROUND

[0011] Over the past few years, the generation, recording, representation, coding, transmission, and rendering of audio are evolving towards enhanced, interactive, and immersive listener experiences. An immersive experience can be described as a state of deep engagement or involvement in a sound scene, for example, when sound comes from all around. In immersive audio (also referred to as 3D audio), sound images are reproduced in all three dimensions of space around the listener, taking into account a wide range of sound features such as timbre, directionality, reverberation, transparency, and (auditory) spaciousness, with accuracy. Immersive audio is produced for a given rendering system (i.e., a loudspeaker configuration, an integrated rendering system (soundbar), or headphones). The interactivity of the audio rendering system can then include, for example, the ability to adjust sound levels, change sound locations, or select different languages for reproduction.

[0012] There are three basic approaches (also referred to as audio formats below) that can enable immersive audio experiences.

[0013] The first approach is channel-based audio, where multiple spaced-apart microphones are used to capture sound from different directions, with one microphone corresponding to one audio channel in a specific loudspeaker layout. Each recorded channel is provided to a loudspeaker in a specific position. Examples of channel-based audio include, for example, stereo, 5.1 surround sound, 5.1+4, etc.

[0014] The second approach is scene-based audio, which represents a desired sound field over a local space as a function of time through a combination of dimensional components. The signals representing scene-based audio are independent of the position of the audio sources, while the sound field must be converted to a chosen loudspeaker layout at the rendering rendering system. One example of scene-based audio is ambisonics.

[0015] The third, and last, immersive audio approach is object-based audio, which represents an auditory scene as a collection of individual audio elements (e.g., a singer, a drum, a guitar), accompanied by information about, for example, their position in the audio scene, so that they can be rendered to their intended position at the rendering system. This provides great flexibility and interactivity for object-based audio, as each object is discrete and can be manipulated individually.

[0016] Each of the above audio formats has its advantages and disadvantages. Therefore, it is common that not only one specific format is used in an audio system, but they can be combined in complex audio systems to create immersive listening scenarios. One example can be a system that combines scene-based or channel-based audio with object-based audio (e.g. ambisonics with few discrete audio objects).

[0017] The present disclosure presents in the following description a framework for encoding and decoding object-based audio. This framework can be a standalone system for object-based audio format codec, or it can form part of a complex immersive codec that can contain the codec of other audio formats and / or combinations thereof. SUMMARY

[0018] According to a first aspect, the present disclosure provides a system for coding object-based audio signals comprising audio objects in response to an audio stream having associated metadata, comprising a metadata processor for coding the metadata, the metadata processor generating information on a bit-budget for coding of the metadata for the audio objects. An encoder codes the audio stream, and a bit-budget allocator allocates a bit-rate for coding of the audio stream by the encoder in response to the information on the bit-budget for coding of the metadata for the audio objects from the metadata processor.

[0019] The present disclosure also provides a method for coding object-based audio signals comprising audio objects in response to an audio stream having associated metadata, comprising encoding the metadata, generating information on a bit-budget for coding of the metadata for the audio objects, encoding the audio stream, and allocating a bit-rate for coding of the audio stream in response to the information on the bit-budget for coding of the metadata for the audio objects.

[0020] According to a third aspect, a system for decoding audio objects in response to an audio stream having associated metadata is provided, comprising a metadata processor for decoding the metadata of the audio objects and for providing information on a respective bit-budget of the metadata of the audio objects, a bit-budget allocator determining a core decoder bit-rate of the audio stream in response to the metadata bit-budget of the audio objects, and a decoder of the audio stream using the core decoder bit-rate determined in the bit-budget allocator.

[0021] The present disclosure also provides a method for decoding an audio object in response to an audio stream having associated metadata, comprising: decoding the metadata of the audio object and providing information about a corresponding bit budget of the metadata of the audio object, determining a core decoder bit rate of the audio stream using the metadata bit budget of the audio object, and decoding the audio stream using the determined core decoder bit rate.

[0022] The foregoing and other objects, advantages and features of systems and methods for encoding and decoding object-based audio signals and systems and methods for decoding object-based audio signals will become more apparent upon reading the following non-limiting description of illustrative embodiments thereof, given by way of example only with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In the attached figure:

[0024] Figure 1 is a schematic block diagram illustrating a system for encoding and decoding an object-based audio signal and a corresponding method for encoding and decoding an object-based audio signal;

[0025] Figure 2 is a diagram illustrating different scenarios of bitstream encoding and decoding of a metadata parameter;

[0026] Figure 3 a is an absolute codec flag indicating metadata parameters of three (3) audio objects without using inter-object metadata codec logic abs A graph of the values ​​of , and Figure 3 b is an absolute codec flag showing metadata parameters of three (3) audio objects using inter-object metadata codec logic abs a graph showing the values ​​of , wherein arrows indicate frames where the value of the absolute codec flag is equal to 1;

[0027] Figure 4 is a graph illustrating an example of bitrate adaptation for three (3) core encoders;

[0028] Figure 5 is a graph illustrating an example of bitrate adaptation based on ISm (Independent Audio Stream with Metadata) importance logic;

[0029] Figure 6 It is shown from Figure 1 The codec system is sent to Figure 7 A schematic diagram of the structure of a bit stream of a decoding system;

[0030] Figure 7is a schematic block diagram simultaneously showing a system for decoding audio objects in response to audio streams having associated metadata and a corresponding method for decoding audio objects; and

[0031] Figure 8 is a simplified block diagram of an example configuration of the system and method for coding object-based audio signals and hardware components of the system and method for decoding object-based audio signals. DETAILED DESCRIPTION

[0032] The present disclosure provides examples of mechanisms for coding metadata. The present disclosure also provides a mechanism for flexible intra- and inter-object bitrate adaptation, i.e. a mechanism to distribute available bitrate as efficiently as possible. In the present disclosure, it is further assumed that the bitrate is fixed (constant). However, it is similarly within the scope of the present disclosure to consider adaptive bitrate, e.g. (a) in a codec based on adaptive bitrate, or (b) as a result of coding a combination of audio formats coded at a fixed total bitrate.

[0033] It is not described in the present disclosure how to actually code audio streams in the so-called "core encoder". Typically, the core encoder for coding one audio stream can be any mono codec using adaptive bitrate coding. One example is a codec based on the EVS codec described in reference [1] with a flexible and efficient distribution of fluctuating bit-budgets among the modules of the core encoder, e.g. as described in reference [2]. The entire contents of references [1] and [2] are hereby incorporated by reference.

[0034] 1. Framework for coding audio objects

[0035] As a non-limiting example, the present disclosure considers a framework that supports the simultaneous coding of several audio objects (e.g. up to 16 audio objects) while a fixed constant ISm total bitrate, referred to as ism_total_brate, is considered for coding the audio objects including audio streams having associated metadata. It should be noted that for at least some audio objects, metadata does not necessarily have to be transmitted, e.g. in case of non-diegetic content. Non-diegetic sound in movies, TV series and other videos is sound that the characters cannot hear. Music is one example of non-diegetic sound, as the listener is the only one to hear the music.

[0036] In case of coding a combination of audio formats in the framework, for example, an ambisonic audio format with two (2) audio objects, the constant total codec bit rate referred to as codec_total_brate then represents the sum of the ambisonic audio format bit rate (i.e. the bit rate at which the ambisonic audio format is encoded) and the ISm total bit rate ism_total_brate (i.e. the sum of the bit rates at which the audio objects (i.e. the audio streams with associated metadata) are coded).

[0037] The present disclosure considers the basic non-limiting example of input metadata comprising two parameters, namely the azimuth and the elevation, which are stored per audio frame for each object. In this example, an azimuth range of [-180°, 180°] and an elevation range of [-90°, 90°] are considered. However, considering only one or more than two (2) metadata parameters is also within the scope of the present disclosure.

[0038] 2. Object-based coding

[0039] Figure 1 is a schematic block diagram simultaneously showing a system 100 for coding an object-based audio signal comprising several processing blocks and a corresponding method 150 for coding an object-based audio signal.

[0040] 2.1 Input buffering

[0041] With reference to Figure 1 The method 150 for coding an object-based audio signal comprises an operation 151 of input buffering. To perform the operation 151 of input buffering, the system 100 for coding an object-based audio signal comprises an input buffer 101.

[0042] The input buffer 101 buffers a number N of input audio objects 102, i.e. a number N of audio streams and the associated respective N metadata. The N input audio objects 102 comprising the N audio streams and the N metadata associated with each of these N audio streams are buffered for one frame, for example a 20 milliseconds long frame. As it is well known in the field of sound signal processing, sound signals are sampled at a given sampling frequency and are processed by successive blocks of samples called “frames”, each frame being divided into a number of “sub-frames”.

[0043] 2.2 Audio stream analysis and pre-processing

[0044] Still with reference to Figure 1The method 150 for encoding and decoding an object-based audio signal includes an operation 153 of analyzing and pre-processing N audio streams. To perform operation 153, the system 100 for encoding and decoding an object-based audio signal includes an audio stream processor 103 for analyzing and pre-processing the buffered N audio streams respectively sent from the input buffer 101 to the audio stream processor 103 through the N transmission channels 104, for example, in parallel.

[0045] The operation 153 performed by the audio stream processor 103 may include, for example, at least one of the following sub-operations: time domain transient detection, spectrum analysis, long-term prediction analysis, pitch tracking and voicing analysis, voice / sound activity detection (VAD / SAD), bandwidth detection, noise estimation, and signal classification (which, in a non-limiting embodiment, may include (a) core encoder selection between, for example, an ACELP core encoder, a TCX core encoder, an HQ core encoder, etc., (b) signal type classification between, for example, an inactive core encoder type, a silent core encoder type, a voiced core encoder type, a universal core encoder type, a conversion core encoder type, and an audio core encoder type, (c) speech / music classification, etc.). Information obtained from the operation 153 is provided to the configuration and decision processor 106 via line 11a 121. Examples of the aforementioned sub-operations related to the EVS codec are described in reference [1] and will not be further described in this disclosure.

[0046] 2.3 Metadata Analysis, Quantization, and Encoding

[0047] For encoding and decoding object-based audio signals Figure 1 The method 150 includes operations 155 of metadata analysis, quantization, and encoding. To perform the operations 155, the system 100 for encoding and decoding an object-based audio signal includes a metadata processor 105.

[0048] 2.3.1 Metadata Analysis

[0049] Signal classification information 120 from the audio stream processor 103 (e.g., the VAD or localVAD flag used in the EVS codec (see reference [1])) is provided to the metadata processor 105. The metadata processor 105 includes an analyzer (not shown) of the metadata for each of the N audio objects to determine whether the current frame is inactive (e.g., VAD = 0) or active (e.g., VAC ≠ 0) with respect to that particular audio object. In inactive frames, the metadata associated with the object is not encoded or decoded by the metadata processor 105. In active frames, the metadata for the audio object is quantized and encoded using a variable bit rate. More details on metadata quantization and encoding are provided in Sections 2.3.2 and 2.3.3 below.

[0050] 2.3.2 Metadata Quantification

[0051] In the described non-limiting illustrative embodiment, Figure 1 The metadata processor 105 sequentially quantizes and encodes metadata of N audio objects in a loop, and a certain correlation may be adopted between the quantization of the audio objects and the metadata parameters of these audio objects.

[0052] As mentioned above, in the present disclosure, two metadata parameters, azimuth and elevation, are considered (included in the N input metadata). As a non-limiting example, the metadata processor 105 includes a quantizer (not shown) indexed by the following metadata parameters, which uses the following example resolutions to reduce the number of bits being used:

[0053] - Azimuth parameter: The 12-bit azimuth parameter index from the input metadata file is quantized to B az Bit index (e.g., B az = 7). Given the minimum and maximum azimuth limits (-180° and +180°), (B az = 7) The quantization step size of the bit-uniform scalar quantizer is 2.835°.

[0054] - Elevation parameter: The 12-bit elevation parameter index from the input metadata file is quantized to B el Bit index (e.g., B el = 6). Given the minimum and maximum elevation limits (-90° and +90°), (B el = 6) The quantization step size of the bit-uniform scalar quantizer is 2.857°.

[0055] The total metadata bit-budget for coding N metadata and the total number of quantization bits used for quantizing the metadata parameter indices (i.e. the quantization index granularity, hence the resolution) can be determined depending on the bit-rate codec_total_brate, ism_total_brate and / or element_brate (the latter resulting from the sum of the metadata bit-budget and / or the core encoder bit-budget related to one audio object).

[0056] Azimuth and elevation parameters can be represented as one parameter, e.g. by a point on a sphere. It is within the scope of the present disclosure to implement different metadata comprising two or more parameters.

[0057] 2.3.3 Metadata coding

[0058] Both azimuth and elevation indices, once quantized, can be coded by the metadata encoder (not shown) of the metadata processor 105 using absolute or differential coding. Absolute coding means, as known, that the current value of the parameter is coded. Differential coding means that the difference between the current and previous value of the parameter is coded. Since the indices of azimuth and elevation parameters usually evolve smoothly (i.e. the change of azimuth or elevation position can be considered as continuous and smooth), differential coding is used by default. However, absolute coding can be used, for example, in the following cases:

[0059] - the difference between the current and previous value of the parameter index is too large, which would result in a higher or equal number of bits using differential coding compared to absolute coding (exception can occur);

[0060] - no metadata was coded and sent in the previous frame;

[0061] - too many consecutive frames are coded with differential coding. To control the noise in the channel (bad frame indicator, BFI = 1). For example, if the number of consecutive frames coded with differential is higher than the maximum number of consecutive frames coded with a different coding, the metadata encoder codes the metadata parameter indices with absolute coding. The maximum number of consecutive frames is set to β. In a non-limiting illustrative example, β = 10 frames.

[0062] The metadata encoder generates a 1-bit absolute coding flag flag abs to distinguish between absolute and differential coding.

[0063] In the case of absolute coding, the coding flag flag abs is set to 1, followed by the B az its (or B elB az and B el refer respectively to the index of the azimuth and elevation parameters to be coded / decoded.

[0064] In the case of differential coding, a 1-bit coding flag flag abs is set to 0, followed by a 1-bit zero coding flag flag zero signaling that the difference Δ between the B az bit index in the current and previous frames (in the other case the B el bit index) is equal to 0. If the difference Δ is not equal to 0, the metadata encoder continues the coding by producing a 1-bit sign flag flag sign followed by a difference index whose number of bits is adaptive, in the form of for example a unary code indicating the difference Δ.

[0065] Figure 2 is a diagram illustrating different scenarios of bitstream coding of one metadata parameter.

[0066] Referring to Figure 2 , it is noted that not all metadata parameters are always transmitted in each frame. Some can be transmitted only in every y-th frame, some are not transmitted at all, for example when they do not evolve, they are not important or the available bit budget is low. Referring to Figure 2 , for example:

[0067] - in the case of absolute coding (first line of Figure 2 ), the absolute coding flag flag abs and the B az bit index (in the other case the B el bit index) are transmitted;

[0068] - in the case of differential coding where the difference Δ between the B az bit index in the current and previous frames (in the other case the B el bit index) is equal to 0 (second line of Figure 2 ), the absolute coding flag flag abs = 0 and the zero coding flag flag zero = 1 are transmitted;

[0069] - in the case of differential coding where the positive difference Δ between the B az bit index in the current and previous frames (in the other case the B el bit index) (third line of Figure 2 ), the absolute coding flag flag abs = 0, the zero coding flag flag zero = 0, the sign flag flagsign = 0 and difference index (1 to (B az - 3) bit index (in other cases 1 to (B el - 3) bit index) are sent; and

[0070] - in case of differential coding of the negative difference D between the B az bit index (in other cases B el bit index) in the current frame and the previous frame, Figure 2 the last row of the absolute coding flag flag abs = 0, the zero coding flag flag zero = 0, the sign flag flag sign = 1 and difference index (1 to (B az - 3) bit index (in other cases 1 to (B el - 3) bit index) are sent.

[0071] 2.3.3.1 Intra-object metadata coding logic

[0072] The logic for setting absolute or differential coding can be further extended by the intra-object metadata coding logic. In particular, to limit the range of metadata coding bit budget fluctuations between frames, thereby avoiding too low bit budget remaining for the core encoder 109, the metadata encoder limits the absolute coding to one, or generally as few as possible, metadata parameters in a given frame.

[0073] In a non-limiting example of azimuth and elevation metadata parameter coding, the metadata encoder uses a logic that avoids absolute coding of the elevation index in the same frame if the azimuth bit index has been coded using absolute coding in the given frame. In other words, the azimuth and elevation parameters of one audio object are (in practice) never coded using absolute coding in the same frame. Thus, if the absolute coding flag flag abs.azi for the azimuth parameter is equal to 1, the absolute coding flag flag abs.ele for the elevation parameter is not sent in the audio object bitstream.

[0074] It is also within the scope of the present disclosure to make the intra-object metadata coding logic dependent on the bit rate. For example, if the bit rate is large enough, the absolute coding flag flag abs.ele for the elevation parameter and the absolute coding flag flag abs.azi for the azimuth parameter can be sent in the same frame.

[0075] 2.3.3.2 Inter-object metadata coding logic

[0076] The metadata encoder can apply similar logic to the metadata coding of different audio objects. The implemented inter-object metadata coding logic minimizes the number of metadata parameters of different audio objects coded using absolute coding in the current frame. This is achieved by the metadata encoder mainly by controlling the frame counter of metadata parameters coded using absolute coding selected according to robustness purposes and denoted by the parameter β. As a non-limiting example, consider a scenario where the metadata parameters of audio objects evolve slowly and smoothly. To control the decoding in the noise channel, the index is coded using absolute coding every β frames, the azimuth B az of audio object #1 is coded using absolute coding in frame M, the index is coded using absolute coding in frame M+1, the elevation B el of audio object #1 is coded using absolute coding in frame M+2, the index is coded using absolute coding in frame M+3, the azimuth B az of audio object #2 is coded using absolute coding in frame M+4, the index is coded using absolute coding in frame M+5, the elevation B el of audio object #2 is coded using absolute coding in frame M+6, etc.

[0077] Figure 3 a is a graph showing the values of the absolute coding flag flag abs of the metadata parameters of three (3) audio objects without using inter-object metadata coding logic, and Figure 3 b is a graph showing the values of the absolute coding flag flag abs of the metadata parameters of three (3) audio objects using inter-object metadata coding logic. In Figure 3 a, the arrows indicate the frames where the value of several absolute coding flags is equal to 1.

[0078] More specifically, Figure 3 a shows the values of the absolute coding flag flag abs of two metadata parameters (in this particular example, the azimuth and the elevation) of audio objects without using inter-object metadata coding logic, while Figure 3 b shows the same values but with the inter-object metadata coding logic implemented. Figure 3 The graphs of a and Figure 3 b correspond to (from top to bottom):

[0079] - the audio stream of audio object #1;

[0080] - the audio stream of audio object #2;

[0081] - the audio stream of audio object #3,

[0082] - the absolute coding flag flag of the azimuth parameter of audio object #1abs,azi ;

[0083] - Absolute coding flag flag of the elevation parameter of audio object #1 abs,ele ;

[0084] - Absolute coding flag flag of the azimuth parameter of audio object #2 abs,azi ;

[0085] - Absolute coding flag flag of the elevation parameter of audio object #2 abs,ele ;

[0086] - Absolute coding flag flag of the azimuth parameter of audio object #3 abs,azi ; and

[0087] - Absolute coding flag flag of the elevation parameter of audio object #3 abs,ele .

[0088] From Figure 3 a, it can be seen that, when the inter-object metadata coding logic is not used, in the same frame, several absolute coding flags flag abs may have a value equal to 1 (see arrows). Conversely, Figure 3 b shows that, when the inter-object metadata coding logic is used, in a given frame, only one absolute coding flag flag abs may have a value equal to 1.

[0089] The inter-object metadata coding logic can also rely on the bitrate. In this case, for example, if the bitrate is large enough, then even when the inter-object metadata coding logic is used, in a given frame, more than one absolute coding flag flag abs may have a value equal to 1.

[0090] The technical advantages of the inter-object metadata coding logic and the intra-object metadata coding logic are to limit the range of fluctuation of the metadata coding bit budget between frames. Another technical advantage is to increase the robustness of the codec in noisy channels; when a frame is lost, then only a limited number of metadata parameters from the audio objects coded using absolute coding are lost. Therefore, any error propagated from the lost frame only affects a small number of metadata parameters on the audio objects, and therefore does not affect the entire audio scene (or several different channels).

[0091] As mentioned above, the overall technical advantage of analyzing, quantizing and coding the metadata separately from the audio stream is that it is possible to carry out a processing that is particularly adapted to the metadata and that is more efficient in terms of metadata coding bitrate, metadata coding bit budget fluctuation, robustness in noisy channels and error propagation due to lost frames.

[0092] The quantized and coded metadata 112 from the metadata processor 105 is provided to the multiplexer 110 for insertion into the output bitstream 111 sent to the decoding system 700 Figure 7 ).

[0093] Once the metadata of the N audio objects is analyzed, quantized and encoded, the information 107 from the metadata processor 105 about the coded bit-budget of the metadata per audio object is provided to the configuration and decision processor 106 (bit-budget allocator), which will be described in more details in section 2.4 below. When the configuration and bit-rate distribution between the audio streams is done in the configuration and decision processor 106 (bit-budget allocator), the coding is continued by further pre-processing 158, which will be described later. Finally, the N audio streams are encoded using an encoder, which comprises for example N fluctuating bit-rate core encoders 109, such as single core encoders.

[0094] 2.4 Bit-rate configuration and decision per channel

[0095] Method for coding object-based audio signals Figure 1 The method 150 for coding object-based audio signals comprises an operation 156 of configuration and decision of the bit-rate per each of the N transport channels 104. To perform the operation 156, the system 100 for coding object-based audio signals comprises a configuration and decision processor 106 forming a bit-budget allocator.

[0096] The configuration and decision processor 106 (hereafter in the bit-budget allocator 106) uses a bit-rate adaptation algorithm to distribute the available bit-budget for the core encoding of the N audio streams in the N transport channels 104.

[0097] The bit-rate adaptation algorithm of the operation 156 comprises the following sub-operations 1-6 performed by the bit-budget allocator 106:

[0098] 1. The ISm total bit-budget per frame is calculated from the ISm total bit-rate ism_total_brate (or the codec total bit-rate codec_total_brate if only audio objects are coded) using for example the following relation:

[0099]

[0100] The denominator 50 corresponds to the number of frames per second, assuming a frame length of 20 milliseconds. If the size of the frame is different from 20 milliseconds, the value 50 will be different.

[0101] 2. The above element bitrate element_brate defined for N audio objects (resulting from the sum of the metadata bit-budget and the core encoder bit-budget associated with one audio object) should be constant during a session at a given codec total bitrate, and approximately the same for the N audio objects. A "session" is defined as for example a phone call or offline compression of an audio file. The corresponding element bit-budget bits element :

[0102]

[0103] where represents the largest integer less than or equal to x. In order to spend all available ISm total bit-budget bits ism , the element bit-budget bits element of the last audio object is for example adjusted using the following relation:

[0104]

[0105] where "mod" represents the remainder modulo operation. Finally, the element bit-budget bits element of the N audio objects is used to set the value element_brate of the audio objects n = 0,..., N-1 using for example the following relation:

[0106]

[0107] where the number 50 corresponds to the number of frames per second as already mentioned, assuming a frame length of 20 milliseconds.

[0108] 3. The metadata bit-budget bits meta of each frame for the N audio objects is summed using the following relation:

[0109]

[0110] and the resulting value bits metal_all is added to the ISm common signaling bit-budget bits Ism_signalling resulting in the codec side bit-budget:

[0111]

[0112] 4. The codec side bit-budget bits side of each frame is split equally among the N audio objects and used to calculate the core encoder bit-budget bits CoreCoder of each of the N audio streams using for example the following relation:

[0113]

[0114] While the core encoder bit-budget of e.g. the last audio stream can eventually be adjusted to spend all available core encoding bit-budget using e.g. the following relation:

[0115]

[0116] Then, for n = 0,..., N-1, the corresponding total bitrate total_brate, i.e. the bitrate at which an audio stream is coded in the core encoder, is obtained using e.g. the following relation:

[0117]

[0118] where the number 50 again corresponds to the number of frames per second, assuming a frame length of 20 ms.

[0119] 5. The total bitrate total_brate in inactive frames (or frames with very low energy or otherwise without meaningful content) can be reduced and set to a constant value across the relevant audio streams. The bit-budget thus saved is then re-distributed evenly across the audio streams that have active content in the frame. This re-distribution of bit-budget will be further described in section 2.4.1 below.

[0120] 6. The total bitrate total_brate in active frames (with active content) is further adjusted across the audio streams based on the ISm importance classification. This adjustment of bitrate will be further described in section 2.4.2 below.

[0121] The last two sub-operations 5 and 6 above can be skipped when all audio streams are in inactive segments (or without meaningful content). Thus, the bitrate adaptation algorithms described in sections 2.4.1 and 2.4.2 below are employed when at least one audio stream has active content.

[0122] 2.4.1 Bitrate adaptation based on signal activity

[0123] In inactive frames (VAD = 0), the total bitrate total_brate is reduced and the saved bit-budget is re-distributed, e.g. evenly across the audio streams in active frames (VAD ≠ 0). It is assumed that no waveform coding is required for the audio streams in frames that are classified as inactive; the audio objects can be muted. The logic used in each frame can be represented by the following sub-operations 1-3:

[0124] 1. For a particular frame, set a lower core encoder bit-budget for each audio stream n that has inactive content:

[0125]

[0126] wherein is a lower, constant core encoder bit budget to be set in inactive frames; for example, = 140 (for 20 ms frames, corresponding to 7 kbps) or = 49 (for 20 ms frames, corresponding to 2.45 kbps).

[0127] 2. Next, the saved bit budget is calculated using, for example, the following relationship:

[0128]

[0129] 3. Finally, the saved bit budget is redistributed, for example, evenly distributed among the core encoder bit budgets of the audio streams having active content in the given frame, using, for example, the following relationship:

[0130]

[0131] wherein is the number of audio streams having active content. The core encoder bit budget of the first audio stream having active content is finally increased using, for example, the following relationship:

[0132]

[0133] For each audio stream n = 0,..., N-1, the corresponding core encoder total bit rate total_brate is finally obtained as follows:

[0134]

[0135] Figure 4 is a plot showing an example of bit rate adaptation of three (3) core encoders. In particular, in Figure 4 , the first row shows the core encoder total bit rate total_brate of audio stream #1, the second row shows the core encoder total bit rate total_brate of audio stream #2, the third row shows the core encoder total bit rate total_brate of audio stream #3, the fourth row is audio stream #1, the fifth row is audio stream #2, and the fourth row is audio stream #3.

[0136] In the example of Figure 4 , the adaptation of the total bit rate total_brate of the three (3) core encoders is based on VAD activity (active / inactive frames). From Figure 4It can be seen that most of the time, due to fluctuating side bit budgets bits side , the core encoder total bit rate total_brate has a small fluctuation. Then, due to VAD activity, the core encoder total bit rate total_brate has a rare substantial change.

[0137] For example, referring to Figure 4 , instance A) corresponds to a frame where the VAD activity of audio stream #1 changes from 1 (active) to 0 (inactive). According to this logic, the minimum core encoder total bit rate total_brate is assigned to audio object #1, while the core encoder total bit rates total_brate of active audio objects #2 and #3 are increased. Instance B) corresponds to a frame where the VAD activity of audio stream #3 changes from 1 (active) to 0 (inactive) while the VAD activity of audio stream #1 remains 0. According to this logic, the minimum core encoder total bit rate total_brate is assigned to audio streams #1 and #3, while the core encoder total bit rate total_brate of active audio stream #2 is further increased.

[0138] The above logic of section 2.4.1 can rely on the total bit rate ism_total_brate. For example, for higher total bit rates ism_total_brate, the bit budget in sub-operation 1 above may be set higher, while for lower total bit rates ism_total_brate, it can be set lower.

[0139] 2.4.2 ISm importance based bit rate adaptation

[0140] The logic described in section 2.4.1 above leads to approximately the same core encoder bit rate in each audio stream that has active content (VAD = 1) in a given frame. However, ISm importance based classification (or more generally, based on a metric that indicates how critical the coding of a particular audio object in the current frame is for obtaining a given (adequate) quality of the decoded synthesis), it can be beneficial to introduce inter-object core encoder bit rate adaptation.

[0141] ISm importance classification can be based on several parameters and / or parameter combinations, e.g., core encoder type (coder_type), FEC (forward error correction), sound signal class (class), speech / music classification decision, and / or SNR (signal-to-noise ratio) estimates from the open-loop ACELP / TCX (Algebraic Code-Excited Linear Prediction / Transform Coded Excitation) core decision module (snr_celp, snr_tcx) described in reference [1]. Other parameters can be used to determine the ISm importance classification.

[0142] In one non-limiting example, the simple classification of ISm importance is based on the core coder type defined in reference [1]. To this end, Figure 1 The bit-budget allocator 106 comprises a classifier (not shown) for assessing the importance of a particular ISm stream. As a result, four (4) different ISm importance classes class ISm are defined:

[0143] - No metadata class, ISM_NO_META: frames without metadata coding, e.g. inactive frames with VAD = 0;

[0144] - Low importance class, ISM_LOW_IMP: frames with coder_type = UNVOICED or INACTIVE;

[0145] - Medium importance class, ISM_MEDIUM_IMP: frames with coder_type = VOICED;

[0146] - High importance class, ISM_HIGH_IMP: frames with coder_type = GENERIC.

[0147] The bit-budget allocator 106 then uses the ISm importance classes in the bit-rate adaptation algorithm (see section 2.4 above, sub-operation 6) to assign higher bit-budgets to audio streams with higher ISm importance and lower bit-budgets to audio streams with lower ISm importance. Thus, for each audio stream n, n = 0,..., N-1, the bit-budget allocator 106 uses the following bit-rate adaptation algorithm:

[0148] 1. In frames classified as class ISm = ISM_NO_META, a constant low bit-rate is assigned.

[0149] 2. In frames classified as class ISm = ISM_LOW_IMP, the total bit-rate total_brate is reduced, e.g. to:

[0150]

[0151] where the constant is set to a value lower than 1.0, e.g. 0.6. Then, the constant represents the minimum bit rate threshold supported by a particular configured codec, which can depend on, for example, the internal sampling rate of the codec, the coded audio bandwidth, etc. (see reference [1] for more details on these values).

[0152] 3. In frames classified as class ISm = ISM_MEDIUM_IMP: the core encoder total bit rate total_brate is reduced, for example to:

[0153]

[0154] where the constant is set to a value lower than 1.0 but higher than , for example 0.8.

[0155] 4. In frames classified as class ISm = ISM_HIGH_IMP: no bit rate adaptation is used;

[0156] 5. Finally, the saved bit budget (sum of the difference between the old (total_brate) total bit rate and the new (total_brate new ) total bit rate) is re-distributed evenly among the audio streams that have active content in the frame. The same bit budget re-distribution logic as described in sub-operations 2 and 3 of section 2.4.1 can be used.

[0157] Figure 5 is a graph illustrating an example of bit rate adaptation based on ISm importance logic. From top to bottom, Figure 5 the graph illustrates, over time:

[0158] - active speech segments of the audio stream of audio object #1 ;

[0159] - active speech segments of the audio stream of audio object #2;

[0160] - total bit rate total _ brate of the audio stream of audio object #1 without using the bit rate adaptation algorithm;

[0161] - total bit rate total _ brate of the audio stream of audio object #2 without using the bit rate adaptation algorithm;

[0162] - total bit rate total _ brate of the audio stream of audio object #1 when using the bit rate adaptation algorithm; and

[0163] - total bit rate total _ brate of the audio stream of audio object #2 when using the bit rate adaptation algorithm.

[0164] In Figure 5 the non-limiting example of two audio objects (N = 2) and a fixed and constant total bitrate ism_total_brate equal to 48 kbps, the core encoder total bitrate total_brate in the active frames of audio object #1 fluctuates between 23.45 kbps and 23.65 kbps when the bitrate adaptation algorithm is not used, while it fluctuates between 19.15 kbps and 28.05 kbps when the bitrate adaptation algorithm is used. Similarly, the core encoder total bitrate total_brate in the active frames of audio object #2 fluctuates between 23.40 kbps and 23.65 kbps when the bitrate adaptation algorithm is not used, while it fluctuates between 19.10 kbps and 28.05 kbps when the bitrate adaptation algorithm is used. A better, more efficient distribution of the available bit-budget between the audio streams is thus obtained.

[0165] 2.5 Pre-processing

[0166] With reference Figure 1 to Figure 1, the method 150 for coding object-based audio signals comprises a pre-processing 158 of the N audio streams transmitted from the configuration and decision processor 106 (bit-budget allocator) through the N transmission channels 104. To perform the pre-processing 158, the system 100 for coding object-based audio signals comprises a pre-processor 108.

[0167] Once the configuration and bitrate distribution between the N audio streams is completed by the configuration and decision processor 106 (bit-budget allocator), the pre-processor 108 performs a further pre-processing 158 on each of the N audio streams. This pre-processing 158 can comprise, for example, a further signal classification, a further core encoder selection (e.g. a selection between ACELP core, TCX core and HQ core), a different internal sampling frequency F s other resamplings, etc. Examples of such pre-processing can be found, for example, in the reference [1] related to the EVS codec, and will therefore not be further described in the present disclosure.

[0168] 2.6 Core encoding

[0169] With reference Figure 1The method 150 for coding object-based audio signals comprises an operation 159 of core encoding. To perform operation 159, the system 100 for coding object-based audio signals comprises the above-mentioned N audio stream encoders, including for example the number N of core encoders 109 to respectively encode the N audio streams conveyed from the pre-processor 108 through the N transport channels 104.

[0170] In particular, the N audio streams are encoded using the N fluctuating bitrate core encoders 109, for example single-core encoders. The bitrate used by each of the N core encoders is the bitrate selected for the corresponding audio stream by the configuration and decision processor 106 (bit-budget allocator). For example, the core encoders described in reference [1] can be used as core encoders 109.

[0171] 3.0 Bitstream structure

[0172] With reference to Figure 1 The method 150 for coding object-based audio signals comprises an operation 160 of multiplexing. To perform operation 160, the system 100 for coding object-based audio signals comprises the multiplexer 110.

[0173] Figure 6 is a schematic diagram showing the structure of the bitstream 111 generated by the multiplexer 110 and sent from the Figure 1 coding system 100 to the Figure 7 decoding system 700 for one frame. The structure of the bitstream 111 can be structured as shown in Figure 6 regardless of whether the metadata is present and sent or not.

[0174] With reference to Figure 6 The multiplexer 110 writes the indices of the N audio streams from the beginning of the bitstream 111, while the indices from the ISm common signaling 113 from the configuration and decision processor 106 (bit-budget allocator) and the metadata 112 from the metadata processor 105 are written from the end of the bitstream 111.

[0175] 3.1 ISm common signaling

[0176] The multiplexer writes the ISm common signaling 113 from the end of the bitstream 111. The ISm common signaling is generated by the configuration and decision processor 106 (bit-budget allocator) and comprises a variable number of bits representing:

[0177] (a) the number N of audio objects: the signaling that the number N of coded audio objects is present in the bitstream 111 is in the form of, for example, a unary code with a stop bit (for example, for N = 3 audio objects, the first 3 bits of the ISm common signaling will be "110").

[0178] (b) Metadata presence flag flag meta : When using signal activity based bitrate adaptation as described in section 2.4.1, the flag flag meta is present and each audio object includes one bit to indicate whether the metadata for that particular audio object is present in the bitstream 111 (flag meta = 1) or not (flag meta = 0), or (c) ISm importance class: When using ISm importance based bitrate adaptation as described in section 2.4.2, the signaling is present and each audio object includes two bits to indicate the ISm importance class class ISm (ISM_NO_META, ISM_LOW_IMP, ISM_MEDIUM_IMP and ISM_HIGH_IMP) as defined in section 2.4.2.

[0179] (d) ISm VAD flag flag VAD : The ISm VAD flag is sent when flag meta = 0, in addition to class ISm = ISM_NO_META, and distinguishes between the following two cases:

[0180] 1) The input metadata is not present or the metadata is not coded, so the audio stream needs to be coded by the active coding mode (flag VAD = 1); and

[0181] 2) The input metadata is present and sent, so the audio stream can be coded by the inactive coding mode (flag VAD = 0).

[0182] 3.2 Coded metadata payload

[0183] The multiplexer 110 is provided with the coded metadata 112 from the metadata processor 105 and writes the metadata payload from the end of the bitstream in the order for the audio objects in the current frame for which metadata is coded (flag meta = 1, in addition to class ISm ≠ ISM_NO_META). The metadata bit budget per audio object is not constant, but is object- and frame-adaptive. Different metadata format scenarios are shown as Figure 2 .

[0184] In case that for at least some of the N audio objects, the metadata is not present or not transmitted, for these audio objects, the metadata flag is set to 0, i.e. flag meta = 0, and in addition class ISm = ISM_NO_META. Then, no metadata index related to those audio objects (i.e. ) is transmitted.

[0185] 3.3 Audio stream payload

[0186] The multiplexer 110 receives the N audio streams 114 coded by the N core encoders 109 through the N transport channels 104 and writes the audio stream payloads in time order for the N audio streams from the beginning of the bitstream 111 (see Figure 6 ). Due to the bit rate adaptation algorithm described in section 2.4, the bit budget of the N audio streams is fluctuating individually.

[0187] 4.0 Decoding of audio objects

[0188] Figure 7 is a schematic block diagram simultaneously showing a decoding system 700 for decoding audio objects in response to audio streams having associated metadata and a corresponding method 750 for decoding audio objects.

[0189] 4.1 Demultiplexing

[0190] With reference to Figure 7 , the method 750 for decoding audio objects in response to audio streams having associated metadata comprises an operation 755 of demultiplexing. In order to perform the operation 755, the decoding system 700 for decoding audio objects in response to audio streams having associated metadata comprises a demultiplexer 705.

[0191] The demultiplexer receives the bitstream 701 transmitted from the coding system 100 of Figure 1 to the decoding system 700 of Figure 7 . In particular, the bitstream 701 of Figure 7 corresponds to the bitstream 111 of Figure 1 .

[0192] The demultiplexer 110 extracts from the bitstream 701 (a) the N encoded audio streams 114, (b) the encoded metadata 112 of the N audio objects, and (c) the ISm common signaling 113 read from the end of the received bitstream 701.

[0193] 4.2 Metadata decoding and dequantization

[0194] With reference to Figure 7The method 750 for decoding audio objects in response to audio streams with associated metadata comprises an operation 756 of metadata decoding and dequantization. To perform operation 756, the decoding system 700 for decoding audio objects in response to audio streams and associated metadata comprises a metadata decoding and dequantization processor 706.

[0195] The metadata decoding and dequantization processor 706 is provided with the transmitted encoded metadata 112, ISm common signaling 113 and output settings 709 of the audio objects to decode and dequantize the metadata of the audio streams / objects with active content. The output settings 709 are command line parameters on the decoded audio objects / transport channels and / or the number M of audio formats, which can be equal or different from the number N of encoded audio objects / transport channels. The metadata decoding and dequantization processor 706 produces decoded metadata 704 of M audio objects / transport channels and provides information on the respective bit budgets of the M decoded metadata on line 708. Obviously, the decoding and dequantization performed by processor 706 are the inverse processes of the quantization and coding performed by the metadata processor 105 of Figure 1

[0196] 4.3 Configuration and decision on bitrates

[0197] With reference to Figure 7 The method 750 for decoding audio objects in response to audio streams with associated metadata comprises an operation 757 of configuration and decision on bitrates per channel. To perform operation 757, the decoding system 700 for decoding audio objects in response to audio streams and associated metadata comprises a configuration and decision processor 707 (bit-budget allocator).

[0198] The bit-budget allocator 707 receives from the common signaling 113 (a) information on the respective bit budgets of the M decoded metadata on line 708 and (b) the ISm importance classes class ISm and determines the core decoder bitrate per audio stream total_brate[n]. The bit-budget allocator 707 uses the same procedure as in the bit-budget allocator 106 of Figure 1 (see section 2.4) to determine the core decoder bitrate.

[0199] 4.4 Core decoding

[0200] With reference to Figure 7 ​The method 750 for decoding audio objects in response to audio streams with associated metadata comprises the operation 760 of core decoding. To perform the operation 760, the decoding system 700 for decoding audio objects in response to audio streams with associated metadata comprises N audio stream 114 decoders, including N core decoders 710, e.g. N fluctuating bitrate core decoders.

[0201] The N audio streams 114 from the demultiplexer 705 are decoded, e.g. sequentially in the number N of fluctuating bitrate core decoders 710 at their respective core decoder bitrates determined by the bit budget allocator 707. When the number M of decoded audio objects requested by the output settings 709 is lower than the number of transport channels, i.e. M < N, a lower number of core decoders is used. Similarly, not all metadata payloads can be decoded in this case.

[0202] In response to the N audio streams 114 from the demultiplexer 705, the core decoder bitrates determined by the bit budget allocator 707, and the output settings 709, the core decoders 710 produce M decoded audio streams 703 on the respective M transport channels.

[0203] 5.0 Audio Channel Rendering

[0204] In the operation of audio channel rendering 761, the renderer 711 of audio objects converts the M decoded metadata 704 and the M decoded audio streams 703 into a plurality of output audio channels 702, while taking into account the output settings 712 indicating the number and content of output audio channels to be produced. Again, the number of output audio channels 702 can be equal to or different from the number M.

[0205] The renderer 711 can be designed in various different structures to obtain the desired output audio channels. To this end, the renderer will not be described further in this disclosure.

[0206] 6.0 Source Code

[0207] According to non-limiting illustrative embodiments, the system and method for coding object-based audio signals as disclosed in the foregoing description can be implemented by the following source code (expressed in C code) given below as additional disclosure.

[0208] void ism_metadata_enc(

[0209] const long ism_total_brate, / * i : ISm total bitrate * /

[0210] const short n_ISms, / * i : number of objects* /

[0211] ISM_METADATA_HANDLE hIsmMeta[], / * i / o: ISM metadata handle* /

[0212] ENC_HANDLE hSCE[], / * i / o: element encoder handle* /

[0213] BSTR_ENC_HANDLE hBstr, / * i / o: bitstream handle* /

[0214] short nb_bits_metadata[], / * o : number of metadata bits* /

[0215] short localVAD[] )

[0217] {

[0218] short i, ch, nb_bits_start, diff;

[0219] short idx_azimuth, idx_azimuth_abs, flag_abs_azimuth[MAX_NUM_OBJECTS], nbits_diff_azimuth;

[0220] short idx_elevation, idx_elevation_abs, flag_abs_elevation[MAX_NUM_OBJECTS], nbits_diff_elevation;

[0221] float valQ;

[0222] ISM_METADATA_HANDLE hIsmMetaData;

[0223] long element_brate[MAX_NUM_OBJECTS], total_brate[MAX_NUM_OBJECTS];

[0224] short ism_metadata_flag_global;

[0225] short ism_imp[MAX_NUM_OBJECTS];

[0226] / * Initialization * /

[0227] ism_metadata_flag_global = 0;

[0228] set_s( nb_bits_metadata, 0, n_ISms );

[0229] set_s( flag_abs_azimuth, 0, n_ISms );

[0230] set_s( flag_abs_elevation, 0, n_ISms );

[0231] / *----------------------------------------------------------------*

[0232] * Set metadata presence / significance flags

[0233] *----------------------------------------------------------------* /

[0234] for( ch = 0; ch < n_ISms; ch++ )

[0235] {

[0236] if( hIsmMeta[ch]->ism_metadata_flag )

[0237] {

[0238] hIsmMeta[ch]->ism_metadata_flag = localVAD[ch];

[0239] }

[0240] else

[0241] {

[0242] hIsmMeta[ch]->ism_metadata_flag = 0;

[0243] }

[0244] if (hSCE[ch]->hCoreCoder[0]->tcxonly )

[0245] {

[0246] / * Send metadata at highest bitrate (TCX core only) in every frame * /

[0247] hIsmMeta[ch]->ism_metadata_flag = 1;

[0248] }

[0249] }

[0250] rate_ism_importance( n_ISms, hIsmMeta, hSCE, ism_imp );

[0251] / *----------------------------------------------------------------*

[0252] * Write ISM common signaling

[0253] *----------------------------------------------------------------* /

[0254] / * Write number of objects - unary coding * /

[0255] for( ch = 1; ch < n_ISms; ch++ )

[0256] {

[0257] push_indice( hBstr, IND_ISM_NUM_OBJECTS, 1, 1 );

[0258] }

[0259] push_indice( hBstr, IND_ISM_NUM_OBJECTS, 0, 1 );

[0260] / * Write ISM metadata flag (one per object) * /

[0261] for( ch = 0; ch < n_ISms; ch++ )

[0262] {

[0263] push_indice( hBstr, IND_ISM_METADATA_FLAG, ism_imp[ch], ISM_METADATA_FLAG_BITS );

[0264] ism_metadata_flag_global |= hIsmMeta[ch]->ism_metadata_flag;

[0265] }

[0266] / * Write VAD flags * /

[0267] for( ch = 0; ch < n_ISms; ch++ )

[0268] {

[0269] if( hIsmMeta[ch]->ism_metadata_flag == 0 )

[0270] {

[0271] push_indice( hBstr, IND_ISM_VAD_FLAG, localVAD[ch], VAD_FLAG_BITS );

[0272] }

[0273] }

[0274] if( ism_metadata_flag_global )

[0275] {

[0276] / *----------------------------------------------------------------*

[0277] * Metadata quantization and coding, loop over all objects

[0278] *----------------------------------------------------------------* /

[0279] for( ch = 0; ch < n_ISms; ch++ )

[0280] {

[0281] hIsmMetaData = hIsmMeta[ch];

[0282] nb_bits_start = hBstr->nb_bits_tot;

[0283] if( hIsmMeta[ch]->ism_metadata_flag )

[0284] {

[0285] / *----------------------------------------------------------------*

[0286] * Azimuth quantization and coding

[0287] *----------------------------------------------------------------* /

[0288] / * Azimuth quantization * /

[0289] idx_azimuth_abs = usquant( hIsmMetaData->azimuth, &valQ, ISM_AZIMUTH_MIN, ISM_AZIMUTH_DELTA, (1 << ISM_AZIMUTH_NBITS) );

[0290] idx_azimuth = idx_azimuth_abs;

[0291] nbits_diff_azimuth = 0;

[0292] flag_abs_azimuth[ch] = 0; / * Default to differential coding * /

[0293] if( hIsmMetaData->azimuth_diff_cnt == ISM_FEC_MAX / * Use differential coding over ISM_FEC_MAX consecutive frames with the maximum value in order to control the decoding in FEC * /

[0294] || hIsmMetaData->last_ism_metadata_flag == 0 / * Do not use differential coding if the last frame has not been coded with metadata * / )

[0296] {

[0297] flag_abs_azimuth[ch] = 1;

[0298] }

[0299] / * Try differential coding * /

[0300] if( flag_abs_azimuth[ch] == 0 )

[0301] {

[0302] diff = idx_azimuth_abs - hIsmMetaData->last_azimuth_idx;

[0303] if( diff == 0 )

[0304] {

[0305] idx_azimuth = 0;

[0306] nbits_diff_azimuth = 1;

[0307] }

[0308] else if( ABSVAL( diff ) < ISM_MAX_AZIMUTH_DIFF_IDX ) / * When diff bits >= abs bits, prefer abs * /

[0309] {

[0310] idx_azimuth = 1 << 1;

[0311] nbits_diff_azimuth = 1;

[0312] if( diff < 0 )

[0313] {

[0314] idx_azimuth += 1; / * Negative sign * /

[0315] diff *= -1;

[0316] }

[0317] else

[0318] {

[0319] idx_azimuth += 0; / * positive sign * /

[0320] }

[0321] idx_azimuth = idx_azimuth « diff;

[0322] nbits_diff_azimuth++;

[0323] / * unary coding of "diff"

[0324] idx_azimuth += ((1 « diff) - 1);

[0325] nbits_diff_azimuth += diff;

[0326] if( nbits_diff_azimuth < ISM_AZIMUTH_NBITS - 1 )

[0327] {

[0328] / * add stop bit - only for codewords shorter than ISM_AZIMUTH_NBITS

[0329] idx_azimuth = idx_azimuth « 1;

[0330] nbits_diff_azimuth++;

[0331] }

[0332] }

[0333] else

[0334] {

[0335] flag_abs_azimuth[ch] = 1;

[0336] }

[0337] }

[0338] / * update counters

[0339] if( flag_abs_azimuth[ch] == 0 )

[0340] {

[0341] hIsmMetaData->azimuth_diff_cnt++;

[0342] hIsmMetaData->elevation_diff_cnt = min( hIsmMetaData->elevation_diff_cnt, ISM_FEC_MAX );

[0343] }

[0344] else

[0345] {

[0346] hIsmMetaData->azimuth_diff_cnt = 0;

[0347] }

[0348] / * Write Azimuth * /

[0349] push_indice( hBstr, IND_ISM_AZIMUTH_DIFF_FLAG, flag_abs_azimuth[ch],1 );

[0350] if( flag_abs_azimuth[ch] )

[0351] {

[0352] push_indice( hBstr, IND_ISM_AZIMUTH, idx_azimuth, ISM_AZIMUTH_NBITS );

[0353] }

[0354] else

[0355] {

[0356] push_indice( hBstr, IND_ISM_AZIMUTH, idx_azimuth, nbits_diff_azimuth );

[0357] }

[0358] / *----------------------------------------------------------------*

[0359] * Elevation quantization and encoding

[0360] *----------------------------------------------------------------* /

[0361] / * Azimuth quantization * /

[0362] idx_azimuth_abs = usquant( hIsmMetaData->azimuth, &valQ, ISM_AZIMUTH_MIN, ISM_AZIMUTH_DELTA, (1 << ISM_AZIMUTH_NBITS) );

[0363] idx_azimuth = idx_azimuth_abs;

[0364] nbits_diff_azimuth = 0;

[0365] flag_abs_azimuth[ch] = 0; / * Default: differential coding * /

[0366] if( hIsmMetaData->azimuth_diff_cnt == ISM_FEC_MAX / * Differential coding with maximum number of consecutive frames (to control decoding in FEC) * /

[0367] || hIsmMetaData->last_ism_metadata_flag == 0 / * Do not use differential coding if the last frame was not coded with metadata * / )

[0369] {

[0370] flag_abs_azimuth[ch] = 1;

[0371] }

[0372] / * Note: Azimuth is coded from the second frame only (no sense in init_frame) * /

[0373] if( hSCE[0]->hCoreCoder[0]->ini_frame == 0 )

[0374] {

[0375] flag_abs_azimuth[ch] = 1;

[0376] hIsmMetaData->last_elevation_idx = idx_elevation_abs;

[0377] }

[0378] diff = idx_elevation_abs - hIsmMetaData->last_elevation_idx;

[0379] / * Avoid absolute encoding of elevation if absolute encoding is already used for azimuth* /

[0380] if( flag_abs_azimuth[ch] == 1 )

[0381] {

[0382] flag_abs_elevation[ch] = 0;

[0383] if( diff >= 0 )

[0384] {

[0385] diff = min( diff, ISM_MAX_ELEVATION_DIFF_IDX );

[0386] }

[0387] else

[0388] {

[0389] diff = -1 * min( -diff, ISM_MAX_ELEVATION_DIFF_IDX );

[0390] }

[0391] }

[0392] / * Try differential encoding and decoding* /

[0393] if( flag_abs_elevation[ch] == 0 )

[0394] {

[0395] if( diff == 0 )

[0396] {

[0397] idx_elevation = 0;

[0398] nbits_diff_elevation = 1;

[0399] }

[0400] else if( ABSVAL( diff ) < ISM_MAX_ELEVATION_DIFF_IDX ) / * 当diff 比特> = abs比特时,优选abs * /

[0401] {

[0402] idx_elevation = 1 << 1;

[0403] nbits_diff_elevation = 1;

[0404] if( diff < 0 )

[0405] {

[0406] idx_elevation += 1; / * negative sign * /

[0407] diff *= -1;

[0408] }

[0409] else

[0410] {

[0411] idx_elevation += 0; / * positive sign * /

[0412] }

[0413] idx_elevation = idx_elevation << diff;

[0414] nbits_diff_elevation++;

[0415] / * “diff的一元编解码 * /

[0416] idx_elevation += ((1 << diff) - 1);

[0417] nbits_diff_elevation += diff;

[0418] if( nbits_diff_elevation < ISM_ELEVATION_NBITS - 1 )

[0419] {

[0420] / * add stop bit * /

[0421] idx_elevation = idx_elevation << 1;

[0422] nbits_diff_elevation++;

[0423] }

[0424] }

[0425] else

[0426] {

[0427] flag_abs_elevation[ch] = 1;

[0428] }

[0429] }

[0430] / * update counter * /

[0431] if( flag_abs_elevation[ch] == 0 )

[0432] {

[0433] hIsmMetaData->elevation_diff_cnt++;

[0434] hIsmMetaData->elevation_diff_cnt = min( hIsmMetaData->elevation_diff_cnt, ISM_FEC_MAX );

[0435] }

[0436] else

[0437] {

[0438] hIsmMetaData->elevation_diff_cnt = 0;

[0439] }

[0440] / * write elevation * /

[0441] if( flag_abs_azimuth[ch] == 0 ) / * If "flag_abs_azimuth == 1", do not write "Flag_abs_elevation" * / / * VE: 0-> 1 for VAD, TBV * /

[0442] {

[0443] push_indice( hBstr, IND_ISM_ELEVATION_DIFF_FLAG, flag_abs_elevation[ch], 1 );

[0444] }

[0445] if( flag_abs_elevation[ch] )

[0446] {

[0447] push_indice( hBstr, IND_ISM_ELEVATION, idx_elevation, ISM_ELEVATION_NBITS );

[0448] }

[0449] else

[0450] {

[0451] push_indice( hBstr, IND_ISM_ELEVATION, idx_elevation, nbits_diff_elevation );

[0452] }

[0453] / *----------------------------------------------------------------*

[0454] * Update

[0455] *----------------------------------------------------------------* /

[0456] hIsmMetaData->last_azimuth_idx = idx_azimuth_abs;

[0457] hIsmMetaData->last_elevation_idx = idx_elevation_abs;

[0458] / * Save the number of metadata bits written * /

[0459] nb_bits_metadata[ch] = hBstr->nb_bits_tot - nb_bits_start;

[0460] }

[0461] }

[0462] / *----------------------------------------------------------------*

[0463] * Minimize the use of several absolute coding between objects

[0464] * Indexes in the same frame

[0465] *----------------------------------------------------------------* /

[0466] i = 0;

[0467] while( i == 0 || i < n_ISms / INTER_OBJECT_PARAM_CHECK )

[0468] {

[0469] short num, abs_num, abs_first, abs_next, pos_zero;

[0470] short abs_matrice[INTER_OBJECT_PARAM_CHECK * 2];

[0471] num = min( INTER_OBJECT_PARAM_CHECK, n_ISms - i * INTER_OBJECT_PARAM_CHECK );

[0472] i++;

[0473] set_s( abs_matrice, 0, INTER_OBJECT_PARAM_CHECK * ISM_NUM_PARAM );

[0474] for( ch = 0; ch < num; ch++ )

[0475] {

[0476] if( flag_abs_azimuth[ch] == 1 )

[0477] {

[0478] abs_matrice[ch*ISM_NUM_PARAM] = 1;

[0479] }

[0480] if( flag_abs_elevation[ch] == 1 )

[0481] {

[0482] abs_matrice[ch*ISM_NUM_PARAM + 1] = 1;

[0483] }

[0484] }

[0485] abs_num = sum_s( abs_matrice, INTER_OBJECT_PARAM_CHECK * ISM_NUM_PARAM );

[0486] abs_first = 0;

[0487] while( abs_num > 1 )

[0488] {

[0489] / * Find the first "1" entry * /

[0490] while( abs_matrice[abs_first] == ​​0 )

[0491] {

[0492] abs_first++;

[0493] }

[0494] / * Find the next "1" entry * /

[0495] abs_next = abs_first + 1;

[0496] while( abs_matrice[abs_next] == 0 )

[0497] {

[0498] abs_next++;

[0499] }

[0500] / * Find the "0" position * /

[0501] pos_zero = 0;

[0502] while( abs_matrice[pos_zero] == 1 )

[0503] {

[0504] pos_zero++;

[0505] }

[0506] ch = abs_next / ISM_NUM_PARAM;

[0507] if( abs_next % ISM_NUM_PARAM == 0 )

[0508] {

[0509] hIsmMeta[ch]->azimuth_diff_cnt = abs_num - 1;

[0510] }

[0511] if( abs_next % ISM_NUM_PARAM == 1 )

[0512] {

[0513] hIsmMeta[ch]->elevation_diff_cnt = abs_num - 1;

[0514] / *hIsmMeta[ch]->elevation_diff_cnt = min( hIsmMeta[ch]->elevation_diff_cnt, ISM_FEC_MAX );* /

[0515] }

[0516] abs_first++;

[0517] abs_num--;

[0518] }

[0519] }

[0520] }

[0521] / *----------------------------------------------------------------*

[0522] * Configuration and determination of the bit rate per channel

[0523] *----------------------------------------------------------------* /

[0524] ism_config( ism_total_brate, n_ISms, hIsmMeta, localVAD, ism_imp,element_brate, total_brate, nb_bits_metadata );

[0525] for( ch = 0; ch < n_ISms; ch++ )

[0526] {

[0527] hIsmMeta[ch]->last_ism_metadata_flag = hIsmMeta[ch]->ism_metadata_flag;

[0528] hSCE[ch]->hCoreCoder[0]->low_rate_mode = 0;

[0529] if (hIsmMeta[ch]->ism_metadata_flag == 0 && localVAD[ch][0] == 0 &&ism_metadata_flag_global )

[0530] {

[0531] hSCE[ch]->hCoreCoder[0]->low_rate_mode = 1;

[0532] }

[0533] hSCE[ch]->element_brate = element_brate[ch];

[0534] hSCE[ch]->hCoreCoder[0]->total_brate = total_brate[ch];

[0535] / * Write metadata only in active frames * /

[0536] if( hSCE[0]->hCoreCoder[0]->core_brate > SID_2k40 )

[0537] {

[0538] reset_indices_enc( hSCE[ch]->hMetaData, MAX_BITS_METADATA );

[0539] }

[0540] }

[0541] return;

[0542] }

[0543] void rate_ism_importance(

[0544] const short n_ISms, / * i : number of objects* /

[0545] ISM_METADATA_HANDLE hIsmMeta[], / * i / o: ISM metadata handle* /

[0546] ENC_HANDLE hSCE[], / * i / o: element encoder handle* /

[0547] short ism_imp[] / * o : ISM importance flag* / )

[0549] {

[0550] short ch, ctype;

[0551] for( ch = 0; ch < n_ISms; ch++ )

[0552] {

[0553] ctype = hSCE[ch]->hCoreCoder[0]->coder_type_raw;

[0554] if( hIsmMeta[ch]->ism_metadata_flag == 0 )

[0555] {

[0556] ism_imp[ch] = ISM_NO_META;

[0557] }

[0558] else if( ctype == INACTIVE || ctype == UNVOICED )

[0559] {

[0560] ism_imp[ch] = ISM_LOW_IMP;

[0561] }

[0562] else if( ctype == VOICED )

[0563] {

[0564] ism_imp[ch] = ISM_MEDIUM_IMP;

[0565] }

[0566] else / * GENERIC * /

[0567] {

[0568] ism_imp[ch] = ISM_HIGH_IMP;

[0569] }

[0570] }

[0571] return;

[0572] }

[0573] void ism_config(

[0574] const long ism_total_brate, / * i : ISm总比特率 * /

[0575] const short n_ISms, / * i : 对象的数量 * /

[0576] ISM_METADATA_HANDLE hIsmMeta[], / * i / o: ISM metadata handles * /

[0577] short localVAD[],

[0578] const short ism_imp[], / * i : ISM importance flags * /

[0579] long element_brate[], / * o : element bitrate per object * /

[0580] long total_brate[], / * o : total bitrate per object * /

[0581] short nb_bits_metadata[] / * i / o: metadata bit number * / )

[0583] {

[0584] short ch;

[0585] short bits_element[MAX_NUM_OBJECTS], bits_CoreCoder[MAX_NUM_OBJECTS];

[0586] short bits_ism, bits_side;

[0587] long tmpL;

[0588] short ism_metadata_flag_global;

[0589] / * Initialization * /

[0590] ism_metadata_flag_global = 0;

[0591] bits_side = 0;

[0592] if( hIsmMeta!= NULL )

[0593] {

[0594] for( ch = 0; ch < n_ISms; ch++ )

[0595] {

[0596] ism_metadata_flag_global |= hIsmMeta[ch]->ism_metadata_flag;

[0597] }

[0598] }

[0599] / * About per channel bitrate - decision of constant during session (in one ism_total_brate) * /

[0600] bits_ism = ism_total_brate / FRMS_PER_SECOND;

[0601] set_s( bits_element, bits_ism / n_ISms, n_ISms );

[0602] bits_element[n_ISms - 1] += bits_ism % n_ISms;

[0603] bitbudget_to_brate( bits_element, element_brate, n_ISms );

[0604] / * Count ISM common signaling bits * /

[0605] if( hIsmMeta!= NULL )

[0606] {

[0607] nb_bits_metadata[0] += n_ISms * ISM_METADATA_FLAG_BITS + n_ISms;

[0608] for( ch = 0; ch < n_ISms; ch++ )

[0609] {

[0610] if( hIsmMeta[ch]->ism_metadata_flag == 0 )

[0611] {

[0612] nb_bits_metadata[0] += ISM_METADATA_VAD_FLAG_BITS;

[0613] }

[0614] }

[0615] }

[0616] / * Split the metadata bit budget evenly among the channels * /

[0617] if( nb_bits_metadata!= NULL )

[0618] {

[0619] bits_side = sum_s( nb_bits_metadata, n_ISms );

[0620] set_s( nb_bits_metadata, bits_side / n_ISms, n_ISms );

[0621] nb_bits_metadata[n_ISms - 1] += bits_side % n_ISms;

[0622] v_sub_s( bits_element, nb_bits_metadata, bits_CoreCoder, n_ISms );

[0623] bitbudget_to_brate( bits_CoreCoder, total_brate, n_ISms );

[0624] mvs2s( nb_bits_metadata, nb_bits_metadata, n_ISms );

[0625] }

[0626] / * Assign less core coder bit budget to inactive streams (at least one stream must be active) * /

[0627] if( ism_metadata_flag_global )

[0628] {

[0629] long diff;

[0630] short n_higher, flag_higher[MAX_NUM_OBJECTS];

[0631] set_s( flag_higher, 1, MAX_NUM_OBJECTS );

[0632] diff = 0;

[0633] for( ch = 0; ch < n_ISms; ch++ )

[0634] {

[0635] if( hIsmMeta[ch]->ism_metadata_flag == 0 && localVAD[ch] == 0 )

[0636] {

[0637] diff += bits_CoreCoder[ch] - BITS_ISM_INACTIVE;

[0638] bits_CoreCoder[ch] = BITS_ISM_INACTIVE;

[0639] flag_higher[ch] = 0;

[0640] }

[0641] }

[0642] n_higher = sum_s( flag_higher, n_ISms );

[0643] if( diff > 0 && n_higher > 0 )

[0644] {

[0645] tmpL = diff / n_higher;

[0646] for( ch = 0; ch < n_ISms; ch++ )

[0647] {

[0648] if( flag_higher[ch] )

[0649] {

[0650] bits_CoreCoder[ch] += tmpL;

[0651] }

[0652] }

[0653] tmpL = diff % n_higher;

[0654] ch = 0;

[0655] while( flag_higher[ch] == 0 )

[0656] {

[0657] ch++;

[0658] }

[0659] bits_CoreCoder[ch] += tmpL;

[0660] }

[0661] bitbudget_to_brate( bits_CoreCoder, total_brate, n_ISms );

[0662] diff = 0;

[0663] for( ch = 0; ch < n_ISms; ch++ )

[0664] {

[0665] long limit;

[0666] limit = MIN_BRATE_SWB_BWE / FRMS_PER_SECOND;

[0667] if( element_brate[ch] < MIN_BRATE_SWB_STEREO )

[0668] {

[0669] limit = MIN_BRATE_WB_BWE / FRMS_PER_SECOND;

[0670] }

[0671] else if( element_brate[ch] >= SCE_CORE_16k_LOW_LIMIT )

[0672] {

[0673] / *limit = SCE_CORE_16k_LOW_LIMIT;* /

[0674] limit = (ACELP_16k_LOW_LIMIT + SWB_TBE_1k6) / FRMS_PER_SECOND;

[0675] }

[0676] if( ism_imp[ch] == ISM_NO_META && localVAD[ch] == 0 )

[0677] {

[0678] tmpL = BITS_ISM_INACTIVE;

[0679] }

[0680] else if( ism_imp[ch] == ISM_LOW_IMP )

[0681] {

[0682] tmpL = BETA_ISM_LOW_IMP * bits_CoreCoder[ch];

[0683] tmpL = max( limit, bits_CoreCoder[ch] - tmpL );

[0684] }

[0685] else if( ism_imp[ch] == ISM_MEDIUM_IMP )

[0686] {

[0687] tmpL = BETA_ISM_MEDIUM_IMP * bits_CoreCoder[ch];

[0688] tmpL = max( limit, bits_CoreCoder[ch] - tmpL );

[0689] }

[0690] else / * ism_imp[ch] == ISM_HIGH_IMP * /

[0691] {

[0692] tmpL = bits_CoreCoder[ch];

[0693] }

[0694] diff += bits_CoreCoder[ch] - tmpL;

[0695] bits_CoreCoder[ch] = tmpL;

[0696] }

[0697] if( diff > 0 && n_higher > 0 )

[0698] {

[0699] tmpL = diff / n_higher;

[0700] for( ch = 0; ch < n_ISms; ch++ )

[0701] {

[0702] if( flag_higher[ch] )

[0703] {

[0704] bits_CoreCoder[ch] += tmpL;

[0705] }

[0706] }

[0707] tmpL = diff % n_higher;

[0708] ch = 0;

[0709] while( flag_higher[ch] == 0 )

[0710] {

[0711] ch++;

[0712] }

[0713] bits_CoreCoder[ch] += tmpL;

[0714] }

[0715] / * Verify maximum bitrate @ 12.8KHz core * /

[0716] diff = 0;

[0717] for ( ch = 0; ch < n_ISms; ch++ )

[0718] {

[0719] limit_high = STEREO_512k / FRMS_PER_SECOND;

[0720] if ( element_brate[ch] < SCE_CORE_16k_LOW_LIMIT ) / * copy function set_ACELP_flag() -> it is not intended to switch ACELP internal sampling rate within the object

[0721] {

[0722] limit_high = ACELP_12k8_HIGH_LIMIT / FRMS_PER_SECOND;

[0723] }

[0724] tmpL = min( bits_CoreCoder[ch], limit_high );

[0725] diff += bits_CoreCoder[ch] - tmpL;

[0726] bits_CoreCoder[ch] = tmpL;

[0727] }

[0728] if ( diff > 0 )

[0729] {

[0730] ch = 0;

[0731] for ( ch = 0; ch < n_ISms; ch++ )

[0732] {

[0733] if ( flag_higher[ch] == 0 )

[0734] {

[0735] if ( diff > limit_high )

[0736] {

[0737] diff += bits_CoreCoder[ch] - limit_high;

[0738] bits_CoreCoder[ch] = limit_high;

[0739] }

[0740] else

[0741] {

[0742] bits_CoreCoder[ch] += diff;

[0743] break;

[0744] }

[0745] }

[0746] }

[0747] }

[0748] bitbudget_to_brate( bits_CoreCoder, total_brate, n_ISms );

[0749] }

[0750] return;

[0751] }

[0752] 7.0 Hardware Implementation

[0753] Figure 8 is a simplified block diagram of an example configuration of hardware components that form the above-described encoding and decoding systems and methods.

[0754] Each codec and decoding system can be implemented as part of a mobile terminal, part of a portable media player or any similar device. Each codec and decoding system (in Figure 8 1200 ) includes an input 1202 , an output 1204 , a processor 1206 , and a memory 1208 .

[0755] Input 1202 is configured to receive input signal(s), for example, in digital or analog form. Figure 1 N input audio objects 102 (N audio streams and corresponding N metadata) or Figure 7 The bit stream 701. The output 1204 is configured to provide (multiple) output signals, for example, Figure 1 The bit stream 111 or Figure 7 The M decoded audio streams 703 and the M decoded metadata 704. The input 1202 and the output 1204 may be implemented in a common module, for example, a serial input / output device.

[0756] The processor 1206 is operatively connected to the input 1202, the output 1204, and the memory 1208. The processor 1206 is implemented as one or more processors for executing code instructions to support Figure 1 and Figure 7the various processors and other modules of the system.

[0757] Memory 1208 can include non-transitory memory for storing code instructions executable by processor(s) 1206, and in particular, processor-readable memory including non-transitory instructions that, when executed, cause the processor(s) to implement the operations and processes of the encoding and decoding systems and methods described in this disclosure. Memory 1208 can also include random access memory or buffer(s) to store intermediate processing data from the various functions performed by processor(s) 1206.

[0758] Those of ordinary skill in the art will realize that the description of the encoding and decoding systems and methods is merely illustrative and is not intended to be limiting in any way. Other embodiments will readily occur to those of ordinary skill in the art upon reading the present disclosure. Furthermore, the disclosed encoding and decoding systems and methods can be customized to provide valuable solutions to existing needs and problems in encoding and decoding sound.

[0759] In the interest of clarity, not all of the routine features of the embodiments of the encoding and decoding systems and methods are shown and described. It will, of course, be appreciated that in the development of any such actual implementation of the encoding and decoding systems and methods, numerous implementation-specific decisions can be made to achieve the developer's specific goals, such as compliance with application-related, system-related, business-related, and

[0760] In accordance with the present disclosure, the processor / modules, processing operations, and / or data structures described herein can be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general purpose machines. In addition, those of ordinary skill in the art will recognize that devices of a non-generic nature, such as hardwired devices, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc., can also be used. Where a method comprising a series of operations and sub-operations is implemented by a processor, computer, or machine, these operations and sub-operations can be stored as a series of non-transitory code instructions readable by the processor, computer, or machine, which can be stored on a tangible and / or non-transitory medium.

[0761] The encoding and decoding systems and methods described herein can use software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.

[0762] In the coding and decoding systems and methods described herein, various operations and sub-operations can be performed in various orders, and some operations and sub-operations can be optional.

[0763] Although the present disclosure has been described above with reference to non-limiting illustrative embodiments, these embodiments are not intended to limit the scope of the disclosure, but rather are intended to be exemplary embodiments of the disclosure. Various modifications to these embodiments can be made by those skilled in the art without departing from the spirit and scope of the disclosure.

[0764] 8.0 References

[0765] The following references are cited in the present disclosure, and are incorporated by reference herein in their entirety.

[0766] [1] 3GPP Specification TS 26.445: "Codec for Enhanced Voice Services (EVS). Detailed Algorithmic Description", v12.0.0, September 2014.

[0767] [2] V. Ekslade "Method and Device for Allocating a Bit-budget Between Sub-frames in a CELP Codec", PCT Patent Application PCT / CA2018 / 51175.

[0768] 9.0 Other Embodiments

[0769] The following embodiments (Embodiments 1 to 83) are part of the present disclosure in relation to the present invention.

[0770] Embodiment 1. A system for coding an object-based audio signal comprising audio objects in response to an audio stream having associated metadata, comprising:

[0771] an audio stream processor for analyzing the audio stream; and

[0772] a metadata processor for encoding metadata for the input audio stream in response to information about the audio stream from the analysis by the audio stream processor.

[0773] Embodiment 2. The system of embodiment 1, wherein the metadata processor outputs information about a metadata bit-budget for the audio objects, and wherein the system further comprises a bit-budget allocator for allocating a bit-rate to the audio stream in response to the information about the metadata bit-budget for the audio objects from the metadata processor.

[0774] Embodiment 3. The system of Embodiments 1 or 2, comprising an encoder of an audio stream including transcoded metadata.

[0775] Embodiment 4. The system of any of Embodiments 1 to 3, wherein the encoder comprises a plurality of core encoders using a bit rate allocated to the audio stream by a bit budget allocator.

[0776] Embodiment 5. The system of any of Embodiments 1 to 4, wherein the object-based audio signal comprises at least one of speech, music, and general audio sounds.

[0777] Embodiment 6. The system of any of Embodiments 1 to 5, wherein the object-based audio signal represents or encodes a complex audio auditory scene as a cluster of individual elements of the audio objects.

[0778] Embodiment 7. The system of any of Embodiments 1 to 6, wherein each audio object comprises an audio stream with associated metadata.

[0779] Embodiment 8. The system of any of Embodiments 1 to 7, wherein the audio stream is an independent stream with metadata.

[0780] Embodiment 9. The system of any of Embodiments 1 to 8, wherein the audio stream represents an audio waveform and typically comprises one or two channels.

[0781] Embodiment 10. The system of any of Embodiments 1 to 9, wherein the metadata is a set of information describing the audio stream and the artistic intent for translating the original or transcoded audio objects to a final reproduction system.

[0782] Embodiment 11. The system of any of Embodiments 1 to 10, wherein the metadata typically describes spatial properties of each audio object.

[0783] Embodiment 12. The system of any of Embodiments 1 to 11, wherein the spatial properties comprise one or more of a position, a direction, a volume, a width of the audio object.

[0784] Embodiment 13. The system of any of Embodiments 1 to 12, wherein each audio object comprises a set of metadata referred to as input metadata defined as an unquantized metadata representation used as an input to the codec.

[0785] Embodiment 14. The system of any of Embodiments 1 to 13, wherein each audio object comprises a set of metadata referred to as transcoded metadata defined as a quantized and transcoded metadata that is part of a bitstream sent from the encoder to the decoder.

[0786] Embodiment 15. The system according to any one of embodiments 1 to 14, wherein the rendering system is configured to render the audio objects in a 3D audio space around the listener using the transmitted metadata and artistic intent on the rendering side.

[0787] Embodiment 16. The system according to any one of embodiments 1 to 15, wherein the rendering system comprises a head tracking device for dynamically modifying the metadata during rendering of the audio objects.

[0788] Embodiment 17. The system according to any one of embodiments 1 to 16, comprising a framework for simultaneous coding of a number of audio objects.

[0789] Embodiment 18. The system according to any one of embodiments 1 to 17, wherein the simultaneous coding of a number of audio objects uses a fixed constant overall bit rate to encode the audio objects.

[0790] Embodiment 19. The system according to any one of embodiments 1 to 18, comprising a transmitter for transmitting some or all of the audio objects.

[0791] Embodiment 20. The system according to any one of embodiments 1 to 19, wherein in case of coding a combination of audio formats in the framework, the constant overall bit rate represents the sum of the bit rates of the formats.

[0792] Embodiment 21. The system according to any one of embodiments 1 to 20, wherein the metadata comprises two parameters, including an azimuth angle and an elevation angle.

[0793] Embodiment 22. The system according to any one of embodiments 1 to 21, wherein the azimuth and elevation angle parameters are stored per audio frame for each audio object.

[0794] Embodiment 23. The system according to any one of embodiments 1 to 22, comprising an input buffer for buffering at least one input audio stream and input metadata associated with the audio stream.

[0795] Embodiment 24. The system according to any one of embodiments 1 to 23, wherein the input buffer buffers one frame per audio stream.

[0796] Embodiment 25. The system according to any one of embodiments 1 to 24, wherein the audio stream processor analyzes and processes the audio stream.

[0797] Embodiment 26. The system according to any one of Embodiments 1 to 25, wherein the audio stream processor comprises at least one of the following elements: a time domain transient detector, a spectral analyzer, a long term prediction analyzer, a pitch tracker and voicing analyzer, a voice / sound activity detector, a bandwidth detector, a noise estimator, and a signal classifier.

[0798] Embodiment 27. The system according to any one of Embodiments 1 to 26, wherein the signal classifier performs at least one of codec type selection, signal classification, and speech / music classification.

[0799] Embodiment 28. The system according to any one of Embodiments 1 to 27, wherein the metadata processor analyzes, quantizes, and encodes metadata of the audio stream.

[0800] Embodiment 29. The system according to any one of Embodiments 1 to 28, wherein in inactive frames, no metadata is encoded by the metadata processor and transmitted by the system in the bitstream of the corresponding audio object.

[0801] Embodiment 30. The system according to any one of Embodiments 1 to 29, wherein in active frames, metadata is encoded by the metadata processor using a variable bit rate for the corresponding object.

[0802] Embodiment 31. The system according to any one of Embodiments 1 to 30, wherein the bit budget allocator sums the bit budgets of the metadata of the audio objects and adds the sum of the bit budgets to the signaling bit budget in order to allocate a bit rate to the audio stream.

[0803] Embodiment 32. The system according to any one of Embodiments 1 to 31, comprising a pre-processor for further processing the audio streams when the configuration and bit rate distribution between the audio streams have been completed.

[0804] Embodiment 33. The system according to any one of Embodiments 1 to 32, wherein the pre-processor performs at least one of a further classification of the audio streams, a core encoder selection, and a resampling.

[0805] Embodiment 34. The system according to any one of Embodiments 1 to 33, wherein the encoders sequentially encode the audio streams.

[0806] Embodiment 35. The system according to any one of Embodiments 1 to 34, wherein the encoders sequentially encode the audio streams using a plurality of fluctuating bit rate core encoders.

[0807] Embodiment 36. The device according to any one of Embodiments 1 to 35, wherein the metadata processor sequentially encodes the metadata in a loop according to a correlation between the quantization of the audio objects and the metadata parameters of the audio objects.

[0808] Embodiment 37. The system according to any one of embodiments 1 to 36, wherein to encode the metadata parameters, the metadata processor uses a quantization step to quantize the metadata parameter indices.

[0809] Embodiment 38. The system according to any one of embodiments 1 to 37, wherein to encode the azimuth parameter, the metadata processor uses a quantization step size to quantize the azimuth index, and to encode the elevation parameter, the metadata processor uses a quantization step size to quantize the elevation index.

[0810] Embodiment 39. The device according to any one of embodiments 1 to 38, wherein the total metadata bit budget and the number of quantized bits depend on the codec total bit rate, the metadata total bit rate, or the sum of the metadata bit budget related to one audio object and the core encoder bit budget.

[0811] Embodiment 40. The system according to any one of embodiments 1 to 39, wherein the azimuth and elevation parameters are represented as one parameter.

[0812] Embodiment 41. The system according to any one of embodiments 1 to 40, wherein the metadata processor encodes the metadata parameter indices absolutely or differentially.

[0813] Embodiment 42. The system according to any one of embodiments 1 to 41, wherein when there is a difference between the current parameter index and the previous parameter index that results in a number of bits required for differential coding being higher than or equal to the number of bits required for absolute coding, the metadata processor encodes the metadata parameter index using absolute coding.

[0814] Embodiment 43. The system according to any one of embodiments 1 to 42, wherein when there is no metadata in the previous frame, the metadata processor encodes the metadata parameter index using absolute coding.

[0815] Embodiment 44. The system according to any one of embodiments 1 to 43, wherein when the number of consecutive frames using differential coding is higher than the maximum number of consecutive frames coded using differential coding, the metadata processor encodes the metadata parameter index using absolute coding.

[0816] Embodiment 45. The system according to any one of embodiments 1 to 44, wherein when the metadata parameter index is encoded using absolute coding, the metadata processor writes an absolute coding flag after the metadata parameter absolute coding index, the absolute coding flag distinguishing between absolute coding and differential coding.

[0817] Embodiment 46. The system according to any of Embodiments 1 to 45, wherein when using differential coding to encode the metadata parameter index, the metadata processor sets an absolute coding flag to 0 and writes a zero coding flag after the absolute coding flag, signaling if the difference between the current frame index and the previous frame index is 0.

[0818] Embodiment 47. The system according to any of Embodiments 1 to 46, wherein if the difference between the current frame index and the previous frame index is not equal to 0, the metadata processor continues coding by writing a sign flag and subsequently writing a self-adaptive bit difference index.

[0819] Embodiment 48. The system according to any of Embodiments 1 to 47, wherein the metadata processor uses an intra-object metadata coding logic to limit the range of metadata bit budget fluctuations between frames and avoid too low bit budget left for core coding.

[0820] Embodiment 49. The system according to any of Embodiments 1 to 48, wherein the metadata processor limits, according to an intra-object metadata coding logic, the use of absolute coding in a given frame to only one metadata parameter or as few metadata parameters as possible.

[0821] Embodiment 50. The system according to any of Embodiments 1 to 49, wherein the metadata processor avoids, according to an intra-object metadata coding logic, absolute coding of an index of one metadata parameter if an index of another metadata coding logic has already been coded using absolute coding in the same frame.

[0822] Embodiment 51. The system according to any of Embodiments 1 to 50, wherein the intra-object metadata coding logic is bitrate dependent.

[0823] Embodiment 52. The system according to any of Embodiments 1 to 51, wherein the metadata processor uses an inter-object metadata coding logic used between metadata coding of different objects to minimize the number of absolute coding metadata parameters for different audio objects in a current frame.

[0824] Embodiment 53. The system according to any of Embodiments 1 to 52, wherein the metadata processor uses the inter-object metadata coding logic to control a frame counter of absolute coding metadata parameters.

[0825] Embodiment 54. The system according to any of the Embodiments 1 to 53, wherein the metadata processor uses inter-object metadata coding logic that, when metadata parameters of audio objects evolve slowly and smoothly, (a) encodes a first metadata parameter index of a first audio object using absolute coding in frame M, (b) encodes a second metadata parameter index of the first audio object using absolute coding in frame M+1, (c) encodes a first metadata parameter index of a second audio object using absolute coding in frame M+2, and (d) encodes a second metadata parameter index of the second audio object using absolute coding in frame M+3.

[0826] Embodiment 55. The system according to any of the Embodiments 1 to 54, wherein the inter-object metadata coding logic is bitrate-dependent.

[0827] Embodiment 56. The system according to any of the Embodiments 1 to 55, wherein the bitrate adaptation algorithm used by the bit-budget allocator distributes a bit-budget for encoding the audio stream.

[0828] Embodiment 57. The system according to any of the Embodiments 1 to 56, wherein the bitrate adaptation algorithm used by the bit-budget allocator obtains a metadata total bit-budget from a metadata total bitrate or a codec total bitrate.

[0829] Embodiment 58. The system according to any of the Embodiments 1 to 57, wherein the bitrate adaptation algorithm used by the bit-budget allocator computes an element bit-budget by dividing the metadata total bit-budget by a number of audio streams.

[0830] Embodiment 59. The system according to any of the Embodiments 1 to 58, wherein the bitrate adaptation algorithm used by the bit-budget allocator adjusts the element bit-budget of a last audio stream to spend all available metadata bit-budget.

[0831] Embodiment 60. The system according to any of the Embodiments 1 to 59, wherein the bitrate adaptation algorithm used by the bit-budget allocator sums metadata bit-budgets of all audio objects and adds the sum to a metadata common signaling bit-budget, resulting in a core encoder side bit-budget.

[0832] Embodiment 61. The system according to any of the Embodiments 1 to 60, wherein the bitrate adaptation algorithm used by the bit-budget allocator (a) splits the core encoder side bit-budget evenly among audio objects and (b) uses the split core encoder side bit-budget and the element bit-budget to compute a core encoder bit-budget for each audio stream.

[0833] Example 62. The system of any of Examples 1-61, wherein the bit-budget allocator uses a bitrate adaptation algorithm to adjust the core encoder bit-budget of the last audio stream to spend all of the available core encoder bit-budget.

[0834] Example 63. The system of any of Examples 1-62, wherein the bit-budget allocator uses a bitrate adaptation algorithm to calculate a bitrate for encoding one audio stream in the core encoder using the core encoder bit-budget.

[0835] Example 64. The system of any of Examples 1-63, wherein the bit-budget allocator uses a bitrate adaptation algorithm to reduce the bitrate for encoding one audio stream in the core encoder and set it to a constant value in inactive frames or low-energy frames, and redistribute the saved bit-budget among the audio streams in active frames.

[0836] Example 65. The system of any of Examples 1-64, wherein the bit-budget allocator uses a bitrate adaptation algorithm to adjust the bitrate for encoding one audio stream in the core encoder based on the metadata importance class in active frames.

[0837] Example 66. The system of any of Examples 1-65, wherein the bit-budget allocator reduces the bitrate for encoding one audio stream in the core encoder in inactive frames (VAD = 0), and redistributes the bit-budget saved by the bitrate reduction among the audio streams in frames that are classified as active.

[0838] Example 67. The system of any of Examples 1-66, wherein the bit-budget allocator (a) sets a lower, constant core encoder bit-budget for each audio stream with inactive content in a frame, (b) calculates a saved bit-budget as the difference between the lower constant core encoder bit-budget and the core encoder bit-budget, and (c) redistributes the saved bit-budget among the core encoder bit-budgets of the audio streams in active frames.

[0839] Example 68. The system of any of Examples 1-67, wherein the lower, constant bit-budget is dependent on the metadata total bitrate.

[0840] Example 69. The system of any of Examples 1-68, wherein the bit-budget allocator uses the lower, constant core encoder bit-budget to calculate a bitrate for encoding one audio stream in the core encoder.

[0841] Embodiment 70. The system according to any one of Embodiments 1 to 69, wherein the bit-budget allocator uses inter-object core encoder bitrate adaptation based on classification of metadata importance.

[0842] Embodiment 71. The system according to any one of Embodiments 1 to 70, wherein the metadata importance is based on a metric indicating how critical the codec of a particular audio object at the current frame is for obtaining good quality of the decoded synthesis.

[0843] Embodiment 72. The system according to any one of Embodiments 1 to 71, wherein the bit-budget allocator makes the classification of metadata importance based on at least one of the following parameters: coder type (coder_type), FEC signal class, speech / music classification decision, and SNR estimate (snr_celp, snr_tcx) from the open-loop ACELP / TCX core decision module.

[0844] Embodiment 73. The system according to any one of Embodiments 1 to 72, wherein the bit-budget allocator makes the classification of metadata importance based on the coder type (coder_type).

[0845] Embodiment 74. The system according to any one of Embodiments 1 to 73, wherein the bit-budget allocator defines the following four different classes of metadata importance (class ISm ):

[0846] - No metadata class, ISM_NO_META: frames without metadata codec, e.g. in inactive frames with VAD = 0

[0847] - Low importance class, ISM_LOW_IMP: frames with coder_type = UNVOICED or INACTIVE

[0848] - Medium importance class, ISM_MEDIUM_IMP: frames with coder_type = VOICED

[0849] - High importance class, ISM_HIGH_IMP: frames with coder_type = GENERIC.

[0850] Embodiment 75. The system according to any one of Embodiments 1 to 74, wherein the bit-budget allocator uses the class of metadata importance in the bitrate adaptation algorithm to assign higher bit-budget to audio streams with higher importance and lower bit-budget to audio streams with lower importance.

[0851] Embodiment 76. The system according to any one of Embodiments 1 to 75, wherein the bit-budget allocator uses the following logic in a frame:

[0852] 1. class ISm = ISM_NO_META frame: lower constant core encoder bitrate is assigned;

[0853] 2. class ISm = ISM_LOW_IMP frame: the bitrate (total_brate) of one audio stream encoded in the core encoder is reduced to

[0854]

[0855] where the constant is set to a value lower than 1.0 and the constant is the minimum bitrate threshold supported by the core encoder;

[0856] 3. class ISm = ISM_MEDIUM_IMP frame: the bitrate (total_brate) of one audio stream encoded in the core encoder is reduced to

[0857]

[0858] where the constant is set to a value lower than 1.0 but higher than the value ;

[0859] 4. class ISm = ISM_HIGH_IMP frame: no bitrate adaptation is used.

[0860] Embodiment 77. The system of any of embodiments 1 to 76, wherein the bit budget allocator redistributes the saved bit budget represented as the sum of the difference between the previous bitrate and the new bitrate total_brate among the audio streams classified as active.

[0861] Embodiment 78. A system for decoding audio objects in response to audio streams having associated metadata, comprising:

[0862] a metadata processor for decoding metadata of the audio streams having active content;

[0863] a bit budget allocator determining core encoder bitrates of the audio streams in response to the decoded metadata and respective bit budgets of the audio objects; and

[0864] a decoder of the audio streams using the core encoder bitrates determined in the bit budget allocator.

[0865] Embodiment 79. The system of embodiment 78, wherein the metadata processor is responsive to metadata common signaling read from an end of the received bitstream.

[0866] Embodiment 80. The system of embodiment 78 or 79, wherein the decoder comprises core decoders that decode the audio streams.

[0867] Embodiment 81. The system of any of embodiments 78 to 80, wherein the core decoders comprise fluctuating bit rate core decoders that sequentially decode the audio streams at their respective core encoder bit rates.

[0868] Embodiment 82. The system of any of embodiments 78 to 81, wherein the number of decoded audio objects is lower than the number of core decoders.

[0869] Embodiment 83. The system of any of embodiments 78 to 83, comprising a renderer responsive to the decoded audio streams and the audio objects of the decoded metadata.

[0870] Any of embodiments 2 to 77 further describing elements of embodiments 78 to 83 can be implemented in any of these embodiments 78 to 83. As an example, the core encoder bit rate per audio stream in the decoding system is determined using the same procedure as in the codec system.

[0871] The invention also relates to a codec method and a decoding method. In this regard, the system embodiments 1 to 83 can be drafted as method embodiments, wherein the elements of the system embodiments are replaced by the operations performed by such elements.

Claims

1. A system for encoding and decoding an object-based audio signal comprising audio objects in response to an audio stream having associated metadata, comprising: a metadata processor for encoding and decoding the metadata, the metadata processor generating information about a bit budget for encoding and decoding the metadata of the audio object; An encoder, configured to encode and decode the audio stream; and a bit budget allocator that allocates a bit rate for encoding and decoding the audio stream by the encoder in response to information from the metadata processor about a bit budget for encoding and decoding metadata of the audio object, The metadata processor encodes and decodes the metadata before and separately from encoding and decoding the audio stream, and after encoding and decoding the metadata, generates information about a bit budget used by the metadata processor to encode and decode the metadata of the audio object.

2. The system of claim 1, comprising an audio stream processor for analyzing the audio stream and providing information about the audio stream to the metadata processor and the bit budget allocator.

3. The system of claim 2, wherein the audio stream processors analyze the audio streams in parallel.

4. The system according to any one of claims 1 to 3, wherein: The bit budget allocator uses a bit rate adaptation algorithm to distribute the available bit budget for encoding and decoding the audio stream.

5. The system according to claim 4, wherein: The bit budget allocator calculates an audio stream and metadata (ISm) total bit budget from an ISm total bit rate or a codec total bit rate used to encode and decode the audio stream and the associated metadata using the bit rate adaptation algorithm.

6. The system according to claim 5, wherein: The bit budget allocator calculates an element bit budget by dividing the ISm total bit budget by the number of the audio streams using the bit rate adaptation algorithm.

7. The system according to claim 6, wherein: The bit budget allocator uses the bit rate adaptation algorithm to adjust the element bit budget of the last audio object to spend all of the ISm total bit budget.

8. The system according to claim 6, wherein: The element bit budget is constant over an ISm total bit budget.

9. The system according to claim 6, wherein: The bit budget allocator uses the bit rate adaptation algorithm to sum the bit budgets used for encoding and decoding the metadata of the audio object, and adds the sum to the ISm common signaling bit budget to generate a codec side bit budget.

10. The system according to claim 9, wherein: The bit budget allocator uses the bit rate adaptation algorithm to (a) evenly divide the codec-side bit budget among the audio objects, and (b) calculate an encoding bit budget for each audio stream using the divided codec-side bit budget and the element bit budget.

11. The system according to claim 10, wherein: The bit budget allocator uses the bit rate adaptation algorithm to adjust the encoding bit budget of the last audio stream to spend all available encoding bit budget.

12. The system according to claim 10, wherein: The bit budget allocator uses the bit rate adaptation algorithm to calculate a bit rate for encoding and decoding one of the audio streams using the encoding bit budgets of the audio streams.

13. The system according to claim 4, wherein: The bit budget allocator uses the bitrate adaptation algorithm for audio streams with inactive content or without meaningful content, reduces the value of the bitrate used to encode and decode one of the audio streams, and redistributes the saved bit budget among audio streams with active content.

14. The system according to claim 4, wherein: The bit budget allocator uses the bitrate adaptation algorithm for audio streams with active content to adjust a bitrate for encoding and decoding one of the audio streams based on audio stream and metadata (ISm) importance classification.

15. The system according to claim 13, wherein: The bit budget allocator uses the bit rate adaptation algorithm for an audio stream having inactive content or no meaningful content, reduces a bit budget for encoding and decoding the audio stream, and sets the bit budget to a constant value.

16. The system of claim 13, wherein: The bit budget allocator calculates the saved bit budget as a difference between a lower value of the bit budget for encoding and decoding the audio stream and a non-lower value of the bit budget for encoding and decoding the audio stream.

17. The system according to claim 15, wherein: The bit budget allocator calculates a bit rate for encoding and decoding the audio stream using a lower value of the bit budget.

18. The system of claim 14, wherein: The bit budget allocator classifies the ISm importance based on a metric indicating how critical it is to encode and decode an audio object to obtain a given quality of decoding synthesis.

19. The system of claim 14, wherein: The bit budget allocator classifies the ISm importance based on one or more of the following parameters: audio stream encoder type, FEC (forward error correction), voice signal classification, speech / music classification, and SNR (signal-to-noise ratio) estimation.

20. The system of claim 19, wherein: The bit budget allocator classifies the ISm importance based on the audio stream encoder type (coder_type).

21. The system of claim 20, wherein: The bit budget allocator defines the following ISm importance classes: ISm ): - No metadata class, ISM_NO_META: frames without metadata codec; - Low importance class, ISM_LOW_IMP: frames with coder_type = UNVOICED or INACTIVE; - Medium importance class, ISM_MEDIUM_IMP: frames with coder_type = VOICED; and - High importance class, ISM_HIGH_IMP: frames with coder_type = GENERIC.

22. The system of claim 14, wherein: The bit budget allocator uses the ISm importance classification in the bit rate adaptation algorithm to increase the bit budget for encoding and decoding audio streams with higher ISm importance and decrease the bit budget for encoding and decoding audio streams with lower ISm importance.

23. The system of claim 21, wherein: For each audio stream in a frame, the bit budget allocator uses the following logic: 1.class ISm = ISM_NO_META frame: assigns a constant low bit rate for encoding and decoding the audio stream; 2.class ISm = ISM _ LOW _ IMP or class ISm = ISM_MEDIUM_IMP frame: reduces the bit rate used to encode and decode the audio stream using the given relationship; and 3.class ISm = ISM_HIGH_IMP frame: bit rate adaptation is not used.

24. The system of claim 23, wherein: The bit budget allocator redistributes the saved bit budget among the audio streams with active content in the frame.

25. The system according to any one of claims 1 to 3, comprising a pre-processor for further processing the audio streams once the bit budget allocator has completed the bit rate distribution among the audio streams.

26. The system of claim 25, wherein: The pre-processor performs at least one of further classification of the audio stream, core encoder selection, and resampling.

27. The system according to any one of claims 1 to 3, wherein: The encoder of the audio stream includes a plurality of core encoders for encoding and decoding the audio stream.

28. The system of claim 27, wherein: The core encoder is a fluctuating bit rate core encoder that sequentially encodes and decodes the audio stream.

29. A method for encoding and decoding an object-based audio signal comprising audio objects in response to an audio stream having associated metadata, comprising: encoding and decoding the metadata; generating information about a bit budget for encoding and decoding metadata of the audio object; encoding the audio stream; and allocating a bit rate for encoding the audio stream in response to information about a bit budget for encoding and decoding metadata of the audio object, The metadata is encoded and decoded before and separately from encoding and decoding of the audio stream, and information about a bit budget for encoding and decoding the metadata of the audio object is generated after encoding and decoding the metadata.

30. The method of claim 29, comprising analyzing the audio stream and providing information about the audio stream for encoding and decoding the metadata and information allocating a bit rate for encoding and decoding the audio stream.

31. The method according to claim 30, wherein The audio streams are analyzed in parallel.

32. The method according to any one of claims 29 to 31, wherein Allocating a bit rate for encoding and decoding the audio stream includes using a bit rate adaptation algorithm to distribute an available bit budget for encoding and decoding the audio stream.

33. The method according to claim 32, wherein Allocating a bitrate for encoding and decoding the audio stream using the bitrate adaptation algorithm includes calculating an audio stream and metadata (ISm) total bitrate or a codec total bitrate for encoding and decoding the audio stream and the associated metadata.

34. The method according to claim 33, wherein Allocating a bit rate for encoding and decoding the audio stream using the bit rate adaptation algorithm includes calculating an element bit budget by dividing the ISm total bit budget by the number of the audio streams.

35. The method according to claim 34, wherein Allocating a bit rate for encoding and decoding the audio stream using the bit rate adaptation algorithm includes adjusting the element bit budget of the last audio object to spend all of the ISm total bit budget.

36. The method of claim 34, wherein: The element bit budget is constant over an ISm total bit budget.

37. The method of claim 34, wherein: Allocating a bit rate for encoding and decoding the audio stream using the bit rate adaptation algorithm includes summing the bit budgets used for encoding and decoding metadata of the audio objects and adding the sum to the ISm common signaling bit budget to generate a codec side bit budget.

38. The method of claim 37, wherein: Allocating a bit rate for encoding and decoding the audio stream using the bit rate adaptation algorithm includes (a) dividing the codec-side bit budget evenly among the audio objects, and (b) calculating an encoding bit budget for each audio stream using the divided codec-side bit budget and the element bit budget.

39. The method according to claim 38, wherein Allocating a bit rate for encoding and decoding the audio streams using the bit rate adaptation algorithm includes adjusting the encoding bit budget of the last audio stream to spend all available encoding bit budget.

40. The method of claim 38, wherein Allocating a bit rate for encoding and decoding the audio streams using the bit rate adaptation algorithm includes using encoding bit budgets of the audio streams to calculate a bit rate for encoding and decoding one of the audio streams.

41. The method of claim 32, wherein: Allocating bitrates for encoding and decoding audio streams using the bitrate adaptation algorithm for audio streams with inactive content or without meaningful content includes reducing the value of the bitrate used for encoding and decoding one of the audio streams and redistributing the saved bit budget among audio streams with active content.

42. The method of claim 32, wherein: Allocating bitrates for encoding and decoding audio streams using the bitrate adaptation algorithm for audio streams having active content includes adjusting a bitrate for encoding and decoding one of the audio streams based on audio stream and metadata (ISm) importance classification.

43. The method according to claim 41, wherein Allocating a bit rate for encoding and decoding an audio stream having inactive content or no meaningful content using the bit rate adaptation algorithm includes reducing a bit budget for encoding and decoding the audio stream and setting the bit budget to a constant value.

44. The method of claim 41, wherein Allocating a bit rate for encoding and decoding the audio stream includes calculating a saved bit budget as a difference between a lower value of the bit budget for encoding and decoding the audio stream and a non-lower value of the bit budget for encoding and decoding the audio stream.

45. The method of claim 43, wherein Allocating a bit rate for encoding and decoding the audio stream includes using a lower value of the bit budget to calculate a bit rate for encoding and decoding the audio stream.

46. ​​The method of claim 42, wherein Allocating a bit rate for encoding and decoding the audio stream includes classifying the ISm importance based on a metric indicating how critical a decoding synthesis is to encoding and decoding an audio object to obtain a given quality.

47. The method of claim 42, wherein Allocating a bit rate for encoding and decoding the audio stream includes classifying the ISm importance based on one or more of the following parameters: audio stream encoder type, FEC (forward error correction), sound signal classification, speech / music classification, and SNR (signal-to-noise ratio) estimation.

48. The method of claim 47, wherein Allocating a bit rate for encoding and decoding the audio stream includes classifying the ISm importance based on the audio stream encoder type (coder_type).

49. The method according to claim 48, wherein Classifying ISm importance includes defining the following ISm importance classes: ISm ): - No metadata class, ISM_NO_META: frames without metadata codec; - Low importance class, ISM_LOW_IMP: frames with coder_type = UNVOICED or INACTIVE; - Medium importance class, ISM_MEDIUM_IMP: frames with coder_type = VOICED; and - High importance class, ISM_HIGH_IMP: frames with coder_type = GENERIC.

50. The method of claim 42, wherein Allocating bit rates for encoding and decoding the audio streams includes using the ISm importance classification in the bitrate adaptation algorithm to increase the bit budget for encoding and decoding audio streams with higher ISm importance and reduce the bit budget for encoding and decoding audio streams with lower ISm importance.

51. The method of claim 49, wherein Allocating bitrates for encoding and decoding the audio streams involves using the following logic for each audio stream in a frame: 1.class ISm = ISM_NO_META frame: assigns a constant low bit rate for encoding and decoding the audio stream; 2.class ISm = ISM _ LOW _ IMP or class ISm = ISM_MEDIUM_IMP frame: reduces the bit rate used to encode and decode the audio stream using the given relationship; and 3.class ISm = ISM_HIGH_IMP frame: bit rate adaptation is not used.

52. The method of claim 51, wherein Allocating a bit rate for encoding and decoding the audio stream includes redistributing saved bit budget among audio streams with active content in the frame.

53. A method according to any one of claims 29 to 31 , comprising pre-processing the audio streams once bit rate distribution between the audio streams is completed by allocating bit rates for encoding and decoding the audio streams.

54. The method according to claim 53, wherein The pre-processing includes performing at least one of further classification, core encoder selection, and resampling of the audio stream.

55. The method according to any one of claims 29 to 31, wherein Encoding the audio stream includes using a plurality of core encoders to encode and decode the audio stream.

56. The method of claim 55, wherein: The core encoder is a fluctuating bit rate core encoder that sequentially encodes and decodes the audio stream.

57. A system for decoding an audio object in response to an audio stream having associated metadata, comprising: a metadata processor for decoding metadata of the audio object and for providing information about a corresponding bit budget of the decoded metadata of the audio object; a bit budget allocator that determines a core decoder bit rate for the audio stream in response to the information about the corresponding bit budget of the decoded metadata of the audio object; and a decoder of said audio stream using a core decoder bit rate determined in said bit budget allocator, The metadata processor decodes the metadata before and separately from decoding of the audio stream, and obtains the information about the corresponding bit budget of the decoded metadata of the audio object after decoding the metadata.

58. The system of claim 57, wherein: The metadata processor is responsive to common signaling read from the received bitstream.

59. The system of claim 57 or 58, wherein: The decoder includes multiple core decoders to decode audio streams.

60. The system of claim 59, wherein: The core decoder includes a fluctuating bitrate core decoder that sequentially decodes the audio streams at their respective core decoder bit rates.

61. The system of claim 57 or 58, wherein: The bit budget allocator distributes the available bit budget for decoding the audio stream using a bit rate adaptation algorithm.

62. The system of claim 61, wherein: The bit budget allocator calculates an audio stream and metadata (ISm) total bit budget from an ISm total bit rate or a codec total bit rate for decoding the audio stream and the associated metadata using the bit rate adaptation algorithm.

63. The system of claim 62, wherein: The bit budget allocator calculates an element bit budget by dividing the ISm total bit budget by the number of the audio streams using the bit rate adaptation algorithm.

64. The system of claim 63, wherein: The bit budget allocator uses the bit rate adaptation algorithm to adjust the element bit budget of the last audio object to spend all of the ISm total bit budget.

65. The system of claim 63, wherein: The bit budget allocator uses the bit rate adaptation algorithm to sum the bit budgets used for decoding the metadata of the audio objects and adds the sum to the ISm common signaling bit budget to generate a codec side bit budget.

66. The system of claim 65, wherein: The bit budget allocator uses the bitrate adaptation algorithm to (a) evenly divide the codec-side bit budget among the audio objects, and (b) calculate a decoding bit budget for each audio stream using the divided codec-side bit budget and the element bit budget.

67. The system of claim 66, wherein: The bit budget allocator uses the bit rate adaptation algorithm to adjust the decoding bit budget of the last audio stream to spend all available decoding bit budget.

68. The system of claim 66, wherein: The bit budget allocator uses the bit rate adaptation algorithm to calculate a bit rate for decoding one of the audio streams using the decoding bit budgets of the audio streams.

69. The system of claim 61, wherein: The bit budget allocator uses the bitrate adaptation algorithm for audio streams with inactive content or without meaningful content, reduces the value of the bitrate used to decode one of the audio streams, and redistributes the saved bit budget among audio streams with active content.

70. The system of claim 61, wherein The bit budget allocator uses the bitrate adaptation algorithm for audio streams with active content to adjust a bitrate for decoding one of the audio streams based on audio stream and metadata (ISm) importance classification.

71. The system of claim 70, wherein: The bit budget allocator uses the bit rate adaptation algorithm for audio streams with inactive content or without meaningful content, reduces the bit budget for decoding the audio streams and sets the bit budget to a constant value.

72. The system of claim 69, wherein: The bit budget allocator calculates the saved bit budget as a difference between a lower value of the bit budget for decoding the audio stream and a non-lower value of the bit budget for decoding the audio stream.

73. The system of claim 71, wherein: The bit budget allocator uses a lower value of the bit budget to calculate a bit rate for decoding the audio stream.

74. The system of claim 57 or 58, wherein: The bit budget allocator uses audio stream and metadata (ISm) importance read from common signaling in the received bitstream to indicate how critical it is to decode the audio object to obtain a given quality of decoded composition.

75. The system of claim 74, wherein: The bit budget allocator defines the following ISm importance classes: ISm ): - No metadata class, ISM_NO_META: frames without metadata codec; - Low importance class, ISM_LOW_IMP: frames with audio stream decoder type (coder_type) = UNVOICED or INACTIVE; - Medium importance class, ISM_MEDIUM_IMP: frames with coder_type = VOICED; and - High importance class, ISM_HIGH_IMP: frames with coder_type = GENERIC.

76. The system of claim 70, wherein: The bit budget allocator uses the ISm importance classification in the bit rate adaptation algorithm to increase the bit budget for decoding audio streams with higher ISm importance and decrease the bit budget for decoding audio streams with lower ISm importance.

77. The system of claim 75, wherein: For each audio stream in a frame, the bit budget allocator uses the following logic: 1.class ISm = ISM_NO_META frame: assigns a constant low bit rate for decoding the audio stream; 2.class ISm = ISM _ LOW _ IMP or class ISm = ISM_MEDIUM_IMP frame: reduces the bit rate used to decode the audio stream using the given relationship; and 3.class ISm = ISM_HIGH_IMP frame: bit rate adaptation is not used.

78. The system of claim 77, wherein: The bit budget allocator redistributes the saved bit budget among the audio streams with active content in the frame.

79. A method for decoding an audio object in response to an audio stream having associated metadata, comprising: decoding metadata of the audio object and providing information about a corresponding bit budget of the decoded metadata of the audio object; determining a core decoder bitrate for the audio stream using the corresponding bit budget of the decoded metadata of the audio object; and decoding the audio stream using the determined core decoder bit rate, The metadata of the audio object is decoded prior to and separately from decoding of the audio stream, and the information about the corresponding bit budget of the decoded metadata of the audio object is generated after decoding the metadata.

80. The method of claim 79, wherein Decoding metadata of the audio object is responsive to common signaling read from the received bitstream.

81. The method of claim 79 or 80, wherein Decoding the audio stream includes using a plurality of core decoders to decode the audio stream.

82. The method of claim 81, wherein Decoding the audio streams includes sequentially decoding the audio streams at their respective core decoder bit rates using a fluctuating bitrate core decoder as a core decoder.

83. The method of claim 79 or 80, wherein Determining a core decoder bit rate for the audio stream includes using a bit rate adaptation algorithm to distribute an available bit budget for decoding the audio stream.

84. The method of claim 83, wherein Determining a core decoder bitrate for the audio stream using the bitrate adaptation algorithm includes calculating an audio stream and metadata (ISm) total bitrate or a codec total bitrate for decoding the audio stream and the associated metadata.

85. The method of claim 84, wherein Determining a core decoder bit rate for the audio stream using the bit rate adaptation algorithm includes calculating an element bit budget by dividing the ISm total bit budget by the number of audio streams.

86. The method of claim 85, wherein Determining the core decoder bit rate of the audio stream using the bit rate adaptation algorithm includes adjusting the element bit budget of the last audio object to spend all of the ISm total bit budget.

87. The method of claim 85, wherein Determining the core decoder bit rate of the audio stream using the bit rate adaptation algorithm includes summing the bit budgets of the decoded metadata of the audio objects and adding the sum to the ISm common signaling bit budget to produce a codec side bit budget.

88. The method of claim 87, wherein Determining a core decoder bit rate for the audio stream using the bitrate adaptation algorithm includes (a) dividing the codec-side bit budget equally among audio objects, and (b) calculating a decoding bit budget for each audio stream using the divided codec-side bit budget and the element bit budget.

89. The method of claim 88, wherein Determining the core decoder bit rate of the audio stream using the bit rate adaptation algorithm includes adjusting the decoding bit budget of the last audio stream to spend all available decoding bit budget.

90. The method of claim 88, wherein Determining the core decoder bitrates for the audio streams using the bitrate adaptation algorithm includes calculating a bitrate for decoding one of the audio streams using decoding bit budgets for the audio streams.

91. The method of claim 83, wherein Using the bitrate adaptation algorithm to determine the core decoder bitrate for audio streams with inactive content or without meaningful content includes reducing the value of the bitrate used to decode one of the audio streams and redistributing the saved bit budget among audio streams with active content.

92. The method of claim 83, wherein Using the bitrate adaptation algorithm for audio streams with active content to determine a core decoder bitrate for the audio streams includes adjusting a bitrate for decoding one of the audio streams based on audio stream and metadata (ISm) importance classification.

93. The method of claim 92, wherein: Using the bitrate adaptation algorithm to determine a core decoder bitrate for an audio stream having inactive content or no meaningful content includes reducing a bit budget for decoding the audio stream and setting the bit budget to a constant value.

94. The method of claim 91, wherein Determining a core decoder bit rate for the audio stream includes calculating a saved bit budget as a difference between a lower value of a bit budget for decoding the audio stream and a non-lower value of a bit budget for decoding the audio stream.

95. The method of claim 93, wherein Determining a core decoder bit rate for the audio stream includes using a lower value of the bit budget to calculate a bit rate for decoding the audio stream.

96. The method of claim 79, wherein Determining a core decoder bitrate for the audio stream includes using audio stream and metadata (ISm) importance read from common signaling in the received bitstream to indicate how critical it is to decode audio objects to obtain a given quality of decoding composition.

97. The method of claim 96, wherein Determining the core decoder bitrate of the audio stream includes defining the following ISm importance classes: ISm ): - No metadata class, ISM_NO_META: frames without metadata codec; - Low importance class, ISM_LOW_IMP: frames with audio stream decoder type (coder_type) = UNVOICED or INACTIVE; - Medium importance class, ISM_MEDIUM_IMP: frames with coder_type = VOICED; and - High importance class, ISM_HIGH_IMP: frames with coder_type = GENERIC.

98. The method of claim 92, wherein Determining a core decoder bit rate for the audio stream includes using the ISm importance classification in the bitrate adaptation algorithm to increase a bit budget for decoding audio streams with higher ISm importance and to decrease a bit budget for decoding audio streams with lower ISm importance.

99. The method of claim 97, wherein Determining the core decoder bitrate for the audio stream includes, for each audio stream in a frame, the following logic: 1.class ISm = ISM_NO_META frame: assigns a constant low bit rate for decoding the audio stream; 2.class ISm = ISM _ LOW _ IMP or class ISm = ISM_MEDIUM_IMP frame: reduces the bit rate used to decode the audio stream using the given relationship; and 3.class ISm = ISM_HIGH_IMP frame: bit rate adaptation is not used.

100. The method of claim 99, wherein Determining a core decoder bit rate for the audio stream includes redistributing saved bit budget among audio streams with active content in frames.

Citation Information

Patent Citations

  • Method and system for coding metadata in audio streams and for flexible intra-object and inter-object bitrate adaptation

    CN114097028A

  • Post-encoding bitrate reduction of multiple object audio

    US20150255076A1

  • Audio encoding device and audio decoding device

    US20160225377A1