Method and device for flexible mixed format bitrate adaptation in an audio codec

The composite format encoding and decoding method addresses bitrate inefficiencies in audio codecs by dynamically adapting bitrates for multiple audio formats, enhancing encoding efficiency and flexibility in immersive audio systems.

JP2026503560APending Publication Date: 2026-01-29VOICEAGE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025542072
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-20
Filing Date
2024-01-22
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing audio codecs struggle to efficiently adapt bitrates for combinations of audio formats, such as object-based audio and immersive audio, leading to inefficiencies in encoding and decoding processes.

Method used

A composite format encoding and decoding method that adapts bitrates dynamically for both ISM and immersive audio formats using a hybrid format bit rate adaptation device, allowing flexible allocation of bitrates based on audio object importance and scene characteristics.

Benefits of technology

Enhances encoding efficiency by dynamically adjusting bitrates, improving flexibility and adaptability in handling multiple audio formats, thereby optimizing resource utilization and quality of immersive audio experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026503560000001_ABST
    Figure 2026503560000001_ABST
Patent Text Reader

Abstract

A composite format method and encoder for encoding a first number of audio object channels and a second number of MASA audio channels in an ISM format using an ISM format encoder including an audio object channel front pre-processor for generating ISM pre-processing parameters and a core encoder section for coding the audio object channels, the core encoder section being sensitive to an adapted ISM gross bit rate. The MASA format encoder is sensitive to the adapted MASA gross bit rate for coding the MASA channels. A device for composite format bit rate adaptation uses at least one ISM pre-processing parameter to (a) adapt an initial ISM gross bit rate to result in the adapted ISM gross bit rate, and (b) adapt the initial MASA gross bit rate to result in the adapted MASA gross bit rate. A corresponding composite format decoder is also proposed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a hybrid format encoding method using hybrid format bit rate adaptation, a hybrid format encoder with a device for hybrid format bit rate adaptation, a hybrid format decoding method, and a hybrid format decoder. [Background technology]

[0002] In this disclosure and the accompanying claims: (a) The term "sound" may relate to voice, audio, and any other sound. (b) The term "multi-channel" may relate to two or more channels. (c) The term "stereo" is an abbreviation of "stereophonic." (d) The term "mono" is an abbreviation for "monophonic." (e) The term "object-based audio" is intended to represent an auditory scene as a collection of individual elements, also called audio objects, and may comprise, for example, speech, music, and any other sound, including general audio sounds. (f) The term "audio object" is intended to designate an audio stream with associated metadata. For example, in this disclosure, an "audio object" is referred to as an independent audio stream with metadata (ISM). (g) The term "audio stream" is intended to denote, in a bitstream, an audio waveform, e.g., speech, music, or any other sound, including general audio sounds, which may consist of one channel (mono), although multi-channel, including two or more channels (stereo), may also be considered. (h) The term "metadata" is intended to denote a set of information, e.g., describing an audio stream and the artistic intent used to transform original or coded audio objects into a playback system. The metadata typically describes the spatial characteristics of each individual audio object, such as its position, orientation, volume, width, etc. (i) The term "audio format" is intended to designate a technique for achieving an immersive audio experience. (j) The term "playback system" is intended to designate an element in a decoder that is capable of using the transmitted metadata and artistic intent on the playback side to render audio objects, for example, but not limited to, in a 3D (three-dimensional) audio space around the listener. Rendering may be performed to a target loudspeaker layout (e.g., 5.1 surround) or to headphones, while metadata may be dynamically modified, for example, in response to head-tracking device feedback. Other types of rendering may be contemplated.

[0003] Historically, conversational telephone communication has been carried out using mono handsets that have only one transducer to output sound to only one of the user's ears. In recent decades, users have begun to use their portable handsets with headphones, receiving sound through their two ears and primarily listening to music, but occasionally listening to voices as well. However, when a portable handset is used to send and receive conversational audio, the content is still mono but is presented to the user's two ears when headphones are used.

[0004] The 3GPP® (Third Generation Partnership Project) speech coding standard, i.e., the codec for Enhanced Voice Services (EVS) as described in Reference [1], the entire contents of which are incorporated herein by reference, has significantly improved the quality of coded sound, e.g., voice and / or audio, transmitted and received through portable handsets. The next logical step is to transmit stereo information that allows the receiver to match as closely as possible the real audio scene captured at the other end of the communication link.

[0005] Furthermore, in recent years, audio generation, recording, presentation, coding, transmission, and playback have been moving toward enhanced, interactive, and immersive experiences for listeners. An immersive experience can be described, for example, as a state of being deeply engaged and involved in a sound scene with sounds coming from all directions. In immersive audio (also called 3D (three-dimensional) audio), sound images are reproduced around the listener in all three dimensions, taking into account a wide range of sound characteristics, such as the accuracy of timbre, directionality, reverberation, transparency, and (aural) spaciousness. Immersive audio is created for a specific sound playback or reproduction system, such as a loudspeaker-based system, an integrated reproduction system (sound bar), or headphones. The interactivity of the sound reproduction system may then include, for example, the ability to adjust sound levels, change the location of the sound, or select different languages ​​for playback.

[0006] There are three basic approaches (hereafter also referred to as audio formats) to achieving an immersive audio (IA) experience:

[0007] The first IA approach is channel-based audio, in which multiple spaced microphones are used to capture sound from different directions, with one microphone corresponding to one audio channel in a particular loudspeaker layout. Each recorded channel is fed to a loudspeaker in a specific location. Examples of channel-based audio include, for example, stereo, 5.1 surround, 5.1+4, etc.

[0008] The second IA approach is scene-based audio (SBA), which represents a desired sound field over a localized space as a function of time by a combination of dimensional components. Although the signals representing scene-based audio are independent of the sound source positions, the sound field must be transformed into a chosen loudspeaker layout in a rendering playback system. An example of scene-based audio is Ambisonics.

[0009] A third immersive audio (IA) approach is object-based audio, which represents an auditory scene as a set of individual audio elements (e.g., speakers, singers, drums, guitars) accompanied by information about their position within the audio scene so that such audio elements can be rendered in their intended locations in a playback system. This gives object-based audio great flexibility and interactivity, as each object can be kept separate and manipulated individually. An example of an object coding system is described, for example, in Reference [3], the entire contents of which are incorporated herein by reference.

[0010] Beyond the three basic approaches mentioned above, new multi-channel IA coding techniques are under development, such as Metadata-Assisted Spatial Audio (MASA), as described in Reference [4], the entire contents of which are incorporated herein by reference. In the MASA approach, MASA audio channels are treated as (multi-)mono or (multi-)stereo transport signals that are coded by a core-encoder, while MASA metadata (e.g., direction, energy ratio, spread coherence, distance, surround coherence, all within several time-frequency slots) are generated, quantized, coded, and passed into the bitstream in a MASA analyzer. In a MASA decoder, the MASA metadata then guides the decoding and rendering processes to recreate the output spatial sound.

[0011] Each of the above-mentioned audio approaches to achieving an immersive experience presents pros and cons. It is thus common for several audio approaches to be combined in a complex audio system to create an immersive auditory scene, rather than just one audio approach. One example would be an audio system that combines scene-based audio (SBA) or MASA with object-based audio, e.g., SBA or MASA with several individual audio objects.

[0012] Recently, 3GPP® has started work on defining a 3D sound codec for immersive services called IVAS (Immersive Voice and Audio Services) as described in Reference [2], the entire contents of which are incorporated herein by reference, which is based on the EVS codec as described in Reference [1]. The IVAS codec specifies several audio formats in which audio scenes are captured, transmitted, decoded, and rendered: a stereo format, an object-based audio format, a multi-channel (MC) audio format, a scene-based audio (SBA) format, and a MASA format.

[0013] Beyond coding various audio formats, the IVAS codec can support audio format combinations. Advantages of such an approach include, for example, a single coding instance, smaller memory demands, better coding efficiency than encoding audio formats separately, and better control of the overall audio scene representation and playback. One approach to achieving this is to jointly capture input format combinations, for example, capturing object-based audio (e.g., voice) in combination with a spatial audio representation of an audio scene (e.g., ambience, or ambience with dominant speakers or instruments). While this disclosure considers combinations of audio objects with MASA, other combinations, such as SBA with audio objects or stereo with audio objects, may also be implemented. Similarly, combinations of three or more basic audio formats may also be implemented. Summary of the Invention [Means for solving the problem]

[0014] According to a first aspect, the present disclosure relates to a composite format method for encoding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the composite format method comprising the steps of ISM format encoding of the audio object channels, comprising front pre-processing the audio object channels to generate ISM pre-processing parameters and core-encoding the audio object channels in response to an adapted ISM total bitrate; IA format encoding of the second audio channels in response to the adapted IA total bitrate; and composite format bitrate adaptation using at least one ISM pre-processing parameter from the front pre-processing of the audio object channels to (a) adapt an initial ISM total bitrate to result in the adapted ISM total bitrate, and (b) adapt the initial IA total bitrate to result in the adapted IA total bitrate.

[0015] According to a second aspect, the present disclosure provides a composite format encoder for coding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the composite format encoder comprising: an ISM format encoder including a first front pre-processor for the audio object channels for generating ISM pre-processing parameters and a first core encoder section for coding the audio object channels that is sensitive to an adapted ISM total bit rate; an IA format encoder for coding the second audio channels that is sensitive to the adapted IA total bit rate; and a device for composite format bit rate adaptation using at least one ISM pre-processing parameter from the first front pre-processor to (a) adapt an initial ISM total bit rate to result in the adapted ISM total bit rate, and (b) adapt the initial IA total bit rate to result in the adapted IA total bit rate.

[0016] According to a third aspect, the present disclosure relates to a composite format method for decoding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the composite format method comprising the steps of receiving a bitstream, decoding from the bitstream information about a codec total bitrate, information about the number of audio object channels, audio object channel coding information, second audio channel coding information, and information about an ISM importance class for each audio object channel, and decoding the audio object channel coding information, the second audio channel coding information, and information about an ISM importance class for each audio object channel to result in an adapted ISM total bitrate and an adapted IA total bitrate. the composite format bitrate adaptation step using the number of audio object channels, the codec total bitrate, and the ISM importance class per audio object channel; core-decoding the audio object channels in response to audio object channel coding information from the bitstream, comprising configuring an ISM core-decoder in response to the adapted ISM total bitrate; and core-decoding the second audio channel in response to second audio channel coding information from the bitstream, comprising configuring core-decoding of the second audio channel in response to the adapted IA total bitrate.

[0017] According to a fourth aspect, there is provided a composite format decoder for decoding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the composite format decoder comprising: a bitstream receiver; a bitstream decoder for decoding from the bitstream information about a codec total bitrate, information about a number of audio object channels, audio object channel coding information, second audio channel coding information, and information about an ISM importance class for each audio object channel; a device for composite format bitrate adaptation using the number of audio object channels, the codec total bitrate, and the ISM importance class for each audio object channel to result in an adapted ISM total bitrate and an adapted IA total bitrate; an ISM core decoder for decoding the audio object channels in response to the audio object channel coding information from the bitstream, and a composer of the ISM core decoder responsive to the adapted ISM total bitrate; and an IA core decoder for decoding the second audio channels in response to the second audio channel coding information from the bitstream, and a composer of the IA core decoder responsive to the adapted IA total bitrate.

[0018] These and other objects, advantages and features of the hybrid format encoding method using hybrid format bit rate adaptation, the hybrid format encoder comprising a device for hybrid format bit rate adaptation, the hybrid format decoding method and the hybrid format decoder will become more apparent upon reading the following non-limiting description of exemplary embodiments thereof, given by way of example only with reference to the accompanying drawings, in which like elements are identified by like reference numerals, in the various figures of the drawings. [Brief explanation of the drawings]

[0019] [Figure 1]1 is a schematic block diagram illustrating an example of a combined mixed-format encoder and encoding method; [Figure 2] 1 is a schematic block diagram illustrating simultaneously a hybrid format encoder using a device for hybrid format bit rate adaptation according to the present disclosure and a corresponding hybrid format encoding method using hybrid format bit rate adaptation according to the present disclosure; [Figure 3] 10 is a graphical representation of an example of the impact of combined format bitrate adaptation on ism_total_brate and masa_total_brate bitrates. [Figure 4] 1 is a schematic block diagram showing simultaneously a hybrid format encoder using a device for hybrid format bit rate adaptation according to the present disclosure and a corresponding hybrid format encoding method using hybrid format bit rate adaptation according to the present disclosure, where the hybrid format bit rate adaptation depends on parameters from different format encoder parts. [Figure 5] FIG. 1 is a simplified block diagram of an exemplary configuration of hardware components forming a hybrid format encoder with a device for hybrid format bit rate adaptation, a corresponding hybrid format encoding method, a hybrid format decoding method, and a hybrid format decoder using hybrid format bit rate adaptation. DETAILED DESCRIPTION OF THE INVENTION

[0020] Mixed format bit rate adaptation in an audio codec is described by way of non-limiting example only, with reference to the IVAS coding framework, referred to throughout this disclosure as the IVAS codec (or IVAS sound codec). However, it is within the scope of this disclosure to (a) incorporate such techniques of mixed format bit rate adaptation into any other sound codec that supports a combination of at least two audio formats, as well as (b) use any immersive audio (IA) format other than the MASA format.

[0021] 1. Introduction By way of non-limiting example, this disclosure contemplates a framework that supports simultaneous coding of: (a) Several audio objects (e.g., up to four audio objects) containing audio streams along with their associated metadata (ISM format). Note that, for example, in the case of non-diegetic content, metadata is not necessarily transmitted for at least some of the audio objects. Non-diegetic sounds in movies, TV shows, and other videos are sounds that cannot be heard by the characters in the film. A soundtrack is an example of non-diegetic sound because only the audience hears the music. (b) MASA formats with their associated metadata. The MASA metadata is provided to the input of the codec as it is generated by the user device or to a MASA analyzer. A description of MASA metadata and the MASA analyzer can be found in reference [5], the entire contents of which are incorporated herein by reference.

[0022] The combined coding of audio objects in the ISM format and the MASA format is further referred to as the Complex Object MASA (OMASA) format.

[0023] Codecs such as the IVAS codec support simultaneous coding of several transport channels at a fixed total codec bit rate, ivas_total_rate. In IVAS, the total codec bit rate is constant at some value between 13.2 kbps and 512 kbps. Note that other constant values ​​of the total codec bit rate as well as the adaptive total bit rate may be considered without departing from the scope of this disclosure.

[0024] In the case of coding a combination of audio formats in the IVAS framework, e.g., OMASA, a constant total codec bitrate represents the sum of the MASA format bitrate masa_total_brate (i.e., the bitrate for encoding the MASA format part of the OMASA channel) and the ISM total bitrate ism_total_brate (i.e., the sum of the bitrates for encoding all audio objects with their metadata related to the ISM part of the OMASA channel). ivas_total_brate=masa_total_brate+ism_total_brate (1)

[0025] In a basic and simple implementation, at a given ivas_total_brate bitrate, both the masa_total_brate bitrate and the ism_total_brate bitrate are constant, and their actual values ​​can be predefined, for example, in a ROM table. For example, in a scenario with two audio objects in OMASA format and ivas_total_brate=96 kbps, the audio object channels can be coded at ism_total_brate=40 kbps, but then masa_total_brate=56 kbps. It should be pointed out that while the ism_total_brate bitrate is constant, the bitrates allocated to encoding individual audio object channels (ISM audio channels) can be variable, for example, based on the method described in reference [3]. The bit rates allocated to individual audio object channels (ISM audio channels) 1, 2, ..., N are denoted as ism_brate(n), i.e., ism_brate(1), ism_brate(2), ..., ism_brate(N), and they are

[0026]

number

[0027] where N is the number of audio objects to be coded separately.

[0028] FIG. 1 is a schematic block diagram illustrating an example of a combined mixed-format encoder 100 and an encoding method 150.

[0029] The compound format encoder 100 uses compound format encoding with constant masa_total_brate and ism_total_brate bitrates. In FIG. 1, N+2 input audio channels are considered, where N is the number of input audio object channels (audio streams with metadata), and the additional two "+2" channels correspond to input MASA audio channels (MASA audio streams with metadata). MASA typically encodes audio using one channel (mono MASA) or two channels (stereo MASA). For simplicity, this disclosure considers stereo MASA with two input audio channels as a non-limiting embodiment, but mono MASA with one input audio channel may also be implemented. Both audio objects (as described above) and MASA have their associated metadata and signaling. Also, both audio objects and MASA are processed by frames (e.g., signal segments of length 20 ms).

[0030] 1, the composite format encoder 100 comprises an OMASA configurator 101 that performs an OMASA configuration operation 151 of a composite format encoding method 150. The configurator 101 configures the OMASA composite format by setting high-level parameters, such as the number of transport channels, the OMASA mode, and / or the nominal (initial) masa_total_brate and nominal (initial) ism_total_brate bitrates, where the OMASA mode is set depending on the ivas_total_brate bitrate and the number of input audio object channels.

[0031] The combined format encoder 100 includes an OMASA analyzer and ISM / MASA metadata coder 102 that performs the OMASA analysis and ISM / MASA metadata coding operations 152 of the combined format encoding method 150. The analyzer / coder 102 (a) analyzes the MASA audio channels and audio object channels using their respective metadata, (b) quantizes and codes the ISM metadata of the N audio object channels from the composer 101, (c) quantizes and codes the MASA metadata, and (d) may possibly downmix at least a portion of the N+2 audio channels from the composer 101, where the analysis, coding, and downmixing depend on the OMASA mode. Downmixing is typically used at a lower total codec bitrate (ivas_total_bitrate) when the available bit budget is too small to encode all audio objects and MASA individually. In these cases, one, several or even all audio object channels are mixed with the MASA audio channel to obtain M+2 transport channels, where M is the number of audio object channels to be coded separately and M≦N. The coded ISM metadata (line 175) and MASA metadata (line 121) are then routed to a bitstream writer 113, which performs a bitstream writing operation 163 for transmission of the resulting bitstream to a distant composite format decoder through a transmitter and communication channel (not shown).

[0032] Next, using a procedure such as that described in Reference [3], the M audio object channels 103 (audio stream without metadata) are analyzed and processed using an ISM format encoder 104 of the compound format encoder 100, which performs an ISM format encoding operation 154 of the compound format encoding method 150. The ISM format encoder 104 comprises M single channel elements (SCEs), where all M audio object channels 103 are analyzed and processed in parallel in a front pre-processor 105 of the ISM format encoder 104, which performs a front pre-processing operation 155 of the ISM format encoding operation 154. Although three single channel elements (SCEs) and three corresponding audio object channels 103 are shown in Figure 1, a number M different from three can obviously be implemented. The front pre-processing operation 155 generates ISM pre-processing parameters, including, for example, time-domain transient detection, spectral analysis, long-term prediction analysis, pitch tracking and voicing analysis, voice / sound activity detection (VAD / SAD), bandwidth detection, and noise estimation, for performing signal classification (coder type selection, signal classification, speech / music classification), as described, for example, in Reference [1].

[0033] Classification information, e.g., a VAD or local VAD flag as defined in EVS (Reference [1]), and / or coder type from the front pre-processor 105, is passed to the ISM classifier 106 of the ISM format encoder 104, which performs an ISM classification operation 156 of the ISM format encoding operation 154. The ISM classifier 106 receives the classification information from the front pre-processor 105 and further classifies the individual audio object channels 107 according to their importance, e.g., using a method based on the method from Reference [3] and further described in the following description. This classification information serves as the basis for a bitrate adaptation algorithm (see Reference [3]), which distributes the available bit budget among all M audio object channels 108 (audio stream without metadata) from the ISM classifier 106 using the core encoder configurator 109 of the ISM format encoder 104, which performs a core encoder configurator operation 159 of the ISM format encoding operation 154. The available bit budget for coding the audio stream is then the bit budget corresponding to the ism_total_brate bitrate minus the ism_metadata_brate bitrate for coding the metadata associated with the N audio object channels and the ism_signaling_brate bitrate for coding the ISM signaling. As explained herein above, the ism_total_brate bitrate 110 is configured in the OMASA configurator 101. The core encoder configurator 109 further configures the high-level parameters of the core encoder in the core encoder section (see 111), e.g., the internal sampling rate or the coded audio bandwidth, based on the actual available bitrate corresponding to ism_total_brate.

[0034] Once the core encoder configuration and bitrate distribution among the audio object channels 108 (audio streams without metadata) has been performed, the ISM format encoding operation 154 continues with a series of further pre-processing (further classification, core selection, further resampling, ...) and core encoding operations 161 performed by the pre-processor of the ISM format encoder 104 and the core encoder 111. The further pre-processing (operation 161) of the M audio object channels 112 (audio streams without metadata) from the core encoder configurer 109 is described, for example, in references [1] and [3]. Finally, the pre-processor and core encoder 111 comprises a core encoder section including M individual variable bit rate mono core encoders for sequentially encoding all M audio object channels 112 (audio streams without metadata), and the core encoder index is sent to a bit stream writer 113 which performs a bit stream writing operation 163, where the resulting bit stream from the bit stream writer 113 is transmitted to a distant composite format decoder via a transmitter and communication channel (not shown).

[0035] The mixed format encoder 100 includes a MASA format encoder 115 that performs the MASA format encoding operation 165 of the mixed format encoding method 150 .

[0036] In turn, the MASA format encoder 115 comprises a front preprocessor 116 that performs front preprocessing operations 166 of the MASA format encoding operation 165, a core encoder configurator 117 that performs core encoder configurator operations 167 of the MASA format encoding operation 165, and a preprocessor and core encoder 118 that performs further preprocessing (further classification, core selection, other resampling, ...) and core encoding operations 168 of the MASA format encoding operation 165.

[0037] By default, the stereo MASA audio channels 119 (audio streams without metadata) from the OMASA analyzer and ISM / MASA metadata coder 102 are coded using a channel pair element (CPE) MASA format encoder 115. Similar to the ISM format encoder 104, the MASA format encoder 115 starts with a front preprocessing operation 166 to generate MASA preprocessing parameters. Next, a core encoder configurator 117 receives information about the masa_total_bitrate 129 from the OMASA configurator 101 and sets high-level core encoder parameters. Finally, further preprocessing and core encoding operations 168 are performed on the two MASA audio channels 120 (audio streams without metadata), and a core encoder index is sent to a bitstream writer 113 that performs a bitstream writing operation 163, where the resulting bitstream from the bitstream writer 113 is transmitted to a remote composite format decoder via a transmitter and communication channel (not shown).

[0038] It should be pointed out that in the exemplary implementation described, the front pre-processor 116 is similar to the front pre-processor 105, the core encoder composer 117 is similar to the core encoder composer 109, and the pre-processor and core encoder 118 is similar to the pre-processor and core encoder 111.

[0039] Although this is not shown in the figures, the ISM signaling coded using the ism_signaling_brate bitrate and the MASA signaling coded using the masa_signaling_brate bitrate are forwarded to the writer 113 for insertion into the bitstream and transmission to a distant composite format decoder. Obviously, the masa_metadata_brate bitrate for coding metadata in the OMASA analyzer and MASA metadata coder 102, as well as the masa_signaling_brate bitrate, form part of the masa_total_brate bitrate. Similarly, the ism_metadata_brate bitrate for coding ISM metadata and the ism_signaling_brate bitrate form part of the ism_total_brate bitrate.

[0040] 2. Flexible mixed format bit rate adaptation Coding the audio objects and MASA that make up the OMASA composite format at constant bit rates masa_total_brate and ism_total_brate is usually not the most efficient way to allocate the available ivas_total_brate bitrate. In a typical scenario, an audio scene (e.g., ambience or main speaker) is coded by MASA, with additional individual speakers coded as separate audio objects. When one of the speakers is not speaking, the bit rate associated with this speaker can be reduced or set to 0, and the saved bit budget can be relinquished to coding the active speaker voice or ambience.

[0041] This disclosure thus extends the composite format coding method of [3] to create the masa_total_brate and ism_total_brate bitrates in one composite format coding variable, the ivas_total_brate bitrate, making the codec framework more flexible, adaptive, and efficient.

[0042] Referring to FIG. 2, the present disclosure thus introduces into the hybrid format encoder 100 a hybrid format bit rate adaptation device 201 that performs a hybrid format bit rate adaptation operation 251 (forming part of the hybrid format encoding method 150).

[0043] 2, the composite format bitrate adaptation device 201 receives information about the classification of the M audio object channels 107 (e.g., one parameter per audio object channel), and as explained above, the ISM classification of the M audio object channels 107 is based on the ISM pre-processing parameters from the front pre-processing operation 155. In the composite format bitrate adaptation device 201, the nominal (initial) bitrates of the audio object channels are adapted based on the classification information, resulting in a variable adapted bitrate ism_total_brate. Thus, the adapted masa_total_brate bitrate changes accordingly.

[0044] In the following, the present disclosure takes into account that the number M of separately coded audio object channels 103 is equal to the number N of input audio object channels, i.e., M=N, however, the disclosed algorithm is general such that M can be smaller than N without departing from the disclosed logic.

[0045] As mentioned in the above description, the classification in the ISM classifier 106 can be performed using, for example, the method from Reference [3]. Alternatively, the ISM classification can be based on one or more other front-end pre-processing parameters. An example of such an alternative is the combination of a coder type parameter and long-term noise, as described in Reference [1].

[0046] Therefore, the ISM classification can be based on several parameters and / or their combinations, such as coder type (coder_type), VAD, Forward Erasure Concealment (FEC) signal classification (class), speech / music classification decision, long-term signal-to-noise ratio (SNR) estimation from the open-loop ACELP / TCX core decision module (snr_celp, snr_tcx) of Reference [1], etc. In a non-limiting example, a simple ISM classification based on coder type as specified in Reference [3] is implemented. Thus, the ISM classifier 106 ranks the importance of the audio object channels 107 to the core encoder constructor 109. As a result, four distinct ISM importance classes are generated: ISM is prescribed. (a) Inactive class ISM_INACTIVE: For example, frames with VAD=0. (b) Low importance class ISM_LOW_IMP: Frames with coder_type=UNVOICED or INACTIVE. (c) Medium importance class ISM_MEDIUM_IMP: Frames with coder_type=VOICED. (d) High importance class ISM_HIGH_IMP: Frames with coder_type=GENERIC.

[0047] Thus, the output from the ISM classifier 106 are ISM importance flags (one per audio object channel), which further serve as driving parameters for setting the bitrates ism_brate(n) (n=1, 2, ..., N) and the MASA bitrate ism_masa_brate for all audio object channels in the combined format bitrate adaptation device 201. It should be pointed out that the combined format bitrate adaptation (operation 251) in this disclosure differs from the teaching of reference [3] in that the final ISM total bitrate ism_total_brate typically changes from frame to frame and is therefore variable according to this disclosure.

[0048] ISM importance class ISM is sent over line 114 from the ISM classifier 106 to the bitstream writer 113, where it is written into the bitstream and sent along with the bitstream to the far decoder, where it serves as a driving parameter for setting the bitrates ism_brate(n) and MASA bitrate ism_masa_brate for all audio object channels. Thus, the same hybrid format bitrate adaptation algorithm is used in both the encoder 100 and the far decoder.

[0049] In general, the hybrid format bit rate adaptation device 201 uses the following hybrid format bit rate adaptation logic to allocate higher bit rates to audio object channels with higher importance and lower bit rates to audio object channels with lower importance. (1) The initial ism_total_brate bitrate divided by the number of audio object channels N, i.e., For n=1, ..., N, ism_brate(n)=(ism_total_brate) / N (3) The initial ism_brate(n) bitrates for all N audio object channels are set as new 205 represents the "nominal" bit rate around which it fluctuates. (2) class ISM = ISM_INACTIVE frame: A constant low bit rate B as the ism_brate(n) bit rate for audio object channel n in this class VAD0 For example, a low bit rate B VAD0 may correspond to a low rate core coder mode within IVAS that encodes audio at 2.45 kbps. (3) class ISM =ISM_LOW_IMP frame: The initial ism_brate(n) for audio object channel n in this class is adapted using the following relation (4): ism_brate new (n)=γ low *ism_brate(n) (4) However, the weighting constant γ low is typically set to a value less than 1.0, for example, a value of 0.8. (4) class ISM =ISM_MEDIUM_IMP frame: The initial ism_brate(n) for audio object channel n in this class is adapted using the following relation (5): ism_brate new (n)=γ med *ism_brate(n) (5) However, the weighting constant γ med is γ low is set to a value greater than , for example, a value of 1.0. (5) class ISM=ISM_HIGH_IMP frame: The initial ism_brate(n) for audio object channel n in this class is adapted using the following relation (6): ism_brate new (n)=γ high *ism_brate(n) (6) However, the weighting constant γ high is usually greater than 1.0 (γ med is set to a value greater than 1.4, for example. (6) Adapted bitrate ism_brate new (n) is checked against the minimum and maximum thresholds supported by the codec for the particular configuration (which depends, for example, on the core encoder internal sampling rate, coded audio bandwidth, etc.). (7) Repeat steps (2) to (6) for every audio object channel n (n=1, ..., N).

[0050] Adapted ism_brate new After the (n) bit rates are calculated for all N audio object channels, the composite format bit rate adaptation device 201 calculates the adapted ISM total bit rate ism_total_bitrate using the following relation (7): new Calculate 205.

[0051]

number

[0052] Next, the core encoder configurator 109 configures the parameters of the core encoders in the core encoder section (see 111). For example, the internal sampling rate or the coded audio bandwidth of each core encoder is set based on the initial ism_brate(n) bitrate. Meanwhile, the adapted individual ism_brate(n) bitrates from the device for composite format bitrate adaptation 201 are configured. newThe bitrate is used by the core encoder configurator 109 to specify the individual bitrates to be attributed to each core encoder in the core encoder section (see 111) for coding the different audio object channels (without metadata and signaling bits). new The bitrate is a driving parameter for setting other core encoder parameters of the core encoder in the core encoder section (see 111), such as the core mode (eg, ACELP or TCX), the coder type, the BWE bitrate, etc.

[0053] Finally, in the composite format bit rate adaptation device 201, the adapted MASA total bit rate masa_total_bitrate new 210 is calculated using the following relation (8): masa_total_brate new =ivas_total_brate-ism_total_brate new (8)

[0054] An example of variable composite format bitrate adaptation (operation 251) in OMASA format coding is shown in Figure 3. In this example, a combination of two audio objects coded at 80 kbps and the MASA format was used. From the top of the figure, a first input audio object 1, a second input audio object 2, input audio MASA 3, output sound (binaural output) 4, reference ism_total_brate 5, reference masa_total_brate 6, new (adapted) ism_total_brate 7, and new (adapted) masa_total_brate 8 are shown, where "reference" corresponds to the codec variant without the composite format bitrate adaptation disclosed herein and "new" corresponds to the codec variant of which the composite format bitrate adaptation disclosed herein is a part.

[0055] The following can be understood from FIG. (a) In the reference variant (see 5 and 6 in FIG. 3), the ism_total_brate bitrate (5) is constant at 32 kbps and the masa_total_brate bitrate (6) is constant at 48 kbps. (b) In a new variant (using hybrid format bitrate adaptation, see 7 and 8 in Figure 3), the ism_total_brate bitrate (7) is variable and varies between 4.9 kbps and 44.8 kbps, and the masa_total_brate bitrate (8) is also variable and varies between 35.2 kbps and 75.1 kbps.

[0056] 3. Flexible mixed format bit rate adaptation variants The schematic block diagram of Figure 2 assumes that the hybrid format bitrate adaptation logic (device 201) relies on the ISM importance classification from ISM classifier 106. Note that the hybrid format bitrate adaptation logic can similarly rely on classifications from other format parameters, or from format parameters from both pre-processors 105 and 116. An example of hybrid format bitrate adaptation logic that relies on parameters from both the ISM front pre-processor 105 and the MASA front pre-processor 116 is shown in Figure 4.

[0057] In FIG. 4, an example of a MASA parameter that may be employed in the composite format bit rate adaptation logic (device 201) is the low-pass filtered (LP) long-term (LT) noise energy value lp from the MASA front pre-processor 116. noise 220. The idea is to calculate the LT noise energy lp of the MASA audio channel (i.e., the scene ambience or main speaker). noise (MASA) is the LT noise energy value lp of the audio object channel noiseIt is based on the assumption that audio objects can be coded using the low bitrate core coder mode in IVAS, which encodes audio at 2.45 kbps, which is high compared to (ISM(n)). Therefore, the lp of the scene ambience coded by MASA noise (MASA)220 and audio object lp noise (ISM(n)) 230 is calculated and compared to a threshold. In other words, an audio object channel is coded in a low bitrate core coder mode if its background noise would be "masked" by the scene ambience audio. Therefore, in this example, the hybrid format bitrate adaptation (operation 251) relies on both the front pre-processor 105 of the ISM format encoder 104 and the front pre-processor 116 of the MASA format encoder 115.

[0058] Therefore, the ISM classification differs from the previous description, and the inactive class ISM_INACTIVE is set under the condition of relation (9). (lp noise (MASA)-lp noise If (ISM(n)) ≥ δ), then class ISM =ISM_INACTIVE (9)

[0059] Relation (9) is applied in a loop for all N audio objects, where δ is the threshold mentioned above, e.g., δ=30. Note that the processing in the ISM format encoder 104 and the MASA format encoder 115 is performed serially, and that the parameters of the current frame for the ISM and MASA formats may not be available when performing the combined format bit rate adaptation operation 251. To get around this limitation, parameters from the previous frame can be used instead. For example, lp from the current frame noise (ISM(n)) value and lp from the previous frame noiseThe (MASA) value can be used in relation (9).

[0060] In the previous description, the ISM classification was considered to be independent of the ism_total_brate bitrate. However, in other instances, it may be advantageous to modify the classification at a higher initial (non-adapted) ism_total_brate bitrate, where the bitrates per audio object and MASA are all sufficiently high. For example, B VAD0 The ISM_INACTIVE class in and the corresponding low bitrate coding is not used at higher bitrates, for example in the case of IVAS when the initial bitrate ism_brate(n)>48 kbps.

[0061] According to another example, the ISM classification of audio object channels into one of a plurality of ISM importance classes (operation 156) may be adapted depending on the number of separately coded audio object channels and the initial ISM ism_total_brate bitrate. To that end, the device for combined format bitrate adaptation 201 (a) sets initial bitrates ism_brate(n) for the audio object channels as described above, (b) adapts the initial bitrates by multiplying these initial bitrates ism_brate(n) by a weighting constant associated with each ISM importance class of the ISM importance classes, respectively, as described above, and (c) adjusts the weighting constant γ depending on the number N of separately coded audio object channels and the initial ISM bitrate ism_brate(n) due to the audio object channels. low , γ med , and γ high Modify the.

[0062] Similarly, in a further example, it is not necessary to employ a low-rate core coder mode at higher ism_total_rate bit rates.ISM ISM_INACTIVE is omitted in these cases, and class ISM ISM_LOW_IMP is used instead.

[0063] 4. Combined Format Decoder A composite format decoder (not shown) receives a bitstream from the bitstream writer 113, which typically includes audio format signaling, ISM or audio object transport channels (M×SCE index), ISM metadata, MASA transport audio channels (SCE or CPE index), MASA metadata, and indices related to several codec modules, including the composite format signaling. First, the decoder reads and decodes information about the audio format from the received bitstream. In the case of a composite format, the decoder then reads the information needed to set the individual format bitrates, including the number N of audio objects and the ISM importance class for every audio object channel.

[0064] In the case of OMASA, the decoder thus reads the number N of audio objects and their ISM importance classes (one parameter per audio object channel). These parameters are then used to implement the combined format bitrate adaptation logic in the same manner as in the encoder. The output from this logic are the ISM and MASA total bitrate parameters ism_total_brate and masa_total_brate, which are further used to configure the core decoder for ISM and MASA decoding. The core decoders for the ISM and MASA portions then output N compositions corresponding to the N audio objects plus two MASA compositions, which are processed consecutively. Finally, all N+2 compositions with decoded ISM and MASA metadata are fed to a renderer that generates the final spatial sound in the desired output audio format (e.g., binaural, multi-channel, etc.).

[0065] In one example implementation, a composite format decoder (and corresponding composite format decoding method) is provided for decoding a first number N of audio object channels and a second number "+2" of MASA audio transport channels in an ISM format. The composite format decoder (not shown) includes: a receiver of the bitstream; - a bitstream decoder for decoding from the bitstream information about the codec total bit rate, information about the codec format (e.g., OMASA format in IVAS), information about the first number N of ISM audio object channels, ISM audio channel coding information, MASA audio channel coding information, and information about the ISM importance class per audio object channel; - Adapted ISM total bitrate ism_total_brate new and the adapted MASA total bitrate masa_total_bratenew a device for combined format bitrate adaptation using a first number N of ISM audio object channels, a codec total bitrate ivas_total_brate, and an ISM importance class per audio object channel to provide: an ISM core decoder for decoding audio object channels in response to ISM audio channel coding information from the bitstream, and an adapted ISM total bitrate ism_total_bitrate; new an ISM core decoder constructor responsive to a MASA core decoder for decoding MASA audio channels in response to MASA audio channel coding information from the bitstream, and an adapted MASA total bitrate masa_total_bitrate new and a MASA core decoder constructor responsive to the

[0066] 5. Exemplary Configurations of Hardware Components FIG. 5 is a simplified block diagram of an exemplary configuration of hardware components forming the above-described compound format encoder 100 using a device for compound format bit rate adaptation, compound format encoding method 150 using compound format bit rate adaptation, compound format decoder, and compound format decoding method (hereinafter, “compound format encoder / decoder and encoding / decoding method”).

[0067] The mixed-format encoder / decoder and encoding / decoding method may be implemented as part of a mobile terminal, as part of a portable media player, or in any similar device. The mixed-format encoder / decoder (identified as 500 in FIG. 5) comprises an input section 502, an output section 504, a processor 506, and a memory 508.

[0068] The input unit 502 is configured to receive input signal information. The output unit 504 is configured to provide output signal information. The input unit 502 and the output unit 504 may be implemented in a common module, for example, a serial input / output device.

[0069] The processor 506 is operatively connected to the input 502, to the output 504, and to the memory 508. The processor 506 may be implemented as one or more processors for executing code instructions supporting the functions of the various operations and elements of the above-described combined format encoder / decoder and encoding / decoding methods as shown in the accompanying drawings and / or described in this disclosure.

[0070] The memory 508 may comprise non-transitory memory for storing code instructions executable by the processor 506, in particular processor-readable memory that comprises / stores non-transitory instructions that, when executed, cause the processor to perform operations and elements of the combined format encoder / decoder and encoding / decoding method. The memory 508 may also comprise random access memory or buffers for storing intermediate processed data from various functions performed by the processor 508.

[0071] Those skilled in the art will appreciate that the description of the compound format encoder / decoder and encoding / decoding method is merely exemplary and is not intended to be limiting in any way. Other embodiments will readily suggest themselves to such skilled artisans having the benefit of this disclosure. Furthermore, the disclosed compound format encoder / decoder and encoding / decoding method may be customized to provide useful solutions to existing needs and problems of encoding and decoding sound.

[0072] For the sake of clarity, not all of the conventional features of implementations of compound format encoders / decoders and encoding / decoding methods are shown and described. Of course, it will be appreciated that in the development of any such actual implementation of compound format encoders / decoders and encoding / decoding methods, numerous implementation-specific decisions may need to be made to achieve the specific goals of the developer, such as compliance with application-related, system-related, network-related, and business-related constraints, and that these specific goals will vary from implementation to implementation and from developer to developer. Moreover, it will be appreciated that the development effort may be complex and time-consuming, but will nevertheless be a routine undertaking of engineering for one of ordinary skill in the art of sound processing having the benefit of this disclosure.

[0073] In accordance with this disclosure, the elements, processing operations, and / or data structures described herein may be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. In addition, those skilled in the art will recognize that devices of a less general-purpose nature, such as hardwired devices, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., may also be used. When a method comprising a series of operations and sub-operations is performed by a processor, the computer or machine, and the operations and sub-operations may be stored as a series of non-transitory code instructions readable by the processor, computer, or machine, which may be stored on a tangible and / or non-transitory medium.

[0074] The processing operations and elements of the hybrid format encoders / decoders and encoding / decoding methods as described herein may comprise software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.

[0075] In the mixed-format encoder / decoder and encoding / decoding methods, various processing operations and sub-operations may be performed in various orders, and some of the processing operations and sub-operations may be optional.

[0076] Although the present disclosure has been described above by way of non-limiting exemplary embodiments thereof, these embodiments may be freely modified within the scope of the appended claims without departing from the spirit and essence of the present disclosure.

[0077] 7. References This disclosure refers to the following references, the entire contents of which are incorporated herein by reference: [1] 3GPP TS 26.445, v.17.0.0, "Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description", April 2022 [2] 3GPP SA4 Contribution S4-170749, “New WID on EVS Codec Extension for Immersive Voice and Audio Services,” SA4 Meeting #94, June 26-30, 2017, http: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_94 / Docs / S4-170749.zip [3] V. Eksler, “Method and System for Coding Metadata in Audio Streams and for Efficient Bitrate Allocation to Audio Streams Coding,” U.S. Patent Application No. 17 / 596,567, filed December 13, 2021, and published under US20220319524 A1. [4] 3GPP SA4 contribution S4-180462, “On spatial metadata for IVAS spatial audio input format”, SA4 meeting #98, April 9-13, 2018, https: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_98 / Docs / S4-180462.zip [5] 3GPP SA4 contribution S4-220443, "MASA format updates", SA4 meeting #118-e, April 6-14, 2022, https: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_118-e / Docs / S4-220443.zip

[0078] 8. Source Code The ISM classification algorithm used by the ISM classifier 106 may be implemented, for example, using the following pseudocode:

[0079]

number

number

[0080] An algorithm for mixed format bitrate adaptation in an audio codec that adapts the bitrate of an audio object channel can be implemented, for example, using the following pseudocode:

[0081]

number

number

[0082] 100 Mixed Format Encoder 101 OMASA Configurator 102 OMASA analyzer and ISM / MASA metadata coder 103 Audio Object Channels 104 ISM Format Encoder 105 Front Preprocessor 106 ISM classifier 107 Audio Object Channels 108 audio object channels 109 Core Encoder Configuration 110 ism_total_brate bitrate 111 Preprocessor and Core Encoder 112 audio object channels 113 Bitstream Writer 115 MASA format encoder 116 Front Preprocessor 117 Core Encoder Configuration 118 Preprocessor and Core Encoder 119 stereo MASA audio channels 120 MASA audio channels 129 masa_total_brate bitrate 150 Combined Format Encoding Method 151 OMASA configuration operations 152 OMASA Analysis and ISM / MASA Metadata Coding Operations 154 ISM format encoding operation 155 Front Pre-processing Operation 156 ISM classification behavior 159 Core Encoder Configuration Operation 161 Preprocessing and Core Coding Operations 163 Bitstream Write Operations 165 MASA format encoding operation 166 Front Preprocessing Operation 167 Core Encoder Configuration Operation 168 Preprocessing and Core Coding Operations 201 Multi-format bit rate adaptation device 205 Adapted ISM Total Bitrate 210 Adapted MASA total bitrate 220 Low-pass filtered long-term noise energy values 230 Low-pass filtered long-term noise energy values 251 Mixed Format Bit Rate Adaptation 500 Mixed Format Encoder / Decoder 502 Input section 504 Output Section 506 processor 508 memory

Claims

1. 1. A hybrid format encoder for coding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the encoder comprising:

1. An ISM format encoder, comprising: - a first front pre-processor of the audio object channel for generating ISM pre-processing parameters; and - an adapted ISM total bit rate sensitive first core encoder section for coding said audio object channels; an ISM format encoder comprising: an adapted IA aggregate bit rate sensitive IA format encoder for coding the second audio channel; a device for hybrid format bitrate adaptation using at least one ISM preprocessing parameter from the first front preprocessor to (a) adapt an initial ISM aggregate bitrate to result in the adapted ISM aggregate bitrate, and (b) adapt an initial IA aggregate bitrate to result in the adapted IA aggregate bitrate; and A composite format encoder comprising:

2. 2. The composite format encoder of claim 1, wherein the first core encoder section comprises a first core encoder for coding the audio object channels and a composer of the first core encoder responsive to the adapted ISM gross bit rate.

3. the IA format encoder - a second front pre-processor for the second audio channel for generating IA pre-processing parameters; - a second core encoder section sensitive to the adapted IA aggregate bit rate for coding the second audio channel; 3. A compound format encoder according to claim 1, comprising:

4. 4. The composite format encoder of claim 3, wherein the second core encoder section comprises a second core encoder for coding the second audio channel and a composer of the second core encoder responsive to the adapted IA aggregate bit rate.

5. said device for mixed format bit rate adaptation comprising: (a) adapting the initial ISM aggregate bitrate to result in the adapted ISM aggregate bitrate; and (b) adapting the initial IA aggregate bitrate to result in the adapted IA aggregate bitrate. - at least one ISM preprocessing parameter from said first front preprocessor; or at least one ISM preprocessing parameter from the first front preprocessor and at least one IA preprocessing parameter from the second front preprocessor; 4. The composite format encoder of claim 3, wherein:

6. 6. A composite format encoder according to claim 1, comprising an ISM classifier for the audio object channels into one of a plurality of ISM importance classes, using the at least one ISM pre-processing parameter from the first front pre-processor.

7. The ISM importance class is: - Inactive class for frames with Voice Activity Detection (VAD) flag equal to 0, - low importance class for silent or inactive frames, - a medium importance class for voiced frames, and - High importance class for general frames 7. The composite format encoder of claim 6, selected from the group consisting of:

8. 8. A compound format encoder as claimed in claim 6 or 7, wherein the device for compound format bit rate adaptation (a) sets initial bit rates for coding the audio object channels, and (b) adapts a portion of the initial bit rates by multiplying these initial bit rates by weighting constants associated with each ISM importance class of the ISM importance classes, respectively.

9. 8. The compound format encoder of claim 7, wherein the device for compound format bit rate adaptation (a) sets initial bit rates for coding the audio object channels, and (b) adapts the initial bit rates for the low importance class, the medium importance class, and the high importance class by multiplying the initial bit rates by weighting constants associated with each importance class of the low importance class, the medium importance class, and the high importance class, respectively.

10. 10. A compound format encoder as claimed in claim 8 or 9, wherein the weighting constants comprise at least one weighting constant less than 1.0 and at least one weighting constant greater than 1.

0.

11. 8. The compound format encoder of claim 7, wherein when the ISM classifier classifies one of the audio object channels into the inactive class, the device for compound format bit rate adaptation allocates a constant bit rate for coding one of the audio object channels.

12. 10. A compound format encoder according to claim 8 or 9, wherein the device for compound format bit rate adaptation calculates the adapted ISM total bit rate as the sum of adapted bit rates of the individual audio object channels including metadata of these individual audio object channels.

13. 13. A compound format encoder according to claim 1, wherein the device for compound format bitrate adaptation adapts the IA aggregate bitrate by subtracting the adapted ISM aggregate bitrate from a total codec bitrate.

14. 6. The composite format encoder of claim 5, wherein the at least one ISM pre-processing parameter comprises a low-pass filtered long-term noise energy value of one of the audio object channels and / or the at least one IA pre-processing parameter comprises the low-pass filtered long-term noise energy value of the second audio channel.

15. The compound format encoder of claim 5 , wherein the at least one ISM pre-processing parameter or the at least one IA pre-processing parameter used in a current frame is a parameter from a previous frame.

16. The composite format encoder of claim 5 , wherein the at least one IA pre-processing parameter used in a current frame is a parameter from a previous frame.

17. 12. The compound format encoder of claim 6, wherein the ISM classifier for the audio object channels adapts the classification depending on an initial ISM gross bitrate.

18. 10. The compound format encoder of claim 8, wherein the ISM classifier for the audio object channels adapts the classification depending on an initial ISM total bitrate, and the weighting constants depend on the initial ISM total bitrate.

19. 19. A compound format encoder according to claim 1, further comprising an ISM classifier for the audio object channels into one of a plurality of ISM importance classes, the ISM classifier for the audio object channels adapting the classification depending on the number of audio object channels to be coded separately and on an initial ISM bitrate by audio object channels.

20. 20. The compound format encoder of claim 19, wherein the device for compound format bit rate adaptation (a) sets initial bit rates for coding the audio object channels, (b) adapts the initial bit rates by multiplying these initial bit rates by weighting constants associated with each ISM importance class of the ISM importance classes, respectively, and (c) modifies the weighting constants depending on the number of separately coded audio object channels and the initial ISM bit rates by audio object channels.

21. 1. A hybrid format encoder for coding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the encoder comprising: at least one processor; a memory coupled to the processor for storing non-transitory instructions; the non-transient instructions, when executed, cause the processor to:

1. An ISM format encoder, comprising: - a first front pre-processor of the audio object channel for generating ISM pre-processing parameters; and - an adapted ISM total bit rate sensitive first core encoder section for coding said audio object channels; an ISM format encoder comprising: an adapted IA aggregate bit rate sensitive IA format encoder for coding the second audio channel; a device for hybrid format bitrate adaptation using at least one ISM preprocessing parameter from the first front preprocessor to (a) adapt an initial ISM aggregate bitrate to result in the adapted ISM aggregate bitrate, and (b) adapt an initial IA aggregate bitrate to result in the adapted IA aggregate bitrate; and A composite format encoder that implements the above.

22. 1. A hybrid format encoder for coding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the encoder comprising: at least one processor; a memory coupled to the processor for storing non-transitory instructions; the non-transient instructions, when executed, cause the processor to: encoding the audio object channels, - front-preprocessing the audio object channels to generate ISM preprocessing parameters; and - core-encoding the audio object channels in response to an adapted ISM aggregate bit rate. and encoding the second audio channel in response to the adapted IA aggregate bit rate; (a) adapting an initial ISM aggregate bitrate to result in the adapted ISM aggregate bitrate; and (b) adapting an initial IA aggregate bitrate to result in the adapted IA aggregate bitrate using at least one ISM pre-processing parameter from the front pre-processing of the audio object channels. A mixed format encoder.

23. 1. A hybrid format decoder for decoding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the decoder comprising: a bitstream receiver; a bitstream decoder for decoding from said bitstream information about the codec total bitrate, information about the number of audio object channels, audio object channel coding information, second audio channel coding information, and information about the ISM importance class for each audio object channel; a device for hybrid format bitrate adaptation using the number of audio object channels, the codec total bitrate, and the ISM importance class per audio object channel to result in an adapted ISM total bitrate and an adapted IA total bitrate; an ISM core decoder for decoding the audio object channel in response to the audio object channel coding information from the bitstream, and a composer of the ISM core decoder in response to the adapted ISM aggregate bitrate; an IA core decoder for decoding the second audio channel in response to the second audio channel coding information from the bitstream, and a configurer of the IA core decoder in response to the adapted IA aggregate bitrate; A composite format decoder comprising:

24. 1. A hybrid format decoder for decoding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the decoder comprising: at least one processor; a memory coupled to the processor for storing non-transitory instructions; the non-transient instructions, when executed, cause the processor to: a bitstream receiver; a bitstream decoder for decoding from said bitstream information about the codec total bitrate, information about the number of audio object channels, audio object channel coding information, second audio channel coding information, and information about the ISM importance class for each audio object channel; a device for hybrid format bitrate adaptation using the number of audio object channels, the codec total bitrate, and the ISM importance class per audio object channel to result in an adapted ISM total bitrate and an adapted IA total bitrate; an ISM core decoder for decoding the audio object channel in response to the audio object channel coding information from the bitstream, and a composer of the ISM core decoder in response to the adapted ISM aggregate bitrate; an IA core decoder for decoding the second audio channel in response to the second audio channel coding information from the bitstream, and a configurer of the IA core decoder in response to the adapted IA aggregate bitrate; A composite format decoder that implements the above.

25. 1. A hybrid format decoder for decoding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the decoder comprising: at least one processor; a memory coupled to the processor for storing non-transitory instructions; the non-transient instructions, when executed, cause the processor to: receiving a bitstream; decoding information about the codec total bit rate, information about the number of audio object channels, audio object channel coding information, second audio channel coding information, and information about the ISM importance class for each audio object channel from the bitstream; using the number of audio object channels, the codec total bit rate, and the ISM importance class per audio object channel to yield an adapted ISM total bit rate and an adapted IA total bit rate; core decoding the audio object channels in response to the audio object channel coding information from the bitstream, and configuring the core decoding of the audio object channels in response to the adapted ISM aggregate bitrate; core decoding the second audio channel in response to the second audio channel coding information from the bitstream, and configuring the core decoding of the second audio channel in response to the adapted IA aggregate bitrate; A composite format decoder.

26. 1. A composite format method for encoding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the method comprising: encoding the audio object channel in an ISM format, - front-preprocessing the audio object channels to generate ISM preprocessing parameters; and - core-encoding the audio object channels in response to an adapted ISM aggregate bitrate. and IA format encoding the second audio channel in response to the adapted IA total bit rate; a composite format bitrate adaptation step using at least one ISM pre-processing parameter from the front pre-processing of the audio object channels to (a) adapt an initial ISM gross bitrate to result in the adapted ISM gross bitrate, and (b) adapt an initial IA gross bitrate to result in the adapted IA gross bitrate; A composite formatting method comprising:

27. 27. The hybrid format encoding method of claim 26, wherein core encoding the audio object channels comprises using a first core encoder to code the audio object channels and configuring the first core encoder in response to the adapted ISM aggregate bit rate.

28. encoding the second audio channel in an IA format, - front-preprocessing the second audio channel to generate IA preprocessing parameters; - core-encoding the second audio channel in response to the adapted IA aggregate bitrate; 28. The method of claim 26 or 27, comprising:

29. 30. The hybrid format encoding method of claim 28, wherein core encoding the second audio channel comprises using a second core encoder to code the second audio channel and configuring the second core encoder in response to the adapted IA aggregate bit rate.

30. (a) adapting the initial ISM aggregate bitrate to result in the adapted ISM aggregate bitrate; and (b) adapting the initial IA aggregate bitrate to result in the adapted IA aggregate bitrate. said hybrid format bit rate adaptation comprising: - at least one ISM pre-processing parameter from the front pre-processing of the second audio channel, or - at least one ISM pre-processing parameter from the front pre-processing of the audio object channel and at least one IA pre-processing parameter from the front pre-processing of the second audio channel; 29. The mixed-format encoding method of claim 28, wherein:

31. 31. A mixed-format encoding method according to any one of claims 26 to 30, comprising classifying the audio object channels into one of a plurality of ISM importance classes.

32. The ISM importance class is: - Inactive class for frames with Voice Activity Detection (VAD) flag equal to 0, - low importance class for silent or inactive frames, - a medium importance class for voiced frames, and - High importance class for general frames 32. The composite format encoding method of claim 31, wherein the composite format encoding method is selected from the group consisting of:

33. 33. The method of claim 31 or 32, wherein the hybrid format bit rate adaptation comprises: (a) setting initial bit rates for coding the audio object channels; and (b) adapting a portion of the initial bit rates by multiplying these initial bit rates by weighting constants associated with each of the ISM importance classes, respectively.

34. 33. The hybrid format encoding method of claim 32, wherein the hybrid format bit rate adaptation comprises: (a) setting an initial bit rate for coding the audio object channels; and (b) adapting the initial bit rate for the low importance class, the medium importance class, and the high importance class by multiplying the initial bit rate by a weighting constant associated with each importance class of the low importance class, the medium importance class, and the high importance class, respectively.

35. 35. A method of hybrid format encoding according to claim 33 or 34, wherein the weighting constants comprise at least one weighting constant less than 1.0 and at least one weighting constant greater than 1.

0.

36. 33. The hybrid format encoding method of claim 32, wherein the hybrid format bit rate adaptation allocates a constant bit rate for coding one of the audio object channels without metadata when the classification of the audio object channels classifies the one audio object channel into the inactive class.

37. 35. A mixed format encoding method according to claim 33 or 34, wherein the mixed format bit rate adaptation calculates the adapted ISM total bit rate as the sum of the adapted bit rates of the individual audio object channels including metadata of these individual audio object channels.

38. 38. The hybrid format encoding method of claim 26, wherein the hybrid format bitrate adaptation adapts the IA aggregate bitrate by subtracting the adapted ISM aggregate bitrate from a total codec bitrate.

39. 31. The composite format encoding method of claim 30, wherein the at least one ISM pre-processing parameter comprises a low-pass filtered long-term noise energy value of one of the audio object channels and / or the at least one IA pre-processing parameter comprises the low-pass filtered long-term noise energy value of the second audio channel.

40. 31. The mixed-format encoding method of claim 30, wherein the at least one ISM pre-processing parameter or the at least one IA pre-processing parameter used in a current frame is a parameter from a previous frame.

41. 31. The mixed-format encoding method of claim 30, wherein the at least one IA pre-processing parameter used in a current frame is a parameter from a previous frame.

42. 35. A mixed-format encoding method according to any one of claims 31 to 34, wherein classifying the audio object channels comprises modifying the classification depending on an initial ISM gross bitrate.

43. 35. A method for encoding audio object channels in a mixed format according to claim 33 or 34, wherein classifying the audio object channels comprises modifying the classification depending on an initial ISM total bit rate, and wherein the weighting constants depend on the initial ISM total bit rate.

44. 44. A method for encoding a composite format according to any one of claims 26 to 43, comprising the step of classifying the audio object channels into one of a plurality of ISM importance classes, and modifying the classification of the audio object channels depending on the number of audio object channels to be coded separately and on an initial ISM bit rate by the audio object channels.

45. 45. The method of claim 44, wherein the hybrid format bit rate adaptation comprises: (a) setting initial bit rates for coding the audio object channels; (b) adapting the initial bit rates by multiplying these initial bit rates by weighting constants associated with each ISM importance class of the ISM importance classes, respectively; and (c) modifying the weighting constants depending on the number of separately coded audio object channels and the initial ISM bit rates by audio object channels.

46. 1. A composite format method for decoding a number of first audio object channels in an ISM format and a number of second audio channels in an immersive audio (IA) format, the method comprising: receiving a bitstream; decoding from the bitstream information about the codec total bit rate, information about the number of ISM audio object channels, audio object channel coding information, second audio channel coding information, and information about the ISM importance class for each audio object channel; hybrid format bitrate adaptation using the number of audio object channels, the codec total bitrate, and the ISM importance class per audio object channel to result in an adapted ISM total bitrate and an adapted IA total bitrate; core decoding the audio object channels in response to the audio channel coding information from the bitstream, comprising configuring an ISM core decoder in response to the adapted ISM aggregate bitrate; core decoding the second audio channel in response to the second audio channel coding information from the bitstream, including configuring core decoding of the second audio channel in response to the adapted IA aggregate bitrate; A composite formatting method comprising:

47. 23. The composite format encoder of claim 1, wherein the immersive audio (IA) format is the MASA format.

48. 26. A composite format decoder according to any one of claims 23 to 25, wherein the immersive audio (IA) format is the MASA format.

49. 46. ​​The mixed format encoding method of claim 26, wherein the immersive audio (IA) format is the MASA format.

50. 47. The mixed format encoding method of claim 46, wherein the immersive audio (IA) format is the MASA format.