Method and device for discontinuous transmission in an object-based audio codec
The method and device for DTX in object-based audio codecs synchronize SID frame signaling across channels using a global counter and adaptive metadata encoding, addressing the inefficiencies in multi-channel codecs by reducing bit rates and enhancing immersive audio experiences.
Patent Information
- Application Number
- JP2025528795
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-18
- Filing Date
- 2023-11-14
- Publication Date
- 2025-10-30
AI Technical Summary
Existing DTX methods in multi-channel codecs face a trade-off between maintaining a low SID bit rate and representing a large number of channels, leading to inefficient coding.
A method and device for discontinuous transmission (DTX) in object-based audio codecs that analyze audio streams to detect DTX signal segments and SID frames, using a global SID counter for synchronized signaling and efficient coding of SID frames, along with adaptive metadata encoding to manage bit rates.
This approach reduces bit rates while maintaining high-quality immersive audio experiences by efficiently encoding and decoding audio objects during silence periods, ensuring accurate recreation of background noise and spatial metadata.
Smart Images

Figure 2025536102000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to audio coding, and particularly, but not exclusively, to methods and devices for discontinuous transmission (DTX) in object-based audio codecs.
[0002] In this disclosure and the accompanying claims: (a) The term "audio" may refer to speech, music, and any other sound. (b) The term "multi-channel" may refer to two or more channels. (c) The term "stereo" is a contraction of "stereophonic." (d) The term "mono" is an abbreviation of "monophonic." (e) The term "object-based audio" is intended to refer to an auditory scene as a collection of individual elements, also known as audio objects, and may comprise, for example, speech, music, and any other sound, including commonly audible sounds. (f) The term "audio object" is intended to refer to an audio stream with associated metadata. For example, in this disclosure, an "audio object" is referred to as an independent audio stream with metadata (ISM). (g) The term "audio stream" is intended to denote an audio waveform, e.g., speech, music, and / or any other sound, including general audible sounds, in a bitstream, which may consist of one channel (mono), although multi-channel (stereo) including two channels may also be considered. (h) The term "metadata" is intended to denote a set of information, e.g., describing an audio stream and artistic intent, used to transform an original or coded audio object into a playback system. The metadata typically describes spatial properties of each individual audio object, such as position, orientation, volume, width, etc. As a non-limiting example, two sets of metadata are considered in the context of this disclosure: Input Metadata: unquantized metadata representation used as input to the codec; this disclosure is not limited to a particular format of the input metadata; and - Coded Metadata: quantized and coded metadata that forms part of the bitstream transmitted from the encoder to the decoder. (i) The term "audio format" is intended to refer to a method for achieving an immersive audio experience. (j) The term "playback system" is intended to refer, for example, but not exclusively, to an element capable of rendering audio objects at a decoder in a 3D (three-dimensional) audio space around a listener using transmitted metadata and artistic intent at the playback side. Rendering may be performed for a target loudspeaker layout (e.g., 5.1 surround) or headphones, although metadata may be dynamically modified in response to feedback from, for example, a head-tracking device. Other types of rendering may also be contemplated. [Background technology]
[0003] Discontinuous transmission (DTX) is used in mobile communication systems to turn off radio transmitters during speech or general audio pauses. The use of DTX saves power in mobile stations and increases the time required between battery recharges. It also reduces the overall interference level, thus improving transmission quality. However, if the channel is completely disconnected during speech or general audio pauses, the background noise that is normally transmitted along with the speech or general audio also disappears. This results in an audio signal (silence) that sounds unnatural at the receiving end of the communication.
[0004] Instead of turning off transmission completely during speech or general audio pauses, several techniques have been developed in which parameters characterizing the background noise are generated and transmitted at a low bit rate in a silence insertion descriptor (SID) frame bitstream. These parameters, often called comfort noise (CN) parameters, can then be used at the receiving end (decoder) to recreate background noise that emphasizes as much as possible the spatial and temporal content of the background noise at the transmitting end (encoder). This process for recreating background noise is known as comfort noise generation (CNG).
[0005] Historically, conversational telephony has been implemented using mono handsets, which have only one earpiece for outputting sound to only one ear of the user. As a result, mono codec SIDs can achieve low bit rates. In the past decade, users have begun to use portable handsets with headphones to listen to sound with both ears, primarily for listening to music, but also occasionally for listening to speech. Nevertheless, when a portable handset is used to transmit and receive conversations, the content is still mono, but when headphones are used, it is delivered to both ears of the user.
[0006] The 3GPP (Third Generation Partnership Project) speech coding standard, which implements the Codec for Enhanced Voice Services (EVS) as described in Reference [1], the entire contents of which are incorporated herein by reference, has greatly improved the quality of coded audible sounds, such as speech, music, and any other sounds transmitted and received through portable handsets. The next natural step is to transmit stereo information so that the receiver gets as close as possible to the real-world audio scene captured at the other end of the communication link.
[0007] Moreover, over the past few years, audio generation, recording, presentation, coding, transmission, and playback have been moving toward enhanced, interactive, and immersive experiences for listeners. An immersive experience can be described, for example, as a state of being deeply drawn into or absorbed in an audio scene while sounds are coming from all directions. Immersive audio (also known as 3D audio) recreates sound images in all three dimensions around the listener, taking into account various sound characteristics such as timbre, directionality, reverberation, clarity, and accuracy of (auditory) spaciousness. Immersive audio is created for specific audio playback or reproduction systems, such as loudspeaker-based systems, integrated playback systems (sound bars), or headphones. Interactivity with an audio playback system can include, for example, the ability to adjust the volume, change the location of the sound, or select different languages for playback.
[0008] There are three basic approaches (hereafter also referred to as audio formats) to achieving an immersive audio experience:
[0009] The first approach is channel-based audio, where multiple spaced microphones are used to capture sound from different directions, with one microphone corresponding to one audio channel in a particular loudspeaker layout. Each recorded channel is fed to a loudspeaker in a specific location. Examples of channel-based audio include, for example, stereo, 5.1 surround, 5.1+4, etc.
[0010] The second approach is scene-based audio (SBA), which represents the desired sound field over a local space as a function of time by a combination of dimensional components. The signals representing scene-based audio are independent of the position of the sound sources, but the sound field must be transformed into a chosen loudspeaker layout in the rendering playback system. An example of scene-based audio is Ambisonics.
[0011] The third and final immersive audio approach is object-based audio, which represents an auditory scene as a set of audio elements (e.g., singers, drums, guitars) along with information about their positions in the audio scene, so that the individual audio elements can be rendered by the playback system in their intended positions. This gives object-based audio great flexibility and interactivity, as each object remains distinct and can be manipulated individually.
[0012] Beyond the basic approach, new multi-channel coding techniques have been developed, such as Metadata-Assisted Spatial Audio (MASA), as described in Reference [5], the entire contents of which are incorporated herein by reference. In the MASA approach, MASA metadata (e.g., direction, energy ratio, spread coherence, distance, surround coherence, all within a few time-frequency slots) is generated, quantized, coded into a bitstream in a MASA analyzer, while the MASA audio channels are treated as (multi-)mono or (multi-)stereo transport signals that are coded by a core encoder. In a MASA decoder, the MASA metadata then guides the decoding and rendering process to reproduce the output spatial sound.
[0013] Each of the above-described audio techniques for achieving an immersive experience has advantages and disadvantages. Therefore, in complex audio systems, it is common to combine several audio techniques to create an immersive auditory scene, rather than using only one audio technique. An example of this would be an audio system that combines scene-based audio (SBA) or MASA with object-based audio, e.g., combining SBA or MASA with a small number of distinct audio objects.
[0014] Recently, 3GPP® has begun work on developing a 3D audio codec for immersive services called IVAS (Immersive Voice and Audio Services) as described in Reference [2], the entire contents of which are incorporated herein by reference, which is based on the EVS codec as described in Reference [1]. The IVAS codec is a multi-channel codec where the bitrate requirements are usually more stringent as the number of channels coded and transmitted increases. Summary of the Invention [Problem to be solved by the invention]
[0015] Therefore, DTX operation in a multi-channel codec must address a trade-off between (a) keeping the SID bit rate low and (b) using a large number of channels to be represented. For example, if each channel were represented by its own SID, the SID bit rate for the overall codec would be too high. As a result, efficient DTX methods and SID coding are needed. [Means for solving the problem]
[0016] According to a first aspect, the present disclosure relates to a method for discontinuous transmission (DTX) of audio objects in an object-based audio codec, the audio objects including respective audio streams, the method comprising the steps of: analyzing the audio streams to yield voice or signal activity information for the audio objects; detecting DTX signal segments of the audio objects and SID frames within the DTX signal segments in response to the activity information for the audio objects, the segment and frame detection comprising: (a) updating a global SID counter for invalid frames; and (b) signaling the detected SID frames within the DTX signal segments according to a value of the global SID counter; and encoding the signaled and detected SID frames using SID frame coding.
[0017] According to another aspect, the present disclosure relates to a device for discontinuous transmission (DTX) of audio objects in an object-based audio codec, the audio objects including respective audio streams, the device comprising: an analyzer of the audio stream for producing voice or signal activity information for the audio objects; a DTX controller for detecting DTX signal segments of the audio objects and SID frames within the DTX signal segments in response to the activity information for the audio objects, the DTX controller (a) updating a global SID counter for invalid frames and (b) signaling detected SID frames within the DTX signal segments according to the value of the global SID counter; and an encoder of the signaled detected SID frames using SID frame coding.
[0018] According to a further aspect, the present disclosure describes a method for decoding audio objects during discontinuous transmission (DTX) operation, each of the audio objects including an audio stream with an MD including at least one metadata (MD) parameter, the method comprising: decoding the metadata, the MD parameter comprising adjusting a value of the MD parameter to a smaller difference of the MD parameter between frames; and decoding the audio stream.
[0019] According to a fourth aspect, the present disclosure discloses a device for decoding audio objects during discontinuous transmission (DTX) operation, each of the audio objects including an audio stream with an MD including at least one metadata (MD) parameter, the device comprising: a metadata decoder for decoding the metadata, wherein the metadata decoder adjusts a value of the MD parameter to a smaller difference of the MD parameter between frames; and an audio stream decoder for decoding the audio stream.
[0020] The above and other objects, advantages, and features of (a) a method and device for discontinuous transmission (DTX) of audio objects in an object-based audio codec, and (b) a method and device for decoding audio objects during discontinuous transmission (DTX) operations, will become more apparent from the following non-limiting description of exemplary embodiments thereof, given by way of example only with reference to the accompanying drawings.
[0021] In the accompanying figure: [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a schematic block diagram illustrating a DTX transmission method and device simultaneously implemented in an ISM encoder and a corresponding ISM encoding method; [Figure 2] 2 is a flowchart illustrating the SID / DTX logic used in the DTX transmission method and device implemented in the ISM encoder and corresponding ISM encoding method of FIG. 1. [Figure 3] FIG. 1 is a diagram of a non-limiting example of the structure of a SID bitstream in an object-based audio codec. [Figure 4] 1 is a schematic block diagram illustrating an ISM decoder and a corresponding ISM decoding method simultaneously; [Figure 5] 10 is a graph of an example of metadata (azimuth angle) adjustment in DTX operation. [Figure 6] FIG. 1 is a simplified block diagram of an exemplary configuration of hardware components forming (a) a method and device for discontinuous transmission (DTX) of audio objects in an object-based audio codec, (b) an ISM encoding method and encoder, (c) a method and device for decoding audio objects during discontinuous transmission (DTX) operations, and / or (d) an ISM decoding method and decoder. DETAILED DESCRIPTION OF THE INVENTION
[0023] This disclosure describes methods and devices for discontinuous transmission (DTX) of audio objects in object-based audio codecs, as well as methods and devices for decoding audio objects during discontinuous transmission (DTX) operations.
[0024] Non-limiting, exemplary embodiments of methods and devices for discontinuous transmission (DTX) of audio objects in object-based audio codecs and methods and devices for decoding audio objects during discontinuous transmission (DTX) operations, as described below, involve concepts such as identifying invalid signal segments, decisions on coding SID frames and associated CN parameters, and coding, quantization, and recovery of metadata within invalid signal segments in object-based audio codecs.
[0025] In this disclosure, methods and devices for discontinuous transmission (DTX) of audio objects in object-based audio codecs, and methods and devices for decoding audio objects during discontinuous transmission (DTX) operations, are described with reference to the IVAS coding framework, referred to throughout this disclosure as the IVAS codec (or IVAS audio codec), by way of non-limiting example only. Specifically, the IVAS format below refers to the ISM format, the OMASA format, and the OSBA format. However, it is within the scope of this disclosure to incorporate such DTX techniques into any other audio codec that supports object-based audio.
[0026] 1. Introduction By way of non-limiting example, this disclosure considers a framework in which a fixed and constant codec total bitrate is considered for coding audio objects, including audio streams with their associated metadata, while supporting simultaneous coding of several audio objects (e.g., up to four audio objects). Note that, for non-narrative content, for example, metadata is not necessarily transmitted for at least some of the audio objects. Non-narrative sounds in movies, TV shows, and other videos are sounds that cannot be heard by the characters. A soundtrack is an example of non-narrative sounds because only the audience hears the music.
[0027] This disclosure also considers a basic, non-limiting example of input metadata consisting of two metadata (MD) parameters, namely, azimuth and elevation, stored per audio frame for each audio object. In this example, an azimuth range of [-180°, 180°) and an elevation range of [-90°, 90°) are considered. However, it is within the scope of this disclosure to consider only one or more than two MD parameters with different ranges.
[0028] 2. Object-Based Coding FIG. 1 is a schematic block diagram illustrating simultaneously a method 150 and a device 100 for discontinuous transmission (DTX) of audio objects in an object-based audio codec implemented within an ISM encoding method and a corresponding ISM encoder.
[0029] 2.1 Input Buffering 1, the ISM encoding method comprises an input buffering operation 151. To implement the input buffering operation 151, the ISM encoder comprises an input buffer 101.
[0030] The input buffer 101 buffers N input audio objects 102, i.e., N audio streams with their respective N associated metadata. The N input audio objects 102, including the N audio streams and the N metadata associated with each of these N audio streams, are buffered for one frame, e.g., a frame of length 20 ms. As is well known in the art of audio signal processing, audio signals are sampled at a given sampling frequency and processed by successive blocks of these samples, called "frames", which are each divided into a certain number of "sub-frames".
[0031] 2.2 Audio Stream Analysis and Front Pre-Processing 1, the DTX transmission method 150 comprises an operation of analyzing and pre-emptively pre-processing the N audio streams 153. To perform operation 153, the DTX transmission device 100 comprises an audio stream processor (analyzer) 103 for analyzing and pre-emptively pre-processing, e.g., sequentially, the N buffered audio streams transmitted from the input buffer 101 to the audio stream processor (analyzer) 103 over the N transport channels 104, respectively.
[0032] The analysis and look-ahead pre-processing operations 153 performed by the audio stream processor (analyzer) 103 may comprise at least one of the following sub-operations: time-domain transient detection, spectral analysis, long-term prediction analysis, pitch tracking and voicing analysis, voice or signal activity detection (VAD / SAD), bandwidth detection, noise estimation, and signal classification (which may include, in a non-limiting embodiment, (a) core encoder selection, e.g., from an ACELP core encoder, a TCX core encoder, an HQ core encoder, etc.; (b) signal type classification, e.g., among a disabled core encoder type, an unvoiced core encoder type, a voiced core encoder type, a generic core encoder type, a transition core encoder type, and an audio core encoder type, etc.; (c) speech / music classification, etc.). Information obtained from the analysis and look-ahead pre-processing 153 is supplied to the configuration and decision processor 106 via line 121. Examples of the above-mentioned sub-operations are described in Reference [1] with respect to the EVS codec and will not be described further in this disclosure.
[0033] 2.2.1 Overview of DTX / CNG operation As indicated herein above, discontinuous transmission (DTX) and comfort noise generation (CNG) are used in the method 150 and corresponding device 100 to reduce the transmission bit rate by simulating background noise during periods of inactive signal. In the EVS codec as described in Reference [1], the conventional DTX / CNG scheme is supported for bit rates up to 24.4 kbps. For higher bit rates, the EVS codec supports a less aggressive DTX / CNG scheme that switches to comfort noise generation (CNG) only for low input signal powers.
[0034] The reduction of the transmission bit rate during the invalid signal periods is achieved by coding parameters called comfort noise (CN) parameters. The CN parameters can be transmitted at a fixed or adaptive bit rate during the invalid signal periods. As an example, in the EVS codec as described in reference [1], the default transmission rate of CNG updates (also known as SID update rate) is set to 8 frames.
[0035] When the EVS codec operates in the DTX / CNG mode, a signal activity detector (SAD) is used to analyze the input audio signal and determine whether the input audio signal is valid or invalid (see SAD decision in section 5.1.12 of [1]). Based on that analysis, the SAD detector generates the SAD flag f SAD , whose state is determined when the input audio signal is valid (f SAD =1) or just background noise (f SAD = 0). SAD When f = 1, normal encoding and decoding is performed, as in the default option. SAD = 0, the DTX function is performed in the encoder that transmits either Silence Insertion Descriptor (SID) frames or NO_DATA frames. In the following description, it is assumed that the input audio signal is valid (f SAD =1) frames are called "valid frames" and are accompanied by only background noise (f SAD =0) frames are called "invalid frames".
[0036] SID frames contain CN parameters, which are used to update the characteristics of the background noise in the decoder, while NO_DATA frames are empty. SID frames in EVS are always coded using 48 bits (which corresponds to a bit rate of 2.4 kbps).
[0037] As mentioned above, the multi-channel IVAS codec is based on the EVS codec. Specifically, a variable bitrate version of the EVS encoder is used as the so-called core encoder. The SID bitrate of the core encoder in IVAS is the same as that of EVS, i.e., 2.4 kbps, but the CNG algorithm of the EVS codec is reused as much as possible in the core encoder of the IVAS codec.
[0038] Generally, the number of core encoders in IVAS is at most the number of input audio channels (i.e., every single audio channel is coded by an associated core encoder). In the case of ISM formats, the number of core encoders N is equal to the number of independent audio streams (ISMs), usually accompanied by metadata, N, and it is then called "discrete ISM" coding. Note that "parametric ISM" techniques, in which the number of core encoders is less than the number of audio objects, can also be used. These "parametric ISM" techniques are usually used at low bitrates where effective channel coding of all audio objects is not possible due to bitrate constraints.
[0039] Coding all ISMs in a SID frame results in a bitrate of N x (2.4 kbps + metadata), which obviously also results in a too high IVAS SID bitrate. To keep it reasonably low, the IVAS SID bitrate is set to 5.2 kbps and an efficient coding of SID frames is introduced.
[0040] 2.2.2 Classification of Invalid Frames in ISM Encoders As indicated herein above, the audio stream processor (analyzer) 103 analyzes the audio stream carried by each channel 104 from the input buffer 101, classifies the input audio objects, and generates voice or signal activity detection (VAD / SAD) flags f, one for each audio object. SAD produces f SAD =1 to f SAD =0 indicates the start of an invalid signal segment for a particular audio object. Some variants of the VAD / SAD flags can be used, for example the "local VAD" flag as described in paragraph 5.1.12 of reference [1]. It is clear that the start of an invalid signal segment usually occurs at different times for different audio objects.
[0041] By default, in an object-based audio system, when an invalid signal segment is declared for all audio objects, a DTX signal segment is declared. An additional classification stage is then used to classify the invalid signal segment. Also, the value of the metadata influences whether a frame is declared as an invalid frame. The logic for classifying an invalid (SID or NO_DATA) frame or a valid frame is shown in Figure 2.
[0042] 2.2.3 Global SID counter, DTX flag, SID flag FIG. 2 is a flow chart illustrating simultaneously a DTX control operation 250 and corresponding DTX controller 200 used in the DTX transmission method 150, and a device 100 implemented in the ISM encoder and corresponding ISM encoding method of FIG.
[0043] Referring to FIG. 2, to control the SID / NO_DATA frame decision and the start of the DTX signal segment, the DTX controller 200 counts a global SID counter cnt for invalid frames. SID Use this global SID counter cntSID controls the SID update rate and allows synchronizing SID / NO_DATA frames across all audio objects, resulting in a global SID counter cnt SID allows efficient adjustment of possible residual (hysteresis) DTX logic, allowing other than the default SID update rate, for example 8 frames or SID adaptive update rate.
[0044] In the disclosed logic, therefore, the global SID counter cnt SID (parameter "cnt_SID_ISM" in the source code at the end of this disclosure) is superior to the SID counter per audio object, which essentially guarantees that the SID counters per individual audio objects are synchronized.
[0045] 2, the DTX controller 200 receives information from line 121 resulting from the analysis and pre-processing operation 153, including VAD information. The DTX controller 200 then sets the DTX flag DTX and SID flag SID is initialized to "0" (block 201).
[0046] By default, the DTX Controller 200 sets the VAD flag for all audio objects. VAD When the DTX flag is equal to 0, the DTX controller 200 detects a DTX signal segment (SID or NO_DATA frame) of the audio object (block 202). DTX =1 (parameter "dtx_flag" in the source code at the end of this disclosure) VAD is equal to 0 (block 203). This can be expressed using the relationship (1).
[0047]
number
[0048] where flag DTX is the DTX flag as mentioned before. The VAD flag for all audio objects VAD When ≠ 0 (block 202), the DTX controller 200 signals a valid frame (block 210) and selects a valid frame coding.
[0049] Furthermore, the DTX controller 200 sets the SID flag SID (The parameter "sid_flag" in the source code at the end of this disclosure) is used to signal the SID frame within the DTX signal segment. SID flag SID (a) global SID counter cnt SID is set to 1 (block 205) when it is equal to 0 (block 204), in which case the DTX controller 200 signals an SID frame in the DTX signal segment (block 207), selects an SID frame coding for the signaled SID frame, and (b) sets the global SID counter cnt SID is set to 0 when it is not equal to 0 (block 204), in which case the DTX controller 200 signals a NO_DATA frame in the DTX signal segment (block 209) and selects a NO_DATA frame coding for the signaled NO_DATA frame. This can be expressed using relationship (2).
[0050]
number
[0051] In block 206 of FIG. 2, the DTX controller 200 counts the SID counter cnt SIDThe counter cnt is reset to -1 and is incremented by 1 for every invalid frame up to a value corresponding to the SID update rate (which is 8 frames by default in the example implementation above). SID When the value of reaches the SID update rate, it is reset to 0 and the DTX controller 200 signals an SID frame within the DTX signal segment (blocks 204, 205, and 207) and selects an SID frame coding for the signaled SID frame.
[0052] The DTX controller 200 uses another classification stage to determine the DTX flags based on the core encoder preprocessing values for all audio objects. DTX can be changed (block 208).
[0053] As a non-limiting example, the DTX controller 200 may calculate: 1) the average value of the LT (long-term) background noise across all audio objects, mean is higher than the first threshold β1, or 2) the average value of the LT background noise across all audio objects mean is higher than the second threshold β2, and the LT background noise variation noise_var across all audio objects mean is higher than the third threshold β3, the valid frame coding (flag DTX =0) and logic to force the selection of a valid frame coding (block 210). This can be expressed using relationship (3): noise mean >β1 or (noise mean > β2 and noise_var mean >β3), flag DTX =0 (3)
[0054] The thresholds are found empirically and can be set, for example, as β1 = 50, β2 = 10, and β3 = 2. To check the similarity of the background noise among all audio objects, an estimate of the LT background noise variance is introduced. The parameter noisemean represents the average value of the long-term background noise energy of all audio objects (see paragraph 5.1.11 in [1]). Then, the parameter noise_var mean represents the variation of the long-term background noise energy values of all audio objects.
[0055] Also, similar logic to that in the EVS codec is used in the IVAS codec; for higher bit rates, the IVAS codec supports a less aggressive DTX / CNG scheme, switching to CNG only for low input signal powers.
[0056] As another non-limiting example of determining an invalid signal segment, the DTX controller 200 includes logic to evaluate the energy of background noise in all audio channels 104 (audio streams). Typically, when there is one audio object dominated by high-energy background noise and other audio objects have low-energy background noise, the DTX flag is set to 1 (flag DTX =1), otherwise it is set to 0 (flag DTX =0), which means that when there is more than one audio object with high energy noise, DTX is not triggered (effective frame coding (block 210) is selected).
[0057] 2.3 Metadata Analysis, Quantization, and Coding 1 for coding an object-based audio signal further comprises a metadata analysis, quantization, and coding operation 155. To perform operation 155, the ISM encoder for coding an object-based audio signal comprises a metadata processor 105.
[0058] The analysis, quantization, and coding of metadata in non-DTX operations may be performed, for example, as described in reference [3], the entire contents of which are incorporated herein by reference. In the described non-limiting exemplary embodiment, the metadata processor 105 of Figure 1 quantizes and codes the metadata 140 of N audio objects 102 sequentially in a loop, while it should be noted that certain dependencies may be exploited between the quantization of the audio objects and the metadata parameters of these audio objects.
[0059] As indicated herein above, in this disclosure, two metadata parameters, azimuth angle and elevation angle (as contained in the N input metadata), are considered in an exemplary implementation. By way of non-limiting example, the metadata processor 105 includes a quantizer (not shown) for the following metadata parameter indexes that uses the following exemplary strategy to reduce the number of bits used: - Azimuth parameter: The 12-bit azimuth parameter index from the input metadata file is B az The bit index (for example, B az =7). Given the lower and upper bounds of the azimuth angle (-180° and +180°), (B az The quantization step of the (=7)-bit uniform scalar quantizer is 2.835°. - Elevation parameters: The 12-bit elevation parameter index from the input metadata file is el The bit index (for example, B el =6). Given the lower and upper bounds of the elevation angle (-90° and +90°), (B el The quantization step of a 6-bit uniform scalar quantizer is 2.857°.
[0060] Once both the azimuth and elevation parameter indices are quantized, they can be coded by a metadata encoder (not shown) of the metadata processor 105 (112 in FIG. 1) using either absolute coding or differential coding. As known in the art, absolute coding means that the current value of the parameter is coded. Differential coding means that the difference between the current and previous value of the parameter is coded. Since the indices of the azimuth and elevation parameters usually evolve smoothly (i.e., the change in azimuth or elevation position can be considered to be continuous and smooth), differential coding is used by default because it consumes fewer bits compared to absolute coding.
[0061] 2.3.1 Metadata Analysis, Quantization, and Coding in SID Frames As indicated herein above, the amount of bits in an SID frame is relatively small. For the IVAS codec, the SID bit rate is 5.2 kbps, of which roughly half is reserved for the metadata (MD) payload. Some compromise is made in order to transmit as many MD values as possible.
[0062] This disclosure is based on forcing the MD coding by the metadata encoder (not shown) of the metadata processor 105 in SID frames to be an absolute coding method. Note that this is different from the MD coding in valid frames, where absolute / differential coding is utilized. The motivation for using only absolute coding in SID frames is to prevent possible degradation in long segments of invalid frames due to loss of SID frames in the case of a noisy channel when differential coding is used.
[0063] Then, one possibility to meet the bit amount constraints is to reduce the resolution of the MD value. However, this means that the reconstruction of MD at the decoder is not accurate enough, causing subjective quality degradation. Therefore, this disclosure introduces several mechanisms to overcome these constraints.
[0064] First, the DTX controller 200 decomposes the MD value according to the number of audio objects. Specifically, when the number of audio objects is small, the amount of available bits for coding metadata (MD) is relatively generous, so the resolution of the MD is kept relatively high (e.g., the same as that in a valid frame). On the other hand, when the number of coded audio objects is large, the resolution of the MD value is relatively low.
[0065] In an exemplary implementation where azimuth and elevation MD parameters are considered, the metadata processor 105: a) For example, in a system with one or two audio objects, B az = 8 bits and B el = 7 bits for azimuth and elevation indexes, b) For example, in a system with three or four audio objects, B az = 6 bits and B el = 5 bits for azimuth and elevation indexes , but the actual number of coded bits, and therefore the resolution, is explicitly known from the ISM common signaling (see 113 in Figure 1 and paragraph 2.8 below).
[0066] Second, to keep the amount of SID bits as small as possible, a saving in the amount of MD bits is achieved by computing one flag per audio object that indicates that the MD parameters have not changed (or have not changed significantly) since the last frame, and consequently indicates that no metadata (MD) parameters for that particular audio object will be coded or transmitted. Similarly, this flag serves as an indication that no input MD parameters are present for that particular audio object.
[0067] Referring back to FIG. 2, in one example implementation, the DTX controller 200 sets the MD on / off flag for all MD parameter values. MD Calculate (block 213) the flag (parameter "diff_flag" in the source code at the end of this disclosure). MD,θ is calculated for the azimuthal MD parameters using the following relation (4):
[0068]
number
[0069] where θ is the azimuth angle of the current frame, and θ last is the azimuth angle of the last frame, and δ θ is the maximum difference value of the azimuth angle, for example, δ θ = 15. In the same manner (see relation (4)), the DTX controller (block 213) sets the flag
[0070]
number
[0071] Calculate the last MD on / off flag for one audio object. MDcan be obtained as the logical OR between all the specific MD flags (the symbol V in relation (5) below).
[0072]
number
[0073] Finally, a flag is given for each audio object, represented as a single bit of information. MD 214 is inserted into the bitstream in the SID frame 207 immediately after the ISM signaling of the N coded audio objects 301 (see paragraph 2.8.1 below).
[0074] Third, the DTX controller determines the amount of bits for metadata quantization. MD It should be noted that at this stage, only the amount of bits for quantization of the metadata values is estimated (pre-calculated), and the quantization itself may only be performed later. MD is the maximum amount of available bits for MD coding. available If it is higher (block 212), the flag DTX is reset to 0 (block 215) and valid frame coding (block 210) is performed. Meanwhile, the estimated MD bit amount bits MD is the maximum amount of available bits for MD coding. available If (block 212) the flag DTX remains unchanged and the DTX coding segment continues, while the maximum available bit amount for MD coding is available is calculated as the difference between the codec SID bit amount and the amount of bits required for non-MD coding in the SID frame (e.g., SID frame signaling, core encoder SID bit amount, spatial information bit amount, ISM signaling).
[0075] This logic comes from the premise that CNG frames do not represent important audio details, but metadata (MD) is considered to be important details for representing the output audio and is therefore transmitted to the decoder. When there are more audio objects, the probability of switching to valid frame coding (block 210) is obviously higher. This logic also becomes more important when more metadata (not just azimuth and elevation) exists and needs to be coded.
[0076] 2.4.1 Bitrate and Decision for Each Channel Configuration 1, the ISM encoding method comprises an operation 156 of configuring and deciding on a bit rate per transport channel 104. To perform operation 156, the ISM encoder comprises a configuration and decision processor 106 that forms a bit budget allocator.
[0077] The configuration and decision processor 106 (here after the bit amount allocator 106) can use a bit rate adaptation algorithm to distribute the available bit amount for core encoding the N audio streams over the N transport channels 104. Details of the bit adaptation algorithm for distributing the available bit amount for core encoding can be found in reference [3].
[0078] 2.5 Preprocessing 1, the ISM encoding method comprises an operation 158 of pre-processing N audio streams carried from the configuration and decision processor 106 (bit budget allocator) over N transport channels 104. To perform operation 158, the ISM encoder comprises a pre-processor 108.
[0079] Once the configuration and bit rate distribution among the N audio streams has been completed by the configuration and decision processor 106 (bit budget allocator), the pre-processor 108 performs sequential further pre-processing 158 on each of the N audio streams. Such pre-processing 158 may include, for example, further signal classification, further core encoder selection (e.g., selection between an ACELP core, a TCX core, and an HQ core), a different internal sampling frequency F adapted to the bit rate to be used for core encoding, s Other resampling at 100 kHz, etc. Examples of such pre-processing can be found, for example, in reference [1] for the EVS codec, and will not be described further in this disclosure.
[0080] 2.6 Core Encoding 1, the ISM encoding method comprises a core encoding operation 159. To perform operation 159, the ISM encoder 100 comprises N core encoders 109 for coding the N audio streams, e.g., for coding the N audio streams carried from the pre-processor 108 over the N transport channels 104, respectively.
[0081] Specifically, for valid frame coding, the N audio streams are encoded using N variable bit rate core encoders 109, e.g., mono core encoders. The bit rate used by each of the N core encoders is the bit rate selected by the configuration and decision processor 106 (bit budget allocator) for the corresponding audio stream. For example, an EVS-based core encoder as described in Reference [1] may be used as the core encoder 109.
[0082] In case of DTX operation (SID frame coding), the SID of an audio object is coded by one of the core encoders 109, while processing in the other core encoders 109 may be done in NO_DATA frame operation (the other encoders only update the core encoder state parameters). Alternatively, no processing is performed in said other core encoders 109, i.e. pre-processing and core coding are completely skipped in said other core encoders 109.
[0083] 2.7 Coding spatial information 1, the ISM encoding method comprises an operation 180 of coding spatial information, and the ISM encoder comprises a corresponding coding module 130. The spatial information analyzes similarities between audio objects and may be based on inter-channel cues, such as inter-channel level differences or inter-channel coherence, for example.
[0084] Spatial information is usually estimated and coded in parametric ISM coding techniques or in SID frame coding. An example can be found in the IVAS framework described in reference [4], the entire contents of which are incorporated herein by reference.
[0085] 2.8 SID Bitstream Structure 1, the ISM encoding method comprises a multiplexing operation 160. To perform operation 160, the ISM encoder comprises a multiplexer 110.
[0086] Figure 3 is a schematic diagram illustrating the structure of the SID bitstream 111 produced by the multiplexer 110 and transmitted for a frame from the ISM encoder of Figure 1 to the ISM decoder 400 of Figure 4. The structure of the SID bitstream 111 may be structured as shown in Figure 3, regardless of whether metadata is present and transmitted.
[0087] Referring to Figure 3, the multiplexer 110 writes the index 302 of the SID format signaling followed by the index 114 of one core encoder SID from the beginning of the bitstream 111, while the index of the ISM common signaling 113 from the configuration and decision processor 106 (bit budget allocator), spatial information 131 from the coding module 130, and metadata 112 from the metadata processor 105 are written from the end of the bitstream 111.
[0088] 2.8.1 ISM common signaling in SID frames The multiplexer writes the ISM common signaling 113 from the end of the bitstream 111. The ISM common signaling in the SID frame is produced by the configuration and decision processor 106 (bit budget allocator) and comprises a variable number of bits representing:
[0089] (a) N Audio Objects: The signaling for the N coded audio objects present in the bitstream 111 is, for example, in the form of a unary code 301 with stop bits (e.g., for N=3 audio objects, the first three bits of the ISM common signaling are “110” written in reverse order).
[0090] (b) a metadata on / off flag, one per audio object, as described in relation (5) and represented by line 214; MD .
[0091] 2.8.2 Metadata Payload In the SID frame immediately following the ISM common signaling 113, in the reverse order, a) the spatial information index 131 and finally b) the metadata value 112 as quantized in paragraph 2.3.1 above are written into the bitstream 111.
[0092] 2.8.3 Audio Stream Payload In a valid frame, the multiplexer 110 receives the N audio streams 114 coded by the N core encoders 109 over the N transport channels 104 and writes them from the beginning of the bitstream 111 immediately after the IVAS format bits 302 (see FIG. 3).
[0093] In a SID frame, the multiplexer 110 receives one audio stream 114 coded by one of the core encoders 109 through one transport channel 104 and writes it from the beginning of the bitstream 111 immediately after the IVAS SID format bit 302 (signaling the SID mode for the IVAS format) (see Figure 3).
[0094] 3. Decoding Audio Objects FIG. 4 is a schematic block diagram illustrating simultaneously an ISM decoding method 450 implementing a method for decoding audio objects during discontinuous transmission (DTX) operation and a corresponding ISM decoder 400 implementing a device for decoding audio objects during discontinuous transmission (DTX) operation.
[0095] 3.1 Demultiplexing 4, the ISM decoding method 450 comprises a demultiplexing operation 451. To perform operation 451, the ISM decoder 400 comprises a demultiplexer 401.
[0096] Demultiplexer 401 receives bitstream 402 transmitted from the ISM encoder of Figure 1 to ISM decoder 400 of Figure 4. Specifically, bitstream 402 of Figure 4 corresponds to bitstream 111 of Figure 1.
[0097] The demultiplexer 401 extracts from the bitstream 402 (a) the N coded audio streams 114 in the case of valid frame coding, or one audio stream in the case of SID frames, (b) the coded metadata 112 for the N audio objects, (c) the spatial information 131, and (d) the ISM common signaling 113 read from the end of the received bitstream 402.
[0098] 3.2 Metadata Dequantization and Decoding 4, the ISM decoding method 450 comprises a metadata decoding and inverse quantization operation 453. To perform operation 453, the ISM decoder 400 comprises a metadata decoding and inverse quantization processor (metadata decoder) 403.
[0099] The Metadata Decoding and Inverse Quantization Processor (Metadata Decoder) 403 provides an output setup 404 for decoding and inverse quantizing the coded metadata 112 for the N transmitted audio objects, the ISM common signaling 113, and the metadata for audio streams / objects with valid content. The output setup 404 is command line parameters for the M decoded audio objects / transport channels and / or audio formats, which may be equal to or different from the N coded audio objects / transport channels. The Metadata Decoding and Inverse Quantization Processor (Metadata Decoder) 403 produces decoded metadata 405 for the M audio objects / transport channels and provides information about the decoded metadata and their respective bit amounts on line 406. Obviously, the decoding and inverse quantization performed by the processor (Metadata Decoder) 103 is the inverse of the quantization and coding performed by the Metadata Processor 105 of FIG. 1.
[0100] 3.2.1 Metadata adjustment during DTX operation In the ISM decoder 400, SID frames are received at a certain rate (e.g., 8 frames by default in IVAS), which means that the received MD parameter values may change in large steps between SID frames. These large steps may cause subjective artifacts. For example, the position of an audio object may suddenly change from one location to another. To avoid these artifacts, the metadata decoding and dequantization processor (metadata decoder) 403 adjusts the MD parameter values in the ISM decoder 400 so that the difference in MD parameter values between frames is small. For example, interpolation between the true decoded and dequantized MD parameter values of the current frame and the MD parameter values of the previous frame may be applied in several frames after the SID frame. This allows the MD parameter values to evolve more smoothly while smoothing is applied in several CNG frames, or several valid frames, or several CNG and valid frames after the SID frame.
[0101] For example, adjustment or smoothing may be applied to each MD parameter so that the maximum difference (step) of the MD parameter between two adjacent frames is not larger (smaller) than a given threshold. dec The maximum smoothing step of the difference in azimuth angle is Δ θ Then,
[0102]
number
[0103] is.
[0104] where θ true is the quantized azimuth angle (generally the quantized MD parameter value) transmitted in the SID frame, and θ last is in the previous frame and θ last =θ decis the azimuth angle (generally, the MD parameter value) updated at the end of each frame decoding as Δ θ = 5. Furthermore, the action sgn(x) in relation (6) is expressed mathematically as
[0105]
number
[0106] is.
[0107] Furthermore, the maximum number of frames for applying the smoothing step is given by the threshold D max For example, in one example implementation, it is not limited in invalid segments, but in valid segments, in that example implementation, D max = 5 (constant "IVAS_ISM_DTX_HO_MAX" in the source code at the end of this disclosure) frames.
[0108] Moreover, the quantized MD parameter values θ of the current frame true and the decoded MD parameter value θ of the previous frame last The absolute value of the difference between θ Threshold D max For example, for the azimuth parameter, it is |θ true -θ last |>D max Δ θ If , skip smoothing (8) means.
[0109] Relationship (8) also applies to MD parameters other than azimuth. Relationship (8) is valid when the true value of the MD parameter changes significantly from frame to frame and the number of frames between the last SID frame and the valid frame is small (typically D max is used to prevent erroneous smoothed MD parameter estimation when
[0110] The effect of smoothing on the MD azimuth angle parameter of the first audio object when coding two audio objects with DTX operation in the IVAS codec can be seen in the graph in Figure 5. From top to bottom, for the second segment of 1.1, there are: a) the noisy segment of the input audio signal, b) the CNG synthesis, c) the IVAS total bit rate (48 kbps) indicating a valid frame, IVAS total bit rate (5.2 kbps) indicating a SID frame, and IVAS total bit rate (0 bps) indicating a NO_DATA frame, d) the encoder input azimuth angle, e) the decoder (quantized) azimuth angle for valid frame coding (without DTX), f) the decoder azimuth angle for the CNG segment without smoothing, and g) the decoder azimuth angle for the CNG segment with smoothing.
[0111] 3.3 Bitrate Configuration and Decisions 4, the ISM decoding method 450 comprises a configuration and decision operation for per-channel bitrate 457. To perform operation 457, the ISM decoder 400 comprises a configuration and decision processor 407 (bit budget allocator).
[0112] The bit budget allocator 407 receives information about the respective bit budgets for the M decoded metadata on line 406 to determine the core decoder bit rate for each audio stream. The bit budget allocator 407 uses the same procedure as in the bit budget allocator 106 of Figure 1 to determine the core decoder bit rate (see section 2.4). In the case of an SID frame, the core decoder SID bit rate is allocated to one core decoder, while a bit rate of 0 kbps is allocated to the other core decoders. Obviously, in the case of a NO_DATA frame, a bit rate of 0 kbps is allocated to all core decoders.
[0113] 3.4 Decoding spatial information 4, the ISM decoding method 450 comprises an operation of decoding spatial information 458. The ISM decoder 400 comprises a spatial information decoding module 408 for performing operation 458.
[0114] Operation 458 is responsible for decoding spatial information 409 used in the parametric ISM technique or in the SID frame. Decoding 458 is the inverse of coding 180 in FIG.
[0115] 3.5 Core Decryption 4, the ISM decoding method 450 comprises a core decoding operation 460. To perform operation 460, the ISM decoder 400 comprises N audio stream decoders 410, including N core decoders 410, for example N variable bit rate core decoders.
[0116] N audio streams 114 from the demultiplexer 401 are decoded and sequentially decoded in N variable bit rate core decoders 410 at respective core decoder bit rates 411, such as determined by the bit rate allocator 407. When the number M of decoded audio objects, as required by the output setup 404, is less than the number of transport channels, i.e., when M < N, fewer core decoders 410 are used. Similarly, in such cases, not all metadata payloads may be decoded.
[0117] In response to the N audio streams 114 from the demultiplexer 401, the core decoder bit rate 411, such as determined by the bit rate allocator 407, and the output setup 404, the core decoder 410 produces M decoded audio streams 412 on respective M transport channels.
[0118] In the case of an SID frame, one audio stream 114 from the demultiplexer 401 is decoded and supplied to one of the core decoders 410. The other core decoders 410 are supplied with the SID and spatial information 409 of the first core encoder.
[0119] 3.6 Audio Channel Rendering The ISM decoding method 450 may include an operation 463 for audio channel rendering. To perform the operation 463, the ISM decoder 400 includes a renderer 413 for audio objects.
[0120] Considering the output setup 415 indicating the number and content of the output audio channels to be produced, the renderer 413 converts the M decoded metadata 405 and the M decoded audio streams 412 into a number of output audio channels 414. Again, the number of output audio channels 414 may be equal to or different from the number M.
[0121] The renderer 413 can be designed in a variety of different configurations to obtain the desired output audio channels, for which reason the renderer will not be further described in this disclosure.
[0122] 4. Exemplary Configurations of Hardware Components FIG. 6 is a simplified block diagram of an exemplary configuration of hardware components that form (a) a method and device for discontinuous transmission (DTX) of audio objects in an object-based audio codec, (b) an ISM encoding method and encoder, (c) a method and device for decoding audio objects during discontinuous transmission (DTX) operations, and / or (d) an ISM decoding method and decoder (hereinafter collectively "encoding and decoding methods and devices").
[0123] The encoding and decoding method and device may be implemented as part of a mobile terminal, as part of a portable media player, or in any similar device. The encoding and decoding method and device (identified as 600 in FIG. 6) comprises an input 602, an output 603, a processor 601, and a memory 604.
[0124] The input 602 is configured to receive an input signal. The output 603 is configured to provide an output signal. The input 602 and the output 603 may be implemented in a common module, for example, a serial input / output device.
[0125] The processor 601 is operatively connected to an input 602, an output 603, and a memory 604. The processor 601 is implemented as one or more processors for executing code instructions that underpin the above-described encoding and decoding methods and functions of the various elements and operations of the device as shown in the accompanying drawings and / or as described in this disclosure.
[0126] Memory 604 may comprise non-transitory memory for storing code instructions executable by processor 601, specifically processor-readable memory that stores non-transitory instructions that, when executed, cause the processor to perform encoding and decoding methods and device elements and operations. Memory 604 may also comprise random access memory or buffers for storing intermediate processing data from various functions performed by processor 601.
[0127] Those skilled in the art will recognize that the descriptions of the encoding and decoding methods and devices are merely exemplary and are not intended to be limiting in any way. Other embodiments will readily suggest themselves to those skilled in the art given the benefit of this disclosure. Furthermore, the disclosed encoding and decoding methods and devices can be customized to provide valuable solutions to existing needs and problems of encoding and decoding audio signals.
[0128] For the sake of clarity, not all routine features of implementations of encoding and decoding methods and devices are shown and described. Of course, it will be understood that in developing any such actual implementation of encoding and decoding methods and devices, numerous implementation-specific decisions may need to be made to achieve the developer's specific goals, such as meeting application-related, system-related, network-related, and business-related constraints, and that these specific goals will vary from implementation to implementation and from developer to developer. Moreover, it will be understood that the developer's efforts may be complex and time-consuming, but would nevertheless be a routine undertaking for one of ordinary skill in the art of audio processing having the benefit of this disclosure.
[0129] In accordance with this disclosure, the elements, processing operations, and / or data structures described herein may be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. In addition, those skilled in the art will recognize that less general-purpose devices, such as hardwired devices, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., may also be used. When a method comprising a series of operations and sub-operations is implemented by a processor, computer, or machine, and the operations and sub-operations may be stored as a series of non-transitory code instructions readable by a processor, computer, or machine, they may be stored on a tangible and / or non-transitory medium.
[0130] The elements and processing operations of the encoding and decoding methods and devices as described herein may comprise software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.
[0131] In the encoding and decoding methods and devices, the various processing operations and sub-operations may be performed in different orders, and some of the processing operations and sub-operations may be optional.
[0132] Although the present disclosure has been described hereinabove by way of its non-limiting exemplary embodiments, these embodiments may be freely modified within the scope of the appended claims without departing from the spirit and nature of the present disclosure.
[0133] 5. References This disclosure refers to the following references, the entire contents of which are incorporated herein by reference: (References)
[0134] Source code Below is a code excerpt implementing the present disclosure in the IVAS audio codec framework: / *---------------------------------------------------------------------* * ivas_ism_enc() * * ISM metadata + CoreCorders encoding routines *-------------------------------------------------------------------* / ivas_ism_enc( ) { ... / *------------------------------------------------------------------* *DTX analysis *---------------------------------------------------------------* / if ( st_ivas->hEncoderConfig->Opt_DTX_ON ) { / * Analysis and decision on DTX * / dtx_flag = ivas_ism_dtx_enc( ... ); } / *------------------------------------------------------------------* * Analysis of objects, configurations, and decisions about bitrate per channel * Metadata quantization and encoding *-----------------------------------------------------------------* / if ( dtx_flag ) { ivas_ism_metadata_sid_enc( ... ); } else if ( st_ivas->ism_mode == ISM_MODE_PARAM ) { ivas_ism_compute_noisy_speech_flag( ... ); ivas_ism_metadata_enc( ... ); } else / * ISM_MODE_DISC * / { ivas_ism_metadata_enc( ... ); } update_last_metadata( ... ); / *----------------------------------------------------------------* * Write IVAS format signaling into SID frame *----------------------------------------------------------------* / if ( sid_flag ) { ivas_write_format_sid ... ); } ... } / *-------------------------------------------------------------------* * ivas_ism_get_dtx_enc() * * Analysis and judgment on DTX in ISM format *-------------------------------------------------------------------* / / *! r: DTX frame indication * / int16_t ivas_ism_dtx_enc( ISM_DTX_HANDLE hISMDTX, SCE_ENC_HANDLE hSCE[MAX_SCE], const int16_t num_obj, const int16_t nchan_transport, int16_t vad_flag[MAX_NUM_OBJECTS], ISM_METADATA_HANDLE hIsmMeta[], int16_t md_diff_flag[], int16_t *sid_flag ) { ... / *------------------------------------------------------------------* * Calculate global ISM DTX flag *-----------------------------------------------------------------* / / * Calculate global ISM based on local VAD * / dtx_flag = 1; for ( ch = 0; ch < num_obj; ch++ ) { dtx_flag &= !vad_flag[ch]; } / * Calculate global ISM based on long-term background noise * / / * One of the channels is enabled -> no DTX * / for ( ch = 0; ch < num_obj; ch++ ) { lp_noise[ch] = hSCE[ch]->hCoreCoder[0]->lp_noise; } noise_var = var( lp_noise, num_obj ); noise_mean = mean( lp_noise, num_obj ); if( noise_mean > BETA1 || (noise_mean > BETA2 && noise_var > BETA3)) { dtx_flag = 0; } / *------------------------------------------------------------------* * Reset the bitstream *-----------------------------------------------------------------* / if ( dtx_flag ) { reset_indices_enc( hSCE[0]->hCoreCoder[0]->hBstr, MAX_NUM_IND ); } / *------------------------------------------------------------------* * Determine if SID metadata should be sent (per object) * Estimate MD bit consumption *-----------------------------------------------------------------* / if ( dtx_flag ) { ivas_get_ism_sid_quan_bitbudget( num_obj, &nBits_azimuth, &nBits_elevation, &nBits_ener, &nBits_coh ); nBits = 0; for ( ch = 0; ch < num_obj; ch++ ) { / * Check the difference between the current metadata and the last metadata * / md_diff_flag[ch] = 0; if ( fabsf( hIsmMeta[ch]->azimuth - hIsmMeta[ch]->last_azimuth ) > MD_MAX_DIFF_AZIMUTH ) { md_diff_flag[ch] = 1; } if ( fabsf( hIsmMeta[ch]->elevation - hIsmMeta[ch]->last_elevation ) > MD_MAX_DIFF_ELEVATION ) { md_diff_flag[ch] = 1; } / * Estimate the amount of bits in SID metadata * / nBits++; / * Number of objects * / nBits++; / * SID metadata flags * / if ( md_diff_flag[ch] == 1 ) { nBits += nBits_azimuth; nBits += nBits_elevation; } } / * Calculate the maximum amount of available MD bits * / nBits_MD_max = ( IVAS_SID_5k2 - SID_2k40 ) / FRAMES_PER_SEC; nBits_MD_max -= SID_FORMAT_NBITS; for ( ch = 0; ch < nchan_transport - 1; ch++ ) { nBits_MD_max -= nBits_ener; nBits_MD_max -= nBits_coh; } / * Too many metadata bits -> switch to valid coding * / if ( nBits > nBits_MD_max ) { dtx_flag = 0; } } / *------------------------------------------------------------------* * Set core_brate for all channels * Get the "sid_flag" value *-----------------------------------------------------------------* / *sid_flag = 0; if ( !dtx_flag ) { / * At least one of the channels is enabled -> no DTX * / for ( ch = 0; ch < num_obj; ch++ ) { hSCE[ch]->hCoreCoder[0]->core_brate = -1; } hISMDTX->cnt_SID_ISM = -1; / * IVAS format signaling was cleared in dtx() * / if ( hSCE[0]->hCoreCoder[0]->hBstr->nb_bits_tot == 0 ) { push_indice( hSCE[0]->hCoreCoder[0]->hBstr, IND_IVAS_FORMAT, 2 / * == ISM format* / , IVAS_FORMAT_SIGNALING_NBITS ); } } else / * ism_dtx_flag == 1 * / { for ( ch = 0; ch < num_obj; ch++ ) { hSCE[ch]->hCoreCoder[0]->cng_type = FD_CNG; } / * * Update the global SID counter * / hISMDTX->cnt_SID_ISM++; if ( hISMDTX->cnt_SID_ISM >= hSCE[0]->hCoreCoder[0]->hDtxEnc->max_SID ) { / * Adaptive SID Update Interval * / hSCE[0]->hCoreCoder[0]->hDtxEnc->max_SID = hSCE[0]->hCoreCoder[0]->hDtxEnc->interval_SID; hISMDTX->cnt_SID_ISM = 0; } / * Encode the SID in only one channel * / for ( ch = 0; ch < num_obj; ch++ ) { hSCE[ch]->hCoreCoder[0]->core_brate = FRAME_NO_DATA; } if ( hISMDTX->cnt_SID_ISM == 0 ) { hSCE[hISMDTX->sce_id_dtx]->hCoreCoder[0]->core_brate = SID_2k40; *sid_flag = 1; } } if ( dtx_flag == 1 && *sid_flag == 0 ) { set_s( md_diff_flag, 0, num_obj ); } return dtx_flag; } / *-------------------------------------------------------------------* * ivas_ism_metadata_sid_enc() * * Quantize and encode ISM metadata in SID frames *-------------------------------------------------------------------* / void ivas_ism_metadata_sid_enc( ISM_DTX_HANDLE hISMDTX, const int16_t num_obj, const int16_t nchan_transport, ISM_METADATA_HANDLE hIsmMeta[], const int16_t sid_flag, const int16_t md_diff_flag[], BSTR_ENC_HANDLE hBstr, int16_t nb_bits_metadata[] ) { ... if ( sid_flag ) { nBits = ( IVAS_SID_5k2 - SID_2k40 ) / FRAMES_PER_SEC; nBits -= SID_FORMAT_NBITS; nBits_start = hBstr->nb_bits_tot; / *-------------------------------------------------------------* * Write ISM common signaling *-----------------------------------------------------------* / / * Write the number of objects - unary coding * / for ( ch = 1; ch < num_obj; ch++ ) { push_indice( hBstr, IND_ISM_NUM_OBJECTS, 1, 1 ); } push_indice( hBstr, IND_ISM_NUM_OBJECTS, 0, 1 ); / * Write SID metadata flags (one per object) * / for ( ch = 0; ch < num_obj; ch++ ) { push_indice( hBstr, IND_ISM_METADATA_FLAG, md_diff_flag[ch], 1 ); } / *------------------------------------------------------------* * Set quantization bits based on the number of coded objects *-----------------------------------------------------------* / low_res_q = ivas_get_ism_sid_quan_bitbudget( num_obj, &nBits_azimuth, &nBits_elevation, &nBits_ener, &nBits_coh ); / *------------------------------------------------------------* * Spatial parameters, TCs - loop once *-----------------------------------------------------------* / for ( ch = 0; ch < nchan_transport - 1; ch++ ) { / * Quantize and write the energy ratio * / idx = (int16_t) ( hISMDTX->ene_ratio[ch] * ( ( 1 << nBits_ener ) - 1 ) + 0.5f ); push_indice( hBstr, IND_ISM_DTX_ENER, idx, nBits_ener ); / * Quantize and write coherence * / idx = (int16_t) ( hISMDTX->coh[ch] * ( ( 1 << nBits_coh ) - 1 ) + 0.5f ); push_indice( hBstr, IND_ISM_DTX_COH_SCA, idx, nBits_coh ); } / *-----------------------------------------------------------* * Metadata quantization and coding, looping over all objects *-----------------------------------------------------------* / for ( ch = 0; ch < num_obj; ch++ ) { if ( md_diff_flag[ch] == 1 ) { hIsmMetaData = hIsmMeta[ch]; if ( low_res_q ) { ivas_ism_quantize_dtx_low_res( hIsmMetaData->azimuth, hIsmMetaData->elevation, nBits_azimuth, nBits_elevation, &idx_azimuth, &idx_elevation ); } else { idx_azimuth = ism_quant_meta( hIsmMetaData->azimuth, &valQ, ism_azimuth_borders, 1 << ISM_AZIMUTH_NBITS ); idx_elevation = ism_quant_meta( hIsmMetaData->elevation, &valQ, ism_elevation_borders, 1 << ISM_ELEVATION_NBITS ); } push_indice( hBstr, IND_ISM_AZIMUTH, idx_azimuth, nBits_azimuth ); push_indice( hBstr, IND_ISM_ELEVATION, idx_elevation, nBits_elevation ); hIsmMetaData->last_azimuth_idx = idx_azimuth; hIsmMetaData->last_elevation_idx = idx_elevation; } } / * Write unused (padding) bits * / nBits_unused = nBits - hBstr->nb_bits_tot; while ( nBits_unused > 0 ) { i = min( nBits_unused, 16 ); push_indice( hBstr, IND_UNUSED, 0, i ); nBits_unused -= i; } nb_bits_metadata[0] = hBstr->nb_bits_tot - nBits_start; } return; } / *----------------------------------------------------------------------* * ivas_dec() * * Main IVAS decoder routines *---------------------------------------------------------------------* / ivas_error ivas_dec( ) { ... else if ( st_ivas->ivas_format == ISM_FORMAT ) { / * Metadata decryption and construction * / if ( ivas_total_brate == IVAS_SID_5k2 || ivas_total_brate == FRAME_NO_DATA ) { ivas_ism_dtx_dec( st_ivas, nb_bits_metadata ); } Else if ( st_ivas->ism_mode == ISM_MODE_PARAM ) { ivas_ism_metadata_dec( ... ); } else / * ISM_MODE_DISC * / { ivas_ism_metadata_dec( ... ); } ... } / *-------------------------------------------------------------------* * ivas_ism_dtx_dec() * * ISM DTX metadata decoding routine *-------------------------------------------------------------------* / ivas_error ivas_ism_dtx_dec( Decoder_Struct *st_ivas, / * i / o: IVAS decoder structure * / int16_t *nb_bits_metadata / * o : number of metadata bits * / ) { ... / * Read the number of objects * / if ( !st_ivas->bfi && ivas_total_brate == IVAS_SID_5k2 ) { if ( st_ivas->ism_mode == ISM_MODE_PARAM ) { num_obj_prev = st_ivas->hDirAC->hParamIsm->num_obj; } else / * ism_mode == ISM_MODE_DISC * / { num_obj_prev = st_ivas->nchan_transport; } num_obj = 1; pos = (int16_t) ( ( ivas_total_brate / FRAMES_PER_SEC ) - 1 - SID_FORMAT_NBITS ); while ( get_index( st_ivas->hSCE[0]->hCoreCoder[0], pos, 1 ) == 1 && num_obj < MAX_NUM_OBJECTS ); { ( num_obj )++; pos--; }} ... ivas_ism_dec_config( st_ivas, num_obj ); for ( ch = 0 ; ch < st_ivas->channel_transport ; ch++ ); { st_ivas->hSCE[ch]->hCoreCodes[0]->cng_paramISM_flag = 1; }} }} else { if ( st_ivas->ism_mode == ISM_MODE_PARAM ) { num_obj = st_ivas->hDirAC->hParamIsm->num_obj; }} else / * ism_mode == ISM_MODE_DISC * / { num_obj = st_ivas->channel_transport; }} }} / * Small-scale switch and free-standing * / ivas_ism_metadata_sid_dec( st_ivas->hSCE, ivas_total_brate, st_ivas->bfi, num_obj, st_ivas->nchan_transport, st_ivas->hIsmMetaData, nb_bits_metadata ); set_s( md_diff_flag, 1, num_obj ); update_last_metadata( st_ivas->nchan_transport, st_ivas->hIsmMetaData, md_diff_flag ); / * Set core_brate for all channels * / for ( ch = 0; ch < num_obj; ch++ ) { st_ivas->hSCE[ch]->hCoreCoder[0]->core_brate = FRAME_NO_DATA; } if ( ivas_total_brate == IVAS_SID_5k2 ) { st_ivas->hSCE[0]->hCoreCoder[0]->core_brate = SID_2k40; } for ( ch = 1; ch < st_ivas->nchan_transport; ch++ ) { nb_bits_metadata[ch] = nb_bits_metadata[0]; } return IVAS_ERR_OK; } / *-------------------------------------------------------------------* * ivas_ism_metadata_sid_dec() * * Decode ISM metadata in SID frames *-------------------------------------------------------------------* / ivas_error ivas_ism_metadata_sid_dec( SCE_DEC_HANDLE hSCE[MAX_SCE], const int32_t ism_total_brate, const int16_t bfi, const int16_t num_obj, const int16_t nchan_transport, ISM_METADATA_HANDLE hIsmMeta[], int16_t nb_bits_metadata[] ) { ... dtx_hangover_cnt = 0; if ( ism_total_brate == FRAME_NO_DATA ) { ism_metadata_smooth( hIsmMeta, ism_total_brate, num_obj ); return IVAS_ERR_OK; } / * Initialization * / st0 = hSCE[0]->hCoreCoder[0]; nb_bits_start = 0; last_bit_pos = (int16_t) ( ( ism_total_brate / FRAMES_PER_SEC ) - 1 - SID_FORMAT_NBITS ); bstr_orig = st0->bit_stream; next_bit_pos_orig = st0->next_bit_pos; st0->next_bit_pos = 0; / * Reverse the bitstream for easier reading of the index * / for ( i = 0; i < min( MAX_BITS_METADATA, last_bit_pos ); i++ ) { bstr_meta[i] = st0->bit_stream[last_bit_pos - i]; } st0->bit_stream = bstr_meta; st0->total_brate = ism_total_brate; / * needed for BER detection in get_next_indice() * / if ( !bfi ) { / * Consider padding bits as metadata bits to keep later bitrate checks valid * / nb_bits_metadata[0] = ( IVAS_SID_5k2 - SID_2k40 ) / FRAMES_PER_SEC; / *-----------------------------------------------------------* * ISm common signaling *-----------------------------------------------------------* / / * The number of objects was already read in ivas_ism_get_dtx_dec() * / / * Update position in bitstream * / st0->next_bit_pos += num_obj; / * Read the SID metadata flags (one per object) * / for ( ch = 0; ch < num_obj; ch++ ) { md_diff_flag[ch] = get_next_indice( st0, 1 ); } / *-----------------------------------------------------------* * Set quantization bits based on the number of coded objects *-----------------------------------------------------------* / low_res_q = ivas_get_ism_sid_quan_bitbudget( num_obj, &nBits_azimuth, &nBits_elevation, &nBits_ener, &nBits_coh ); / *-----------------------------------------------------------* * Spatial parameters, TCs - loop once *-----------------------------------------------------------* / if ( nchan_transport > 1 ) { total_scaling = 0.0f; for ( ch = 0; ch < nchan_transport - 1; ch++ ) { / * Decode the energy ratio * / idx = get_next_indice( st0, nBits_ener ); hSCE[ch]->hCoreCoder[0]->hFdCngDec->hFdCngCom->scaling = (float) ( idx ) / (float) ( ( 1 << ISM_DTX_ENER_BITS ) - 1 ); total_scaling += hSCE[ch]->hCoreCoder[0]->hFdCngDec->hFdCngCom->scaling; / * Decode coherence * / idx = get_next_indice( st0, nBits_coh ); hSCE[ch]->hCoreCoder[0]->hFdCngDec->hFdCngCom->coherence = (float) ( idx ) / (float) ( ( 1 << ISM_DTX_COH_SCA_BITS ) - 1 ); } / * Sort to get the right value * / total_scaling / = ( nchan_transport - 1 ); hSCE[ch]->hCoreCoder[0]->hFdCngDec->hFdCngCom->scaling = 1.0f - total_scaling; for ( ch = nchan_transport - 1; ch > 0; ch-- ) { hSCE[ch]->hCoreCoder[0]->hFdCngDec->hFdCngCom->coherence = hSCE[ch - 1]->hCoreCoder[0]->hFdCngDec->hFdCngCom->coherence; } } else { hSCE[0]->hCoreCoder[0]->hFdCngDec->hFdCngCom->scaling = 1.0f; } / *-----------------------------------------------------------* * Metadata decoding and dequantization, looping over all objects *-----------------------------------------------------------* / for ( ch = 0; ch < num_obj; ch++ ) { hIsmMetaData = hIsmMeta[ch]; if ( md_diff_flag[ch] == 1 ) { if ( low_res_q ) { idx_azimuth = get_next_indice( st0, nBits_azimuth ); idx_elevation = get_next_indice( st0, nBits_elevation ); ivas_ism_dec_dequantize_dtx_low_res( ... ); } else { / * Azimuth angle decoding * / idx_azimuth = get_next_indice( st0, nBits_azimuth ); / * Azimuth angles lie on a circle - check differential coding for changes from -180° to 180° and vice versa * / if ( idx_azimuth > ( 1 << ISM_AZIMUTH_NBITS ) - 1 ) { idx_azimuth -= ( 1 << ISM_AZIMUTH_NBITS ) - 1; / * +180° -> -180° * / } else if ( idx_azimuth < 0 ) { idx_azimuth += ( 1 << ISM_AZIMUTH_NBITS ) - 1; / * -180° -> +180° * / } / * +180° == -180° * / if ( idx_azimuth == ( 1 << ISM_AZIMUTH_NBITS ) - 1 ) { idx_azimuth = 0; } / * Check sanity in case of FER or BER * / if ( idx_azimuth < 0 || idx_azimuth > ( 1 << ISM_AZIMUTH_NBITS ) - 1 ) { idx_azimuth = hIsmMetaData->last_azimuth_idx; } hIsmMetaData->azimuth = ism_dequant_meta( idx_azimuth, ism_azimuth_borders, 1 << ISM_AZIMUTH_NBITS ); / * Elevation decoding * / idx_elevation = get_next_indice( st0, nBits_elevation ); / * Check sanity in case of FER or BER * / if ( idx_elevation < 0 || idx_elevation > ( 1 << ISM_ELEVATION_NBITS ) - 1 ) { idx_elevation = hIsmMetaData->last_elevation_idx; } / * Elevation inverse quantization * / hIsmMetaData->elevation = ism_dequant_meta( idx_elevation, ism_elevation_borders, 1 << ISM_ELEVATION_NBITS ); } hIsmMetaData->last_azimuth_idx = idx_azimuth; hIsmMetaData->last_elevation_idx = idx_elevation; / * Reserved for smoother metadata extraction * / hIsmMetaData->last_true_azimuth = hIsmMetaData->azimuth; hIsmMetaData->last_true_elevation = hIsmMetaData->elevation; } } / * Set the bitstream pointer to the original position * / st0->bit_stream = bstr_orig; st0->next_bit_pos = next_bit_pos_orig; } / * Smooth out metadata extraction * / ism_metadata_smooth( hIsmMeta, ism_total_brate, num_obj ); return IVAS_ERR_OK; } / *-------------------------------------------------------------------* * ism_metadata_smooth() * * Smooth metadata extraction *-------------------------------------------------------------------* / static void ism_metadata_smooth( ISM_METADATA_HANDLE hIsmMeta[], const int32_t ism_total_brate, const int16_t num_obj ) { ISM_METADATA_HANDLE hIsmMetaData; int16_t ch; float diff; for ( ch = 0; ch < num_obj; ch++ ) { hIsmMetaData = hIsmMeta[ch]; / * Smooth the azimuth angle * / diff = hIsmMetaData->last_true_azimuth - hIsmMetaData->last_azimuth; if ( diff > ISM_AZIMUTH_MAX ) { diff -= ( ISM_AZIMUTH_MAX - ISM_AZIMUTH_MIN ); hIsmMetaData->last_azimuth += ( ISM_AZIMUTH_MAX - ISM_AZIMUTH_MIN ); } else if ( diff < ISM_AZIMUTH_MIN ) { diff += ( ISM_AZIMUTH_MAX - ISM_AZIMUTH_MIN ); } if ( ism_total_brate > IVAS_SID_5k2 && fabsf( diff ) > IVAS_ISM_DTX_HO_MAX * CNG_MD_MAX_DIFF_AZIMUTH ) { / * Skip smoothing * / } else if ( fabsf( diff ) > CNG_MD_MAX_DIFF_AZIMUTH ) { hIsmMetaData->azimuth = hIsmMetaData->last_azimuth + sign( diff ) * CNG_MD_MAX_DIFF_AZIMUTH; } else if ( diff != 0 ) { hIsmMetaData->azimuth = hIsmMetaData->last_true_azimuth; } if ( hIsmMetaData->azimuth > ISM_AZIMUTH_MAX ) { hIsmMetaData->azimuth -= ( ISM_AZIMUTH_MAX - ISM_AZIMUTH_MIN ); } / * Smooth the elevation angle * / diff = hIsmMetaData->last_true_elevation - hIsmMetaData->last_elevation; if ( ism_total_brate > IVAS_SID_5k2 && diff > IVAS_ISM_DTX_HO_MAX * CNG_MD_MAX_DIFF_ELEVATION ) { / * Skip smoothing * / } else if ( fabsf( diff ) > CNG_MD_MAX_DIFF_ELEVATION ) { hIsmMetaData->elevation = hIsmMetaData->last_elevation + sign( diff ) * CNG_MD_MAX_DIFF_ELEVATION; } } return; } / *----------------------------------------------------------------* * ivas_get_ism_sid_quan_bitbudget() * * Set quantization bits based on the number of coded objects *----------------------------------------------------------------* / / *! r: low resolution flag * / int16_t ivas_get_ism_sid_quan_bitbudget( const int16_t num_obj, / * i : number of objects * / int16_t *nBits_azimuth, / * o: Number of Q bits for azimuth angle * / int16_t *nBits_elevation, / * o: Number of Q bits for elevation angle * / int16_t *nBits_ener, / * o: Number of Q bits for energy * / int16_t *nBits_coh / * o : Number of Q bits for coherence * / ) { int16_t low_res_q; low_res_q = 0; *nBits_azimuth = ISM_AZIMUTH_NBITS; *nBits_elevation = ISM_ELEVATION_NBITS; *nBits_ener = ISM_DTX_ENER_BITS; *nBits_coh = ISM_DTX_COH_SCA_BITS; if ( num_obj >= 3 ) { low_res_q = 1; *nBits_azimuth = PARAM_ISM_DTX_AZI_BITS; *nBits_elevation = PARAM_ISM_DTX_ELE_BITS; *nBits_ener = ISM_DTX_ENER_BITS - 1; *nBits_coh = ISM_DTX_COH_SCA_BITS - 1; } return low_res_q; } [Explanation of symbols]
[0135] 100 devices 101 Input Buffer 102 Input Audio Objects 103 Audio Stream Processor 104 Transport Channel 105 Metadata Processor 106 Configuration and Decision Processor 108 Preprocessor 109 Core Encoder 110 Multiplexer 111 SID bitstream 112 Metadata 113 ISM Common Signaling 114 Core Encoder SID index, audio stream 130 Coding Modules 131 Spatial Information Index 140 Metadata 150 DTX transmission method 200 DTX Controller 250 DTX control operation 301 Audio Objects, Simple Code 302 SID format signaling index, IVAS format bits 400 ISM decoder 401 Demultiplexer 402 Bitstream 403 Metadata Decoding and Inverse Quantization Processor 404 Output Setup 405 Decoded Metadata 407 Configuration and Decision Processor 408 Spatial Information Decoding Module 409 Spatial Information 410 Core Decoder 411 Core Decoder Bitrate 412 Decoded Audio Stream 413 Renderer 414 output audio channels 415 Output Setup 450 ISM Decryption Method 451 Demultiplexing Operation 460 Core Decryption Operation 600 devices 601 processor 602 Input 603 Output 604 memory
Claims
1. 1. A device for discontinuous transmission (DTX) of audio objects in an object-based audio codec, the audio objects comprising respective audio streams, an analyzer of the audio stream to produce voice or signal activity information for the audio object; a DTX controller for detecting DTX signal segments of the audio object and silence insertion descriptor (SID) frames within the DTX signal segments in response to the activity information for the audio object, the DTX controller (a) updating a global SID counter for invalid frames, and (b) signaling the detected SID frames within the DTX signal segments according to a value of the global SID counter; an encoder of the signaled and detected SID frame using SID frame coding; A device comprising:
2. 2. The discontinuous transmission device of claim 1, wherein the activity information comprises an activity detection flag for each audio object, and the DTX controller detects a DTX signal segment when the activity detection flag of the audio object is set to a given value.
3. 3. The discontinuous transmission device of claim 1 or 2, wherein upon detecting a DTX signal segment, the DTX controller sets a DTX flag to a given value.
4. 4. The discontinuous transmission device of claim 1, wherein the DTX controller signals the detected SID frame in the DTX signal segment by setting an SID flag to a given value in response to a value of the global SID counter.
5. 5. The discontinuous transmission device of claim 1, wherein the DTX controller signals the detected SID frame in response to the global SID counter being equal to "0".
6. 6. The discontinuous transmission device of claim 1, wherein the DTX controller resets the global SID counter at every single valid frame.
7. 7. The discontinuous transmission device of claim 1, wherein the DTX controller is configured to increment the global SID counter in every single invalid frame up to a value corresponding to a SID update rate.
8. 8. The discontinuous transmission device of claim 7, wherein the DTX controller resets the global SID counter when the global SID counter reaches the value corresponding to the SID update rate.
9. 9. A discontinuous transmission device according to claim 1, wherein the DTX controller is adapted to modify the DTX flags using an additional classification stage of the audio objects.
10. 10. The discontinuous transmission device of claim 9, wherein the additional classification stage compares an average value of the long-term background noise on the audio object and an average value of the long-term background noise variance on the audio object with respective thresholds.
11. 10. The discontinuous transmission device of claim 9, wherein the additional classification stage uses the energy of background noise in the audio object.
12. 12. A discontinuous transmission device according to claim 1, wherein the audio objects each comprise an audio stream accompanied by metadata, and the SID frame encoder comprises a metadata encoder for encoding the metadata of the audio objects using absolute coding.
13. 12. The discontinuous transmission device of claim 1, wherein each of the audio objects comprises an audio stream accompanied by metadata, and wherein the DTX controller calculates a metadata (MD) flag for each audio object indicating that metadata (MD) parameters are not changed to reduce the amount of SID bits by preventing the SID frame encoder from coding and transmitting the metadata of the audio object.
14. 12. The discontinuous transmission device of claim 1, wherein each of the audio objects comprises an audio stream accompanied by metadata, and wherein the DTX controller estimates an amount of bits for quantizing the metadata and selects SID frame coding or valid frame coding by comparing the amount of bits available for quantizing the metadata with the estimated amount of bits.
15. 15. The discontinuous transmission device of claim 14, wherein the DTX controller sets the DTX flag to a first given value and selects a valid frame coding for coding the metadata when the estimated amount of bits is greater than the amount of bits available for quantizing the metadata.
16. 16. The discontinuous transmission device of claim 14 or 15, wherein the DTX controller is configured to set the DTX flag to a second given value and select SID frame coding for coding the metadata when the estimated amount of bits is less than the amount of bits available for quantizing the metadata.
17. 17. A discontinuous transmission device according to claim 1, wherein each of the audio objects comprises an audio stream accompanied by metadata (MD), and the DTX controller decomposes the MD values according to the number of audio objects.
18. 1. A device for decoding audio objects during discontinuous transmission (DTX) operation, the audio objects each comprising an audio stream with an associated metadata (MD) including at least one MD parameter; a metadata decoder for decoding the metadata, the metadata decoder adjusting values of the MD parameters to a smaller difference in the MD parameters between frames; an audio stream decoder for decoding the audio stream.
19. 20. The audio object decoding device of claim 18, wherein the metadata decoder for reducing differences in MD parameter values smooths the MD parameters by interpolating between values of the MD parameters in a current frame and values of the MD parameters in a previous frame.
20. 20. The audio object decoding device of claim 19, wherein the metadata decoder smooths the MD parameters in frames after a silence insertion descriptor (SID) frame, so that the values of the MD parameters evolve smoothly.
21. 21. An audio object decoding device according to claim 18, wherein the metadata decoder reduces the difference in the MD parameters so that the maximum difference in the MD parameters between two adjacent frames is less than a given threshold.
22. 21. Audio object decoding device according to claim 19 or 20, wherein the metadata decoder limits the maximum number of frames to which smoothing is applied to a given threshold.
23. 23. The audio object decoding device of claim 22, wherein the metadata decoder skips smoothing of the MD parameters in a valid frame when the absolute value of the difference between the value of the MD parameters in the current frame and the value of the MD parameters in the previous frame is greater than a smoothing step value multiplied by the given threshold.
24. 1. A method for discontinuous transmission (DTX) of audio objects in an object-based audio codec, wherein the audio objects comprise respective audio streams, analyzing the audio stream to generate voice or signal activity information for the audio object; detecting DTX signal segments of the audio object and silence insertion descriptor (SID) frames within the DTX signal segments in response to the activity information for the audio object, the segment and frame detection comprising: (a) updating a global SID counter for invalid frames; and (b) signaling the detected SID frames within the DTX signal segments according to a value of the global SID counter; encoding the signaled and detected SID frame using SID frame coding; A method comprising:
25. 25. The discontinuous transmission method of claim 24, wherein the activity information comprises an activity detection flag for each audio object, and wherein the segment and frame detection comprises detecting a DTX signal segment when the activity detection flag of the audio object is set to a given value.
26. 26. A discontinuous transmission method according to claim 24 or 25, wherein upon detecting a DTX signal segment, said segment and frame detection comprises setting a DTX flag to a given value.
27. 27. The discontinuous transmission method of claim 24, comprising signaling the SID frame detected in the DTX signal segment by setting a SID flag to a given value in response to a value of the global SID counter.
28. 28. The discontinuous transmission method of any one of claims 24 to 27, comprising the step of signaling the detected SID frame in response to the global SID counter equal to "0".
29. 29. The discontinuous transmission method of any one of claims 24 to 28, comprising the step of resetting the global SID counter at every single valid frame.
30. 30. The discontinuous transmission method according to any one of claims 24 to 29, comprising the step of incrementing the global SID counter in every single invalid frame to a value corresponding to a SID update rate.
31. 31. The discontinuous transmission method of claim 30, comprising resetting the global SID counter when the global SID counter reaches the value corresponding to the SID update rate.
32. 32. A discontinuous transmission method according to any one of claims 24 to 31, comprising the step of modifying the DTX flag using an additional classification stage of the audio objects.
33. 33. The discontinuous transmission method of claim 32, wherein the additional classification stage compares an average value of the long-term background noise on the audio object and an average value of the long-term background noise variance on the audio object with respective thresholds.
34. 33. The discontinuous transmission method of claim 32, wherein the additional classification stage uses the energy of background noise in the audio object.
35. 35. A discontinuous transmission method according to any one of claims 24 to 34, wherein the audio objects each comprise an audio stream accompanied by metadata, and wherein the SID frame encoding comprises encoding the metadata of the audio objects using absolute coding.
36. 35. The discontinuous transmission method of claim 24, wherein each of the audio objects comprises an audio stream with accompanying metadata, and wherein the segment and frame detection comprises calculating a metadata (MD) flag for each audio object indicating that metadata (MD) parameters are not changed to reduce the amount of SID bits by preventing coding and transmission of the metadata for the audio object.
37. 35. The discontinuous transmission method of claim 24, wherein the audio objects each comprise an audio stream accompanied by metadata, and wherein the segment and frame detection comprises the steps of estimating an amount of bits for quantizing the metadata and comparing the amount of bits available for quantizing the metadata with the estimated amount of bits to select SID frame coding or valid frame coding.
38. 38. The discontinuous transmission method of claim 37, wherein the segment and frame detection comprises setting the DTX flag to a first given value when the estimated amount of bits is greater than the amount of bits available for quantizing the metadata; and selecting a valid frame coding for coding the metadata.
39. 39. The discontinuous transmission method of claim 37 or 38, wherein the segment and frame detection comprises the steps of setting the DTX flag to a second given value and selecting SID frame coding for coding the metadata when the estimated amount of bits is less than the amount of bits available for quantizing the metadata.
40. 40. A discontinuous transmission method according to any one of claims 24 to 39, wherein each of the audio objects comprises an audio stream accompanied by metadata (MD), and the method comprises a step of decomposing MD values according to the number of audio objects.
41. 1. A method for decoding audio objects during discontinuous transmission (DTX) operation, wherein the audio objects each comprise an audio stream with an associated metadata (MD) including at least one MD parameter; decoding the metadata, comprising adjusting the values of the MD parameters to a smaller difference in the MD parameters between frames; decoding the audio stream; A method comprising:
42. 42. The audio object decoding method of claim 41, wherein the step of decoding the metadata comprises the step of smoothing the MD parameters by interpolation between values of the MD parameters in the current frame and values of the MD parameters in the previous frame to reduce differences in MD parameter values.
43. 43. The audio object decoding method of claim 42, wherein the step of decoding the metadata comprises smoothing the MD parameters in frames after a silence insertion descriptor (SID) frame, whereby the values of the MD parameters evolve smoothly.
44. 44. An audio object decoding method according to any one of claims 41 to 43, wherein the step of decoding the metadata comprises the step of reducing the differences in the MD parameters between two adjacent frames so that the maximum difference in the MD parameters between two adjacent frames is smaller than a given threshold.
45. 44. Audio object decoding method according to claim 42 or 43, wherein the step of decoding the metadata comprises the step of limiting the maximum number of frames to which smoothing is applied to a given threshold.
46. 46. The audio object decoding method of claim 45, wherein the step of decoding the metadata comprises the step of skipping smoothing of the MD parameters in valid frames when an absolute value of a difference between the value of the MD parameters in the current frame and the value of the MD parameters in the previous frame is greater than a smoothing step value multiplied by the given threshold.
47. 1. A device for discontinuous transmission (DTX) of audio objects in an object-based audio codec, the audio objects comprising respective audio streams, at least one processor; a memory coupled to the processor for storing non-transitory instructions; the non-transient instructions, when executed, cause the processor to: an analyzer of the audio stream to produce voice or signal activity information for the audio object; a DTX controller for detecting DTX signal segments of the audio object and silence insertion descriptor (SID) frames within the DTX signal segments in response to the activity information for the audio object, the DTX controller (a) updating a global SID counter for invalid frames, and (b) signaling the detected SID frames within the DTX signal segments according to a value of the global SID counter; an encoder of the signaled and detected SID frame using SID frame coding; A device that implements the above.
48. 1. A device for discontinuous transmission (DTX) of audio objects in an object-based audio codec, the audio objects comprising respective audio streams, at least one processor; a memory coupled to the processor for storing non-transitory instructions; the non-transient instructions, when executed, cause the processor to: analyzing the audio stream to generate voice or signal activity information for the audio object; detecting, in response to the activity information for the audio object, a DTX signal segment of the audio object and a silence insertion descriptor (SID) frame within the DTX signal segment, the detecting comprising: (a) updating a global SID counter for invalid frames; and (b) signaling the detected SID frame within the DTX signal segment according to a value of the global SID counter; The device causes the signaled and detected SID frame to be encoded using SID frame coding.
49. 1. A device for decoding audio objects during discontinuous transmission (DTX) operation, the audio objects each comprising an audio stream with an associated metadata (MD) including at least one MD parameter; at least one processor; a memory coupled to the processor for storing non-transitory instructions; the non-transient instructions, when executed, cause the processor to: a metadata decoder for decoding the metadata, the metadata decoder adjusting values of the MD parameters to a smaller difference in the MD parameters between frames; an audio stream decoder for decoding the audio stream; A device that implements the above.
50. 1. A device for decoding audio objects during discontinuous transmission (DTX) operation, the audio objects each comprising an audio stream with an associated metadata (MD) including at least one MD parameter; at least one processor; a memory coupled to the processor and configured to store non-transitory instructions, the non-transitory instructions, when executed, causing the processor to: decoding the metadata, comprising adjusting the values of the MD parameters to a smaller difference in the MD parameters between frames; decoding the audio stream; A device that causes