Apparatus, method and computer program for encoding an audio signal or for decoding an encoded audio scene
Patent Information
- Application Number
- CN202180067397.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-30
- Filing Date
- 2021-05-31
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2041-05-31
AI Technical Summary
[0305] Embodiments of the present invention allow for the efficient extension of DTX to parameter-space audio coding. Even for inactive frames, this enables the recovery of background noise with high perceptual fidelity, and transmission can be interrupted to save communication bandwidth.
Smart Images

Figure CN116348951B_ABST
Abstract
Description
[0001] This document relates in particular to an apparatus for generating encoded audio scenes, and an apparatus for decoding and / or processing encoded audio scenes. It also relates to related methods and a non-transitory storage unit for storing instructions that, when executed by a processor, cause the processor to perform the related methods.
[0002] This paper discusses methods for discontinuous transmission mode (DTX) and comfort noise generation (CNG) for audio scenes. For audio scenes, spatial images are parametrically encoded using the DirAC paradigm or transmitted in the metadata-assisted spatial audio (MASA) format.
[0003] The embodiments relate to the discontinuous transmission of parametrically encoded spatial audio, such as the DTX mode for DirAC and MASA.
[0004] Embodiments of the present invention relate to the efficient transmission and rendering of conversational speech, for example, captured using a sound field microphone. The captured audio signal is thus often referred to as three-dimensional (3D) audio, which enhances immersion and improves intelligibility and user experience because sound events can be confined to three-dimensional space.
[0005] For example, transmitting audio scenes in three dimensions requires handling multiple channels that typically cause large data transmissions. For instance, the Directional Audio Coding (DirAC) technique [1] can be used to reduce large raw data rates. DirAC is considered an efficient method for analyzing and parametrically representing audio scenes. It is perceptually excited and represents the sound field by means of the direction of arrival (DOA) and diffusion measured per frequency band. It is based on the assumption that, at an instant and for a critical frequency band, the spatial resolution of the auditory system is limited to decoding one cue for direction and another cue for interaural coherence. The spatial sound is then reproduced in the frequency domain by crossfading the two streams, a non-directional diffuse stream and a directional non-diffused stream.
[0006] Furthermore, in a typical session, each speaker is silent for approximately 60 percent of the time. By distinguishing between frames of audio signals containing speech (“active frames”) and frames containing only background noise or silence (“inactive frames”), the speech encoder can save effective data rate. Inactive frames are typically perceived as carrying very little or no information, and speech encoders are often configured to reduce their bit rate for such frames, or even not transmit any information at all. In this case, the encoder operates in a so-called discontinuous transmission (DTX) mode, an efficient way to significantly reduce the transmission rate of a communication codec in the absence of voice input. In this mode, most frames determined to consist only of background noise are discarded from transmission and replaced by some comfort noise generation (CNG) in the decoder. For these frames, the extremely low rate parameter of the signal indicates transmission via silence insertion descriptor (SID) frames sent periodically, but not at every frame. This allows the CNG in the decoder to generate artificial noise similar to actual background noise.
[0007] Embodiments of the present invention relate to DTX systems for, for example, 3D audio scenes captured by a sound field microphone and parametrically encoded using encoding schemes based on the DirAC paradigm and similar methods, as well as, in particular, SID and CNG. The present invention allows for a dramatic reduction in the bit rate requirements for transmitting conversational immersive speech. Existing technology
[0008] [1] V. Pulkki, MV. Laitinen, J. Vilkamo, J. Ahonen, T. Lokki, and T.Pihlajamäki, ''Directional audio coding - perception-based reproduction of spatial sound'', International Workshop on the Principles and Application onSpatial Hearing, November 2009, Zao; Miyagi, Japan.
[0009] [2] 3GPP TS 26.194; Voice Activity Detector (VAD); - 3GPP technical specification Searched on 2009-06-17.
[0010] [3] 3GPP TS 26.449, "Codec for Enhanced Voice Services (EVS); ComfortNoise Generation (CNG) Aspects".
[0011] [4] 3GPP TS 26.450, "Codec for Enhanced Voice Services (EVS); Discontinuous Transmission (DTX)"
[0012] [5] A. Lombard, S. Wilde, E. Ravelli, S. Döhla, G. Fuchs and M. Dietz, "Frequency-domain Comfort Noise Generation for Discontinuous Transmission in EVS," 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, QLD, 2015, pp. 5893-5897, doi: 10.1109 / ICASSP.2015.7179102.
[0013] [6] V. Pulkki, "Virtual source positioning using vector base amplitude panning", J. Audio Eng. Soc., 45(6): 456-466, June 1997.
[0014] [7] J. Ahonen and V. Pulkki, "Diffuseness estimation using temporal variation of intensity vectors", in Workshop on Applications of Signal Processing to Audio and Acoustics WASPAA, Mohonk Mountain House, New Paltz, 2009.
[0015] [8] T. Hirvonen, J. Ahonen, and V. Pulkki, ''Perceptual compression methods for metadata in Directional Audio Coding applied to audiovisualteleconference'', AES 126th Convention, May 7-10, 2009, Munich, Germany.
[0016] [9] Vilkamo, Juha & Bäckström, Tom & Kuntz, Achim. (2013). OptimizedCovariance Domain Framework for Time--Frequency Processing of Spatial Audio. Journal of the Audio Engineering Society. 61.
[0017]
[10] M. Laitinen and V. Pulkki, "Converting 5.1 audio recordings to B-format for directional audio coding reproduction," 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Prague, 2011, pp. 61-64, doi: 10.1109 / ICASSP.2011.5946328. Summary of the Invention
[0018] According to one aspect, an apparatus is provided for generating an encoded audio scene from an audio signal having a first frame and a second frame, comprising:
[0019] A sound field parameter generator is used to determine a first sound field parameter representation for the first frame from the audio signal in the first frame, and to determine a second sound field parameter representation for the second frame from the audio signal in the second frame;
[0020] An activity detector is used to analyze audio signals to determine the first frame as an active frame and the second frame as an inactive frame based on the audio signals.
[0021] An audio signal encoder is used to generate an encoded audio signal for the first frame (active frames) and to generate parameter descriptions for the second frame (inactive frames); and
[0022] The encoded signal shaper is used to construct an encoded audio scene by combining a first sound field parameter representation for a first frame, a second sound field parameter representation for a second frame, an encoded audio signal for the first frame, and a parameter description for the second frame.
[0023] The sound field parameter generator can be configured to generate a first sound field parameter representation or a second sound field parameter representation, such that the first sound field parameter representation or the second sound field parameter representation includes parameters indicating the characteristics of the audio signal relative to the listener's position.
[0024] The first sound field parameter representation or the second sound field parameter representation may include one or more direction parameters indicating the direction of the sound in the first frame relative to the listener's position, or one or more diffusion parameters indicating the portion of the diffused sound in the first frame relative to the direct sound, or one or more energy ratio parameters indicating the energy ratio of the direct sound to the diffused sound in the first frame, or inter-channel / surround coherence parameters in the first frame.
[0025] The sound field parameter generator can be configured to identify multiple individual sound sources from the first or second frame of an audio signal and determine parameter descriptions for each sound source.
[0026] The sound field generator is configured to decompose a first frame or a second frame into multiple frequency ranges, each frequency range representing an individual sound source, and to determine at least one sound field parameter for each frequency range. The sound field parameter includes, for example, a direction parameter, a direction of arrival parameter, a diffusion parameter, an energy ratio parameter, or any parameter representing the characteristics of the sound field represented by the first frame of the audio signal relative to the listener's position.
[0027] The audio signals used for the first and second frames may contain an input format having multiple components representing the sound field relative to the listener.
[0028] The sound field parameter generator is configured, for example, to use downmixing of multiple components to calculate one or more transport channels for the first and second frames, and to analyze the input format to determine a first parameter representation associated with one or more transport channels, or
[0029] The sound field parameter generator is configured, for example, to calculate one or more transmission channels using downmixing of multiple components, and
[0030] The activity detector is configured to analyze one or more transmission channels derived from the audio signal in the second frame.
[0031] The audio signal used for the first or second frame may contain an input format, which, for each of the first and second frames, has one or more transport channels and metadata associated with each frame.
[0032] The sound field parameter generator is configured to read metadata from a first frame and a second frame, use or process the metadata for the first frame as a first sound field parameter representation, and process the metadata for the second frame to obtain a second sound field parameter representation, wherein the processing to obtain the second sound field parameter representation reduces the amount of information units required to transmit the metadata for the second frame relative to the amount required before processing.
[0033] The sound field parameter generator can be configured to process the metadata used for the second frame to reduce the number of information items in the metadata or to resample the information items in the metadata to a lower resolution, such as temporal resolution or frequency resolution, or to requantize the information units of the metadata used for the second frame into a coarser representation relative to the case before requantization.
[0034] The audio signal encoder can be configured to determine the silence information description used for inactive frames as a parameter description.
[0035] The silent information description exemplarily includes amplitude-related information such as energy, power, or loudness for the second frame and shaping information such as spectral shaping information, or amplitude-related information such as energy, power, or loudness for the second frame and linear prediction coding (LPC) parameters for the second frame, or scale parameters for the second frame with varying associated frequency resolution, such that different scale parameters refer to frequency bands with different widths.
[0036] An audio signal encoder can be configured to encode an audio signal for a first frame using a time-domain or frequency-domain coding mode. The encoded audio signal includes, for example, encoded time-domain samples, encoded frequency-domain samples, encoded LPC-domain samples, and side information obtained from the components of the audio signal or from one or more transport channels, for example, through downmixing.
[0037] The audio signal may include an input format, such as a first-order ambisonics format, a higher-order ambisonics format, a multichannel format associated with a given speaker setup such as 5.1, 7.1, or 7.1+4, or one or more audio channels representing one or more different audio objects located in a space as indicated by information included in the associated metadata, or an input format that is a metadata-associated spatial audio representation.
[0038] The sound field parameter generator is configured to determine a first sound field parameter representation and a second sound field representation, such that the parameters represent the sound field relative to a defined listener position, or
[0039] The audio signal includes microphone signals such as those obtained from a real microphone or a virtual microphone, or synthesized microphone signals, for example, in a first-order stereo reverb format or a higher-order stereo reverb format.
[0040] The activity detector can be configured to detect inactive phases on the second frame and one or more frames thereafter, and
[0041] The audio signal encoder is configured to generate another parameter description for inactive frames only for another third frame, wherein, in terms of frame timing, the other third frame is separated from the second frame by at least one frame, and
[0042] The sound field parameter generator is configured to determine an alternative sound field parameter representation only for frames for which the audio signal encoder has already defined parameters, or
[0043] The activity detector is configured to determine an inactive phase including the second frame and the eight frames following the second frame, and the audio signal encoder is configured to generate a parameter description for the inactive frame only at every eighth frame, and the sound field parameter generator is configured to generate a sound field parameter representation for every eighth inactive frame, or
[0044] The sound field parameter generator is configured to generate a sound field parameter representation for each inactive frame, even when the audio signal encoder has not generated a parameter description for inactive frames.
[0045] The sound field parameter generator is configured to determine the parameter representation at a higher frame rate than the parameter descriptions generated by the audio signal encoder for one or more inactive frames.
[0046] The sound field parameter generator can be configured to determine a second sound field parameter representation for a second frame using spatial parameters for one or more directions in the frequency band and the associated energy ratio in the frequency band corresponding to the ratio of a directional component to the total energy.
[0047] The sound field parameter generator is configured to determine a second sound field parameter representation for the second frame to determine a diffusion parameter indicating the ratio of diffused sound to direct sound, or
[0048] The sound field parameter generator is configured to determine a second sound field parameter representation for the second frame to determine directional information using a coarser quantization scheme compared to the quantization in the first frame, or
[0049] The sound field parameter generator is configured to use averaging of directions over time or frequency to obtain a coarser time or frequency resolution, to determine a second sound field parameter representation for the second frame, or
[0050] The sound field parameter generator is configured to determine a second sound field parameter representation for the second frame to determine a sound field parameter representation for one or more inactive frames, the sound field parameter representation for the one or more inactive frames having the same frequency resolution as the first sound field parameter representation for the active frame, and the direction information in the sound field parameter representation for the inactive frames having a lower temporal occurrence rate compared to the temporal occurrence rate for the active frame, or
[0051] The sound field parameter generator is configured to determine a second sound field parameter representation for the second frame, wherein the second sound field parameter representation has a diffusion parameter, and the diffusion parameter is transmitted with the same time or frequency resolution as the active frame but after coarser quantization, or
[0052] The sound field parameter generator is configured to determine a second sound field parameter representation for the second frame to quantize a diffusion parameter for the second sound field representation with a first number of bits, and wherein only a second number of bits for each quantization index is transmitted, the second number of bits being less than the first number of bits, or
[0053] The sound field parameter generator is configured to determine a second sound field parameter representation for the second frame, thereby determining inter-channel coherence for the second sound field parameter representation if the audio signal has an input channel corresponding to a channel located in the spatial domain, or determining inter-channel sound level difference for the second sound field parameter representation if the audio signal has an input channel corresponding to a channel located in the spatial domain.
[0054] The sound field parameter generator is configured to determine a second sound field parameter representation for the second frame to determine surround coherence, which is defined as the ratio of coherent diffuse energy in the sound field represented by the audio signal.
[0055] According to one aspect, an apparatus is provided for processing a coded audio scene, the coded audio scene including a first sound field parameter representation and a coded audio signal in a first frame, wherein a second frame is an inactive frame, the apparatus comprising:
[0056] An activity detector is used to detect if the second frame is an inactive frame.
[0057] A synthesized signal synthesizer is used to synthesize a synthesized audio signal for a second frame using a parameter description for the second frame.
[0058] An audio decoder, used to decode the encoded audio signal for the first frame; and
[0059] A spatial renderer for spatially rendering an audio signal for the first frame using a first sound field parameter representation and a synthesized audio signal for the second frame, or a transcoder for generating a metadata-assisted output format containing the audio signal for the first frame, a first sound field parameter representation for the first frame, a synthesized audio signal for the second frame, and a second sound field parameter representation for the second frame.
[0060] The encoded audio scene may include a second sound field parameter description for a second frame, and wherein the device includes a sound field parameter processor for deriving one or more sound field parameters from the second sound field parameter representation, and wherein a spatial renderer is configured to use one or more sound field parameters for the second frame to render a synthesized audio signal for the second frame.
[0061] The device may include a parameter processor for deriving one or more sound field parameters for the second frame.
[0062] The parameter processor is configured to store a sound field parameter representation for a first frame and use the stored first sound field parameter representation for the first frame to synthesize one or more sound field parameters for a second frame, wherein the second frame is temporally after the first frame, or
[0063] The parameter processor is configured to store one or more sound field parameter representations for several frames that occur temporally before or after the second frame, to extrapolate or interpolate using at least two of the sound field parameter representations for the several frames to determine one or more sound field parameters for the second frame, and
[0064] The spatial renderer is configured to use one or more sound field parameters for the second frame to render the synthesized audio signal for the second frame.
[0065] The parameter processor can be configured to perform jitter using directions included in at least two sound field parameter representations that occur before or after the second frame when extrapolating or interpolating to determine one or more sound field parameters for the second frame.
[0066] An encoded audio scene may contain one or more transport channels for the first frame.
[0067] The synthesized signal generator is configured to generate one or more transport channels as a synthesized audio signal for the second frame, and
[0068] The spatial renderer is configured to render one or more transport channels in space for the second frame.
[0069] The synthesized signal generator can be configured to generate multiple synthesized component audio signals for the second frame, which are individual components related to the audio output format of the spatial renderer, as a synthesized audio signal.
[0070] The synthesized signal generator can be configured to generate individual synthesized component audio signals for each of at least a subset of at least two individual components related to the audio output format.
[0071] The first other synthesized component audio signal is decorrelated with the second other synthesized component audio signal, and
[0072] The spatial renderer is configured to render the audio output format components using a combination of a first other synthesized component audio signal and a second other synthesized component audio signal.
[0073] The spatial renderer can be configured to apply the covariance method.
[0074] The spatial renderer can be configured not to use any decorrelation processing or control decorrelation processing, such that when generating components of the audio output format, only a certain amount of decorrelation signal generated by decorrelation processing as indicated by the covariance method is used.
[0075] The synthesized signal generator is a comfort noise generator.
[0076] The synthesized signal generator may include a noise generator, and a first additional synthesized component audio signal is generated by a first sample of the noise generator, and a second additional synthesized component audio signal is generated by a second sample of the noise generator, wherein the second sample is different from the first sample.
[0077] The noise generator may include a noise table, wherein a first additional synthesized component audio signal is generated by taking a first portion of the noise table, and wherein a second additional synthesized component audio signal is generated by taking a second portion of the noise table, wherein the second portion of the noise table is different from the first portion of the noise table, or
[0078] The noise generator includes a pseudo-noise generator, and the first additional synthesized component audio signal is generated using a first seed for the pseudo-noise generator, and the second additional synthesized component audio signal is generated using a second seed for the pseudo-noise generator.
[0079] An encoded audio scene may contain two or more transport channels for the first frame, and
[0080] The synthesized signal generator includes a noise generator and is configured to generate a first transmission channel by sampling the noise generator and a second transmission channel by sampling the noise generator, wherein the first and second transmission channels determined by sampling the noise generator are weighted using the same parameter description for the second frame.
[0081] The space renderer can be configured as
[0082] The mixture of a direct signal and a diffused signal generated from the direct signal under the control of a decorrelation unit representing the first sound field parameters operates in a first mode for the first frame.
[0083] The second mode for the second frame is operated by mixing a first synthesized component signal with a second synthesized component signal, wherein the first synthesized component signal and the second synthesized component signal are generated by a synthesized signal synthesizer through different implementations of noise processing or pseudo-noise processing.
[0084] The spatial renderer can be configured to control the blending in the second mode by using diffusion parameters, energy distribution parameters, or coherence parameters derived by the parameter processor for the second frame.
[0085] The synthesized signal generator can be configured to generate a synthesized audio signal for the first frame using a parameter description for the second frame, and
[0086] The spatial renderer is configured to perform a weighted combination of the audio signal for the first frame and the synthesized audio signal for the first frame before or after spatial rendering, wherein in the weighted combination, the intensity of the synthesized audio signal for the first frame is reduced relative to the intensity of the synthesized audio signal for the second frame.
[0087] The parameter processor can be configured to determine surround coherence for a second inactive frame, which is defined as the ratio of coherent diffuse energy in the sound field represented by the second frame, wherein the spatial renderer is configured to redistribute the energy between the direct signal and the diffuse signal in the second frame based on the acoustic coherence, wherein the energy of the acoustic surround coherence component is removed from the diffuse energy to be redistributed to the directional component, and wherein the directional component is translated in the reproduction space.
[0088] The device may include an output interface for converting an audio output format generated by a spatial renderer into a transcoded output format, such as an output format containing multiple output channels dedicated to speakers to be placed in a predetermined location, or a transcoded output format containing FOA or HOA data, or
[0089] The alternative spatial renderer provides a transcoder for generating a metadata-assisted output format, which includes an audio signal for the first frame, first sound field parameters for the first frame, a synthesized audio signal for the second frame, and a representation of second sound field parameters for the second frame.
[0090] The activity detector can be configured to detect the second frame as an inactive frame.
[0091] According to one aspect, a method for generating an encoded audio scene from an audio signal having a first frame and a second frame is provided, comprising:
[0092] A first sound field parameter representation for the first frame is determined from the audio signal in the first frame, and a second sound field parameter representation for the second frame is determined from the audio signal in the second frame;
[0093] Analyze the audio signal to determine the first frame as the active frame and the second frame as the inactive frame based on the audio signal;
[0094] Generate the encoded audio signal for the first frame as the active frame and generate the parameter description for the second frame as the inactive frame; and
[0095] An encoded audio scene is constructed by combining the first sound field parameter representation for the first frame, the second sound field parameter representation for the second frame, the encoded audio signal for the first frame, and the parameter description for the second frame.
[0096] According to one aspect, a method for processing a coded audio scene is provided, the coded audio scene including a first sound field parameter representation and a coded audio signal in a first frame, wherein a second frame is an inactive frame, the method comprising:
[0097] The second frame is detected as an inactive frame, and a parameter description for the second frame is provided.
[0098] Use the parameter description for the second frame to synthesize the synthesized audio signal for the second frame;
[0099] Decode the encoded audio signal used for the first frame; and
[0100] The audio signal for the first frame is spatially rendered using the first sound field parameter representation and the synthesized audio signal for the second frame, or a metadata-assisted output format is generated that includes the audio signal for the first frame, the first sound field parameter representation for the first frame, the synthesized audio signal for the second frame, and the second sound field parameter representation for the second frame.
[0101] The method may include a parameter description that is provided for the second frame.
[0102] According to one aspect, an encoded audio scene is provided, comprising:
[0103] The first sound field parameters are used for the first frame;
[0104] The second sound field parameter representation used for the second frame;
[0105] Encoded audio signal for the first frame; and
[0106] Parameter description used for the second frame.
[0107] According to one aspect, a computer program is provided for performing the above or below methods when running on a computer or processor.
[0108] Attached Figure
[0109] Figure 1 (divided into) Figure 1a and Figure 1b (This shows examples based on existing technologies that can be used for analysis and synthesis based on examples.)
[0110] Figure 2 Examples of decoders and encoders are shown based on the example.
[0111] Figure 3 This shows an example of an encoder based on the example.
[0112] Figure 4 and 5 Examples of components are shown.
[0113] Figure 5 Examples of components based on the example are shown.
[0114] Figures 6 to 11 An example of a decoder is shown. Example
[0115] First, some discussions of known paradigms (DTX, DirAC, MASA, etc.) are provided, and the descriptions of some of these techniques can be implemented, at least in some cases, in the examples of this invention.
[0116] DTX
[0117] Comfort noise generators are commonly used in discontinuous transmission of speech (DTX). In this mode, speech is first classified into active and inactive frames by a voice activity detector (VAD). An example of a VAD can be found in [2]. Based on the VAD results, only active speech frames are encoded and transmitted at the nominal bit rate. During long pauses where only background noise exists, the bit rate is reduced or zeroed, and the background noise is encoded sporadically and parametrically. The average bit rate is then significantly reduced. Noise is generated by a comfort noise generator (CNG) on the decoder side during inactive frames. For example, both the speech encoder AMR-WB [2] and 3GPP EVS [3,4] can potentially operate in DTX mode. An example of an efficient CNG is given in [5].
[0118] Embodiments of the present invention extend this principle in a way that applies the same principle to immersive conversational speech by utilizing the spatial localization of sound events.
[0119] DirAC
[0120] DirAC is a perceptually stimulated reproduction of spatial sound. It is assumed that at any given moment and for a critical frequency band, the spatial resolution of the auditory system is limited to decoding a cue for one direction and decoding another cue for interaural coherence.
[0121] Based on these assumptions, DirAC represents spatial sound in a frequency band through crossfading of two streams: a non-directional diffuse stream and a directional non-diffusion stream. DirAC processing is performed in two stages: analysis and synthesis, as depicted in Figure 1. Figure 1a The synthesis is shown. Figure 1b (Analysis shown).
[0122] During the DirAC analysis phase, a first-order coincident microphone in B format is considered as input, and the sound dispersion and direction of arrival are analyzed in the frequency domain.
[0123] In the DirAC synthesis stage, the sound is split into two streams, a non-diffuse stream and a diffuse stream. The non-diffuse stream is reproduced as a point source using amplitude translation, which can be done using vector-based amplitude translation (VBAP) [6]. The diffuse stream is largely responsible for the sense of immersion and is generated by sending mutually decorrelated signals to the loudspeaker.
[0124] The DirAC parameters, also referred to below as spatial metadata or DirAC metadata, consist of tuples of diffusivity and direction. Direction can be represented in spherical coordinates by two angles (azimuth and elevation), while diffusivity can be a scalar factor between 0 and 1.
[0125] Some work has been done to reduce the size of metadata so that the DirAC paradigm can be used in spatial audio coding and teleconference scenarios [8].
[0126] To the inventors' knowledge, no system has been built or proposed around a parameter-space audio codec, and even fewer have been built or proposed based on the DirAC paradigm. This is the subject of embodiments of the present invention.
[0127] MASA
[0128] Metadata-Assisted Spatial Audio (MASA) is a spatial audio format derived from the DirAC principle. It can be calculated directly from the raw microphone signal and transmitted to the audio codec without intermediate formats such as stereo reverberation. A set of parameters, such as directional parameters in the frequency band and / or energy ratio parameters in the frequency band (e.g., the proportion of directional sound energy), can also be used as spatial metadata for the audio codec or renderer. These parameters can be estimated from the audio signal captured by the microphone array; for example, mono or stereo signals can be generated from the microphone array signal to be transmitted along with the spatial metadata. Mono or stereo signals can be encoded, for example, using a core encoder such as 3GPP EVS or its derivatives. The decoder decodes the audio signal into sound in the frequency band and processes it (using the transmitted spatial metadata) to obtain a spatial output, which can be a binaural output, a speaker multichannel signal, or a multichannel signal in a stereo reverberation format.
[0129] motivation
[0130] Immersive voice communication is a new research area with very few systems in existence, and there are no DTX systems designed for such applications.
[0131] However, existing solutions can be easily combined. For example, DTX can be applied independently to each individual multichannel signal. This minimalist approach faces several problems. For these, each individual channel needs to be transmitted separately, which is incompatible with low bit-rate communication constraints and therefore almost incompatible with DTX, which is designed for low bit-rate communication situations. Furthermore, VAD decisions need to be synchronized across channels to avoid unusual events and unmasking effects, and the bit-rate reduction of the DTX system also needs to be fully utilized. In fact, to interrupt transmission and benefit from it, it is necessary to ensure that voice activity decisions are synchronized across all channels.
[0132] Another problem arises on the receiver side when lost background noise is generated during inactive frames using one or more comfort noise generators. For immersive communication, especially when DTX is applied directly to individual channels, a generator is required for each channel. If these generators, which typically sample random noise, are used independently, the coherence between channels will be zero or close to zero, potentially resulting in a perceptual deviation from the original soundscape. On the other hand, if only one generator is used and the resulting comfort noise is replicated across all output channels, the coherence will be extremely high, and the immersion will be significantly reduced.
[0133] These problems can be partially addressed by not applying DTX directly to the system's input or output channels, but instead applying it after a parametric spatial audio coding scheme such as DirAC to the resulting transport channels, which are typically downmixed or reduced versions of the original multichannel signal. In this case, it is necessary to define how inactive frames are parameterized and then spatialized via the DTX system. This is not insignificant and is the subject of embodiments of the invention. The spatial image must be consistent between active and inactive frames and must be perceptually as faithful as possible to the original background noise.
[0134] Figure 3 An encoder 300 is shown according to an example. Encoder 300 can generate an encoded audio scene 304 from an audio signal 302.
[0135] Audio signal 304 (bitstream) or audio scene 304 (and other audio signals disclosed below) may be divided into frames (e.g., a sequence of frames). Frames may be associated with time slots, which may subsequently define one another (in some examples, a previous aspect may overlap with a subsequent frame). For each frame, values in the time domain (TD) or frequency domain (FD) may be written into bitstream 304. In TD, values may be provided for each sample (each frame having, for example, a discrete sample sequence). In FD, values may be provided for each frequency interval. As explained later, each frame may be classified (e.g., by an activity detector) as an active frame 306 (e.g., a non-empty frame) or an inactive frame 308 (e.g., an empty frame, or a silent frame, or a noise-only frame). Active frames 306 and inactive frames 308 may also be associated to provide different parameters (e.g., active spatial parameter 316 or inactive spatial parameter 318) (in the absence of data, reference numeral 319 indicates no data provided).
[0136] Audio signal 302 may be, for example, a multi-channel audio signal (e.g., having two or more channels). Audio signal 302 may be, for example, a stereo audio signal. Audio signal 302 may be, for example, a stereo reverb signal in A or B format. Audio signal 302 may have, for example, a metadata-assisted spatial audio (MASA) format. Audio signal 302 may have an input format that is a first-order stereo reverb format, a higher-order stereo reverb format, a multi-channel format associated with a given speaker setup such as 5.1, 7.1, or 7.1+4, or one or more audio channels representing one or more different audio objects located in a space as indicated by information included in the associated metadata, or an input format that is a metadata-associated spatial audio representation. Audio signal 302 may contain microphone signals, such as those picked up by a real or virtual microphone. Audio signal 302 may contain synthesized microphone signals (e.g., in a first-order or higher-order stereo reverb format).
[0137] Audio scene 304 may include at least one or a combination of the following:
[0138] The first sound field parameter representation (e.g., active space parameter) 316 for the first frame 306;
[0139] The second sound field parameter representation (e.g., inactive spatial parameters) 318 is used for the second frame 308;
[0140] Encoded audio signal 346 for the first frame 306; and
[0141] Parameter description 348 for the second frame 308 (in some examples, inactive spatial parameter 318 may be included in parameter description 348, but parameter description 348 may also include other parameters that are not spatial parameters).
[0142] Active frames 306 (first frames) may be those frames containing speech (or, in some examples, other audio sounds that are different from pure noise). Inactive frames 308 (second frames) may be understood as those frames that do not contain speech (or, in some examples, other audio sounds that are different from pure noise) and may be understood as those frames that contain only noise.
[0143] A sound field parameter generator 310 may be provided, for example, to generate a transmission channel version 324 (subdivided among 326 and 328) of the audio signal 302. Reference may be made here to one or more transmission channels 326 for each first frame 306 and / or one or more transmission channels 328 for each second frame 308 (the one or more transmission channels 328 can be understood as providing, for example, a parameter description of silence or noise). One or more transmission channels 324 (326, 328) may be a downmixed version of the input format 302. Generally, if the input audio signal 302 is a stereo channel, each of the transmission channels 326, 328 may be, for example, a mono channel. If the input audio signal 302 has more than two channels, the downmixed version 324 of the input audio signal 302 may have fewer channels than the input audio signal 302, but in some examples, it may still have more than one channel (for example, if the input audio signal 302 has four channels, the downmixed version 324 may have one, two or three channels).
[0144] The sound field parameter generator 310 may additionally or alternatively provide sound field parameters (spatial parameters) indicated by 314. Specifically, the sound field parameters 314 may include an active spatial parameter (first spatial parameter or representation of first spatial parameter) 316 associated with the first frame 306, and an inactive spatial parameter (second spatial parameter or representation of second spatial parameter) 318 associated with the second frame 308. Each active spatial parameter 314 (316, 318) may include (e.g., may be) a parameter indicating spatial characteristics of the audio signal (302) relative to the listener's position. In some other examples, the active spatial parameters 314 (316, 318) may at least partially include (e.g., may be) a parameter indicating characteristics of the audio signal 302 relative to the speaker's position. In some examples, the active spatial parameters 314 (316, 318) may be or at least partially include characteristics such as those of an audio signal taken from a signal source.
[0145] For example, spatial parameter 314 (316, 318) may include diffusion parameters: such as one or more diffusion parameters indicating the ratio of diffused signal to sound in the first frame 306 and / or the second frame 308, or one or more energy ratio parameters indicating the ratio of direct sound to diffused sound energy in the first frame 306 and / or the second frame 308, or inter-channel / surround coherence parameters in the first frame 306 and / or the second frame 308, or one or more coherent diffusion power ratios in the first frame 306 and / or the second frame 308, or one or more signal diffusion ratios in the first frame 306 and / or the second frame 308.
[0146] In the example, one or more active spatial parameters (represented by the first sound field parameter) 316 and / or one or more inactive spatial parameters 318 (represented by the second sound field parameter) can be obtained from the input signal 302 in the form of its complete channel version or a subset thereof (such as the first-order component of the higher-order stereo reverberation input signal).
[0147] Device 300 may include an activity detector 320. The activity detector 320 may analyze an input audio signal (either in the form of its input version 302 or its downmixed version 324) to determine, based on the audio signal (302 or 324), whether a frame is an active frame 306 or an inactive frame 308, thereby performing frame classification. (The last sentence appears to be incomplete and possibly refers to a different device or system.) Figure 3 As can be seen, the activity detector 320 can be assumed to control (e.g., via control 321) the first offset 322 and the second offset 322a. The first offset 322 can select between an active spatial parameter 316 (represented by a first sound field parameter) and an inactive spatial parameter 318 (represented by a second sound field parameter). Thus, the activity detector 320 can decide whether to output (e.g., signal in bitstream 304) the active spatial parameter 316 or the inactive spatial parameter 318. The same control 321 can control the second offset 322a, which can select between outputting a first frame 326 (306) in the transmission channel 324 or outputting a second frame 328 (308) in the transmission channel 326 (e.g., parameter description). The activities of the first offset device 322 and the second offset device 322a are coordinated with each other: when the active spatial parameter 316 is output, the transmission channel 326 of the first frame 306 is also output subsequently, and when the inactive spatial parameter 318 is output, the transmission channel 328 of the transmission channel of the first frame 306 is also output subsequently. This is because the active spatial parameter 316 (represented by the first sound field parameter) describes the spatial characteristics of the first frame 306, while the inactive spatial parameter 318 (represented by the second sound field parameter) describes the spatial characteristics of the second frame 308.
[0148] The activity detector 320 can thus essentially determine which of the following outputs: the first frame 306 (326, 346) and its associated parameters (316), and the second frame 308 (328, 348) and its associated parameters (318). The activity detector 320 can also control the encoding of some signaling in the bitstream that signals whether a frame is active or inactive (other techniques may be used).
[0149] The activity detector 320 can perform processing on each frame 306 / 308 of the input audio signal 302 (e.g., by measuring the energy in all or at least several frequency ranges of the frame, for example, a specific frame of the audio signal), and can classify the specific frame as either a first frame 306 or a second frame 308. Generally, the activity detector 320 can determine a single classification result for a single complete frame, without distinguishing between different frequency ranges and different samples within the same frame. For example, a classification result could be "speech" (which would correspond to the first frames 306, 326, 346 spatially described by the active spatial parameter 316) or "silence" (which would correspond to the second frames 308, 328, 348 spatially described by the inactive spatial parameter 318). Therefore, based on the classification imposed by the activity detector 320, the biasers 322 and 322a can perform their exchange, and the result is, in principle, valid for all frequency ranges (and samples) of the classified frame.
[0150] Device 300 may include an audio signal encoder 330. The audio signal encoder 330 may generate an encoded audio signal 344. Specifically, the audio signal encoder 330 may provide an encoded audio signal 346 for a first frame (306, 326), for example, generated by a transport channel encoder 340, which may be a portion of the audio signal encoder 330. The encoded audio signal 344 may be or include a silent parameter description 348 (e.g., a noise parameter description) and may be generated by a transport channel SI descriptor 350, which may be a portion of the audio signal encoder 330. The generated second frame 348 may correspond to at least one second frame 308 of the original audio input signal 302 and to at least one second frame 328 of the downmixed signal 324, and may be spatially described by an inactive spatial parameter 318 (represented by a second sound field parameter). It is noteworthy that the encoded audio signal 344 (whether 346 or 348) may also be in the transport channel (and may therefore be the downmixed signal 324). Encoded audio signals 344 (whether 346 or 348) can be compressed to reduce their size.
[0151] Device 300 may include an encoded signal former 370. The encoded signal former 370 can write an encoded version of at least an encoded audio scene 304. The encoded signal former 370 operates by combining a first (active) sound field parameter representation 316 for a first frame 306, a second (inactive) sound field parameter representation 318 for a second frame 308, an encoded audio signal 346 for the first frame 306, and a parameter description 348 for the second frame 308. Therefore, the audio scene 304 can be a bitstream that can be transmitted or stored (or both) and used by a general-purpose decoder to generate an output audio signal that is a copy of the original input signal 302. In the audio scene (bitstream) 304, a sequence of "first frame" / "second frame" can thus be obtained to allow for the reproduction of the input signal 306.
[0152] Figure 2 Examples of encoder 300 and decoder 200 are shown. In some examples, encoder 300 can be used with... Figure 3 The encoder (or a variant thereof) is the same (in some other examples, it may be a different embodiment). Encoder 300 may receive an audio signal 302 as input (which may be, for example, in B format) and may have a first frame 306 (which may be, for example, an active frame) and a second frame 308 (which may be, for example, an inactive frame). The audio signal 302 may be provided to audio signal encoder 330 as a signal 324 (e.g., as an encoded audio signal 326 for the first frame and an encoded audio signal 328 for the second frame, or a parameter representation) after selection within selector 320 (which may include audio associated with biasers 322 and 322a). Notably, block 320 may also have the capability to downmix the input signals 302 (306, 308) to the transmission channels 324 (326, 328). Essentially, block 320 (beamforming / signal selection block) can be understood to include Figure 3 The activity detector 320 has the functionality, but Figure 3 Some other functions performed by block 310 (such as generating space parameters 316 and 318) can be performed by... Figure 2 The "DirAC analysis block" 310 is executed. Therefore, channel signals 324 (326, 328) can be downmixed versions of the original signal 302. However, in some cases, it is also possible that no downmixing is performed on signal 302, and signal 324 is only a selection between the first and second frames. Audio signal encoder 330 may include at least one of blocks 340 and 350, as explained above. Audio signal encoder 330 may output encoded audio signal 344 for the first frame 346 or for the second frame 348. Figure 2 The encoded signal former 370 is not shown, but it may be present.
[0153] As shown, block 310 may include a DirAC analysis block (or more generally, a sound field parameter generator 310). Block 310 (sound field parameter generator) may include a filter bank analysis 390. The filter bank analysis 390 may subdivide each frame of the input signal 302 into multiple frequency intervals, which may be the output 391 of the filter bank analysis 390. The diffusion estimation block 392a may, for example, provide a diffusion parameter 314a for each of the multiple frequency intervals 391 output by the filter bank analysis 390 (which may be one of one or more active spatial parameters 316 for active frame 306 or one of one or more inactive spatial parameters 318 for inactive frame 308). The sound field parameter generator 310 may include a direction analysis block 392b, the output of which 314b may be, for example, a direction parameter for each of a plurality of frequency intervals 391 output by the filter bank analysis 390 (which may be a direction parameter for one or more active spatial parameters 316 for active frame 306 or a direction parameter for one or more inactive spatial parameters 318 for inactive frame 308).
[0154] Figure 4 An example of block 310 (sound field parameter generator) is shown. Sound field parameter generator 310 can be used with... Figure 2 The sound field parameter generator is the same as and / or can be the same as... Figure 3 Block 310 is the same as or at least implements the function of block 310, although the fact is Figure 3 Block 310 can also perform downmixing of the input signal 302, but this fact is not shown (or not implemented) in Figure 4 In the sound field parameter generator 310.
[0155] Figure 4 The sound field parameter generator 310 may include a filter bank analysis block 390 (which can be used with...) Figure 2 (The filter bank analysis block 390 is the same). The filter bank analysis block 390 can provide frequency domain information 391 for each frame and each interval (frequency block). The frequency domain information 391 can be provided to... Figure 3The diffusion analysis blocks 392a and / or direction analysis blocks 392b shown are included. The diffusion analysis blocks 392a and / or direction analysis blocks 392b provide diffusion information 314a and / or direction information 314b. This information can be provided for each first frame 306 (346) and for each second frame 308 (348). In summary, the information provided by blocks 392a and 392b is considered as sound field parameters 314, which include first sound field parameters 316 (active space parameters) and second sound field parameters 318 (inactive space parameters). The active space parameters 316 can be provided to the active space metadata encoder 396, and the inactive space parameters 318 can be provided to the inactive space metadata encoder 398. The result is a representation of the first and second sound field parameters (316, 318, collectively indicated by 314) that can be encoded in a bitstream 304 (e.g., via encoder signal shaper 370) and stored for subsequent playback by a decoder. Whether the active spatial metadata encoder 396 or the inactive spatial parameter 318 will encode the frame can be determined by factors such as... Figure 3 Control 321 controls (deviation 322 not shown in) Figure 2 (This can be controlled, for example, by classification via an activity detector. Note that in some examples, encoders 396 and 398 can also perform quantization.)
[0156] Figure 5 Another example of a possible sound field parameter generator 310 is shown, which can be replaced Figure 4 A sound field parameter generator, and it can also be implemented in Figure 2 and Figure 3 In this example, the input audio signal 302 may already be in MASA format, where spatial parameters are, for example, a portion of the input audio signal 302 for each of multiple frequency ranges (e.g., as spatial metadata). Therefore, there is no need for a diffusion analysis block and / or a direction block, which can be replaced by a MASA reader 390M. The MASA reader 390M can read specific data fields in the audio signal 302 that already contain information such as one or more active spatial parameters 316 and one or more inactive spatial parameters 318 (depending on whether the frame of signal 302 is the first frame 306 or the second frame 308). Examples of parameters that can be encoded in signal 302 (and can be read by the MASA reader 390M) may include at least one of direction, energy ratio, surround coherence, dispersion coherence, etc. Downstream of the MASA reader 390M, an active spatial metadata encoder 396 (e.g., such as...) can be provided. Figure 4 The one in the middle) and the inactive spatial metadata encoder 398 (e.g., as Figure 4The first sound field parameter representation 316 and the second sound field parameter representation 318 are output respectively. If the input audio signal 302 is a MASA signal, the activity detector 320 can be implemented as an element that reads the determined data field in the input MASA signal 302 and classifies it into an active frame 306 or an inactive frame 308 based on the value encoded in the data field. Figure 5 The example can be generalized to an audio signal 302 in which spatial information has been encoded, which can be encoded as an active spatial parameter 316 or an inactive spatial parameter 318.
[0157] Embodiments of the present invention can be applied to, for example... Figure 2 The spatial audio coding system shown depicts a DirAC-based spatial audio encoder and decoder. The discussion follows.
[0158] The encoder 300 can typically analyze spatial audio scenes in B-format. Alternatively, DirAC analysis can be adapted to analyze different audio formats, such as audio objects or multichannel signals, or any combination of spatial audio formats.
[0159] DirAC analysis (e.g., performed at any of stages 392a, 392b) extracts a parameter representation from the input audio scene 302 (input signal). Direction of arrival (DOA) 314b and / or spread 314a measured per unit of time frequency form one or more parameters 316, 318. The DirAC analysis (e.g., performed at any of stages 392a, 392b) can be followed by a spatial metadata encoder (e.g., 396 and / or 398), which quantizes and / or encodes the DirAC parameters to obtain a low bit rate parameter representation (in the figures, the low bit rate parameter representations 316, 318 are indicated by the same reference numerals indicating the parameter representations upstream of the spatial metadata encoders 396 and / or 398).
[0160] Together with parameters 316 and / or 318, a downmixed signal 324 (326) derived from one or more different sources (e.g., different microphones) or one or more audio input signals (e.g., different components of a multichannel signal) 302 can be encoded by a conventional audio core encoder (e.g., for transmission and / or for storage). In a preferred embodiment, the EVS audio encoder (e.g., 330, Figure 2This can be preferably used to encode downmixed signals 324 (326, 328), but embodiments of the invention are not limited to this core encoder and can be applied to any audio core encoder. Downmixed signals 324 (326, 328) can consist of, for example, different channels also referred to as transport channels: signal 324 can be, for example, or contain four coefficient signals constituting a B-format signal, a stereo pair, or a mono tone downmix, depending on the target bit rate. The encoded spatial parameters 328 and the encoded audio bitstream 326 can be multiplexed before transmission (or storage) via the communication channels.
[0161] In the decoder (see below), transmit channel 344 is decoded by the core decoder, while DirAC metadata (e.g., spatial parameters 316, 318) can be decoded before being transmitted to the DirAC synthesizer along with the decoded transmit channels. The DirAC synthesizer uses the decoded metadata to control the reproduction of the direct sound stream and its mixture with the diffuse sound stream. The reproduced sound field can be reproduced on any speaker layout or generated in any order in a stereo reverberation format (HOA / FOA).
[0162] DirAC parameter estimation
[0163] This section explains non-limiting techniques for estimating spatial parameters 316, 318 (e.g., diffusivity 314a, orientation 314b). An example in B format is provided.
[0164] In each frequency band (e.g., obtained from filter bank analysis 390), the direction of arrival 314a of the sound, along with the sound dispersion 314b, can be estimated. From the input B-format components... From time-frequency analysis, the pressure and velocity vectors can be determined as follows:
[0165]
[0166]
[0167] in For the index of input 302, and and For time and frequency blocks, the time and frequency indices are provided, and This represents a Cartesian unit vector. In some examples, it may be necessary to... and The DirAC parameters (316, 318), i.e., DOA 314a and diffusivity 314a, can be calculated, for example, by calculating the intensity vector:
[0168] ,
[0169] in Indicates the complex conjugate. The diffusion of the combined sound field is given by the following formula:
[0170]
[0171] in Indicator time averaging operator, Represents the speed of sound and the energy of the sound field. It is given by the following formula:
[0172]
[0173] The diffusion of a sound field is defined as the ratio between sound intensity and energy density, with a value between 0 and 1.
[0174] Direction of Arrival (DOA) using unit vectors It indicates that it is limited to:
[0175]
[0176] The direction of arrival 314b can be determined by energy analysis of the input signal 302 in B format (e.g., at 392b) and can be defined as the relative direction of the intensity vector. The direction is defined in Cartesian coordinates but can be easily transformed, for example, in spherical coordinates defined by unit radius, azimuth, and elevation.
[0177] In the case of transmission, parameters 314a and 314b (316 and 318) need to be transmitted to the receiver (e.g., the decoder) via a bitstream (e.g., 304). For more robust transmission over networks with limited capacity, a low bitrate bitstream is preferred or even necessary, which can be achieved by designing an efficient coding scheme for DirAC parameters 314a and 314b (316 and 318). For example, techniques such as band grouping, prediction, quantization, and entropy coding by averaging parameters over different frequency bands and / or time units can be used. At the decoder, the transmitted parameters can be decoded for each time / frequency unit (k,n) assuming no errors occur in the network. However, if network conditions are not good enough to guarantee proper packet transmission, packets may be lost during transmission. Embodiments of the present invention aim to provide a solution to the latter situation.
[0178] decoder
[0179] Figure 6 An example of a decoder device 200 is shown. The decoder device may be a device for processing an encoded audio scene (304), the encoded audio scene including a first sound field parameter representation (316) and an encoded audio signal (346) in a first frame (346), wherein a second frame (348) is an inactive frame. The decoder device 200 may include at least one of the following:
[0180] An activity detector (2200) is used to detect that the second frame (348) is an inactive frame and to provide a parameter description (328) for the second frame (308).
[0181] A synthesized signal synthesizer (210) is used to synthesize a synthesized audio signal (228) for the second frame (308) using a parameter description (348) for the second frame (308).
[0182] An audio decoder (230) is used to decode the encoded audio signal (346) for the first frame (306); and
[0183] A spatial renderer (240) is used to render the audio signal (202) for the first frame (306) in space using the first sound field parameter representation (316) and the synthesized audio signal (228) for the second frame (308).
[0184] It is worth noting that the activity detector (2200) can issue a command 221' that determines whether an input frame is classified as an active frame 346 or an inactive frame 348. The activity detector 2200 can determine the classification of the input frame, for example, based on information 221, which is determined by signaling or from the length of the acquired frame.
[0185] The synthesized signal synthesizer (210) may, for example, use information obtained from parameter representation 348 (e.g., parameter information) to generate noise 228. The spatial renderer 220 may generate the output signal 202 by processing the inactive frame 228 (obtained from the encoded frame 348) through inactive spatial parameters 318 to obtain a 3D spatial impression of the source of noise for a human listener.
[0186] It should be noted that, Figure 6 In the text, numbers 314, 316, 318, 344, 346, and 348 are... Figure 3 The labels are the same because they correspond to each other due to being derived from bitstream 304. Nevertheless, some slight differences may exist (e.g., due to quantization).
[0187] Figure 6Also shown is control 221', which controls bias 224', allowing signal 226 (output by synthesizer 210) or audio signal 228 (output by audio decoder 230) to be selected, for example, by classification operated by activity detector 2200. Notably, signal 224 (226 or 228) can still be a downmixed signal, which can be provided to spatial renderer 220 so that the spatial renderer generates output signal 202 using active or inactive spatial parameters 314 (316, 318). In some examples, signal 224 (226 or 228) can also be upmixed so that the number of channels of signal 224 is increased relative to the encoded version 344 (346, 348). In some examples, although upmixed, the number of channels of signal 224 may be less than the number of channels of output signal 202.
[0188] Other examples of decoder devices 200 are provided below. Figures 7 to 10 Examples of decoder devices 700, 800, 900, and 1000 that can embody decoder device 200 are shown.
[0189] Even in Figures 7 to 10 Some components are shown as being inside the spatial renderer 220, but in some examples they may also be outside the spatial renderer 220. For example, the synthesized signal synthesizer 210 may be partially or completely outside the spatial renderer 220.
[0190] In those examples, a parameter processor 275 may be included (which may be inside or outside the spatial renderer 220). Although not shown, the parameter processor 275 may also be considered to exist within... Figure 6 In the decoder.
[0191] Figures 7 to 10 The parameter processor 275 of any of them may include, for example, an inactive spatial parameter decoder 278 and / or block 279 (“decoder for recovering spatial parameters in untransmitted frames”) for providing inactive frames that may be parameters 318 (e.g., obtained from signaling in bitstream 304), the block providing inactive spatial parameters that are not read in bitstream 304 but are obtained, for example, by extrapolation (e.g., recovery, reconstruction, extrapolation, inference, etc.) or synthesized.
[0192] Therefore, the second sound field parameter can also be represented as the generated parameter 219, which is not present in bitstream 304. As will be explained later, the recovered (reconstructed, extrapolated, inferred, etc.) spatial parameter 219 can be obtained, for example, through extrapolation from a "maintenance strategy" to a "direction strategy" and / or through "direction jitter" (see below). Therefore, the parameter processor 275 can extrapolate or obtain the spatial parameter 219 from previous frames in any way. Figures 6 to 9As can be seen, switching 275' can be selected between inactive spatial parameters 318 signaled in bitstream 304 and recovered spatial parameters 219. As explained above, the encoding of the silent frame 348 (SID) (and the encoding of inactive spatial parameters 318) is updated at a lower bit rate than the encoding of the first frame 346: inactive spatial parameters 318 are updated at a lower frequency relative to active spatial parameters 316, and some strategies are performed by parameter processor 275 (1075) to recover the unsignaled spatial parameters 219 used for inactive frames that have not been transmitted. Therefore, switching 275' can be selected between signaled inactive spatial parameters 318 and unsignaled (but recovered or otherwise reconstructed) inactive spatial parameters 219. In some cases, parameter processor 275' may store one or more sound field parameter representations 318 for several frames that appear before or after the second frame, extrapolating (or interpolating) the sound field parameters 219 used for the second frame. Generally, the spatial renderer 220 can use one or more sound field parameters 318 for the second frame 219 to render the synthesized audio signal 202 for the second frame 308. Alternatively, the parameter processor 275 can store a sound field parameter representation 316 for the active spatial parameters. Figure 10 (As shown in the diagram) and using the stored first sound field parameter representation 316 (active frame), sound field parameters 219 for the second frame (inactive frame) are synthesized to generate the restored spatial parameters 319. For example... Figure 10 As shown (and can also be implemented in) Figures 6 to 9 (In any of the), it may also include an active spatial parameter decoder 276, the active spatial parameter 316 being obtained from the bitstream 304 via the active spatial parameter decoder. This may be performed by extrapolation or interpolation to determine one or more sound field parameters for the second frame (308), wherein the direction of the jitter is included in at least two sound field parameter representations that occur temporally before or after the second frame (308).
[0193] The synthesizer 210 may be internal to or external to the spatial renderer 220, or in some cases, the synthesizer may have both internal and external portions. The synthesizer 210 may operate on the downmixed channels of the transmit channel 228 (which are fewer than the output channels) (note that M is the number of downmixed channels and N is the number of output channels). The synthesizer generator 210 (another name for the synthesizer) may generate, for the second frame, multiple synthesized component audio signals (in at least one of the transmit signal channels or in at least one individual component of the output audio format) as synthesized audio signals for individual components related to the external format of the spatial renderer. In some cases, this may be in the downmixed channel of the signal 228, and in some cases, it may be in one of the internal channels of the spatial renderer.
[0194] Figure 7 An example is shown in which at least K channels 228a obtained from the synthesized audio signal 228 (e.g., downstream of filter bank analysis 720 in version 228b of the synthesized audio signal) can be decorrelated. This is achieved, for example, when the synthesized signal synthesizer 210 generates the synthesized audio signal 228 in at least one of the M channels of the synthesized audio signal 228. This correlation processing 730 can be applied downstream of filter bank analysis block 720 to signal 228b (or at least one or some of its components) to obtain at least K channels (where K ≥ M and / or K ≤ N, where N is the number of output channels). Subsequently, the K decorrelated channels 228a and / or the M channels of signal 228b can be provided to block 740 to generate a mixing gain / matrix, which can provide a mixed signal 742 via spatial parameters 318, 219 (see above). The mixed signal 742 can be filtered and combined into block 746 to obtain output signals in N output channels 202. basically, Figure 7 Reference numeral 228a may be an individual synthesized component audio signal decoupled from individual synthesized component audio signals 228b, so that the spatial renderer (and block 740) utilizes the combination of component 228a and component 228b. Figure 8 This shows an example where all 228 audio channels are generated across K channels.
[0195] In addition, Figure 7 In this configuration, a decorrelation 730 applied to K decorrelation channels 228b is downstream of filter bank analysis block 720. This can be performed, for example, for a diffuse field. In some cases, M channels of signal 228b are downstream of feedback analysis block 720 and can be provided to block 744 for generating a mixing gain / matrix. Covariance methods can be used, for example, to reduce the problems of decorrelation 730 by scaling channels 228b with values associated with complementary values of covariance between the different channels.
[0196] Figure 8 An example of a synthesized signal synthesizer 210 in the frequency domain is shown. The covariance method can be used... Figure 8 The synthesized signal synthesizer 210 (810) is noteworthy. The synthesized audio synthesizer 210 (810) provides its output 228c in K channels (where K ≥ M), while the transmission channel 228 will be in M channels.
[0197] Figure 9 An example of decoder 900 (an embodiment of decoder 200) is shown, which can be understood as utilizing... Figure 8 decoder 800 and Figure 7The decoding 700 employs a mixing technique. As can be seen here, the synthesized signal synthesizer 210 includes a first section 210 (710) that generates a synthesized audio signal 228 in M channels of the downmixed signal 228. The signal 228 can be input to a filter bank analysis block 730, which provides an output 228b, wherein multiple filter bands are distinguished from each other. At this time, channels 228b can be decorrelated to obtain a decorrelated signal 228a in K channels. Simultaneously, the output 228b of the filter bank analysis in the M channels is provided to block 740 for generating a mixed gain matrix that provides a mixed version of the mixed signal 742. The mixed signal 742 may take into account the inactive spatial parameters 318 and / or the recovered (reconstructed) spatial parameters for the inactive frame 219. It should be noted that the output 228a of the decorrelation unit 730 can also be added at adder 920 to the output 228d of the second part 810 of the synthesizer 210, which provides the synthesized signal 228d across K channels. At adder 920, signal 228d can be summed to decorrelation signal 228a to provide summed signal 228e to mixer 740. Therefore, the final output signal 202 can be rendered using a combination of components 228b and 228e, where component 228e takes into account both decorrelation component 228a and the generated component 228d. Figure 8 and Figure 7 The components 228b, 228a, 228d, and 228e (existing) can be understood as, for example, the diffused and non-diffuse components of the synthesized signal 228. Specifically, refer to... Figure 9 In the decoder 900, the low-frequency band of signal 228e is basically obtained from the transmission channel 710 (and from 228a) and the high-frequency band of signal 228e is generated in the synthesizer 810 (and in the channel 228d). The addition of the low-frequency band and the high-frequency band at the adder 920 allows both to be present in signal 228e.
[0198] It is worth noting that, in the above Figures 7 to 10 The transmit channel decoder used for active frames is not shown in the diagram.
[0199] Figure 10 An example of decoder 1000 (an embodiment of decoder 200) is shown, wherein an audio decoder 230 (which provides decoded channels 226) and a synthesized signal synthesizer 210 are shown (here considered to be divided into a first external portion 710 and a second internal portion 810). A switch 224' is shown, which may be similar to... Figure 6The switching is controlled (e.g., by control or command 221' provided by activity detector 2200). Essentially, a choice can be made between a mode that provides the decoded audio scene 226 to the spatial renderer 220 and another mode that provides the synthesized audio signal 228. The downmixed signals 224 (226, 228) are in M channels, typically fewer than the N output channels of the output signal 202.
[0200] Signals 224 (226, 228) can be input to filter bank analysis block 720. The output 228b of filter bank analysis 720 (in multiple frequency ranges) can be input to upmixing adder block 750, which can also input signal 228d provided by the second part 810 of synthesizer 210. The output 228f of upmixing adder block 750 can be input to correlator processing 730. The output 228a of decorrelationer processing 730, together with the output 228f of upmixing adder block 750, can be provided to block 740 for generating mixing gain and matrix. Upmixing adder block 750 can, for example, increase the number of channels from M to K (and in some cases, it can multiply these channels, for example, by a constant factor) and can add K channels to K channels 228d generated by synthesizer 210 (e.g., the second internal part 810). To render the first (active) frame, the blend block 740 may consider at least one of the active space parameters 316 provided in the bitstream 304, or the recovered (reconstructed) space parameters 210 obtained by extrapolation or other means (see above).
[0201] In some examples, the output of filter bank analysis block 720 can be in M channels, but different frequency bands can be considered. For the first frame (and if located in...) Figure 10 The decoded signal 226 (in at least two channels) can be provided to filter bank analysis 720 via switches 224' and 222, and can thus be weighted at upmixing adder block 750 by K noise channels 228d (synthesized signal channels) to obtain signal 228f in K channels. It should be remembered that K ≥ M and may include, for example, diffuse and directional channels. In particular, the diffuse channel can be decorrelated by decorrelation 730 to obtain decorrelated signal 228a. Therefore, the decoded audio signal 224 can be weighted (e.g., at block 750) with the synthesized audio signal 228d, which masks the transition between active and inactive frames (first frame and second frame). Subsequently, the second part 810 of synthesized signal synthesizer 210 is used not only for active frames but also for inactive frames.
[0202] Figure 11Another example of decoder 200 is shown, which may include a first sound field parameter representation (316) and an encoded audio signal (346) in a first frame (346), wherein a second frame (348) is an inactive frame. The device includes: an activity detector (2200) for detecting that the second frame (348) is an inactive frame and for providing a parameter description (328) for the second frame (308); a synthesizer (210) for synthesizing a synthesized audio signal (228) for the second frame (308) using the parameter description (348) for the second frame (308); and an audio decoder (230) for decoding the first frame (306). Encoded audio signal (346); and spatial renderer (240) for spatially rendering audio signal (202) for first frame (306) using first sound field parameter representation (316) and synthesized audio signal (228) for second frame (308), or transcoder for generating metadata auxiliary output format containing audio signal (346) for first frame (306), first sound field parameter representation (316) for first frame (306), synthesized audio signal (228) for second frame (308) and second sound field parameter representation (318) for second frame (308).
[0203] Referring to the synthesized signal synthesizer 210 in the above example, as explained above, it may include (or even) a noise generator (e.g., a comfort noise generator). In the example, the synthesized signal generator (210) may include a noise generator, and a first additional synthesized component audio signal is generated by a first sample of the noise generator, and a second additional synthesized component audio signal is generated by a second sample of the noise generator, wherein the second sample is different from the first sample.
[0204] Alternatively, the noise generator includes a noise table, wherein a first additional synthesized component audio signal is generated by taking a first portion of the noise table, and wherein a second additional synthesized component audio signal is generated by taking a second portion of the noise table, wherein the second portion of the noise table is different from the first portion of the noise table.
[0205] In the example, the noise generator includes a pseudo-noise generator, and a first additional synthesized component audio signal is generated using a first seed for the pseudo-noise generator, and a second additional synthesized component audio signal is generated using a second seed for the pseudo-noise generator.
[0206] Generally speaking, in Figure 6 , Figure 7 , Figure 9 , Figure 10 and Figure 11In the example, the spatial renderer 220 can use a mixture of a direct signal and a diffused signal generated from the direct signal by the decorrelation unit (730) under the control of the first sound field parameter representation (316) to operate in a first mode for the first frame (306), and use a mixture of a first synthesized component signal and a second synthesized component signal to operate in a second mode for the second frame (308), wherein the first synthesized component signal and the second synthesized component signal are generated by the synthesized signal synthesizer (210) through different implementations of noise processing or pseudo-noise processing.
[0207] As explained above, the spatial renderer (220) can be configured to control the blending (740) in the second mode by using the diffusion parameter, energy distribution parameter or coherence parameter derived by the parameter processor for the second frame (308).
[0208] The above example also relates to a method for generating an encoded audio scene from an audio signal having a first frame (306) and a second frame (308), comprising: determining a first sound field parameter representation (316) for the first frame (306) from the audio signal in the first frame (306), and determining a second sound field parameter representation (318) for the second frame (308) from the audio signal in the second frame (308); analyzing the audio signal to determine, based on the audio signal, that the first frame (306) is an active frame and the second frame (308) is an inactive frame; generating an encoded audio signal for the first frame (306) as an active frame and generating a parameter description (348) for the second frame (308) as an inactive frame; and constructing an encoded audio scene by combining the first sound field parameter representation (316) for the first frame (306), the second sound field parameter representation (318) for the second frame (308), the encoded audio signal for the first frame (306), and the parameter description (348) for the second frame (308).
[0209] The above examples also relate to a method for processing an encoded audio scene, wherein the encoded audio scene includes a first sound field parameter representation (316) and an encoded audio signal in a first frame (306), and a second frame (308) is an inactive frame, the method comprising: detecting that the second frame (308) is an inactive frame and providing a parameter description (348) for the second frame (308); synthesizing a synthesized audio signal for the second frame (308) using the parameter description (348) for the second frame (308) (228); and decoding the audio signal for the first frame (306). Encoded audio signal; and spatially rendering audio signal for first frame (306) using first sound field parameter representation (316) and synthesized audio signal (228) for second frame (308), or generating metadata auxiliary output format containing audio signal for first frame (306), first sound field parameter representation (316) for first frame (306), synthesized audio signal (228) for second frame (308) and second sound field parameter representation (318) for second frame (308).
[0210] It also provides an encoded audio scene (304) comprising: a first sound field parameter representation (316) for a first frame (306); a second sound field parameter representation (318) for a second frame (308); an encoded audio signal for the first frame (306); and a parameter description (348) for the second frame (308).
[0211] In the above example, spatial parameters 316 and / or 318 can be transmitted for each frequency band (sub-band).
[0212] Based on some examples, this silent parameter description 348 may contain this part parameter 318, which may therefore be part of SID 348.
[0213] Spatial parameter 318 for inactive frames can be valid for each sub-band (or frequency band or frequency).
[0214] The spatial parameters 316 and / or 318 transmitted or encoded in SID 348 during the active phase 346 discussed above may have different frequency resolutions, and additionally or alternatively, the spatial parameters 316 and / or 318 transmitted or encoded in SID 348 during the active phase 346 discussed above may have different time resolutions, and additionally or alternatively, the spatial parameters 316 and / or 318 transmitted or encoded in SID 348 during the active phase 346 discussed above may have different quantization resolutions.
[0215] It should be noted that the decoding and encoding devices can be devices such as CELP, DCX, or bandwidth extension modules.
[0216] Alternatively, an MDCT-based coding scheme (improved discrete cosine transform) can be used.
[0217] In this example of decoder device 200 (in any of its embodiments, for example) Figures 6 to 11 In those embodiments), an audio decoder 230 and a spatial renderer 240 can be replaced by a transcoder for generating a metadata auxiliary output format that includes an audio signal for a first frame, a first sound field parameter representation for the first frame, a synthesized audio signal for a second frame, and a second sound field parameter representation for the second frame.
[0218] Discussion
[0219] Embodiments of the present invention propose a method for extending DTX to parametric spatial audio coding. Therefore, it is proposed to apply conventional DTX / CNG to the downmixing / transmit channels (e.g., 324, 224) to extend the downmixing / transmit channels using spatial parameters (referred to as rear spatial SIDs) such as 316, 318, and to apply spatial rendering to inactive frames (e.g., 308, 328, 348, 228) on the decoder side. To recover the spatial image of the inactive frames (e.g., 308, 328, 348, 228), the transmit channel SIDs 326, 226 are corrected using some specially designed spatial parameters (spatial SIDs) 319 (or 219) that are related to immersive background noise. Embodiments of the present invention (discussed below and / or above) cover at least two aspects:
[0220] • Extend the transmit channel SID for spatial rendering. For this purpose, spatial parameters 318 are modified using, for example, spatial parameters derived from the DirAC paradigm or MASA format. At least one of the parameters 318, such as diffusion 314a and / or one or more directions of arrival 314b and / or inter-channel / surround coherence and / or energy ratio, may be transmitted along with the transmit channel SID 328 (348). In some cases and under certain assumptions, some parameters 318 may be discarded. For example, if it is assumed that background noise is fully diffused, the subsequent transmission of directions 314b, which are meaningless, may be discarded.
[0221] • Spatialization of inactive frames at the receiver side by rendering the transmit channel CNG in space: This can be guided by the DirAC synthesis principle or its derivatives, based on the spatial parameters 318 of the final transmission within the spatial SID descriptor of the background noise. At least two options exist, which can even be combined: transmit channel comfort noise generation can be performed only for transmit channel 228 (this is...) Figure 7 In the case where comfort noise 228 is generated by synthesizer 710; or transport channel CNG (this is) can be generated for the transport channel and additional channels used for upmixing in the renderer. Figure 9 In some cases, some comfort noise 228 is generated by the first part 710 of the synthesized signal synthesizer, while other comfort noise 228d is generated by the second part 810 of the synthesized signal synthesizer. In more recent cases, for example, the second part 710 of CNG sampling random noise 228d using different seeds can automatically decorrelate the generated channels 228d and minimize the use of the decorrelator 730, which can be a typical pseudo-sound source. In addition, CNG can also be used in active frames (such as... Figure 10 (as shown in the example), but in some examples, the transition between active and inactive phases (frames) is smoothed with reduced intensity, and the final pseudo-voice from the transport channel encoder and the parameter DirAC paradigm is also masked.
[0222] Figure 3 An overview of an embodiment of encoder device 300 is provided. At the encoder side, the signal can be analyzed using DirAC analysis. DirAC can analyze signals such as B format or first-order stereo reverberation (FOA). However, the principle can be extended to higher-order stereo reverberation (HOA), and even to multi-channel signals associated with a given loudspeaker setup such as 5.1, 7.1, or 7.1+4 as proposed in
[10] . Input format 302 can also be an individual audio channel representing one or more different audio objects located in a space indicated by information included in the associated metadata. Alternatively, input format 302 can be metadata-associated spatial audio (MASA). In this case, spatial parameters and the transmitted channel are directly transmitted to encoder device 300. Audio scene analysis (e.g., as presented in
[10] ) can then be skipped. Figure 5 As shown in the figure, the final spatial parameter (re)quantization and resampling only need to be performed on either the inactive set 318 of the spatial parameters or on either the active and inactive sets 316 and 318 of the spatial parameters.
[0223] Audio scene analysis can be performed on active and inactive frames 306, 308, generating two sets 316, 318 of spatial parameters. A first set 316 is generated for active frame 308, and another set (318) is generated for inactive frame 308. Inactive spatial parameters may not be present, but in a preferred embodiment of the invention, inactive spatial parameters 318 are less and / or more coarsely quantized compared to active spatial parameters 316. Thereafter, two versions of spatial parameters (also referred to as DirAC metadata) are obtained. Importantly, embodiments of the invention can primarily concern the spatial representation of the audio scene from the listener's perspective. Therefore, spatial parameters such as DirAC parameters 318, 316 are considered, including one or more directions along with a final diffusion factor or one or more energy ratios. Unlike inter-channel parameters, these spatial parameters from the listener's perspective offer significant advantages for agnostic sound capture and reproduction systems. This parameterization is not specific to any particular microphone array or speaker layout.
[0224] A voice activity detector (or more generally, an activity detector) 320 may then be applied to the input signal 302 and / or the transport channel 326 generated by the audio scene analyzer. The transport channel is fewer than the number of input channels; it is typically a mono downmixed, stereo downmixed, A-format, or first-order stereo reverberation signal. Based on VAD decisions, the current frame under processing is defined as active (306, 326) or inactive (308, 328). In the case of active frames (306, 326), conventional speech or audio coding of the transport channel is performed. The resulting code data is then combined with the active spatial parameter 316. In the case of inactive frames (308, 328), typically during the inactive phase, at regular frame intervals, such as every 8 active frames (306, 326, 346), a silence information description 328 for the transport channel 324 is occasionally generated. The transmit channel SIDs (328, 348) can then be corrected in the multiplexer (via the encoded signal shaper) 370 using the inactive space parameter. When the inactive space parameter 318 is empty, only the transmit channel SID 348 is subsequently transmitted. The total SID can typically be described by a very low bit rate, such as as low as 2.4 or 4.25 kbps. During the inactive phase, the average bit rate is even lower because no transmission occurs and no data is sent for most of the time.
[0225] In a preferred embodiment of the invention, the transmit channel SID 348 has a size of 2.4 kbps, and the total SID including spatial parameters has a size of 4.25 kbps. For DirAC with a multi-channel signal such as FOA as input, the calculation of inactive spatial parameters is... Figure 4 As described in [the document], for the MASA input format, in [the document] Figure 5In this context, inactive spatial parameters can be directly derived from higher-order stereo reverberation (HOA). As previously mentioned, inactive spatial parameters 318 can be derived in parallel with active spatial parameters 316, thereby averaging and / or requantizing the encoded active spatial parameters 318. In the case of a multi-channel signal such as FOA as input format 302, filter bank analysis of the multi-channel signal 302 can be performed for each time and frequency block before calculating spatial parameters, direction, and spread. Metadata encoders 396 and 398 can average parameters 316 and 318 across different frequency bands and / or time slots before applying quantizers and encoding quantized parameters. Other inactive spatial metadata encoders can inherit some of the quantized parameters derived in the active spatial metadata encoder to directly use them in the inactive spatial parameters or requantize them. In the case of MASA format (e.g.) Figure 5 The input metadata can first be read and provided to the metadata encoders 396 and 398 at a given time frequency and bit depth resolution. One or more metadata encoders 396 and 398 will then further process it by: finally transforming some parameters, adjusting its resolution (i.e., reducing the resolution, for example, averaging it), and requantizing these parameters before encoding them, for example, by an entropy coding scheme.
[0226] For example Figure 6 As depicted, on the decoder side, VAD information 221 (e.g., whether the frame is classified as active or inactive) is first recovered by detecting the size of the transmitted packets (e.g., frames) or by detecting untransmitted packets. In active frames 346, the decoder operates in active mode, and the transmit channel encoder payload and active spatial parameters are decoded. Spatial renderer 220 (DirAC synthesizer) then upmixes / spatializes the decoded transmit channels using the decoded spatial parameters 316, 318 in output spatial format. In inactive frames, the transmit channel CNG portion 810 (e.g., in...) can be used to... Figure 10 Comfort noise is generated in the transmit channel. The CNG is guided by the transmit channel SID for general adjustments to energy and spectral shape (e.g., by scaling factors applied in the frequency domain or linear prediction coding coefficients applied to the time-domain synthesis filter). One or more comfort noises 228d, 228a, etc., are then rendered / spatialized in a spatial renderer (DirAC synthesis) 740, guided at this time by the inactive spatial parameter 318. The output spatial format 202 can be a binaural signal (2 channels), a multi-channel signal for a given speaker layout, or a multi-channel signal in a stereo reverberation format. In an alternative embodiment, the output format can be metadata-assisted spatial audio (MASA), which means that the decoded transmit channel or transmit channel comfort noise, along with the active or inactive spatial parameters, is directly output for rendering by an external device.
[0227] Encoding and Decoding of Inactive Spatial Parameters
[0228] The inactive spatial parameter 318 may consist of one of multiple directions in the frequency band and an associated energy ratio in the frequency band corresponding to the ratio of a directional component to the total energy. In the case of one direction, as in a preferred embodiment, the energy ratio may be replaced by spread, which is complementary to the energy ratio and subsequently follows the original DirAC set of parameters. Since one or more directional components are generally expected to be less correlated with the spread portion in inactive frames, they may be transmitted on fewer bits, such as by using a coarser quantization scheme in active frames and / or by averaging the direction over time or frequency to obtain a coarser time and / or frequency resolution. In a preferred embodiment, the direction may be transmitted every 20 ms instead of 5 ms for active frames, but using the same frequency resolution across 5 non-uniform frequency bands.
[0229] In a preferred embodiment, spread 314a can be transmitted at the same time / frequency as in the active frame but on fewer bits, thereby forcing the implementation of a minimum quantization index. For example, if spread 314a is quantized on 4 bits in the active frame, it is subsequently transmitted on only 2 bits, thus avoiding the transmission of the original index from 0 to 3. The decoded index is then incremented by an offset of +4.
[0230] In some examples, the direction 314b can be avoided entirely, or alternatively, the spread 314a can be avoided and replaced with the default value or estimate at the decoder.
[0231] Furthermore, if the input channel corresponds to a channel located in the spatial domain, then the coherence between transmission channels can be considered. The sound level difference between channels is also an alternative for directionality.
[0232] More relevant is the transmission of surround coherence, which is defined as the ratio of coherent diffuse energy in the sound field. This surround coherence can be utilized, for example, at the spatial renderer (DirAC synthesizer) by redistributing energy between the direct and diffuse signals. The energy of the surround coherent component is removed from the diffuse energy to be redistributed to the directional component, which will then be translated more uniformly in space.
[0233] Naturally, for inactive space parameters, any combination of the parameters listed above can be considered. For bit-saving purposes, it is also conceivable that no parameters are sent during the inactive phase.
[0234] Example pseudocode for an inactive spatial metadata encoder is given below:
[0235]
[0236]
[0237]
[0238]
[0239]
[0240] Recovering spatial parameters without transmission at the decoder side
[0241] In the case of SID during the inactive phase, the spatial parameters can be fully or partially decoded and subsequently used for DirAC synthesis.
[0242] In the absence of data transmission, or in the absence of spatial parameter 318 transmitted along with the transmission channel 348, it may be necessary to recover spatial parameter 219. This can be achieved by synthesizing the lost parameter 219 (e.g., considering previously received parameters (e.g., 316 and 7 or 318)). Figures 7 to 10 This is achieved through [various methods]. Unstable spatial images can be perceptually uncomfortable, especially in the face of background noise that is perceived as stable and not rapidly evolving. On the other hand, absolutely constant spatial images can be perceived as unnatural. Different strategies can be applied:
[0243] Maintenance strategy
[0244] It is generally safe to assume that spatial images must remain relatively stable over time, and spatial images can be translated based on DirAC parameters, i.e., DOA and diffusion that do not change significantly between frames. For this reason, a simple but effective approach is to retain the last received spatial parameters 316 and / or 318 as the recovered spatial parameters 219. This is a highly robust approach, at least for diffusion, which has long-term characteristics. However, for orientation, different strategies can be envisioned, as listed below.
[0245] extrapolation of direction:
[0246] Alternatively or additionally, it is conceivable to estimate the trajectories of sound events in the audio scene and then attempt to extrapolate the estimated trajectories. This is particularly meaningful when the sound events are well-located in space as point sources, which are reflected by low diffusion in the DirAC model. The estimated trajectories can be computed from observations in past directions and by fitting curves to these points, which can evolve into interpolation or smoothing. Regression analysis can also be employed. Extrapolation of parameter 219 can then be performed by estimating a fitted curve that exceeds the range of the observed data (e.g., including the previously mentioned parameters 316 and / or 318). However, this method can result in low relevance to inactive frames 348, which are useless for background noise and are expected to have extremely high diffusion.
[0247] Directional jitter:
[0248] When sound events are more diffuse (especially in the case of background noise), direction becomes less meaningful and can be considered a randomized implementation. Jitter can then help make the rendered sound field more natural and pleasant by injecting random noise into the previous direction before transmitting frames. The injected noise and its variance can depend on the diffuseness. For example, the variance of the injected noise at azimuth and elevation angles... and A simple model function that can be followed for the diffusion degree Ψ is as follows:
[0249]
[0250] Comfort noise generation and spatialization (decoder side)
[0251] The examples provided above will now be discussed.
[0252] In the first embodiment, in such Figure 7 The core decoder described herein implements a comfort noise generator 210 (710). The resulting comfort noise is injected into the transmit channel and then spatialized in DirAC synthesis using the transmitted inactive spatial parameter 318 or, in the absence of transmission, the spatial parameter 219 derived as previously described. Spatialization can then be achieved, for example, by generating two streams, a directional stream and a non-directional stream derived from the decoded transmit channel, and, in the case of an inactive frame, the comfort noise from the transmit channel. The two streams are then upmixed and combined at block 740 according to the spatial parameter 318.
[0253] Alternatively, comfort noise, or a portion thereof, can be generated directly within the DirAC synthesis in the filter bank domain. In practice, DirAC can control the coherence of the reconstructed scene by means of the transport channel 224, spatial parameters 318, 316, 319, and some decorrelation (e.g., 730). The decorrelation 730 reduces the coherence of the synthesized sound field. The spatial image is then perceived with more width, depth, diffusion, reverberation, or externalization in the case of headphone reproduction. However, decorrelation often tends to be typical audible artifacts, and its use is desirable to be reduced. This can be achieved, for example, by using the pre-existing incoherent components of the transport channel through the so-called covariance synthesis method [5]. However, this method can have limitations, especially in the case of a single-tone transport channel.
[0254] When comfort noise is generated from random noise, it is advantageous to generate dedicated comfort noise for each output channel or at least a subset thereof. More specifically, it is advantageous to apply comfort noise generation not only to the transport channels but also to the intermediate audio channels used in the spatial renderer (DirAC synthesizer) 220 (and in the mixing block 740). The decorrelation of the diffusion field is then directly provided by using a different noise generator instead of using a decorrelator 730, which reduces the amount of artifacts and the overall complexity. In fact, by definition, different implementations of random noise are decorrelated. Figure 8 and Figure 9 This illustrates two methods of achieving this by generating comfortable noise, either entirely or partially, within the spatial renderer 220. Figure 8 In this context, CN is performed in the frequency domain as described in [5], and it can be generated directly using the filter bank domain of the spatial renderer, thus avoiding filter bank analysis 720 and decorrelation 730. Here, the number of channels K for generating comfortable noise is equal to or greater than the number of transmission channels M, and less than or equal to the number of output channels N. In the simplest case, K=N.
[0255] Figure 9 An alternative approach is shown that includes comfort noise generation 810 within the renderer. Comfort noise generation is divided into internal (at 710) and external (at 810) stages within the spatial renderer 220. Comfort noise 228d within the renderer 220 is added (at adder 920) to the final decorrelation output 228a. For example, low-frequency bands can be generated outside the same domain as in the core encoder to facilitate easy updates to the required memory. On the other hand, for high frequencies, comfort noise generation can be performed directly within the renderer.
[0256] In addition, comfort noise generation can be applied during active frame 346. Instead of completely disabling comfort noise generation during active frame 346, it can be kept active by reducing its intensity. It is then worthwhile to mask the transition between active and inactive frames, as well as mask artifacts and defects in the core encoder and the parametric spatial audio model. This is proposed for monophonic speech coding in
[11] . The same principle can be extended to spatial speech coding. Figure 10 The implementation method is illustrated. Comfort noise generation in the spatial renderer 220 is activated during both the active and inactive phases. In the inactive phase 348, comfort noise generation in the renderer is complementary to comfort noise generation performed in the transport channels. In the renderer, comfort noise is achieved on K channels equal to or greater than M transport channels, aiming to reduce the use of decorrelation. The comfort noise generation in the spatial renderer 220 is added to the upmix version 228f of the transport channels, which can be achieved through a simple copy from M channels to K channels.
[0257] aspect
[0258] For encoders:
[0259] 1. An audio encoder device (300) for encoding a spatial audio format having multiple channels or one or more audio channels using metadata describing an audio scene, comprising at least one of the following:
[0260] a. A sound field parameter generator (310) for a spatial audio input signal (302) is configured to generate a first set or a first set and a second set of spatial parameters (318, 319) describing a spatial image and downmixed version (326) of an input signal (202) containing one or more transmission channels, wherein the number of transmission channels is less than the number of input channels;
[0261] b. A transmission channel encoder device (340) is configured to generate encoded data (346) during an active phase (306) by encoding a downmixed signal (326) containing the transmission channel.
[0262] c. A silent insertion descriptor (350) for generating a silent insertion description (348) of the background noise of the transmission channel (328) during the inactive phase (308).
[0263] d. A multiplexer (370) for combining a first set of spatial parameters (318) with encoded data (344) into a bit stream (304) during an active phase (306), and for not transmitting data or transmitting a silent insert description (348) during an inactive phase (308), or for combining a silent insert description (348) with a second set of spatial parameters (318).
[0264] 2. The audio encoder as described in 1, wherein the sound field parameter generator (310) follows the principle of directional audio coding (DirAC).
[0265] 3. The audio encoder as described in 1, wherein the sound field parameter generator (310) interprets the input metadata along with one or more transmission channels (348).
[0266] 4. The audio encoder as described in 1, wherein the sound field parameter generator (310) derives one or two sets of parameters (316, 318) from input metadata and derives a transmission channel from one or more input audio channels.
[0267] 5. The audio encoder as described in 1, wherein the spatial parameters are one or more directions of arrival (DOA) (314b), or diffusion (314a), or one or more coherences.
[0268] 6. The audio encoder as described in 1, wherein spatial parameters are derived for different sub-bands.
[0269] 7. The audio encoder as described in 1, wherein the transmission channel encoding device follows the CELP principle, or is an MDCT-based encoding scheme or a combination of the two schemes.
[0270] 8. The audio encoder as described in 1, wherein the active phase (306) and the inactive phase (308) are determined by a voice activity detector (320) performed on the transmission channel.
[0271] 9. The audio encoder as described in 1, wherein the first and second sets of spatial parameters (316, 318) differ in terms of time or frequency resolution, or quantization resolution, or the nature of the parameters.
[0272] 10. The audio encoder as described in 1, wherein the spatial audio input format (202) is a stereo reverb format or B format, or a multichannel signal associated with a given speaker setup, or a multichannel signal derived from a microphone array, or a set of individual audio channels along with metadata, or metadata-assisted spatial audio (MASA).
[0273] 11. The audio encoder as described in 1, wherein the spatial audio input format consists of two or more audio channels.
[0274] 12. The audio encoder as described in 1, wherein the number of transmission channels is 1, 2 or 4 (other numbers may be selected).
[0275] For the decoder:
[0276] 1. An audio decoder device (200) for decoding a bitstream (304) to generate a bitstream from a spatial audio output signal (202), the bitstream (304) comprising at least an active phase (306) followed by at least an inactive phase (308), wherein the bitstream has encoded therein at least a silent insert descriptor frame S1D (348), the silent insert descriptor frame describing background noise characteristics and / or spatial image information of a transport / downmixing channel (228), the audio decoder device (200) comprising at least one of the following:
[0277] a. A silent insert descriptor decoder (210) is configured to decode a silent S1D (348) to reconstruct background noise in the transmit / downmix channel (228);
[0278] b. Decoding device (230) is configured to reconstruct the transmission / downmixing channel (226) from the bitstream (304) during the active phase (306);
[0279] c. A spatial rendering apparatus (220) is configured to reconstruct the spatial output signal (740) (202) from the decoded transmit / downmix channel (224) and the transmitted spatial parameters (316) during an active phase (306), and to reconstruct the spatial output signal from the reconstructed background noise in the transmit / downmix channel (228) during an inactive phase (308).
[0280] 2. The audio decoder as described in 1, wherein the spatial parameters (316) transmitted during the active phase consist of diffusion or direction of arrival or coherence.
[0281] 3. The audio decoder as described in 1, wherein spatial parameters (316, 318) are transmitted via sub-bands.
[0282] 4. The audio decoder as described in 1, wherein the silent insertion description (348) contains, in addition to the background noise characteristics of the transmission / downmixing channel (228), spatial parameters (318).
[0283] 5. The audio decoder as described in 4, wherein the parameters (318) transmitted in the SID (348) may consist of diffusion or direction of arrival or coherence.
[0284] 6. The audio decoder as described in 4, wherein the spatial parameters (318) transmitted in SID (348) are transmitted via sub-bands.
[0285] 7. The audio decoder as described in 4, wherein the spatial parameters (316, 318) transmitted or encoded during the active phase (346) and in the SID (348) have different frequency resolutions or time resolutions or quantization resolutions.
[0286] 8. The audio decoder as described in 1, wherein the spatial renderer (220) may be composed of:
[0287] a. Decorrelation unit (730) for obtaining a decorrelation version (228b) of one or more decoded transmit / downmixed channels (226) and / or reconstructed background noise (228);
[0288] b. Upmixer for deriving an output signal from one or more decoded transmit / downmixed channels (226) or reconstructed background noise (228) and its decorrelated version (228b) and from spatial parameters (348).
[0289] 9. The audio decoder as described in 8, wherein the upmixer of the spatial renderer comprises:
[0290] a. At least two noise generators (710, 810) for generating at least two decorrelational background noises (228, 228a, 228d) having the characteristics described in the silent descriptor (448) and / or the characteristics given by noise estimation applied to the active phase (346).
[0291] 10. The audio decoder as described in 9, wherein the generated decorrelational background noise in the upmixer is mixed with the decoded transmit channel or the reconstructed background noise in the transmit channel, taking into account the spatial parameters transmitted during the active phase and / or the spatial parameters included in the SID.
[0292] 11. An audio decoder as described in any of the preceding aspects, wherein the decoding device comprises a voice encoder such as CELP or a general audio encoder or bandwidth extension module such as TCX.
[0293] Other features of the attached figures
[0294] Figure 1: DirAC analysis and synthesis from [1].
[0295] Figure 2 Detailed block diagram of DirAC analysis and synthesis in a low bit-rate 3D audio encoder.
[0296] Figure 3 : Block diagram of the decoder.
[0297] Figure 4 : Block diagram of the audio scene analyzer in DirAC mode.
[0298] Figure 5 : Block diagram for audio scene analyzer using MASA input format.
[0299] Figure 6 : Block diagram of the decoder.
[0300] Figure 7 : A block diagram of the spatial renderer (DirAC composition), where CNG in the transmission channels is outside the renderer.
[0301] Figure 8 : A block diagram of the spatial renderer (DirAC composition), in which CNG is directly performed on K channels in the renderer's filter bank domain, where K>=M transport channels.
[0302] Figure 9 : A block graph of the spatial renderer (DirAC composition), in which CNG is performed both outside and inside the spatial renderer.
[0303] Figure 10: A block graph of the spatial renderer (DirAC composition), in which CNG is performed both outside and inside the spatial renderer and CNG is enabled for active and inactive frames.
[0304] Advantages
[0305] Embodiments of the present invention allow for the efficient extension of DTX to parameter-space audio coding. Even for inactive frames, this enables the recovery of background noise with high perceptual fidelity, and transmission can be interrupted to save communication bandwidth.
[0306] To this end, the SID of the transport channel is extended by inactive spatial parameters associated with a spatial image describing the background noise. The generated comfort noise is applied to the transport channel before being spatialized by the renderer (DirAC synthesis). Alternatively, to improve quality, CNG can be applied to more channels than the transport channel within the render. This allows for reduced complexity and minimizes the annoyance of decorrelation artifacts.
[0307] Other aspects
[0308] It should be mentioned here that all the alternatives or aspects discussed above, as well as all aspects defined by independent aspects, can be used individually, i.e., without any other alternatives or objects other than the intended alternatives, objects, or independent aspects. However, in other embodiments, two or more of the alternatives or aspects or independent aspects may be combined with each other, and in other embodiments, all aspects or alternatives and all independent aspects may be combined with each other.
[0309] The encoded signal of the present invention can be stored on a digital storage medium or a non-transitory storage medium, or can be transmitted on a transmission medium, such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0310] Although some aspects are described in the context of the device, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of the method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block, item, or feature of the corresponding device.
[0311] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software. Implementations may be performed using digital storage media, such as floppy disks, DVDs, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory, which store electronically readable control signals that cooperate (or are capable of cooperating) with a programmable computer system.
[0312] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0313] Generally, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, is operatively used to perform one of the methods. The program code may, for example, be stored on a machine-readable medium.
[0314] Other embodiments include a computer program for performing one of the methods described herein, the computer program being stored on a machine-readable carrier or a non-transitory storage medium.
[0315] In other words, therefore, an embodiment of the method of the present invention is a computer program having program code that, when run on a computer, performs one of the methods described herein.
[0316] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) containing a computer program recorded thereon for performing one of the methods described herein.
[0317] Therefore, another embodiment of the method of the present invention is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection, such as via the Internet.
[0318] Another embodiment includes a processing component, such as a computer or programmable logic device configured or adapted to perform one of the methods described herein.
[0319] Another embodiment includes a computer having a computer program installed on it for performing one of the methods described herein.
[0320] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0321] The above embodiments are merely illustrative of the principles of the invention. It should be understood that modifications and variations to the arrangements and details described herein will be apparent to those skilled in the art. Therefore, it is intended that the scope be limited only by the following patent aspects and not by the specific details presented in the description and explanation of the embodiments herein.
[0322] The subsequent defining aspects of the first set of embodiments and the second set of embodiments can be combined such that certain features of one set of embodiments can be included in another set of embodiments.
Claims
1. An apparatus (300) for generating an encoded audio scene (304) from an audio signal (302) having a first frame (306) and a second frame (308), comprising: A sound field parameter generator (310) is configured to determine a first sound field parameter representation (316) for the first frame (306) from the audio signal (302) in the first frame (306) and a second sound field parameter representation (318) for the second frame (308) from the audio signal (302) in the second frame (308); and An activity detector (320) is used to analyze the audio signal (302) to determine, based on the audio signal (302), that the first frame is an active frame (306) and the second frame is an inactive frame (308). The sound field parameter generator (310) is configured to determine one or more individual sound sources from the second frame (308) of the audio signal and determine a parameter description (328) for each sound source for the second frame. The sound field parameter generator (310) is configured to decompose the second frame (308) into one or more frequency intervals, each frequency interval representing an individual sound source among one or more individual sound sources, and to determine at least one inactive spatial parameter for each frequency interval as a second sound field parameter representation (318) for the second frame (308), the at least one inactive spatial parameter including a direction parameter, a direction of arrival parameter, a diffusion parameter, or an energy ratio parameter; The equipment also includes: An audio signal encoder (330) is used to generate an encoded audio signal (344), which provides an encoded version (348) of the parameters of an encoded audio signal (346) for a first frame as an active frame (306) and a second frame as an inactive frame (308). and The encoded signal former (370) is used to form an encoded audio scene (304) by combining a first sound field parameter representation (316) for the first frame (306), a second sound field parameter representation (318) for the second frame (308), an encoded audio signal (346) for the first frame (306), and an encoded version (348) of the parameter description for the second frame (308).
2. The device of claim 1, wherein the sound field parameter generator (310) is configured to determine a plurality of individual sound sources from a second frame (308) of an audio signal and determine a parameter description (328) for the second frame for each sound source, each frequency range representing an individual sound source among the plurality of individual sound sources.
3. The device of claim 1, wherein the sound field parameter generator (310) is configured to generate a second sound field parameter representation (318) such that the second sound field parameter representation (318) includes parameters indicating the characteristics of the audio signal (302) relative to the listener's position.
4. The device of claim 1, wherein the first sound field parameter representation (316) includes one or more orientation parameters indicating the direction of the sound in the first frame (306) relative to the listener's position, or one or more diffusion parameters indicating the portion of the diffused sound relative to the direct sound in the first frame (306), or one or more energy ratio parameters indicating the energy ratio of the direct sound to the diffused sound in the first frame (306), or inter-channel / surround coherence parameters in the first frame (306).
5. The device of claim 1, wherein the audio signals for the first frame (306) and the second frame (308) comprise an input format having multiple components representing a sound field relative to the listener. The sound field parameter generator (310) is configured to use downmixing of multiple components to calculate one or more transport channels (324, 326, 328) for the first frame (306) and the second frame (308), and analyze the input format to determine a first parameter representation associated with one or more transport channels, or The sound field parameter generator (310) is configured to use downmixing of multiple components to calculate one or more transmission channels (324, 326, 328), and The activity detector (320) is configured to analyze one or more transmission channels (328) derived from the audio signal in the second frame (308).
6. The device as described in claim 1, The audio signal used for the first frame (306) or the second frame (308) includes an input format, which for each frame in the first and second frames has one or more transport channels and metadata associated with each frame. The sound field parameter generator (310) is configured to read metadata from the first frame (306) and the second frame (308), and use or process the metadata for the first frame as a first sound field parameter representation (316) and process the metadata of the second frame (308) to obtain a second sound field parameter representation (318), wherein the processing to obtain the second sound field parameter representation (318) reduces the amount of information units required to transmit the metadata for the second frame (308) relative to the amount required before processing.
7. The device as described in claim 6, The sound field parameter generator (310) is configured to process the metadata for the second frame (308) to reduce the number of information items in the metadata or to resample the information items in the metadata to a lower resolution, which is a time resolution or a frequency resolution, or to requantize the information units of the metadata for the second frame (308) into a coarser representation relative to the case before requantization.
8. The device as described in claim 1, The audio signal encoder (330) is configured to determine the silence information description used for inactive frames as a parameter description. The silent information description includes Amplitude-related information and shaping information for the second frame (308), wherein the amplitude-related information for the second frame (308) is energy, power, or loudness, and the shaping information is spectral shaping information; or Amplitude-related information for the second frame (308) and linear predictive coding (LPC) parameters for the second frame (308), wherein the amplitude-related information for the second frame (308) is energy, power, or loudness, or The scale parameter used for the second frame (308) has a varying associated frequency resolution, such that different scale parameters refer to frequency bands with different widths.
9. The device as described in claim 1, The audio signal encoder (330) is configured to encode an audio signal for the first frame (306) using a time-domain or frequency-domain coding mode. The encoded audio signal includes encoded time-domain samples, encoded frequency-domain samples, encoded LPC-domain samples, and side information obtained from the components of the audio signal or from one or more transport channels, which are derived from the components of the audio signal through a downmixing operation.
10. The device as claimed in claim 1, The audio signal (302) includes an input format, which may be a first-order stereo reverberation format, a higher-order stereo reverberation format, a multichannel format associated with a given speaker setup of 5.1, 7.1, or 7.1+4, or one or more audio channels representing one or more different audio objects located in a space indicated by information included in the associated metadata, or an input format that is a metadata-associated spatial audio representation. The sound field parameter generator (310) is configured to determine a first sound field parameter representation (316) and a second sound field representation, such that the parameters represent the sound field relative to a defined listener position.
11. The device as claimed in claim 1, The audio signal (302) includes an input format, which may be a first-order stereo reverberation format, a higher-order stereo reverberation format, a multichannel format associated with a given speaker setup of 5.1, 7.1, or 7.1+4, or one or more audio channels representing one or more different audio objects located in a space indicated by information included in the associated metadata, or an input format that is a metadata-associated spatial audio representation. The audio signal includes microphone signals acquired from real or virtual microphones, or synthesized microphone signals in first-order or higher-order stereo reverberation formats.
12. The device as claimed in claim 1, The activity detector (320) is configured to detect inactive phases on one or more frames after the second frame (308), and The activity detector (320) is configured to determine an inactive phase comprising eight frames following the second frame (308), and the audio signal encoder (330) is configured to generate a parameter description for the inactive frame only at every eighth frame, and the sound field parameter generator (310) is configured to generate a sound field parameter representation for every eighth inactive frame.
13. The device as claimed in claim 1, The activity detector (320) is configured to detect inactive phases on one or more frames after the second frame (308), and The sound field parameter generator (310) is configured to generate a sound field parameter representation for each inactive frame, even when the audio signal encoder (330) has not generated a parameter description for the inactive frame.
14. The device as claimed in claim 1, The sound field parameter generator (310) is configured to determine a second sound field parameter representation (318) for the second frame (308) using spatial parameters for one or more directions in the frequency band and associated energy ratios in the frequency band corresponding to the ratio of a directional component to the total energy.
15. The device as claimed in claim 1, The sound field parameter generator (310) is configured to determine a second sound field parameter representation (318) for the second frame (308) to determine a diffusion parameter indicating the ratio of diffused sound to direct sound.
16. The device as claimed in claim 1, The sound field parameter generator (310) is configured to determine a second sound field parameter representation (318) for the second frame (308) to determine directional information using a coarser quantization scheme compared to the quantization in the first frame (306).
17. The device as claimed in claim 1, The sound field parameter generator (310) is configured to use the average of the direction over time or frequency to obtain a coarser time or frequency resolution to determine a second sound field parameter representation (318) for the second frame (308).
18. The device as claimed in claim 1, The sound field parameter generator (310) is configured to determine a second sound field parameter representation (318) for a second frame (308) to determine a sound field parameter representation for one or more inactive frames, the sound field parameter representation for one or more inactive frames having the same frequency resolution as the first sound field parameter representation (316) for an active frame, and having a lower temporal occurrence rate for directional information in the sound field parameter representation for inactive frames compared to the temporal occurrence rate for active frames.
19. The device as claimed in claim 1, The sound field parameter generator (310) is configured to determine a second sound field parameter representation (318) for the second frame (308) to determine a second sound field parameter representation (318) with a diffusion parameter, wherein the diffusion parameter is transmitted with the same time or frequency resolution as the active frame but with coarser quantization.
20. The device as claimed in claim 1, The sound field parameter generator (310) is configured to determine a second sound field parameter representation (318) for the second frame (308) to quantize the diffusion parameter for the second sound field representation with a first number of bits, and wherein only a second number of bits for each quantization index is transmitted, the second number of bits being less than the first number of bits.
21. The device as claimed in claim 1, The sound field parameter generator (310) is configured to determine a second sound field parameter representation (318) for the second frame (308), thereby determining inter-channel coherence for the second sound field parameter representation (318) if the audio signal has an input channel corresponding to a channel located in the spatial domain, or determining inter-channel sound level difference for the second sound field parameter representation (318) if the audio signal has an input channel corresponding to a channel located in the spatial domain.
22. The device as claimed in claim 1, The sound field parameter generator (310) is configured to determine a second sound field parameter representation (318) for the second frame (308) to determine surround coherence, which is defined as the ratio of coherent diffuse energy in the sound field represented by the audio signal.
23. An apparatus (200) for processing an encoded audio scene (304), the encoded audio scene comprising a first sound field parameter representation (316) and an encoded audio signal (346) in a first frame and an inactive frame in a second frame, the second frame being decomposed into one or more frequency intervals, and at least one inactive spatial parameter being determined for each frequency interval as a second sound field parameter representation (318) for the second frame (308), the at least one inactive spatial parameter comprising a direction parameter, a direction of arrival parameter, a diffusion parameter, or a power ratio parameter. The equipment includes: An activity detector (2200) is used to detect that the second frame is an inactive frame; A synthesized signal synthesizer (210) is used to synthesize a synthesized audio signal (228) for the second frame (308) using a parameter description for the second frame (308). An audio decoder (230) is used to decode the encoded audio signal (346) for the first frame (306); and A spatial renderer (220) is used to spatially render the audio signal for the first frame (306) using a first sound field parameter representation (316) and a synthesized audio signal (228) for the second frame (308) and a second sound field parameter representation (318). The synthesizer (210) is configured to generate one or more transmission channels (228) for the second frame (308) as a synthesized audio signal (228), and The spatial renderer (220) is configured to render one or more transmission channels (228) in space for the second frame (308).
24. The device of claim 23, wherein for a second frame (308) of an audio signal, one or more individual sound sources are determined, and for each sound source, a parameter description for the second frame is determined, each frequency range representing an individual sound source.
25. The device of claim 23, wherein the encoded audio scene (304) includes a second sound field parameter description (318) for a second frame (308), and wherein the device includes a parameter processor (275, 1075) for deriving one or more sound field parameters (219, 318) from the second sound field parameter description (318), and wherein a spatial renderer (220) is configured to use one or more sound field parameters for the second frame (308) to render a synthesized audio signal (228) for the second frame (308).
26. The device as claimed in claim 23, The parameter processor (275, 1075) is configured to store one or more sound field parameter representations for multiple frames that occur temporally before or after the second frame (308), to extrapolate or interpolate using at least two of the one or more sound field parameter representations for the multiple frames to determine one or more sound field parameters for the second frame (308), and The spatial renderer is configured to use one or more sound field parameters for the second frame (308) to render the synthesized audio signal (228) for the second frame (308).
27. The device as claimed in claim 26, The parameter processor (275) is configured to perform jitter in directions included in at least two sound field parameter representations that occur before or after the second frame (308) when extrapolating or interpolating to determine one or more sound field parameters for the second frame (308).
28. The device as claimed in claim 23, The synthesized signal synthesizer (210) is configured to generate multiple synthesized component audio signals as synthesized audio signals (228) for the second frame (308) for individual components related to the audio output format of the spatial renderer.
29. The apparatus of claim 28, wherein the synthesized signal synthesizer (210) is configured to generate individual synthesized component audio signals for each of at least a subset of at least two individual components (228a, 228b) associated with the audio output format. The first other synthesized component audio signal (228a) is decorrelated with the second other synthesized component audio signal (228b), and The spatial renderer (220) is configured to render components of the audio output format using a combination of a first additional synthesized component audio signal (228a) and a second additional synthesized component audio signal (228b).
30. The device as claimed in claim 29, The spatial renderer (220) is configured to apply the covariance method.
31. The device as claimed in claim 30, The spatial renderer (220) is configured not to use any decorrector or control decorrector, such that when generating components of the audio output format, only multiple decorrector signals generated by the decorrector indicated by the covariance method are used (228a).
32. The device of claim 23, wherein the synthesized signal synthesizer (210, 710) is a comfort noise generator.
33. The device of claim 29, wherein the synthesized signal synthesizer (210) includes a noise generator, and a first additional synthesized component audio signal is generated by a first sample of the noise generator and a second additional synthesized component audio signal is generated by a second sample of the noise generator, wherein the second sample is different from the first sample.
34. The apparatus of claim 33, wherein the noise generator comprises a noise table, and wherein a first additional synthesized component audio signal is generated by taking a first portion of the noise table, and wherein a second additional synthesized component audio signal is generated by taking a second portion of the noise table, wherein the second portion of the noise table is different from the first portion of the noise table.
35. The apparatus of claim 33, wherein the noise generator includes a pseudo-noise generator, and wherein the first additional synthesized component audio signal is generated using a first seed for the pseudo-noise generator, and wherein the second additional synthesized component audio signal is generated using a second seed for the pseudo-noise generator.
36. The device as claimed in claim 23, The encoded audio scene (304) includes two or more transmission channels (326) for the first frame (306), and The synthesized signal synthesizer (210, 710) includes a noise generator (810) and is configured to generate a first transmission channel by sampling the noise generator (810) and a second transmission channel by sampling the noise generator (810) using the parameter description for the second frame (308), wherein the first and second transmission channels determined by sampling the noise generator (810) are weighted using the same parameter description for the second frame (308).
37. The device of claim 23, wherein the spatial renderer (220) is configured to The mixture of a direct signal and a diffused signal generated from the direct signal under the control of a decorrelation unit (730) in a first sound field parameter representation (316) operates in a first mode for the first frame (306). The second mode for the second frame (308) is operated by mixing the first and second composite component signals, wherein the first and second composite component signals are generated by the composite signal synthesizer (210) through different implementations of noise processing or pseudo-noise processing.
38. The device of claim 37, wherein the spatial renderer (220) is configured to control the mixing (740) in the second mode by means of a diffusion parameter, an energy distribution parameter or a coherence parameter derived by the parameter processor for the second frame (308).
39. The device as claimed in claim 23, The synthesized signal synthesizer (210) is configured to generate a synthesized audio signal (228) for the first frame (306) using a parameter description for the second frame (308), and The spatial renderer is configured to perform a weighted combination of the audio signal for the first frame (306) and the synthesized audio signal (228) for the first frame (306) before or after spatial rendering, wherein in the weighted combination, the intensity of the synthesized audio signal (228) for the first frame (306) is reduced relative to the intensity of the synthesized audio signal (228) for the second frame (308).
40. The device as claimed in claim 23, The parameter processor (275, 1075) is configured to determine the surround coherence for a second inactive frame (308), the surround coherence being defined as the ratio of coherent diffuse energy in the sound field represented by the second frame (308), wherein the spatial renderer is configured to redistribute the energy between the direct signal and the diffuse signal in the second frame (308) based on the sound coherence, wherein the energy of the sound surround coherence component is removed from the diffuse energy to be redistributed to the directional component, and wherein the directional component is translated in the reproduction space.
41. The device of claim 23, further comprising an output interface for converting an audio output format generated by a spatial renderer into a transcoded output format, said transcoded output format being an output format comprising multiple output channels dedicated to a speaker to be placed at a predetermined location, or a transcoded output format comprising FOA or HOA data.
42. The apparatus of claim 23, further comprising: The parameter processor (275, 1075) is configured to derive one or more second sound field parameters (219, 318) for the second frame (308), wherein the parameter processor (275, 1075) is configured to store a first sound field parameter representation for the first frame (306) and synthesize one or more second sound field parameters for the second frame (308) using the stored first sound field parameter representation (316) for the first frame (306), wherein the second frame (308) is temporally after the first frame (306).
43. A method for processing an encoded audio scene, wherein the encoded audio scene includes a first sound field parameter representation (316) and an encoded audio signal in a first frame (306), and includes an inactive frame in a second frame (308), the encoded audio scene (304) including one or more transport channels (326) for the first frame (306), the second frame being decomposed into one or more frequency intervals, and at least one inactive spatial parameter being determined for each frequency interval as a second sound field parameter representation (318) for the second frame (308), the method comprising: The second frame (308) was detected as an inactive frame; The synthesized audio signal (228) for the second frame (308) is synthesized using the parameter description for the second frame (308). Decode the encoded audio signal used for the first frame (306); and The first sound field parameters are used (316) and the synthesized audio signal (228) for the second frame (308) is used to spatially render the audio signal for the first frame (306); The method also includes: generating one or more transport channels (228) for the second frame (308) as a synthesized audio signal (228) and spatially rendering one or more transport channels (228) for the second frame (308). The method further includes: deriving one or more second sound field parameters (219, 318) for the second frame (308), wherein the parameter processor (275, 1075) is configured to store a first sound field parameter representation for the first frame (306) and synthesize one or more second sound field parameters for the second frame (308) using the stored first sound field parameter representation (316) for the first frame (306), wherein the second frame (308) is temporally after the first frame (306).
44. The method of claim 43, further comprising providing a parameter description for the second frame (308).
45. A computer program product comprising a computer program for performing the method of claim 43 when executed on a computer or processor.
Citation Information
Patent Citations
Method, apparatus and system for processing multi-channel audio signal
EP3511934A1
Speech encoding / Decoding system, speech encoding device, and speech decoding device
JP1999003099A
Systems, methods, and apparatus for voice activity detection
US20120130713A1