Acoustic Scene Encoder, Acoustic Scene Decoder, and Method Thereof Using Hybrid Encoder / Decoder Spatial Analysis
The hybrid encoding/decoding scheme for spatial acoustic encoding improves acoustic quality and flexibility by estimating and encoding spatial parameters at the encoder for some parts and the decoder for others, addressing inaccuracies in 3D scene reconstruction.
Patent Information
- Application Number
- JP2023063771
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-07-26
- Filing Date
- 2023-04-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-01-31
AI Technical Summary
Existing spatial acoustic encoding methods at low bitrates suffer from interference in spatial cues due to non-waveform-preserving parametric encoding, leading to inaccuracies in reconstructing 3D acoustic scenes.
A hybrid encoding/decoding scheme where spatial parameters are estimated and encoded by the encoder for some parts of the time-frequency representation, while others are estimated by the decoder, allowing for higher flexibility and accuracy in reconstructing 3D acoustic scenes.
This approach improves acoustic quality and flexibility by maintaining high-quality parametric information at the encoder and enabling high-resolution spatial parameter estimation at the decoder, reducing bitrate requirements.
Smart Images

Figure 0007711124000001 
Figure 0007711124000002 
Figure 0007711124000003
Abstract
Description
Technical Field
[0001] The present invention relates to the encoding or decoding of audio, and more particularly to hybrid encoder / decoder parametric spatial audio encoding.
Background Art
[0002] To transmit a three-dimensional acoustic scene, it is usually necessary to process a plurality of channels that transmit a large amount of data. Further, 3D sound is: conventional channel-based sound where each transmission channel is associated with the position of a speaker; sound carried through acoustic objects arranged three-dimensionally independently of the position of the speaker; scene-based (or ambisonics) where the acoustic scene is represented by a set of coefficient signals that are linear weights of spherical harmonic basis functions in space; etc. are expressed in various ways. In contrast to channel-based representations, scene-based representations are independent of a specific speaker setup and require an additional rendering process at the decoder, but can be reproduced with any speaker setup.
[0003] For each of these formats, dedicated encoding methods have been developed to efficiently store or transmit acoustic signals at low bitrates. For example, MPEG Surround is a parametric encoding method for channel-based surround sound, and MPEG Spatial Audio Object Coding (SAOC) is a parametric encoding method dedicated to object-based audio. The recent standard MPEG-H phase 2 also provides parametric encoding technology for higher-order ambisonics.
[0004] In this transmission scenario, the spatial parameters for all signals are always estimated, encoded, based on the signals to be encoded and transmitted, i.e., the signals available in the encoder for all possible 3D sound scenes, and then decoded and used in the decoder for the reconstruction of the acoustic scene. Due to the speed constraints for transmission, the time-frequency resolution of the parameters usually transmitted is limited and made lower than the time-frequency resolution of the acoustic data transmitted.
[0005] As another possibility for creating a tertiary acoustic scene, a low-dimensional representation, such as a 2-channel stereo or a first-order ambisonics representation, may be upmixed to the desired dimension using cues and parameters directly inferred from the low-dimensional representation. In this case, the time-frequency resolution can be finely selected as needed. On the other hand, the low-dimensional and perhaps encoded representation of the acoustic scene used leads to a quasi-optimal estimation of the spatial cues and parameters. In particular, when the analyzed acoustic scene is encoded and transmitted using parametric acoustic encoding tools and semi-parametric acoustic encoding tools, the spatial cues of the original signal are subject to more interference than would occur with just the low-dimensional representation.
[0006] Low-rate acoustic encoding using parametric encoding tools has shown progress in recent years. Such progress in encoding acoustic signals at very low bitrates has led to the widespread use of so-called parametric encoding tools, guaranteeing high quality. Using waveform-preserving encoding, i.e., encoding where only quantization noise is added to the encoded acoustic signal, for example, quantization noise time-frequency transform-based encoding and shaping using perceptual models such as MPEG-2 AAC and MPEG-1 MP3, causes audible quantization noise, especially for low bitrates.
[0007] In the parametric encoding tools developed to overcome this problem, a part of the signal is not directly encoded, but the decoder uses the parametric description of the desired acoustic signal for playback. The parametric description requires a lower transmission rate than waveform-preserving encoding. These methods do not attempt to retain the waveform of the signal, but rather generate an acoustic signal that is perceptually equivalent to the original signal. Examples of such parametric encoding tools include bandwidth expansion such as Spectral Band Replication (SBR). In SBR, the high-band portion of the spectral representation of the decoded signal is generated by copying the waveform-encoded low-band spectral signal portion and adapting it according to the above parameters. Another method is Intelligent Gap Filling (IGF). In IGF, some bands of the spectral representation are directly encoded, while the bands quantized to zero at the encoder are replaced with other bands of the spectrum that have already been decoded and reselected and adjusted according to the transmitted parameters. The third parametric encoding tool used is noise filling. In noise filling, a part of the signal or spectrum is quantized to zero, filled with random noise, and adjusted according to the transmitted parameters. Recent acoustic coding standards used for coding at medium to low bitrates use such parametric tools in combination to improve the perceptual quality at these bitrates. Examples of such standards include xHE-AAC, MPEG4-H, and EVS.
[0008] DirAC spatial parameter estimation and blind upmix are further procedures. DirAC is perceptually motivated spatial sound reproduction. Here, as an assumption, at a certain point in a certain critical band, it is assumed that the spatial resolution of the auditory system is limited to the decoding of one cue about direction and another cue about the coherence or diffusibility between hearings.
[0009] Based on these assumptions, in DirAC, the spatial sound in one frequency band is represented by cross-fading the spatial sound in one frequency band into two streams: an omnidirectional diffusion stream and a directional non-diffusion stream. The DirAC process is executed in two phases of analysis and synthesis shown in FIGS. 5a and 5b.
[0010] In the DirAC analysis stage shown in FIG. 5a, the B-format primary coincident microphones are regarded as inputs, and the diffusion and arrival direction of the sound are analyzed in the frequency domain. In the DirAC synthesis stage shown in FIG. 5b, the sound is divided into two streams: a non-diffusion stream and a diffusion stream. The non-diffusion stream is reproduced as a point source using amplitude panning, which is performed using vector base amplitude panning (VBAP) (Patent Document 2). The diffusion stream provides a sense of envelopment and is generated by transmitting uncorrelated signals to the speakers.
[0011] In the analysis stage of FIG. 5a, a band filter 1000, an energy estimator 1001, an intensity estimator 1002, time averaging units 999a and 999b, a diffuseness calculator 1003, and a direction calculator 1004 are provided. The calculated spatial parameters are diffuseness values (diffuseness) between 0 and 1 for each time / frequency tile. In FIG. 5a, the direction parameters include azimuth and elevation angles. These azimuth and elevation angles indicate the arrival direction of the sound from a reference point or listening position, particularly the position where the microphone is placed. From the microphone, a four-component signal of the input to the band filter 1000 is collected. These component signals (component signals) are primary ambisonics components including an omnidirectional component W, a directional component X, another directional component Y, and a further directional component Z, as shown in FIG. 5a.
[0012] The DirAC synthesis stage shown in Fig. 5b includes a band filter 1005 that generates time-frequency representations of the B-format microphone signals W, X, Y, Z. The signals corresponding to the individual time / frequency tiles are input to a virtual microphone stage 1006 that generates virtual microphone signals for each channel. In particular, for example, to generate a virtual microphone signal for the center channel, the virtual microphone is directed in the direction of the center channel, and the resulting signal becomes the component signal corresponding to the center channel. This signal is processed via a direct signal branch 1015 and a diffuse signal branch 1014. Both branches have corresponding gain adjusters or amplifiers, which are controlled by diffusion values derived from the original diffuseness parameters within blocks 1007, 1008, and are further processed in blocks 1009, 1010 to obtain a predetermined microphone correction.
[0013] The component signals within the direct signal branch 1015 are also gain-adjusted using gain parameters derived from direction parameters consisting of azimuth and elevation angles. In particular, these angles are input to a VBAP (Vector Base Amplitude Panning) gain table 1011. The result is input to a speaker gain averaging stage 1012 for each channel, and further via a normalization circuit 1013, and the resulting gain parameters are sent to the amplifier or gain adjuster within the direct signal branch 1015. The diffuse signal generated at the output of the decorrelator 1016 and the direct signal, i.e., the non-diffuse stream, are combined by a combiner 1017, and then other subbands are added by another combiner 1018. The combiner 1018 is, for example, a synthesis filter bank. Thus, a loudspeaker signal for one loudspeaker is generated, and the same procedure is executed for the other channels for the other loudspeakers 1019 in that loudspeaker setup.
[0014] The high-quality version of DirAC synthesis is shown in Fig. 5b. Here, the synthesizer receives all B-format signals and calculates each microphone signal for each speaker direction. The directivity pattern utilized is typically a dipole. Next, the virtual microphone signals are non-linearly modified according to the metadata, as described with respect to branches 1016 and 1015. The low-bitrate version of DirAC is not shown in Fig. 5b. However, in this low-bitrate version, only a single channel of audio is transmitted. The difference in processing is that all virtual microphone signals are replaced by the single received audio channel. The virtual microphone signals are split into two streams, a diffuse stream and a non-diffuse stream, and processed separately. The non-diffuse sound is reproduced as a point source using vector-based amplitude panning (VBAP). In panning, a monophonic sound signal is multiplied by a gain factor specific to the loudspeaker and then applied to a subset of speakers. The gain factor is calculated using information on the speaker setup and the specified pan direction. In the low-bitrate version, the input signal is simply panned in the direction indicated by the metadata. In the high-quality version, each virtual microphone signal is multiplied by the corresponding gain factor. This results in the same effect as panning, while at the same time making it less likely for non-linear artifacts to occur.
[0015] The purpose of synthesizing diffuse sound is to create the perception of sound surrounding the listener. In the low-bitrate version, the diffuse stream is reproduced by decorrelating the input signal and playing it back from all speakers. In the high-quality version, the virtual microphone signals of the diffuse stream need only be made somewhat less coherent and slightly decorrelated.
[0016] The DirAC parameters, also called spatial metadata, are composed of a tuple of diffuseness and direction. In spherical coordinates, they are represented by two angles, the azimuth angle and the elevation angle. When both the analysis and synthesis stages are executed on the decoder side, the time-frequency resolution of the DirAC parameters is chosen to be the same as the unique set of parameters for all time slots and frequency bins of the filterbank used for DirAC analysis and synthesis, i.e., the filterbank representation of the acoustic signal.
[0017] The problem when performing analysis in a spatial acoustic coding system only on the decoder side is, as described above, that parametric tools from medium to low bitrates are used. Due to the non-waveform-preserving characteristics of these tools, in the spatial analysis of the spectral part where parametric coding is mainly used, it is possible to derive values that are very different from the spatial parameters that the analysis of the original signal should generate. Figures 2a and 2b show such a mismatch scenario. Here, the DirAC analysis is performed on an uncoded signal (a) and a low-bitrate B-format transmission signal (b) using a coder that uses partial waveform preservation and parametric coding. In particular, a large difference is seen with respect to diffuseness.
[0018] Recently, a spatial acoustic encoding method that uses DirAC analysis in an encoder and transmits encoded spatial parameters to a decoder has been disclosed in Non-Patent Documents 1 and 2. Figure 3 shows an overview of a system of an encoder and a decoder that combines DirAC spatial sound processing with an acoustic coder. An input signal such as an object-encoded signal composed of a multi-channel input signal, a first-order ambisonics (FOA), or a higher-order ambisonics (HOA) signal or a downmix of an object, and one or more transport signals corresponding to object metadata such as energy metadata and / or correlation data is input to a format conversion / merger 900. The format conversion / merger 900 is configured to convert each of the input signals into a corresponding B-format signal, and further combines the streams received in different representations by adding the corresponding B-format components together, or by other combination techniques including weighted addition or selection of different information of different input data.
[0019] The resulting B-format signal is introduced into a DirAC analyzer 210 to derive DirAC metadata such as arrival direction metadata and diffuseness metadata, and the obtained signal is encoded using a spatial metadata encoder 220. Further, the B-format signal is sent to a beamformer / signal selector to downmix the B-format signal to a transport channel or several transport channels, and then encoded using an EVS-based core encoder 140.
[0020] The outputs of one block 220 and the other block 140 represent an encoded acoustic scene. The encoded acoustic scene is sent to a decoder, where a spatial metadata decoder 700 receives the encoded spatial metadata and an EVS-based core decoder 500 receives the encoded transport channel. The decoded spatial metadata obtained by block 700 is sent to a DirAC synthesis stage 800, and the decoded one or more transport channels at the output of block 500 are subjected to frequency analysis at block 860. The resulting time / frequency decomposition is also sent to the DirAC synthesizer 800, where a loudspeaker signal or a first-order ambisonics or higher-order ambisonics component or any other representation of the acoustic scene is generated as the decoded acoustic scene.
[0021] In the procedures disclosed in Patent Documents 1 and 2, DirAC metadata, i.e., spatial parameters, are estimated, encoded at a low bit rate, and transmitted to a decoder. At the decoder, the spatial parameters are used to reconstruct a 3D acoustic scene together with a low-dimensional representation of the acoustic signal.
[0022] In the present invention, DirAC metadata, i.e., spatial parameters, are estimated and encoded at a low bit rate and transmitted to a decoder, where they are used to reconstruct a 3D acoustic scene together with a low-dimensional representation of the acoustic signal.
[0023] To achieve a low bitrate for metadata, the time-frequency resolution is made smaller than the time-frequency resolution of the filter bank used in the analysis and synthesis of 3D acoustic scenes. Figures 4a and 4b show a comparison between the uncoded and ungrouped spatial parameters of the DirAC analysis (a) and the coded and grouped parameters of the same signal using the DirAC metadata that has been coded and transmitted by the DirAC spatial acoustic coding system disclosed in Patent Document 1. Comparing Figure 2a and Figure 2b, it can be seen that the parameters used in the decoder (b) are close to the parameters estimated from the original signal, but the time-frequency resolution is lower than the estimation by the decoder alone. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0024] An object of the present invention is to provide an improved concept for processing such as encoding or decoding of an acoustic scene. MEANS FOR SOLVING THE PROBLEMS
[0025] The present invention is based on the discovery that improved acoustic quality, higher flexibility, and generally improved performance can be obtained by applying a hybrid encoding / decoding scheme. Here, the spatial parameters used in the decoder to generate the decoded two-dimensional or three-dimensional acoustic scene are estimated in the decoder based on typically low-order acoustic representations that have been encoded and transmitted for some parts of the time-frequency representation of the scene, and for other parts, they are estimated, quantized, and encoded in the encoder and transmitted to the decoder.
[0026] Depending on the implementation, the separation between the estimation region on the encoder side and the estimation region on the decoder side may vary depending on the various spatial parameters used in the generation of the three-dimensional or two-dimensional acoustic scene in the decoder.
[0027] In an embodiment, the different parts or preferably the partitioning into the time-frequency domain can be arbitrary. However, in a preferred embodiment, the decoder estimates parameters for the part of the spectrum that is mainly encoded in a way that maintains the waveform, while for the part of the spectrum where parametric encoding tools are mainly used, it is advantageous to encode and transmit the parameters calculated by the encoder.
[0028] Embodiments of the present invention propose a low bitrate encoding solution for transmitting a 3D acoustic scene by using a hybrid encoding system in which spatial parameters used for reconstructing a 3D acoustic scene estimated and encoded by an encoder are partly estimated and encoded by the encoder and transmitted to a decoder, and the remaining part is directly estimated by the decoder.
[0029] The present invention discloses 3D acoustic reproduction based on a hybrid approach for a decoder that only estimates parameters for a part of a signal. Here, the spatial representation is brought into a low dimension within an acoustic encoder, the low-dimensional representation is encoded, estimated within the encoder, encoded at a low encoder level, and the spatial cue and parameters are transmitted from the encoder to the decoder as part of the spectrum, and yet the spatial cue is well maintained. Here, the low dimensionality associated with the encoding of the low-dimensional representation is thought to lead to a quasi-optimal estimation of the spatial parameters.
[0030] In one embodiment, an acoustic scene encoder is configured to encode an acoustic scene. The acoustic scene includes at least two component signals. The acoustic scene encoder includes a core encoder configured to core-encode at least two component signals, the core encoder generating a first encoded representation for a first portion of at least two component signals and a second encoded representation for a second portion of at least two component signals. A spatial analyzer analyzes the acoustic scene to derive one or more spatial parameters or one or more sets of spatial parameters for the second portion, and an output interface then forms an encoded acoustic scene signal including the first encoded representation, the second encoded representation, and one or more spatial parameters or one or more sets of spatial parameters for the second portion. Typically, any spatial parameters for the first portion are not included in the encoded acoustic signal. This is because these spatial parameters are estimated by the decoder from the decoded first representation within the decoder. On the other hand, the spatial parameters for the second portion have already been calculated within the acoustic scene encoder based on the original acoustic scene or an acoustic scene that has already been processed and whose dimension and thus bitrate have been reduced.
[0031] Therefore, the parameters calculated by the encoder can carry high-quality parametric information. The reason is that these parameters are calculated by the encoder from very accurate data that are not affected by the distortion of the core encoder and can even be available in a very high-dimensional manner like the signals obtained from a high-quality microphone array. Due to the fact that such very high-quality parametric information is preserved, it becomes possible to core-encode the second part with lower accuracy or usually lower resolution. Thus, by core-encoding the second part quite coarsely, bits can be saved and thus can be given to the representation of the encoded space metadata. The bits saved by the very coarse encoding of the second part can also be used for the high-resolution encoding of the first part of at least two component signals. The high-resolution or high-quality encoding of at least two component signals is useful. The reason is that on the decoder side, the parametric space data do not exist in the first part and are derived within the decoder by spatial analysis. Therefore, instead of calculating all the spatial metadata by the encoder, by core-encoding at least two component signals, it is possible to ensure all the bits that would otherwise be required for the encoding metadata and to core-encode at least two component signals in the first part with high quality.
[0032] Accordingly, according to the present invention, the separation of the first part and the second part of the acoustic scene can be performed in a very flexible manner, for example, according to bit rate requirements, acoustic quality requirements, processing requirements, i.e., whether more processing resources are available in the encoder or decoder, etc. In a preferred embodiment, the separation into the first part and the second part is performed based on the functionality of the core encoder. In particular, in the case of a high-quality and low-bitrate core encoder that applies parametric encoding operations to specific bands, such as spectral band replication processing, intelligent gap filling processing, noise filling processing, etc., the separation regarding the spatial parameters is performed such that the non-parametrically encoded part of the signal forms the first part and the parametrically encoded part of the signal forms the second part. Accordingly, a more accurate representation of the spatial parameters can be obtained for the parametrically encoded second part, which is usually the low-resolution encoded part of the audio signal, while high-quality parameters can be obtained for better encoding, i.e., the high-resolution encoded first part. The reason is that very high-quality parameters can be estimated at the decoder side using the decoded representation of the first part.
[0033] In a further embodiment, in order to further reduce the bit rate, the spatial parameters of the second part are calculated in the encoder with a certain time-frequency resolution. This time-frequency resolution can be high or low. In the case of a high time-frequency resolution, the calculated parameters are grouped in a specific way to obtain low time-frequency resolution spatial parameters. These low time-frequency resolution spatial parameters are nevertheless high-quality spatial parameters with only low resolution. However, the low resolution has the advantage that bits are saved for transmission because the number of spatial parameters in its time duration and frequency band decreases. However, since the spatial data does not change very much with respect to time and frequency, usually, even if the number of spatial parameters is reduced, it does not cause much of a problem. Accordingly, a good-quality representation with a low bit rate of the spatial parameters for the second part can be obtained.
[0034] The spatial parameters for the first part are calculated on the decoder side and do not need to be transmitted anywhere, so there is no need to make a compromise regarding resolution. Therefore, a fast and high-frequency resolution estimation of the spatial parameters can be performed on the decoder side, and this high-resolution parametric data helps to provide a good spatial representation of the first part of the acoustic scene. Thus, the "drawback" of calculating the spatial parameters on the decoder side based on at least two transmitted components for the first part can be reduced or eliminated by calculating the spatial parameters with high time-frequency resolution and by using these parameters in the spatial rendering of the acoustic scene. This has no adverse effect on the transmission bitrate between the encoder / decoder as any processing performed on the decoder side does not affect the bitrate at all.
[0035] A further embodiment of the present invention depends on the situation where, for the first part, at least two components are encoded and transmitted, and parametric data estimation can be performed on the decoder side based on the at least two components. However, in one embodiment, it is preferable to encode only a single transport channel for the second representation, so that the second part of the acoustic scene can be encoded at a substantially low bitrate. This transport channel, i.e., the downmix channel, is represented at a very low bitrate compared to the first part. The reason is that in the first part, two or more components are required for encoding and sufficient data for spatial analysis on the decoder side is needed, while in the second part, only a single channel or component is encoded.
[0036] Therefore, the present invention provides additional flexibility regarding the bitrate, acoustic quality, and processing requirements available on the encoder or decoder side.
[0037] Preferred embodiments of the present invention will be described below with reference to the accompanying drawings.
Brief Description of the Drawings
[0038]
Fig. 1a
Fig. 1b
Fig. 2
Fig. 3
Fig. 4
Fig. 5a
Fig. 5b
Fig. 6a
Fig. 6b
Fig. 7a
Fig. 7b
Fig. 8a
Fig. 8b
Fig. 9a
Fig. 9b
Fig. 10a
Fig. 10b
Fig. 11
Embodiments for Carrying Out the Invention
[0039] FIG. 1a shows an acoustic scene encoder for encoding an acoustic scene 110 that includes at least two component signals. The acoustic scene encoder includes a core encoder 100 for core encoding at least two component signals. Specifically, the core encoder 100 is configured to generate a first encoded representation 310 for a first portion of at least two component signals and a second encoded representation 320 for a second portion of at least two component signals. The acoustic scene encoder includes a spatial analyzer that analyzes the acoustic scene to derive one or more spatial parameters or one or more sets of spatial parameters for the second portion. The acoustic scene encoder includes an output interface 300 for forming an encoded acoustic scene signal 340. The encoded acoustic scene signal 340 has a first encoded representation 310 representing a first portion of at least two component signals, a second encoder representation 320, and parameters 330 for the second portion. The spatial analyzer 200 is configured to apply spatial analysis to a first portion of at least two component signals using the original acoustic scene 110. Alternatively, the spatial analysis can be performed based on a reduced dimensionality representation of the acoustic scene. For example, if the acoustic scene 110 includes recordings of several microphones arranged in a microphone array, the spatial analysis 200 will of course be performed based on this data. However, the core encoder 100 is configured to reduce the dimensionality of the acoustic scene to at least two components, for example, consisting of an omnidirectional component and at least one directional component such as X, Y, or Z of a B-format representation. However, other representations such as higher order representations or A-format representations can be used as well. The first encoder representation of the first portion will then consist of at least two different components that are decodable and typically consist of the encoded acoustic signals of each component.
[0040] The second encoder representation for the second part can consist of the same number of components or can have a lower number, such as consisting of a single omnidirectional component encoded by the core encoder of the second part. In an implementation where the core encoder 100 reduces the dimension of the original acoustic scene 110, the reduced-dimension acoustic scene can optionally be transferred to the spatial analyzer via line 120 instead of the original acoustic scene.
[0041] FIG. 1b shows an acoustic scene decoder comprising an input interface 400 for receiving an encoded acoustic scene signal 340. This encoded acoustic scene signal includes a first encoded representation 410, a second encoded representation 420, and one or more spatial parameters of the second part. The encoded representation of the second part can also be a single encoded acoustic channel or can include two or more encoded acoustic channels. On the other hand, the first encoded representation of the first part includes at least two different encoded acoustic signals. The acoustic signals in the first encoded representation, or different encoded acoustic signals in the second encoded representation if available, are either encoded signals together, such as an encoded stereo signal, or more preferably, individually encoded monoaural acoustic signals.
[0042] An encoded representation including a first encoded representation 410 of a first part and a second encoded representation 420 of a second part is input to a core decoder for decoding the first and second encoded representations to obtain at least two decoded representations and obtaining a decoded representation consisting of at least two component signals representing an acoustic scene. The decoded representation includes a first decoded representation of the first part shown at 810 and a second decoded representation of the second part shown at 820. The first decoded representation is transferred to a spatial analyzer 600 to analyze a part of the decoded representation corresponding to the first part of the at least two component signals and obtain one or more spatial parameters 840 for the first part of the at least two component signals. The acoustic scene decoder also includes, in the embodiment of FIG. 1b, a spatial renderer 800 for spatially rendering a decoded representation including the first decoded representation of the first part 810 and the second decoded representation of the second part 820. The spatial renderer 800 is configured to use, for the purpose of acoustic rendering, the parameter 840 derived from the spatial analyzer for the first part and the parameter 830 derived from the parameter decoded via the parameter / metadata decoder 700 for the second part. If the representation of the parameter in the encoded signal is in an unencoded format, the parameter / metadata decoder 700 is not required, and one or more spatial parameters of the second part of the at least two component signals are sent directly from the input interface 400, after demultiplexing or a specific processing operation, to the spatial renderer 800 as data 830.
[0043] FIG. 6a shows a schematic view of different typically overlapping time frames F1 to F4. The core encoder 100 of FIG. 1a is configured to form such subsequent time frames from at least two component signals. In such a situation, the first time frame can be taken as the first part and the second time frame as the second part. Thus, according to an embodiment of the present invention, the first part can be the first time frame, the second part can be another time frame, and the switching between the first and second parts can be performed over time. Although FIG. 6a shows overlapping time frames, non-overlapping time frames can be used in the same way. FIG. 6a shows time frames having equal lengths, but the switching can also be performed using time frames having different lengths. Thus, for example, if time frame F2 is smaller than time frame F1, this will result in an increased temporal resolution of the second time frame F2 with respect to the first time frame F1. And the second time frame F2 having the increased resolution preferably corresponds to the first part that is encoded with respect to its components, while the first time part, i.e., the low-resolution data, will correspond to the second part that is encoded at low resolution, but the spatial parameters for this second part can be calculated at any resolution since the overall acoustic scene is obtained by the encoder.
[0044] FIG. 6b shows an alternative implementation where the spectra of at least two component signals are shown as having a specific number of bands B1, B2, …, B6, …. Preferably, the bands are separated into bands having different bandwidths increasing from the lowest to the highest center frequency for performing a perceptually motivated spectral band splitting. The first part of the at least two component signals can consist of, for example, the first four bands, and the second part, for example, can consist of bands B5 and B6. This is consistent with a situation where the core encoder performs spectral band replication and the crossover frequency between the non-parametrically encoded low-frequency part and the parametrically encoded high-frequency part is at the boundary between band B4 and band B5.
[0045] Separately from this, in the case of intelligent gap filling (IGF) or noise filling (NF), since the bands are arbitrarily selected according to signal analysis, the first part consists of, for example, bands B1, B2, B4, B6, and the second part consists of B3, B5, and possibly another higher frequency band. Therefore, as shown in FIG. 6b, regardless of whether the band is a typical scale factor band with a bandwidth increasing from the lowest to the highest frequency, or whether the bands are of the same size, a very flexible separation of the acoustic signal into bands can be performed. The boundary between the first part and the second part does not necessarily coincide with the scale factor band typically used in the core encoder, but it is desirable for the boundary between the first part and the second part to coincide with the boundary between the scale factor band and the adjacent scale factor band.
[0046] FIG. 7a shows a preferred implementation of the acoustic scene encoder. In particular, the acoustic scene is preferably input into a signal separator 140 that is part of the core encoder 100 of FIG. 1a. The core encoder 100 of FIG. 1a includes both parts, namely, dimension reducers 150a and 150b for the first part and the second part of the acoustic scene. At the output of the dimension reducer 150a, there are at least two component signals that are encoded by an acoustic encoder 160a for the first part. The dimension reducer 150b for the second part of the acoustic scene can include the same configuration as the dimension reducer 150a. However, alternatively, the reduced dimension obtained by the dimension reducer 150b can be a single transport channel that is then encoded by an acoustic encoder 160b to obtain a second encoded representation 320 of at least one transport / component signal.
[0047] The acoustic encoder 160a for the first encoded representation can include an encoder that maintains the waveform, is non-parametric, or has high temporal or high-frequency resolution. On the other hand, the acoustic encoder 160b is a parametric encoder such as an SBR encoder, an IGF encoder, a noise-filling encoder, or others with low temporal or frequency resolution. Thus, the acoustic encoder 160b typically results in a lower-quality output representation compared to the acoustic encoder 160a. This "drawback" is addressed by performing spatial analysis via the spatial data analyzer 210 on the original audio scene, or the dimensionality-reduced audio scene if the dimensionality-reduced audio scene still includes at least two component signals. The spatial data obtained by the spatial data analyzer 210 is transferred to a metadata encoder 220 that outputs encoded low-resolution spatial data. Blocks 210 and 220 are both preferably subsumed within the spatial analyzer block 200 of FIG. 1a.
[0048] Preferably, the spatial data analyzer performs spatial data analysis at a high resolution such as high-frequency resolution or high-temporal resolution, and groups the high-resolution spatial data and entropy encodes it by the metadata encoder to obtain encoded low-resolution spatial data in order to keep the bitrate required for the encoded metadata within a reasonable range. For example, if the spatial data analysis is performed for, say, 8 time slots per frame and 10 bands per time slot, the spatial data can be grouped into 1 spatial parameter per frame and, for example, 5 bands per parameter.
[0049] On the one hand, it is preferable to calculate direction data, and on the other hand, to calculate diffusivity data. At this time, the metadata encoder 220 is configured to output encoded data for the directionality data and the diffusivity data at different time / frequency resolutions. Usually, the directivity data requires a higher resolution than the diffusivity data. A preferred method for calculating parametric data at different resolutions is to perform a spatial analysis at a high resolution, usually the same resolution, for both parametric types, and then perform grouping in time and / or frequency using different parametric information in different ways for different parameter types. For example, it has an encoded low-resolution spatial data output 330 with a medium time and / or frequency resolution for the directionality data and a low resolution for the diffusivity data.
[0050] FIG. 7b shows the decoder-side implementation of the corresponding acoustic scene decoder.
[0051] In the embodiment of FIG. 7b, the core decoder 500 of FIG. 1b has a first acoustic decoder instance 510a and a second acoustic decoder instance 510b. Preferably, the first acoustic decoder instance 510a is a non-parametric or waveform-preserving or high-resolution (in time and / or frequency) encoder, and generates a first decoded portion of at least two component signals at the output. This data 810 is sent, on the one hand, to the spatial renderer 800 of FIG. 1b and further input to the spatial analyzer 600. Preferably, the spatial analyzer 600 is a high-resolution spatial analyzer that preferably calculates high-resolution spatial parameters for the first portion. Usually, the resolution of the spatial parameters of the first portion is higher than the resolution associated with the encoded parameters input to the parameter / metadata decoder 700. However, the entropy-decoded low-time or frequency-resolution spatial parameters output by block 700 are input to a parameter de-grouping unit for resolution improvement 710. Such de-grouping (un-grouping) of the parameters can be performed by copying the transmitted parameters to specific time-frequency tiles, and the de-grouping is performed according to the corresponding grouping performed by the encoder-side metadata encoder 220 of FIG. 7a. Naturally, along with the de-grouping, further processing or smoothing operations can be performed as needed.
[0052] At this time, the result of block 710 is a collection of preferably high-resolution parameters decoded for the second portion, which usually has the same resolution as the parameters 840 for the first portion. Also, the encoded representation of the second portion is decoded by the acoustic decoder 510b to obtain a second decoded portion 820 of a signal that usually has at least one or at least two components.
[0053] FIG. 8a shows a preferred implementation of an encoder that depends on the functions discussed with respect to FIG. 3. In particular, multi-channel input data, or first-order ambisonics or higher-order ambisonics input data, or object data, is input to a B-format converter. The B-format converter converts and combines the individual input data to produce, for example, four B-format components, typically omnidirectional acoustic signals, and three directional acoustic signals such as X, Y, and Z.
[0054] Alternatively, the signal input to the format converter or core encoder may be a signal captured by an omnidirectional microphone disposed in a first portion and another signal captured by an omnidirectional microphone disposed in a second portion different from the first portion. Further, the acoustic scene may include, as a first component signal, a signal captured by a directional microphone directed in a first direction and, as a second component, at least one signal captured by another directional microphone directed in a second direction different from the first direction. These “directional microphones” do not necessarily have to be actual microphones and may be virtual microphones.
[0055] As the acoustic input to block 900, or the output by block 900, or generally as the acoustics used as the acoustic scene, A-format component signals, B-format component signals, first-order ambisonics component signals, higher-order ambisonics component signals, or component signals captured by a microphone array having at least two microphone capsules or component signals calculated from virtual microphone processing can be used.
[0056] The output interface 300 of FIG. 1a is configured such that, for the second portion of the encoded acoustic scene signal, it does not include any spatial parameters of the same parameter type as the one or more spatial parameters generated by the spatial analyzer.
[0057] Thus, when the parameters 330 of the second part are arrival direction data and diffusivity data, the first encoded representation of the first part does not include the arrival direction data and the diffusivity data, but of course can include any other parameters, such as scale factors, LPC coefficients, etc., which are calculated by the core encoder.
[0058] Furthermore, the band separation performed by the signal separator 140 can be implemented such that when different parts are in different bands, the start band of the second part is lower than the bandwidth extension start band. Moreover, the core noise filling does not necessarily need to apply a fixed crossover band, but can be gradually used for more parts of the core spectrum as the frequency increases.
[0059] Furthermore, the parametric or largely parametric processing for the second frequency subband of the time frame includes the calculation of the amplitude-related parameters of the second frequency subband and the quantization and entropy coding of these amplitude-related parameters instead of the individual spectral lines of the second frequency subband. Such amplitude-related parameters forming the low-resolution representation of the second part are given, for example, for each scale factor band, by a spectral envelope representation having only one scale factor or energy value, while the high-resolution first part depends on individual MDCT or FFT or general individual spectral lines.
[0060] Thus, the first part of at least two component signals is given by a specific frequency band of each component signal, and the specific frequency band of each component signal is encoded using several spectral lines to obtain the encoded representation of the first part. However, for the second part, the sum of the individual spectral lines of the second part, the sum of the squared spectral lines representing the energy of the second part, or the sum of the cubed spectral lines representing the loudness measurement for the spectral part can also be used for the parametric encoded representation of the second part.
[0061] Referring again to FIG. 8a, the core encoder 160, which includes the individual core encoder branches 160a, 160b, can include beamforming / signal selection procedures for the second part. Thus, the core encoders shown as 160a, 160b in FIG. 8b output, on the one hand, the first part of the encoding of all four B-format components, the second part of the encoding of a single transport channel, and the spatial metadata for the second part generated by the DirAC analysis 210 depending on the second part, and are connected to the subsequent spatial metadata encoder 220.
[0062] On the decoder side, the encoded spatial metadata is input to the spatial metadata decoder 700, and the parameters of the second part shown as 830 are generated. The core decoders 510a, 510b, which are preferably implemented as an EVS-based core decoder composed of elements, output a decoded representation consisting of both parts, but the two parts are not yet separated. The decoded representation is input to the frequency analysis block 860, and the frequency analyzer 860 generates the component signals of the first part and transfers it to the DirAC analyzer 600 to generate the parameters 840 for the first part. The transport channel / component signals of the first and second parts are transferred from the frequency analyzer 860 to the DirAC synthesizer 800. The DirAC synthesizer operates as usual in this embodiment because it has no knowledge and actually requires no specific knowledge. This is regardless of whether the parameters for the first and second parts are generated on the encoder side or the decoder side. Instead, both the DirAC synthesizer 800 and the DirAC synthesizer can generate the "same" parameters based on the frequency representation of the decoded representation of at least two component signals representing the acoustic scene shown as 862, and the parameters for both parts, the loudspeaker output, the first-order ambisonics (FOA), the higher-order ambisonics (HOA), or the binaural output.
[0063] Figure 9a shows another preferred embodiment of the acoustic scene encoder. Here, the core encoder 100 of Figure 1a is implemented as a frequency domain encoder. In this implementation, the signal encoded by the core encoder is preferably input to an analysis filter bank 164 that applies a time-frequency transform or decomposition, typically to overlapping time frames. The core encoder comprises a waveform maintenance encoder processor 160a and a parametric encoder processor 160b. The distribution of the spectral portions to the first and second portions is controlled by a mode controller 166. The mode controller 166 can depend on signal analysis, bitrate control, or apply fixed settings. Usually, the acoustic scene encoder can be configured to operate at different bitrates, in which case a predetermined boundary frequency between the first and second portions depends on the selected bitrate, and the predetermined boundary frequency is low for low bitrates and high for high bitrates.
[0064] Alternatively, the mode controller comprises a tonal mask processing function known from intelligent gap filling that analyzes the spectrum of the input signal, determines the bands that need to be encoded with high spectral resolution, and this ultimately becomes the first encoded portion. Also, it determines the bands that can be encoded in a parametric way, and this ultimately becomes the second decoded portion. The mode controller 166 also controls the spatial analyzer 200 on the encoder side and is preferably configured to control the band separator 230 of the spatial analyzer or the parameter separator 240 of the spatial analyzer. Thereby, ultimately, only the spatial parameters of the second portion rather than the first portion are generated and output to the encoded scene signal.
[0065] In particular, when the spatial analyzer 200 receives directly either before the acoustic scene signal is input to the analysis filter bank or after being input to the filter bank, the spatial analyzer 200 analyzes the first and second parts as a whole, and subsequently the parameter separator 240 selects the parameters for the second part for output to the encoded scene signal. Alternatively, when the spatial analyzer 200 receives the input data from the band separator and the band separator 230 has already sent only the second part, the parameter separator 240 no longer requires anything. The reason is that the spatial analyzer 200 simply receives only the second part anyway and outputs the spatial data for the second part.
[0066] Therefore, the selection of the second part can be performed before or after the spatial analysis, and preferably is controlled by the mode controller 166 or can also be implemented fixedly. The spatial analyzer 200 relies on the analysis filter bank of the encoder or, although not shown in FIG. 9a, uses its own individual filter bank as shown, for example, as the implementation of the DirAC analysis stage at 1000 in FIG. 5a.
[0067] FIG. 9b shows a time domain encoder in contrast to the frequency domain encoder of FIG. 9a. A band separator 168 is provided instead of the analysis filter bank 164. This band separator 168 is controlled by the mode controller 166 of FIG. 9a (not shown in FIG. 9b) or is fixed. When controlled, the control can be performed based on the bit rate, signal analysis, or other procedures useful for this purpose. The typically M components input to the band separator 168 are processed on the one hand by the low-band time domain encoder 160a and on the other hand by the time domain bandwidth expansion parameter calculator 160b. Preferably, the low-band time domain encoder 160a outputs a first encoded representation in which the M individual components are encoded. In contrast, the second encoded representation generated by the time domain bandwidth expansion parameter calculator 160b contains only N component / transport signals, where N is smaller than M and N is 1 or more.
[0068] Depending on whether the spatial resolver 200 depends on the band separator 168 of the core encoder, a separate band separator 230 is not required. However, if the spatial resolver 200 depends on the band separator 230, the connection between block 168 and block 200 in FIG. 9b is not necessary. If neither the band separator 168 nor 230 is connected to the input of the spatial resolver 200, the spatial resolver performs full-band analysis, and the band separator 240 separates only the second part of the spatial parameters to be transferred to the output, which is sent to the output interface or becomes the encoded acoustic scene.
[0069] Thus, FIG. 9a shows a waveform-preserving encoder processor 160a or a spectral encoder for quantizing entropy coding, while the corresponding block 160a in FIG. 9b is any time-domain encoder such as an EVS encoder, an ACELP encoder, an AMR encoder, or a similar encoder. While block 160b shows a frequency-domain parametric encoder or a general parametric encoder, block 160b in FIG. 9b is basically a time-domain bandwidth expansion parameter calculator that can calculate the same parameters as block 160 or different parameters in some cases.
[0070] Figure 10a shows a frequency domain decoder. This frequency domain decoder typically corresponds to the frequency domain encoder of Figure 9a. The spectral decoder that receives the first part of the encoding has, as shown at 160a, an entropy decoder, an inverse quantizer, and any other elements known, for example, from AAC encoding or any other arbitrary spectral domain encoding. The parametric decoder 160b that receives parametric data such as per-band energy as the second encoded representation of the second part typically operates as an SBR decoder, an IGF decoder, a noise filling decoder, or other parametric decoder. The spectral values of the first part and the spectral values of the second part are input into a synthesis filter bank 169 to obtain an encoded representation. The obtained encoded representation is typically transferred to a spatial renderer for the purpose of spatial rendering.
[0071] The first part may be transferred directly to the spatial analyzer 600 or may also be derived from the decoded representation at the output of the synthesis filter bank 169 via the band separator 630. Depending on the situation, the parameter separator 640 may or may not be present. If the spatial analyzer 600 receives only the first part, the band separator 630 and the parameter separator 640 are not required. If the spatial analyzer 600 receives the decoded representation and there is no band separator, the parameter separator 640 is required. If the decoded representation is input into the band separator 630, the spatial analyzer 600 outputs only the spatial parameters of the first part, so the spatial analyzer does not need to have the parameter separator 640.
[0072] FIG. 10b shows a time-domain decoder corresponding to the time-domain encoder of FIG. 9b. In particular, the first encoded representation 410 is input to the low-band time-domain decoder 160a, and the decoded first part is input to the combiner 167. The bandwidth expansion parameter 420 is input to a time-domain bandwidth expansion processor that outputs a second part. The second part is also input to the combiner 167. Depending on the implementation, a combiner is implemented to combine spectral values if the first and second parts are spectral values, or to combine their time-domain samples if the first and second parts are already obtained as time-domain samples. The output of the combiner 167 is a decoded representation that can be processed by the spatial analyzer 600, regardless of the presence or absence of the band separator 630 or the parameter separator 640, as described above with respect to FIG. 10a.
[0073] FIG. 11 shows a preferred implementation of the spatial renderer. However, implementations are also possible for those that depend on the DirAC parameter or parameters other than the DirAC parameter, or that generate a rendering signal representation different from a direct loudspeaker representation such as an HOA representation. Usually, the data 862 input to the DirAC synthesizer 800 is composed of several components such as B-format for the first and second parts, as shown in the upper left corner of FIG. 11. Also, there may be a case where the second part is not obtained as a plurality of components but only as a single component. Such a situation is shown in the lower left of FIG. 11. In particular, for example, when the first and second parts have all components, that is, when the signal 862 in FIG. 8b includes all components of the B-format, the entire spectrum of all components is available, and processing can be performed for each individual time-frequency tile by time-frequency decomposition. This processing is performed by the virtual microphone processor 870a to calculate the loudspeaker components from the decoded representation for each loudspeaker in the loudspeaker arrangement.
[0074] Alternatively, if the second part is only available as a single component, the time - frequency tiles of the first part are input to the virtual microphone processor 870a, while the time / frequency part for the single or fewer components of the second part can be input to the processor 870b. The processor 870b, for example, only performs a copy operation. That is, it copies a single transport channel to the output signals for each loudspeaker signal. Thus, the processing of the virtual microphone processor 870a in this alternative configuration is replaced by a simple copy operation.
[0075] Next, the outputs of block 870a in the first embodiment, i.e., 870a for the first part and block 870b for the second part, are input to a gain processor 872 to modify the output component signals using one or more spatial parameters. This data is also input to a weight / decorrelator processor 874 to generate decorrelated output component signals using one or more spatial parameters. The output of block 872 and the output of block 874 are combined in a combiner 876 that operates on each component, whereby the output of block 876 provides the frequency - domain representation of each loudspeaker signal.
[0076] Next, the all - frequency - domain loudspeaker signals are converted to the time - domain representation by the synthesis filter bank 878, and the generated time - domain loudspeaker signals are digitally - to - analog - converted and can be used to drive the corresponding loudspeakers placed at the defined loudspeaker positions.
[0077] Typically, the gain processor 872 operates based on spatial parameters, and preferably direction parameters such as the direction of arrival data, and optionally diffuseness parameters. Further, the weight / decorrelator processor operates based on spatial parameters and preferably also based on diffuseness parameters.
[0078] Thus, in an implementation, gain processor 872 generates the non-diffused stream of FIG. 5b, shown at 1015, and weighting / decorrelating processor 874 generates a diffused stream, such as that shown by upper branch 1014 of FIG. 5b. However, other implementations that depend on different procedures, different parameters, and different ways of generating direct and diffused signals are equally possible.
[0079] Exemplary benefits and advantages of the preferred embodiments over the prior art are as follows. Embodiments of the present invention provide better time-frequency resolution for a selected portion of a signal chosen to have spatial parameters estimated on the decoder side than a system that uses parameters estimated and encoded on the encoder side for the entire signal. Embodiments of the present invention provide better spatial parameter values for a reconstructed signal portion by analysis, encoding of parameters at the encoder, and transmission of the parameters to the decoder than a system in which the spatial parameters are estimated at the decoder using decoded lower-order acoustic signals. Embodiments of the present invention allow for a more flexible trade-off between time-frequency resolution, transmission rate, and parameter accuracy than either a system that uses coded parameters for the entire signal or a system that uses decoder-side estimated parameters for the entire signal. Embodiments of the present invention provide better parameter accuracy by selecting encoder-side estimation and encoding of some or all of the spatial parameters for a signal portion encoded primarily using parametric encoding tools, and encoding some or all of the spatial parameters for those portions, and provide better time-frequency resolution by using waveform-preserving encoding tools for signal portions encoded primarily and delegating the estimation of spatial parameters for those signal portions to the decoder side.
Prior Art Documents
Patent Documents
[0080] [Non-Patent Document 1] V. Pulkki, M-V Laitinen, J Vilkamo, J Ahonen, T Lokki and T Pihlajamaeki, “Directional audio coding - perception-based reproduction of spatial sound”, International Workshop on the Principles and Application on Spatial Hearing, Nov. 2009, Zao; Miyagi, Japan. [Non-Patent Document 2] Ville Pulkki. “Virtual source positioning using vector base amplitude panning”. J. Audio Eng. Soc., 45(6):456{466, June 1997.
[0081] [Patent Document 1] European Patent Application No. 17202393.9, “EFFICIENT CODING SCHEMES OF DIRAC METADATA”. [Patent Document 2] European Patent Application No. 17194816.9 “Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to DirAC based spatial audio coding”
[0082] The encoded audio signal of the present invention can be stored in a digital storage medium or a non-transitory storage medium, or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0083] Although several aspects have been described as devices, it should be apparent that these aspects also represent descriptions of corresponding methods. In that case, a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent descriptions of corresponding blocks or items or features of a corresponding device.
[0084] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. This implementation can be carried out using a digital storage medium, such as a flexible disk, DVD, CD, ROM, PROM, EPROM, EEPROM or flash memory, having electronically readable control signals stored thereon and cooperating or capable of cooperating with a programmable computer system so that respective methods are executed.
[0085] Some embodiments according to the present invention include a data carrier having electronically readable control signals capable of cooperating with a programmable computer system so that the methods described herein are executed.
[0086] Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to execute one of the methods when the computer program product runs on a computer. The program code can be stored, for example, on a machine-readable carrier.
[0087] Other embodiments include a computer program for executing one of the methods described herein, stored on a machine-readable carrier or a non-transitory storage medium.
[0088] In other words, an embodiment of a method according to the present invention is a computer program having program code for executing the method described herein when the computer program runs on a computer.
[0089] A further embodiment of the method of the present invention is a data carrier (i.e., a digital storage medium or a computer-readable medium) having recorded thereon a computer program for carrying out the method described herein.
[0090] A further embodiment of the method of the present invention is a data stream or a sequence of signals representing a computer program for carrying out the method described herein. The data stream or sequence of signals can be configured to be transferred via a data communication connection, such as the Internet.
[0091] A further embodiment includes processing means, such as a computer or a programmable logic device, configured or adapted to carry out one of the methods described herein.
[0092] A further embodiment includes a computer having installed thereon a computer program for carrying out one of the methods described herein.
[0093] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to carry out some or all of the functions of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to carry out the methods described herein. Generally, these methods are preferably carried out by any hardware device.
[0094] The above embodiments are only for explaining the principle of the present invention. It will be understood that changes and modifications to the configurations and details described herein will be apparent to those skilled in the art. Therefore, the present invention is limited only by the claims, and is not limited by the specific details presented in the description and explanation of the embodiments herein.
Claims
1. An acoustic scene encoder for encoding an acoustic scene (110) containing at least two component signals, a core encoder (160) for core encoding the at least two component signals, the core encoder (160) being configured to generate a first encoded representation (310) for a first portion of the at least two component signals and a second encoded representation (320) for a second portion of the at least two component signals, the first encoded representation (310) including M encoded component signals, the second encoded representation (320) including N encoded component signals, M being greater than N, and N being 1 or more, the core encoder (160); a spatial analyzer (200) for analyzing the acoustic scene (110) to derive one or more spatial parameters (330) or one or more sets of spatial parameters for the second portion of the at least two component signals; an output interface (300) for forming an encoded acoustic scene signal (340), the encoded acoustic scene signal (340) including the first encoded representation (310), the second encoded representation (320), and the one or more spatial parameters (330) or one or more sets of spatial parameters in the second portion of the at least two component signals, the output interface (300); the core encoder (160) is configured to perform parametric processing (160b) on a second frequency subband of a time frame corresponding to the second portion of the at least two component signals, the parametric processing (160b) calculating parameters related to amplitude for the second frequency subband, quantizing and entropy coding the parameters related to amplitude instead of individual spectral lines of the second frequency subband, and the core encoder (160) is configured to quantize and entropy code (160a) individual spectral lines of a first frequency subband of the time frame corresponding to the first portion of the at least two component signals; the first portion of the at least two component signals is a first frequency subband of the at least two component signals, and the second portion of the at least two component signals is a second frequency subband of the at least two component signals; The spatial analyzer (200) is configured to calculate, for the second frequency sub-band, at least one of a directional parameter and an omnidirectional parameter as the one or more spatial parameters (330). An acoustic scene encoder characterized by the above. **Claim 2** The acoustic scene (110) includes an omnidirectional acoustic signal as a first component signal of at least two component signals and at least one directional acoustic signal as a second component signal of the at least two component signals, or The acoustic scene (110) includes a signal captured by an omnidirectional microphone located at a first position as a first component signal of at least two component signals and at least one signal captured by an omnidirectional microphone located at a second position different from the first position as a second component signal of the at least two component signals, or The acoustic scene encoder according to claim 1, wherein the acoustic scene (110) includes at least one signal captured by a directional microphone directed in a first direction as a first component signal of at least two component signals and at least one signal captured by a directional microphone directed in a second direction different from the first direction as a second component signal of the at least two component signals. **Claim 3** The acoustic scene (110) includes an A-format component signal, a B-format component signal, a first-order ambisonics component signal, a higher-order ambisonics component signal, or a component signal determined by a virtual microphone algorithm from an acoustic scene captured by a microphone array having at least two microphone capsules or previously recorded or synthesized, the acoustic scene encoder according to claim 1, characterized in that it comprises. **Claim 4** The output interface (300) is configured such that any spatial parameter from a specific parameter type as the one or more spatial parameters (330) generated by the spatial analyzer (200) for the second part of the at least two component signals is not included in the encoded acoustic scene signal (340), only the second part of the at least two component signals has the specific parameter type, and any parameter of the specific parameter type is not included for the first part of the at least two component signals of the encoded acoustic scene signal (340). The acoustic scene encoder according to claim 1, characterized in that.
5. The core encoder (160) is configured to perform an encoding operation (160a) on the first part of the at least two component signals, or The start band of the second part of the at least two component signals is lower than the bandwidth extension start band, and the core noise filling operation performed by the core encoder (160) has no arbitrary fixed crossover band and is gradually used for more parts of the core spectrum as the frequency increases. The acoustic scene encoder according to claim 1, characterized in that.
6. The second frequency subband of the time frame is a high-frequency subband, the first frequency subband is a low-frequency subband of the time frame, and the core encoder (160) is configured to perform a time-domain encoding operation. The acoustic scene encoder according to claim 1, characterized in that.
7. The time-domain encoding operation includes LPC coding, LPC / TCX coding, EVS coding, AMR Wideband coding or AMR Wideband+ coding. The acoustic scene encoder according to claim 6, characterized in that.
8. The parametric processing (160b) includes a spectral band replication (SBR) process, an intelligent gap filling (IGF) process or a noise filling process. The acoustic scene encoder according to claim 1, characterized in that.
9. The core encoder (160) is configured to use a predetermined boundary frequency between the first frequency subband and the second frequency subband, or The core encoder (160) includes a dimensionality reducer (150a) that reduces the dimension of the acoustic scene (110) equal to M to obtain a low-dimensional acoustic scene equal to N, and the core encoder (160) is configured to calculate the first encoded representation (310) for the first part of the at least two component signals from the low-dimensional acoustic scene equal to N. The spatial analyzer (200) is configured to derive the spatial parameter (330) from the acoustic scene (110) having the dimension that is equal to M and higher than the dimension N of the low-dimensional acoustic scene. The acoustic scene encoder according to claim 1, characterized in that.
10. It is configured to operate at different bitrates, and a predetermined boundary frequency between the first part of the at least two component signals and the second part of the at least two component signals depends on the selected bitrate. The predetermined boundary frequency is small for a smaller bitrate, or the predetermined boundary frequency is large for a larger bitrate. The acoustic scene encoder according to claim 1, characterized in that.
11. The acoustic scene encoder according to claim 1, characterized in that the omnidirectional parameter is a diffusion parameter.
12. The core encoder (160) is a time-frequency converter (164) that converts a sequence of time frames of the at least two component signals into a sequence of spectral frames for the at least two component signals, a spectral encoder (160a) that quantizes and entropy-codes the spectral values of the frames of the sequence of spectral frames within a first frequency subband of the spectral frames, a parametric encoder (160b) that parametrically encodes the spectral values of the spectral frames within a second frequency subband of the spectral frames, or The core encoder (160) includes a time-domain or hybrid time-domain frequency-domain core encoder that performs an encoding operation in the time domain or hybrid time domain frequency domain of the low-band portion of the time frame, or The spatial analyzer (200) is configured to subdivide the second portion of the at least two component signals into analysis bands, and the bandwidth of the analysis bands is greater than or equal to the bandwidth associated with two adjacent spectral values processed by the spectral encoder (160a) within the first portion of the at least two component signals, or lower than the bandwidth of the low-band portion representing the first portion of the at least two component signals. The spatial analyzer (200) is configured to calculate at least one of the directional parameter and the non-directional parameter for each analysis band of the second portion of the at least two component signals, or The acoustic scene encoder according to claim 1, characterized in that the core encoder (160) and the spatial analyzer (200) are configured to use a common filter bank (164) or different filter banks (164, 1000) having different characteristics. **Claim 13** The non-directional parameter is a diffusion parameter, and the spatial analyzer (200) is configured to use an analysis band narrower than the analysis band used to calculate the diffusion parameter to calculate the directional parameter. The acoustic scene encoder according to claim 12. **Claim 14** The core encoder (160) includes a multi-channel encoder that generates an encoded multi-channel signal for the at least two component signals, or The core encoder (160) includes a multi-channel encoder that generates two or more encoded multi-channel signals when the number of component signals of the at least two component signals is three or more, or The core encoder (160) is configured to generate the first encoded representation (310) having a first resolution and the second encoded representation (320) having a second resolution, and the second resolution is lower than the first resolution, or The core encoder (160) is configured to generate the first encoded representation (310) having a first time or first frequency resolution and the second encoded representation (320) having a second time or second frequency resolution, and the second time or frequency resolution is lower than the first time or frequency resolution, or The output interface (300) is configured such that any spatial parameters (330) for the first part of the at least two component signals are not included in the encoded acoustic scene signal (340), or such that a smaller number of spatial parameters for the first part of the at least two component signals are included in the encoded acoustic scene signal (340) compared to the number of the spatial parameters (330) for the second part of the at least two component signals. The acoustic scene encoder according to claim 1, characterized in that.
15. An input interface (400) for receiving an encoded acoustic scene signal (340) comprising a first encoded representation (410) of a first part of at least two component signals, a second encoded representation (420) of a second part of the at least two component signals, and one or more spatial parameters (430) for the second part of the at least two component signals; A core decoder (500) for decoding the first encoded representation (410) and the second encoded representation (420) to obtain decoded representations (810, 820) of the at least two component signals representing an acoustic scene; A spatial analyzer (600) for analyzing a portion (810) of the decoded representation corresponding to the first part of the at least two component signals and deriving one or more spatial parameters (840) for the first part of the at least two component signals, wherein the first part of the at least two component signals includes a first part of the time-frequency representation of the at least two component signals; Using the one or more spatial parameters (840) for the first part of the at least two component signals and the one or more spatial parameters (830) for the second part of the at least two component signals, spatially render the decoded representations (810, 820) to be included in the encoded acoustic scene signal (340), wherein the second part of the at least two component signals includes a second part of the time-frequency representation of the at least two component signals, and the second part includes a spatial renderer (800) different from the first part. An acoustic scene decoder characterized by that.
16. The core decoder (500) is configured to provide a sequence of decoded frames, the first part of the at least two component signals being the first frame of the sequence of decoded frames, the second part of the at least two component signals being the second frame of the sequence of decoded frames, the core decoder (500) further comprising an overlap adder that overlaps and adds subsequent decoded time frames to obtain the decoded representation, or, The acoustic scene decoder according to claim 15, wherein the core decoder (500) performs ACELP-based system operation without an overlapping addition operation.
17. The spatial renderer (800) is configured to operate in a bandwise manner, the first part of the at least two component signals being a first frequency sub-band, the first frequency sub-band being subdivided into a plurality of first bands, the second part of the at least two component signals being a second frequency sub-band, the second frequency sub-band being subdivided into a plurality of second bands, The spatial renderer (800) is configured to render output component signals for each first band using corresponding spatial parameters derived by the spatial analyzer (600). The spatial renderer (800) is configured to render output component signals for each second band using corresponding spatial parameters included in the encoded acoustic scene signal (340), the second bands of the plurality of second bands being larger than the first bands of the plurality of first bands. The spatial renderer (800) is configured to combine (878) the output component signals for the first band and the second band to obtain a rendered output signal, the rendered output signal being a loudspeaker signal, an A-format signal, a B-format signal, a first-order ambisonics signal, a higher-order ambisonics signal, or a binaural signal. The acoustic scene decoder according to claim 15.
18. The core decoder (500) is configured to perform a parametric decoding operation (510b) on the second part of the at least two component signals and a waveform-maintaining decoding operation (510a) on the first part of the at least two component signals, the acoustic scene decoder according to claim 15.
19. The encoded acoustic scene signal (340) includes an encoded multi-channel signal for the at least two component signals, or the encoded acoustic scene signal (340) includes at least two encoded multi-channel signals for a number of component signals greater than two. The core decoder (500) includes a multi-channel decoder that core-decodes the encoded multi-channel signal or the at least two encoded multi-channel signals, the acoustic scene decoder according to claim 15.
20. A method for encoding an acoustic scene (110), the acoustic scene (110) including at least two component signals. A step of core-encoding the at least two component signals, including a step of generating a first encoded representation (310) for a first part of the at least two component signals and a step of generating a second encoded representation (320) for a second part of the at least two component signals, the first encoded representation (310) including M encoded component signals, the second encoded representation (320) including N encoded component signals, M being greater than N, and N being 1 or more, the step of core-encoding the at least two component signals. A step of analyzing the acoustic scene (110) to derive one or more spatial parameters (330) or one or more sets of spatial parameters for the second part of the at least two component signals. A step of generating an encoded acoustic scene signal including the first encoded representation (310), the second encoded representation (320), and the one or more spatial parameters (330) or the one or more sets of spatial parameters for the second part of the at least two component signals. The step of encoding the core performs parametric processing (160b) on a second frequency sub-band of a time frame corresponding to the second portion of the at least two component signals, the parametric processing (160b) calculates parameters related to amplitude for the second frequency sub-band, quantizes and entropy-codes the parameters related to amplitude instead of individual spectral lines of the second frequency sub-band, and quantizes and entropy-encodes (160a) the individual spectral lines of a first frequency sub-band of the time frame corresponding to the first portion of the at least two component signals. The first portion of the at least two component signals is a first frequency sub-band of the at least two component signals, the second portion of the at least two component signals is a second frequency sub-band of the at least two component signals, and the step of analyzing calculates, for the second frequency sub-band, at least one of a directional parameter and a non-directional parameter as the one or more spatial parameters (330). A method for encoding an acoustic scene, characterized in that. [
21. ] A method for decoding an acoustic scene, comprising: Receiving an encoded acoustic scene signal (340) including a first encoded representation (410) of a first portion of at least two component signals, a second encoded representation (420) of a second portion of the at least two component signals, and one or more spatial parameters (430) for the second portion of the at least two component signals; Decoding the first encoded representation (410) and the second encoded representation (420) to obtain a decoded representation of the at least two component signals representing the acoustic scene; Analyzing a portion of the decoded representation corresponding to the first portion of the at least two component signals to derive one or more spatial parameters for the first portion of the at least two component signals, the step of deriving spatial parameters for deriving spatial parameters including a first portion of a time-frequency representation of the at least two component signals. A step of spatially rendering such that the encoded representation is included in the encoded acoustic scene signal (340), using the one or more spatial parameters (840) for the first portion of the at least two component signals and the one or more spatial parameters (430) for the second portion of the at least two component signals, wherein the second portion of the at least two component signals includes a second portion of the time-frequency representation of the at least two component signals, and the second portion is different from the first portion. A method for decoding an acoustic scene, characterized by including the above. Claim 22 A computer program that, when operating on a computer or a processor, executes the method according to claim 20 or the method according to claim 21.
Citation Information
Patent Citations
EP17194816.9
EP17202393.9,
Concepts for Bridging the Gap Between Parametric Multichannel Audio Coding and Matrix Surround Multichannel Coding
JP2009501948A
Generation of multi-channel audio signals
JP2009501957A
Audio encoder for encoding a multichannel signal and audio decoder for decoding an encoded audio signal
US20170365263A1