Parameter Encoding and Decoding

By using a synthesis processor and a mixing rule calculator in an audio synthesizer, a synthetic signal is generated based on channel level and covariance information, and the problem of difficulty in encoding and decoding multi-channel audio content in the prior art at low bit rates is solved, and high-quality and highly adaptable multi-channel audio output is achieved.

CN114270437BActive Publication Date: 2025-05-30FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080057545.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-14
Filing Date
2020-06-15
Publication Date
2025-05-30
Estimated Expiration
2040-06-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively encode and decode multi-channel audio content at low bit rates, especially when using DirAC frameworks, which cannot meet the requirements of high-quality output.

Method used

An audio synthesizer is proposed to receive a downmix signal through an input interface and to generate a synthetic signal based on channel level, related information and covariance information using a synthesis processor. The system includes a prototype signal calculator, a hybrid rule calculator and a synthesis processor, which can reconstruct the target covariance information and number of channels at low bit rates to adapt to different speaker settings.

Benefits of technology

It realizes high-quality multi-channel audio output at low-bit rates, adapts to different speaker settings, keeps the sound quality close to the original signal, and saves the spatial characteristics of the multi-channel signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114270437B_ABST
    Figure CN114270437B_ABST
Patent Text Reader

Abstract

Several examples of encoding and decoding techniques are disclosed. In particular, an audio synthesizer (300) for generating a synthesized signal (336, 340, y) from a downmixed signal (246, x R ) includes: an input interface (312) for receiving the downmixed signal (246, x), the downmixed signal (246, x) having a plurality of downmixed channels and side information (228), the side information (228) including channel levels and correlation information (314, ξ, χ) of an original signal (212, y), the original signal (212, y) having a plurality of original channels; and a synthesis processor (404) for generating the synthesized signal (336, 340, y R ) according to at least one mixing rule using: the channel levels and correlation information (220, 314, ξ, χ) of the original signal (212, y); and covariance information (C x ) associated with the downmixed signal (324, 246, x).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] 1. Introduction

[0002] Here, several examples of encoding and decoding techniques are disclosed. In particular, one invention is directed to encoding and decoding multi-channel audio content at low bitrates, for example using the DirAC framework. This method can achieve high-quality output while using low bitrates. This can be used in many applications, including artworks, communications, and virtual reality. Background Art

[0003] 1.1 Prior Art

[0004] This section briefly describes the prior art.

[0005] 1.1.1 Discrete Coding of Multichannel Content

[0006] The most straightforward method of encoding and transmitting multi-channel content is to directly quantize and encode the waveform of the multi-channel audio signal without any prior processing or assumptions. Although this method can work perfectly in theory, there is a major drawback, namely the bit consumption required to encode the multi-channel content. Therefore, the other methods (and the proposed invention) to be described are so-called "parametric methods" because they use meta-parameters to describe and transmit the multi-channel audio signal rather than the original multi-channel audio signal itself.

[0007] 1.1.2 MPEG Surround

[0008] MPEG Surround is an ISO / MPEG standard completed in 2006 for parametric encoding of multi-channel sound [1]. This method mainly relies on two parameter sets:

[0009] - Interchannel coherences (ICC), which describe the coherence between each channel of a given multi-channel audio signal.

[0010] - Channel Level Difference (CLD), corresponding to the level difference between two input channels of a multi-channel audio signal.

[0011] A particularity of MPEG Surround is the use of so-called "tree structures" that allow "describing two input channels through a single output channel" (quoted from [1]).

[0012] As an example, the encoder scheme for a 5.1 multi-channel audio signal using MPEG Surround can be found below. In this figure, six input channels (labeled "L", "L" in the figure S”, “R”, “R S ”, “C” and “LFE”) are successively processed by tree structure elements (marked as “R_OTT” on the figure). Each of these tree structure elements will generate a set of parameters such as the previously mentioned ICC and CLD and a residual signal, and the residual signal will be processed again by another tree structure and generate another set of parameters. Once reaching the end of the tree, the different parameters previously calculated are transmitted to the decoder, just like the downmixed signal. These elements are used by the decoder to generate an output multi-channel signal, and the decoder processes essentially the inverse tree structure used by the encoder.

[0013] The main advantage of MPEG Surround depends on this structure and the use of the parameters mentioned previously. However, one of the disadvantages of MPEG Surround is due to the lack of flexibility of the tree structure. Also due to the particularity of the processing, quality deterioration may occur in some specific projects.

[0014] Among other things, see Figure 7, which shows an overview of the MPEG Surround encoder for 5.1 signals extracted from [1].

[0015] 1.2 Directional Audio Coding

[0016] Directional Audio Coding (abbreviated as “DirAC”) [2] is also a parametric method for reproducing spatial audio, which was developed by Ville Pulkki at Aalto University in Finland. DirAC relies on band processing, and the band processing uses two sets of parameters to describe spatial sound:

[0017] - Direction of Arrival (DOA), which is an angle in degrees that describes the direction of arrival of the predominant sound in the audio signal.

[0018] - Degree of diffuseness, which is a value between 0 and 1 and is used to describe how “diffuse” the sound is. If the value is 0, the sound is non-diffuse and can be assimilated to a point source from an exact angle; if the value is 1, the sound is fully diffuse and is assumed to come from “every” angle.

[0019] To synthesize the output signal, DirAC assumes that it is decomposed into diffuse and non-diffuse parts. The synthesis of the diffuse sound aims to generate the perception of a surrounding sound, while the synthesis of the direct sound aims to generate the predominant sound.

[0020] While DirAC provides high-quality output, it has a major drawback: it is not applicable to multi-channel audio signals. Therefore, the DOA and diffusion parameters are not very suitable for describing multi-channel audio inputs, and as a result, the output quality is affected.

[0021] 1.3 Binaural Cue Coding

[0022] Binaural Cue Coding (BCC) [3] is a parametric method developed by Christof Faller. This method relies on a similar set of parameters as those described for MPEG Surround (see 1.1.2), namely:

[0023] - Inter-channel level difference (ICLD), which is a measure of the energy ratio between two channels of a multi-channel input signal.

[0024] - Inter-channel time difference (ICTD), which is a measure of the delay between two channels of a multi-channel input signal.

[0025] - Inter-channel correlation (ICC), which is a measure of the correlation between two channels of a multi-channel input signal.

[0026] Compared with the novel invention to be described later, the BCC method has very similar characteristics in terms of the calculation of the transmitted parameters, but it lacks the flexibility and scalability of the transmitted parameters.

[0027] 1.4 MPEG Spatial Audio Object Coding

[0028] Spatial Audio Object Coding [4] will be briefly mentioned here. This is an MPEG standard for encoding so-called audio objects, which is related to multi-channel signals to a certain extent. It uses parameters similar to those of MPEG Surround. Summary of the Invention

[0029] 1.5 Incentives / Disadvantages of the Prior Art

[0030] 1.5.1 Incentives

[0031] 1.5.1.1 Using the DirAC Framework

[0032] One aspect that must be mentioned about the present invention is that the current invention must be suitable for the DirAC framework. Nevertheless, as previously mentioned, the parameters of DirAC are not applicable to multi-channel audio signals. More explanations should be given on this topic.

[0033] The original DirAC processing uses microphone signals or ambisonics signals. From these signals, parameters are calculated, namely the direction of arrival (DOA) and diffuseness.

[0034] To use DirAC with multichannel audio signals, the first approach tried was to use a method proposed by Ville Pulkki to convert the multichannel signals into ambisonic content, as described in [5]. Then, once these ambisonic signals are derived from the multichannel audio signals, the conventional DirAC processing can be used with the DOA and diffuseness. The result of the first attempt was that the quality and spatial characteristics of the output multichannel signals deteriorated and could not meet the requirements of the target applications.

[0035] Therefore, the motivation behind this novel invention is to use a set of parameters that effectively describe the multichannel signals and also use the DirAC framework, and further explanations will be given in Section 1.1.2.

[0036] 1.5.1.2 Providing a System for Low Bit-Rate Operation

[0037] One of the aims and objectives of the present invention is to propose a method that allows for low bitrate applications. This requires finding the optimal dataset to describe the multichannel content between the encoder and the decoder. This also requires finding the best trade-off in terms of the number of transmission parameters and the output quality.

[0038] 1.5.1.3 Provide a flexible system

[0039] Another important objective of the present invention is to propose a flexible system that can accept any multichannel audio format intended to be reproduced on any speaker setup. Depending on the input settings, the output quality should not be impaired.

[0040] 1.5.2 Disadvantages of the Prior Art

[0041] Several of the disadvantages of the aforementioned prior art are listed in the following table.

[0042]

[0043] 2. Description of the Invention 2.1 Summary of the Invention

[0045] According to one aspect, there is provided an audio synthesizer (encoder) for generating a synthesized signal having a plurality of synthesized channels from a downmixed signal, the audio synthesizer comprising:

[0046] An input interface, configured to receive the downmixed signal, the downmixed signal having a plurality of downmixed channels and side information, the side information including the channel levels and correlation information of the original signal, the original signal having a plurality of original channels; and

[0047] A synthesis processor, configured to generate the synthesized signal according to at least one mixing rule, using:

[0048] The channel levels and correlation information of the original signal; and

[0049] Covariance information associated with the downmixed signal.

[0050] The audio synthesizer may include:

[0051] A prototype signal calculator, configured to calculate a prototype signal from the downmixed signal, the prototype signal having the plurality of synthesized channels;

[0052] A mixing rule calculator, configured to calculate at least one mixing rule using:

[0053] The channel levels and correlation information of the original signal; and

[0054] The covariance information associated with the downmixed signal;

[0055] Wherein the synthesis processor is configured to generate the synthesized signal using the prototype signal and the at least one mixing rule.

[0056] The audio synthesizer may be configured to reconstruct the target covariance information of the original signal.

[0057] The audio synthesizer may be configured to reconstruct the target covariance information adapted to the number of channels of the synthesized signal.

[0058] The audio synthesizer may be configured to reconstruct the covariance information adapted to the number of channels of the synthesized signal by allocating groups of original channels to a single synthesized channel, or vice versa, such that the reconstructed target covariance information is reported to the plurality of channels of the synthesized signal.

[0059] The audio synthesizer may be configured to reconstruct the covariance information adapted to the number of channels of the synthesized signal by generating target covariance information for the number of original channels and then applying a downmixing rule or an upmixing rule and energy compensation to obtain the target covariance for the synthesized channels.

[0060] The audio synthesizer may be configured to reconstruct a target version of the covariance information based on an estimated version of the original covariance information, wherein the estimated version of the original covariance information is reported to the plurality of synthesized channels or the plurality of original channels.

[0061] The audio synthesizer may be configured to obtain the estimated version of the original covariance information from covariance information associated with the downmixed signal.

[0062] The audio synthesizer may be configured to obtain the estimated version of the original covariance information by applying an estimation rule to the covariance information associated with the downmixed signal, the estimation rule being associated with a prototype rule for calculating a prototype signal.

[0063] The audio synthesizer may be configured to, for at least one channel pair, normalize the estimated version of the original covariance information (C y ) to the square root of the level of the channels in the channel pair.

[0064] The audio synthesizer may be configured to construct a matrix using the normalized estimated version of the original covariance information.

[0065] The audio synthesizer may be configured to complete the matrix by inserting terms obtained in the side information of the bitstream.

[0066] The audio synthesizer may be configured to denormalize the matrix by scaling the estimated version of the original covariance information by the square root of the level of the channels forming the channel pair.

[0067] The audio synthesizer may be configured to retrieve channel levels and related information among the side information of the downmixed signal, and the audio synthesizer is further configured to reconstruct the target version of the covariance information from the estimated versions of the original channel levels and related information from:

[0068] Covariance information for at least one first channel or channel pair; and

[0069] Channel levels and related information for at least one second channel or channel pair.

[0070] The audio synthesizer may be configured to preferably obtain the channel levels and related information describing the channel or channel pair from the side information of the bitstream rather than the covariance information reconstructed from the downmixed signal for the same channel or channel pair.

[0071] The reconstructed target version of the original covariance information can be understood as describing the energy relationship between a pair of sound channels, or being at least partially based on the levels associated with each of the pair of sound channels.

[0072] The audio synthesizer can be configured to obtain a frequency domain FD version of the downmixed signal, the frequency domain version of the downmixed signal being divided into frequency bands or groups of frequency bands, where different channel levels and associated information are associated with different frequency bands or groups of frequency bands,

[0073] where the audio synthesizer is configured to operate differently for different frequency bands or groups of frequency bands to obtain different mixing rules for different frequency bands or groups of frequency bands.

[0074] The downmixed signal is divided into time slots, where different channel levels and associated information are associated with different time slots, and the audio synthesizer is configured to operate differently for different time slots to obtain different mixing rules for different time slots.

[0075] The downmixed signal is divided into frames, and each frame is divided into time slots, where when the presence and position of a transient in a frame are signaled as being in a transient time slot, the audio synthesizer is configured to:

[0076] associate the current channel level and associated information with the transient time slot and / or the time slots subsequent to the transient time slot of the frame; and

[0077] associate the time slots prior to the transient time slot of the frame with the channel level and associated information of the prior time slots.

[0078] The audio synthesizer can be configured to select a prototype rule, the prototype rule being configured to calculate a prototype signal based on the plurality of synthesized channels.

[0079] The audio synthesizer can be configured to select the prototype rule among a plurality of pre-stored prototype rules.

[0080] The audio synthesizer can be configured to define a prototype rule on a manually selected basis.

[0081] The prototype rule can be based on or include a matrix having a first dimension and a second dimension, where the first dimension is associated with the number of downmixed channels and the second dimension is associated with the number of synthesized channels.

[0082] The audio synthesizer can be configured to operate at a bit rate equal to or lower than 160 kbit / s.

[0083] The audio synthesizer may further include an entropy decoder for obtaining the downmixed signal having the side information.

[0084] The audio synthesizer further includes a decorrelation module to reduce the amount of correlation between different channels.

[0085] The prototype signal may be directly provided to the synthesis processor without performing decorrelation.

[0086] At least one of the channel levels and correlation information of the original signal, the at least one mixing rule, and the covariance information associated with the downmixed signal is in matrix form.

[0087] The side information includes the identification of the original channels;

[0088] Wherein the audio synthesizer may further be configured to calculate the at least one mixing rule using at least one of the channel levels and correlation information of the original signal, the covariance information associated with the downmixed signal, the identification of the original channels, and the identification of the synthesized channels.

[0089] The audio synthesizer may be configured to calculate at least one mixing rule by singular value decomposition (SVD).

[0090] The downmixed signal may be partitioned into frames, and the audio synthesizer is configured to smooth the received parameters, estimated or reconstructed values, or mixing matrix using a linear combination of parameters, estimated or reconstructed values, or mixing matrices obtained for previous frames.

[0091] The audio synthesizer may be configured to deactivate the smoothing of the received parameters, the estimated or reconstructed values, or the mixing matrix when the presence and / or location of a transient in a frame is signaled.

[0092] The downmixed signal may be partitioned into frames, and the frames are partitioned into time slots, wherein the channel levels and correlation information of the original signal are obtained from the side information of the bitstream on a frame-by-frame basis, the audio synthesizer is configured to use a mixing matrix (or mixing rule) for the current frame, the audio synthesizer is configured to use a mixing rule for the current frame, and the mixing rule is obtained by scaling the mixing matrix (or mixing rule) calculated for the current frame by coefficients increasing along the subsequent time slots of the current frame and adding the mixing matrix (or mixing rule) for the previous frame by a scaled version with coefficients decreasing along the subsequent time slots of the current frame.

[0093] The number of the synthesized channels may be greater than the number of the original channels. The number of the synthesized channels may be less than the number of the original channels. The number of the synthesized channels and the number of the original channels may be greater than the number of the downmixed channels.

[0094] At least one or all of the number of the synthesized channels, the number of the original channels, and the number of the downmixed channels are a plural number.

[0095] The at least one mixing rule may include a first mixing matrix and a second mixing matrix, and the audio synthesizer includes:

[0096] A first path, including:

[0097] A first mixing matrix block configured to synthesize a first component of the synthesized signal according to the first mixing matrix calculated from the following:

[0098] A covariance matrix associated with the synthesized signal, the covariance matrix being reconstructed from the channel levels and related information; and

[0099] A covariance matrix associated with the downmixed signal,

[0100] A second path for synthesizing a second component of the synthesized signal, the second component being a residual component, and the second path includes:

[0101] A prototype signal block configured to upmix the downmixed signal from the number of the downmixed channels to the number of the synthesized channels;

[0102] A decorrelator configured to decorrelate the upmixed prototype signal;

[0103] A second mixing matrix block configured to synthesize the second component of the synthesized signal from the decorrelated version of the downmixed signal according to a second mixing matrix, the second mixing matrix being a residual mixing matrix,

[0104] wherein the audio synthesizer is configured to estimate the second mixing matrix from the following:

[0105] A residual covariance matrix provided by the first mixing matrix block; and

[0106] An estimate of the covariance matrix of the decorrelated prototype signal obtained from the covariance matrix associated with the downmixed signal,

[0107] wherein the audio synthesizer further includes an adder block for summing the first component of the synthesized signal and the second component of the synthesized signal.

[0108] According to one aspect, there is provided an audio synthesizer for generating a synthesized signal from a downmixed signal having a plurality of downmixed channels, the synthesized signal having a plurality of synthesized channels, the downmixed signal being a downmixed version of an original signal having a plurality of original channels, the audio synthesizer comprising:

[0109] A first path, comprising:

[0110] A first mixing matrix block configured to synthesize a first component of the synthesized signal according to a first mixing matrix calculated from:

[0111] The covariance matrix associated with the synthesized signal; and

[0112] The covariance matrix associated with the downmixed signal;

[0113] A second path for synthesizing a second component of the synthesized signal, wherein the second component is a residual component, the second path comprising:

[0114] A prototype signal block configured to upmix the downmixed signal from the number of downmixed channels to the number of synthesized channels;

[0115] A decorrelator configured to decorrelate the upmixed prototype signal (613c);

[0116] A second mixing matrix block configured to synthesize the second component of the synthesized signal from the decorrelated version of the downmixed signal according to a second mixing matrix, the second mixing matrix being a residual mixing matrix,

[0117] wherein the audio synthesizer is configured to calculate the second mixing matrix from:

[0118] The residual covariance matrix provided by the first mixing matrix block; and

[0119] An estimate of the covariance matrix of the decorrelated prototype signal obtained from the covariance matrix associated with the downmixed signal,

[0120] wherein the audio synthesizer further comprises an adder block for summing the first component of the synthesized signal with the second component of the synthesized signal.

[0121] The residual covariance matrix is obtained by subtracting the matrix obtained by applying the first mixing matrix to the covariance matrix associated with the downmixed signal from the covariance matrix associated with the synthesized signal.

[0122] The audio synthesizer may be configured to define the second mixing matrix from:

[0123] A second matrix obtained by decomposing the residual covariance matrix associated with the synthesized signal;

[0124] A first matrix that is the inverse matrix or regularized inverse matrix of a diagonal matrix obtained from the estimate of the covariance matrix of the decorrelated prototype signals.

[0125] The diagonal matrix can be obtained by applying the square root function to the main diagonal elements of the covariance matrix of the decorrelated prototype signals.

[0126] The second matrix can be obtained by applying singular value decomposition to the residual covariance matrix associated with the synthesized signal.

[0127] The audio synthesizer can be configured to define the second mixing matrix by multiplying the second matrix with the inverse matrix or regularized inverse matrix of the diagonal matrix obtained from the estimate of the covariance matrix of the decorrelated prototype signals and a third matrix.

[0128] The audio synthesizer can be configured to obtain the third matrix by applying singular value decomposition to a matrix obtained from a normalized version of the covariance matrix of the decorrelated prototype signals, where the normalization is with respect to the main diagonals of the residual covariance matrix, the diagonal matrix, and the second matrix.

[0129] The audio synthesizer can be configured to define the first mixing matrix from the second matrix and the inverse matrix or regularized inverse matrix of the second matrix,

[0130] where the second matrix is obtained by decomposing the covariance matrix associated with the downmixed signal, and the second matrix is obtained by decomposing the reconstructed target covariance matrix associated with the downmixed signal.

[0131] The audio synthesizer can be configured to estimate the covariance matrix of the decorrelated prototype signals from the diagonal terms of a matrix obtained by applying a prototype rule used at the prototype block to upmix the downmixed signal from the number of downmixed channels to the number of synthesized channels to the covariance matrix associated with the downmixed signal.

[0132] The frequency bands are aggregated into aggregated frequency band groups, where information about the aggregated frequency band groups is provided in the side information of the bitstream, and the channel levels and related information of the original signal are provided for each frequency band group to calculate the same at least one mixing matrix for different frequency bands of the same aggregated frequency band group.

[0133] According to one aspect, there is provided an audio encoder for generating a downmixed signal from an original signal having a plurality of original channels, the downmixed signal having a plurality of downmixed channels, the audio encoder comprising:

[0134] a parameter estimator configured to estimate channel levels and correlation information of the original signal, and

[0135] a bitstream writer for encoding the downmixed signal into a bitstream such that the downmixed signal is encoded in the bitstream to have side information including the channel levels and correlation information of the original signal.

[0136] The audio encoder may be configured to provide the channel levels and correlation information of the original signal as normalized values.

[0137] The channel levels and correlation information of the original signal encoded in the side information represent at least channel level information associated with the total number of the original channels.

[0138] The channel levels and correlation information of the original signal encoded in the side information represent at least correlation information describing an energy relationship between at least a pair, but less than the total number of, different original channels.

[0139] The channel levels and correlation information of the original signal include at least one coherence value describing the coherence between two channels in a pair of original channels.

[0140] The coherence value may be normalized. The coherence value may be

[0141]

[0142] where C yi,j is the covariance between channels i and j, C yi,i and C yj,j are the levels associated with channels i and j, respectively.

[0143] The channel levels and correlation information of the original signal include at least one inter-channel level difference (ICLD).

[0144] The at least one ICLD may be provided as a logarithmic value. The at least one ICLD is normalized. The at least one ICLD may be

[0145]

[0146] where

[0147] -χ iis the inter-channel level difference for channel i,

[0148] -P i is the power of the current channel i,

[0149] -P dmx,i is a linear combination of values of the covariance information of the downmixed signal.

[0150] The audio encoder may be configured to select, based on state information, whether to encode at least a part of the channel levels and associated information of the original signal or not encode it, so as to include an increased amount of channel levels and associated information in the side information in the case of a relatively low payload.

[0151] The audio encoder may be configured to select, based on a measure regarding the channels, which part of the channel levels and associated information of the original signal to encode in the side information, so as to include in the side information the channel levels and associated information associated with a more sensitive measure.

[0152] The channel levels and associated information of the original signal may be in the form of terms of a matrix.

[0153] The matrix may be a symmetric matrix or a Hermitian matrix, where the terms of the channel levels and associated information are provided for all terms in the diagonal of the matrix or less than the total number of terms and / or for less than half of the non-diagonal elements of the matrix.

[0154] The bitstream writer is configured to encode the identification of at least one channel.

[0155] The original signal or its processed version may be divided into a plurality of subsequent frames having equal time lengths.

[0156] The audio encoder may be configured to encode in the side information the channel levels and associated information of the original signal specific to each frame.

[0157] The audio encoder may be configured to encode in the side information the same channel levels and associated information of the original signal that are commonly associated with a plurality of consecutive frames.

[0158] The audio encoder may be configured to select the number of consecutive frames for which the same channel levels and associated information of the original signal are selected such that:

[0159] A relatively high bit rate or a high payload implicitly indicates an increase in the number of consecutive frames associated with the same channel levels and associated information of the original signal, and vice versa.

[0160] The audio encoder may be configured to reduce the number of consecutive frames associated with the same channel level and related information of the original signal when a transient is detected.

[0161] Each frame may be subdivided into an integer number of consecutive time slots.

[0162] The audio encoder may be configured to estimate the channel level and related information for each time slot, and encode in the side information the sum, or the average, or another predetermined linear combination of the channel levels and related information estimated for different time slots.

[0163] The audio encoder may be configured to perform a transient analysis on a time-domain version of the frame to determine the occurrence of a transient within the frame.

[0164] The audio encoder may be configured to determine in which time slot of the frame the transient has occurred, and:

[0165] encode the channel level and related information of the original signal associated with the time slot in which the transient has occurred and / or subsequent time slots in the frame,

[0166] and not encode the channel level and related information of the original signal associated with time slots prior to the transient.

[0167] The audio encoder may be configured to signal in the side information that the occurrence of the transient has occurred in one time slot of the frame.

[0168] The audio encoder may be configured to signal in the side information in which time slot of the frame the transient has occurred.

[0169] The audio encoder may be configured to estimate the channel level and related information of the original signal associated with multiple time slots of the frame, and sum them, or average them, or linearly combine them, to obtain the channel level and related information associated with the frame.

[0170] The original signal may be transformed into a frequency-domain signal, wherein the audio encoder is configured to encode the channel level and related information of the original signal in the side information in a per-frequency-band manner.

[0171] The audio encoder may be configured to aggregate multiple frequency bands of the original signal into a reduced number of frequency bands, in order to encode the channel level and related information of the original signal in the side information in a per-aggregated-frequency-band manner.

[0172] The audio encoder may be configured to further aggregate the frequency bands in case a transient in the frame is detected, such that:

[0173] the number of frequency bands is reduced; and / or

[0174] the width of at least one frequency band is increased by aggregating it with another frequency band.

[0175] The audio encoder may also be configured to code at least one channel level and associated information of a frequency band as an increment relative to a previously coded channel level and associated information in the bitstream.

[0176] The audio encoder may be configured to code an incomplete version of the channel level and associated information relative to the channel level and associated information estimated by the estimator in the side information of the bitstream.

[0177] The audio encoder may be configured to adaptively select the selected information to be coded in the side information of the bitstream among the overall channel levels and associated information estimated by the estimator, such that the remaining unselected information of the channel levels and / or associated information estimated by the estimator is not coded.

[0178] The audio encoder may be configured to reconstruct the channel level and associated information from the selected channel level and associated information, thereby simulating the estimation of the unselected channel level and associated information at the decoder, and calculating error information between:

[0179] the unselected channel level and associated information estimated by the encoder; and

[0180] the unselected channel level and associated information reconstructed by simulating the estimation of the uncoded channel level and associated information at the decoder; and

[0181] such that a distinction is made based on the calculated error information between:

[0182] channel levels and associated information that can be correctly reconstructed; and

[0183] channel levels and associated information that cannot be correctly reconstructed,

[0184] to determine:

[0185] select the channel levels and associated information that cannot be correctly reconstructed to be coded in the side information of the bitstream; and

[0186] do not select the channel levels and associated information that can be correctly reconstructed, thereby avoiding coding the channel levels and associated information that can be correctly reconstructed in the side information of the bitstream.

[0187] The channel levels and associated information can be indexed according to a predetermined ordering, wherein the encoder is configured to signal in the side information of the bitstream an index associated with the predetermined ordering, the index indicating which of the channel levels and associated information are encoded. The index is provided by a bitmap. The index is defined according to a combined numbering system that associates a one-dimensional index with the entries of a matrix.

[0188] The audio encoder can be configured to select between:

[0189] an adaptive provision of the channel levels and associated information, wherein the index associated with the predetermined ordering is encoded in the side information of the bitstream; and

[0190] a fixed provision of the channel levels and associated information such that the encoded channel levels and associated information are predetermined and sorted according to a predetermined fixed order without providing an index.

[0191] The audio encoder can be configured to signal in the side information of the bitstream whether the channel levels and associated information are provided according to an adaptive provision or a fixed provision.

[0192] The audio encoder can also be configured to encode the current channel levels and associated information as deltas relative to previous channel levels and associated information in the bitstream.

[0193] The audio encoder can also be configured to generate the downmix signal according to a static downmix.

[0194] According to one aspect, there is provided a method for generating a synthesis signal from a downmix signal, the synthesis signal having a plurality of synthesis channels, the method comprising:

[0195] receiving a downmix signal and side information, the downmix signal having a plurality of downmix channels, the side information including:

[0196] channel levels and associated information of an original signal, the original signal having a plurality of original channels;

[0197] generating the synthesis signal using the channel levels and associated information of the original signal and covariance information associated with the signal.

[0198] The method can include:

[0199] calculating a prototype signal from the downmix signal, the prototype signal having the plurality of synthesis channels;

[0200] Calculate a mixing rule using the channel levels and correlation information of the original signal and the covariance information associated with the downmixed signal; and

[0201] Generate the synthesized signal using the prototype signal and the mixing rule.

[0202] According to one aspect, there is provided a method for generating a downmixed signal from an original signal, the original signal having a plurality of original channels, the downmixed signal having a plurality of downmixed channels, the method comprising:

[0203] Estimate the channel levels and correlation information of the original signal,

[0204] Encode the downmixed signal into a bitstream such that the downmixed signal is encoded in the bitstream to have side information including the channel levels and correlation information of the original signal.

[0205] According to one aspect, there is provided a method for generating a synthesized signal from a downmixed signal having a plurality of downmixed channels, the synthesized signal having a plurality of synthesized channels, the downmixed signal being a downmixed version of an original signal having a plurality of original channels, the method comprising the following stages:

[0206] A first stage, comprising:

[0207] Synthesize a first component of the synthesized signal according to a first mixing matrix calculated from:

[0208] The covariance matrix associated with the synthesized signal; and

[0209] The covariance matrix associated with the downmixed signal,

[0210] A second stage for synthesizing a second component of the synthesized signal, wherein the second component is a residual component, the second stage comprising:

[0211] A prototype signal step of upmixing the downmixed signal from the number of downmixed channels to the number of synthesized channels;

[0212] A decorrelator step of decorrelating the upmixed prototype signal;

[0213] A second mixing matrix step of synthesizing the second component of the synthesized signal from the decorrelated version of the downmixed signal according to a second mixing matrix, the second mixing matrix being a residual mixing matrix,

[0214] wherein the method calculates the second mixing matrix from:

[0215] The residual covariance matrix provided by the first mixing matrix step; and

[0216] An estimate of the covariance matrix of the decorrelated prototype signals obtained from the covariance matrix associated with the downmixed signal

[0217] wherein the method further includes an adder step of summing the first component of the synthesized signal and the second component of the synthesized signal, thereby obtaining the synthesized signal.

[0218] According to one aspect, there is provided an audio synthesizer for generating a synthesized signal from a downmixed signal, the synthesized signal having a number of synthesized channels greater than one or greater than two, the audio synthesizer including at least one of the following:

[0219] An input interface configured to receive the downmixed signal, the downmixed signal having at least one downmixed channel and side information, the side information including at least one of the following:

[0220] Channel levels and correlation information of the original signal, the original signal having a plurality of original channels, the number of the original channels being greater than one or greater than two;

[0221] A component, such as a prototype signal calculator [e.g., "prototype signal calculation"], configured to calculate a prototype signal from the downmixed signal, the prototype signal having the number of synthesized channels;

[0222] A component, such as a mixing rule calculator [e.g., "parameter reconstruction"], configured to calculate one or more mixing rules using the channel levels and correlation information of the original signal and covariance information associated with the downmixed signal; and

[0223] A component, such as a synthesis processor [e.g., "synthesis engine"], configured to generate the synthesized signal using the prototype signal and the mixing rules.

[0224] The number of synthesized channels may be greater than the number of original channels. Alternatively, the number of synthesized channels may be less than the number of original channels.

[0225] The audio synthesizer (in particular, in some aspects, the mixing rule calculator) may be configured to reconstruct a target version of the original channel levels and correlation information.

[0226] The audio synthesizer (in particular, in some aspects, the mixing rule calculator) may be configured to reconstruct a target version of the original channel levels and correlation information, the correlation information being adapted to the plurality of channels of the synthesized signal.

[0227] The audio synthesizer (in particular, in some aspects, the mixing rule calculator) may be configured to reconstruct a target version of the original channel levels and associated information, the associated information being based on an estimated version of the original channel levels and associated information.

[0228] The audio synthesizer (in particular, in some aspects, the mixing rule calculator) may be configured to obtain the estimated version of the original channel levels and associated information from covariance information associated with the downmixed signal.

[0229] The audio synthesizer (in particular, in some aspects, the mixing rule calculator) may be configured to, for the prototype signal, obtain the estimated version of the original channel levels and associated information by applying an estimated rule associated with the prototype rule used by the prototype signal calculator to the covariance information associated with the downmixed signal.

[0230] The audio synthesizer (in particular, in some aspects, the mixing rule calculator) may be configured to retrieve, among the side information of the downmixed signal:

[0231] Covariance information associated with the downmixed signal, describing the level of a first channel in the downmixed signal or the energy relationship between channel pairs; and

[0232] The channel levels and associated information of the original signal, describing the level of a first channel in the original signal or the energy relationship between channel pairs,

[0233] such that the target version of the original channel levels and associated information is reconstructed by using at least one of:

[0234] The covariance information of the original channels for at least one first channel or channel pair; and

[0235] Describing the channel levels and associated information of the at least one first channel or channel pair.

[0236] The audio synthesizer (in particular, in some aspects, the mixing rule calculator) may be configured to prefer that the channel levels and associated information describe the channel or channel pair rather than the covariance information of the original channels for the same channel or channel pair.

[0237] The reconstructed target version of the original channel levels and associated information describes the energy relationship between channel pairs at least partially based on the levels associated with each channel in the channel pair.

[0238] The downmixed signal can be divided into frequency bands or groups of frequency bands: different channel levels and associated information can be associated with different frequency bands or groups of frequency bands; the audio synthesizer (the prototype signal calculator, in particular, in some aspects, at least one of the mixing rule calculator and the synthesis processor) is configured to operate differently for different frequency bands or groups of frequency bands to obtain different mixing rules for different frequency bands or groups of frequency bands.

[0239] The downmixed signal can be divided into time slots, where different channel levels and associated information are associated with different time slots, and at least one component of the audio synthesizer (such as the prototype signal calculator, the mixing rule calculator, the synthesis processor, or other elements of the synthesizer) is configured to operate differently for different time slots to obtain different mixing rules for different time slots.

[0240] The audio synthesizer (such as the prototype signal calculator) can be configured to select a prototype rule that is configured to calculate a prototype signal based on the number of synthesized channels.

[0241] The audio synthesizer (such as the prototype signal calculator) can be configured to select the prototype rule among pre-stored prototype rules.

[0242] The audio synthesizer (such as the prototype signal calculator) can be configured to define a prototype rule based on a manual selection.

[0243] The prototype rule (such as the prototype signal calculator) can include a matrix that has a first dimension and a second dimension, where the first dimension is associated with the number of downmixed channels and the second dimension is associated with the number of synthesized channels.

[0244] The audio synthesizer (such as the prototype signal calculator) can be configured to operate at a bit rate equal to or lower than 160 kbit / s.

[0245] The side information can include the identification of the original channels [such as L, R, C, etc.].

[0246] The audio synthesizer (in particular, in some aspects, the mixing rule calculator) can be configured to calculate [such as "parameter reconstruction"] a mixing rule [such as a mixing matrix] using the channel levels and associated information of the original signal, the covariance information associated with the downmixed signal, the identification of the original channels, and the identification of the synthesized channels.

[0247] The audio synthesizer can select the number of channels for the synthesized signal [e.g., by a selection such as a manual selection, or by a pre-selection, or automatically e.g., by identifying the number of speakers], the number of channels being independent of at least one of the channel levels and the associated information of the original channels in the side information.

[0248] In some examples, the audio synthesizer can select different prototype rules for different selections. The mixing rule calculator can be configured to calculate the mixing rule.

[0249] According to one aspect, there is provided a method for generating a synthesized signal from a downmixed signal, the synthesized signal having a plurality of synthesized channels, the number of synthesized channels being greater than one or greater than two, the method comprising:

[0250] Receiving the downmixed signal, the downmixed signal having at least one downmixed channel and side information, the side information including:

[0251] Channel levels and associated information of the original signal, the original signal having a plurality of original channels, the number of original channels being greater than one or greater than two;

[0252] Calculating a prototype signal from the downmixed signal, the prototype signal having the number of synthesized signals;

[0253] Calculating a mixing rule using the channel levels and associated information of the original signal and covariance information associated with the downmixed signal; and

[0254] Generating the synthesized signal using the prototype signal and the mixing rule [e.g., rule].

[0255] According to one aspect, there is provided an audio encoder for generating a downmixed signal from an original signal [e.g., y], the original signal having at least two channels, the downmixed signal having at least one downmixed channel, the audio encoder comprising at least one of the following:

[0256] A parameter estimator configured to estimate the channel levels and associated information of the original signal,

[0257] A bitstream writer for encoding the downmixed signal into a bitstream such that the downmixed signal is encoded in the bitstream with side information, the side information including the channel levels and associated information of the original signal.

[0258] The channel levels and associated information of the original signal encoded in the side information represent channel level information associated with less than the total number of channels of the original signal.

[0259] The channel levels and associated information of the original signal encoded in the side information represent associated information that describes the energy relationship between at least one pair of different channels in the original channels, but less than the total number of channels of the original signal.

[0260] The channel levels and associated information of the original signal may include at least one coherence value that describes the coherence between two channels in a channel pair.

[0261] The channel levels and associated information of the original signal may include at least one inter-channel level difference (ICLD) between two channels of a channel pair.

[0262] The audio encoder may be configured to select whether to encode or not encode at least a portion of the channel levels and associated information of the original signal based on the state information, to include an increased amount of channel levels and associated information in the side information in the case of a relatively low payload.

[0263] The audio encoder may be configured to select which portion of the channel levels and associated information of the original signal is to be encoded in the side information based on a measure of the channels, to include in the side information the channel levels and associated information associated with a more sensitive measure [e.g., a measure associated with a perceptually more significant covariance].

[0264] The channel levels and associated information of the original signal may be in the form of a matrix.

[0265] The bitstream writer is configured to encode the identification of at least one channel.

[0266] According to one aspect, a method is provided for generating a downmixed signal from an original signal having at least two channels, the downmixed signal having at least one downmixed channel.

[0267] The method may include:

[0268] estimating the channel levels and associated information of the original signal,

[0269] encoding the downmixed signal into a bitstream such that the downmixed signal is encoded in the bitstream to have side information that includes the channel levels and associated information of the original signal.

[0270] The audio encoder may be agnostic to the decoder. The audio synthesizer may be agnostic to the decoder.

[0271] According to one aspect, a system is provided that includes the audio synthesizer as described above or below and one of the audio encoders as described above or below.

[0272] According to one aspect, a non - transitory storage unit storing instructions is provided, which when executed by a processor causes the processor to perform a method as described above or below. BRIEF DESCRIPTION OF THE DRAWINGS

[0273] 3. Examples

[0274] 3.1 Drawings

[0275] FIG. 1 shows a simplified overview of the processing according to the present invention.

[0276] Figure 2a An audio encoder according to the present invention is shown.

[0277] FIG. 2b shows another view of the audio encoder according to the present invention.

[0278] Figure 2c Another view of the audio encoder according to the present invention is shown.

[0279] Figure 2d Another view of the audio encoder according to the present invention is shown.

[0280] Figure 3a An audio synthesizer (decoder) according to the present invention is shown.

[0281] FIG. 3b shows another view of the audio synthesizer (decoder) according to the present invention.

[0282] Figure 3c Another view of the audio synthesizer (decoder) according to the present invention is shown.

[0283] Figures 4a - 4d An example of covariance synthesis is shown.

[0284] Figure 5 An example of a filter bank for an audio encoder according to the present invention is shown.

[0285] Figures 6a - 6c An example of the operation of an audio encoder according to the present invention is shown.

[0286] FIG. 7 shows an example of the prior art.

[0287] Figures 8a - 8c An example of how to obtain covariance information according to the present invention is shown.

[0288] Figures 9a - 9d An example of an inter - channel coherence matrix is shown.

[0289] Figures 10a - 10b An example of a frame is shown.

[0290] Figure 11 A scheme used by the decoder to obtain a mixing matrix is shown. Detailed Description

[0291] 3.2 Regarding the Concept of the Invention

[0292] It will be shown that the example is based on the encoder downmixing the signal 212 and providing the decoder with channel level and correlation information 220. The decoder can generate a mixing rule (e.g., a mixing matrix) from the channel level and correlation information 220. Information important for generating the mixing rule can include covariance information (e.g., covariance matrix C y ) of the original signal 212 and covariance information of the downmixed signal (e.g., covariance matrix C x ). Although the covariance matrix C x can be directly estimated by the decoder by analyzing the downmixed signal, the covariance matrix C y of the original signal 212 is difficult for the decoder to estimate. The covariance matrix C y of the original signal 212 is usually a symmetric matrix (e.g., a 5x5 matrix in the case of a 5-channel original signal 212): while the matrix shows the level of each channel on the diagonal, it presents the covariance between channels at the non-diagonal entries. The matrix is a diagonal matrix because the covariance between general channels i and j is the same as the covariance between j and i. Therefore, in order to provide the decoder with the entire covariance information, it is necessary to signal to the decoder the 5 levels at the diagonal entries and the 10 covariances at the non-diagonal entries. However, it will be shown that it is possible to reduce the amount of information to be encoded.

[0293] In addition, it will be shown that in some cases, instead of providing levels and covariances, normalized values can be provided. For example, an inter-channel coherence value (ICC, inter channel coherence, also denoted by ξ i,j ) indicating the energy value and an inter-channel level difference (ICLD, inter channel level difference, also denoted by χ i ) can be provided. The ICC can be, for example, a correlation value provided instead of the matrix C yCovariance of the off - diagonal terms. An example of relevant information can be in the form of. In some examples, only a part of ξ i,j is actually encoded.

[0294] In this way, an ICC matrix is generated. The diagonal terms of the ICC matrix will, in principle, be equal to 1 and thus do not have to be encoded in the bitstream. However, it has been understood that it is feasible for the encoder to provide the ICLD to the decoder, for example in the form of (see also below). In some examples, all χ i are actually encoded.

[0295] Figures 9a to 9d Fig. shows an example of an ICC matrix 900, where the diagonal value "d" can be the ICLD χ i and the off - diagonal values are indicated by 902, 904, 905, 906, 907 (see below), which can be the ICC ξ i,j .

[0296] In this document, the product between matrices is indicated in an unsigned manner. For example, the product between matrix A and matrix B is indicated by AB. The conjugate transpose of a matrix is indicated by an asterisk (*).

[0297] When referring to the diagonal, it means the main diagonal.

[0298] 3.3 The Invention

[0299] FIG. 1 shows an audio system 100 having an encoder side and a decoder side. The encoder side can be implemented by an encoder 200 and can obtain an audio signal 212, for example, from an audio sensor unit (such as a microphone), or can be obtained from a storage unit or from a remote unit (such as via radio transmission). The decoder side can be implemented by an audio decoder (audio synthesizer) 300, which can provide audio content to an audio reproduction unit (such as a speaker). The encoder 200 and the decoder 300 can communicate with each other, for example, through a communication channel, which can be wired or wireless (such as via radio frequency waves, light, or ultrasonic waves, etc.). The encoder and / or the decoder can thus include or be connected to a communication unit (such as an antenna, a transceiver, etc.) for transmitting the encoded bitstream 248 from the encoder 200 to the decoder 300. In some cases, the encoder 200 can store the encoded bitstream 248 in a storage unit (such as a RAM memory, a FLASH memory, etc.) for future use. Similarly, the decoder 300 can read the bitstream 248 stored in the storage unit. In certain examples, the encoder 200 and the decoder 300 can be the same device: after the bitstream 248 has been encoded and stored, the device may need to read it to play back the audio content.

[0300] Figure 2a , 2b, 2c, and 2d show examples of the encoder 200. In certain examples, Figure 2a the encoders of 2b, 2c, and 2d can be the same and differ from each other only by the absence of certain elements in one and / or the other figure.

[0301] The audio encoder 200 can be configured to generate a downmix signal 246 from an original signal 212 (an original signal 212 having at least two (such as three or more) channels and a downmix signal 246 having at least one downmixed channel).

[0302] The audio encoder 200 can include a parameter estimator 218 configured to estimate the channel levels and related information 220 of the original signal 212. The audio encoder 200 can include a bitstream writer 226 for encoding the downmix signal 246 into the bitstream 248. Thus, the downmix signal 246 is encoded in the bitstream 248 in such a way that it has side information 228 including the channel levels and related information of the original signal 212.

[0303] In particular, in some examples, the input signal 212 may be understood as a time-domain audio signal, such as for example a time series of audio samples. The original signal 212 has at least two channels, which may for example correspond to different microphones (e.g., for stereo audio positions, or however, multi-channel audio positions), or for example correspond to different speaker positions of an audio reproduction unit. The input signal 212 may be downmixed at the downmixer computation block 244 to obtain a downmixed version 246 of the original signal 212 (also denoted as x). This downmixed version of the original signal 212 is also referred to as the downmixed signal 246. The downmixed signal 246 has at least one downmixed channel. The downmixed signal 246 has fewer channels than the original signal 212. The downmixed signal 212 may be in the time domain.

[0304] The downmixed signal 246 is encoded in the bitstream 248 by a bitstream writer 226 (e.g., including an entropy encoder, or a multiplexer, or a core encoder) for storing or transmitting the bitstream to a receiver (e.g., associated with the decoder side). The encoder 200 may include a parameter estimator (or parameter estimation block) 218. The parameter estimator 218 may estimate the channel levels and associated information 220 associated with the original signal 212. The channel levels and associated information 220 may be encoded in the bitstream 248 as side information 228. In an example, the channel levels and associated information 220 are encoded by the bitstream writer 226. In an example, even if Figure 2b does not show the bitstream writer 226 downstream of the downmixer computation block 235, the bitstream writer 226 may still be present. In Figure 2c it is shown that the bitstream writer 226 may include a core encoder 247 to encode the downmixed signal 246 to obtain an encoded version of the downmixed signal 246. Figure 2c It is also shown that the bitstream writer 226 may include a multiplexer 249 that multiplexes both the encoded downmixed signal 246 and the channel levels and associated information 220 (e.g., as encoded parameters) in the side information 228 in the bitstream 228.

[0305] As shown in Figure 2b (missing in Figure 2a and 2c ), the original signal 212 may be processed (e.g., by a filter bank 214, see below) to obtain a frequency-domain version 216 of the original signal 212.

[0306] An example of parameter estimation is shown in Figure 6c where the parameter estimator 218 defines parameters ξ i,j and χ i (e.g., normalized parameters) to be subsequently encoded in the bitstream. Covariance estimators 502 and 504 estimate the covariance C for the downmixed signal 246 and the input signal 212 to be encoded, respectively.x and C y Then, at the ICLD block 506, the ICLD parameter χ i is computed and provided to the bitstream writer 246. At the covariance-to-coherence block 510, the ICC ξ i,j (412) is obtained. At block 250, only some of the ICCs are selected to be encoded.

[0307] The parameter quantization block 222 (FIG. 2b) may allow obtaining the channel levels and correlation information 220 in the quantized version 224.

[0308] The channel levels and correlation information 220 of the original signal 212 may generally include information about the energy (or level) of the channels of the original signal 212. Additionally or alternatively, the channel levels and correlation information 220 of the original signal 212 may include correlation information between channel pairs, such as the correlation between two different channels. The channel levels and correlation information may include information associated with the covariance matrix C y (e.g., in its normalized form, such as correlation or ICC), where each column and each row is associated with a particular channel of the original signal 212, and the diagonal elements of the matrix C y and the correlation information are used to describe the channel levels, and the off-diagonal elements of the matrix C y are used to describe the correlation information. The matrix C y may be a symmetric matrix (i.e., it is equal to its transpose matrix) or a Hermitian matrix (i.e., it is equal to its conjugate transpose). C y is generally positive semidefinite. In some examples, the correlation may be replaced by covariance (and the correlation information is replaced by covariance information). It has been understood that it is feasible to encode information associated with less than the total number of channels of the original signal 212 in the side information 228 of the bitstream 248. For example, it is not necessary to provide the channel levels and correlation information about all channels or all channel pairs. For example, it may be possible to encode only a reduced set of information about the correlation between channel pairs of the downmixed signal 212 in the bitstream 248, while the remaining information may be estimated at the decoder side. Generally, it is feasible to encode fewer elements than the diagonal elements of C y and it is feasible to encode fewer elements than the elements outside the diagonal of C y .

[0309] For example, the channel levels and correlation information may include the covariance matrix C of the original signal 212 y(Channel levels and associated information 220 of the original signal) and / or covariance matrix C of the downmixed signal 246 x Entries of (covariance information of the downmixed signal), e.g., in a normalized form. For example, the covariance matrix can associate each row and each column with each channel to represent the covariance between different channels, and the diagonal of the matrix represents the level of each channel. In some examples, the channel levels and associated information 220 of the original signal 212 encoded in the side information 228 can include only channel level information (e.g., only the diagonal values of the correlation matrix C y ) or only the associated information (e.g., only the values outside the diagonal of the correlation matrix C y ). The same applies to the covariance information of the downmixed signal.

[0310] As will be shown later, the channel levels and associated information 220 can include at least one coherence value (ξ i,j ), describing the coherence between two channels i and j in the channel pair i, j. Additionally or alternatively, the channel levels and associated information 220 can include at least one inter-channel level difference ICLD (χ i ). In particular, it is feasible to define a matrix with ICLD values or ICC values. Thus, the above examples of the transmission of the elements of matrices C y and C x can be generalized for other values to be encoded (e.g., transmitted) for implementing the channel levels and associated information 220 and / or the coherence information of the downmixed channels.

[0311] The input signal 212 can be subdivided into a plurality of frames. Different frames can have, for example, the same time length (e.g., each frame can be constructed from the same number of samples in the time domain during the time period of one frame). Thus, different frames generally have equal time lengths. In the bitstream 248, the downmixed signal 246 (which can be a time-domain signal) can be encoded on a per-frame basis (or in any case, it can be determined by the decoder to be subdivided into frames). As encoded as side information 228 in the bitstream 248, the channel level and associated information 220 can be associated with each frame (e.g., parameters of the channel level and associated information 220 can be provided for each frame or for a plurality of consecutive frames). Accordingly, for each frame of the downmixed signal 246, the associated side information 228 (e.g., parameters) can be encoded in the side information 228 of the bitstream 248. In some cases, a plurality of consecutive frames can be associated with the same channel level and associated information 220 (e.g., the same parameters) encoded in the side information 228 of the bitstream 248. Accordingly, one parameter can result in being commonly associated with a plurality of consecutive frames. In certain examples, this can occur when two consecutive frames have similar attributes, or when the bit rate needs to be reduced (e.g., due to the necessity of reducing the payload). For example:

[0312] In the case of a high payload, increase the number of consecutive frames associated with the same specific parameter to reduce the number of bits written to the bitstream;

[0313] In the case of a low payload, reduce the number of consecutive frames associated with the same specific parameter to improve the mixing quality.

[0314] In other cases, when the bit rate is reduced, increase the number of consecutive frames associated with the same specific parameter to reduce the number of bits written to the bitstream, and vice versa.

[0315] In certain cases, it is feasible to use a linear combination of the parameters (or reconstructed or estimated values, such as covariance) of the previous frame of the current frame, e.g., by addition, averaging, etc., to smooth the parameters (or reconstructed or estimated values, such as covariance).

[0316] In certain examples, a frame can be divided among a plurality of subsequent time slots. Figure 10a Frame 920 (subdivided into four consecutive time slots 921 to 924) is shown, Figure 10b Frame 930 (subdivided into four consecutive time slots 931 to 934) is shown. The time lengths of different time slots can be the same. If the length of a frame is 20 ms and the time slot size is 1.25 ms, then there are 16 time slots in one frame (20 / 1.25 = 16).

[0317] The time slot subdivision can be performed in a filter bank (e.g., 214), as discussed below.

[0318] In the example, the filter bank is a complex modulated low-delay filter bank (CLDFB), the frame size is 20 ms, the slot size is 1.25 ms, resulting in 16 filter banks per frame and the number of frequency bands per slot depending on the input sampling frequency, and where the frequency bands have a width of 400 Hertz (Hz). Thus, for example, for an input sampling frequency of 48 kilohertz (kHz), the length of a frame in samples is 960, the slot length is 60 samples, and the number of filter bank samples per slot is also 60.

[0319]

[0320] Even though each frame (and each slot) can be encoded in the time domain, a per-band analysis can be performed. In the example, multiple frequency bands are analyzed for each frame (or slot). For example, the filter bank can be applied to the time signal and the resulting sub-band signals can be analyzed. In some examples, the channel levels and associated information 220 are also provided in a per-band manner. For example, for each frequency band of the input signal 212 or the downmixed signal 246, the associated channel levels and associated information 220 (e.g., C y or ICC matrix) can be provided. In some examples, the number of frequency bands can be modified based on the properties of the signal and / or the properties of the requested bit rate, or the properties measured on the current payload. In some examples, the more slots required, the fewer frequency bands used to maintain a similar bit rate.

[0321] Since the slot size is smaller than the frame size (in terms of time length), in the case where a transient in the original signal 212 is detected within a frame, the slots can be used opportunely: the encoder (especially the filter bank 214) can identify the presence of the transient, signal its presence in the bitstream, and indicate in the side information 228 of the bitstream 248 in which slot of the frame the transient has occurred. Additionally, the parameters of the channel levels and associated information 220 encoded in the side information 228 of the bitstream 248 can thus be associated only with the slots following the transient and / or the slot in which the transient has occurred. Thus, the decoder will determine the presence of the transient and will associate the channel levels and associated information 220 only with the slots following the transient and / or the slot in which the transient has occurred (for the slots prior to the transient, the decoder will use the channel levels and associated information 220 of the previous frame). In Figure 10a none, no transient has occurred, and the parameters 220 encoded in the side information 228 can thus be understood as being associated with the entire frame 920. In Figure 10bIn [the figure], a transient has occurred at time slot 932: thus, the parameter 220 encoded in the side information 228 will refer to time slots 932, 933, and 934, while the parameter associated with time slot 931 will be assumed to be the same as the parameter of the frame before frame 930.

[0322] In view of the above, for each frame (or time slot) and each frequency band, specific channel levels and associated information 220 related to the original signal 212 can be defined. For example, the elements of the covariance matrix C y (such as covariance and / or level) can be estimated for each frequency band.

[0323] If the detection of a transient occurs while multiple frames are commonly associated with the same parameter, it is feasible to reduce the number of frames commonly associated with the same parameter, thereby increasing the mixing quality.

[0324] Figure 10a Frame 920 is shown (herein indicated as a "normal frame"), for which eight frequency bands are defined in the original signal 212 (the eight frequency bands 1...8 are shown on the vertical axis, while time slots 921 to 924 are shown on the horizontal axis). The parameters of the channel level and associated information 220 can theoretically be encoded in the side information 228 of the bitstream 248 in a per-frequency-band manner (for example, there will be one covariance matrix for each original frequency band). However, to reduce the amount of side information 228, the encoder can aggregate multiple original frequency bands (such as consecutive frequency bands) to obtain at least one aggregated band formed by the multiple original frequency bands. For example, in Figure 10a [the figure], the eight original frequency bands are grouped to obtain four aggregated bands (aggregated band 1 is associated with original frequency band 1; aggregated band 2 is associated with original frequency band 2; aggregated band 3 groups original frequency bands 3 and 5; aggregated band 4 groups original frequency bands 5...8). Matrices of covariance, correlation, ICC, etc. can be associated with each of the aggregated bands. In some examples, the parameter encoded in the side information 228 of the bitstream 248 is obtained from the sum (or average or another linear combination) of the parameters associated with each aggregated band. Thus, the size of the side information 228 of the bitstream 248 is further reduced. Hereinafter, "aggregated band" is also referred to as "parameter band" because it means those frequency bands used to determine the parameter 220.

[0325] Figure 10bShow the frame 931 in which the transient occurs (subdivided into four consecutive time slots 931 to 934, or another integer). Here, the transient occurs in the second time slot 932 ("transient slot"). In this case, the decoder can determine to direct the parameters of the channel level and related information 220 only to the transient time slot 932 and / or the subsequent time slots 933 and 934. The channel level and related information 220 of the previous time slot 931 will not be provided: It is understood that the channel level and related information of time slot 931 will be particularly different in principle from the channel level and related information of the time slot, but may be more similar to the channel level and related information of the frame before frame 930. Therefore, the decoder applies the channel level and related information of the frame before frame 930 to time slot 931, and the channel level and related information of frame 930 are only applied to time slots 932, 933, and 934.

[0326] Since the presence and location of the time slot 931 with the transient can be signaled in the side information 228 of the bitstream 248 (e.g., in 261, as shown later), a technique has been developed to avoid or reduce the increase in the size of the side information 228: The grouping between the aggregated bands can be changed: For example, aggregated band 1 groups the original bands 1 and 2, and aggregated band 2 groups the original bands 3…8. Therefore, relative to Figure 10a the situation, the number of bands is further reduced, and parameters will only be provided for two aggregated bands.

[0327] Figure 6a Show that the parameter estimation block (parameter estimator) 218 is capable of retrieving a certain number of channel levels and related information 220.

[0328] Figure 6a Show that the parameter estimator 218 is capable of retrieving a certain number of parameters (channel levels and related information 220), which can be Figures 9a to 9d the ICC of the matrix 900 of

[0329] However, in fact, only a part of the estimated parameters are submitted to the bitstream writer 226 for encoding the side information 228. This is because the encoder 200 can be configured to select (at a determination block 250 not shown in FIGS. 1 to 5) whether to encode at least a part of the channel levels and related information 220 of the original signal 212.

[0330] This is illustrated in Figure 6a as a plurality of switches 254s, which are controlled by the selection (command) 254 from the determination block 250. If each of the outputs 220 of the block parameter estimation 218 is Figure 9cFor the ICC of matrix 900, not all the parameters estimated by parameter estimation block 218 are actually encoded in side information 228 of bitstream 248: In particular, while terms 908 (ICC between channels: R and L; C and L; C and R; RS and CS) are actually encoded, term 907 is not encoded (i.e., determination block 250, which can be the same as Figure 6c can be considered to have a switch 254s that is open for unencoded term 907 but closed for term 908 to be encoded in side information 228 of bitstream 248. It should be noted that information 254’ (terms 908) regarding which parameters have been selected for encoding can be encoded (e.g., as a bitmap or other information regarding which terms 908 are encoded). In fact, information 254’ (which can be an ICC map for example) can include indices of encoded terms 908 (schematically shown in Figure 9d ). Information 254’ can be in the form of a bitmap: For example, information 254’ can be composed of fixed-length fields, each position associated with an index according to a predetermined ordering, and the value of each bit provides information on whether the parameter associated with that index is actually provided.

[0331] Generally, determination block 250 can, for example, select whether to encode at least a portion of channel levels and related information 220 (i.e., decide whether the terms of matrix 900 are to be encoded), for example, based on status information 252. Status information 252 can be based on payload status: For example, in the case of a highly loaded transmission, it will be possible to reduce the amount of side information 228 to be encoded in bitstream 248. For example, and with reference to Figure 9c :

[0332] In the case of a high payload, reduce the number of terms 908 of matrix 900 in side information 228 actually written to bitstream 248;

[0333] In the case of a lower payload, reduce the number of terms 908 of matrix 900 in side information 228 actually written to bitstream 248.

[0334] Alternatively or additionally, metric 252 can be evaluated to determine which parameters 220 are to be encoded in side information 228 (e.g., which terms of matrix 900 are designated as encoded terms 908 and which terms are to be discarded). In this case, perhaps only parameters 220 associated with more sensitive metrics (e.g., metrics associated with perceptually more important covariances) that are to be selected as encoded terms 908 are encoded in the bitstream.

[0335] It should be noted that this process can be repeated for each frame (or for multiple frames in the case of downsampling) and for each frequency band.

[0336] Thus, in addition to state metrics and the like, the determination block 250 can also be controlled by the parameter estimator 218 via Figure 6a the command 251 in

[0337] In some examples (e.g., Figure 6b ), the audio encoder can be further configured to encode the current channel level and associated information 220t in the bitstream 248 as an increment 220k relative to the previous channel level and associated information 220(t - 1). Thus, what the bitstream writer 226 encodes in the side information 228 can be the increment 220k associated with the current frame (or time slot) relative to the previous frame. This is shown in Figure 6b . The current channel level and associated information 220t are provided to the storage element 270 such that the storage element 270 stores the value of the current channel level and associated information 220t for subsequent frames. At the same time, the current channel level and associated information 220t can be compared with the previously obtained channel level and associated information 220(t - 1). (This is shown as the subtractor 273 in Figure 6b ). Thus, the subtraction result 220Δ can be obtained by the subtractor 273. The difference 220Δ can be used at the scaler 220s to obtain the relative increment 220k between the previous channel level and associated information 220(t - 1) and the current channel level and associated information 220t. For example, if the current channel level and associated information 220t are 10% greater than the previous channel level and associated information 220(t - 1), the increment 220 encoded by the bitstream writer 226 in the side information 228 will indicate information about a 10% increment. In some examples, instead of providing the relative increment 220k, the difference 220Δ can simply be encoded.

[0338] Among the parameters such as ICC and ICLD discussed above and below, the selection of the parameters to be actually encoded can be adapted to a particular situation. For example, in some examples:

[0339] For a first frame, only the Figure 9c ICC 908 is selected to be encoded in the side information 228 of the bitstream 248, while the ICC 907 is not encoded in the side information 228 of the bitstream 248;

[0340] For a second frame, a different ICC is selected to be encoded, while the different unselected ICCs are not encoded.

[0341] It may also be effective for time slots and frequency bands (and for different parameters, such as ICLD). Thus, the encoder (especially block 250) can determine which parameters are to be encoded and which are not, thus adapting the selection of the parameters to be encoded to a specific situation (e.g., state, selection, etc.). Therefore, the "feature for importance" can be analyzed to select which parameters are to be encoded and which are not. The feature for importance can be, for example, a measure associated with the result obtained in the simulation of the operations performed by the decoder. For example, the encoder can simulate the reconstruction of the uncoded covariance parameter 907 by the decoder, and the feature for importance can be a measure indicating the absolute error between the uncoded covariance parameter 907 and the same parameter presumably reconstructed by the decoder. By measuring the errors in different simulation scenarios (e.g., each simulation scenario is associated with the transmission of certain coded covariance parameters 908 and the measurement of the error affecting the reconstruction of the uncoded covariance parameter 907), it is feasible to determine the simulation scenario least affected by the error (e.g., the measure of all errors in the reconstruction in the simulation scenario) to distinguish the coded covariance parameter 908 to be encoded from the uncoded covariance parameter 907 based on the least affected simulation scenario. In the case of the least affected scenario, the unselected parameter 907 is the parameter most easily reconstructed, while the selected parameter 908 tends to be the parameter with the largest measure associated with the error.

[0342] The same can be done by simulating the reconstruction of the decoder or estimating the covariance, or by simulating the mixing characteristics or mixing results, rather than simulating parameters such as ICC and ICLD. It is worth noting that the simulation can be performed for each frame or each time slot, and can be performed for each frequency band or aggregated frequency bands.

[0343] An example can start with the parameters encoded in the side information 228 of the bitstream 248 and simulate the reconstruction of the covariance using formula (4) or (6) (see below).

[0344] More generally, it is feasible to reconstruct the channel levels and related information from the selected channel levels and related information, thus simulating the estimation of the unselected channel levels and related information (220, C y ) at the decoder (300), and calculating the error information between:

[0345] The unselected channel levels and related information (220) estimated by the encoder; and

[0346] The unselected channel levels and related information reconstructed by simulating the estimation of the uncoded channel levels and related information (220) at the decoder (300); and

[0347] To distinguish, based on the computed error information:

[0348] Channel levels and associated information that can be correctly reconstructed; and

[0349] Channel levels and associated information that cannot be correctly reconstructed,

[0350] To determine:

[0351] Select the channel levels and associated information that cannot be correctly reconstructed to be encoded in the side information (228) of the bitstream (248); and

[0352] Do not select the channel levels and associated information that can be correctly reconstructed, thereby avoiding encoding the channel levels and associated information that can be correctly reconstructed in the side information (228) of the bitstream (248).

[0353] Generally, the encoder can simulate any operation of the decoder and evaluate an error metric based on the simulation result.

[0354] In some examples, the importance metric can be different from (or can include other metrics that are different from) the evaluation of the metric associated with the error. In some cases, the feature of importance can be associated with a manual selection or be based on importance according to psychoacoustic criteria. For example, the most important channels can be selected for encoding (908) even without simulation.

[0355] Now, some additional discussion is provided to explain how the encoder signals which parameters 908 are actually encoded in the side information 220 of the bitstream 248.

[0356] Referring Figure 9d , the parameters on the diagonal of the ICC matrix 900 are associated with the ordered indices 1...10 (the order is predetermined and known to the decoder). In Figure 9c , it is shown that the selected parameters 908 to be encoded are the ICCs for L-R, L-C, R-C, LS-RS indexed by indices 1, 2, 5, 10 respectively. Thus, in the side information 228 of the bitstream 248, an indication of indices 1, 2, 5, 10 will also be provided (e.g., in Figure 6ain the side information 254'). Accordingly, by virtue of the information on indices 1, 2, 5, 10 provided by the encoder in the side information 228, the decoder will understand that the four ICCs provided in the side information 228 of the bitstream 248 are L-R, L-C, R-C, LS-RS. The indices can be provided, for example, by associating the position of each bit in the bitmap with a predefined position. For example, in order to signal the indices 1, 2, 5, 10, "1100100001" can be written (in field 254' of the side information 228), since the first, second, fifth, and tenth bits refer to the indices 1, 2, 5, 10 (other possibilities are at the disposal of the person skilled in the art). This is the so-called one-dimensional indexing, but other indexing strategies are also possible. For example, a combinatorial number technique according to which a number N is encoded (in field 254' of the side information 228), the number N being unambiguously associated with a specific channel pair (see also https: / / en.wikipedia.org / wiki / Combinatorial_number_system). When the bitmap points to ICCs, it can also be referred to as an ICC map.

[0357] It should be noted that, in some cases, a non-adaptive (fixed) parameter provision is used. This means that, in Figure 6a the example of, the selection 254 among the parameters to be encoded is fixed and there is no need to indicate in field 254' the selected parameters. Figure 9b Example showing a fixed parameter provision: The selected ICCs are L-C, L-LS, R-C, C-RS, and there is no need to signal their indices since the decoder already knows which ICCs are encoded in the side information 228 of the bitstream 248.

[0358] However, in some cases, the encoder can choose between a fixed parameter provision and an adaptive provision of the parameters. The encoder can signal said choice in the side information 228 of the bitstream 248 so that the decoder can know which parameters are actually encoded.

[0359] In some cases, at least some parameters can be provided without modification: for example:

[0360] ICDL can be encoded in any case without indicating them in the bitmap; and

[0361] ICCs may be subject to an adaptive provision.

[0362] The explanation involves each frame, or time slot, or frequency band. For subsequent frames, time slots or frequency bands, different parameters 908 are provided to the decoder, associating different indices with the subsequent frames, time slots or frequency bands; and different selections (e.g., fixed and adaptive) can be made. Figure 5 Shows an example of the filter bank 214 of the encoder 200, which can be used to process the original signal 212 to obtain the frequency-domain signal 216. From Figure 5 It can be seen that the time-domain (TD) signal 212 can be analyzed by the transient analysis block 258 (transient detector). In addition, the filter 263 (which can implement, for example, a Fourier filter, a short Fourier filter, an orthogonal mirror, etc.) provides the conversion of the frequency-domain (FD) version 264 of the input signal 212 in multiple frequency bands. The frequency-domain version 264 of the input signal 212 can be analyzed, for example, at the frequency band analysis block 267, which can determine (command 268) the specific frequency band group to be executed at the partition grouping block 265. Thereafter, the FD signal 216 will be a signal with a reduced number of aggregated frequency bands. The aggregation of frequency bands has been described above with respect to Figure 10a and Figure 10b The partition grouping block 267 can also be adjusted by the transient analysis performed by the transient analysis block 258. As described above, in the case of a transient, it is possible to further reduce the number of aggregated frequency bands: thus, the information 260 about the transient can adjust the partition grouping. Additionally or alternatively, the information 261 about the transient is encoded in the side information 228 of the bitstream 248. When the information 261 is encoded in the side information 228, the information 261 can include, for example, a flag (such as: "1", meaning "there is a transient in the frame", and "0" meaning: "there is no transient in the frame") indicating whether a transient has occurred and / or an indication of the position of the transient in the frame (such as a field indicating in which time slot the transient has been observed). In some examples, when the information 261 indicates that there is no transient ("0") in the frame, no indication of the position of the transient is encoded in the side information 228 to reduce the size of the bitstream 248. The information 261 is also referred to as "transient parameter", and as Figure 2d and 6b shown, is encoded in the side information 228 of the bitstream 246.

[0363] In some examples, the partition grouping at the block 265 can also be adjusted by external information 260', such as information about the state of the transmission (e.g., measurements associated with the transmission, error rate, etc.). For example, the higher the payload (or the greater the error rate), the greater the aggregation (the tendency is that fewer aggregated frequency bands are wider), so that a smaller amount of side information 228 has to be encoded in the bitstream 248. In some examples, the information 260' can be similar to Figure 6a the information or metric 252.

[0364] It is not feasible to send parameters for each frequency band / slot combination typically. However, the filter bank samples are grouped both over multiple slots and over multiple frequency bands to reduce the number of parameter sets sent per frame. Along the frequency axis, grouping the frequency bands into parameter bands uses a non-constant partitioning in the parameter bands, where the number of frequency bands in a parameter band is not constant but attempts to follow a psychoacoustically motivated parameter band resolution, i.e., at lower frequency bands, a parameter band includes only one or a few filter bank bands, and for higher parameter bands, a larger (and steadily increasing) number of filter bank bands are grouped into one parameter band.

[0365] Thus, for example, for a case where the input sampling rate is 48 kHz and the number of parameter bands is set to 14, the following vector grp 14 describes the filter bank indices that give the band boundaries for the parameter bands (indexing starts from 0):

[0366] grp 14 = [0, 1, 2, 3, 4, 5, 6, 8, 10, 13, 16, 20, 28, 40, 60]

[0367] Parameter band j includes the filter bank bands [grp 14 [j], grp 14 [j + 1]

[0368] Note that by simply truncating the bands, the bands grouped at 48 kHz can also be directly used for other possible sampling rates, since the grouping follows a psychoacoustically motivated frequency scale and has certain band boundaries corresponding to the number of bands for each sampling frequency (Table 1).

[0369] If the frame is non-transient or no transient processing is implemented, the grouping along the time axis will traverse all the slots in the frame so that one parameter set is available per parameter band.

[0370] Nevertheless, the number of parameter sets is still large, but the temporal resolution can be lower than a 20ms frame (40ms on average). Therefore, in order to further reduce the number of parameter sets sent per frame, only a subset of the parameter bands is used to determine and encode the parameters for sending to the decoder in the bitstream. The subsets are fixed and known to both the encoder and the decoder. The specific subset sent in the bitstream is signaled by a field in the bitstream to indicate to the decoder which subset of parameter bands the transmitted parameters belong to, and the decoder then replaces the parameters for this subset with the transmitted parameters (ICC, ICLD) and keeps the parameters (ICC, ICLD) from the previous frame for all parameter bands that are not in the current subset.

[0371] In an example, the parameter bands may be divided into two subsets that include roughly half of the total parameter bands and a contiguous subset for the lower parameter band and one contiguous subset for the upper parameter band. Since we have two subsets, the bitstream field used to signal the subsets is a single bit, and an example of a subset for 48kHz and 14 parameter bands is:

[0372] s 14 =[1,1,1,1,1,1,1,0,0,0,0,0,0,0]

[0373] where s 14 [j] indicates which subset of parameter band j it belongs to.

[0374] It is to be noted that the downmix signal 246 may be a signal in the time domain actually encoded in the bitstream 248: simply, the subsequent parameter estimator 218 will estimate the parameters 220 (eg ξ i,j and / or x i ) (and the decoder 300 will use the parameters 220 for preparing a mixing rule (eg, a mixing matrix) 403, which will be explained below.

[0375] Figure 2d An example of an encoder 200 is shown, which may be one of the encoders described above or may include elements of the encoders discussed previously. A TD input signal 212 is input to the encoder, and a bitstream 248 is output, which includes a downmix signal 246 (e.g., encoded by a core encoder 247) and correlation and level information 220 encoded in side information 228.

[0376] from Figure 2d It can be seen that the filter bank 214 (in Figure 5Examples of filter banks are provided in). In block 263, a frequency domain (FD) conversion (frequency domain DMX) is provided to obtain an FD signal 264, which is the FD version of the input signal 212. The FD signal 264 (also denoted as X) in multiple frequency bands is obtained. A frequency band / time slot grouping block 265 (which can be implemented as Figure 5 The grouping block 265) can be provided to obtain the FD signal 216 in the aggregated frequency band. In some examples, the FD signal 216 can be a version of the FD signal 264 in fewer frequency bands. Next, the signal 216 can be provided to the parameter estimator 218, which includes covariance estimation blocks 502, 504 (shown here as a single block), and downstream parameter estimation and coding blocks 506, 510 (embodiments of elements 502, 504, 506, and 510 are shown in Figure 6c ). The parameter estimation and coding blocks 506, 510 can also provide the parameter 220 to be encoded in the side information 228 of the bitstream 248. The transient detector 258 (which can be implemented as Figure 5 The transient analysis block 258) can find the transient and / or the position of the transient within a frame (e.g., in which time slot the transient has been identified). Thus, the information 261 about the transient (e.g., transient parameters) can be provided to the parameter estimator 218 (e.g., to determine which parameters are to be encoded). The transient detector 258 can also provide information or a command (268) to the block 265 to perform grouping by considering the presence and / or position of the transient in the frame.

[0377] Figure 3a , FIGS. 3b, 3c show examples of an audio decoder 300 (also referred to as an audio synthesizer). In the example, Figure 3a , the decoders in FIGS. 3b, 3c can be the same decoder, with some differences just to avoid different elements. In the example, the decoder 300 can be the same as the decoder in FIGS. 1 and 4. In the example, the decoder 300 can also be the same device as the encoder 200.

[0378] The decoder 300 can be configured to generate a synthesized signal (336, 340, y) from the downmixed signal x in TD (246) or FD (314) R)。The audio synthesizer 300 may include an input interface 312 configured to receive a downmixed signal 246 (e.g., the same downmixed signal encoded by the encoder 200) and side information 228 (e.g., encoded in the bitstream 248). As explained above, the side information 228 may include the channel levels and correlation information (220, 314) of the original signal (which may be the original input signal 212, y on the encoder side), such as ξ, χ, etc. or at least one of its elements (as will be explained below). In certain examples, all ICLDs (χ) outside the diagonal of the ICC matrix 900 and some (but not all) of the terms 906 or 908 (ICC or ξ values) are obtained by the decoder 300.

[0379] The decoder 300 may be configured (e.g., by a prototype signal calculator or prototype signal calculation module 326) to calculate a prototype signal 328 from the downmixed signal (324, 246, x), the prototype signal 328 having a plurality of channels (more than one) of the synthesized signal 336.

[0380] The decoder 300 may be configured (e.g., by a mixing rule calculator 402) to calculate a mixing rule 403 using at least one of the following:

[0381] The channel levels and correlation information of the original signal (212, y) (e.g., 314, C y , ξ, χ or its elements); and

[0382] The covariance information associated with the downmixed signal (324, 246, x) (e.g., C x or its elements).

[0383] The decoder 300 may include a synthesis processor 404 configured to use the prototype signal 328 and the mixing rule 403 to generate a synthesized signal (336, 340, y R ).

[0384] The synthesis processor 404 and the mixing rule calculator 402 may be incorporated in a synthesis engine 334. In certain examples, the mixing rule calculator 402 may be external to the synthesis engine 334. In certain examples, Figure 3a the mixing rule calculator 402 of may be integrated with the parameter reconstruction module 316 of FIG. 3b.

[0385] The synthesized signal (336, 340, y R) The number of synthesized channels is greater than 1 (in some cases greater than 2 or greater than 3), and can be greater than, less than, or equal to the number of original channels of the original signal (212, y), where the number of original channels is also greater than 1 (in some cases greater than 2 or greater than 3). The number of channels of the downmixed signal (246, 216, x) is at least one or two, and less than the number of original channels of the original signal (212, y) and the number of synthesized channels of the synthesized signal (336, 340, y R ) of the synthesized channels.

[0386] The input interface 312 can read the encoded bitstream 248 (e.g., the same bitstream 248 encoded by the encoder 200). The input interface 312 can be or include a bitstream reader and / or an entropy decoder. As described above, the bitstream 248 can encode the downmixed signal (246, x) and the side information 228 as described above. The side information 228 can include, for example, the original channel levels and associated information 220, in a form output by the parameter estimator 218 or any element downstream of the parameter estimator 218 (e.g., the parameter quantization block 222, etc.). The side information 228 can include encoded values or indexed values or both. Even though the input interface 312 is not shown for the downmixed signal (346, x) in FIG. 3b, the input interface 312 can be applied to the downmixed signal as Figure 3a shown. In some examples, the input interface 312 can quantize the parameters obtained from the bitstream 248.

[0387] Thus, the decoder 300 can obtain the downmixed signal (246, x), and the downmixed signal (246, x) can be in the time domain. As described above, the downmixed signal 246 can be partitioned into frames and / or time slots (see above). In an example, the filter bank 320 can transform the downmixed signal 246 in the time domain to obtain a version 324 of the downmixed signal 246 in the frequency domain. As described above, the frequency bands of the frequency domain version 324 of the downmixed signal 246 can be grouped into band groups. In an example, the same grouping as performed at the filter bank 214 (see above) can be carried out. The parameters for grouping (e.g., which frequency bands and / or how many frequency bands are to be grouped...) can be based, for example, on signaling from the partitioner / grouping unit 265 or the band analysis block 267, which is encoded in the side information 228.

[0388] The decoder 300 may include a prototype signal calculator 326. The prototype signal calculator 326 may calculate a prototype signal 328 from a downmixed signal (such as one of versions 324, 246, x), for example by applying a prototype rule (such as matrix Q). The prototype rule may be implemented by a prototype matrix (Q) having a first dimension and a second dimension, where the first dimension is associated with the number of downmixed channels and the second dimension is associated with the number of synthesized channels. Thus, the prototype signal has a plurality of channels of the synthesized signal 340 to be ultimately produced.

[0389] The prototype signal calculator 326 may apply a so-called upmix to the downmixed signal (324, 246, x) as it simply produces a version of the downmixed signal (324, 246, x) with an increased number of channels (the number of channels of the synthesized signal to be produced) without applying excessive "intelligence". In an example, the prototype signal calculator 326 may simply apply a fixed predetermined prototype matrix (identified in this document as "Q") to the FD version 324 of the downmixed signal 246. In an example, the prototype signal calculator 326 may apply different prototype matrices to different frequency bands. For example, based on a specific number of downmixed channels and a specific number of synthesized channels, a prototype rule (Q) may be selected from a plurality of pre-stored prototype rules.

[0390] The prototype signal 328 may be decorrelated at the decorrelation module 330 to obtain a decorrelated version 332 of the prototype signal 328. However, in some examples, advantageously, the decorrelation module 330 is absent because the present invention has proven to be sufficiently effective to allow its avoidance.

[0391] The prototype signal (in any of its versions 328, 332) may be input to the synthesis engine 334 (and in particular the synthesis processor 404). Here, the prototype signal (328, 332) is processed to obtain a synthesized signal (336, y R )). The synthesis engine 334 (and in particular the synthesis processor 404) may apply a mixing rule 403 (in some examples, discussed below, there are two mixing rules, for example one for the main component of the synthesized signal and one for the residual component). The mixing rule 403 may be implemented, for example, by a matrix. The matrix 403 may be produced, for example, by the mixing rule calculator 402 based on the channel levels and correlation information (314, such as ξ, χ or its elements) of the original signal (212, y).

[0392] The synthesized signal 336 output by the synthesis engine 334 (specifically by the synthesis processor 404) may optionally be filtered at the filter bank 338. Additionally or alternatively, the synthesized signal 336 may be converted to the time domain at the filter bank 338. Thus, a version 340 of the synthesized signal 336 (in the time domain or after filtering) can be used for audio reproduction (e.g., via a speaker).

[0393] To obtain the mixing rules (e.g., mixing matrix) 403, the channel levels and associated information of the original signals (e.g., C y , etc.) and the covariance information associated with the downmixed signals (e.g., C x ) can be provided to the mixing rule calculator 402. For this purpose, it is feasible to encode the channel levels and associated information 220 in the side information 228 using the encoder 200.

[0394] However, in some cases, to reduce the amount of information encoded in the bitstream 248, not all parameters are encoded by the encoder 200 (e.g., not the entire channel levels and associated information of the original signal 212 and / or not the entire covariance information of the downmixed signal 246). Therefore, some parameters 318 will be estimated at the parameter reconstruction module 316.

[0395] The parameter reconstruction module 316 can be fed, for example, at least one of the following:

[0396] A copy 322 of the downmixed signal 246(x), which can be, for example, a filtered version or FD version of the downmixed signal 246; and

[0397] The side information 228 (including the channel levels and associated information 228).

[0398] The side information 228 may include information associated with the correlation matrix C y associated with the original signals (212, y) (as the levels and associated information of the input signals): However, in some cases, not all elements of the correlation matrix C y are actually encoded. Therefore, estimation and reconstruction techniques have been developed to reconstruct a version y of the correlation matrix C (e.g., through an intermediate step of obtaining an estimated version ).

[0399] The parameters 314 provided to the module 316 can be obtained by the entropy decoder 312 (input interface) and can be quantized, for example.

[0400] Figure 3cAn example of decoder 300 is shown, and the decoder can be an embodiment of one of the decoders in FIGS. 1 to 3b. Here, decoder 300 includes an input interface 312 represented by a demultiplexer. Decoder 300 outputs a composite signal 340, which can be, for example, in TD (signal 340) to be played back by a speaker or in FD (signal 336). Figure 3c Decoder 300 can include a core decoder 347, and core decoder 347 can also be part of input interface 312. Core decoder 347 can thus provide a downmixed signal x, 246. Filter bank 320 can convert the downmixed signal 246 from TD to FD. The FD version of the downmixed signal x, 246 is indicated at 324. The FD downmixed signal 324 can be provided to covariance synthesis block 388. Covariance synthesis block 388 can provide a composite signal 336 (Y) in FD. Inverse filter bank 338 can convert the audio signal 314 in its TD version 340. The FD downmixed signal 324 can be provided to band / time slot grouping block 380. Band / time slot grouping block 380 can perform the same operations as those already performed by the Figure 5 and 2d partitioning and grouping block 265 in the encoder. In the encoder, as Figure 5 and 2d the bands of the downmixed signal 216 have been grouped or aggregated in a few bands (with a wider width), and the parameters 220 (ICC, ICLD) have been associated with the aggregated band groups, it is now necessary to aggregate the decoded downmixed signal in the same way and associate each aggregated band with the relevant parameters. Thus, reference numeral 385 means the downmixed signal X that has been aggregated B . It should be noted that the filter provides an unaggregated FD representation so that the bands / time slots can be grouped in the decoder (380) in the same way as in the encoder to process the parameters, perform the same aggregation on the bands / time slots as in the encoder, and provide the aggregated downmixed X B .

[0401] Band / time slot grouping block 380 can also aggregate over different time slots in a frame, such that signal 385 is also aggregated in a time slot size similar to that of the encoder. Band / time slot grouping block 380 can also receive information 261 encoded in side information 228 in bitstream 248, where information 261 indicates the presence of a transient and, optionally, also indicates the position of the transient within the frame.

[0402] At covariance estimation block 384, the covariance C of the downmixed signal 246 (324) is estimated x . At covariance calculation block 386, the covariance C y is obtained, for example, by using formulas (4) to (8) which can be used for this purpose. Figure 3cShows "multichannel parameters", which can be, for example, parameter 220 (ICC and ICLD). Then covariance C y and C x are provided to covariance synthesis block 388 to synthesize synthesized signal 388. In some examples, when blocks 384, 386, and 388 are implemented together, both parameter reconstruction 316 and mixing will be calculated 402, and synthesis processor 404 will be as discussed above and below.

[0403] 4 Discussion

[0404] 4.1 Overview

[0405] The novel method of this example is particularly aimed at encoding and decoding multichannel content at a low bitrate (meaning equal to or lower than 160 kbits / sec), while maintaining the sound quality as close as possible to the original signal and preserving the spatial characteristics of the multichannel signal. One function of the novel method is also to be suitable for the aforementioned DirAC framework. The output signal can be rendered on the same speaker setup as input 212, or on a different speaker setup (which can be larger or smaller in terms of speakers). Similarly, the output signal can be rendered on speakers using binaural rendering.

[0406] The current section will provide an in-depth description of the present invention and the different modules that make up the present invention.

[0407] The proposed system consists of two main parts:

[0408] 1. Encoder 200, which derives the necessary parameters 220 from input signal 212, quantizes them (at 222) and encodes them (at 226). Encoder 200 can also calculate the downmixed signal 246 to be encoded in bitstream 248 (and can be sent to decoder 300).

[0409] 2. Decoder 300, which uses the encoded (e.g., sent) parameters and downmixed signal 246 to produce a multichannel output with a quality as close as possible to the original signal 212.

[0410] Figure 1 shows an overview of the novel method proposed according to an example. Note that some examples will only use a subset of the building blocks shown in the overall drawings and discard some processing blocks depending on the application scenario.

[0411] The input 212 (y) of the present invention is a multichannel audio signal 212 (also referred to as "multichannel stream") in the time domain or time-frequency domain (e.g., signal 216), for example, a set of audio signals generated by a set of speakers or intended to be played.

[0412] The first part of the processing is the encoding part; from the multichannel audio signal, the so-called "downmix" signal 246 (see 4.2.6) will be calculated together with a parameter set or side information 228 (see 4.2.2 and 4.2.3 ), which is derived from the input signal 212 in the time domain or the frequency domain. These parameters will be encoded (see 4.2.5) and, if appropriate, sent to the decoder 300.

[0413] Then the downmix signal 246 and the encoded parameters 228 can be sent to the core encoder and the transmission canal, which links the encoder side and the decoder side of the processing.

[0414] On the decoder side, the downmix signal is processed (4.3.3 and 4.3.4), and the transmitted parameters are decoded (see 4.3.2). The decoded parameters will be used to synthesize the output signal using covariance synthesis (see 4.3.5), which will result in the final multichannel output signal in the time domain.

[0415] Before going into details, some general characteristics need to be established, and at least one of these general characteristics is valid:

[0416] - The processing can be used with any speaker setup. Keep in mind that when increasing the number of speakers, the complexity of the processing and the bits required to encode the transmitted parameters will also increase.

[0417] - The whole processing can be done on a frame basis, i.e., the input signal 212 can be divided into frames that are processed independently. On the encoder side, each frame will generate a parameter set, and these parameters will be transmitted to the decoder side for processing.

[0418] - A frame can also be divided into time slots; these time slots then exhibit statistical properties that cannot be obtained at the frame scale. A frame can be divided into, for example, eight time slots, and the length of each time slot will be equal to 1 / 8 of the frame length.

[0419] 4.2 Encoder

[0420] The purpose of the encoder is to extract appropriate parameters 220 to describe the multichannel signal 212, quantize them (at 222), encode them as side information 228 (at 226), and then, if appropriate, send them to the decoder side. Here, the parameters 220 and how to calculate them will be described in detail.

[0421] A more detailed scheme of the encoder 200 can be found in Figures 2a to 2d This overview highlights the two main outputs 228 and 246 of the encoder.

[0422] The first output of the encoder 200 is the downmixed signal 228 calculated from the multichannel audio input 212; the downmixed signal 228 is a representation of the original multichannel stream (signal) on fewer channels than the original content (212). For more information on its calculation, see Section 4.2.6.

[0423] The second output of the encoder 200 is the encoded parameters 220 represented as side information 228 in the bitstream 248; these parameters 220 are the key points of this example: they are the parameters that will be used to effectively describe the multichannel signal on the decoder side. These parameters 220 provide a good trade-off between the quality and the number of bits required to encode them in the bitstream 248. On the encoder side, the parameter calculation can be done in several steps; the process will be described in the frequency domain, but it can also be done in the time domain. The parameters 220 are first estimated from the multichannel input signal 212, then they are quantized at the quantizer 222, and then they can be converted into a digital bitstream 248 as side information 228. For more information on these steps, see Sections 4.2.2, 4.2.3 and 4.2.5.

[0424] 4.2.1 Filter Bank and Partitioning into Groups

[0425] The filter banks are discussed for the encoder side (e.g., filter bank 214) or the decoder side (e.g., filter banks 320 and / or 338).

[0426] The present invention can use filter banks at various points during processing. These filter banks can transform the signal from the time domain to the frequency domain (so-called aggregated frequency bands or parametric bands), in which case they are called "analysis filter banks", or from the frequency domain to the time domain (e.g., 338), in which case they are called "synthesis filter banks".

[0427] The choice of the filter bank must meet the required performance and optimization requirements, but the rest of the processing can be carried out independently of the specifically chosen filter bank. For example, a filter bank based on quadrature mirror filters or a filter bank based on the short-time Fourier transform is used.

[0428] Refer to Figure 5 , the output of the filter bank 214 of the encoder 200 will be the signal 216 in the frequency domain represented on a certain number of frequency bands (266 as opposed to 264). Performing the rest of the processing for all the frequency bands (264) can be understood as providing better quality and better frequency resolution, but it also requires a more significant bit rate to transmit all the information. Therefore, a so-called "partitioning and grouping" (265) is performed together with the filter bank processing, which corresponds to grouping certain frequencies together in order to represent the information 266 on smaller groups of frequency bands.

[0429] For example, the output 264 of filter 263 ( Figure 5 ) can be represented on 128 frequency bands, and the partitioning and grouping at 265 can result in signal 266 (216) having only 20 frequency bands. There are several ways to group the frequency bands together, and a meaningful way could be, for example, to try to approximate an equivalent rectangular bandwidth. The equivalent rectangular bandwidth is a frequency band partitioning of psychoacoustic excitation that attempts to model how the human auditory system processes audio events, i.e., the aim is to group the filter banks in a way that is suitable for human hearing.

[0430] 4.2.2 Parameter Estimation (e.g., Estimator 218)

[0431] Aspect 1: Describing and synthesizing multichannel content using covariance matrices

[0432] The parameter estimation at 218 is one of the key points of the present invention; they are used on the decoder side to synthesize the output multichannel audio signal. Those parameters 220 (encoded as side information 228) have been chosen because they effectively describe the multichannel input stream (signal) 212 and they do not require the transmission of a large amount of data. These parameters 220 are calculated on the encoder side and are later used jointly with the synthesis engine on the decoder side to calculate the output signal.

[0433] Here, a covariance matrix can be calculated between the channels of the multichannel audio signal and the downmixed signal. That is:

[0434] - C y : The covariance matrix of the multichannel stream (signal), and / or

[0435] - C x : The covariance matrix of the downmixed stream (signal) 246

[0436] The processing can be carried out on the basis of parameter bands, so that one parameter band is independent of another parameter band, and the formula for describing a given parameter band can be given without loss of generality.

[0437] For a given parameter band, the covariance matrix is defined as follows:

[0438]

[0439] where

[0440] - denotes the real part operator.

[0441] - Instead of the real part, it can be any other operation that produces a real value that is related to the complex value from which it is derived (e.g., the absolute value).

[0442] - * represents the conjugate transpose operator.

[0443] - B represents the relationship between the original multiple frequency bands and the grouped frequency bands (see 4.2.1 for partition grouping).

[0444] - Y and X are the original multichannel signal 212 and the downmixed signal 246 in the frequency domain, respectively.

[0445] C y (or its elements, or values obtained from C y or from its elements) is also indicated as the channel levels and related information of the original signal 212. C x (or its elements, or values obtained from C y or from its elements) is also indicated as the covariance information associated with the downmixed signal 212.

[0446] For a given frame (and frequency band), only one or two covariance matrices C y and / or C x , can be output, for example, by the estimator block 218. The process is slot-based rather than frame-based, and different implementations can be adopted regarding the relationship between a given slot and the matrix for the entire frame. As an example, the covariance matrices can be calculated for each slot within a frame and summed to output the matrix for one frame. Note that the definitions for calculating the covariance matrices are mathematical definitions, but it is also feasible to calculate or at least modify those matrices in advance if an output signal with specific characteristics is desired.

[0447] As described above, not all elements of the matrices C y and / or C x need to be actually encoded in the side information 228 of the bitstream 248. For C x , it is feasible to simply estimate it from the downmixed signal 246 encoded by applying formula (1), and thus the encoder 200 can easily avoid encoding any elements of C x (or more generally, the covariance information associated with the downmixed signal). For C y (or for the channel levels and related information associated with the original signal), it is feasible to estimate at least one of the elements of C y using the techniques discussed below at the decoder side.

[0448] Aspect 2a: Transmitting covariance matrices and / or energy to describe and reconstruct multichannel audio signals

[0449] As mentioned before, covariance matrices are used for synthesis. It is feasible to directly transmit those covariance matrices (or a subset thereof) from the encoder to the decoder.

[0450] In some examples, matrix C x does not necessarily have to be transmitted, since the matrix can be recalculated at the decoder side using the downmixed signal 246, but depending on the application scenario, this matrix may be required as a transmitted parameter.

[0451] From an implementation perspective, not all values in those matrices C x , C y have to be encoded or transmitted, for example, to meet certain specific requirements regarding the bit rate. The values that are not transmitted can be estimated at the decoder side (see 4.3.2).

[0452] Aspect 2b: Transmitting inter-channel coherence and inter-channel level differences to describe and reconstruct a multichannel signal

[0453] A set of alternative parameters can be defined from the covariance matrices C x , C y and used to reconstruct the multichannel signal 212 at the decoder side. These parameters can be, for example, inter-channel coherence (ICC) and / or inter-channel level difference (ICLD).

[0454] Inter-channel coherence describes the coherence between each pair of channels in a multichannel stream. The parameter can be derived from the covariance matrix C y and calculated as follows (for a given parameter band and for two given channels i and j):

[0455]

[0456] where

[0457] - ξ i,j is the ICC between channels i and j of the input signal 212

[0458] - C yi,j is the value in the covariance matrix of the multichannel signal between channels i and j of the input signal 212 - previously defined in formula (1) -

[0459] ICC values can be calculated between each pair of channels of the multichannel signal, which may result in a large amount of data as the size of the multichannel signal grows. In practice, a reduced set of ICCs can be encoded and / or transmitted. In some examples, the values to be encoded and / or transmitted must be defined according to performance requirements.

[0460] As an example, when processing a signal generated by a 5.1 (or 5.0) speaker setup as defined by the speaker setup defined in the ITU recommendation "ITU-R BS.2159-4", it is feasible to select to transmit only four ICCs. These four ICCs can be one of the following:

[0461] - Center and right channels

[0462] - Center and left channels

[0463] - Left and left surround channels

[0464] - Right and right surround channels

[0465] Typically, the indices of the ICCs selected from the ICC matrix are described by an ICC map.

[0466] Typically, for each speaker setup, a fixed set of ICCs that averages to the best quality can be selected to be encoded and / or transmitted to the decoder. The number of ICCs and which ICCs are to be transmitted can depend on the speaker setup and / or the total available bitrate, and are available both at the encoder and the decoder without transmitting the ICC map in the bitstream 248. In other words, a fixed set of ICCs and / or the corresponding fixed ICC map can be used, e.g., depending on the speaker setup and / or the total bitrate.

[0467] This fixed set may not be suitable for a particular material, and in some cases, using the fixed set of ICCs results in a quality that is significantly worse than the average quality for all materials. To overcome this for each frame (or time slot) in another example, the optimal set of ICCs and the corresponding ICC map can be estimated based on the importance characteristics of a certain ICC. Then, the ICC map for the current frame is explicitly encoded and / or transmitted in the bitstream 248 together with the quantized ICCs.

[0468] For example, similar to a decoder using formulas (4) and (6) from 4.3.2, the downmix covariance C from formula (1) can be used x to generate an estimate of the covariance or an estimate of the ICC matrix to determine the importance characteristics of the ICCs. Depending on the selected characteristics, for each ICC or the corresponding entry in the covariance matrix, for each frequency band, for which parameters will be sent in the current frame and combined for all frequency bands, the characteristics are calculated. Then, the combined characteristic matrix is used to determine the most important ICCs, thus determining the set of ICCs to be used and the ICC map to be sent.

[0469] For example, the importance characteristic of an ICC is the difference between the estimated covariance and the actual covariance C ythe absolute error between terms, and the combined feature matrix is the sum of the absolute errors of each ICC to be transmitted over all frequency bands in the current frame. From the combined feature matrix, n terms are selected, where the summed absolute error is the highest, and n is the number of ICCs to be transmitted for the speaker / bitrate combination, and an ICC map is constructed from these terms.

[0470] In addition, in another example as Figure 6b shown, to avoid the ICC map changing too much between frames, the feature matrix can be emphasized for each term in the selected ICC map of the previous parametric frame. For example, in the case of the absolute error of covariance, by applying a coefficient > 1 (220k) to the terms of the ICC map of the previous frame.

[0471] In addition, in another example, a flag transmitted in the side information 228 of the bitstream 248 can indicate whether a fixed ICC map or the best ICC map is used in the current frame, and if the flag indicates a fixed group, the ICC map is not transmitted in the bitstream 248.

[0472] The best ICC map is, for example, encoded and / or sent as a bitmap (e.g., the ICC map can implement Figure 6a the information 254’).

[0473] Another example for transmitting the ICC map is to transmit an index into a table of all possible ICC maps, where the index itself is, for example, additionally entropy - encoded. For example, the table of all possible ICC maps is not stored in memory, but the ICC map indicated by the index is directly calculated from the index.

[0474] A second parameter that can be transmitted jointly (or separately) with the ICC is the ICLD. “ICLD” stands for inter - channel level difference, and it describes the energy relationship between each channel of the input multichannel signal 212. There is no unique definition of the ICLD; the important aspect of this value is that it describes the energy ratio within the multichannel stream.

[0475] As an example, the conversion from C y to ICLD can be obtained as follows:

[0476]

[0477] where:

[0478] -χ i ICLD for channel i.

[0479] -P i The power of the current channel i, which can be obtained from the diagonal of C y : P i = C yi,iExtraction

[0480] -P dmx,i Depends on channel i, but will always be a linear combination of the values in C x and also depends on the original speaker setup.

[0481] In the example, P dmx,i is not the same for every channel, but depends on the mapping related to the downmix matrix (which is also the prototype matrix for the decoder), which is typically mentioned in one of the points under formula (3). Depending on whether only channel i is downmixed to one of the downmix channels or to more than one of them. In other words, in the presence of non-zero elements in the downmix matrix, P dmx,i may be or include the sum of all diagonal elements of C x so formula (3) can be rewritten as:

[0482]

[0483]

[0484] P i = C yi,i

[0485] where α i is a weighting factor related to the expected energy contribution of the channel for downmixing. This weighting factor is fixed for a specific input speaker configuration and is known at both the encoder and the decoder. The concept of matrix Q will be provided below. Some values of α i and matrix Q are also provided in the last part of the document.

[0486] In the case of an implementation that defines a mapping for each input channel i, where the mapping index is the downmix channel j to which input channel i is only mixed, or if the mapping index is greater than the number of downmix channels. So, we have a mapping index m ICLD,i which is used to determine P dmx,i as follows:

[0487]

[0488] 4.2.3 Parameter Quantization

[0489] To obtain the quantization parameter 224, an example of the quantization of parameter 220 can be performed, for example, by the quantization parameter module 222 of FIGS. 2b and 4.

[0490] Once the parameter set 220 is calculated, meaning the covariance matrices {C x , C y} or ICC and ICLD {ξ, χ}, which are quantized. The choice of quantizer can be a trade-off between quality and the amount of data to be transmitted, but there is no limitation on the quantizer to be used.

[0491] As an example, in the case of using ICC and ICLD; a non-linear quantizer for ICC with 10 quantization steps over the interval [-1, 1] can be provided, and another non-linear quantizer for ICLD with 20 quantization steps over the interval [-30, 30].

[0492] Likewise, as an implementation of an optimization scheme, it is feasible to downsample the parameters to be transmitted, meaning that the quantized parameter 224 is used by two or more frames in a row.

[0493] In one aspect, a subset of the parameters transmitted in the current frame is signaled by a parameter frame index in the bitstream.

[0494] 4.2.4 Transient Processing, Downsampling Parameters

[0495] As can be understood from certain examples discussed below, as shown in Figure 5 which can in turn be an example of block 214 of FIGS. 1 and 2d.

[0496] In the case of a downsampled parameter set (e.g., obtained at block 265 in Figure 5 ), i.e., the parameter set 220 for a subset of the parameter bands can be used for more than one processed frame, transients that appear in more than one subset cannot be preserved in terms of localization and coherence. Therefore, it may be advantageous to transmit the parameters of all bands in such frames. This special type of parameter frame can be signaled, for example, by a flag in the bitstream.

[0497] In one aspect, transient detection at 258 is used to detect such transients in the signal 212. The position of the transient in the current frame can also be detected. The time granularity can be advantageously linked to the time granularity of the filter bank 214 used, such that each transient position can correspond to a time slot or a set of time slots of the filter bank 214. Then, based on the transient position, the time slots for calculating the covariance matrices C y and C x are selected, for example, using only the time slots from the time slot including the transient to the end of the current frame.

[0498] The transient detector (or transient analysis block 258) can be the transient detector that is also used for encoding the downmixed signal 212, e.g., the time-domain transient detector of the IVAS core encoder. Therefore, Figure 5 the example of

[0499] In one example, a bit is used to encode the occurrence of a transient (such as: "1" meaning "there is a transient in the frame" and "0" meaning "there is no transient in the frame"). If a transient is detected, additionally the position of the transient is encoded and / or sent as an encoded field 261 (information about the transient) in the bitstream 248 to allow similar processing in the decoder 300.

[0500] If a transient is detected and transmission of all bands is signaled (for example), using the normal partitioned packet transmission parameters 220 may result in a spike in the data rate required for the side information 228 in the bitstream 248 for the transmission parameters 220. Additionally, time resolution is more important than frequency resolution. Therefore, at block 265, it may be advantageous to change the partitioned packet for such a frame to have fewer bands to transmit (for example, from many bands in the signal version 264 to fewer bands in the signal version 266). One example employs such a different partitioned packet, for example, by combining two adjacent bands across all bands with a normal downsampling factor of 2 for the parameters. Generally, the occurrence of a transient implies that the covariance matrix itself can be expected to be very different before and after the transient. To avoid artifacts in the time slots prior to the transient, only the transient time slot itself and all subsequent time slots until the end of the frame may be considered. This is also based on the assumption that the signal was stable enough beforehand and that it is possible to use the information and mixing rules that were derived for the previous frame and also apply to the time slots prior to the transient.

[0501] In general, the encoder can be configured to determine in which time slot of a frame a transient has occurred and encode the channel levels and associated information (220) of the original signal (212, y) associated with the time slot in which the transient has occurred and / or subsequent time slots in the frame, without encoding the channel levels and associated information (220) of the original signal (212, y) associated with the time slots prior to the transient.

[0502] Similarly, when the presence and position of a transient in a frame are signaled (261), the decoder can (for example, at block 380):

[0503] Associate the current channel levels and associated information (220) with the time slot in which the transient has occurred and / or subsequent time slots in the frame; and

[0504] Associate the time slots of the frame prior to the time slot in which the transient has occurred with the channel levels and associated information (220) of the previous time slots.

[0505] Another important aspect of a transient is that, in the case where it is determined that there is a transient in the current frame, no smoothing operation is performed on the current frame. In the case of a transient, no Cy and C x is smoothed, but C from the current frame yR and C x are used for the calculation of the mixing matrix.

[0506] 4.2.5 Entropy Coding

[0507] The entropy coding module (bitstream writer) 226 can be the module of the final encoder; its purpose is to convert the previously obtained quantized values into a binary bitstream, which will also be referred to as "side information".

[0508] The method used to encode the values can be, for example, Huffman coding [6] or delta coding. The coding method is not crucial and will only affect the final bit rate. A person should adapt the coding method depending on the bit rate he wants to achieve.

[0509] Several implementation optimization schemes can be executed to reduce the size of the bitstream 248. As an example, a switching mechanism can be implemented, which depends on which is more efficient from the perspective of the bitstream size to switch from one coding scheme to another.

[0510] For example, these parameters can be delta-coded along the frequency axis of a frame, and the resulting sequence of delta-index entropies is encoded by a range encoder.

[0511] Similarly, in the case of parameter downsampling, also as an example, a mechanism can be implemented to send only a subset of the parameter bands per frame in order to send data continuously.

[0512] These two examples require signaling bits to signal specific processing aspects to the decoder on the encoder side.

[0513] 4.2.6 Downmix Calculation

[0514] The downmixing part 244 of the processing can be simple, but in some examples it is crucial. The downmixing used in the present invention can be passive downmixing, which means that its calculation method remains the same during processing and is independent of the signal or its characteristics at a given time. However, it has been understood that the downmixing calculation at 244 can be extended to active downmixing calculations (such as those described in [7]).

[0515] The downmixed signal 246 can be calculated at two different positions:

[0516] - The first time is for parameter estimation on the encoder side (see 4.2.2 ), because it may be necessary (in some examples) to calculate the covariance matrix C x .

[0517] - Second, on the encoder side, between the encoder 200 and the decoder 300 (in the time domain), the downmixed signal 246 is encoded and / or transmitted to the decoder 300 and is used as the basis for synthesis at module 334.

[0518] As an example, for a 5.1 input stereo downmix, the downmixed signal can be calculated as follows:

[0519] - The downmixed left channel is the sum of the left channel, the left surround channel, and the center channel.

[0520] The downmixed right channel is the sum of the right channel, the right surround channel, and the center channel. Alternatively, in the case of a 5.1 input monophonic downmix, the downmixed signal is calculated as the sum of each channel in the multichannel stream.

[0521] In the example, each channel of the downmixed signal 246 can be obtained as a linear combination of the channels of the original signal 212, for example with constant parameters, thus achieving passive downmixing.

[0522] Depending on the processing requirements, the calculation of the downmixed signal can be extended and adapted to other speaker setups.

[0523] Aspect 3: Low-latency processing using passive downmixing and a low-latency filter bank

[0524] The present invention can provide low-latency processing by using passive downmixing such as that previously described for a 5.1 input and a low-latency filter bank. Using these two elements, it is possible to achieve a latency of less than 5 milliseconds between the encoder 200 and the decoder 300.

[0525] 4.3 Decoder

[0526] The purpose of the decoder is to synthesize an audio output signal (336, 340, y R ) using the encoded (e.g., transmitted) downmixed signal (246, 324) and the encoded side information 228, on a given speaker setup. The decoder 300 can render the output audio signal (334, 240, y R ) on the same speaker setup as that used for the input (212, y) or on a different speaker setup. Without loss of generality, it will be assumed that the input and output speaker setups are the same (but in the example, they may be different). In this section, the different modules that can make up the decoder 300 will be described.

[0527] Figure 3aFigures 3a and 3b depict a detailed overview of possible decoder processing. It is important to note that depending on the needs and requirements of a given application, at least some of the modules in Figure 3b (specifically those with dashed borders, such as 320, 330, 338) may be discarded. The decoder 300 may input (e.g., receive) two sets of data from the encoder 200:

[0528] - Side information 228 with encoded parameters (as described in 4.2.2)

[0529] - The downmixed signal (246, y) may be in the time domain (as described in 4.2.6).

[0530] The encoded parameters 228 may need to be decoded first (e.g., by the input unit 312), for example, using an inverse coding method previously employed. Once this step is completed, the relevant parameters for synthesis, such as the covariance matrix, can be reconstructed. In parallel, the downmixed signal (246, x) can be processed through several modules: First, an analysis filter bank 320 (see 4.2.1 ) can be used to obtain the frequency-domain version 324 of the downmixed signal 246. Then, a prototype signal 328 can be calculated (see 4.3.3 ), and an additional decorrelation step can be performed (at 330) (see 4.3.4 ). The key point of synthesis is the synthesis engine 334, which uses the covariance matrix (e.g., reconstructed at block 316) and the prototype signal (328 or 332) as inputs and produces the final signal 336 as output (see 4.3.5 ). Finally, the last step at the synthesis filter bank 338 can be completed, for example, if the analysis filter bank 320 was previously used, to produce the output signal 340 in the time domain.

[0531] 4.3.1 Entropy Decoding (e.g., Block 312)

[0532] The entropy decoding at block 312 (input interface) may allow obtaining the quantization parameters 314 previously obtained in 4 . The decoding of the bitstream 248 can be understood as a straightforward operation; the bitstream 248 can be read according to the coding method used in 4.2.5 and then decoded.

[0533] From the perspective of the implementation, the bitstream 248 may include signaling bits, which are not data but indicate certain peculiarities of the processing on the encoder side.

[0534] For example, in the case where the encoder 200 has the possibility to switch between several coding methods, the two first bits used can indicate which coding method has been used. The following bits can also be used to describe which parameter bands are currently being transmitted.

[0535] Other information that can be encoded in the side information of the bitstream 248 can include flags that indicate transients and a field 261 that indicates in which time slot of the frame the transient has occurred.

[0536] 4.3.2 Parameter Reconstruction

[0537] Parameter reconstruction can be performed, for example, by block 316 and / or the hybrid rule calculator 402.

[0538] The purpose of this parameter reconstruction is to reconstruct the covariance matrices C x and C y (or more generally, the covariance information associated with the downmixed signal 246 and the level and correlation information of the original signal) from the downmixed signal 246 and / or from the side information 228 (or from its version represented by the quantized parameters 314). These covariance matrices C x and C y may be necessary for synthesis because they are matrices that effectively describe the multichannel signal 246.

[0539] The parameter reconstruction at module 316 can be a two-step process:

[0540] First, the matrix C x (or more generally, the covariance information associated with the downmixed signal 246) is recalculated from the downmixed signal 246 (this step can be avoided if the covariance information associated with the downmixed signal 246 is actually encoded in the side information 228 of the bitstream 248); and

[0541] Then, the matrix C y (or more generally, the level and correlation information of the original signal 212) can be recovered, for example, at least in part using the transmitted parameters and C x or more generally the covariance information associated with the downmixed signal 246 (this step can be avoided if the level and correlation information of the original signal 212 is actually encoded in the side information 228 of the bitstream 248).

[0542] Note that in some examples, for each frame, it is feasible to use a linear combination of the reconstructed covariance matrices from previous frames, for example, by addition, averaging, etc., to smooth the covariance matrix C x of the current frame. For example, at the t-th frame, the final covariance to be used in formula (4) can be considered as the target covariance reconstructed from previous frames, for example

[0543] C x,t = C x,t + C x,t-1 .

[0544] However, in the case where a transient is present in the current frame, no smoothing operation is performed on the current frame. In the case of a transient, the current frame is not used for any smoothing of C x .

[0545] An overview of the process can be found below.

[0546] Note: As for the encoder, the processing here can be done independently for each frequency band on the basis of the parametric bands. For the sake of clarity, the processing will be described only for one specific frequency band and the notation will be adapted accordingly.

[0547] Aspect 4a: Reconstructing parameters in the case where the covariance matrix is transmitted

[0548] For this aspect, it is assumed that the parameter encoded (e.g., transmitted) in the side information 228 (the covariance matrix associated with the downmixed signal 246 and the channel levels and correlation information of the original signal 212) is the covariance matrix (or a subset thereof), as defined in aspect 2a. However, in some examples, the covariance matrix associated with the downmixed signal 246 and / or the channel levels and correlation information of the original signal 212 can be implemented by other information.

[0549] If the complete covariance matrix C x and C y are encoded (e.g., transmitted), then there is no further processing to be done at block 318 (so block 318 can be avoided in such examples). If only a subset of at least one of those matrices is encoded (e.g., transmitted), then the missing values must be estimated. The final covariance matrix as used in the synthesis engine 334 (or more specifically in the synthesis processor 404) will consist of the encoded (e.g., transmitted) values 228 and the estimated values on the decoder side. For example, if only some elements of matrix C y are encoded in the side information 228 of the bitstream 248, then the remaining elements of C y are estimated here.

[0550] For the covariance matrix C x of the downmixed signal 246, it is feasible to calculate the missing values by using the downmixed signal 246 on the decoder side and applying formula (1).

[0551] In one aspect, where the occurrence and location of a transient are transmitted or encoded, the same time slots are used on the encoder side for calculating the covariance matrix C of the downmixed signal 246x 。

[0552] For the covariance matrix C y , the missing values can be calculated with a first estimate in the following manner:

[0553]

[0554] Where:

[0555] - Estimate of the covariance matrix of the original signal 212 (this is an example of an estimated version of the original channel levels and correlation information)

[0556] - Q, the so-called prototype matrix (prototype rule, estimation rule), which describes the relationship between the downmixed signal and the original signal (see 4.3.3 )(this is an example of a prototype rule)

[0557] - C x Covariance matrix of the downmixed signal (this is an example of the covariance information of the downmixed signal 212)

[0558] - * Denotes conjugate transpose

[0559] Once these steps are completed, the covariance matrix will be obtained again and can be used for the final synthesis.

[0560] Aspect 4b: Reconstruction parameters in the case where ICC and ICLD are transmitted

[0561] For this aspect, it can be assumed that the encoded (e.g., transmitted) parameters in the side information 228 are the ICC and ICLD (or a subset thereof) defined in aspect 2b.

[0562] In this case, it may first be necessary to recalculate the covariance matrix C x . This can be done using the downmixed signal 212 on the decoder side and applying formula (1).

[0563] In one aspect, where the occurrence and location of the transients are transmitted, as in the encoder, the same time slots are used to calculate the covariance matrix C of the downmixed signal x . Then, the covariance matrix C y can be recalculated from the ICC and ICLD; this operation can be carried out as follows:

[0564] The energy (also referred to as the level) of each channel of the multichannel input can be obtained. These energies are derived using the transmitted inter-channel level differences and the following formula

[0565]

[0566] Where

[0567]

[0568] P i = C yi,i

[0569] where the weighting factor relates to the expected energy contribution of the channel pair to the downmixing, this weighting factor being fixed for certain input loudspeaker configurations and being known both at the encoder and at the decoder. In one implementation where a mapping is defined for each input channel i, where the mapping index is the downmixed channel j into which only the input channel i is mixed, or if the mapping index is greater than the number of downmixed channels. Thus, we have the mapping index m ICLD,i which is used to determine P in the following way dmx,i :

[0570]

[0571] These symbols are the same as those used in the parameter estimation in 4.2.3 .

[0572] These energies can be used to normalize the estimated C y . In the case where not all ICCs are transmitted from the encoder side, an estimate of C y can be calculated for the values that are not transmitted. The estimated covariance matrix can be obtained using formula (4) with the prototype matrix Q and the covariance matrix C x .

[0573] This estimate of the covariance matrix leads to an estimate of the ICC matrix, for which the entry at index (i,j) can be given by

[0574]

[0575] Thus, the "reconstruction" matrix can be defined as follows

[0576]

[0577] where

[0578] - the subscript R indicates the reconstruction matrix (which is an example of a reconstructed version of the original levels and correlation information)

[0579] - the ensemble {transmitted indices} corresponds to all ( i ,j) pairs that have been decoded in the side information 228 (e.g., transmitted from the encoder to the decoder).

[0580] In the example, since Less accurate than the encoded value ξ i,j Accurate, so ξ i,j May be more preferable than More preferable.

[0581] Finally, from the reconstructed ICC matrix thus obtained, the reconstructed covariance matrix can be inferred This matrix can be obtained by applying the energy obtained in formula (5) to the reconstructed ICC matrix, so for the index ( i , j):

[0582]

[0583] In the case where the complete ICC matrix is transmitted, only formulas (5) and (8) are required. The previous paragraphs describe a method for reconstructing missing parameters, other methods can be used, and the proposed method is not unique.

[0584] From the example of aspect 1b using the 5.1 signal, it can be noted that the values not transmitted are the values that need to be estimated on the decoder side.

[0585] Now the covariance matrix C x And It is important to interpret the reconstructed matrix Can be the estimated covariance matrix C of the input signal 212 y Of. The trade-off of the present invention can be to make the estimation of the covariance matrix on the decoder side close enough to the original, but also to transmit as few parameters as possible. These matrices may be essential for the final synthesis described in 4.3.5.

[0586] Note that in some examples, for each frame, a linear combination of the reconstructed covariance matrix prior to the current frame can be used to smooth the reconstructed covariance matrix of the current frame, for example, by addition, averaging, etc. For example, at frame t, the final covariance to be used for synthesis can be considered as the target covariance reconstructed from the previous frame, for example

[0587]

[0588] However, in the case of transients, no smoothing is done, and the C for the current frame yR Is used for the calculation of the mixing matrix.

[0589] It should also be noted that in some examples, for each frame, the unsmoothed covariance matrix of the downmixed channels C x Is used for parameter reconstruction, while the smoothed covariance matrix C as in Section 4.2.3 x,t Is used for the synthesis.

[0590] Figure 8a At the decoder 300, operations for obtaining the covariance matrix C are restored x and (e.g., performed at block 386 or 316...). In the Figure 8a block, the formula adopted by a specific block is also indicated between parentheses. It can be seen that through formula (1), the covariance estimator 384 allows the covariance C of the downmixed signal 324 (or its downsampled version 385) to be achieved x . By using formula (4) and an appropriate type of rule Q, the first covariance estimator block 384' allows the first estimate of the covariance C y to be achieved Subsequently, by applying formula (6), the covariance coherence block 390 obtains the coherence Subsequently, the ICC replacement block 392 selects between the estimated ICC and the ICC signaled in the side information 228 of the bitstream 348 by adopting formula (7). Then the selected coherence ξ R is input to the energy application block 394, and the energy application block 394 applies energy according to the ICLD (χ i ). Then, the target covariance matrix is provided to Figure 3a the mixer rule calculator 402 of Figure 3c or the covariance synthesis block 388, or

[0591] 4.3.3 Prototype Signal Calculation (Block 326)

[0592] The purpose of the prototype signal module 326 is to shape the downmixed signal 212 (or its frequency domain version 324) in a way that can be used by the synthesis engine 334 (see 4.3.5 ). The prototype signal module 326 can perform upmixing of the downmixed signal. The prototype signal module 326 can complete the calculation of the prototype signal 328 by multiplying the downmixed signal 212 (or 324) by a so-called prototype matrix Q:

[0593] Y p = XQ (9)

[0594] where

[0595] - Q is the prototype matrix (which is an example of a prototype rule)

[0596] - X is the downmixed signal (212 or 324)

[0597] - Y p is the prototype signal (328).

[0598] The way of establishing the prototype matrix may be processing-dependent and can be defined to meet the requirements of the application. The only limitation may be that the number of channels of the prototype signal 328 must be the same as the number of desired output channels; this directly limits the size of the prototype matrix. For example, Q can be a matrix, where the number of rows of the matrix is the number of channels of the downmixed signal (212, 324), and the number of columns is the number of channels of the final synthesized output signal (332, 340).

[0599] As an example, in the case of a 5.1 or 5.0 signal, the prototype matrix can be established as follows:

[0600]

[0601] Note that the prototype matrix can be predetermined and fixed. For example, for all frames, Q can be the same, but can be different for different frequency bands. In addition, there are different Qs for different relationships between the number of channels of the downmixed signal and the number of channels of the synthesized signal. For example, based on the number of specific downmixed channels and the number of specific synthesized channels, Q can be selected from multiple pre-stored Qs.

[0602] Aspect 5: Re-parameterize in the case where the output speaker setup is different from the input speaker setup:

[0603] One application of the proposed invention is to generate an output signal 336 or 340 that is different from the original signal 212 in terms of the speaker setup (for example, meaning having a different number of speakers).

[0604] To this end, the prototype matrix must be modified accordingly. In this case, the prototype signal obtained by formula (9) will include as many channels as the output speaker setup. For example, if we have a 5-channel signal as the input (on the signal 212 side) and want to obtain a 7-channel signal as the output (on the signal 336 side), the prototype signal will already include 7 channels.

[0605] In this way, the estimation of the covariance matrix in formula (4) still holds and will still be used to estimate the covariance parameters of the channels that do not exist in the input signal 212.

[0606] The parameters 228 transmitted between the encoder and the decoder are still relevant, and formula (7) can still be used. More precisely, the encoded (e.g., transmitted) parameters must be assigned to channel pairs that are geometrically as close as possible to the original setup. Basically, an adaptation operation is required.

[0607] For example, if the ICC value between one loudspeaker on the right and one loudspeaker on the left is estimated on the encoder side, this value can be assigned to the channel pair of the output setting with the same left and right positions; in the case of different geometries, this value can be assigned to the loudspeaker pair with positions as close as possible to the original loudspeakers.

[0608] Then, once the target covariance matrix C for the new output setting is obtained y , the remaining processing remains unchanged.

[0609] Therefore, in order to adapt the target covariance matrix to the number of synthesized channels, it is feasible to:

[0610] Use the prototype matrix Q, which is converted from the number of downmixed channels to the number of synthesized channels; this can be done by

[0611] adapting formula (9) so that the prototype signal has the number of synthesized channels;

[0612] adapting formula (4) so as to estimate with the number of synthesized channels

[0613] Keep formulas (5) to (8), which can thus obtain the number of original channels;

[0614] but assign the original channel groups (e.g., original channel pairs) to a single synthesized channel (e.g., select the assignment according to the geometry), and vice versa.

[0615] In Figure 8b an example is provided, which is Figure 8a a version of

[0616] where the ICC (obtained from the side information 228 of the bitstream 348) is applied to the ICC matrix at 392, moving the original channel groups (e.g., pairs of original channels) to a single synthesized channel (selecting the assignment according to the geometry), and vice versa. Then in a second step this matrix Applied to the input channel power (ICLD) to be transmitted and obtain a channel power vector for the number of output (synthesized) channels, and adjust the first target covariance matrix according to the vector to obtain a second target covariance matrix with the desired number of synthesized channels. The adjusted second target covariance matrix can now be used in synthesis. In Figure 8c An example is provided therein, Figure 8c is Figure 8a where blocks 390 to 394 operate to reconstruct the target covariance matrix to have a version with the number of original channels of the original signal 212. After that, at block 395, the prototype signal QN (to be converted to the number of synthesized channels) and the vector ICLD can be applied. It is noted that, Figure 8c block 386 of Figure 8a is the same as block 386 of Figure 8c except for the fact that in Figure 8a the number of channels of the reconstructed target covariance is exactly the same as the number of original channels of the input signal 212 (and in

[0617] 4.3.4 Decorrelation

[0618] For the purpose of the decorrelation module 330, it is to reduce the number of correlations between each channel of the prototype signal. Highly correlated speaker signals may result in phantom sources and degrade the quality and spatial characteristics of the output multichannel signal. This step is optional and can be performed or not according to the application requirements. In the present invention, decorrelation is used before the synthesis engine. As an example, an all-pass frequency decorrelator can be used.

[0619] Notes on MPEG Surround:

[0620] In MPEG Surround according to the prior art, the so-called "mixing matrix" (denoted as M in the standard 1 and M 2 ) is used. Matrix M 1 controls how the available downmixed signals are input to the decorrelator. Matrix M 2 describes how the direct signal and the decorrelated signal should be combined to produce the output signal.

[0621] Although it may be similar to the prototype matrix defined in 4.3.3 and the usage of the decorrelator described in this section, it is important to note that:

[0622] - The function of the prototype matrix Q is completely different from the matrix used in MPEG surround. The key point of this matrix is to generate a prototype signal. The purpose of the prototype signal is to be input into the synthesis engine.

[0623] - The prototype matrix is not intended to prepare a downmix signal for a decorrelator and can be adapted depending on the requirements and the application purpose. For example, the prototype matrix can generate a prototype signal for an output speaker setup that is greater than the input speaker setup.

[0624] - In the proposed invention, the use of a decorrelator is not mandatory; the processing relies on the use of the covariance matrix within the synthesis engine (see 5.1).

[0625] - The proposed invention does not generate an output signal by combining a direct signal and a decorrelated signal.

[0626] - M 1 and M 2 The calculation of highly depends on the tree structure. From a structural point of view, the different coefficients of these matrices vary depending on the situation. This is not the case in the proposed invention, where the processing is independent of the downmix calculation (see 5.2 ), and conceptually, the proposed processing aims to consider the relationships between each channel, rather than just considering channel pairs, as can be done using a tree structure.

[0627] Therefore, the present invention is different from MPEG surround according to the prior art.

[0628] 4.3.5 Synthesis Engine, Matrix Calculation

[0629] The last step of the decoder includes the synthesis engine 334 or the synthesis processor 402 (if needed, also including the synthesis filter bank 338). The purpose of the synthesis engine 334 is to generate the final output signal 336 relative to certain constraints. The synthesis engine 334 can calculate the output signal 336, and the characteristics of the output signal 336 are constrained by the input parameters. In the present invention, in addition to the prototype signal 328 (or 332), the input parameters 318 of the synthesis engine 338 are the covariance matrices C x and C y . Since the characteristics of the output signal should be as close as possible to the target covariance matrix defined by C y , it is especially called the target covariance matrix (it will be shown that an estimated version and a pre-built version of the target covariance matrix are discussed).

[0630] The available synthesis engine 334 is not the only one. As an example, the covariance synthesis of the prior art can be used [8], which is incorporated herein by reference. Another synthesis engine 333 that can be used will be the synthesis engine described in the DirAC processing of [2].

[0631] The output signal of the synthesis engine 334 may need to be further processed by the synthesis filter bank 338.

[0632] As a final result, the output multi-channel signal 340 is obtained in the time domain.

[0633] Aspect 6: High-quality output signal using "covariance synthesis"

[0634] As described above, the used synthesis engine 334 is not the only one, and any engine that uses the transmitted parameters or a subset thereof can be used. However, an aspect of the present invention can provide a high-quality output signal 336, for example, by using covariance synthesis [8].

[0635] This synthesis method aims to calculate the output signal 336, and the characteristics of the output signal 336 are defined by the covariance matrix For this purpose, the so-called optimal mixing matrices are calculated, which will mix the prototype signals 328 into the final output signal 336 and provide the best results from a mathematical point of view given the target covariance matrix

[0636] The mixing matrix M is the matrix that transforms the prototype signal x R into the output signal y P via the relation y P = Mx R (336).

[0637] The mixing matrix can also be the matrix that transforms the downmixed signal x into the output signal via the relation y R = Mx. From this relation, we can also infer

[0638] in the presented processing and C x and may be known in some examples (since they are the target covariance matrix of the downmixed signal 246 and the covariance matrix C x ) respectively.

[0639] From a mathematical point of view, one solution is given by where K y and are obtained by performing operations on C x and ​All matrices obtained by performing singular value decomposition. For P, it is an open parameter here, but with respect to the constraints governed by the prototype matrix Q, an optimal solution (from the perspective of the listener's perception) can be found. The mathematical proof presented here can be found in [8].

[0640] The synthesis engine 334 provides high-quality output 336 because the method is designed to provide an optimal mathematical solution for the reconstruction of the output signal problem.

[0641] In less mathematical terms, it is very important to understand that the covariance matrix represents the energy relationship between different channels of a multi-channel audio signal. The matrix C for the original multi-channel signal 212 y and the matrix C for the downmixed multi-channel signal 246 x . Each value of these matrices reflects the energy relationship between two channels of the multi-channel stream.

[0642] Therefore, the philosophy behind covariance synthesis is to generate a signal whose characteristics are driven by the target covariance matrix . This matrix is calculated in such a way as to describe the original input signal 212 (or, in the case of being different from the input signal, the output signal we want to obtain). Then, with these elements, covariance synthesis will optimally mix the prototype signals to generate the final output signal.

[0643] On the other hand, the mixing matrix for the synthesis of time slots is a combination of the mixing matrix M of the current frame and the mixing matrix M p of the previous frame to ensure smooth synthesis, such as linear interpolation based on the time slot index within the current frame.

[0644] On the other hand, where the occurrence and location of the transient are transmitted, before the transient position, the previous mixing matrix M p is used for all time slots, and the mixing matrix M is used for the time slot including the transient position and all subsequent time slots in the current frame. Note that in some examples, for each frame or time slot, a linear combination with the mixing matrix for the previous frame or time slot can be used to smooth the mixing matrix for the current frame or time slot, such as by addition, averaging, etc. Let's assume that for the current frame t, the time slot s frequency band i of the output signal is obtained by Y s,i = M s,i X s,i where M s,i is a combination of the mixing matrix M t-1,i for the previous frame, and M t,i is the mixing matrix calculated for the current frame, for example, a linear interpolation between them:

[0645]

[0646] where n s is the number of time slots in a frame (e.g., 16), and t - 1 and t indicate the previous frame and the current frame. More generally, the mixing matrix M calculated for the current frame is scaled by increasing factors along subsequent time slots of the current frame t t,i , and the mixing matrix M scaled by decreasing factors is added along subsequent time slots of the current frame t t-1,i to obtain the mixing matrix M associated with each time slot s,i . The factors can be linear.

[0647] It can be provided that in the case of a transient (signaled, for example, in message 261), the current mixing matrix and the past mixing matrix are not combined, but rather the previous time slots up to and including the transient and the current ones for the time slot including the transient and all subsequent time slots until the end of the frame

[0648]

[0649] where s is the time slot index, i is the frequency band index, t and t - 1 indicate the current frame and the previous frame, and s y is the time slot including the transient.

[0650] Differences from the Prior Art Document [8]

[0651] It is also important to note that the proposed invention goes beyond the scope of the method proposed in [8]. The significant differences are in particular:[[]]

[0652] - The target covariance matrix is calculated on the encoder side of the proposed processing.

[0653] - The target covariance matrix can also be calculated in a different way (in the proposed invention, the covariance matrix is not a sum of diffusion direct parts).

[0654] - The processing is not carried out separately for each frequency band, but rather for grouped parametric frequency bands (as described in 0 ).

[0655] - From a more global view: covariance synthesis is only one block of the whole process here and must be used together with all other elements on the decoder side.

[0656] 4.3 Preferred Aspects as a List

[0657] At least one of the following aspects can characterize the present invention:[[]]

[0658] 1. On the encoder side

[0659] a. Input the multi-channel audio signal 246.

[0660] b. Use the filter bank 214 to transform the signal 212 from the time domain to the frequency domain (216)

[0661] c. Calculate the downmixed signal 246 at block 244

[0662] d. Estimate the first parameter set from the original signal 212 and / or the downmixed signal 246 to describe the multi-channel stream (signal)

[0663] 246: Covariance matrix C x and / or C y

[0664] e. Transmit and / or encode the covariance matrix C x and / or C y Directly or calculate the ICC and / or ICLD and transmit them

[0665] f. Use an appropriate coding scheme to encode the transmitted parameter 228 in the bitstream 248

[0666] g. Calculate the downmixed signal 246 in the time domain

[0667] h. Transmit the side information (i.e., parameters) and the downmixed signal 246 in the time domain

[0668] 2. On the decoder side

[0669] a. Decode the bitstream 248 including the side information 228 and the downmixed signal 246

[0670] b. (Optional) Apply the filter bank 320 to the downmixed signal 246 to obtain a version 324 of the downmixed signal in the frequency domain

[0671] c. Reconstruct the covariance matrix C from the previously decoded parameter 228 and the downmixed signal 246 x and

[0672] d. Calculate the prototype signal 328 (324) from the downmixed signal 246

[0673] e. (Optional) Decorrelate the prototype signal (at block 330)

[0674] f. Use as the reconstructed C x and Apply the synthesis engine 334 to the prototype signal

[0675] g. (Optional) Apply the synthesis filter bank 338 to the output 336 of the covariance synthesis 334

[0676] h. Obtain the output multichannel signal 340

[0677] 4.5 Covariance Synthesis

[0678] In this section, some techniques that can be implemented in the systems of FIGS. 1 to 3d are discussed. However, these techniques can also be implemented independently: for example, in some examples, covariance calculations such as those performed in connection with Figures 8a to 8c and Formulas (1) to (8) are not required. Thus, in some examples, when referring to (the reconstructed target covariance), it can also be replaced by C y (which can also be provided directly without reconstruction). Nevertheless, the techniques of this section can be advantageously used in conjunction with the above techniques.

[0679] Now refer to Figures 4a to 4d . Here, examples of the covariance synthesis blocks 388a to 388d are discussed. Blocks 388 to 388d can be implemented as, for example, Figure 3c block 388 for covariance synthesis. Blocks 388a to 388d can be, for example, Figure 3a a part of the synthesis processor 404 and the mixing rule calculator 402 of the synthesis engine 334 and / or the synthesis processor 404 and the mixing rule calculator 402 of the parameter reconstruction block 316. In Figures 4a to 4d , the downmixed signal 324 is in the frequency domain FD (i.e., downstream of the filter bank 320) and is indicated by X, while the synthesized signal 336 is also in FD and is indicated by Y. However, it is feasible to generalize these results in the time domain. Note that Figures 4a to 4d each of the covariance synthesis blocks 388a to 388d in x and (or other reconstructed information) can be referred to as a single frequency band (e.g., decomposed once in 380), and the covariance matrix C x and (or other reconstructed information) can thus be associated with a specific frequency band. For example, covariance synthesis can be performed on a frame-by-frame basis, and in that case, the covariance matrix C x and (or other reconstructed information) is associated with a single frame (or with multiple consecutive frames): thus, covariance synthesis can be performed on a frame-by-frame basis or on a multiple-frame-by-multiple-frame basis.

[0680] In Figure 4a , the covariance synthesis block 388a can be composed of an energy-compensated optimal mixing block 600a and a missing correlator block. Basically, a single mixing matrix M is found, and the only significant operation performed additionally is the calculation of the energy-compensated mixing matrix M'.

[0681] Figure 4b A covariance synthesis block 388b inspired by [8] is shown. The covariance synthesis block 388b may allow obtaining the synthesized signal 336 as a synthesized signal having a first principal component 336M and a second residual component 336R. Although the principal component 336M may be obtained at the optimal principal component mixing matrix 600b, for example by x and Find the mixing matrix M in M , and no decorrelator is used, but the residual component 336R can be obtained in another way. R In principle, the relationship Usually, the obtained mixing matrix does not fully meet the requirements and can be used Find the residual target covariance. It can be seen that the downmix signal 324 can be derived on the path 610b (path 610b can be referred to as the second path, the second path is parallel to the first path 610b', the first path 610b' includes block 600b). The prototype version 613b of the downmix signal 324 (with Y pR (represented) can be obtained at the prototype signal block (upmix block) 612b. For example, a formula such as formula (9) can be used, that is,

[0682] Y pR =XQ

[0683] Examples of Q (prototype matrix or upmix matrix) are provided in this document. Downstream of block 612b, a decorrelator 614b is provided so that the prototype signal 613b is decorrelated to obtain a decorrelated signal 615b (also referred to as Indicated). At block 616b, from the decorrelated signal 615b, the decorrelated signal is estimated. The covariance matrix of (615b) By using C as the principal component mixture x The equivalent value of is the decorrelated signal The covariance matrix of and C as the target covariance in another optimal mixing block r , the residual component 336R of the synthesized signal 336 can be obtained at the optimal residual component mixing matrix block 618b. The optimal residual component mixing matrix block 618b can be implemented in such a way that the mixing matrix M is generated. R , to mix the decorrelated signal 615b and obtain a residual component 336R (for a particular frequency band) of the synthesized signal 336. At adder block 620b, the residual component 336R is added to the main component 336M (so paths 610b and 610b' are connected together at adder block 620b).

[0684] Figure 4c Show alternative Figure 4b 388c of the covariance synthesis 388b. The covariance synthesis block 388c allows obtaining the synthesized signal 336 as a signal Y having a first principal component 336M' and a second residual component 336R'. Although the principal component 336M' can be obtained at the optimal principal component mixing matrix 600c, for example by x and (or C y Other information 220) find the mixing matrix M M , and no correlator is used, but the residual component 336R' can be obtained in another way. The downmix signal 324 can be derived to the path 610c (path 610c can be called the second path, the second path is parallel to the first path 610c', and the first path 610c' includes block 600c). By applying the prototype matrix Q (for example, a matrix that upmixes the downmix signal 234 to the version 613c of the downmix signal 234 with the number of channels, i.e., the number of synthesized channels), the prototype version 613c of the downmix signal 324 can be obtained at the downmix block (upmix block) 612c. For example, a formula such as formula (9) can be used. This document provides an example of Q. Downstream of block 612c, a decorrelator 614c can be provided. In some examples, the first path does not have a decorrelator, while the second path has a decorrelator.

[0685] Decorrelator 614c can provide a decorrelated signal 615c (also referred to as instructions). However, unlike Figure 4b The covariance synthesis block 388b is the opposite of the technique used in Figure 4c The covariance synthesis block 388c does not decorrelate the signal 615c from Estimate the covariance matrix of the decorrelated signal 615c In contrast, the covariance matrix of the decorrelated signal 615c is is obtained (at block 616c) from:

[0686] The covariance matrix C of the downmix signal 324 is x (For example, in Figure 3c at block 384 and / or estimated using equation (1); and

[0687] Prototype matrix Q.

[0688] By using the covariance matrix C from the downmix signal 324 x The estimated covariance matrix C as the principal component mixing matrix x and Cr As an equivalent of the target covariance matrix, the residual component 336R' of the synthesized signal 336 is obtained at the optimal residual component mixing matrix block 618c. The optimal residual component mixing matrix block 618c can be implemented in a manner that generates the residual component mixing matrix M R to mix the decorrelated signal 615c according to the residual component mixing matrix M R to obtain the residual component 336R'. At the adder block 620c, the residual component 336R' is added to the principal component 336M' to obtain the synthesized signal 336 (paths 610c and 610c' are thus joined together at the adder block 620c).

[0689] In some examples, the residual component 336R or 336R' is not always or need not be calculated (and paths 610b or 610c are not always used). In some examples, although covariance synthesis is performed for some frequency bands without calculating the residual signal 336R or 336R', for other frequency bands of the same frame, the residual signal 336R or 336R' is also considered to handle covariance synthesis. Figure 4d An example showing the covariance synthesis block 388d, which can be a specific case of the covariance synthesis block 388b or 388c: Here, the band selector 630 can select or deselect (in the manner indicated by the switch 631) the calculation of the residual signal 336R or 336R'. For example, paths 610b or 610c can be selectively enabled by the selector 630 for some frequency bands and disabled for other frequency bands. In particular, paths 610b or 610c can be disabled for frequency bands above a predetermined threshold (such as a fixed threshold), and the predetermined threshold (such as a maximum value) can be used to distinguish between frequency bands where the human ear is insensitive to phase (frequency bands above the threshold) and frequency bands where the human ear is sensitive to phase (frequency bands below the threshold). Thus, the residual component 336R or 336R' is not calculated for frequency bands below the threshold, and the residual component 336R or 336R' is calculated for frequency bands above the threshold.

[0690] Figure 4d Examples of can also be obtained by replacing block 600b or 600c with block 600a of Figure 4a and replacing block 610b or 610c with the covariance synthesis block 388b of Figure 4b or the covariance synthesis block 388c of Figure 4c .

[0691] Some indications on how to obtain the mixing rules (matrices) at blocks 338, 402 (or 404), 600a, 600b, 600c, etc. are provided here. As mentioned above, there are many ways to obtain the mixing matrix, but some of them will be discussed in more detail here.

[0692] Specifically, first, refer to Figure 4b the covariance synthesis block 388b. At the optimal principal component mixing matrix block 600c, for example, the mixing matrix M of the principal component 336M of the synthesized signal 336 can be obtained from the following formula:

[0693] The covariance matrix C of the original signal 212 y (C y can be estimated using at least some of the formulas (6) to (8) discussed above, for example, see Figure 8; it can be in the form of a so-called "target version" e.g., the value estimated according to formula (8)); and

[0694] the covariance matrix C of the downmixed signals 246, 324 x (C y can be estimated using, for example, formula (1)).

[0695] For example, as proposed in [8], according to the following factorization, it is recognized to factorize the covariance matrix C x and C y , which are Hermitian matrices and positive semi - definite matrices:

[0696]

[0697]

[0698] K x and K y can be obtained, for example, by applying the singular value decomposition (SVD) twice to C x and C y . For example:

[0699] C x The SVD of can provide the matrix U of the singular vectors (e.g., left singular vectors) Cx ; and

[0700] the diagonal matrix S of the singular values Cx ;

[0701] Therefore, K x can be obtained by multiplying U Cx by a diagonal matrix that has in its terms the square roots of the values in the corresponding terms of S Cx .

[0702] In addition, regarding the SVD of C y can provide:

[0703] the matrix V of the singular vectors (e.g., right singular vectors) Cy ; and

[0704] The diagonal matrix S of singular values Cy

[0705] Therefore, K y can be obtained by multiplying U Cy by a diagonal matrix that has in its terms the square roots of the values in the corresponding terms of S Cy .

[0706] Then, it is possible to obtain the principal component mixing matrix M_M which, when applied to the downmixed signal 324, will allow the principal components 336M of the synthesized signal 336 to be obtained. The principal component mixing matrix M M can be obtained as follows:

[0707]

[0708] If K x is a non-invertible matrix, a regularized inverse matrix can be obtained using known techniques and substituted instead of

[0709] The parameter P is usually free, but it can be optimized. To derive P, SVD can be applied to:

[0710] C x (the covariance matrix of the downmixed signal 324); and

[0711] (the covariance matrix of the prototype signal 613b).

[0712] Once the SVD is performed, it is possible to obtain P, as

[0713] P = VΛU *

[0714] Λ is a matrix that has the same number of rows as the number of synthesized channels and the same number of columns as the number of downmixed channels. Λ is the identity in its first square block and is completed with zeros in the remaining terms. Now, how V and U are obtained from C x and is described. V and U are matrices of the singular vectors obtained from the SVD:

[0715]

[0716] S is the diagonal matrix of singular values usually obtained by SVD. is a diagonal matrix that normalizes the per-channel energy of the prototype signal (615b) to the energy of the synthesized signal y. To obtain it is first necessary to calculate i.e., the prototype signal The covariance matrix (614b). Then, in order to obtain from obtain by normalizing the diagonal values of to C y to the corresponding diagonal values of, thereby providing An example is the diagonal terms of are calculated as where c yii is the value of the diagonal term of C y and is the value of the diagonal term of.

[0717] Once obtained the covariance matrix C of the residual components r can be obtained from

[0718]

[0719] Once C r is obtained, it is possible to obtain the mixing matrix for mixing the decorrelated signals 615b to obtain the residual signal 336R, where in the case of the same optimal mixing C r having the same effect as the main optimal mixing the covariance of the decorrelation prototype acts as the input signal covariance C x having the main optimal mixing.

[0720] However, it has been understood that, compared with the Figure 4b technique of Figure 4c the technique of has some advantages. In some examples, Figure 4c the technique of is the same as the Figure 4c technique of, at least for calculating the main matrix and for generating the main components of the synthetic signal. On the contrary, Figure 4c the technique of is different from the Figure 4b technique of in the calculation of the residual mixing matrix, and more generally, for generating the residual components of the synthetic signal. Now referring to Figure 11 in combination with Figure 4c for calculating the residual mixing matrix. In the Figure 4c example of, a decorrelator 614c in the frequency domain is used, which ensures the decorrelation of the prototype signal 613c, but retains the energy of the prototype signal 613b itself.

[0721] In addition, in the Figure 4c example of, we can assume (at least approximately) that the decorrelated channels of the decorrelated signal 615c are mutually incoherent, so all non-diagonal elements of the covariance matrix of the decorrelated signal are zero. With these two assumptions, we can simply by in Cx Apply Q on it to estimate the covariance of the decorrelated prototype, and only use the main diagonal of the covariance (i.e., the energy of the prototype signal). Starting from the decorrelated signal 615b, Figure 4c The technique is more efficient than Figure 4b the example of x where we need to perform the same frequency band / slot aggregation as has been done for C Figure 4c In the example of x we can simply apply the already aggregated C

[0722] Therefore, the covariance of the decorrelated signal can be estimated at 710 using the following 711

[0723] P decorr = diag(QC x Q * )

[0724] As the main diagonal of a matrix with all non - diagonal elements set to zero, it is used as the input signal covariance In the example C x is smoothed for performing the synthesis of the principal components 336M' of the synthetic signal, and the technique can be used to calculate P x using C decorr as non - smoothed C x .

[0725] Now, the prototype matrix QR should be used. However, it has been noted that for the residual signal, QR is the identity matrix. Knowing (diagonal matrix) and QR (identity matrix) properties can further simplify the calculation of the mixing matrix (at least one SVD can be omitted), see the following techniques and Matlab listings.

[0726] First, similar to Figure 4b the example of r the residual target covariance matrix C of the input signal 212 (Hermitian, positive semi - definite) can be decomposed into The matrix K can be obtained by SVD(702) r : SVD 702 is used for C_r to produce:

[0727] The matrix U of the singular vectors (e.g., left singular vectors) Cr ;

[0728] The diagonal matrix S of the singular values Cr ;

[0729] Thus K r is obtained (at 706) by multiplying U in a diagonal matrix Cr by a diagonal matrix that has in its entries the square roots of the values in the corresponding elements of S Cr (the latter having been obtained at 704).

[0730] At this point, in theory, another SVD could be applied to the covariance of the decorrelated prototype

[0731] However, in this example ( Figure 4c ), a different path has been chosen to reduce the computational load. From P decorr = diag(QC x Q * ), the estimated is a diagonal matrix, so no SVD is needed (the SVD of a diagonal matrix gives the singular values as a sorted vector of the diagonal elements, while the left and right singular vectors only indicate the sorting indices). By calculating (at 712) the square root of each value at the diagonal entries of , the diagonal matrix The diagonal matrix is such that has the advantage that no SVD is needed to obtain . From the diagonal covariance of the decorrelated signals , the estimated covariance matrix of the decorrelated signals 615c But since the prototype matrix is Q r (i.e., a homogeneous matrix), can be directly used to formulate as where is the value of the diagonal entry of C r , and is the value of the diagonal entry of is a diagonal matrix (obtained at 722) that normalizes the per-channel energy of the decorrelated signals (615b) to the desired energy of the synthesized signal y

[0732] At this point, it is possible (at 734) to multiply by )(also called the result 735 of multiplication 734). Then (736), multiply K r by to get K′ y (i.e., ). From K′ y, SVD(738) can be performed to obtain the left singular vector matrix U and the right singular vector matrix V. By multiplying V and U* (740), the matrix P (P = VU H ) is obtained. Finally (742), the mixing matrix M of the residual signal can be obtained by applying the following R :

[0733]

[0734] where (obtained at 745) can be replaced by the regularized inverse. M R Therefore, it can be used at block 618c for residual mixing.

[0735] The Matlab code for performing covariance synthesis as described above is provided here. Note that the asterisk (*) in the code represents multiplication, and the prime (') represents the Hermitian matrix.

[0736] % Calculate the residual mixing matrix

[0737] function [M]=

[0738] ComputeMixingMatrixResidual(C_hat_y,Cr,reg_sx,reg_ghat)

[0739] EPS_ = single(1e - 15); % Epsilon to avoid division by zero

[0740] num_outputs = size(Cr,1);

[0741] % Decomposition of Cy

[0742] [U_Cr,S_Cr]=svd(Cr);

[0743] Kr = U_Cr*sqrt(S_Cr);

[0744] % The singular value decomposition of a diagonal matrix is the sorted diagonal elements,

[0745] % We can skip the sorting and directly obtain Kx from Cx

[0746] K_hat_y = sqrt(diag(C_haty));

[0747] limit = max(K_hat_y)*reg_sx + EPS_;

[0748] S_hat_y_reg_diag = max(K_hat_y,limit);

[0749] % Formulated regularized Kx

[0750] K_hat_y_reg_inverse = 1. / S_hat_y_reg_diag;

[0751] % Formulated normalization matrix G_hat

[0752] % Q is the identity matrix in case of the residual / diffuse part so

[0753] % Q*Cx*Q' = Cx

[0754] Cy_hat_diag = diag(C_hat_y);

[0755] limit = max(Cy_hat_diag)*reg_ghat + EPS_;

[0756] Cy_hat_diag = max(Cy_hat_diag, limit);

[0757] G_hat = sqrt(diag(Cr). / Cy_hat_diag);

[0758] % Formulated optimal P

[0759]

[0760] % Formulated M

[0761]

[0762]

[0763] A discussion on the covariance synthesis of Figure 4b and 4c is provided here. In some examples, two synthesis methods can be considered for each frequency band. For some frequency bands, the frequency bands above a specific frequency where the human ear is generally insensitive to phase include the complete synthesis of the residual path from Figure 4b to achieve the required energy for applying energy compensation in the vocal tract.

[0764] Therefore, also in the examples of Figure 4b for frequency bands below a certain (fixed, known to the decoder) frequency band boundary (threshold), a complete synthesis according to Figure 4b can be performed (for example, in the case of Figure 4d ). In Figure 4bIn the example of, the covariance of the decorrelated signal 615b is derived from the decorrelated signal 615b itself. In contrast, in Figure 4c the example of, a decorrelator 614c in the frequency domain is used, which ensures the decorrelation of the prototype signal 613c but retains the energy of the prototype signal 613b itself.

[0765] Further considerations:

[0766] · In both the examples of Figure 4b and 4c : At the first path (610b’, 610c’), by relying on the covariance C y of the original signal 212 and the covariance C x of the downmixed signal 324 to generate a mixing matrix M M (at blocks 600b, 600c);

[0767] · In both the examples of Figure 4b and 4c : At the second path (610b, 610c), there are decorrelators (614b, 614c), and a mixing matrix M R (at blocks 618b, 618c) is generated, which should consider the covariance of the decorrelated signals (616b, 616c) However

[0768] o In the example of Figure 4b , the decorrelated signals (616b, 616c) are used to intuitively calculate the covariance of the decorrelated signals (616b, 616c) and are weighted in the energy of the original channel y.

[0769] o In the example of Figure 4c , the covariance of the decorrelated signals (616b, 616c) is estimated from the matrix C x and back-calculated in an intuitive way, and is weighted in the energy of the original channel y.

[0770] Note that the covariance matrix can be the reconstruction target matrix discussed above (e.g., obtained from the channel levels and correlation information 220 written in the side information 228 of the bitstream 248), and can thus be considered associated with the covariance of the original signal 212. In any case, since it will be used for synthesizing the signal 336, the covariance matrix can also be considered as the covariance associated with the synthesized signal. The same applies to the residual covariance matrix C r , which can also be understood as the residual covariance matrix associated with the synthesized signal (C r) and the main covariance matrix can also be understood as the main covariance matrix associated with the synthesized signal.

[0771] 5. Advantages

[0772] 5.1 Reducing the Use of Decorrelation and Optimizing the Use of the Synthesis Engine

[0773] Given the proposed technology, as well as the parameters used for processing and the way these parameters are combined with the synthesis engine 334, it is shown that the need for strong decorrelation of the audio signal (e.g., in its version 328) is reduced. Even in the absence of the decorrelation module 330, if not removed, the decorrelated effects (e.g., artifacts or degradation of spatial characteristics or degradation of signal quality) can also be reduced.

[0774] More precisely, as previously mentioned, the decorrelation part 330 of the processing is optional. In fact, the synthesis engine 334 uses the target covariance matrix C y (or a subset thereof) to decorrelate the signal 328 and ensure that the channels constituting the output signal 336 are decorrelated appropriately among themselves. C y The values in the covariance matrix represent the energy relationships between the different channels of our multichannel audio signal, which is why it is used as the target for synthesis.

[0775] In addition, the encoded (e.g., transmitted) parameters 228 (e.g., in their versions 314 or 318) combined with the synthesis engine 334 can ensure a high-quality output 336, given the fact that the synthesis engine 334 uses the target covariance matrix C y to reproduce the output multichannel signal 336, and the spatial characteristics and sound quality of the output multichannel signal 336 are as close as possible to the input signal 212.

[0776] 5.2 Downmix-Agnostic Processing

[0777] Given the proposed technology, as well as the way the prototype signals 328 are calculated and how they are used with the synthesis engine 334, it is shown here that the proposed decoder is independent of the way the downmixed signal 212 is calculated at the encoder.

[0778] This means that the proposed invention can be executed at the decoder 300 independently of the way the downmixed signal 246 is calculated at the encoder, and the output quality of the signal 336 (or 340) does not depend on a specific downmixing method.

[0779] 5.3 Scalability of Parameters

[0780] Given the proposed technology, and the way in which the parameters (28, 314, 318) are calculated and their use with the synthesis engine 334, and the way in which they are estimated on the decoder side, this illustrates that the parameters used to describe the multichannel audio signal are scalable both in number and in use.

[0781] Typically, only a subset of the parameters estimated only on the encoder side (e.g., a subset of C y and / or C x , such as its elements) are encoded (e.g., transmitted): this allows reducing the bit rate used by the processing. Thus, given the fact that the non-transmitted parameters are reconstructed on the decoder side, the number of parameters (e.g., elements of C y and / or C x ) that are encoded (e.g., transmitted) can be scalable. This gives the opportunity to scale the whole processing in terms of output quality and bit rate, the more parameters are transmitted, the better the output quality, and vice versa.

[0782] Moreover, those parameters (e.g., a subset of C y and / or C x or its elements) are scalable in purpose, which means that they can be controlled by user input to modify the characteristics of the output multichannel signal. In addition, those parameters can be calculated for each frequency band, and thus allow scalable frequency resolution.

[0783] For example, it can be decided to cancel one speaker in the output signal (336, 340), and thus the parameters can be directly manipulated on the decoder side to achieve such a transformation.

[0784] 5.4 Flexibility of Output Settings

[0785] Given the proposed technology, and the flexibility of the synthesis engine 334 and the parameters (e.g., a subset of C y and / or C x or its elements) used, it is illustrated here that the proposed invention allows a wide range of rendering possibilities regarding the output settings.

[0786] More precisely, the output settings do not have to be the same as the input settings. It is feasible to manipulate the reconstructed target covariance matrix fed into the synthesis engine to produce the output signal 340 on the speaker settings, where the speaker settings are greater than or less than or only have a geometry different from the original speaker settings. This is possible because the parameters to be transmitted and the proposed system are independent of the downmixed signal (see 5.2).

[0787] For these reasons, it is flexible to interpret the proposed invention from the viewpoint of the output speaker settings.

[0788] 5.Some Examples of Prototype Matrices

[0789] Here, the following table is for 5.1, but the LFE is excluded. Thereafter, we also include the LFE in the processing (only one ICC for the relationship LFE / C and the ICLD for the LFE are sent only in the lowest parametric band and are set to 1 and 0 respectively for all other bands in the synthesis at the decoder side). The channel naming and order follow the CICP in ISO / IEC 23091-3 "Information technology – Coding independent code points – Part 3: Audio". Q is always used as the prototype matrix in the decoder and the downmix matrix in the encoder. 5.1 (CICP6). α i To be used for calculating the ICLD.

[0790]

[0791] α i = [0.4444 0.4444 0.2 0.2 0.4444 0.4444]

[0792] 7.1 (CICP12)

[0793]

[0794] α i = [0.2857 0.2857 0.5714 0.5714 0.2857 0.2857 0.2857 0.2857]

[0795] 5.1 + 4 (CICP16)

[0796]

[0797] α i = [0.1818 0.1818 0.3636 0.3636 0.1818 0.1818 0.1818 0.1818 0.18180.1818]

[0799] 7.1 + 4 (CICP19)

[0800]

[0801] αi = [0.1538 0.1538 0.3077 0.3077 0.1538 0.1538 0.1538 0.1538 0.15380.1538 0.1538

[0803] 6. Method

[0804] Although the above technologies have been mainly discussed as components or functional devices, the present invention can also be implemented as a method. The blocks and elements discussed above can also be understood as steps and / or stages of a method.

[0805] For example, there is provided a decoding method for generating a synthesized signal from a downmixed signal, the synthesized signal having a plurality of synthesized channels, the method comprising:

[0806] Receiving a downmixed signal (246, x) having a plurality of downmixed channels and side information (228), the side information (228) including:

[0807] Channel levels and correlation information (220) of an original signal (212, y) having a plurality of original channels;

[0808] Using the channel levels and correlation information (220) of the original signal (212, y) and covariance information (C x ) associated with the signal (246, x) to generate the synthesized signal.

[0809] The decoding method may include at least one of the following steps:

[0810] Calculating a prototype signal from the downmixed signal (246, x), the prototype signal having the number of synthesized channels;

[0811] Using the channel levels and correlation information (212, y) of the original signal and covariance information associated with the downmixed signal (246, x) to calculate a mixing rule; and

[0812] Using the prototype signal and the mixing rule to generate the synthesized signal.

[0813] There is also provided a decoding method for generating a synthesized signal (336) from a downmixed signal (324, x) having a plurality of downmixed channels, the downmixed signal (336) having a plurality of synthesized channels, the downmixed signal (324, x) being a downmixed version of an original signal (212) having a plurality of original channels, the method comprising the following stages:

[0814] A first stage (610c'), including:

[0815] Synthesizing a first component (336M') of the synthesized signal according to a first mixing matrix (M M ) calculated from the following:

[0816] A covariance matrix associated with the synthesized signal (e.g., the reconstructed target version of the covariance of the original signal); and

[0817] a covariance matrix (C associated with the downmixed signal (324) x ).

[0818] A second stage (610c) for synthesizing a second component (336R') of the synthesized signal, where the second component (336R') is a residual component, and the second stage (610c) includes:

[0819] A prototype signal step (612c) for upmixing the downmixed signal (324) from the number of downmixed channels to the number of synthesized channels;

[0820] A decorrelator step (614c) for decorrelating the upmixed prototype signal (613c);

[0821] A second mixing matrix step (618c) for synthesizing the second component (336R') of the synthesized signal according to a second mixing matrix (M from the decorrelated version (615c) of the downmixed signal (324), and the second mixing matrix (M R ) is a residual mixing matrix, R )

[0822] wherein the method calculates the second mixing matrix (M R ) from the following:

[0823] The residual covariance matrix (C provided by the first mixing matrix step (600c); and r )

[0824] An estimate of the covariance matrix of the decorrelated prototype signal obtained from the covariance matrix (C x ) associated with the downmixed signal (324) ;

[0825] wherein the method further includes an adder step (620c) for adding the first component (336M') of the synthesized signal

[0826] to the second component (336R') of the synthesized signal to obtain the synthesized signal (336).

[0827] In addition, an encoding method is provided for generating a downmixed signal (246, x) from an original signal (212, y), where the original signal (212, y) has a plurality of original channels and the downmixed signal (246, x) has a plurality of downmixed channels, and the method includes:

[0828] Estimate the channel levels and related information (220) of the original signal (212, y).

[0829] Encode the downmixed signal (246, x) into a bitstream (248) such that the downmixed signal (246, x) is encoded in the bitstream (248) with side information (228) that includes the channel levels and related information (220) of the original signal (12, y).

[0830] These methods can be implemented in any of the encoders and decoders discussed above.

[0831] 7. Storage Unit

[0832] In addition, the present invention can be implemented in a non-transitory storage unit storing instructions that, when executed by a processor, cause the processor to perform the methods described above.

[0833] In addition, the present invention can be implemented in a non-transitory storage unit storing instructions that, when executed by the processor, cause the processor to control at least one of the functions of the encoder or the decoder.

[0834] The storage unit can be, for example, part of encoder 200 or decoder 300.

[0835] 8. Other Aspects

[0836] Although some aspects have been described in the context of apparatus, it is clear that these aspects also represent a description of the corresponding methods, where a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of the corresponding blocks or items or features of an apparatus. Some or all of the method steps can be performed by (or using) hardware devices such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some aspects, such a device can perform one or more of the most important method steps.

[0837] Depending on certain implementation requirements, aspects of the present invention can be implemented in hardware or software. The implementation can be carried out using a digital storage medium, such as a floppy disk, a DVD, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a FLASH memory, on which are stored electronically readable control signals that cooperate (or are capable of cooperating) with a programmable computer system such that the corresponding methods are carried out. Thus, the digital storage medium can be computer-readable.

[0838] Some aspects in accordance with the present invention include a data carrier having electronically readable control signals which are capable of cooperating with a programmable computer system such that one of the methods described herein is performed.

[0839] Generally speaking, aspects of the present invention may be implemented as a computer program product having program code which, when the computer program product is run on a computer, is operable to perform one of the methods. The program code may be stored, for example, on a machine-readable carrier.

[0840] Other aspects include a computer program for performing one of the methods described herein, stored on a machine-readable carrier.

[0841] In other words, thus, one aspect of the method of the present invention is a computer program having program code which, when the computer program is run on a computer, is for performing one of the methods described herein.

[0842] Thus, another aspect of the method of the present invention is a data carrier (or a digital storage medium or a computer-readable medium) including the computer program recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recording medium is generally tangible and / or non-transitory.

[0843] Thus, another aspect of the method of the present invention is a data stream or a signal sequence representing the computer program for performing one of the methods described herein. The data stream or the signal sequence may be configured, for example, to be transmitted via a data communication connection, such as via the Internet.

[0844] Another aspect includes a processing device, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.

[0845] Another aspect includes a computer on which the computer program has been installed for performing one of the methods described herein.

[0846] Another aspect in accordance with the present invention includes an apparatus or a system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a storage device or the like. The apparatus or the system may include, for example, a file server for transferring the computer program to the receiver.

[0847] In some aspects, a programmable logic device (e.g., a programmable logic array) can be used to perform some or all of the functions of the methods described herein. In some aspects, a programmable logic array can cooperate with a microprocessor to execute one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0848] The devices described herein can be implemented using a hardware device or using a computer, or using a combination of a hardware device and a computer.

[0849] The methods described herein can be performed using a hardware device or using a computer, or using a combination of a hardware device and a computer.

[0850] The aspects described above are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to those of ordinary skill in the art. Accordingly, the intention of the present invention is limited only by the scope of the appended patent claims and not by the specific details presented in the descriptions and explanations of the aspects herein.

[0851] 9. Bibliography

[0852] [1] J.Herre, K. J.Breebart, C.Faller, S.Disch, H.Purnhagen, J.Koppens, J.Hilpert, J. W.Oomen, K.Linzmeier and K.S.Chong, “MPEG Surround—The ISO / MPEG Standard for Efficient and Compatible Multichannel Audio Coding,” Audio Engineering Society, vol. 56, no. 11, pp. 932 - 955, 2008.

[0853] [2] V.Pulkki, “Spatial Sound Reproduction with Directional Audio Coding,” Audio Engineering Society, vol. 55, no. 6, pp. 503 - 516, 2007.

[0854] [3] C. Faller and F. Baumgarte, “Binaural Cue Coding - Part II: Schemes and Applications,” IEEE Transactions on Speech and Audio Processing, vol. 11, no. 6, pp. 520 - 531, 2003.

[0855] [4] O. Hellmuth, H. Purnhagen, J. Koppens, J. Herre, J. J. Hilpert, L. Villemoes, L. Terentiv, C. Falch, A. M. L. Valero, B. Resch, H. Mundt and H.-O. Oh, “MPEG Spatial Audio Object Coding – The ISO / MPEG Standard for Efficient Coding of Interactive Audio Scenes,” in AES, San Fransisco, 2010.

[0856] [5] L. Mikko-Ville and V. Pulkki, “Converting 5.1 Audio Recordings to B-Format for Directional Audio Coding Reproduction,” in ICASSP, Prague, 2011.

[0857] [6] D. A. Huffman, “A Method for the Construction of Minimum-Redundancy Codes,” Proceedings of the IRE, vol. 40, no. 9, pp. 1098 - 1101, 1952.

[0858] [7] A. Karapetyan, F. Fleischmann and J. Plogsties, “Active Multichannel Audio Downmix,” in 145th Audio Engineering Society, New York, 2018.

[0859] [8] J. Vilkamo, T. and A.Kuntz, “Optimized Covariance Domain Framework for Time-Frequency Processing of Spatial Audio,” Journal of the Audio Engineering Society, vol.61, no.6, pp.403-411, 2013.

Claims

1. An audio synthesizer (300) for generating a synthesized signal (336, 340, y) from a downmixed signal (246, x R ), the synthesized signal (336, 340, y R ) having a plurality of synthesized channels, the audio synthesizer (300) Comprising: An input interface (312) configured to receive the downmixed signal (246, x), the downmixed signal (246, x) having a plurality of downmixed channels and side information (228), the side information (228) including the channel levels and correlation information (314, ξ, χ) of the original signal (212, y), the original signal (212, y) having a plurality of original channels; and A synthesis processor (404), configured to generate the synthesized signal (336, 340, y R ) according to at least one mixing rule that is a mixing matrix The channel levels and correlation information (220, 314, ξ, χ) of the original signal (212, y); and Covariance information (C associated with the downmixed signal (324, 246, x) x ); wherein the audio synthesizer is configured to reconstruct (386) a target version of the covariance information (C y ) of the original signal based on an estimated version of the original covariance information (C ) y ) wherein the estimated version of the original covariance information (C y ) is reported to the plurality of synthesized channels; wherein, the audio synthesizer is configured to obtain the estimated version of the original covariance information from covariance information (C x ) associated with the downmixed signal (324, 246, x) wherein, the audio synthesizer is configured to obtain the estimated version of the original covariance information (220) by applying an estimation rule (Q) to the covariance information (C x ) associated with the downmixed signal (324, 246, x) The estimation rule (Q) is a prototype rule for calculating a prototype signal (326) or is associated with a prototype rule for calculating a prototype signal (326).

2. The audio synthesizer (300) according to claim 1, Comprising: A prototype signal calculator (326) configured to calculate a prototype signal (328) from the downmixed signal (324, 246, x), the prototype signal (328) having the plurality of synthesized channels; A mixing rule calculator (402) configured to calculate at least one mixing rule (403) using: The channel levels and correlation information (314, ξ, χ) of the original signal (212, y); and The covariance information (C associated with the downmixed signal (324, 246, x) x ); wherein the synthesis processor (404) is configured to generate the synthesized signal (336, 340, y) using the prototype signal (328) and the at least one hybrid rule (403) R ) 3. The audio synthesizer according to claim 1, configured to reconstruct the target covariance information (C R ) adapted to the number of channels of the synthesized signal (336, 340, y y ).

4. The audio synthesizer according to claim 3, configured to reconstruct covariance information (C R ) adapted to the number of channels of the synthesized signal (336, 340, y y ) by assigning a group of original channels to a single synthesized channel, or vice versa, such that the reconstructed target covariance information is reported to multiple channels of the synthesized signal (336, 340, y R ).

5. The audio synthesizer according to claim 4, configured to reconstruct the covariance information (C y ) for the number of channels of the synthesized signal (336, 340, y R ) adapted to the synthesized signal by generating target covariance information for the number of original channels and then applying a downmix rule or an upmix rule and energy compensation to derive the target covariance for the synthesized channels. R ) of the number of channels of y ) 6. The audio synthesizer according to claim 1, configured to normalize, for at least one channel pair, the estimated version y of the original covariance information (C ) to the square root of the level of the channels in the channel pair.

7. The audio synthesizer according to claim 6, configured to construct a matrix using the estimated version of the normalized original covariance information (C y ) thereof.

8. The audio synthesizer according to claim 7, configured to complete the matrix by inserting an item (908) obtained from the side information (228) of the bitstream (248).

9. The audio synthesizer according to claim 6, configured to denormalize the matrix by scaling the estimated version of the original covariance information (C y ) by the square root of the levels of the channels forming the channel pair .

10. The audio synthesizer according to claim 1, configured to retrieve channel levels and associated information (ξ, χ) among the side information (228) of the downmixed signal (324, 246, x), the audio synthesizer further configured to reconstruct the target version of the covariance information (C ) from an estimated version of the original channel levels and associated information (220) from both of the following y ​ Covariance information (C for at least one first channel or channel pair x ); and The channel levels and correlation information (ξ, χ) for at least one second channel or channel pair.

11. The audio synthesizer according to claim 10, configured to preferably obtain the channel levels and correlation information (ξ, χ) of the described channel or channel pair from the side information (228) of the bitstream (248), rather than preferably the covariance information (Cy) reconstructed from the downmixed signal (324, 246, x) for the same channel or channel pair.

12. The audio synthesizer according to claim 1, wherein the reconstructed target version of the original covariance information (C y ) describes the energy relationship between a pair of channels or is at least partially based on the levels associated with each of the pair of channels. ​ 13. The audio synthesizer according to claim 1, configured to obtain a frequency domain FD version (324) of the downmixed signal (246, x), the FD version (324) of the downmixed signal (246, x) being divided into frequency bands or groups of frequency bands, wherein different channel levels and correlation information (220) are associated with different frequency bands or groups of frequency bands, wherein the audio synthesizer is configured to operate differently for different frequency bands or groups of frequency bands to obtain different mixing rules (403) for different frequency bands or groups of frequency bands.

14. The audio synthesizer according to claim 1, wherein the downmixed signal (324, 246, x) is divided into time slots, wherein different channel levels and correlation information (220) are associated with different time slots, and the audio synthesizer is configured to operate differently for different time slots to obtain different mixing rules (403) for different time slots.

15. The audio synthesizer according to claim 1, wherein the downmixed signal (324, 246, x) is divided into frames, and each frame is divided into time slots, wherein the audio synthesizer is configured to when the presence and location of a transient in a frame are signaled (261) to be in a transient time slot: Associate the current channel level and related information (220) with the transient time slot and / or the time slots subsequent to the transient time slot of the frame; and Associate the time slots prior to the transient time slot of the frame with the channel level and related information (220) of the previous frame.

16. The audio synthesizer according to claim 1, configured to select a prototype rule (Q), the prototype rule (Q) being configured to calculate a prototype signal (328) based on the plurality of synthesized channels.

17. The audio synthesizer according to claim 16, configured to select the prototype rule (Q) among a plurality of pre-stored prototype rules.

18. The audio synthesizer according to claim 1, configured to define a prototype rule (Q) based on a manual selection.

19. The audio synthesizer according to claim 17, wherein the prototype rule includes a matrix (Q), the matrix (Q) having a first dimension and a second dimension, wherein the first dimension is associated with the plurality of downmixed channels, and the second dimension is associated with the plurality of synthesized channels.

20. The audio synthesizer according to claim 1, configured to operate at a bit rate equal to or lower than 160 kbit / s.

21. The audio synthesizer according to claim 1, further comprising an entropy decoder (312) for obtaining the downmixed signal (246, x) with the side information (314).

22. The audio synthesizer according to claim 1, further comprising a decorrelation module (614b, 614c, 330) to reduce the amount of correlation between different channels.

23. The audio synthesizer according to claim 1, wherein the prototype signal (328) is directly provided to the synthesis processor (600a, 600b, 404) without performing decorrelation.

24. The audio synthesizer according to claim 1, wherein at least one of the channel levels and associated information (ξ, χ) of the original signal (212, y), the at least one mixing rule (403), and the covariance information (C x ) associated with the downmixed signal (246, x) is in matrix form.

25. The audio synthesizer according to claim 1, wherein the side information (228) includes an identification of the original channels; wherein the audio synthesizer is further configured to calculate the at least one mixing rule (403) using at least one of the channel levels and associated information (ξ, χ) of the original signal (212, y), covariance information (C x ) associated with the downmixed signal (246, x), the identity of the original channels, and the identity of the synthesized channels.

26. The audio synthesizer according to claim 1, configured to calculate at least one mixing rule by singular value decomposition SVD.

27. The audio synthesizer according to claim 1, wherein the downmixed signal is divided into frames, and the audio synthesizer is configured to smooth the received parameters, estimated or reconstructed values, or mixing matrix using a linear combination of parameters, estimated or reconstructed values, or mixing matrix obtained for a previous frame.

28. The audio synthesizer according to claim 27, configured to deactivate the smoothing of the received parameters, estimated or reconstructed values, or mixing matrix when the presence and / or position of a transient in a frame is signaled (261).

29. The audio synthesizer according to claim 1, wherein the downmixed signal is divided into frames, and the frames are divided into time slots, wherein the channel levels and correlation information (220, ξ, χ) of the original signal (212, y) are obtained from side information (228) of a bitstream (248) on a per-frame basis, and the audio synthesizer is configured to use a mixing rule for a current frame, the mixing rule being obtained by scaling a mixing rule calculated for the current frame by coefficients increasing along subsequent time slots of the current frame and adding a mixing rule for a previous frame in a scaled version scaled by coefficients decreasing along the subsequent time slots of the current frame.

30. The audio synthesizer according to claim 1, wherein the plurality of synthesized channels is greater than the number of original channels.

31. The audio synthesizer according to claim 1, wherein the plurality of synthesized channels is less than the number of original channels.

32. The audio synthesizer according to claim 1, wherein the at least one mixing rule includes a first mixing matrix (M M ) and a second mixing matrix (M R ), and the audio synthesizer Comprising: A first path (610c'), comprising: The first mixing matrix block (600c), configured to synthesize a first component (336M') of the synthesized signal according to the first mixing matrix (M M ) calculated from the following: The covariance matrix associated with the synthesized signal (212) The covariance matrix is reconstructed from the channel levels and associated information (220); and Covariance matrix (C associated with the downmixed signal (324) x ) A second path (610c) for synthesizing a second component (336R') of the synthesized signal, the second component (336R') being a residual component, the second path (610c) comprising: A prototype signal block (612c) configured to upmix the downmixed signal (324) from the plurality of downmixed channels to the plurality of synthesized channels; A decorrelator (614c) configured to decorrelate the upmixed prototype signal (613c); The second mixing matrix block (618c), configured to synthesize the second component (336R') of the synthesized signal from the decorrelated version (615c) of the downmixed signal (324) according to a second mixing matrix (M R ) where the second mixing matrix (M R ) is a residual mixing matrix, wherein the audio synthesizer (300) is configured to estimate (618c) the second mixing matrix (M R ) from the following: The residual covariance matrix (C provided by the first mixing matrix block (600c) r ); and An estimate of the covariance matrix of the decorrelated prototype signal obtained from the covariance matrix (C x ) associated with the downmixed signal (324) ​ Wherein the audio synthesizer (300) further comprises an adder block (620c) for summing the first component (336M') of the synthesized signal and the second component (336R') of the synthesized signal.

33. A method for generating a synthesized signal from a downmixed signal, the synthesized signal having a plurality of synthesized channels, the method Comprising: Receiving a downmixed signal (246, x) and side information (228), the downmixed signal (246, x) having a plurality of downmixed channels, the side information (228) comprising: Channel levels and correlation information of an original signal (212, y), the original signal (212, y) having a plurality of original channels; and Covariance information (C associated with the downmixed signal (324, 246, x) x ); Using the channel levels and associated information (220) of the original signal (212, y) and covariance information (C) associated with the signal (246, x) according to at least one mixing rule that is a mixing matrix to generate the synthesized signal; the method includes: reconstructing (386) a target version of the covariance information (C) of the original signal based on an estimated version of the original covariance information (C x ) y ) wherein the estimated version of the original covariance information (C y ) is reported to the plurality of synthesized channels; y and the target version of the covariance information (C) of the original signal is reconstructed based on the estimated version of the original covariance information (C ) The method includes: obtaining an estimated version of the original covariance information from covariance information (C x ) associated with the downmixed signal (324, 246, x) by applying an estimation rule (Q) to the covariance information (C x ) associated with the downmixed signal (324, 246, x) The estimation rule (Q) is a prototype rule for calculating a prototype signal (326) or is associated with a prototype rule for calculating a prototype signal (326).

34. The method according to claim 33, the method Comprising: Calculating a prototype signal from the downmixed signal (246, x), the prototype signal having the plurality of synthesized channels; Calculating a mixing rule using the channel levels and correlation information of the original signal (212, y) and covariance information associated with the downmixed signal (246, x); and Generating the synthesized signal using the prototype signal and the mixing rule.

35. A non-transitory storage unit storing instructions which, when executed by a processor, cause the processor to execute the method according to claim 33.

Citation Information

Patent Citations

  • Method and arrangement for a decoder for multi-channel surround sound

    US20090110203A1