Encoding and Decoding of Parameters
The audio synthesizer addresses the limitations of existing methods by using channel levels, correlation information, and covariance information to efficiently encode and decode multi-channel audio within the DirAC framework, achieving high-quality output and flexibility at low bitrates.
Patent Information
- Application Number
- JP2023215842
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-14
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2040-06-15
AI Technical Summary
Existing methods for encoding and decoding multi-channel audio content at low bit rates, such as DirAC, are not well-suited for multi-channel signals due to the use of parameters like DOA and diffuseness, which result in quality degradation and lack of flexibility.
An audio synthesizer that generates a synthesized signal from a downmix signal using channel levels and correlation information of the original signal, along with covariance information, to efficiently describe multi-channel audio signals within the DirAC framework, while operating at low bitrates and maintaining flexibility.
The proposed solution effectively reconstructs target covariance information adapted to the number of channels of the synthesis signal, ensuring high-quality output and flexibility across various loudspeaker settings, even at low bitrates.
Smart Images

Figure 0007696986000191 
Figure 0007696986000192 
Figure 0007696986000193
Abstract
Description
Technical Field
[0001] 1. Introduction Here, some examples of encoding and decoding techniques are disclosed. Specifically, for example, an invention for encoding and decoding multi-channel audio content at a low bit rate using the DirAC framework. This method enables obtaining high-quality output while using a low bit rate. This can be used in many applications including artworks, communications, and virtual reality.
Background Art
[0002] 1.1. Prior Art In this section, the prior art will be briefly described.
[0003] 1.1.1 Discrete Coding of Multi-Channel Content The simplest approach for coding and transmitting multi-channel content is to directly quantize and encode the waveforms of the multi-channel audio signals without any preprocessing or assumptions. This method theoretically functions perfectly, but it has one major drawback in that it requires a large bit consumption to encode multi-channel content. Therefore, the other methods described (and the proposed invention) are so-called "parametric approaches" that use meta-parameters to describe and transmit multi-channel audio signals instead of the original audio multi-channel signals themselves.
[0004] 1.1.2 MPEG Surround MPEG Surround is an ISO / MPEG standard for parametric coding of multi-channel sound finalized in 2006 [1]. This method mainly depends on two parameter sets. - Interchannel coherence (ICC), which represents the coherence between all channels of a given multi-channel audio signal. - Channel Level Difference (CLD), which corresponds to the level difference between two input channels of a multi-channel audio signal.
[0005] One of the special features of MPEG Surround is the use of the so-called "tree structure", which enables the description of "two input channels using a single output channel" (quoted from [1]). As an example, the following shows how to find the encoder method for a 5.1 multi-channel audio signal using MPEG Surround. In this figure, six input channels (denoted as "L", "L S ", "R", "R S ", "C", and "LFE") are processed sequentially through tree structure elements (denoted as "R_OTT" in the figure). Each of these tree structure elements creates a set of parameters (the aforementioned ICC and CLD) and a residual signal, and these sets of parameters and the residual signal are processed again through another tree structure to generate another set of parameters. When the end of the tree is reached, various parameters calculated so far are sent to the decoder, similar to the downmixed signal. These elements are used by the decoder to generate the output multi-channel signal, and the decoder processing is basically a tree structure opposite to the one used by the encoder.
[0006] The main strength of MPEG Surround depends on the use of this structure and the use of the aforementioned parameters. However, one of the drawbacks of MPEG Surround is the lack of flexibility due to the tree structure. Also, due to the specificity of the processing, there may be a quality degradation in some specific items.
[0007] In particular, refer to FIG. 7 showing an overview of an MPEG surround encoder for 5.1 signals excerpted from [1].
[0008] 1.2. Directional Audio Coding Directional Audio Coding (abbreviated as "DirAC: Directional Audio Coding") [2] is also a parametric method for reproducing spatial audio and was developed by Ville Pulkki of Aalto University in Finland. DirAC relies on frequency band processing that describes spatial sound using two parameter sets. - Direction Of Arrival (DOA), which is the angle in degrees representing the arrival direction of the main sound in the audio signal. - Diffuseness, which is a value between 0 and 1 representing how much the sound "diffuses". When the value is 0, the sound has no diffuseness and can be captured as a point sound source arriving from an exact angle. When the value is 1, the sound is assumed to be sufficiently diffuse and arrive from "all" angles.
[0009] In DirAC, it is assumed that the sound is decomposed into a diffuse part and a non-diffuse part in order to synthesize the output signal. Diffuse sound synthesis aims to create the perception of the surrounding sound, and direct sound synthesis aims to generate the main sound.
[0010] DirAC provides high-quality output but has one major drawback. It was not targeted at multi-channel audio signals. Therefore, the DOA and diffuseness parameters are not very suitable for describing multi-channel audio inputs, and as a result, the quality of the output is affected.
[0011] 1.3. Binaural Cue Coding Binaural Cue Coding (BCC) [3] is a parametric technique developed by Christof Faller. This method relies on a parameter set similar to that described for MPEG Surround (see 1.1.2). - Interchannel Level Difference (ICLD), a measure of the energy ratio between two channels of a multi-channel input signal. - Interchannel Time Difference (ICTD), a measure of the delay between two channels of a multi-channel input signal. - Interchannel Correlation (ICC), a measure of the correlation between two channels of a multi-channel input signal.
[0012] The BCC technique has very similar characteristics regarding the calculation of the parameters to be transmitted compared to the novel invention described later, but the flexibility and scalability of the transmitted parameters are not sufficient.
[0013] 1.4. MPEG Spatial Audio Object Coding Here, Spatial Audio Object Coding [4] is briefly described. Spatial Audio Object Coding is an MPEG standard for coding so-called audio objects, which are somewhat related to multi-channel signals. Spatial Audio Object Coding uses parameters similar to those of MPEG Surround.
[0014] 1.5 Motivation / Disadvantages of the Prior Art 1.5.1 Motivation 1.5.1.1 Using the DirAC Framework One aspect of the present invention that must be mentioned is that the present invention must conform to the DirAC framework. Nevertheless, it has also been mentioned above that the parameters of DirAC are not suitable for multi-channel audio signals. This topic will be further explained.
[0015] The original DirAC process uses either a microphone signal or an ambisonics signal. From these signals, parameters, namely the direction of arrival (DOA) and diffuseness, are calculated.
[0016] One of the first approaches tried to use DirAC with multi-channel audio signals was to convert the multi-channel signal into ambisonics content using the method proposed by Ville Pulkki described in [5]. Then, once these ambisonic signals were derived from the multi-channel audio signal, the normal DirAC process was performed using the DOA and diffuseness. The results of this first trial were that the quality and spatial characteristics of the output multi-channel signal deteriorated and did not meet the requirements of the target application.
[0017] Therefore, the main motivation behind this new invention is to use a parameter set that efficiently describes multi-channel signals and to use the DirAC framework. Details will be explained in Section 1.1.2.
[0018] 1.5.1.2 Provide a system that operates at low bitrates One of the goals and objectives of the present invention is to propose a method that enables low-bitrate applications. This method requires finding an optimal dataset for describing multi-channel content between the encoder and the decoder. This method also requires finding an optimal trade-off from the perspective of the number of parameters transmitted and the output quality.
[0019] 1.5.1.3 Provide a flexible system Another important objective of the present invention is to propose a flexible system that can tolerate any multi - channel audio format intended to be reproduced in any loudspeaker setting. The output quality should not be impaired according to the input settings.
[0020] 1.5.2 Disadvantages of the prior art The prior art described above as some disadvantages are listed in the following Table (Table 1).
[0021]
Table 1
Prior art documents
Non - patent documents
[0022]
Non - patent document 1
Non - patent document 2
Non - patent document 3
Non-Patent Document 8
Non-Patent Document 9
Summary of the Invention
Means for Solving the Problems
[0023] 2. Description of the Invention 2.1 Summary of the Invention According to one aspect, an audio synthesizer (encoder) for generating a synthesized signal from a downmix signal, the synthesized signal having a plurality of synthesis channels, the audio synthesizer comprising an input interface configured to receive the downmix signal, the downmix signal having a plurality of downmix channels and side information, the side information including channel levels and correlation information of the original signal, the original signal having a plurality of original channels, the input interface; and channel levels and correlation information of the original signal, and covariance information related to the downmix signal a synthesis processor configured to generate the synthesized signal according to at least one mixing rule using is provided. The audio synthesizer includes
[0024] The audio synthesizer A prototype signal calculator configured to calculate a prototype signal from a downmix signal, the prototype signal having a plurality of synthesis channels, the prototype signal calculator; The channel levels and correlation information of the original signal, and Covariance information related to the downmix signal may include a mixing rule calculator (402) configured to calculate at least one mixing rule using; The synthesis processor is configured to generate a synthesis signal using the prototype signal and at least one mixing rule.
[0025] The audio synthesizer may be configured to reconstruct the target covariance information of the original signal.
[0026] The audio synthesizer may be configured to reconstruct target covariance information adapted to the number of channels of the synthesis signal.
[0027] The audio synthesizer may reconstruct covariance information (C y ) adapted to the number of channels of the synthesis signal by assigning a group of original channels to a single synthesis channel or vice versa, such that the reconstructed target covariance information is reported to several channels of the synthesis signal.
[0028] The audio synthesizer may be configured to reconstruct covariance information adapted to the number of channels of the synthesis signal by generating target covariance information for several original channels and subsequently applying a downmixing or upmixing rule and energy compensation to reach the target covariance of the synthesis channels.
[0029] The audio synthesizer may be configured to reconstruct a target version of the covariance information based on an estimated version of the original covariance information, the estimated version of the original covariance information being reported to several synthesis channels or several original channels.
[0030] The audio synthesizer may be configured to obtain an estimated version of the original covariance information from the covariance information associated with the downmix signal.
[0031] The audio synthesizer may be configured to obtain an estimated version of the original covariance information by applying an estimated rule associated with a prototype rule for calculating a prototype signal to the covariance information associated with the downmix signal.
[0032] The audio synthesizer, for at least one pair of channels, may be configured to normalize an estimated version of the original covariance information (C y ) by the square root of the level of the channel among the pair of channels.
[0033]
Number
[0034] The audio synthesizer may be configured to interpret a matrix having the normalized estimated version of the original covariance information.
[0035] The audio synthesizer may be configured to complete the matrix by inserting an entry obtained in the side information of the bitstream.
[0036] The audio synthesizer may be configured to unnormalize the matrix by scaling the estimated version of the original covariance information by the square root of the level of the channel forming the pair of channels.
[0037] The audio synthesizer may be configured to search from among the side information of the downmix signal, and the audio synthesizer
[0038] with the covariance information of at least one first channel or a pair of channels, and from both the channel levels and correlation information of at least one second channel or pair of channels and is further configured to reconstruct a target version of the covariance information by both estimated versions of the original channel levels and correlation information.
[0039] The audio synthesizer may be configured to prioritize the channel levels and correlation information describing the channels or pair of channels obtained from the side information of the bitstream over the covariance information reconstructed from the downmix signals of the same channels or pair of channels.
[0040] The reconstructed target version of the original covariance information may be understood to describe the energy relationship between a pair of channels and is at least partially based on the levels associated with each channel of the pair of channels.
[0041] The audio synthesizer may be configured to obtain a frequency domain FD version of the downmix signal, the FD version of the downmix signal is divided into bands or groups of bands, and different channel levels and correlation information are associated with different bands or groups of bands, The audio synthesizer is configured to operate in different ways for different bands or groups of bands to obtain different mixing rules for different bands or groups of bands.
[0042] The downmix signal is divided into slots, different channel levels and correlation information are associated with different slots, and the audio synthesizer is configured to operate in different ways for different slots to obtain different mixing rules for different slots.
[0043] The downmix signal is divided into frames, each frame is divided into slots, and when the presence and position of a transient in one frame is signaled as being in one transient slot, the audio synthesizer Associate the current channel level and correlation information with the transient slots and / or the slots following the transient slots of the frame. It is configured to associate the channel level and correlation information of the preceding slot with the slot of the frame preceding the transient slot.
[0044] The audio synthesizer may be configured to select a prototype rule configured to calculate a prototype signal based on the number of synthesis channels.
[0045] The audio synthesizer may be configured to select a prototype rule from among a plurality of pre-stored prototype rules.
[0046] The audio synthesizer may be configured to define a prototype rule based on a manual selection.
[0047] The prototype rule may be based on or include a matrix having a first dimension and a second dimension, the first dimension being associated with the number of downmix channels and the second dimension being associated with the number of synthesis channels.
[0048] The audio synthesizer may be configured to operate at a bit rate of 160 kbit / s or less.
[0049] The audio synthesizer may further include an entropy decoder for obtaining a downmix signal having side information.
[0050] The audio synthesizer further includes a decorrelation module for reducing the amount of correlation between different channels.
[0051] The prototype signal may be provided directly to the synthesis processor without performing decorrelation.
[0052] At least one of the channel level and correlation information of the original signal, at least one mixing rule, and the covariance information related to the downmix signal is in matrix form.
[0053] The side information includes the identification information of the original channel. The audio synthesizer may be further configured to calculate at least one mixing rule using at least one of the channel level and correlation information of the original signal, the covariance information related to the downmix signal, the identification information of the original channel, and the identification information of the synthesis channel.
[0054] The audio synthesizer may be configured to calculate at least one mixing rule by singular value decomposition (SVD).
[0055] The downmix signal may be divided into frames, and the audio synthesizer is configured to smooth the received parameters, or estimated or reconstructed values, or the mixing matrix using the parameters obtained for the previous frame, or estimated or reconstructed values, or a linear combination with the mixing matrix.
[0056] When the presence and / or position of a transient phenomenon in one frame is signaled, the audio synthesizer may be configured to disable the smoothing of the received parameters, or estimated or reconstructed values, or the mixing matrix.
[0057] The downmix signal can be divided into frames, the frames can be divided into slots, the channel levels and correlation information of the original signal are obtained in a per-frame manner from the side information of the bitstream, and the audio synthesizer scales the mixing matrix (or mixing rule) calculated for the current frame by a factor that increases along the subsequent slots of the current frame, and adds a version of the mixing matrix (or mixing rule) used for the previous frame, scaled by a factor that decreases along the subsequent slots of the current frame, to obtain the mixing rule used for the current frame.
[0058] The number of synthesized channels may be greater than the number of original channels. The number of synthesized channels may be less than the number of original channels. The number of synthesized channels and the number of original channels may be greater than the number of downmix channels.
[0059] At least one or all of the number of synthesized channels, the number of original channels, and the number of downmix channels are plural.
[0060] At least one mixing rule may include a first mixing matrix and a second mixing matrix, and the audio synthesizer a covariance matrix related to the synthesized signal reconstructed from the channel levels and correlation information, and a covariance matrix related to the downmix signal a first mixing matrix block configured to synthesize a first component of the synthesized signal according to the first mixing matrix calculated therefrom including a first path, a second path for synthesizing a second component of the synthesized signal, where the second component is a residual component, and the second path is a prototype signal block configured to upmix the downmix signal from the number of downmix channels to the number of synthesized channels, a decorrelator configured to decorrelate the upmixed prototype signal A second mixing matrix block configured to synthesize a second component of a synthesized signal from an uncorrelated version of a downmix signal according to a second mixing matrix, wherein the second mixing matrix is a residual mixing matrix including, a second path comprising, an audio synthesizer a residual covariance matrix provided by a first mixing matrix block, and an estimated value of a covariance matrix of an uncorrelated prototype signal obtained from a covariance matrix associated with a downmix signal configured to estimate a second mixing matrix from The audio synthesizer further comprises an adder block for summing a first component of the synthesized signal with the second component of the synthesized signal.
[0061] According to one aspect, an audio synthesizer for generating a synthesized signal from a downmix signal having a plurality of downmix channels, wherein the synthesized signal has a plurality of synthesis channels, the downmix signal being a downmixed version of an original signal having a plurality of original channels, the audio synthesizer comprising a covariance matrix associated with the synthesized signal reconstructed from channel level and correlation information, and a covariance matrix associated with the downmix signal a first mixing matrix block configured to synthesize a first component of the synthesized signal according to a first mixing matrix calculated from including, a first path a second path for synthesizing a second component of the synthesized signal, the second component being a residual component, the second path comprising a prototype signal block configured to upmix the downmix signal from the number of downmix channels to the number of synthesis channels an uncorrelator configured to decorrelate the upmixed prototype signal A second mixing matrix block configured to synthesize a second component of a synthesized signal from an uncorrelated version of a downmix signal according to a second mixing matrix, wherein the second mixing matrix is a residual mixing matrix including a second path and comprising an audio synthesizer, a residual covariance matrix provided by a first mixing matrix block, and an estimated value of the covariance matrix of an uncorrelated prototype signal obtained from the covariance matrix associated with the downmix signal configured to calculate a second mixing matrix therefrom, wherein the audio synthesizer may be provided with an adder block for summing a first component of the synthesized signal with the second component of the synthesized signal.
[0062] The residual covariance matrix is obtained by subtracting from the covariance matrix associated with the synthesized signal a matrix obtained by applying a first mixing matrix to the covariance matrix associated with the downmix signal.
[0063] The audio synthesizer a second matrix obtained by decomposing a residual covariance matrix associated with the synthesized signal, a first matrix that is the inverse or regularized inverse of a diagonal matrix obtained from an estimated value of the covariance matrix of the uncorrelated prototype signal may be configured to define a second mixing matrix therefrom.
[0064] The diagonal matrix may be obtained by applying a square root function to the main diagonal elements of the covariance matrix of the uncorrelated prototype signal.
[0065] The second matrix may be obtained by applying a singular value decomposition SVD applied to the residual covariance matrix associated with the synthesized signal.
[0066] The audio synthesizer may be configured to define a second mixing matrix by multiplying a diagonal matrix or a regularized inverse matrix obtained from an estimated value of the covariance matrix of the uncorrelated prototype signals and a third matrix by a second matrix.
[0067] The audio synthesizer may be configured to obtain a third matrix by applying SVP to a matrix obtained from a normalized version of the covariance matrix of the uncorrelated prototype signals, where the normalization is performed on the main diagonal, the residual covariance matrix, and the diagonal matrix and the second matrix.
[0068] The audio synthesizer may be configured to define a first mixing matrix from a second matrix and an inverse or regularized inverse of the second matrix. The second matrix is obtained by decomposing the covariance matrix associated with the downmixed signal. The second matrix is obtained by decomposing the reconstructed target covariance matrix associated with the downmixed signal.
[0069] The audio synthesizer may be configured to estimate the covariance matrix of the uncorrelated prototype signals from the diagonal entries of a matrix obtained by applying to the covariance matrix associated with the downmixed signal a prototype rule used to upmix the downmixed signal from the number of downmix channels to the number of synthesis channels in a prototype block.
[0070] The bands are aggregated with each other to form groups of aggregated bands, information regarding the groups of aggregated bands is provided in the side information of the bitstream, and the channel level and correlation information of the original signal are provided for each group of bands such that at least one same mixing matrix is calculated for different bands of the same aggregation group of bands.
[0071] According to one aspect, an audio encoder for generating a downmix signal from an original signal, wherein the original signal has a plurality of original channels, the downmix signal has a plurality of downmix channels, and the audio encoder comprises: a parameter estimator configured to estimate channel levels and correlation information of the original signal; a bitstream writer for encoding the downmix signal into a bitstream such that the downmix signal has side information including the channel levels and correlation information of the original signal; An audio encoder may be provided.
[0072] The audio encoder may be configured to provide the channel levels and correlation information of the original signal as normalized values.
[0073] The channel levels and correlation information of the original signal encoded in the side information represent at least channel level information related to the entirety of the original channels.
[0074] The channel levels and correlation information of the original signal encoded in the side information represent at least correlation information describing an energy relationship between at least one pair of different original channels but less than all of the original channels.
[0075] The channel levels and correlation information of the original signal include at least one coherence value describing the coherence between two channels of a pair of original channels.
[0076] The coherence value may be normalized. The coherence value is
[0077]
Equation
[0078] and may be, where in the formula
[0079] [Mathematics]
[0080] is the covariance between channel i and channel j,
[0081] [Mathematics]
[0082] and
[0083] [Mathematics]
[0084] are the levels related to channel i and channel j, respectively.
[0085] The channel levels and correlation information of the original signal include at least one inter-channel level difference ICLD.
[0086] At least one ICLD can be provided as a logarithmic value. At least one ICLD can be normalized. The ICLD is
[0087] [Mathematics]
[0088] can be, where - χ i is the ICLD of channel i, - P i is the power of the current channel i, - P dmx,i is a linear combination of the values of the covariance information of the downmixed signal.
[0089] The audio encoder may be configured to select whether to encode at least a portion of the channel level and correlation information of the original signal based on status information so as to include the increased amount of channel level and correlation information in the side information when the payload is relatively low.
[0090] The audio encoder may be configured to select which portion of the channel level and correlation information of the original signal to encode in the side information based on a metric on the channel so as to include the channel level and correlation information related to the more susceptible metric in the side information.
[0091] The channel level and correlation information of the original signal may be in the form of entries of a matrix.
[0092] The matrix may be a symmetric matrix or a Hermitian matrix, and entries of the channel level and correlation information are provided for all or less than all of the entries on the diagonal of the matrix and / or less than half of the off-diagonal elements of the matrix.
[0093] The bitstream writer may be configured to encode the identification of at least one channel.
[0094] The original signal or its processed version may be divided into a plurality of subsequent frames of equal time duration.
[0095] The audio encoder may be configured to encode the channel level and correlation information of the original signal specific to each frame in the side information.
[0096] The audio encoder may be configured to encode the same channel level and correlation information of the original signal collectively associated with a plurality of consecutive frames in the side information.
[0097] The audio encoder is such that a relatively high bit rate or payload means an increase in the number of consecutive frames to which the same channel level and correlation information of the original signal are associated, and vice versa, It may be configured to select the number of consecutive frames from which the same channel level and correlation information of the original signal can be selected.
[0098] The audio encoder may be configured to reduce the number of consecutive frames to which the same channel level and correlation information of the original signal are associated when detecting a transient phenomenon.
[0099] Each frame may be subdivided into an integer number of consecutive slots.
[0100] The audio encoder may be configured to estimate the channel level and correlation information for each slot and encode, within the side information, the sum or average or another predefined linear combination of the channel levels and correlation information estimated for different slots.
[0101] The audio encoder may be configured to perform transient analysis on the time-domain version of the frame to determine the occurrence of a transient phenomenon within the frame.
[0102] The audio encoder determines in which slot of the frame a transient phenomenon has occurred, without encoding the channel level and correlation information of the original signal associated with the slot preceding the transient phenomenon, and may be configured to encode the channel level and correlation information of the original signal associated with the slot in which the transient phenomenon has occurred and / or subsequent slots within the frame.
[0103] The audio encoder may be configured to signal in the side information the occurrence of a transient phenomenon within one slot of the frame.
[0104] The audio encoder may be configured to signal in side information in which slot of a frame a transient has occurred.
[0105] The audio encoder may be configured to estimate channel levels and correlation information of an original signal related to a plurality of slots of a frame, and sum them, or average them, or linearly combine them to obtain channel levels and correlation information related to the frame.
[0106] The original signal may be converted into a frequency-domain signal, and the audio encoder is configured to encode channel levels and correlation information of the original signal in side information in a per-band manner.
[0107] The audio encoder may be configured to aggregate some bands of the original signal into a smaller number of bands so as to encode channel levels and correlation information of the original signal in side information in an aggregated-band manner.
[0108] When a transient is detected within a frame, the audio encoder is configured to further aggregate bands such that the number of bands is reduced and / or the width of at least one band is increased by aggregation with another band.
[0109] The audio encoder may be further configured to encode at least one channel level and correlation information of one band in a bitstream as an increment with respect to previously encoded channel levels and correlation information.
[0110] The audio encoder may be configured to encode an incomplete version of channel levels and correlation information in side information of the bitstream as compared to channel levels and correlation information estimated by an estimator.
[0111] The audio encoder may be configured to adaptively select the selected information to be encoded in the side information of the bitstream from among the channel levels and correlation information estimated by the estimator, such that the remaining unselected information channel levels and / or correlation information estimated by the estimator are not encoded.
[0112] The audio encoder reconstructs the channel levels and correlation information from the selected channel levels and correlation information, thereby simulating the estimated values of the unselected channel levels and correlation information in the decoder, the unselected channel levels and correlation information estimated by the encoder, and the unselected channel levels and correlation information reconstructed by simulating the estimated values of the unencoded channel levels and correlation information in the decoder to calculate the error information therebetween, such that based on the calculated error information, distinguish between appropriately reconstructable channel levels and correlation information and inappropriately reconstructable channel levels and correlation information, such that decide on the selection of inappropriately reconstructable channel levels and correlation information to be encoded in the side information of the bitstream and the non-selection of appropriately reconstructable channel levels and correlation information, thereby being configured not to encode the appropriately reconstructable channel levels and correlation information in the side information of the bitstream.
[0113] Channel levels and correlation information may be indexed according to a default order, and the encoder is configured to signal in the side information of the bitstream an index associated with the default order, the index indicating which of the channel levels and correlation information is encoded. The index is provided via a bitmap. The index may be defined according to a combinatorial number system that associates one-dimensional indexes with the entries of a matrix.
[0114] The audio encoder is configured to perform a selection between an adaptive provision of channel levels and correlation information, in which an index associated with a default order is encoded within the side information of the bitstream, and a fixed provision of channel levels and correlation information, in which the channel levels and correlation information to be encoded are predetermined and ordered according to a default fixed order without the provision of an index.
[0115] The audio encoder may be further configured to signal in the side information of the bitstream whether the channel levels and correlation information are provided according to an adaptive provision or a fixed provision.
[0116] The audio encoder may be further configured to encode in the bitstream the current channel levels and correlation information as an increment relative to the previous channel levels and correlation information.
[0117] The audio encoder may be further configured to generate a downmix signal according to a static downmixing.
[0118] According to one aspect, a method for generating a composite signal from a downmix signal, the composite signal having a plurality of composite channels, the method comprising Receiving a downmix signal, the downmix signal having a plurality of downmix channels and side information, the side information being the channel levels and correlation information of the original signal and the original signal having a plurality of original channels; and generating a composite signal using the channel levels and correlation information (220) of the original signal and the covariance information related to the signal. A method is provided that includes.
[0119] The method is calculating a prototype signal from the downmix signal, the prototype signal having a plurality of composite channels; calculating a mixing rule using the channel levels and correlation information of the original signal and the covariance information related to the downmix signal; and generating a composite signal using the prototype signal and the mixing rule. It may include.
[0120] According to one aspect, a method for generating a downmix signal from an original signal, the original signal having a plurality of original channels, the downmix signal having a plurality of downmix channels, the method comprising: estimating the channel levels and correlation information of the original signal; and encoding the downmix signal into a bitstream such that the downmix signal has side information including the channel levels and correlation information of the original signal. A method is provided that includes.
[0121] According to one aspect, a method for generating a composite signal from a downmix signal having a plurality of downmix channels, the composite signal having a plurality of composite channels, the downmix signal being a downmixed version of an original signal having a plurality of original channels, the method comprising the following phases, namely The covariance matrix associated with the synthesized signal, and The covariance matrix associated with the downmix signal Synthesizing a first component of the synthesized signal according to a first mixing matrix calculated from Including a first phase, A second phase for synthesizing a second component of the synthesized signal, where the second component is a residual component, and the second phase is A prototype signal step of upmixing the downmix signal from the number of downmix channels to the number of synthesis channels, An uncorrelator step of decorrelating the upmixed prototype signal, A second mixing matrix step of synthesizing a second component of the synthesized signal from the uncorrelated version of the downmix signal according to a second mixing matrix, where the second mixing matrix is a residual mixing matrix, the second mixing matrix step Including a second phase, and Including, the method being The residual covariance matrix provided by the first mixing matrix step, and An estimated value of the covariance matrix of the uncorrelated prototype signal obtained from the covariance matrix associated with the downmix signal Calculating a second mixing matrix from, The method further includes an adder step of summing the first component of the synthesized signal with the second component of the synthesized signal, thereby obtaining the synthesized signal. A method is provided.
[0122] According to one aspect, an audio synthesizer for generating a synthesized signal from a downmix signal, where the synthesized signal has a plurality of synthesis channels, the number of synthesis channels being greater than 1 or greater than 2, and the audio synthesizer is An input interface configured to receive the downmix signal, where the downmix signal has at least one downmix channel and side information, and the side information is Channel levels and correlation information of the original signal, where the original signal has a plurality of original channels and the number of original channels is greater than 1 or greater than 2, the channel levels and correlation information an input interface including at least one of a part such as a prototype signal calculator [for example, "prototype signal calculation"] configured to calculate a prototype signal from a downmix signal, where the prototype signal has a plurality of composite channels a part such as a mixing rule calculator [for example, "parameter reconstruction"] configured to calculate one (or more) mixing rules [for example, mixing matrix] using the channel levels and correlation information of the original signal and the covariance information related to the downmix signal a part such as a synthesis processor [for example, "synthesis engine"] configured to generate a composite signal using the prototype signal and the mixing rule An audio synthesizer is provided that includes at least one of these.
[0123] The number of composite channels may be greater than the number of original channels. Alternatively, the number of composite channels may be less than the number of original channels.
[0124] The audio synthesizer (specifically, in some embodiments, the mixing rule calculator) may be configured to reconstruct a target version of the original channel levels and correlation information.
[0125] The audio synthesizer (specifically, in some embodiments, the mixing rule calculator) may be configured to reconstruct a target version of the original channel levels and correlation information adapted to the number of channels of the composite signal.
[0126] The audio synthesizer (specifically, in some embodiments, the mixing rule calculator) may be configured to reconstruct a target version of the original channel levels and correlation information based on an estimated version of the original channel levels and correlation information.
[0127] An audio synthesizer (specifically, in some embodiments, a mixing rule calculator) may be configured to obtain an estimated version of the original channel levels and correlation information from the covariance information associated with the downmix signal.
[0128] An audio synthesizer (specifically, in some embodiments, a mixing rule calculator) may be configured to obtain an estimated version of the original channel levels and correlation information by applying an estimated rule related to a prototype rule used by a prototype signal calculator [e.g., "prototype signal calculation"] to calculate a prototype signal, to the covariance information associated with the downmix signal.
[0129] An audio synthesizer (specifically, in some embodiments, a mixing rule calculator) searches among the side information of the downmix signal for both the covariance information associated with the downmix signal that describes the level of the first channel within the downmix signal, or the energy relationship between a pair of channels, and the channel levels and correlation information of the original signal that describes the level of the first channel within the original signal, or the energy relationship between a pair of channels and, as a result, the covariance information of the original channels of at least one first channel or pair of channels, and the channel levels and correlation information that describe at least one second channel or pair of channels and may be configured to reconstruct a target version of the original channel levels and correlation information by using at least one of them.
[0130] An audio synthesizer (specifically, in some embodiments, a mixing rule calculator) may be configured to prioritize the channel levels and correlation information that describe a channel or pair of channels over the covariance information of the original channels of the same channel or pair of channels.
[0131] The original channel levels and the reconstructed target versions of the correlation information that describe the energy relationship between a pair of channels are based at least in part on the levels associated with each channel of the pair of channels.
[0132] The downmix signal can be divided into bands or groups of bands, different channel levels and correlation information can be associated with different bands or groups of bands, and a synthesizer (a prototype signal calculator, specifically, in some embodiments, at least one of a mixing rule calculator and a synthesis processor) operates in different ways for different bands or groups of bands to obtain different mixing rules for different bands or groups of bands.
[0133] The downmix signal can be divided into slots, different channel levels and correlation information are associated with different slots, and at least one of the components of the synthesizer (e.g., a prototype signal calculator, a mixing rule calculator, a synthesis processor, or other elements of the synthesizer) operates in different ways for different slots to obtain different mixing rules for different slots.
[0134] A synthesizer (e.g., a prototype signal calculator) can be configured to select a prototype rule configured to calculate a prototype signal based on the number of synthesis channels.
[0135] A synthesizer (e.g., a prototype signal calculator) can be configured to select a prototype rule from among a plurality of pre-stored prototype rules.
[0136] A synthesizer (e.g., a prototype signal calculator) can be configured to define a prototype rule based on a manual selection.
[0137] A synthesizer (e.g., a prototype signal calculator) may include a matrix having a first dimension and a second dimension, where the first dimension is associated with the number of downmix channels and the second dimension is associated with the number of synthesis channels.
[0138] An audio synthesizer (e.g., a prototype signal calculator) may be configured to operate at a bit rate of 64 kbit / s or less than 160 Kbit / s.
[0139] Side information may include identification information of the original channels [e.g., L, R, C, etc.].
[0140] An audio synthesizer (specifically, in some embodiments, a mixing rule calculator) may be configured to calculate [“parameter reconstruction”] a mixing rule [e.g., a mixing matrix] using the channel levels and correlation information of the original signal, the covariance information related to the downmix signal, and the identification of the original channels and the identification of the synthesis channels.
[0141] The audio synthesizer may select some channels for the synthesized signal [e.g., by selection such as manual selection, or by pre-selection, or automatically by recognizing, for example, the number of loudspeakers] regardless of at least one of the channel levels and correlation information of the original signal in the side information.
[0142] In some examples, the audio synthesizer may select different prototype rules for different selections. The mixing rule calculator may be configured to calculate the mixing rule.
[0143] According to one aspect, a method for generating a synthesized signal from a downmix signal, where the synthesized signal has a plurality of synthesis channels and the number of synthesis channels is greater than 1 or greater than 2, the method includes receiving the downmix signal, where the downmix signal has at least one downmix channel and side information, and the side information includes Channel levels and correlation information of an original signal, the original signal having a plurality of original channels, the number of original channels being greater than 1 or greater than 2, the channel levels and correlation information including steps, and a step of calculating a prototype signal from the downmix signal, the prototype signal having a plurality of composite channels, the step a step of calculating a mixing rule using the channel levels and correlation information of the original signal and the covariance information related to the downmix signal; and a step of generating a composite signal using the prototype signal and the mixing rule [e.g., rule] A method is provided that includes.
[0144] According to one aspect, an audio encoder for generating a downmix signal from an original signal [e.g., y], the original signal having at least two channels, the downmix signal having at least one downmix channel, the audio encoder comprising a parameter estimator configured to estimate channel levels and correlation information of the original signal, and a bitstream writer for encoding the downmix signal into a bitstream such that the downmix signal has side information including the channel levels and correlation information of the original signal An audio encoder is provided that includes at least one of.
[0145] The channel levels and correlation information of the original signal encoded in the side information represent channel level information related to channels fewer than all of the channels of the original signal.
[0146] The channel levels and correlation information of the original signal encoded in the side information represent correlation information that describes an energy relationship between at least one pair of different original channels in the original signal but fewer than all of the channels of the original signal.
[0147] The channel levels and correlation information of the original signal may include at least one coherence value that describes the coherence between two channels of a pair of channels.
[0148] The channel levels and correlation information of the original signal may include at least one inter-channel level difference ICLD between two channels of a pair of channels.
[0149] The audio encoder may be configured to select whether to encode at least a portion of the channel levels and correlation information of the original signal based on status information so as to include an increased amount of the channel levels and correlation information in the side information when the overload is relatively low.
[0150] The audio encoder may be configured to select whether to determine which portion of the channel levels and correlation information of the original signal to encode within the side information based on a metric on the channel so as to include the channel levels and correlation information related to a more susceptible metric [e.g., a metric related to a more perceptually significant covariance] in the side information.
[0151] The channel levels and correlation information of the original signal may be in matrix form.
[0152] The bitstream writer may be configured to encode the identification of at least one channel.
[0153] According to one aspect, a method for generating a downmix signal from an original signal is provided, the original signal having at least two channels and the downmix signal having at least one downmix channel.
[0154] The method includes estimating the channel levels and correlation information of the original signal, and Encoding the downmix signal into a bitstream such that the downmix signal has side information including the channel levels and correlation information of the original signal may be included.
[0155] The audio encoder may be unaware of the decoder. The audio synthesizer may be unaware of the decoder.
[0156] According to one aspect, a system is provided that includes the above or below audio synthesizer and the above or below audio encoder.
[0157] According to one aspect, a non-transitory storage unit is provided that stores instructions that, when executed by a processor, cause the processor to perform the above or below method. BRIEF DESCRIPTION OF THE DRAWINGS
[0158] 3. Example 3.1 Figures
Figure 1
Figure 2a
Figure 2b
Figure 2c
Figure 2d
Figure 3a
Figure 3b
Figure 3c
Figure 4a
Figure 4b
Figure 4c
Figure 4d
Figure 5
Figure 6a
Figure 6b
Figure 6c
Figure 7
Figure 8a
Figure 8b
Figure 8c
Figure 9a
Figure 9b
Figure 9c
Figure 9d
Figure 10a
Figure 10b
Figure 11
Mode for Carrying Out the Invention
[0159] 3.2 Concepts Related to the Invention It can be seen that the example is based on an encoder that downmixes the signal 212 and provides channel level and correlation information 220 to the decoder. The decoder may generate a mixing rule (e.g., a mixing matrix) from the channel level and correlation information 220. Information important for generating the mixing rule may include the covariance information of the original signal 212 (e.g., covariance matrix C y ) and the covariance information of the downmixed signal (e.g., covariance matrix C x ). The covariance matrix C x can be directly estimated by the decoder by analyzing the downmixed signal, and the covariance matrix C y of the original signal 212 can be easily estimated by the decoder. The covariance matrix C y of the original signal 212 is generally a symmetric matrix (e.g., a 5x5 matrix in the case of a 5-channel original signal 212), and the matrix presents the level of each channel on the diagonal and the covariance between channels with off-diagonal entries. Since the covariance between a general channel i and channel j is the same as the covariance between j and i, the matrix is diagonal. Therefore, to provide the decoder with the entire covariance information, it is necessary to signal to the decoder five levels with diagonal entries and ten covariances with off-diagonal entries. However, it is shown that it is possible to reduce the amount of information to be encoded.
[0160] Furthermore, it is shown that in some cases, instead of levels and covariances, normalized values may be provided. For example, inter-channel coherence (ICC, also denoted as ξ i,j ), and inter-channel level difference (ICLD, also denoted as χ i ) indicating the value of energy may be provided. ICC may be, for example, a correlation value provided instead of the covariance of the off-diagonal entries of the matrix C y . An example of correlation information may be in the form of
[0161]
Number
[0162] in some examples, ξi,j Only a part of it is actually encoded.
[0163] In this way, an ICC matrix is generated. The diagonal entries of the ICC matrix are, in principle, equal to 1, and thus there is no need to encode the diagonal entries into the bitstream. However, it is understood that the encoder can provide the ICLD to the decoder in the
[0164]
Number
[0165] form (see below). In some examples, all χ i are actually encoded.
[0166] Figures 9a to 9d show an example of an ICC matrix 900 having diagonal values "d" that can be ICLD χ i and off-diagonal values that can be ICC ξ shown by 902, 904, 905, 906, 907 (see below). i,j
[0167] In this book, the product between matrices is indicated by the absence of a symbol. For example, the product between matrix A and matrix B is indicated by AB. The conjugate transpose of a matrix is indicated by an asterisk (*).
[0168] When referring to the diagonal, the diagonal is intended to be the main diagonal.
[0169] 3.3 The present invention FIG. 1 shows an audio system 100 using an encoder side and a decoder side. The encoder side can be implemented by an encoder 200 and can obtain an audio signal 212, for example, from an audio sensor unit (e.g., a microphone), from a memory unit, or from a remote unit (e.g., via wireless transmission). The decoder side can be implemented by an audio decoder (audio synthesizer) 300 that can provide audio content to an audio reproduction unit (e.g., a loudspeaker). The encoder 200 and the decoder 300 can communicate with each other via a communication channel that can be, for example, wired or wireless (e.g., via radio frequency waves, light, or ultrasonic waves). Thus, the encoder and / or the decoder can include or be connected to a communication unit (e.g., an antenna, a transceiver, etc.) for transmitting the encoded bitstream 248 from the encoder 200 to the decoder 300. In some cases, the encoder 200 can store the encoded bitstream 248 in a memory unit (e.g., a RAM memory, a FLASH memory, etc.) for future use. Similarly, the decoder 300 can read the bitstream 248 stored in the memory unit. In some examples, the encoder 200 and the decoder 300 can be the same device, in which case, after encoding and storing the bitstream 248, the device may need to read the bitstream 248 for playback of the audio content.
[0170] FIGS. 2a, 2b, 2c, and 2d show examples of the encoder 200. In some examples, the encoders of FIGS. 2a, 2b, 2c, and 2d can be the same and may differ from each other only because some elements are absent in one drawing and / or in the other drawing.
[0171] The audio encoder 200 may be configured to generate a downmix signal 246 from the original signal 212 (the original signal 212 has at least two (e.g., three or more) channels, and the downmix signal 246 has at least one downmix channel).
[0172] The audio encoder 200 may include a parameter estimator 218 configured to estimate the channel levels and correlation information 220 of the original signal 212. The audio encoder 200 may include a bitstream writer 226 for encoding the downmix signal 246 into a bitstream 248. Thus, the downmix signal 246 is encoded into the bitstream 248 to have side information 228 including the channel levels and correlation information of the original signal 212.
[0173] In particular, in some examples, the input signal 212 may be understood as a time-domain audio signal such as a time series of audio samples. The original signal 212 may have at least two channels corresponding to, for example, different microphones (e.g., stereo audio positions, or in the case of multi-channel audio positions, stereo audio positions), or corresponding to different loudspeaker positions of an audio reproduction unit. In the downmixer calculation block 244, the input signal 212 can be downmixed to obtain a downmixed version 246 (also denoted as x) of the original signal 212. This downmixed version of the original signal 212 is also referred to as the downmix signal 246. The downmix signal 246 has at least one downmix channel. The downmix signal 246 has fewer channels than the original signal 212. The downmix signal 212 may be in the time domain.
[0174] To store the bitstream or transmit it to a receiver (e.g., related to the decoder side), the downmix signal 246 is encoded into the bitstream 248 by a bitstream writer 226 (e.g., including an entropy encoder, or a multiplexer, or a core coder). The encoder 200 may include a parameter estimator (or parameter estimation block) 218. The parameter estimator 218 may estimate channel level and correlation information 220 related to the original signal 212. The channel level and correlation information 220 may be encoded into the bitstream 248 as side information 228. In an example, the channel level and correlation information 220 are encoded by the bitstream writer 226. In an example, FIG. 2b does not show the bitstream writer 226 downstream of the downmix calculation block 244, but nevertheless, the bitstream writer 226 may exist. In FIG. 2c, it is shown that the bitstream writer 226 may include a core coder 247 for encoding the downmix signal 246 to obtain an encoded version of the downmix signal 246. FIG. 2c also shows that the bitstream writer 226 may include a multiplexer 249, and the multiplexer 249 encodes both the coded downmix signal 246 and the channel level and correlation information 220 (e.g., as coded parameters) within the side information 228 into the bitstream 248.
[0175] As shown by FIG. 2b (not shown in FIGS. 2a and 2c), the original signal 212 may be processed (e.g., by a filter bank 214, see below) to obtain a frequency domain version 216 of the original signal 212.
[0176] The parameter estimator 218 estimates parameters ξ i,j and χ iAn example of parameter estimation that defines (for example, a normalization parameter) is shown in FIG. 6c. Covariance estimators 502 and 504 estimate covariance C x and C y for the downmixed signal 246 and the input signal 212 to be encoded, respectively. Then, in the ICLD block 506, the ICLD parameter χ i is calculated and provided to the bit stream writer 246. In the covariance-to-coherence block 510, the ICC ξ i,j (412) is obtained. In block 250, only a part of the ICC is selected as the encoding target.
[0177] The parameter quantization block 222 (FIG. 2b) may enable the acquisition of channel level and correlation information 220 in the quantized version 224.
[0178] The channel level and correlation information 220 of the original signal 212 may generally include information regarding the energy (or level) of the channels of the original signal 212. Additionally or alternatively, the channel level and correlation information 220 of the original signal 212 may include correlation information between pairs of channels, such as the correlation between two different channels. The channel level and correlation information may include information related to the covariance matrix C y where each column and row is related to a specific channel of the original signal 212 (for example, in a normalized form such as correlation or ICC), the channel level is described by the diagonal elements of the matrix C y and the correlation information, and the correlation information is described by the off-diagonal elements of the matrix C y . The matrix C y may be such that the matrix is symmetric (i.e., the matrix is equal to its transpose) or Hermitian (i.e., the matrix is equal to its conjugate transpose). C yis generally positive semi - definite. In some examples, the correlation can be replaced by the covariance (correlation information can be replaced by covariance information). It is understood that it is possible to encode information related to fewer channels than all channels of the original signal 212 into the side information 228 of the bitstream 248. For example, it is not necessary to provide channel level and correlation information for all channels or all pairs of channels. For example, only a reduced set of information regarding the correlation between pairs of channels of the down - mixed signal 212 can be encoded into the bitstream 248, and the remaining information can be estimated at the decoder side. Generally, C y it is possible to encode fewer elements than the diagonal elements of C y and it is possible to encode fewer elements than the off - diagonal elements of C.
[0179] For example, the channel level and correlation information can include entries of the covariance matrix C y (channel level and correlation information 220 of the original signal) and / or the covariance matrix C x (covariance information of the down - mixed signal) in, for example, a normalized form. For example, the covariance matrix can associate each row and each column with each channel such that it represents the covariance between different channels and the level of each channel on the diagonal of the matrix. In some examples, the channel level and correlation information 220 of the original signal 212 encoded in the side information 228 can include only channel level information (e.g., only the diagonal values of the correlation matrix C y ) or only correlation information (e.g., only the off - diagonal values of the correlation matrix C y ). The same applies to the covariance information of the down - mixed signal.
[0180] As will be shown later, the channel level and correlation information 220 includes at least one coherence value (ξ i,jmay include. Additionally or alternatively, the channel level and correlation information 220 may include at least one inter-channel level difference ICLD (χ i ). In particular, it is possible to define a matrix having ICLD values or inter-channel coherence (ICC) values. Thus, matrix C y and matrix C x The above examples regarding the transmission of elements of can be generalized to other values encoded (e.g., transmitted) to embody the channel level and correlation information 220 and / or the coherence information of the downmix channels.
[0181] The input signal 212 can be subdivided into a plurality of frames. Different frames can have, for example, equal time lengths (e.g., different frames can each be composed of the same number of samples in the time domain during the elapsed time of one frame). Thus, different frames generally have the same time length. In the bitstream 248, the downmix signal 246 (which can be a time-domain signal) can be encoded in a frame-by-frame manner (or, in any case, the subdivision into frames can be determined by the decoder). The channel level and correlation information 220 encoded as side information 228 in the bitstream 248 can be associated with each frame (e.g., the parameters of the channel level and correlation information 220 can be provided for each frame or for a plurality of consecutive frames). Thus, for each frame of the downmix signal 246, the associated side information 228 (e.g., parameters) can be encoded within the side information 228 of the bitstream 248. In some cases, a plurality of consecutive frames can be associated with the same channel level and correlation information 220 (e.g., the same parameters) encoded within the side information 228 of the bitstream 248. Thus, one parameter can result in being collectively associated with a plurality of consecutive frames. This can occur in some examples when two consecutive frames have similar characteristics or when it is necessary to reduce the bitrate (e.g., due to the need to reduce the payload). For example, When the payload is high, the number of consecutive frames associated with the same specific parameters increases, thereby reducing the amount of bits written into the bitstream. When the payload is low, the number of consecutive frames associated with the same specific parameters decreases, thereby improving the mixing quality.
[0182] In other cases, when the bitrate decreases, the number of consecutive frames associated with the same specific parameters increases, thereby reducing the amount of bits written into the bitstream. The reverse is also true.
[0183] In some cases, it is possible to smooth the parameter (or the reconstructed or estimated value such as the covariance) by using a linear combination with the parameter (or the reconstructed or estimated value such as the covariance) preceding the current frame, for example, by addition, averaging, etc.
[0184] In some examples, a frame can be divided among a plurality of subsequent slots. FIG. 10a shows a frame 920 (subdivided into four consecutive slots 921 to 924), and FIG. 10b shows a frame 930 (subdivided into four consecutive slots 931 to 934). The time lengths of different slots can be the same. If the frame length is 20 ms and the slot size is 1.25 ms, there are 16 slots in one frame (20 / 1.25 = 16).
[0185] The subdivision of the slot can be performed in a filter bank (for example, 214) described below.
[0186] In one example, the filter bank is a complex-modulated low-delay filter bank (CLDFB), the frame size is 20 ms, the slot size is 1.25 ms, and as a result, there are 16 filter bank slots per frame. The number of bands in each slot depends on the input sampling frequency, and the bandwidth is 400 Hz. Thus, for example, when the input sampling frequency is 48 kHz, the frame length of the samples is 960, the slot length is 60 samples, and the number of filter bank samples per slot is also 60.
[0187] [Table 2]
[0188] Even when each frame (similarly each slot) can be encoded in the time domain, analysis on a band-by-band basis can be performed. In the example, multiple bands are analyzed for each frame (or slot). For example, a filter bank can be applied to the time signal, and the resulting sub-band signals can be analyzed. In some examples, channel level and correlation information 220 is also provided in a band-by-band fashion. For example, for each band of the input signal 212 or the downmix signal 246, the relevant channel level and correlation information 220 (e.g., C y or ICC matrix) can be provided. In some examples, the number of bands can be changed based on the characteristics of the signal and / or the required bitrate, or the measurements of the current payload. In some examples, the more slots are required to maintain a similar bitrate, the fewer bands are used.
[0189] Since the slot size is smaller than the frame size (time duration), if a transient phenomenon is detected in the original signal 212 detected within the frame, the slot can be appropriately used. The encoder (specifically, the filter bank 214) recognizes the presence of the transient phenomenon, signals its presence in the bitstream, and can indicate in the side information 228 of the bitstream 248 in which slot of the frame the transient phenomenon occurred. Further, the parameters of the channel level and correlation information 220 encoded within the side information 228 of the bitstream 248 can thus be appropriately associated only with the slots following the transient phenomenon and / or the slots in which the transient phenomenon occurred. Accordingly, the decoder will determine the presence of the transient phenomenon and associate the channel level and correlation information 220 only with the slots following the transient phenomenon and / or the slots in which the transient phenomenon occurred (in the case of slots preceding the transient phenomenon, the decoder will use the channel level and correlation information 220 of the previous frame). In FIG. 10a, no transient phenomenon has occurred, and thus it can be understood that the parameters 220 encoded within the side information 228 are associated with the entire frame 920. In FIG. 10b, a transient phenomenon is occurring in slot 932. Accordingly, the parameters 220 encoded within the side information 228 refer to slots 932, 933, and 934, while the parameters associated with slot 931 are assumed to be the same as those of the frame preceding frame 930.
[0190] Taking the above into account, for each frame (or slot) and each band, specific channel level and correlation information 220 related to the original signal 212 can be defined. For example, for each band, the elements of the covariance matrix C y (e.g., covariance and / or level) can be estimated.
[0191] When a transient phenomenon is detected when a plurality of frames are collectively associated with the same parameters, it is possible to reduce the number of frames associated with the same parameters collectively in order to improve the mixed quality.
[0192] FIG. 10a shows a frame 920 (herein shown as a "normal frame") in which eight bands are defined in the original signal 212 (the eight bands 1...8 are shown on the vertical axis and the slots 921 to 924 are shown on the horizontal axis). The parameters of the channel level and correlation information 220 can theoretically be encoded in the side information 228 of the bit stream 248 in a per-band fashion (for example, there is one covariance matrix for each original band). However, in order to reduce the amount of side information 228, the encoder can aggregate a plurality of original bands (for example, consecutive bands) to obtain at least one aggregated band formed by the plurality of original bands. For example, in FIG. 10a, the eight original bands are grouped to obtain four aggregated bands (aggregated band 1 associated with original band 1, aggregated band 2 associated with original band 2, aggregated band 3 grouping original bands 3 and 4, aggregated band 4 grouping original bands 5...8). Matrices such as covariance, correlation, ICC, etc. can be associated with each of the aggregated bands. In some examples, what is encoded within the side information 228 of the bit stream 248 is a parameter obtained from the sum (or average, or another linear combination) of the parameters associated with each aggregated band. Thus, the size of the side information 228 of the bit stream 248 is further reduced. Hereinafter, since an "aggregated band" refers to a band used to determine the parameter 220, it is also called a "parameter band".
[0193] Figure 10b shows a frame 930 in which a transient phenomenon occurs (which is subdivided into four consecutive slots 931 - 934, or another integer). Here, the transient phenomenon occurs in the second slot 932 (the "transient slot"). In this case, the decoder may determine to have the parameters of the channel level and correlation information 220 refer only to the transient slot 932 and / or the subsequent slots 933 and 934. The channel level and correlation information 220 of the preceding slot 931 are not provided. The channel level and correlation information of slot 931 are understood to be specifically different from the channel level and correlation information of the slot, but may be similar to the channel level and correlation information of the frame preceding frame 930. Therefore, the decoder will apply the channel level and correlation information of the frame preceding frame 930 to slot 931, and apply the channel level and correlation information of frame 930 only to slots 932, 933, and 934.
[0194] Since the presence and position of slot 931 with the transient phenomenon can be signaled in the side information 228 of the bitstream 248 (for example, at 261 as described later), techniques have been developed to avoid or reduce the increase in the size of the side information 228. That is, the grouping between the aggregated bands can be changed. For example, aggregated band 1 now groups the original bands 1 and 2, and aggregated band 2 groups the original bands 3... 8. Therefore, the number of bands is further reduced compared to the case of Figure 10a, and the parameters are provided only for the two aggregated bands.
[0195] Figure 6a shows that the parameter estimation block (parameter estimator) 218 can search for a specific number of channel level and correlation information 220.
[0196] Figure 6a shows that the parameter estimator 218 can search for a specific number of parameters (channel level and correlation information 220) that can be the ICC of the matrix 900 in Figures 9a - 9d.
[0197] However, only some of the estimated parameters are actually sent to the bitstream writer 226 to encode the side information 228. The reason is that the encoder 200 can be configured to select whether to encode at least a part of the channel level and correlation information 220 of the original signal 212 (in a decision block 250 not shown in FIGS. 1 to 5).
[0198] This is shown in FIG. 6a as a plurality of switches 254s controlled by a selection (command) 254 from the decision block 250. If each of the outputs 220 of the block parameter estimation 218 is the ICC of the matrix 900 in FIG. 9c, not all of the parameters estimated by the parameter estimation block 218 are actually encoded in the side information 228 of the bitstream 248. Specifically, the entry 908 (the ICC between channels, i.e., between R and L, between C and L, between C and R, between RS and CS) is actually encoded, but the entry 907 is not encoded (i.e., the decision block 250, which may be the same as that in FIG. 6c, may appear to open the switches 254s of the non-encoded entry 907, but closes the switches 254s of the entry 908 encoded in the side information 228 of the bitstream 248). Note that information 254' (entry 908) regarding which parameters are selected for encoding can be encoded (e.g., as a bitmap or other information regarding which entry 908 is encoded). In fact, the information 254', which can be an ICC map for example, can include the index of the encoded entry 908 (formatted in FIG. 9d). The information 254' can be in the form of a bitmap. For example, the information 254' can be composed of fixed-length fields, each position is associated with an index according to a predefined order, and the value of each bit provides information regarding whether the parameter associated with that index is actually provided.
[0199] Generally, decision block 250 may select whether to encode at least a portion of channel level and correlation information 220 (i.e., determine whether to encode the entries of matrix 900), for example, based on status information 252. Status information 252 can be based on payload status, and for example, when transmission is at a high load, it is possible to reduce the amount of side information 228 encoded within bitstream 248. For example, referring to 9c, In the case of a high payload, the number of entries 908 of matrix 900 actually written into side information 228 of bitstream 248 decreases, In the case of a lower payload, the number of entries 908 of matrix 900 actually written into side information 228 of bitstream 248 decreases.
[0200] Alternatively, or in addition, metrics 252 can be evaluated to determine which parameters 220 should be encoded within side information 228 (e.g., which entries of matrix 900 are to be defined as encoded entries 908 and which entries should be discarded). In this case, it is possible to encode only the parameters 220 (associated with metrics that are more susceptible to influence) within the bitstream (e.g., metrics associated with more perceptually significant covariance can be associated with the entries selected as encoded entries 908).
[0201] Note that this process can be repeated for each frame (or for multiple frames in the case of downsampling), for each band.
[0202] Thus, decision block 250 can also be controlled by parameter estimator 218 via command 251 of FIG. 6a, in addition to, for example, status metrics.
[0203] In some examples (e.g., FIG. 6b), the audio encoder may be further configured to encode the current channel level and correlation information 220t into the bitstream 248 as an increment 220k relative to the previous channel level and correlation information 220(t - 1). What is encoded into the side information 228 by this bitstream writer 226 may be the increment 220k relative to the previous frame associated with the current frame (or slot). This is shown in FIG. 6b. The current channel level and correlation information 220t is provided to the storage element 270, such that the storage element 270 stores the value of the current channel level and correlation information 220t for subsequent frames. On the other hand, the current channel level and correlation information 220t may be compared with the previously acquired channel level and correlation information 220(t - 1) (shown as subtractor 273 in FIG. 6b). Thus, the result 220Δ of the subtraction may be obtained by the subtractor 273. At the scaler 220s, the difference 220Δ can be used to obtain the relative increment 220k between the previous channel level and correlation information 220(t - 1) and the current channel level and correlation information 220t. For example, if the current channel level and correlation information 220t is 10% greater than the previous channel level and correlation information 220(t - 1), the increment 220 encoded into the side information 228 by the bitstream writer 226 will indicate information of a 10% increment. In some examples, instead of providing the relative increment 220k, simply the difference 220Δ may be encoded.
[0204] As described above and below, the selection of the parameters to be actually encoded, from among parameters such as ICC and ICLD, can be adapted to a particular situation. For example, in some examples, For one first frame, only the ICC 908 of FIG. 9c is selected to be encoded into the side information 228 of the bitstream 248, and the ICC 907 is not encoded into the side information 228 of the bitstream 248. In the case of the second frame, different ICCs are selected to be encoded, and different ICCs that are not selected are not encoded.
[0205] The same can be true for slots and bands (and various parameters such as ICLD). Thus, the encoder (specifically, block 250) can determine which parameters to encode and which not to encode, thereby adapting the selection of the parameters to be encoded to a specific situation (e.g., status, selection, etc.). Therefore, "importance features" can be analyzed to select which parameters to encode and which not to encode. The importance features can be, for example, metrics related to the results obtained from simulations of the operations performed by the decoder. For example, the encoder can simulate the reconstruction of the decoder for the covariance parameter 907 that is not encoded, and the importance feature can be a metric indicating the absolute error between the covariance parameter 907 that is not encoded and the same parameter that is assumed to be reconstructed by the decoder. To distinguish between the covariance parameter 908 to be encoded and the covariance parameter 907 that is not encoded based on the least influential simulation scenario, by measuring the error in various simulation scenarios (e.g., each simulation scenario is related to the transmission of some encoded covariance parameters 908, and the measurement of the error affects the reconstruction of the covariance parameter 907 that is not encoded), it is possible to determine the simulation scenario with the least influence by the error (e.g., the simulation scenario that includes the metrics for all errors in the reconstruction). In the scenario with the least influence, the parameter 907 that is not selected is the parameter that can be most easily reconstructed, and the parameter 908 that is selected is, tendentially, the parameter with the largest metric related to the error.
[0206] Instead of simulating parameters such as ICC and ICLD, the same can be done by simulating the reconstruction or estimation of the covariance by the decoder, or by simulating the mixing characteristics or mixing results. In particular, the simulation can be performed frame by frame or slot by slot, and can be done band by band or aggregated band by band.
[0207] One example may start from the parameters encoded in the side information 228 of the bit stream 248 and simulate the reconstruction of the covariance using Equation (4) or Equation (6) (see below).
[0208] More generally, reconstruct the channel level and correlation information from the selected channel level and correlation information, thereby simulating the estimated values of the unselected channel level and correlation information (220, C y ) in the decoder (300), between the unselected channel level and correlation information (220) estimated by the encoder, and the unselected channel level and correlation information reconstructed by simulating the estimated values of the unencoded channel level and correlation information (220) in the decoder (300) Calculate the error information, and as a result, Based on the calculated error information, Distinguish between the channel level and correlation information that can be appropriately reconstructed and the channel level and correlation information that cannot be appropriately reconstructed, As a result, Select the channel level and correlation information that cannot be appropriately reconstructed, which is encoded in the side information (228) of the bit stream (248), and Do not select the channel level and correlation information that can be appropriately reconstructed It is possible to decide, and thereby prevent the channel level and correlation information that can be appropriately reconstructed from being encoded in the side information (228) of the bit stream (248).
[0209] Generally, an encoder may simulate any operation of a decoder and evaluate an error metric from the result of the simulation.
[0210] In some examples, an importance feature may be different from (or may include other metrics different from) the evaluation of a metric associated with an error. In some cases, an importance feature may be related to a manual selection or may be based on importance based on psychoacoustic criteria. For example, without simulation, the most important pair of channels can be selected and encoded (908).
[0211] Next, some additional explanations are provided to describe how an encoder may signal which parameter 908 is actually encoded within side information 220 of bitstream 248.
[0212] Referring to FIG. 9d, the parameters on the diagonal of the ICC matrix 900 are associated with the ordered indices 1..10 (the order is pre-determined and recognized by the decoder). In FIG. 9c, it is shown that the parameters 908 selected to be encoded are the ICCs for the pairs L-R, L-C, R-C, LS-RS indexed by indices 1, 2, 5, 10 respectively. Thus, in the side information 228 of the bitstream 248, the indications of indices 1, 2, 5, 10 are also provided (for example, in the information 254' of FIG. 6a). Thus, the decoder understands that the four ICCs provided in the side information 228 of the bitstream 248 are L-R, L-C, R-C, LS-RS, similarly by the information regarding the indices 1, 2, 5, 10 provided in the side information 228 by the encoder. The index can be provided, for example, via a bitmap that associates the position of each bit in the bitmap with a pre-determined one. For example, to signal indices 1, 2, 5, 10, since the first, second, fifth, and tenth bits refer to indices 1, 2, 5, 10, it is possible to write "1100100001" (other possibilities can be freely used by those skilled in the art) in (the field 254' of the side information 228). This is a so-called one-dimensional index, but other indexing strategies are also possible. For example, the combinatorial number technique, according to which a number N uniquely associated with a particular pair of channels is encoded (see also https: / / en.wikipedia.org / wiki / Combinatorial_number_system) in (the field 254' of the side information 228). The bitmap can also be called an ICC map when referring to the ICC.
[0213] Note that in some cases, a non-adaptive (fixed) provision of parameters may be used. This means that in the example of Figure 6a, the selection 254 from among the parameters to be encoded is fixed and there is no need to indicate the selected parameters in field 254'. Figure 9b shows an example of a fixed provision of parameters, where the selected ICCs are L-C, L-LS, R-C, C-RS, and since the decoder already knows which ICCs are encoded in the side information 228 of the bitstream 248, there is no need to signal their indices.
[0214] However, in some cases, the encoder may perform a selection between a fixed provision of parameters and an adaptive provision of parameters. The encoder can signal the selection in the side information 228 of the bitstream 248, so that the decoder can know which parameters are actually encoded.
[0215] In some cases, at least some parameters may be provided without adaptation. For example, ICDL can be encoded in any case without the need to represent ICDL as a bitmap, and ICC can be subject to adaptive provision.
[0216] The description relates to each frame, or slot, or band. In the case of subsequent frames, or slots, or bands, different parameters 908 are provided to the decoder, different indexes are associated with the subsequent frames, or slots, or bands, and various selections (e.g., fixed vs. adaptive) can be performed. FIG. 5 shows an example of a filter bank 214 of an encoder 200 that can be used to process the original signal 212 to obtain the frequency domain signal 216. As seen in FIG. 5, the time domain (TD) signal 212 can be analyzed by a transient analysis block 258 (transient detector). Further, the conversion of the input signal 212 to its frequency domain (FD) version 264 in multiple bands is achieved by a filter 263 (which can implement, for example, a Fourier filter, a short-time Fourier filter, an orthogonal mirror, etc.). The frequency domain version 264 of the input signal 212 can be analyzed, for example, in a band analysis block 267, which can determine (command 268) a specific grouping of bands performed in a partitioning grouping block 265. Thereafter, the FD signal 216 becomes a signal with a reduced number of aggregated bands. The aggregation of bands has been described above with respect to FIGS. 10a and 10b. The partitioning grouping block 265 can also be conditioned by the transient analysis performed by the transient analysis block 258. As described above, in the case of transients, it may be possible to further reduce the number of aggregated bands. Thus, information 260 regarding the transient can condition the partitioning grouping. Additionally, or alternatively, information 261 regarding the transient is encoded within the side information 228 of the bitstream 248. When encoded within the side information 228, the information 261 can include, for example, a flag indicating whether a transient has occurred (e.g., "1" meaning "there was a transient in the frame" vs. "0" meaning "there was no transient in the frame"), and / or an indication of the position of the transient within the frame (such as a field indicating in which slot the transient was observed).In some examples, when information 261 indicates that there is no transient in the frame (“0”), in order to reduce the size of bitstream 248, the indication of the position of the transient is not encoded in side information 228. Information 261 is also referred to as the “transient parameter” and is shown as being encoded in side information 228 of bitstream 248 in FIGS. 2d and 6b.
[0217] In some examples, the partitioning grouping at block 265 can also be conditioned by external information 260' such as information regarding the status of transmission (e.g., measurements related to transmission, error rate, etc.). For example, the higher the payload (or the higher the error rate), the greater the aggregation (there is a tendency for fewer wider aggregation bands), thereby reducing the amount of side information 228 encoded in bitstream 248. In some examples, information 260' may be similar to the information or metrics 252 of FIG. 6a.
[0218] Generally, it is impossible to transmit the parameters of every combination of band / slot, but the samples of the filter bank are grouped together across both the number of slots and the number of bands in order to reduce the number of parameter sets transmitted per frame. To group the bands along the frequency axis into parameter bands, a non-constant partitioning of the parameter bands is used, where the number of bands in the parameter bands is not constant and attempts to follow a psychoacoustically motivated parameter band resolution. That is, at lower bands, the parameter bands include only one or a few filter bank bands, and in the case of higher parameter bands, more (steadily increasing) filter bank bands are grouped into one parameter band.
[0219] Thus, for example, here too, when the input sampling rate is 48 kHz and the number of parameter bands is set to 14, the following vector grp 14 indicates the filter bank indices giving the band boundaries (index starting from 0) of the parameter bands. grp14 =[0,1,2,3,4,5,6,8,10,13,16,20,28,40,60]
[0220] Parameter band j includes the filter bank bands [grp 14 [j], grp 14 [j + 1]].
[0221] Note that the 48 kHz band grouping follows a psychoacoustically motivated frequency scale and has specific band boundaries corresponding to the number of bands for each sampling frequency, so it can be directly used for other possible sampling rates by simply truncating the ends (Table 1).
[0222] If the frame is non-transient or transient processing is not implemented, grouping along the time axis is performed for all slots within the frame so that one parameter set is available for each parameter band.
[0223] Still, the number of parameter sets is large, but the time resolution may be lower than a 20 ms frame (average 40 ms). Therefore, to further reduce the number of parameter sets transmitted per frame, only a subset of the parameter bands is used to determine and encode the parameters for transmission to the decoder in the bitstream. The subset is fixed and recognized by both the encoder and the decoder. The specific subset transmitted in the bitstream is signaled by a field in the bitstream, indicating the decoder to which the subset of the parameter bands to which the transmitted parameters belong belongs. Then, the decoder replaces the parameters of this subset with the transmitted parameters (ICC, ICLD) and retains the parameters (ICC, ICLD) of the previous frame for all parameter bands not present in the current subset.
[0224] In one example, the parameter band can be split into two subsets: a contiguous subset for the lower parameter band that includes approximately half of the full parameter band, and one contiguous subset for the higher parameter band. Since there are two subsets, the bitstream field for signaling the subset is a single bit, and an example of the subset for the case of 48 kHz and 14 parameter bands is s 14 =[1,1,1,1,1,1,1,0,0,0,0,0,0,0] where s 14 [j] indicates to which subset parameter band j belongs.
[0225] Note that the downmix signal 246 can actually be encoded into the bitstream 248 as a signal in the time domain. Briefly, the subsequent parameter estimator 218 estimates the parameters 220 (e.g., ξ i,j and / or χ i ) in the frequency domain (the decoder 300 will use the parameters 220 to prepare the mixing rule (e.g., mixing matrix) 403 as described below).
[0226] Figure 2d shows an example of an encoder 200 that can be one of the previous encoders or can include elements of the aforementioned encoders. The TD input signal 212 is input to the encoder, and the bitstream 248 is output. The bitstream 248 includes the downmix signal 246 (encoded, e.g., by the core coder 247) and the correlation and level information 220 encoded in the side information 228.
[0227] As shown in FIG. 2d, a filter bank 214 may be included (an example of a filter bank is provided in FIG. 5). To obtain the FD signal 264, which is the FD version of the input signal 212, a frequency domain (FD) conversion is provided at block 263 (frequency domain DMX). Multiple bands of the FD signal 264 (also indicated by X) are obtained. To obtain the FD signal 216 in the aggregated band, a band / slot grouping block 265 (which may implement the grouping block 265 of FIG. 5) may be provided. The FD signal 216 may, in some examples, be a version of the FD signal 264 in fewer bands. Subsequently, the signal 216 may be provided to a parameter estimator 218, which includes covariance estimation blocks 502, 504 (shown here as a single block), and downstream parameter estimation and coding blocks 506, 510 (embodiments of elements 502, 504, 506, and 510 are shown in FIG. 6c). The parameter estimation coding blocks 506, 510 may also provide the parameter 220 encoded within the side information 228 of the bitstream 248. A transient detector 258 (which may implement the transient analysis block 258 of FIG. 5) can find the transient and / or the position of the transient within the frame (e.g., in which slot the transient was identified). Thus, information 261 regarding the transient (e.g., transient parameters) may be provided to the parameter estimator 218 (e.g., to determine which parameters to encode). The transient detector 258 may also provide information or a command (268) to block 265 such that grouping is performed by taking into account the presence and / or position of the transient within the frame.
[0228] FIGS. 3a, 3b, and 3c show an example of an audio decoder 300 (also referred to as an audio synthesizer). In the example, the decoders of FIGS. 3a, 3b, and 3c may be the same decoder except for some differences to avoid different elements. In the example, the decoder 300 may be the same as the decoder of FIGS. 1 and 4. In the example, the decoder 300 may also be the same device as the encoder 200.
[0229] The decoder 300 may be configured to generate a composite signal (336, 340, y R ) from the downmix signal x of TD (246) or FD (314). The audio synthesizer 300 may include an input interface 312 configured to receive the downmix signal 246 (e.g., the same downmix signal encoded by the encoder 200) and the side information 228 (e.g., encoded within the bitstream 248). The side information 228 may include at least one of the channel levels and correlation information (220, 314) of the original signal (such as ξ, χ, etc., which may be the original input signal 212, y on the encoder side), or elements thereof (described below). In some examples, all ICLD (χ) outside the diagonal of the ICC matrix 900 (ICC or ξ value) and some entries (not all) 906 or 908 are obtained by the decoder 300.
[0230] The decoder 300 may be configured to calculate a prototype signal 328 from the downmix signal (324, 246, x) (e.g., via a prototype signal calculator or prototype signal calculation module 326), and the prototype signal 328 has several (more than 1) channels of the composite signal 336.
[0231] The decoder 300 is channel levels and correlation information of the original signal (212, y) (e.g., 314, C y , ξ, χ, or elements thereof), and covariance information related to the downmix signal (324, 246, x) (e.g., C x or elements thereof) to calculate the mixing rule 403 (e.g., via a mixing rule calculator 402) using at least one of them.
[0232] The decoder 300 uses the prototype signal 328 and the mixing rule 403 to generate a composite signal (336, 340, y RIt may include a synthesis processor 404 configured to generate
[0233] The synthesis processor 404 and the mixing rule calculator 402 may be incorporated into one synthesis engine 334. In some examples, the mixing rule calculator 402 may be external to the synthesis engine 334. In some examples, the mixing rule calculator 402 of FIG. 3a may be integrated with the parameter reconstruction module 316 of FIG. 3b.
[0234] The number of synthesis channels of the synthesized signal (336, 340, y R ) may be more than one (in some cases, more than two or more than three), and may be more than, less than, or the same as the number of original channels of the original signal (212, y), and the number of original channels is also more than one (in some cases, more than two or more than three). The number of channels of the downmix signal (246, 216, x) is at least one or two, and is less than the number of original channels of the original signal (212, y) and the number of synthesis channels of the synthesized signal (336, 340, y R ).
[0235] The input interface 312 can read the encoded bitstream 248 (e.g., the same bitstream 248 encoded by the encoder 200). The input interface 312 can be or include a bitstream reader and / or an entropy decoder. As described above, the bitstream 248 can encode the downmix signal (246, x) and the side information 228. The side information 228 can include the original channel level and correlation information 220 in any form output by, for example, the parameter estimator 218 or any element downstream of the parameter estimator 218 (such as the parameter quantization block 222, etc.). The side information 228 can include encoded values and / or indexed values, or both. Even if the input interface 312 is not shown for the downmix signal (346, x) in FIG. 3b, the input interface 312 can still be applied to the downmix signal as in FIG. 3a. In some examples, the input interface 312 can quantize the parameters obtained from the bitstream 248.
[0236] Accordingly, the decoder 300 can obtain the downmix signal (246, x) that can be in the time domain. As described above, the downmix signal 246 can be divided into frames and / or slots (see above). In an example, the filter bank 320 can convert the downmix signal 246 in the time domain to obtain a version 324 of the downmix signal 246 in the frequency domain. As described above, the bands of the frequency domain version 324 of the downmix signal 246 can be grouped into groups of bands. In an example, the same grouping (see above) performed by the filter bank 214 can be implemented. The parameters for grouping (e.g., which bands and / or how many bands should be grouped...) can be based on signaling by, for example, the partition grouper 265 or the band analysis block 267, and the signaling is encoded within the side information 228.
[0237] The decoder 300 may include a prototype signal calculator 326. The prototype signal calculator 326 may calculate a prototype signal 328 from a downmix signal (e.g., one of versions 324, 246, x) by applying, for example, a prototype rule (e.g., matrix Q). The prototype rule may be embodied by a prototype matrix (Q) having a first dimension and a second dimension, where the first dimension is associated with the number of downmix channels and the second dimension is associated with the number of synthesis channels. Thus, the prototype signal has some channels of the ultimately generated synthesis signal 340.
[0238] The prototype signal calculator 326 may apply a so-called upmix to the downmix signals (324, 246, x) in the sense of simply generating versions of the downmix signals (324, 246, x) with a larger number of channels (the number of channels of the generated synthesis signal) without applying so much “intelligence”. In an example, the prototype signal calculator 326 may simply apply a fixed default prototype matrix (identified as “Q” in this document) to the FD version 324 of the downmix signal 246. In an example, the prototype signal calculator 326 may apply different prototype matrices to different bands. The prototype rule (Q) may be selected from among a plurality of pre-stored prototype rules based on, for example, a specific number of downmix channels and a specific number of synthesis channels.
[0239] The prototype signal 328 may be decorrelated in a decorrelation module 330 to obtain a decorrelated version 332 of the prototype signal 328. However, in some examples, advantageously, it has been proven that the present invention is sufficiently effective to enable avoidance of the decorrelation module 330, so the decorrelation module 330 does not exist.
[0240] (Either version 328 or 332 of) the prototype signal can be input into the synthesis engine 334 (specifically, into the synthesis processor 404). Here, the prototype signals (328, 332) are processed to obtain the synthesis signals (336, y R ). The synthesis engine 334 (specifically, the synthesis processor 404) can apply the mixing rules 403 (in some examples, there are two mixing rules as described below, for example, one for the main component of the synthesis signal and one for the residual component). The mixing rules 403 can be embodied by, for example, a matrix. The matrix 403 can be generated by the mixing rule calculator 402 based on, for example, the channel levels and correlation information (such as 314, ξ, χ or their elements) of the original signals (212, y).
[0241] The synthesis signal 336 output by the synthesis engine 334 (specifically, by the synthesis processor 404) can optionally be filtered in the filter bank 338. Additionally, or alternatively, the synthesis signal 336 can be converted to the time domain in the filter bank 338. Thus, a version 340 of the synthesis signal 336 (either in the time domain or filtered) can be used for audio reproduction (e.g., by a loudspeaker).
[0242] To obtain the mixing rules (e.g., mixing matrix) 403, the channel levels and correlation information of the original signals (e.g., C y ,
[0243]
Number
[0244] etc.), as well as the covariance information related to the downmix signal (e.g., C x ) can be provided to the mixing rule calculator 402. For this purpose, it is possible to utilize the channel levels and correlation information 220 encoded in the side information 228 by the encoder 200.
[0245] However, in some cases, not all parameters are encoded by the encoder 200 in order to reduce the amount of information encoded within the bitstream 248 (e.g., not all of the channel levels and correlation information of the original signal 212 and / or not all of the covariance information of the downmixed signal 246). Therefore, some parameters 318 will be estimated in the parameter reconstruction module 316.
[0246] The parameter reconstruction module 316 may be supplied, for example, by at least one of a version 322 of the downmixed signal 246(x), which may be, for example, a filtered version or FD version of the downmixed signal 246, and side information 228 (including channel levels and correlation information 220). of which at least one may be supplied.
[0247] The side information 228 may include information related to the correlation matrix C y of the original signal (212, y) (as the level and correlation information of the input signal). However, in some cases, not all elements of the correlation matrix C y are actually encoded. Therefore, techniques for estimation and reconstruction have been developed to reconstruct a version of the correlation matrix C y (
[0248]
Number
[0249] ) (e.g., via an intermediate step of obtaining an estimated version
[0250]
Number
[0251] ).
[0252] The parameter 314 provided to the module 316 can be obtained by the entropy decoder 312 (input interface) and can be quantized, for example.
[0253] FIG. 3c shows an example of a decoder 300 that can be one embodiment of one of the decoders of FIGS. 1 to 3b. Here, the decoder 300 includes an input interface 312 represented by a demultiplexer. The decoder 300 outputs a composite signal 340, and the composite signal 340 can be reproduced at TD by, for example, a loudspeaker (signal 340), or can be reproduced at FD (signal 336). The decoder 300 of FIG. 3c can include a core decoder 347, and the core decoder 347 can also be part of the input interface 312. Thus, the core decoder 347 can provide the downmix signal x, 246. The filter bank 320 can convert the downmix signal 246 from TD to FD. The FD version of the downmix signal x, 246 is indicated by 324. The FD downmix signal 324 can be provided to the covariance synthesis block 388. The covariance synthesis block 388 can provide a composite signal 336 (Y) at FD. The inverse filter bank 338 can convert the audio signal 314 to its TD version 340. The FD downmix signal 324 can be provided to the band / slot grouping block 380. The band / slot grouping block 380 can perform the same operations as those performed by the partitioning grouping block 265 of FIGS. 5 and 2d in the encoder. Since the bands of the downmix signal 216 of FIGS. 5 and 2d are grouped or aggregated into several (wide) bands in the encoder and the parameters 220 (ICC, ICLD) are associated with the groups of the aggregated bands, then it is necessary to aggregate the decoded downmix signal in the same way and associate each aggregated band with the relevant parameters. Thus, the number 385 is the downmix signal X after aggregation Brefers to. The filter provides an unaggregated FD representation, and thus, in order to be able to process the parameters in the same way as the encoder, the band / slot grouping in the decoder (380) performs aggregation across the bands / slots in the same way as the encoder, and the aggregated downmix X B is provided. Note that
[0254] The band / slot grouping block 380 also aggregates across different slots within the frame, and as a result, the signal 385 is also aggregated in the slot dimension in the same way as the encoder. The band / slot grouping block 380 may also receive information 261 encoded in the side information 228 of the bitstream 248 indicating the presence of transients and, in some cases, the position of the transients within the frame.
[0255] In the covariance estimation block 384, the covariance C x of the downmix signal 246(324) is estimated. The covariance C y is obtained in the covariance calculation block 386, for example, by using equations (4) to (8) that can be used for this purpose. FIG. 3c shows, for example, "multichannel parameters" that can be the parameters 220 (ICC and ICLD). Then, the covariance C y and C x are provided to the covariance synthesis block 388 where the synthesized signal 388 is synthesized. In some examples, the blocks 384, 386, and 388, when used together, can embody both the parameter reconstruction 316 and the mixing calculation 402 and synthesis processor 404 as described above and below.
[0256] 4 Discussion 4.1 Overview The novel method of this example aims, in particular, to perform the encoding and decoding of multi-channel content at a low bitrate (meaning 160 kbits / sec or less) while maintaining sound quality as close as possible to the original signal and preserving the spatial characteristics of the multi-channel signal. One of the functions of the novel method is also to fit within the aforementioned DirAC framework. The output signal can be rendered with the same loudspeaker settings as the input 212 or with different settings (which can be louder or quieter depending on the loudspeakers). Also, the output signal can be rendered on the loudspeakers using binaural rendering.
[0257] In this section, the present invention and the various modules that make up the present invention will be described in detail.
[0258] The proposed system consists of two main parts. - Encoder 200. The encoder 200 derives the necessary parameters 220 from the input signal 212, quantizes them (at 222), and encodes them (at 226). The encoder 200 can also calculate a downmix signal 246 that is encoded within the bitstream 248 (and can be sent to the decoder 300). - Decoder 300. The decoder 300 uses the encoded (e.g., transmitted) parameters and the downmixed signal 246 to create a multi-channel output of quality as close as possible to the original signal 212.
[0259] Figure 1 shows an overview of the proposed novel method by way of an example. Note that some examples use only a subset of the building blocks shown in the overall figure and remove specific processing blocks depending on the application scenario.
[0260] The input 212(y) to the present invention is a multi-channel audio signal 212 (also referred to as a "multi-channel stream") in the time domain or the time-frequency domain (e.g., signal 216), and means a set of audio signals, for example, created by a set of loudspeakers or intended to be reproduced by a set of loudspeakers.
[0261] The first part of the processing is the encoding part. A so-called "downmix" signal 246 is calculated from the multi-channel audio signal together with a set of parameters or side information 228 (see 4.2.2 & 4.2.3) derived from the input signal 212 either in the time domain or in the frequency domain (see 4.2.6). These parameters are encoded (see 4.2.5) and are optionally transmitted to the decoder 300.
[0262] Next, the downmix signal 246 and the encoded parameters 228 can be transmitted to a core coder and a transmission canal that connect the encoder side and the decoder side of the process.
[0263] On the decoder side, the downmixed signal is processed (4.3.3 & 4.3.4) and the transmitted parameters are decoded (see 4.3.2). The decoded parameters are used for the synthesis of the output signal using covariance synthesis (see 4.3.5), thereby resulting in the final multi-channel output signal in the time domain.
[0264] Before going into details, there are several general characteristics to be established, and at least one of these general characteristics is effective. - The processing can be used with any loudspeaker setting. Note that increasing the number of loudspeakers increases the complexity of the process and the bits required for encoding the transmitted parameters. - The overall process can be carried out on a frame-by-frame basis. That is, the input signal 212 can be divided into frames that are processed independently. On the encoder side, each frame generates a set of parameters, and the set of parameters is transmitted to the decoder side for processing. - The frame can also be divided into slots. In this case, these slots exhibit statistical characteristics that cannot be obtained at the frame scale. The frame can be divided, for example, into 8 slots, and the length of each slot is equal to 1 / 8 of the length of the frame.
[0265] 4.2 Encoder The purpose of the encoder is to extract appropriate parameters 220 for describing the multi-channel signal 212, quantize them (at 222), encode them as side information 228 (at 226), and then, optionally, transmit them to the decoder side. Here, the parameters 220 and how they can be calculated will be described in detail.
[0266] A more detailed description of the encoder 200 can be found in FIGS. 2a - 2d. In this overview, the focus is on the two main outputs 228 and 246 of the encoder.
[0267] The first output of the encoder 200 is the downmix signal 228 calculated from the multi-channel audio input 212. The downmix signal 228 is a representation of the original multi-channel stream (signal) with fewer channels than the original content (212). Further information about its calculation can be found in Section 4.2.6.
[0268] The second output of the symbolizer 200 is the encoded parameter 220 represented as side information 228 in the bitstream 248. These parameters 220 are the gist of this example. They are parameters used to efficiently describe the multi-channel signal on the decoder side. These parameters 220 provide a good trade-off between the quality and amount of bits necessary to encode the parameters 220 into the bitstream 248. On the symbolizer side, the parameter calculation can be carried out in several steps. Although the process in the frequency domain is described, it can be executed similarly in the time domain. The parameters 220 are first estimated from the multi-channel input signal 212, then they can be quantized by the quantizer 222, and then they can be converted into the digital bitstream 248 as side information 228. Further information about these steps can be found in Sections 4.2.2, 4.2.3, and 4.2.5.
[0269] 4.2.1 Filter Bank & Partition Grouping The filter bank on the symbolizer side (e.g., filter bank 214) or the filter bank on the decoder side (e.g., filter banks 320 and / or 338) will be described.
[0270] The present invention can utilize filter banks at various points in the process. These filter banks can convert a signal from the time domain to the frequency domain (so-called aggregated band or parameter band), in which case it is called an "analysis filter bank", or can convert a signal from frequency to the time domain (e.g., 338), in which case it is called a "synthesis filter bank".
[0271] The selection of the filter bank needs to match the performance and the desired optimization requirements, but the remaining processing can be executed independently of the specific selection of the filter bank. For example, it is possible to use a filter bank based on orthogonal mirror filters or a filter bank based on the short-time Fourier transform.
[0272] Referring to FIG. 5, the output of the filter bank 214 of the encoder 200 is a signal 216 (266 with respect to 264) in the frequency domain represented over a certain number of frequency bands. It can be understood that performing the remaining processing for all frequency bands (264) provides better quality and better frequency resolution, but more important bitrates are also required to transmit all the information. Thus, a so-called “partitioning grouping” (265) corresponding to grouping some frequencies together is performed along with the filter bank process to represent the information 266 in a smaller set of bands.
[0273] For example, the output 264 of the filter 263 (FIG. 5) can be represented in 128 bands, and the partitioning grouping at 265 can result in a signal 266 (216) having only 20 bands. There are several ways to group the bands together, but one meaningful way can be, for example, to attempt to estimate the equivalent rectangular bandwidth. The equivalent rectangular bandwidth is a kind of psychoacoustically motivated band division that attempts to model how the human auditory system processes audio events, i.e., the aim is to group the filter bank in a way suitable for human hearing.
[0274] 4.2.2 Parameter Estimation (e.g., estimator 218) Aspect 1: Use of a covariance matrix for describing and synthesizing multi-channel content
[0275] The parameter estimation at 218 is one of the main points of the present invention. These are used on the decoder side to synthesize the output multi-channel audio signal. These parameters 220 (encoded as side information 228) are selected because they efficiently describe the multi-channel input stream (signal) 212 and there is no need to transmit a large amount of data. These parameters 220 are calculated on the encoder side and are later used jointly with the synthesis engine on the decoder side to calculate the output signal.
[0276] Here, a covariance matrix can be calculated between the channels of the multi-channel audio signal and the channels of the downmix signal. That is, - C y : The covariance matrix of the multi-channel stream (signal), and / or - C x : The covariance matrix of the downmix stream (signal) 246
[0277] The processing can be performed on a parameter band basis, and thus the parameter bands are independent of each other, and the equations for a given parameter band can be described without loss of generality.
[0278] For a given parameter band, the covariance matrix is defined as follows.
[0279]
Equation
[0280]
Equation
[0281] - R denotes the real part operator. - Instead of the real part, it can be any other operation that yields a real value related to the complex-valued source (e.g., the absolute value). - * denotes the conjugate transpose operator. - B indicates the relationship between the original number of bands and the grouped bands (see 4.2.1 for partitioning grouping). - Y and X are the original multi-channel signal 212 and the downmixed signal 246 in the frequency domain, respectively.
[0282] C y (or its elements, or C yOr the value obtained from that element) is also shown as the channel level and correlation information of the original signal 212. C x (Or that element, or C y Or the value obtained from that element) is also shown as the covariance information related to the downmix signal 212.
[0283] For a given frame (and band), for example, by the estimator block 218, one or two covariance matrices C y And / or C x Only can be output. Since the process is slot - based and not frame - based, various implementations can be executed regarding the relationship between the matrix for a given slot and the matrix for the entire frame. As an example, in order to output the matrix for one frame, it is possible to calculate the covariance matrices of each slot within the frame and sum them up. The definitions for calculating the covariance matrices are mathematical, but it should be noted that these matrices can be calculated in advance or at least modified if an output signal with specific characteristics is desired.
[0284] As described above, in practice, not all elements of the matrices C y And / or C x Need to be encoded within the side information 228 of the bitstream 248. In the case of C x , it is possible to simply estimate the elements from the downmix signal 246 encoded by applying Equation (1). Therefore, the encoder 200 can simply refrain from encoding any element of C x (or, more generally, the covariance information related to the downmix signal). In the case of C y (or in the case of the channel level and correlation information related to the original signal), on the decoder side, it is possible to estimate at least one of the elements of C y By using the techniques described below.
[0285] Aspect 2a: Transmission of covariance matrices and / or energy for describing and reconstructing multi-channel audio signals
[0286] As described above, covariance matrices are used for synthesis. It is possible to directly transmit these covariance matrices (or subsets thereof) from the encoder to the decoder. In some examples, the matrix C x need not necessarily be transmitted since it can be recomputed at the decoder side using the downmixed signal 246, although in some application scenarios this matrix may be required as a transmission parameter.
[0287] From an implementation perspective, for example, in order to meet certain requirements regarding bitrate, not all values within these matrices C x C y need to be encoded or transmitted. Values that are not transmitted can be estimated at the decoder side (see 4.3.2).
[0288] Aspect 2b: Transmission of inter-channel coherence and inter-channel level difference for describing and reconstructing multi-channel signals
[0289] From the covariance matrices C x C y an alternative set of parameters can be defined and used to reconstruct the multi-channel signal 212 at the decoder side. That is, these parameters can be, for example, inter-channel coherence (ICC) and / or inter-channel level difference (ICLD).
[0290] Inter-channel coherence represents the coherence between each channel of a multi-channel stream. This parameter is derived from the covariance matrix C y and can be calculated as follows (for a given parameter band and two given channels i and j).
[0291]
Equation
[0292] - ξ i,j is the ICC between channel i and channel j of the input signal 212. -
[0293]
Number
[0294] is a value within the covariance matrix of the multi-channel signal between channel i and channel j of the input signal 212, previously defined by equation (1).
[0295] The ICC values can be calculated between any channels of the multi-channel signal, which can result in a large amount of data as the size of the multi-channel signal increases. In practice, a reduced set of ICCs can be encoded and / or transmitted. In some examples, it may be necessary to define the values to be encoded and / or transmitted according to performance requirements.
[0296] For example, when processing a signal created by 5.1 (or 5.0) as a defined loudspeaker setting as defined in the ITU recommendation "ITU-R BS.2159-4", it is possible to select to transmit only four ICCs. These four ICCs are - between the center channel and the right channel - between the center channel and the left channel - between the left channel and the left surround channel - between the right channel and the right surround channel and can be any of them.
[0297] Generally, the indices of the ICCs selected from the ICC matrix are described by an ICC map.
[0298] Generally, for each loudspeaker setting, a fixed set of ICCs that on average provides the best quality can be selected to be sent to the encoder and / or decoder. The number of ICCs and which ICCs to send may depend on the loudspeaker setting and / or the total available bitrate, and both can be available at the encoder and decoder without the need to send an ICC map in the bitstream 248. In other words, the fixed set of ICCs and / or the corresponding fixed ICC map can be used, for example, depending on the loudspeaker setting and / or the total bitrate.
[0299] This fixed set may not be suitable for a particular material and, in some cases, may result in a significantly worse quality than the average quality of all materials using the fixed set of ICCs. To overcome this, in another example, for every frame (or slot), an optimal set of ICCs and the corresponding ICC map can be estimated based on the importance characteristics of the particular ICCs. The ICC map used for the current frame is then explicitly encoded and / or transmitted together with the quantized ICCs within the bitstream 248.
[0300] For example, the importance characteristics of the ICCs use the downmix covariance C from Equation (1), similar to the decoder using Equations (4) and (6) from 4.3.2 x to use the estimated value of the covariance
[0301]
Number
[0302] or the estimated value of the ICC matrix
[0303]
Number
[0304] It can be determined by generating. Depending on the selected features, for all bands where parameters are transmitted in the current frame and combined for all bands, for the corresponding entries in any ICC or covariance matrix, features are calculated. Then, this combined feature matrix is used to determine the most important ICCs, and thus the set of ICCs to use and the ICC map to transmit.
[0305] For example, the feature of ICC importance is the absolute error between the estimated covariance
[0306]
Number
[0307] entry and the entry of the actual covariance C y and the combined feature matrix is the sum of the absolute errors of all ICCs across all bands transmitted in the current frame. From the combined feature matrix, the n entries with the highest total absolute error are selected, where n is the number of ICCs transmitted for the loudspeaker / bitrate combination, and the ICC map is created from the entries.
[0308] Furthermore, in another example as in Figure 6b, in order to prevent the ICC map from being overly changed between frames, for any entry within the selected ICC map of the previous parameter frame, for example, in the case of the absolute error of covariance, the feature matrix can be emphasized by applying a coefficient >1 (220k) to the entry of the ICC map of the previous frame.
[0309] Furthermore, in another example, a flag transmitted within the side information 228 of the bitstream 248 can indicate whether a fixed ICC map or an optimal ICC map is used in the current frame, and if the flag indicates a fixed set, the ICC map is not transmitted within the bitstream 248.
[0310] The optimal ICC map is encoded and / or transmitted, for example, as a bitmap (e.g., the ICC map may embody the information 254' of FIG. 6a).
[0311] Another example of transmitting the ICC map is to transmit an index to a table of all possible ICC maps, where the index itself is, for example, additionally entropy encoded. For example, the table of all possible ICC maps is not stored in memory, and the ICC map indicated by the index is calculated directly from the index.
[0312] A second parameter that can be transmitted with (or alone) the ICC is the ICLD. "ICLD" represents the inter-channel level difference and represents the energy relationship between the channels of the input multi-channel signal 212. There is no specific definition of the ICLD. An important aspect of this value is that it represents the energy ratio within the multi-channel stream. As an example, the conversion from C y to the ICLD can be obtained as follows.
[0313]
Number
[0314] - χ i is the ICLD of channel i. - P i is the power of the current channel i, and C y of the diagonal, that is,
[0315]
Number
[0316] can be extracted from. - P dmx,i depends on channel i, but is always a linear combination of the values of C x and also depends on the original speaker settings.
[0317] In the example, P dmx,i is not the same for every channel and depends on the mapping associated with the downmix matrix (which is also the prototype matrix of the decoder), which is generally referred to in one of the bullet points below Equation (3). It depends on whether channel i is downmixed to only one of the downmix channels or to two or more of the downmix channels. In other words, P dmx,i is the sum of all the diagonal elements of C x that have non-zero elements in the downmix matrix, or may include that sum, and thus, Equation (3) can be rewritten as
[0318]
Equation
[0319] where α i is a weighting factor related to the expected energy contribution of the channel to the downmix, which is fixed for a particular input loudspeaker configuration and recognized by both the encoder and the decoder. The concept of the matrix Q is provided below. Some values of α i and the matrix Q are also described at the end of this document.
[0320] In the case of an implementation that defines the mapping for all input channels i, the mapping index is either the downmix channel j to which the input channel i is mixed alone or when the mapping index is greater than the number of downmix channels. Thus, there is a mapping index m dmx,i used to determine P ICLD,i .
[0321]
Equation
[0322] 4.2.3 Parameter Quantization An example of the quantization of parameter 220 to obtain the quantized parameter 224 can be performed, for example, by the parameter quantization module 222 of FIGS. 2b and 4.
[0323] Once the set of parameters 220 is calculated, i.e., once either the covariance matrix {C x , C y} or ICC and ICLD {ξ, χ} is calculated, they are quantized. The choice of quantizer can be a trade-off between the quality and amount of data to be transmitted, but there are no restrictions on the quantizer used.
[0324] As an example, when ICC and ICLD are used, one quantizer can be a non-linear quantizer with 10 quantization steps in the interval [-1, 1] of ICC, and another quantizer can be a non-linear quantizer with 20 quantization steps in the interval [-30, 30] of ICLD.
[0325] Also, as an optimization of the implementation, it is possible to choose to downsample the transmitted parameters, i.e., to use the quantized parameter 224 continuously in two or more frames.
[0326] In one aspect, a subset of the parameters transmitted in the current frame is signaled by a parameter frame index in the bitstream.
[0327] 4.2.4 Handling of Transients, Downsampled Parameters Some examples described hereinafter can be understood as those shown in FIG. 5, which can be an example of block 214 of FIGS. 1 and 2d.
[0328] In the case of the downsampled parameter set (e.g., obtained by block 265 in FIG. 5), i.e., the parameter set 220 of the subset of the parameter band, it can be used for two or more processed frames, and the transient phenomena appearing in two or more subsets cannot be preserved from the viewpoints of localization and coherence. Therefore, it may be advantageous to transmit the parameters of all bands within such a frame. Such a special type of parameter frame can be signaled, for example, by a flag in the bitstream.
[0329] In one aspect, in order to detect such transient phenomena in signal 212, transient phenomenon detection at 258 is used. The position of the transient phenomenon in the current frame can also be detected. The time granularity can be advantageously linked to the time granularity of the filter bank 214 used so that each transient phenomenon position can correspond to a slot or a group of slots of the filter bank 214. Then, for example, only the slots from the slot containing the transient phenomenon to the end of the current frame are used to calculate the covariance matrices C y and C x and the slots for calculating are selected.
[0330] The transient phenomenon detector (or transient phenomenon analysis block 258) can be the transient phenomenon detector also used for the coding of the downmixed signal 212, for example, the time-domain transient phenomenon detector of the IVAS core coder. Therefore, the example in FIG. 5 can also be applied upstream of the downmix calculation block 244.
[0331] In one example, the occurrence of the transient phenomenon is encoded using 1 bit (e.g., '1' meaning 'there was a transient phenomenon in the frame', and '0' meaning 'there was no transient phenomenon in the frame', etc.), and when the transient phenomenon is detected, in addition, the position of the transient phenomenon is encoded and / or transmitted as the encoded field 261 (information regarding the transient phenomenon) in the bitstream 248 to enable similar processing in the decoder 300.
[0332] If a transition phenomenon is detected and transmission in all bands is performed (e.g., signaled), when parameter 220 is transmitted using normal partitioning grouping, as a result, the data rate required to transmit parameter 220 as side information 228 within bitstream 248 may increase sharply. Furthermore, temporal resolution is more important than frequency resolution. Therefore, in block 265, it may be advantageous to change the partitioning grouping of such a frame (e.g., from signal version 264 with many bands to signal version 266 with fewer bands) to reduce the transmission band. In the example, for instance, such different partitioning grouping is used by combining two adjacent bands across all bands with respect to the normal downsampling factor 2 of the parameter. Generally, the occurrence of a transition phenomenon means that the covariance matrix itself can be expected to be significantly different before and after the transition phenomenon. To avoid artifacts in the slots before the transition phenomenon, only the transition phenomenon slot itself and all subsequent slots until the end of the frame can be considered. This is also based on the assumption that the signal is sufficiently stationary beforehand, and it is possible to use the information and mixing rules derived for the previous frame for the slots preceding the transition phenomenon.
[0333] In summary, the encoder may be configured to determine in which slot of a frame a transition phenomenon occurs and to encode the channel level and correlation information (220) of the original signal (212, y) associated with the slots occurring after the transition phenomenon and / or subsequent slots within the frame without encoding the channel level and correlation information (220) of the original signal (212, y) associated with the slots preceding the transition phenomenon.
[0334] Similarly, when the presence of a transition phenomenon and the position of the transition phenomenon within one frame are signaled (261), the decoder (e.g., in block 380) Associate the current channel level and correlation information (220) with the slot in which the transient occurred and / or subsequent slots within the frame in which the transient occurred. The channel level and correlation information (220) of the preceding slot can be associated with the slots of the frame preceding the slot in which the transient occurred.
[0335] Another important aspect of the transient phenomenon is that, if it is determined that a transient exists within the current frame, no further smoothing operation is performed on the current frame. If there is a transient, the smoothing of C y and C x is not performed, and C yR and C x from the current frame are used for calculating the mixing matrix.
[0336] 4.2.5 Entropy Coding The entropy coding module (bit stream writer) 226 can be the last encoder module, and its purpose is to convert the previously obtained quantized values into a binary bit stream, also called "side information".
[0337] The method used to encode the values can be, for example, Huffman coding [6] or delta coding. The coding method is not that important and only affects the final bit rate. The coding method should be adapted according to the desired bit rate.
[0338] To reduce the size of the bit stream 248, several implementation optimizations can be performed. For example, a switching mechanism can be implemented to switch from one coding method to another according to which is more efficient from the perspective of the bit stream size.
[0339] For example, the parameters are delta-coded along the frequency axis of one frame, and the resulting series of delta indices is entropy-coded by a range coder.
[0340] Also, in the case of parameter downsampling, as another example, a mechanism may be implemented to transmit only a subset of the parameter band for each frame in order to continuously transmit data.
[0341] In these two examples, signaling bits are required to signal decoder-specific aspects of the processing on the encoder side.
[0342] 4.2.6 Downmix Calculation The downmix section 244 of the processing is simple but can be extremely important in some examples. The downmix used in the present invention can be passive, which means that the way the downmix is calculated remains the same during processing and does not depend on the signal or its characteristics at a given time. Nevertheless, it is understood that the downmix calculation at 244 can be extended to be active (as described, for example, in [7]).
[0343] The downmix signal 246 can be calculated at two different locations. - The first time is calculated on the encoder side for parameter estimation (see 4.2.2), because the downmix signal 246 may be required for the calculation of the covariance matrix C x in some examples. - The second time is calculated on the encoder side, and between the encoder 200 and the decoder 300 (in the time domain), the downmixed signal 246 is transmitted to the encoder and / or decoder 300 and used as the basis for synthesis at module 334.
[0344] As an example, in the case of the 5.1 input stereo downmix, the downmix signal can be calculated as follows. - The left channel of the downmix is the sum of the left channel, the left surround channel, and the center channel.
[0345] The right channel of the downmix is the sum of the right channel, the right surround channel, and the center channel. Alternatively, in the case of a monaural downmix of a 5.1 input, the downmix signal is calculated as the sum of all channels of the multichannel stream.
[0346] In an example, each channel of the downmix signal 246 can be obtained as a linear combination of the channels of the original signal 212, for example using certain parameters, thereby implementing a passive downmix.
[0347] The calculation of the downmixed signal can be extended according to the processing requirements and adapted to further loudspeaker settings.
[0348] Aspect 3: Low-latency processing using passive downmix and low-latency filter bank
[0349] The present invention can provide low-latency processing by using a passive downmix, such as that described above for a 5.1 input, and a low-latency filter bank. Using these two elements, it is possible to achieve a latency of less than 5 milliseconds between the encoder 200 and the decoder 300.
[0350] 4.3 Decoder The purpose of the decoder is to synthesize an audio output signal (336, 340, y R ) using the encoded (e.g., transmitted) downmix signal (246, 324) and the encoded side information 228 for a given loudspeaker setting. The decoder 300 outputs an audio signal (334, 240, y R) can be rendered. Without loss of generality, it is assumed that the input loudspeaker and output loudspeaker settings are the same (although they may be different in the example). In this section, various modules that can constitute the decoder 300 will be described.
[0351] Figures 3a and 3b show a detailed overview of possible decoder processing. It is important to note that at least some of the modules in Figure 3b (specifically, the modules with dashed boundary lines such as 320, 330, 338, etc.) can be removed depending on the needs and requirements of a given application. The decoder 300 can receive two sets of data from the encoder 200, namely, - Side information 228 with encoded parameters (described in 4.2.2) - The downmixed signal (246, y) that can be in the time domain (described in 4.2.6) can be input (e.g., received).
[0352] The encoded parameters 228 may first need to be decoded (e.g., by the input unit 312) using a previously used inverse coding method. When this step is completed, parameters related to synthesis, such as the covariance matrix, can be reconstructed. In parallel, the downmixed signal (246, x) can be processed through several modules. First, the frequency-domain version 324 of the downmixed signal 246 can be obtained using the analysis filter bank 320 (see 4.2.1). Then, the prototype signal 328 can be calculated (see 4.3.3), and an additional decorrelation step (at 330) can be performed (see 4.3.4). The main part of the synthesis is the synthesis engine 334 that uses the covariance matrix (reconstructed, e.g., at block 316) and the prototype signal (328 or 332) as inputs to generate the final signal 336 as the output (see 4.3.5). Finally, the last step in the synthesis filter bank 338 that generates the output signal 340 in the time domain (e.g., if the analysis filter bank 320 was previously used) can be performed.
[0353] 4.3.1 Entropy Decoding (e.g., Block 312) Entropy decoding at block 312 (input interface) may enable obtaining the quantized parameter 314 obtained previously at 4. Decoding of the bitstream 248 can be understood as a simple operation. The bitstream 248 can be read according to the encoding method used at 4.2.5 and then decoded.
[0354] From an implementation perspective, the bitstream 248 may contain signaling bits that indicate some particularities of the processing on the encoder side rather than data.
[0355] For example, the first 2 bits used can indicate which coding method is being used if the encoder 200 may switch between several coding methods. Also, the next bit can be used to describe which parameter band is currently being transmitted.
[0356] Other information that may be encoded within the side information of the bitstream 248 can include a flag indicating a transient phenomenon and a field 261 indicating in which slot of the frame the transient phenomenon occurred.
[0357] 4.3.2 Parameter Reconstruction Parameter reconstruction can be performed, for example, by block 316 and / or the hybrid rule calculator 402.
[0358] The goal of this parameter reconstruction is to construct the covariance matrices C x and C y (or, more generally, the covariance information related to the downmixed signal 246 and the level and correlation information of the original signal) from the downmixed signal 246 and / or from the side information 228 (or its version represented by the quantized parameter 314). These covariance matrices C x and C ySince it efficiently describes the multi-channel signal 246, it may be essential for synthesis.
[0359] Parameter reconstruction at module 316 can be a two-step process. First, matrix C is recalculated from the downmix signal 246 x (or, more generally, the covariance information associated with the downmix signal 246) (this step can be avoided if the covariance information associated with the downmix signal 246 is actually encoded within the side information 228 of the bitstream 248). Then, for example, using at least in part the transmitted parameters and C x , more generally the covariance information associated with the downmix signal 246, matrix C y (or, more generally, the level and correlation information of the original signal 212) can be restored (this step can be avoided if the level and correlation information of the original signal 212 is actually encoded within the side information 228 of the bitstream 248).
[0360] In some examples, for each frame, it should be noted that it is possible to smooth the covariance matrix C x of the current frame, for example by addition, averaging, etc., using a linear combination with the reconstructed covariance matrix of the frames preceding the current frame. For example, in the t-th frame, the final covariance used in equation (4) can take into account the reconstructed target covariance for the preceding frames, for example, C x,t = C x,t + C x,t-1 is used.
[0361] However, if it is determined that there is a transient phenomenon within the current frame, no further smoothing operation is performed on the current frame. If there is a transient phenomenon, no smoothing is performed and C x from the current frame is used.
[0362] The outline of the process can be found below.
[0363] Note: For the encoder, the processing here can be performed independently for each band on a parameter - band basis. For clarity, the processing is described only for one particular band, and the notation is adapted accordingly.
[0364] Aspect 4a: Reconstruction of parameters when the covariance matrix is transmitted
[0365] In this aspect, it is assumed that the encoded (e.g., transmitted) parameters in side information 228 (the covariance matrix related to the down - mixed signal 246, and the channel level and correlation information of the original signal 212) are the covariance matrix (or a subset thereof) defined in aspect 2a. However, in some examples, the covariance matrix related to the down - mixed signal 246 and / or the channel level and correlation information of the original signal 212 may be embodied by other information.
[0366] The complete covariance matrix C x and C y When is encoded (e.g., transmitted), there is no further processing to be done in block 318 (thus, in such examples, block 318 can be avoided). When only a subset of at least one of these matrices is encoded (e.g., transmitted), it is necessary to estimate the missing values. The final covariance matrix used in the synthesis engine 334 (or, more specifically, in the synthesis processor 404) will be composed of the encoded (e.g., transmitted) values 228 and the estimated values on the decoder side. For example, if only some elements of the matrix C y are encoded in the side information 228 of the bitstream 248, the remaining elements of C y are estimated here.
[0367] The covariance matrix C of the down - mixed signal 246 xIn this case, on the decoder side, it is possible to calculate the missing values using the downmixed signal 246 and apply Equation (1).
[0368] In the mode where the occurrence and position of the transient phenomenon are transmitted or encoded, the same slot as that on the encoder side can be used to calculate the covariance matrix C of the downmixed signal 246. x
[0369] Covariance matrix C y In the case of the initial estimation, the missing values can be calculated as follows.
[0370]
Equation
[0371] -
[0372]
Equation
[0373] is an estimated value of the covariance matrix of the original signal 212 (this is an example of an estimated version of the original channel level and correlation information). - Q is a so-called prototype matrix (prototype rule, estimation rule) representing the relationship between the downmixed signal and the original signal (see 4.3.3) (this is an example of a prototype rule). - C x is the covariance matrix of the downmixed signal (this is an example of the covariance information of the downmixed signal 212). - * indicates conjugate transpose.
[0374] When these steps are completed, the covariance matrix can be obtained again and used for the final synthesis.
[0375] Mode 4b: Reconstruction of parameters when ICC and ICLD are transmitted
[0376] In this case, the encoded (e.g., transmitted) parameters within the side information 228 can be assumed to be the ICC and ICLD (or subsets thereof) defined in mode 2b.
[0377] In this case, it may be necessary to recalculate the covariance matrix C x first. This recalculation can be performed by using the downmixed signal 212 at the decoder side and applying Equation (1).
[0378] In the mode in which the occurrence and position of the transient phenomenon are transmitted, the same slot as that at the encoder is used to calculate the covariance matrix C x of the downmixed signal. Then, the covariance matrix C y can be recalculated from the ICC and ICLD. This operation can be performed as follows.
[0379] The energy (also called level) of each channel of the multi-channel input can be obtained. These energies are derived using the transmitted ICLD and the following equation.
[0380]
Number
[0381] where
[0382]
Number
[0383] where α iis a weight factor related to the expected energy contribution of a channel to the downmix, and this weight factor is fixed for a specific input loudspeaker configuration and recognized by both the encoder and the decoder. In the case of an implementation that defines the mapping for all input channels i, the mapping index is the downmix channel j to which input channel i is mixed alone, or the case where the mapping index is greater than the number of downmix channels. Thus, as follows, P dmx,i the mapping index m ICLD,i used to determine
[0384]
Number
[0385] The notation is the same as that used in the parameter estimation of 4.2.2.
[0386] These energies can be used to normalize the estimated C y If not all ICCs are transmitted from the encoder side, for the values not transmitted, the estimated value C y can be calculated. The estimated covariance matrix
[0387]
Number
[0388] can be obtained using Equation (4) with the prototype matrix Q and the covariance matrix C x .
[0389] This estimation of the covariance matrix leads to the estimation of the ICC matrix, in which the term at index (i,j) is
[0390]
Number
[0391] can be given by
[0392] Thus, the "reconstructed" matrix can be defined as follows.
[0393]
Number
[0394] where - The subscript R indicates the reconstructed matrix (this is an example of the reconstructed version of the original level and correlation information). - The set {transmitted indices} corresponds to all (i,j) pairs decoded within the side information 228 (e.g., transmitted from the encoder to the decoder).
[0395] In the example,
[0396]
Number
[0397] is not as accurate as the encoded value ξ i,j Therefore,
[0398]
Number
[0399] ξ i,j can be preferred.
[0400] Finally, from this reconstructed ICC matrix, the reconstructed covariance matrix
[0401]
Number
[0402] can be estimated. This matrix can be obtained by applying the energy obtained in Equation (5) to the reconstructed ICC matrix. Thus, for index (i,j),
[0403]
Number
[0404] it is.
[0405] When the complete ICC matrix is transmitted, only Equations (5) and (8) are required. The previous paragraph shows one method for reconstructing the missing parameters, but other methods can be used, and the proposed method is not unique.
[0406] From the example of Aspect 1b using the 5.1 signal, note that the values not transmitted are the values that need to be estimated on the decoder side.
[0407] Now, the covariance matrix C x and
[0408]
Number
[0409] is obtained. Note that the reconstructed matrix
[0410]
Number
[0411] can be an estimated value of the covariance matrix C y of the input signal 212. The trade-off of the present invention can be to bring the estimated value of the covariance matrix on the decoder side close enough to the original matrix while minimizing the parameters to be transmitted. These matrices can be essential for the final synthesis shown in 4.3.5.
[0412] Note that in some examples, for each frame, it is possible to smooth the reconstructed covariance matrix of the current frame, for example, by addition, averaging, etc., using a linear combination with the reconstructed covariance matrix of the frame preceding the current frame. For example, in the t-th frame, the final covariance used for synthesis can take into account the target covariance reconstructed for the preceding frame, for example,
[0413] [Number]
[0414] is.
[0415] However, in the presence of transient phenomena, smoothing is not performed, and C for the current frame yR is used for the calculation of the mixing matrix.
[0416] In some examples, for each frame, the non-smoothed covariance matrix of the downmix channel C x is used for parameter reconstruction, and the smoothed covariance matrix C x,t described in Section 4.2.3 is used for synthesis.
[0417] Figure 8a resumes the operation for obtaining the covariance matrices C x and
[0418] [Number]
[0419] in the decoder 300 (as executed, for example, in block 386 or 316...). In the blocks of Figure 8a, between the parentheses, the equations adopted by the specific blocks are also shown. As shown in the figure, the covariance estimator 384 calculates the covariance C of the downmix signal 324 (or its reduced-band version 385) via equation (1) xto reach. The first covariance block estimator 384' uses Equation (4) and an appropriate type rule Q to obtain the first estimate of covariance C y of
[0420]
Number
[0421] to reach. Subsequently, the covariance coherence block 390 applies Equation (6) to obtain coherence
[0422]
Number
[0423] . Subsequently, the ICC replacement block 392 adopts Equation (7) to select either the estimated ICC(
[0424]
Number
[0425] ) or the ICC signaled in the side information 228 of the bitstream 348. Then, the selected coherence ξ R is input to the energy application block 394 that applies energy according to ICLD(χ i ). Then, the target covariance matrix
[0426]
Number
[0427] is provided to the mixer rule calculator 402 or the covariance synthesis block 388 of Figure 3a, or the mixer rule calculator of Figure 3c, or the synthesis engine 344 of Figure 3b.
[0428] 4.3.3 Calculation of Prototype Signal (Block 326) The purpose of the prototype signal module 326 is to shape the downmix signal 212 (or its frequency domain version 324) so that it can be used by the synthesis engine 334 (see 4.3.5). The prototype signal module 326 can perform upmixing of the downmixed signal. The calculation of the prototype signal 328 can be performed by the prototype signal module 326 by multiplying the downmixed signal 212 (or 324) by a so-called prototype matrix Q. Y p = XQ (9) - Q is a prototype matrix (an example of a prototype rule). - X is the downmixed signal (212 or 324). - Y p is the prototype signal (328).
[0429] The method of establishing the prototype matrix may depend on the process and can be defined to meet the requirements of the application. The only constraint may be that the number of channels of the prototype signal 328 must be the same as the number of desired output channels. This directly constrains the size of the prototype matrix. For example, Q can be a matrix having the number of rows that is the number of channels of the downmix signal (212, 324) and the number of columns that is the number of channels of the final synthesis output signal (332, 340).
[0430] As an example, in the case of a 5.1 signal or a 5.0 signal, the prototype matrix can be established as follows.
[0431]
Equation
[0432] Note that the prototype matrix can be determined and fixed in advance. For example, Q can be the same for all frames, but can be different for different bands. Furthermore, Q is different when the relationship between the number of channels of the downmix signal and the number of channels of the combined signal is different. Q can be selected from a plurality of pre-stored Qs based on, for example, a specific number of downmix channels and a specific number of combined channels.
[0433] Aspect 5: Reconstruction of parameters when the output loudspeaker setting is different from the input loudspeaker setting
[0434] One application of the present invention proposed is to generate the output signal 336 or 340 with a loudspeaker setting different from that of the original signal 212 (for example, meaning a larger or smaller number of loudspeakers).
[0435] For this purpose, it is necessary to modify the prototype matrix accordingly. In this scenario, the prototype signal obtained in Equation (9) will include the same number of channels as the output loudspeaker setting. For example, if there is a 5-channel signal as input (on the signal 212 side) and a 7-channel signal is to be obtained as output (on the signal 336 side), the prototype signal already includes 7 channels.
[0436] When this is done, the estimation of the covariance matrix in Equation (4) is still valid and will continue to be used to estimate the covariance parameters of the channels that did not exist in the input signal 212.
[0437] The parameter 228 transmitted between the encoder and the decoder is still relevant, and Equation (7) can likewise continue to be used. More precisely, the encoded (for example, transmitted) parameter needs to be assigned to channel pairs as close as possible to the original setting from a geometric point of view. Basically, it is necessary to perform an adaptive operation.
[0438] For example, on the encoder side, if the ICC value is estimated between one right loudspeaker and one left loudspeaker, this value can be assigned to the channel pair of the output setting having the same left and right positions. If the geometries are different, this value can be assigned to the loudspeaker pair at a position as close as possible to the original position.
[0439] Next, when the target covariance matrix C y of the new output setting is obtained, the remaining processing is not changed.
[0440] Therefore, in order to adapt the target covariance matrix (
[0441]
Number
[0442] ) to the number of composite channels, it is possible to use the prototype matrix Q that converts from the number of downmix channels to the number of composite channels, and this prototype matrix Q adapts Equation (9) so that the prototype signal has the number of composite channels, adapts Equation (4), and thus estimates
[0443]
Number
[0444] at the number of composite channels, maintains Equations (5) to (8), thereby obtaining Equations (5) to (8) at the number of original channels, but assigns the group of original channels (for example, the pair of original channels) to a single composite channel (for example, selects the assignment from a geometric perspective), or vice versa to obtain.
[0445] An example is shown in FIG. 8b. FIG. 8b is a version of FIG. 8a and shows the number of channels for several matrices and vectors. When the ICC (obtained from the side information 228 of the bitstream 348) is applied to the ICC matrix at 392, a group of original channels (e.g., a pair of original channels) is applied to a single combined channel (e.g., select an assignment from a geometric perspective), or vice versa.
[0446] Another executable method for generating a target covariance matrix where the number of input channels is different from the number of output channels is to first generate a target covariance matrix for the number of input channels (e.g., the number of original channels of the input signal 212), and then adapt this first target covariance matrix to the number of combined channels to obtain a second target covariance matrix corresponding to the number of output channels. This involves applying a matrix including coefficients for a combination of specific input (original) channels and output channels, e.g., an upmix rule or a downmix rule, to the first target covariance matrix
[0447]
Number
[0448] and in a second step, applying this matrix
[0449]
Number
[0450] This can be done by applying it to the transmitted input channel power (ICLD), obtaining a vector of channel powers for the number of output (synthesized) channels, and adjusting the first target covariance matrix according to the vector to obtain a second target covariance matrix having the required number of synthesized channels. At this point, this adjusted second target covariance matrix can be used in the synthesis. An example of this is shown in FIG. 8c. FIG. 8c is a version of FIG. 8a, and blocks 390 to 394 operate to reconstruct the target covariance matrix
[0451]
Number
[0452] to have the number of original channels of the original signal 212. Then, at block 395, the prototype signal Q N (for conversion to the number of synthesized channels) and the vector ICLD can be applied. In particular, block 386 in FIG. 8c is the same as block 386 in FIG. 8a, except that in FIG. 8c the number of channels of the reconstructed target covariance is exactly the same as the number of original channels of the input signal 212 (in FIG. 8a, generally the reconstructed target covariance has the number of synthesized channels).
[0453] 4.3.4 Decorrelation The purpose of the decorrelation module 330 is to reduce the amount of correlation between each channel of the prototype signal. Highly correlated loudspeaker signals can cause phantom sources and may degrade the quality and spatial characteristics of the output multichannel signal. This step is optional and may or may not be implemented depending on the requirements of the application. In the present invention, decorrelation is used before the synthesis engine. As an example, an all-pass frequency decorrelator can be used.
[0454] Notes on MPEG Surround In MPEG surround according to the prior art, a so-called "mixing matrix" (denoted as M1 and M2 in the standard) is used. Matrix M1 controls how the available downmixed signals are input to the decorrelator. Matrix M2 represents how the direct signal and the uncorrelated signal are combined to generate the output signal.
[0455] There may be similarities with the prototype matrix defined in 4.3.3 and with the use of the decorrelator described in this section, but it is important to note the following points. - The prototype matrix Q has a completely different function from the matrices used in MPEG surround, and the point of this matrix is to generate a prototype signal. The purpose of this prototype signal is to be input to the synthesis engine. - The prototype matrix is not for preparing the downmixed signals of the decorrelator and can be adapted according to requirements and target applications. For example, the prototype matrix can generate a prototype signal for the output loudspeaker setting that is larger than the prototype signal for the input loudspeaker setting. - The use of the decorrelator in the proposed invention is not essential. The processing depends on the use of the covariance matrix within the synthesis engine (see 5.1). - The proposed invention does not generate the output signal by combining the direct signal and the uncorrelated signal. - The calculations of M1 and M2 strongly depend on the tree structure, and the various coefficients of these matrices are case-dependent from a structural perspective. This is not the case in the proposed invention. The processing is not concerned with the downmix calculation (see 5.2), and conceptually, the proposed processing aims to consider the relationships not only between channel pairs but also between all channels so that it can be executed using the tree structure.
[0456] Therefore, the present invention is different from MPEG surround according to the prior art.
[0457] 4.3.5 Synthesis Engine, Matrix Calculation The last step of the decoder involves a synthesis engine 334 or a synthesis processor 402 (optionally including a synthesis filter bank 338 as an addition). The purpose of the synthesis engine 334 is to generate a final output signal 336 based on certain constraints. The synthesis engine 334 can calculate an output signal 336 whose characteristics are constrained by the input parameters. In the present invention, except for the prototype signal 328 (or 332), the input parameters 318 of the synthesis engine 338 are the covariance matrices C x and C y respectively. In particular,
[0458]
Number
[0459] should be made as close as possible to those defined by C y and is thus called the target covariance matrix (it can be seen that the estimated and pre - constructed versions of the target covariance matrix are being described).
[0460] As an example, the synthesis engine 334 that can be used is not unique, and as an example, the covariance synthesis of the prior art [8], which is incorporated herein by reference, can be used. Another synthesis engine 333 that can be used is the one described in the DirAC process of [2].
[0461] The output signal of the synthesis engine 334 may require additional processing via the synthesis filter bank 338.
[0462] As a final result, an output multi - channel signal 340 in the time domain is obtained.
[0463] Aspect 6: High - quality output signal using "covariance synthesis"
[0464] As described above, the synthetic engine 334 used is not proprietary, and any engine that uses the transmitted parameters or a subset thereof can be used. Nevertheless, one aspect of the present invention may be to provide a high-quality output signal 336, for example, by using covariance synthesis [8].
[0465] This synthesis method aims to calculate an output signal 336 whose characteristics are defined by the covariance matrix
[0466]
Number
[0467] To do so, so-called optimal mixing matrices are calculated, and these matrices mix the prototype signal 328 into the final output signal 336, and the target covariance matrix
[0468]
Number
[0469] provides an optimal result from a mathematical point of view when given.
[0470] The mixing matrix M is a matrix that converts the prototype signal x R = Mx P into the output signal y P by the relationship R (336).
[0471] The mixing matrix is also a matrix that converts the downmix signal x into the output signal by the relationship y R = Mx. From this relationship,
[0472]
Number
[0473] can also be estimated.
[0474] In the indicated process,
[0475]
Number
[0476] and C x is, in some examples, (respectively the target covariance matrix
[0477]
Number
[0478] and the covariance matrix C of the downmix signal 246) and may already be recognized. x so).
[0479] One solution from a mathematical perspective is
[0480]
Number
[0481] given by, where K y and
[0482]
Number
[0483] are all C x and
[0484]
Number
[0485] It is a matrix obtained by performing singular value decomposition on [a certain matrix]. Regarding P, although P is a free parameter here, an optimal solution (from the perspective of the listener's perception) for the constraints specified by the prototype matrix Q can be found. The mathematical proof of the content described here can be found in [8].
[0486] Since the method is designed to provide an optimal mathematical solution for reconstructing the output signal problem, this synthesis engine 334 provides a high-quality output 336.
[0487] From a non-mathematical perspective, it is important to understand that the covariance matrix represents the energy relationship between different channels of a multi-channel audio signal. The matrix C of the original multi-channel signal 212 y and the matrix C of the downmixed multi-channel signal 246 x . Each value of these matrices violates the energy relationship between two channels of the multi-channel stream.
[0488] Therefore, the philosophy behind covariance synthesis is to generate a signal whose characteristics are caused by the target covariance matrix
[0489]
Number
[0490] This matrix
[0491]
Number
[0492] is calculated to represent the original input signal 212 (or, if different from the input signal, the output signal to be obtained). Then, covariance synthesis optimally mixes the prototype signals using these elements to generate the final output signal.
[0493] In a further aspect, the mixing matrix used for the synthesis of the slots is a combination of the mixing matrix M of the current frame and the mixing matrix M p of the previous frame, for example, linear interpolation based on the slot index within the current frame.
[0494] In a further aspect in which the occurrence and position of the transient phenomenon are transmitted, the previous mixing matrix M p is used for all slots before the transient phenomenon position, and the mixing matrix M is used for the slots including the transient phenomenon position and all subsequent slots within the current frame. It should be noted that in some examples, for each frame or slot, it is possible to smooth the mixing matrix of the current frame or slot, for example, by addition, averaging, etc., using a linear combination with the mixing matrix used for the preceding frame or slot. For the current frame t, the slot s band i of the output signal is Y s,i =M s,i X s,i is assumed to be obtained by. Wherein M s,i is the mixing matrix M t-1,i used for the previous frame and M t,i is the mixing matrix calculated for the current frame, for example, linear interpolation between them, that is,
[0495]
Number
[0496] where n s is the number of slots within the frame (for example, 16), and t-1 and t indicate the previous frame and the current frame. More generally, the mixing matrix M s,i associated with each slot is to scale the mixing matrix M t,i calculated for the current frame along the subsequent slots of the current frame t by an increasing coefficient, and the scaled mixing matrix M t-1,ican be obtained by adding along subsequent slots of the current frame t by a decreasing coefficient. The coefficient can be linear.
[0497] (For example, signaled in information 261) If there is a transient phenomenon, the current mixing matrix and the past mixing matrix are not combined, and it can be determined that the previous mixing matrix extends up to the slot including the transient phenomenon, and the current mixing matrix extends over all subsequent slots up to the end of the slot and frame including the transient phenomenon.
[0498] [Number]
[0499] where s is the slot index, i is the band index, t and t - 1 indicate the current frame and the previous frame, and s t is the slot including the transient phenomenon.
[0500] Differences from the prior art document [8] It is also important to note that the proposed invention goes beyond the scope of the method proposed in [8]. The notable differences are, among others, as follows. - Target covariance matrix
[0501] [Number]
[0502] is calculated on the encoder side of the proposed process. - Target covariance matrix
[0503] [Number]
[0504] can also be calculated in another way (in the proposed invention, the covariance matrix is not the sum of the diffusion part and the direct part). - The processing is not performed individually for each frequency band, but is grouped for each parameter band (as described in 0). - From a more global perspective, the covariance synthesis is here only one block of the whole process and must be used together with all other elements on the decoder side.
[0505] 4.3. List of Preferred Embodiments At least one of the following embodiments may characterize the present invention. 1. Encoder side a. Input the multi-channel audio signal 246. b. Use the filter bank 214 to convert the signal 212 from the time domain to the frequency domain (216). c. Calculate the downmix signal 246 in block 244. d. From the original signal 212 and / or the downmix signal 246, a first set of parameters for describing the multi-channel stream (signal) 246, namely the covariance matrix C x and / or C y is estimated. e. The covariance matrix C x and / or C y is directly transmitted and / or encoded, or ICC and / or ICLD are calculated and transmitted. f. Using an appropriate coding method, encode the transmitted parameter 228 into the bitstream 248. g. Calculate the downmixed signal 246 in the time domain. h. Transmit the side information (i.e., parameters) and the downmixed signal 246 in the time domain. 2. Decoder side a. Decode the bitstream 248 containing the side information 228 and the downmix signal 246. b. (Optionally) Apply the filter bank 320 to the downmix signal 246 to obtain a version 324 of the downmix signal 246 in the frequency domain. c. From the previously decoded parameter 228 and the downmix signal 246, reconstruct the covariance matrix C x、 and
[0506]
Number
[0507] do so. d. Calculate the prototype signal 328 from the downmix signal 246(324). e. Optionally (in block 330) decorrelate the prototype signal. f. Apply the synthesis engine 334 to the prototype signal using the reconstructed C x and
[0508]
Number
[0509] do so. g. Optionally apply the synthesis filter bank 338 to the output 336 of the covariance synthesis 334. h. Obtain the output multichannel signal 340.
[0510] 4.5 Covariance Synthesis In this section, several techniques that can be implemented within the system of FIGS. 1-3d are described. However, these techniques can also be implemented alone. For example, in some cases, the covariance calculations performed in FIGS. 8a-8c and equations (1)-(8) are not necessary. Thus, in some cases
[0511]
Number
[0512] when referring to (the reconstructed target covariance), this is referred to as C (which could be provided directly similarly without reconstruction) yIt can also be replaced. Nevertheless, the techniques in this section can be advantageously used together with the above techniques.
[0513] Next, refer to FIGS. 4a-4d. Here, examples of the covariance synthesis blocks 388a-388d will be described. The blocks 388a-388d can embody, for example, the block 388 of FIG. 3c for performing covariance synthesis. The blocks 388a-388d can be, for example, part of the synthesis processor 404 and the mixing rule calculator 402 of the synthesis engine 334 of FIG. 3a, and / or the parameter reconstruction block 316. In FIGS. 4a-4d, the downmix signal 324 is in the frequency domain FD (i.e., downstream of the filter bank 320) and is indicated by X, and the synthesis signal 336 is also in the FD and is indicated by Y. However, it is possible to generalize these results, for example, in the time domain. Each of the covariance synthesis blocks 388a-388d of FIGS. 4a-4d can be referenced for one single frequency band (e.g., when decomposed at 380), and thus, the covariance matrix C x and
[0514]
Number
[0515] (or other reconstructed information) can be associated with one specific frequency band. Covariance synthesis can be performed, for example, in a frame-by-frame manner, in which case the covariance matrix C x and
[0516]
Number
[0517] (or other reconstructed information) is associated with one single frame (or multiple consecutive frames). Thus, covariance synthesis can be performed in a frame-by-frame manner or in a multiple-frame-by-frame manner.
[0518] In FIG. 4a, the covariance synthesis block 388a can be constituted by one energy-compensated optimal mixing block 600a, and the correlator block is absent. Basically, one single mixing matrix M is found, and the only important operation additionally performed is the calculation of the energy-compensated mixing matrix M'.
[0519] FIG. 4b shows a covariance synthesis block 388b inspired by [8]. The covariance synthesis block 388b may be capable of obtaining the composite signal 336 as a composite signal having a first principal component 336M and a second residual component 336R. The principal component 336M is, in the optimal principal component mixing matrix 600b, for example, the covariance matrix C x and
[0520]
Number
[0521] from which the mixing matrix M M without decorrelator can be obtained, and the residual component 336R can be obtained in another way. M R should, in principle,
[0522]
Number
[0523] satisfy the relationship of. Usually, the obtained mixing matrix does not fully satisfy this, and the residual target covariance is
[0524]
Number
[0525] can be found by. As shown in the figure, the downmix signal 324 can be directed to path 610b (path 610b can be called a second path parallel to the first path 610b' including block 600b). The (Y pR shown) prototype version 613b can be obtained in the prototype signal block (upmix block) 612b. For example, an equation such as equation (9), that is, Y pR =XQ can be used.
[0526] Examples of Q (prototype matrix or upmixing matrix) are provided in this document. Downstream of block 612b, there is a decorrelator 614b for decorrelating the prototype signal 613b to obtain a decorrelated signal 615b (
[0527]
Number
[0528] also shown by). In block 616b, from the decorrelated signal 615b, the covariance matrix of the decorrelated signal
[0529]
Number
[0530] (615b) is estimated. In the optimal residual component mixing matrix block 618b, the decorrelated signal
[0531]
Number
[0532] is
[0533]
Number
[0534] Covariance matrix
[0535]
Math
[0536] is used as the equivalent of C in the principal component mixture, and by using C x as the target covariance in another optimal mixture block, the residual component 336R of the composite signal 336 can be obtained. The optimal residual component mixing matrix block 618b mixes the uncorrelated signals 615b to obtain the residual component 336R of the composite signal 336 (in a specific band), and the mixing matrix M r is implemented in such a way that it is generated. In the adder block 620b, the residual component 336R is summed with the principal component 336M (thus, paths 610b and 610b' are coupled together in the adder block 620b). R
[0537] FIG. 4c shows an example of a covariance synthesis 388c that is an alternative to the covariance synthesis 388b of FIG. 4b. The covariance synthesis block 388c enables the composite signal 336 to be obtained as a signal Y having a first principal component 336M' and a second residual component 336R'. The principal component 336M' is, for example, in the optimal principal component mixing matrix 600c, the covariance matrix C x and
[0538]
Math
[0539] (or C y , other information 220) to obtain the mixing matrix M without a correlator M can be obtained by finding, and the residual component 336R' can be obtained in another way. The downmix signal 324 can be directed to path 610c (path 610c can be called a second path in parallel with a first path 610c' that includes block 600c). In downmix block (upmix block) 612c, a prototype version 613c of the downmix signal 324 can be obtained by applying a prototype matrix Q (e.g., a matrix that upmixes the downmixed signal 234 to a version 613c of the downmixed signal 234 by the number of channels, which is the number of synthesis channels). For example, an equation such as equation (9) can be used. Examples of Q are provided in this document. Downstream of block 612c, a decorrelator 614c can be provided. In some examples, there is no decorrelator in the first path and there is a decorrelator in the second path.
[0540] The decorrelator 614c can provide an uncorrelated signal 615c (
[0541]
Number
[0542] also indicated by). However, contrary to the technique used in the covariance synthesis block 388b of FIG. 4b, in the covariance synthesis block 388c of FIG. 4c, the covariance matrix of the uncorrelated signal 615c
[0543]
Number
[0544] is not estimated from the uncorrelated signal 615c (
[0545]
Number
[0546] ). In contrast, the covariance matrix of the uncorrelated signal 615c
[0547]
Number
[0548] is (in block 616c) (for example, in block 384 of FIG. 3c and / or estimated using Equation (1)) the covariance matrix C of the downmix signal 324 x , and the prototype matrix Q is obtained from.
[0549] In the optimal residual component mixing matrix block 618c, the covariance matrix C of the downmix signal 324 x the estimated covariance matrix from
[0550]
Number
[0551] is used as the equivalent of C of the principal component mixing matrix, and C x is used as the target covariance matrix, whereby the residual component 336R' of the composite signal 336 is obtained. The optimal residual component mixing matrix block 618c may be implemented in such a way that the residual component mixing matrix M r is generated to obtain the residual component 336R' by mixing the uncorrelated signals 615c according to the residual component mixing matrix M R In the adder block 620c, to obtain the composite signal 336, the residual component 336R' is summed with the principal component 336M' (thus, paths 610c and 610c' are combined together in the adder block 620c). R can be generated.
[0552] In some examples, the residual component 336R or 336R' is not always or necessarily calculated (path 610b or 610c is not always used). In some examples, for some bands, covariance synthesis is performed without calculating the residual signal 336R or 336R', while for other bands in the same frame, covariance synthesis is processed considering the residual signal 336R or 336R'. FIG. 4d shows an example of a covariance synthesis block 388d that can be a specific case of the covariance synthesis block 388b or 388c. Here, the band selector 630 can select or deselect (in the manner represented by the switch 631) the calculation of the residual signal 336R or 336R'. For example, path 610b or 610c can be selectively enabled by the selector 630 for some bands and disabled for other bands. Specifically, path 610b or 610c can be disabled for bands exceeding a predefined threshold (e.g., a fixed threshold) that can be a threshold (e.g., a maximum value) distinguishing between a band where the human ear is less affected by phase (bands with frequencies above the threshold) and a band where the human ear is more affected by phase (bands with frequencies below the threshold), and as a result, the residual component 336R or 336R' is not calculated for bands with frequencies below the threshold and is calculated for bands with frequencies above the threshold.
[0553] The example of FIG. 4d can also be obtained by replacing block 600b or 600c with block 600a of FIG. 4a and replacing block 610b or 610c with the covariance synthesis block 388b of FIG. 4b or the covariance synthesis block 388c of FIG. 4c.
[0554] Here, some instructions regarding how to obtain the mixing rule (matrix) at any of blocks 338, 402 (or 404), 600a, 600b, 600c, etc. are provided. As described above, there are multiple ways to obtain the mixing matrix, and some of them will be described in detail here.
[0555] Specifically, first, refer to the covariance synthesis block 388b in FIG. 4b. In the optimal principal component mixing matrix block 600c, the mixing matrix M of the principal component 336M of the synthesized signal 336 is, for example, the covariance matrix C of the original signal 212 y (C y can be estimated using at least some of the above equations (6) to (8). For example, refer to FIG. 8. This is, for example, the so-called "target version" estimated using equation (8)
[0556]
Number
[0557] can be in the form of), and the covariance matrix C of the downmix signals 246, 324 x (C y can be estimated using, for example, equation (1)) can be obtained from.
[0558] For example, as proposed by [8], the covariance matrices C x and C y which are Hermitian and positive semi-definite, are factorized according to the following factorization, namely
[0559]
Number
[0560] is recognized to be decomposed.
[0561] For example, by applying singular value decomposition (SVD) twice to C x and C y K x and K y can be obtained. For example, the SVD for C x is the matrix U of eigenvectors (for example, left eigenvectors) Cx and the diagonal matrix S of singular valuesCx and can be provided, and as a result, S Cx has a diagonal matrix with the square root of the value in the corresponding entry of S in the entry, and multiplies it by U Cx to obtain K x is obtained.
[0562] Furthermore, the SVD for C y is a matrix V of singular vectors (e.g., right singular vectors) Cy and a diagonal matrix S of singular values Cy and can be provided, and as a result, S Cy has a diagonal matrix with the square root of the value in the corresponding entry of S included in the entry, and multiplies it by U Cy to obtain K y is obtained.
[0563] Next, the principal component mixing matrix M M can be obtained, and the principal component mixing matrix M M makes it possible to obtain the principal component 336M of the composite signal 336 when applied to the downmix signal 324. The principal component mixing matrix M M can be obtained as follows.
[0564]
Equation
[0565] K x If K is a non-invertible matrix, a regularized inverse matrix is obtained using known techniques,
[0566]
Equation
[0567] can be substituted instead.
[0568] The parameter P is generally a free parameter but can be optimized. To reach P, SVD is applied to C x (the covariance matrix of the downmix signal 324), and
[0569] [Number]
[0570] (the covariance matrix of the prototype signal 613b).
[0571] When SVD is executed, P can be obtained as follows. P = VΛU *
[0572] Λ is a matrix having the same number of rows as the number of synthesis channels and the same number of columns as the number of downmix channels. Λ is the identity in the first square block and has zeros entered in the remaining entries. Here, V and U are obtained from C x and
[0573] [Number]
[0574] will be described. V and U are matrices of eigenvectors obtained from SVD, that is,
[0575] [Number]
[0576] and S is a diagonal matrix of singular values typically obtained via SVD.
[0577] [Number]
[0578] is the prototype signal
[0579]
Number
[0580] (615b) is a diagonal matrix that normalizes the energy of each channel of the combined signal y to the energy of the combined signal y.
[0581]
Number
[0582] To obtain
[0583]
Number
[0584] , that is, the prototype signal
[0585]
Number
[0586] (164b), it is necessary to calculate the covariance matrix. Next,
[0587]
Number
[0588] from
[0589]
Number
[0590] to reach
[0591]
Number
[0592] The diagonal values of are normalized to the corresponding diagonal values of Cy, and thus,
[0593]
Number
[0594] are provided. As an example,
[0595]
Number
[0596] the diagonal entries of are
[0597]
Number
[0598] calculated as, where in the formula,
[0599]
Number
[0600] is the value of the diagonal entry of C y and
[0601]
Number
[0602] is
[0603]
Number
[0604] the value of the diagonal entry of.
[0605]
Number
[0606] When [is obtained],
[0607]
Number
[0608] from, the covariance matrix C of the residual component r is obtained.
[0609] C r When [is obtained], it is possible to obtain a mixing matrix for mixing the uncorrelated signal 615b and obtain the residual signal 336R. In the same optimal mixing, C r is in the main optimal mixing
[0610]
Number
[0611] has the same role as, and the covariance of the uncorrelated prototype
[0612]
Number
[0613] is, C x assumes the role of the covariance of the input signal where [had the main optimal mixing].
[0614] However, it is understood that the technique of FIG. 4c presents several advantages compared to the technique of FIG. 4b. In some examples, the technique of FIG. 4c is at least the same as the technique of FIG. 4b for calculating the main matrix and generating the main components of the composite signal. Conversely, the technique of FIG. 4c is different from the technique of FIG. 4b in calculating the residual mixing matrix and more generally in generating the residual components of the composite signal. Next, with reference to FIG. 11 in connection with FIG. 4c, the calculation of the residual mixing matrix will be described. In the example of FIG. 4c, a decorrelator 614c in the frequency domain is used, and the decorrelator 614c ensures the decorrelation of the prototype signal 613c while retaining the energy of the prototype signal 613b itself.
[0615] Furthermore, in the example of FIG. 4c, the uncorrelated channels of the uncorrelated signal 615c are mutually incoherent, and thus it can be assumed (at least approximately) that all off-diagonal elements of the covariance matrix of the uncorrelated signal are zero. Using both assumptions, by applying Q to C x the covariance of the uncorrelated prototype can be easily estimated, and only the main diagonal of the covariance (i.e., the energy of the prototype signal) can be obtained. This technique of FIG. 4c is more efficient than the estimation of the example of FIG. 4b from the uncorrelated signal 615b that requires the same band / slot aggregation as already done for C x . Therefore, in the example of FIG. 4c, the matrix multiplication of the already aggregated C x can be easily applied. Therefore, the same mixing matrix is calculated for all bands in the same group of aggregated bands.
[0616] Therefore, at 710, the covariance 711 of the uncorrelated signal (
[0617]
Number
[0618] ) is P decorr =diag(QC x Q*) is taken as the input signal covariance
[0619] [Number]
[0620] can be estimated by using, as the main diagonal of a matrix in which all off - diagonal elements are set to zero, which is used as such. To execute the synthesis of the main component 336M' of the composite signal, C x In the example where P is smoothed, C decorr used for the calculation of x The version of C, which is not smoothed, x A technique can be used that it is.
[0621] Here, the prototype matrix Q r should be used. However, note that for the residual signal, Q r is the identity matrix.
[0622] [Number]
[0623] (diagonal matrix) and Q r (identity matrix) leads to further simplification in the calculation of the mixing matrix (at least one SVD can be omitted). See the following techniques and Matlab listings.
[0624] First, similar to the example of Figure 4b, the residual target covariance matrix C r (Hermitian, positive semi - definite) of the input signal 212 is
[0625] [Number]
[0626] can be decomposed as. The matrix K r can be obtained via SVD(702). The SVD702 applied to C r is Matrix U of eigenvectors Cr (e.g., left eigenvectors), and Diagonal matrix S of singular values Cr and are generated, and as a result, (at 706) S Cr is multiplied by U to obtain a diagonal matrix (this diagonal matrix is obtained at 704) having the square root of the value in the corresponding entry of Cr in the entry, and K r is obtained.
[0627] At this point, theoretically, it may be possible to apply another SVD. This time, for the covariance of the uncorrelated prototype
[0628]
Number
[0629] is applied.
[0630] However, in this example (Figure 4c), another path is chosen to reduce the computational complexity. P decorr =diag(QC x Q*) is estimated from
[0631]
Number
[0632] is a diagonal matrix, and thus, SVD is not necessary (the SVD of a diagonal matrix gives the singular values as sorted vectors of the diagonal elements, and the left and right eigenvectors only indicate the sorting indices). (at 712)
[0633]
Number
[0634] By calculating the square root of each value at the diagonal entry of
[0635] [Number]
[0636] is obtained. This diagonal matrix
[0637] [Number]
[0638] is
[0639] [Number]
[0640] is of the form such that
[0641] [Number]
[0642] has the advantage that SVD is not required to obtain. The diagonal covariance of uncorrelated signals
[0643] [Number]
[0644] From, the estimated covariance matrix of the uncorrelated signal 615c
[0645] [Number]
[0646] is calculated. However, since the prototype matrix is Q r (i.e., the identity matrix),
[0647] [Number]
[0648] be directly used to
[0649]
Number
[0650] be
[0651]
Number
[0652] can be formulated as, where
[0653]
Number
[0654] is the value of the diagonal entry of C r and
[0655]
Number
[0656] is
[0657]
Number
[0658] the value of the diagonal entry of
[0659]
Number
[0660] is an uncorrelated signal
[0661]
Number
[0662] It is a diagonal matrix that normalizes the energy of each channel of (615b) to the desired energy of the composite signal y (obtained at 722).
[0663] At this point, (in 734)
[0664]
Number
[0665] to
[0666]
Number
[0667] it is possible to multiply (the result 735 of multiplication 734 is
[0668]
Number
[0669] also called). Then (736), K r to
[0670]
Number
[0671] is multiplied by
[0672]
Number
[0673] is obtained. To obtain the left singular vector matrix U and the right singular vector matrix V, SVD (738) can be performed from K'. y By multiplying V and U* (740), the matrix P is obtained (P = VU*)H )。Finally (742),
[0674]
Number
[0675] By applying, the mixing matrix M of the residual signal R can be obtained, where
[0676]
Number
[0677] (obtained at 745) can be replaced by the regularized inverse matrix. Therefore, M R can be used for residual mixing in block 618c.
[0678] Matlab code for performing covariance synthesis as described above is provided here. Note that in the code, the asterisk (*) means multiplication and the apex (') means the Hermitian matrix. %Compute residual mixing matrix function [M]=ComputeMixingMatrixResidual(C_hat_y,Cr,reg_sx,reg_ghat) EPS_=single(1e-15); %Epsilon to avoid divisions by zero num_outputs=size(Cr,1); %Decomposition of Cy [U_Cr, S_Cr]=svd(Cr); Kr=U_Cr*sqrt(S_Cr); %SVD of a diagonal matrix is the diagonal elements ordered, % We can skip the ordering and get Kx directly form Cx K_hat_y = sqrt(diag(C_haty)); limit = max(K_hat_y) * reg_sx + EPS_; S_hat_y_reg_diag = max(K_hat_y, limit); % Formulate regularized Kx K_hat_y_reg_inverse = 1. / S_hat_y_reg_diag; % Formulate normalization matrix G hat % Q is the identity matrix in case of the residual / diffuse part so % Q*Cx*Q' = Cx Cy_hat_diag = diag(C_hat_y); limit = max(Cy_hat_diag) * reg_ghat + EPS_; Cy_hat_diag = max(Cy_hat_diag, limit); G_hat = sqrt(diag(Cr). / Cy_hat_diag); % Formulate optimal P % Kx, G_hat are diagonal matrixes, Q is I... K_hat_y = K_hat_y.*G_hat; for k = 1:num_outputs Ky_dash(k, :) = Kr(k, :) * K_hat_y(k); end [U, ~, V] = svd(Ky_dash); P = V * U'; % Formulate M M = Kr * P; for k = 1:num_outputs M(:,k) = M(:,k) * K_hat_y_reg_inverse(k); end end
[0679] Here, considerations regarding the covariance synthesis in FIGS. 4b and 4c are provided. In some examples, two synthesis methods can be considered for each band. In the case of some bands, a complete synthesis including the residual path in FIG. 4b is applied. Typically, in the case of bands above a certain frequency where the human ear is less affected by phase, energy compensation is applied to reach the desired energy within the channel.
[0680] Therefore, even in the example of FIG. 4b, in the case of bands below a certain (fixed, recognized by the decoder) band boundary (threshold), a complete synthesis according to FIG. 4b can be performed (for example, in the case of the example in FIG. 4d). In the example of FIG. 4b, the covariance
[0681]
Number
[0682] of the uncorrelated signal 615b is derived from the uncorrelated signal 615b itself. In contrast, in the example of FIG. 4c, a decorrelator 614c in the frequency domain is used to ensure the decorrelation of the prototype signal 613c while retaining the energy of the prototype signal 613b itself.
[0683] Further considerations · In both examples of FIGS. 4b and 4c, in the first path (610b', 610c'), the mixing matrix M M is generated (in blocks 600b, 600c) by depending on the covariance C y of the original signal 212 and the covariance C x of the downmix signal 324. · In both the examples of FIGS. 4b and 4c, in the second path (610b, 610c), there is a decorrelator (614b, 614c), and the mixing matrix M R is generated, which takes into account the covariance
[0684]
Number
[0685] of the uncorrelated signals (616b, 616c). However, · In the example of FIG. 4b, the covariance
[0686]
Number
[0687] of the uncorrelated signals (616b, 616c) is calculated intuitively using the uncorrelated signals (616b, 616c) and weighted by the energy of the original channel y. · In the example of FIG. 4c, the covariance of the uncorrelated signals (616b, 616c) is calculated counterintuitively by estimating its covariance from the matrix C x and weighted by the energy of the original channel y.
[0688] The covariance matrix (
[0689]
Number
[0690] ) can be the above reconstructed target matrix (obtained, for example, from the channel level and correlation information 220 written in the side information 228 of the bit stream 248), and thus it should be noted that it can be considered to be associated with the covariance of the original signal 212. In any case, since it will be used for the composite signal 336, the covariance matrix (
[0691]
Number
[0692] ) can also be regarded as the covariance related to the composite signal. The residual covariance matrix C r ) understood as the residual covariance matrix C r and the principal covariance matrix, which can be understood as the principal covariance matrix associated with the composite signal, also apply the same.
[0693] 5. Advantages 5.1 Reduction of the use of decorrelation and optimal use of the synthesis engine Considering the proposed technique, as well as the parameters used in the processing and the way those parameters are combined with the synthesis engine 334, the strong requirement for decorrelation of the audio signal (for example, in its version 328) is reduced, and it is also explained that even in the absence of the decorrelation module 330, the influence of decorrelation (for example, spatial characteristic artifacts or degradation or signal quality degradation) is at least reduced if not eliminated.
[0694] More precisely, as previously mentioned, the decorrelation part 330 of the processing is optional. In fact, the synthesis engine 334 processes the decorrelation of the signal 328 by using the target covariance matrix C y (or a subset thereof) to ensure that the channels that make up the output signal 336 are appropriately decorrelated among themselves. The values within the covariance matrix C y represent the energy relationships between different channels of the multi-channel audio signal and are thus used as the target for synthesis.
[0695] Furthermore, the encoded (for example, transmitted) parameters 228 (for example, in their versions 314 or 318) combined with the synthesis engine 334 enable the synthesis engine 334 to use the target covariance matrix C to reproduce an output multi-channel signal 336 whose spatial characteristics and sound quality are as close as possible to the input signal 212.y Considering the use of it, high-quality output 336 can be guaranteed.
[0696] 5.2 Processing Ignorant of Downmix Considering the proposed techniques, the way the prototype signals 328 are calculated, and how they are used in the synthesis engine 334, here it is explained that the proposed decoder is ignorant of the way the downmixed signal 212 is calculated in the encoder.
[0697] This means that the proposed invention can be executed in the decoder 300 regardless of the way the downmixed signal 246 is calculated in the encoder, and the output quality of the signal 336 (or 340) does not depend on a specific downmixing method.
[0698] 5.3 Parameter Scalability Considering the proposed techniques, the way the parameters (28, 314, 318) are calculated, the way they are used in the synthesis engine 334, and the way they are estimated on the decoder side, it is explained that the number and purpose of the parameters used to describe the multi-channel audio signal are scalable.
[0699] Typically, only a subset of the parameters estimated on the encoder side (e.g., C y and / or C x subset, e.g., its elements) are encoded (e.g., transmitted), thereby reducing the bitrate used in the processing. Therefore, the amount of the encoded (e.g., transmitted) parameters (e.g., elements of C y and / or C x can be scalable considering that the non-transmitted parameters are reconstructed on the decoder side. This gives an opportunity to scale the whole processing from the viewpoints of output quality and bitrate. The more parameters are transmitted, the better the output quality, and vice versa.
[0700] Also, these parameters (e.g., C y and / or C x , or its elements) are scalable in purpose, which means that the parameters can be controlled by user input to modify the characteristics of the output multi-channel signal. Further, these parameters can be calculated for each frequency band, thus enabling scalable frequency resolution.
[0701] For example, it may be possible to determine to disable one loudspeaker in the output signals (336, 340), and thus it may be possible to directly process the parameters on the decoder side to achieve such a conversion.
[0702] 5.4 Flexibility of Output Settings Considering the proposed technique and the flexibility of the synthesis engine 334 and the parameters used (e.g., C y and / or C x , or its elements), it is here explained that the proposed invention enables a wide range of rendering with respect to output settings.
[0703] More precisely, the output settings do not have to be the same as the input settings. It is possible to process the reconstructed target covariance matrix supplied to the synthesis engine to generate the output signal 340 with a loudspeaker setting that is larger, smaller, or simply geometrically different compared to the original loudspeaker setting. This is made possible by the parameters being transmitted and the proposed system being unaware of the downmixed signal (see 5.2).
[0704] For these reasons, the proposed invention is described as being flexible from the perspective of the output loudspeaker settings.
[0705] 5. Some Examples of Prototype Matrices The table regarding 5.1 is shown below. Since the LFE was omitted, the LFE was later included in the processing (only one ICC for the related LFE / C and the ICLD for the LFE are transmitted only in the lowest parameter band, and in the synthesis on the decoder side, they are set to 1 and 0 respectively for all other bands). The naming and order of the channels follow the CICP found in ISO / IEC 23091-3, "Information technology - Coding independent code-points - Part 3: Audio". Q is always used as both the prototype matrix in the decoder and the downmix matrix in the encoder. 5.1 (CICP6). α i is used to calculate the ICLD.
[0706]
Number
[0707] 7.1 (CICP12)
[0708]
Number
[0709] α i =[0.2857 0.2857 0.5714 0.5714 0.2857 0.2857 0.2857 0.2857] 5.1 + 4 (CICP16)
[0710]
Number
[0711] α i =[0.1818 0.1818 0.3636 0.3636 0.1818 0.1818 0.1818 0.1818 0.1818 0.1818] 7.1 + 4 (CICP19)
[0712] [Number]
[0713] α i = [0.1538 0.1538 0.3077 0.3077 0.1538 0.1538 0.1538 0.1538 0.1538 0.1538 0.1538 0.1538]
[0714] 6. Method Regarding the above technology, it has been mainly described as a component or a functional device, but the present invention can also be implemented as a method. The blocks and elements described above can also be understood as method steps and / or phases.
[0715] For example, a decoding method for generating a composite signal from a downmix signal, wherein the composite signal has several composite channels, and the method comprises: Receiving a downmix signal (246, x), wherein the downmix signal (246, x) has several downmix channels and side information (228), and the side information (228) Comprises the channel levels and correlation information (220) of the original signal (212, y) And the original signal (212, y) has several original channels, the step of: Using the channel levels and correlation information (220) of the original signal (212, y), and the covariance information (C x ) related to the signal (246, x) to generate a composite signal A decoding method comprising is provided.
[0716] The decoding method comprises the following steps, namely: Calculating a prototype signal from the downmix signal (246, x), wherein the prototype signal has several composite channels, the step of: Calculating a mixing rule using the channel level and correlation information of the original signal (212, y) and the covariance information related to the downmix signal (246, x); Generating a composite signal using the prototype signal and the mixing rule; may include at least one of the above.
[0717] A decoding method for generating a composite signal (336) from a downmix signal (324, x) having a plurality of downmix channels, wherein the composite signal (336) has a plurality of composite channels and the downmix signal (324, x) is a downmixed version of an original signal (212) having a plurality of original channels, the method comprising the following phases, namely: The covariance matrix associated with the composite signal (
[0718]
Number
[0719] )(e.g., a reconstructed target version of the covariance of the original signal), and The covariance matrix (C x ) associated with the downmix signal (324), and synthesizing a first component (336M') of the composite signal according to a first mixing matrix (M M ) calculated therefrom; comprising a first phase (610c'); A second phase (610c) for synthesizing a second component (336R') of the composite signal, wherein the second component (336R') is a residual component, and the second phase (610c) comprises: A prototype signal step (612c) of upmixing the downmix signal (324) from the number of downmix channels to the number of composite channels; An uncorrelator step (614c) of decorrelating the upmixed prototype signal (613c); From the decorrelated version (615c) of the downmix signal (324), a second mixing matrix step (618c) for synthesizing a second component (336R') of the composite signal according to a second mixing matrix (M R ), wherein the second mixing matrix (M R ) is a residual mixing matrix, the second mixing matrix step (618c) including a second phase (610c) and including, the method being a residual covariance matrix (C r ) provided by a first mixing matrix step (600c), and a covariance matrix (C x ) related to the downmix signal (324), from an estimated value of the covariance matrix (
[0720]
Number
[0721] ) of the decorrelated prototype signal, calculating a second mixing matrix (M R ), The method further includes an adder step (620c) in which the first component (336M') of the composite signal is summed with the second component (336R') of the composite signal, thereby obtaining the composite signal (336). A decoding method is also provided.
[0722] Furthermore, an encoding method for generating a downmix signal (246, x) from an original signal (212, y), wherein the original signal (212, y) has a plurality of original channels and the downmix signal (246, x) has a plurality of downmix channels, the method being estimating the channel level and correlation information (220) of the original signal (212, y) in step (218), and encoding the downmix signal (246, x) into a bit stream (248) such that the downmix signal (246, x) has side information (228) including the channel level and correlation information (220) of the original signal (212, y) in step (226) An encoding method is provided that includes
[0723] These methods can be implemented in either the encoder or decoder described above.
[0724] 7. Memory Unit Furthermore, the present invention can be implemented in a non-transitory memory unit that stores instructions which, when executed by a processor, cause the processor to execute the methods as described above.
[0725] Furthermore, the present invention can be implemented in a non-transitory memory unit that stores instructions which, when executed by a processor, cause the processor to control at least one of the functions of an encoder or a decoder.
[0726] The memory unit can be, for example, part of the encoder 200 or the decoder 300.
[0727] 8. Other Aspects Some aspects are described in the context of an apparatus, but it is clear that these aspects also represent a description of the corresponding method, and that a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of the corresponding block or item, or a feature of the corresponding apparatus. Some or all of the method steps can be executed by (or by using) a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some aspects, any one or more of the most important method steps can be executed by such a device.
[0728] Aspects of the present invention may be implemented in hardware or software, depending on particular implementation requirements. The implementation may be carried out using a digital storage medium having electronically readable control signals stored thereon that cooperate with (or are capable of cooperating with) a programmable computer system so that respective methods are executed, for example, a floppy disk, a DVD, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a FLASH memory. Thus, the digital storage medium may be computer-readable.
[0729] Some aspects according to the present invention include a data carrier having electronically readable control signals capable of cooperating with a programmable computer system so that one of the methods described herein is executed.
[0730] In general, aspects of the present invention may be implemented as a computer program product with program code, which functions to execute one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, on a machine-readable carrier.
[0731] Other aspects include a computer program for executing one of the methods described herein, stored on a machine-readable carrier.
[0732] Thus, in other words, one aspect of the method of the present invention is a computer program having program code for executing one of the methods described herein when the computer program is executed on a computer.
[0733] Thus, a further aspect of the method of the present invention includes a computer program for executing one of the methods described herein, which is a data carrier (or a digital storage medium, or a computer-readable medium) on which it is recorded. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0734] Accordingly, a further aspect of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals may be configured to be transferred via a data communication connection, for example via the Internet.
[0735] A further aspect includes processing means, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.
[0736] A further aspect includes a computer having installed thereon a computer program for performing one of the methods described herein.
[0737] A further aspect according to the present invention includes an apparatus or system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system can include, for example, a file server for transferring the computer program to the receiver.
[0738] In some aspects, a programmable logic device (for example, a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some aspects, the field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the method is preferably performed by any hardware device.
[0739] The apparatus described herein may be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0740] The methods described in this specification can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0741] The above aspects are merely illustrative of the principles of the present invention. It will be understood that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Accordingly, it is intended to be limited only by the scope of the patent claims being filed, rather than by the specific details presented as descriptions and explanations of the aspects herein.
[0742] 9. Related Documents & References
[0743] [Item 1] An audio synthesizer (300) for generating a composite signal (336, 340, y R ) from a downmix signal (246, x), wherein the composite signal (336, 340, y R ) has a plurality of composite channels, and the audio synthesizer (300) An input interface (312) configured to receive the downmix signal (246, x), wherein the downmix signal (246, x) has a plurality of downmix channels and side information (228), the side information (228) includes channel levels and correlation information (314, ξ, χ) of the original signal (212, y), and the original signal (212, y) has a plurality of original channels, the input interface (312); A synthesis processor (404), The channel levels and correlation information (220, 314, ξ, χ) of the original signal (212, y), and Covariance information (C x ) associated with the downmix signal (324, 246, x) R ) to generate the composite signal (336, 340, y An audio synthesizer (300) comprising [Item 2] A prototype signal calculator (326) configured to calculate a prototype signal (328) from the downmix signal (324, 246, x), wherein the prototype signal (328) has a plurality of synthesis channels, the prototype signal calculator (326); A mixing rule calculator (402), The channel levels and correlation information (314, ξ, χ) of the original signal (212, y), and The covariance information (C x ) A mixing rule calculator (402) configured to calculate at least one mixing rule (403) using Comprising, the synthesis processor (404) using the prototype signal (328) and the at least one mixing rule (403) to generate the synthesis signal (336, 340, y R ), the audio synthesizer (300) according to item 1. [Item 3] The audio synthesizer according to item 1 or 2, configured to reconstruct (386) the target covariance information (C y ) of the original signal. [Item 4] The audio synthesizer according to item 3, configured to reconstruct the target covariance information (C R ) adapted to the number of channels of the synthesis signal (336, 340, y y ). [Item 5] By assigning a group of original channels to a single synthesis channel, or vice versa, reconstructing the covariance information (C R ) adapted to the number of channels of the synthesis signal (336, 340, y y ), so that the reconstructed target covariance information (
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Explanation of Signs
[0744] 100 Audio system 200 Encoder 212 Original signal, input signal 214 Filter bank 216 Frequency domain version 218 Parameter estimator 220 Channel level and correlation information 220s Scaler 226 Bit stream writer 228 Side information 244 Downmix section 246 Downmix signal 247 Core coder 248 Bitstream 249 Multiplexer 254s Switch 258 Transient detector 263 Filter 270 Storage 273 Subtractor 300 Decoder 312 Input interface 314 Channel level and correlation information 316 Parameter reconstruction module 320 Filter bank 324 Downmix signal 326 Prototype signal calculator 328 Prototype signal 332 Prototype signal 334 Synthesis engine 336 Synthesis signal 338 Filter bank 340 Synthesis signal 347 Core decoder 384 Covariance estimator 402 Mixing rule calculator 403 Mixing rule 404 Synthesis processor 502 Covariance estimator 504 Covariance estimator 600a Synthesis processor 600b Synthesis processor 614c Decorrelator 616b COV estimator 616c Core estimator 630 Selector 900 ICC matrix
Claims
1. An audio encoder (200) for generating a downmix signal (246, x) from an original signal (212, y), wherein the original signal (212, y) has a plurality of original channels, the downmix signal (246, x) has a plurality of downmix channels, and the audio encoder (200) A parameter estimator (218) configured to estimate channel levels and correlation information (220) of the original signal (212, y), A bitstream writer (226) for encoding the downmix signal (246, x) into a bitstream (248) such that the downmix signal (246, x) has side information (228) including the channel levels and correlation information (220) of the original signal (212, y), comprising the channel levels and correlation information (220) of the original signal (212, y) including at least one inter-channel level difference (ICLD), the channel levels and correlation information (220) of the original signal (212, y) encoded in the side information (228) including at least correlation information (220, 908) describing an energy relationship between at least one pair of different original channels but less than all of the original channels, Audio encoder (200).
2. The audio encoder according to claim 1, configured to provide the channel levels and correlation information (220) of the original signal (212, y) as normalized values.
3. The audio encoder according to claim 1, wherein the channel levels and correlation information (220) of the original signal (212, y) include or represent channel level information related to the entirety of the original channels.
4. The channel level and correlation information (220) of the original signal (212, y) include at least one coherence value (ξ i,j ) that describes the coherence between two channels of the original channel. **Claim 5** The audio encoder according to claim 4, wherein the coherence value is normalized. **Claim 6** The coherence value is **[Equation 1]** wherein, **[Equation 2]** is the covariance between channel i and channel j, **[Equation 3]** and **[Equation 4]** are the levels related to the channel i and the channel j respectively. The audio encoder according to claim 4. **Claim 7** The audio encoder according to claim 1, wherein the at least one ICLD is provided as a logarithmic value. **Claim 8** The audio encoder according to claim 7, wherein the at least one ICLD is normalized. **Claim 9** The ICLD is **[Equation 5]** wherein, -χ i is the ICLD of channel i, -P i is the power of the current channel i, -P dmx,i is a linear combination of the covariance information values of the downmix signal. The audio encoder according to claim 8. **Claim 10** configured to select (250) which portion of the channel level and correlation information (220) of the original signal (212, y) to encode within the side information (228) based on the metrics (252) on the channel so as to include channel level and correlation information (220) related to more susceptible metrics in the side information (228), the audio encoder according to claim 1.
11. wherein the channel level and correlation information (220) of the original signal (212, y) is in the form of entries of a matrix (C y ), the audio encoder according to claim 1.
12. wherein the matrix is a symmetric matrix or a Hermitian matrix, and the entries of the channel level and correlation information (220) are provided for all or less than all of the entries on the diagonal of the matrix (C y ) and / or for less than half of the off-diagonal elements of the matrix (C y ), the audio encoder according to claim 11.
13. wherein the bitstream writer (226) is configured to encode the identification of at least one channel, the audio encoder according to claim 1.
14. wherein the original signal (212, y) or a processed version thereof (216) is divided into a plurality of subsequent frames of equal time duration, the audio encoder according to claim 1.
15. configured to encode the channel level and correlation information (220) of the original signal (212, y) specific to each frame within the side information (228), the audio encoder according to claim 14.
16. configured to encode the same channel level and correlation information (220) of the original signal (212, y) collectively associated with a plurality of consecutive frames within the side information (228), the audio encoder according to claim 15.
17. An audio encoder configured such that the original signal (212, y) is converted (263) into a frequency domain signal (264, 266), and the audio encoder encodes the channel level and correlation information (220) of the original signal (212, y) into the side information (228) in a band-by-band manner. The audio encoder according to claim 1, wherein the audio encoder is configured to aggregate some bands of the original signal (212, y) into a smaller number of bands (266) so as to encode the channel level and correlation information (220) of the original signal (212, y) into the side information (228) in an aggregated band unit manner.
18. The audio encoder according to claim 17, further configured to encode at least one channel level and correlation information (220) of one band into the bitstream (248) as an increment with respect to the previously encoded channel level and correlation information.
19. The audio encoder according to claim 1, configured to encode the correlation information (220) regarding the channel level that is an incomplete version compared to the channel level and correlation information (220) estimated by the estimator (218) into the side information (228) of the bitstream (248).
20. The audio encoder according to claim 19, configured to adaptively select the selected information to be encoded into the side information (228) of the bitstream (248) from among the entire channel level and correlation information (220) estimated by the estimator (218), so that the remaining unselected information channel level and / or correlation information (220) estimated by the estimator (218) are not encoded.
21. The channel levels and correlation information (220) are indexed according to a predefined order, and the encoder is configured to signal an index associated with the predefined order in the side information (228) of the bitstream (228), the index indicating which of the channel levels and correlation information (220) are encoded. The audio encoder according to claim 19.
22. The audio encoder according to claim 21, wherein the index is provided via a bitmap.
23. The audio encoder according to claim 22, wherein the index is defined according to a combination number system that associates a one-dimensional index with entries of a matrix.
24. Adaptive provision of the channel levels and correlation information (220), wherein the index associated with the predefined order is encoded within the side information of the bitstream, and Fixed provision of the channel levels and correlation information (220), wherein the channel levels and correlation information (220) to be encoded are predetermined and ordered according to a predefined fixed order without the provision of an index The audio encoder according to claim 23, configured to perform a selection between.
25. The audio encoder according to claim 24, configured to signal in the side information (228) of the bitstream (248) whether the channel levels and correlation information (220) are provided according to adaptive provision or according to fixed provision.
26. The audio encoder according to claim 1, further configured to encode (226) current channel levels and correlation information (220t) as an increment (220k) relative to previous channel levels and correlation information (220(t-1)) in the bitstream (248).
27. A method for generating a downmix signal (246, x) from an original signal (212, y), wherein the original signal (212, y) has a plurality of original channels, the downmix signal (246, x) has a plurality of downmix channels, and the method comprises: estimating (218) channel levels and correlation information (220) of the original signal (212, y), wherein the channel levels and correlation information (220) of the original signal (212, y) include at least one inter-channel level difference (ICLD), and the channel levels and correlation information (220) of the original signal (212, y) encoded in side information (228) further include correlation information (220, 908) describing an energy relationship between at least one pair of different original channels but fewer than all of the original channels; encoding (226) the downmix signal (246, x) into a bitstream (248) such that the downmix signal (246, x) has the side information (228) including the channel levels and the correlation information (220) of the original signal (212, y); and a method comprising: Claim 28 A non-transitory storage unit that stores instructions which, when executed by a processor, cause the processor to execute the method according to claim 27.
Citation Information
Patent Citations
IEC23091-3、
Method and apparatus for decoder for multi-channel surround sound
JP2009531735A
multichannel audio decoder, multichannel audio encoder, method of using rendered audio signal, computer program and encoded audio representation
JP2016528811A