Audio encoder and decoder

The modulo difference and entropy encoding method addresses inefficiencies in audio encoding by reducing memory and complexity, enhancing encoding efficiency and quality in object-based audio systems.

JP2025100674AActive Publication Date: 2025-07-03DOLBY INTERNATIONAL AB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025064679
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-05-24
Filing Date
2025-04-10
Publication Date
2025-07-03
Estimated Expiration
2034-05-23

AI Technical Summary

Technical Problem

Existing audio encoding systems, particularly object-based audio systems, face challenges in efficiently encoding and decoding audio signals while maintaining quality, as methods like MPEG SAOC rely on complex mathematical processes and assumptions about audio object attributes, leading to inefficiencies in bitrate and quality.

Method used

The proposed solution involves encoding audio parameters using a modulo difference approach and entropy encoding, where each parameter is represented by an index value and associated with a symbol, reducing the number of possible symbols and probability table size, and using a shared probability table for both first and second elements to minimize memory requirements and enhance encoding efficiency.

Benefits of technology

This method reduces memory requirements and encoding complexity, leading to more efficient bitrate usage and improved encoding quality by utilizing a shared probability table and modulo differential encoding, which maintains coding efficiency and reduces the need for expensive memory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100674000001_ABST
    Figure 2025100674000001_ABST
Patent Text Reader

Abstract

To provide methods, devices, and computer program products for encoding and decoding of a vector of parameters in an audio coding system, and to provide a method and an apparatus for reconstructing an audio object in an audio decoding system.SOLUTION: According to the disclosure, a modulo difference approach for coding and encoding a vector of a non-periodic quantity may improve coding efficiency and provide encoders and decoders with reduced memory requirements. Moreover, an efficient method for encoding and decoding a sparse matrix is provided.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of the filing date of U.S. Provisional Patent Application No. 61 / 827,264, filed on May 24, 2013. The content of the application is incorporated herein by reference.

[0002] Technical Field The disclosure herein generally relates to audio encoding. More particularly, it relates to the encoding and decoding of vectors of parameters in an audio encoding system. The disclosure further relates to methods and apparatus for reconstructing audio objects in an audio decoding system.

Background Art

[0003] In a conventional audio system, a channel - based approach is used. Each channel may represent, for example, the content of one speaker or one speaker array. Possible encoding schemes for such a system include discrete multi - channel encoding or parametric encoding such as MPEG Surround.

[0004] More recently, a new approach has been developed. This approach is object - based. In a system using an object - based approach, a three - dimensional audio scene is represented by audio objects with associated position metadata. These audio objects move around within the three - dimensional audio scene during the playback of the audio signal. The system may further include a so - called bed channel. The bed channel may be described as a static audio object that directly maps to the speaker positions of a conventional audio system as described above.

[0005] A problem that can occur in an object-based audio system is how to efficiently encode and decode audio signals while maintaining the quality of the encoded signals. One possible encoding method involves generating, on the encoder side, a downmix signal that includes some channels from the audio object and the bed channel, and on the decoder side, side information that enables regeneration of the audio object and the bed channel.

[0006] MPEG Spatial Audio Object Coding (MPEG SAOC) describes a system for parametric coding of audio objects. This system transmits side information, an upmix matrix reference, which describes the attributes of the object, by parameters such as the level difference and cross-correlation of the objects. These parameters are then used on the decoder side to control the regeneration of the audio object. This process is mathematically complex and often has to rely on assumptions about the attributes of the audio object that are not explicitly described by the parameters. The method presented in MPEG SAOC can reduce the required bitrate for an object-based audio system, but further improvements may be needed to further increase efficiency and quality as described above.

Brief Description of the Drawings

[0007] Exemplary embodiments will now be described with reference to the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

[0008] In view of the above, it is an object to provide an encoder, a decoder and a related method that provide increased efficiency and quality of an encoded audio signal.

[0009] 〈I. Overview - Encoder〉 According to a first aspect, an exemplary embodiment proposes an encoding method, an encoder and a computer program product for encoding. The proposed method, encoder and computer program product may generally have the same features and advantages.

[0010] According to an exemplary embodiment, a method for encoding a vector of parameters in an audio encoding system is provided. Each parameter corresponds to a non-periodic quantity. The vector has a first element and at least one second element. The method includes: representing each parameter in the vector by an index value that can take N values; and associating each of the at least one second element with a symbol, the symbol being: calculating a difference between the index value of the second element and the index value of its preceding element in the vector; and calculating by applying modulo N to the difference. The method further includes encoding each of the at least one second element by entropy encoding the symbol associated with the at least one second element based on a probability table including symbol probabilities.

[0011] The advantage of this method is that the number of possible symbols is reduced by about half compared to a normal differential encoding strategy where the modulo N is not applied to the difference. As a result, the size of the probability table is reduced by about half. As a result, less memory is required to store the probability table, and since the probability table is often stored in expensive memory in the encoder, the encoder can thus be made less expensive. Further, the speed of searching for symbols in the probability table can be increased. A further advantage is that the encoding efficiency can be increased because all symbols in the probability table are possible candidates to be associated with a particular second element. This can be compared to a normal differential encoding strategy where only about half of the symbols in the probability table are candidates for being associated with a particular second element.

[0012] According to various embodiments, the method further includes associating the first element in the vector with a symbol. The symbol is calculated by shifting an index value representing the first element in the vector by an offset value and applying modulo N to the shifted index value. The method further includes encoding the first element by entropy encoding the symbol associated with the first element using the same probability table used to encode the at least one second element.

[0013] This embodiment takes advantage of the fact that the probability distribution of the index value of the first element and the probability distribution of the symbols of the at least one second element are similar, although shifted relative to each other by an offset value. As a result, instead of a dedicated probability table, the same probability table can be used for the first element in the vector. As a result, as described above, it can lead to reduced memory requirements and a less expensive encoder.

[0014] According to an embodiment, the offset value is equal to the difference between the most likely index value for the first element and the most likely symbol for the at least one second element in the probability table. This means that the peaks of their probability distributions are aligned. As a result, for the first element, substantially the same coding efficiency is maintained compared to the case where a dedicated probability table is used for the first element.

[0015] According to embodiments, the first element and the at least one second element of the vector of parameters correspond to different frequency bands used in the audio encoding system in a specific time frame. That is, data corresponding to a plurality of frequency bands can be encoded in the same operation. For example, the vector of parameters may correspond to upmixing or reconstruction coefficients that vary across a plurality of frequency bands.

[0016] According to an embodiment, the first element and the at least one second element of the vector of parameters correspond to different time frames used in the audio encoding system in a specific frequency band. That is, data corresponding to a plurality of time frames can be encoded in the same operation. For example, the vector of parameters may correspond to upmixing or reconstruction coefficients that vary across a plurality of time frames.

[0017] According to embodiments, the probability table is translated into a Huffman codebook. Here, the symbol associated with an element in the vector is used as a codebook index, and the encoding step includes encoding each of the at least one second element by representing the second element with a codeword in the codebook indexed by the codebook index associated with the second element. By using the symbol as a codebook index, the search speed of the codeword representing the element can be improved.

[0018] According to various embodiments, the encoding step includes encoding the first element in the vector using the same Huffman codebook used to encode the at least one second element by representing the first element with a codeword in the Huffman codebook indexed by a codebook index associated with the first element. As a result, only one Huffman codebook needs to be stored in the encoder's memory, which can lead to a less expensive encoder as described above.

[0019] According to a further embodiment, the vector of parameters corresponds to elements in an upmix matrix determined by the audio encoding system. This can reduce the required bitrate in an audio encoding / decoding system since the upmix matrix can be efficiently encoded.

[0020] According to an exemplary embodiment, a computer-readable medium is provided having computer code instructions adapted to execute any method of a first aspect when executed on an apparatus having processing capabilities.

[0021] According to an exemplary embodiment, an encoder is provided that encodes a vector of parameters in an audio encoding system. Each parameter corresponds to an aperiodic quantity. The vector has a first element and at least one second element. The encoder comprises: a receiving component adapted to receive the vector; an indexing component adapted to represent each parameter in the vector by an index value that can take N values; and an associating component adapted to associate each of the at least one second element with a symbol. The symbol is calculated by: calculating a difference between the index value of the second element and the index value of its preceding element in the vector; and applying modulo N to the difference. The encoder further comprises an encoding component that encodes each of the at least one second element by entropy encoding the symbol associated with the at least one second element based on a probability table that includes the probability of the symbol.

[0022] 〈II. Overview - Decoder〉 According to a second aspect, the exemplary embodiment proposes a decoding method, a decoder and a computer program product for decoding. The proposed method, decoder and computer program product may generally have the same features and advantages.

[0023] The features and advantages regarding the setup presented in the overview of the encoder above may generally also be valid for the corresponding features and setup of the decoder.

[0024] According to an exemplary embodiment, a method is provided for decoding a vector of entropy-encoded symbols in an audio decoding system into a vector of parameters related to an aperiodic quantity. The vector of entropy-encoded symbols has a first entropy-encoded symbol and at least one second entropy-encoded symbol, and the vector of parameters has a first element and at least a second element. The method includes: representing each entropy-encoded symbol in the vector of entropy-encoded symbols by a symbol that can take N integer values by using a probability table; associating the first entropy-encoded symbol with an index value; and associating each of the at least one second entropy-encoded symbols with an index value, wherein the index value of the at least one second entropy-encoded symbol is calculated by adding the index value associated with the entropy-encoded symbol preceding the second entropy-encoded symbol in the vector of entropy-encoded symbols and the symbol representing the second entropy-encoded symbol, and applying modulo N to the sum. The method further includes representing the at least one second element of the vector of parameters by a parameter value corresponding to the index value associated with the at least one second entropy-encoded symbol.

[0025] According to an exemplary embodiment, by a symbol, the step of representing each entropy-encoded symbol in the vector of entropy-encoded symbols is performed using the same probability table for all entropy-encoded symbols in the vector of entropy-encoded symbols. The index value associated with the first entropy-encoded symbol is calculated by shifting the symbol representing the first entropy-encoded symbol in the vector of entropy-encoded symbols by an offset value and applying modulo N to the shifted symbol. The method further includes the step of representing the first element of the vector of parameters by a parameter value corresponding to the index value associated with the first entropy-encoded symbol.

[0026] According to an embodiment, the probability table is translated into a Huffman codebook, and each entropy-encoded symbol corresponds to a codeword in the Huffman codebook.

[0027] According to a further embodiment, each codeword in the Huffman codebook is associated with a codebook index, and by a symbol, the step of representing each entropy-encoded symbol in the vector of entropy-encoded symbols includes representing the entropy-encoded symbol by the codebook index associated with the codeword corresponding to the entropy-encoded symbol.

[0028] According to embodiments, each entropy-encoded symbol in the vector of entropy-encoded symbols corresponds to a different frequency band used in the audio decoding system in a particular time frame.

[0029] According to one embodiment, each entropy-encoded symbol in the vector of entropy-encoded symbols corresponds to a different time frame used in the audio decoding system in a particular frequency band.

[0030] According to embodiments, the vector of parameters corresponds to an element in an upmix matrix used by the audio decoding system.

[0031] According to an exemplary embodiment, a computer-readable medium having computer code instructions adapted to execute any method of a second aspect when executed on a device having a processing function is provided.

[0032] According to an exemplary embodiment, a decoder is provided that decodes a vector of entropy - encoded symbols in an audio decoding system into a vector of parameters related to an aperiodic quantity. The vector of entropy - encoded symbols has a first entropy - encoded symbol and at least one second entropy - encoded symbol, and the vector of parameters has a first element and at least a second element. The decoder includes: a receiving component configured to receive a vector of entropy - encoded symbols; an indexing component configured to represent each entropy - encoded symbol in the vector of entropy - encoded symbols by a symbol that can take N integer values by using a probability table; and an associating component configured to associate the first entropy - encoded symbol with an index value, wherein the associating component is further configured to associate each of the at least one second entropy - encoded symbols with an index value, and the index value of the at least one second entropy - encoded symbol is calculated by: calculating a sum of the index value of the entropy - encoded symbol preceding the second entropy - encoded symbol in the vector of entropy - encoded symbols and the symbol representing the second entropy - encoded symbol; and applying modulo N to the sum. The decoder further has a decoding component configured to represent the at least one second element of the vector of parameters by a parameter value corresponding to the index value associated with the at least one second entropy - encoded symbol.

[0033] 〈III. Overview - Sparse Matrix Encoder〉 According to a third aspect, an exemplary embodiment proposes an encoding method, an encoder, and a computer program product for encoding. The proposed method, encoder, and computer program product may generally have the same features and advantages.

[0034] According to an exemplary embodiment, a method for encoding an upmix matrix in an audio encoding system is provided. Each row of the upmix matrix includes M elements that allow reconstruction of a time / frequency tile of an audio object from a downmix signal including M channels. The method includes, for each row of the upmix matrix: selecting a subset of elements from the M elements of that row of the upmix matrix; representing each element in the selected subset of elements by a value and a position in the upmix matrix; and encoding the value and the position in the upmix matrix of each element in the selected subset of elements.

[0035] In the usage in this document, the term "downmix signal including M channels" means a signal including M signals or channels, where each channel is a combination of a plurality of audio objects including the audio object to be reconstructed. The number of channels is typically greater than 1, and in many cases the number of channels is 5 or more.

[0036] In the usage in this document, the term "upmix matrix" refers to a matrix having N rows and M columns that allows N audio objects to be reconstructed from a downmix signal including M channels. The elements of each row of the upmix matrix correspond to one audio object and give the coefficients to be multiplied by the M channels of the downmix to reconstruct the audio object.

[0037] In the usage in this document, the position in the upmix matrix means the row and column indices indicating the row and column of the matrix element. The term "position" may also mean the column index in a given row of the upmix matrix.

[0038] In some cases, sending all elements of the upmix matrix for each time / frequency tile requires an undesirably high bitrate in an audio encoding / decoding system. An advantage of the present method is that only a subset of the upmix matrix elements need to be encoded and transmitted to the decoder. Since less data is transmitted, the required bitrate of the audio encoding / decoding system may be reduced and the data may be encoded more efficiently.

[0039] An audio encoding / decoding system typically divides the time-frequency space into time / frequency tiles by applying, for example, a filter bank suitable for an input audio signal. A time / frequency tile generally means a portion of the time-frequency space corresponding to a certain time interval and a frequency subband. The time interval typically corresponds to the duration of a time frame used in the audio encoding / decoding system. The frequency subband typically corresponds to one or several adjacent frequency subbands defined by the filter bank used in the encoding / decoding system. When the frequency subband corresponds to several adjacent frequency subbands defined by the filter bank, this allows for non-uniform frequency subbands in the decoding process of the audio signal, for example, wider frequency subbands for higher frequencies of the audio signal. In the case of a broadband audio encoding / decoding system that acts on the entire frequency range, the frequency subbands of the time / frequency tiles may correspond to the entire frequency range. The above method discloses various encoding steps for encoding an upmix matrix in an audio encoding system to allow for the reconstruction of an audio object between one such time / frequency tile. However, it is understood that the method may be repeated for each time / frequency tile of the audio encoding system. It is also understood that several time / frequency tiles may be encoded simultaneously. Typically, adjacent time / frequency tiles may overlap slightly in time and / or frequency. For example, the overlap in time may be equivalent to a linear interpolation of the elements of the reconstruction matrix over a certain time interval, i.e., from one time interval to the next. However, the present disclosure also targets other parts of the encoding / decoding system, and any overlap in time and / or frequency between adjacent time / frequency tiles is left to the implementation of those skilled in the art.

[0040] According to various embodiments, for each row in the upmix matrix, the position of the selected subset of elements in the upmix matrix changes across multiple frequency bands and / or across multiple time frames. Thus, the selection of those elements may depend on a particular time / frequency tile, and thus different elements may be selected for different time / frequency tiles. This provides a more flexible encoding method, which enhances the quality of the encoded signal.

[0041] According to various embodiments, the selected subset of elements includes the same number of elements for each row of the upmix matrix. In a further embodiment, the number of elements selected may be exactly 1. This reduces the complexity of the encoder as the algorithm only needs to select the same number of elements (singular or plural) for each row, i.e., the most important element(s) when performing upmix on the decoder side.

[0042] According to various embodiments, for each row in the upmix matrix and for multiple frequency bands or multiple time frames, the values of the elements of the selected subset of elements form one or more vectors of parameters, each parameter in the parameter vector corresponding to one of the multiple frequency bands or the multiple time frames, and the one or more vectors of parameters are encoded using a method based on a first aspect. In other words, the values of the selected elements can be efficiently encoded. The features and advantages regarding the setup presented in the overview of the first aspect above may generally also be valid for this embodiment.

[0043] According to various embodiments, for each row in the upmix matrix and for a plurality of frequency bands or a plurality of time frames, the positions of the elements of a selected subset of the elements form one or more vectors of parameters, each parameter in the parameter vector corresponding to one of the plurality of frequency bands or the plurality of time frames, and the one or more vectors of parameters being encoded using a method based on a first aspect. In other words, the positions of the selected elements can be efficiently encoded. The advantages regarding the features and setup presented in the overview of the first aspect above can generally also be valid for this embodiment.

[0044] According to an exemplary embodiment, there is provided a computer-readable medium having computer code instructions adapted to execute any method of a third aspect when executed on a device having processing capabilities.

[0045] According to an exemplary embodiment, there is provided an encoder for encoding an upmix matrix in an audio encoding system. Each row of the upmix matrix includes M elements that allow reconstruction of a time / frequency tile of an audio object from a downmix signal including M channels. The encoder has: a receiving component adapted to receive each row in the upmix matrix; a selection component adapted to select a subset of elements from the M elements of the row in the upmix matrix; and an encoding component adapted to represent each element in the selected subset of elements by a value and a position in the upmix matrix, the encoding component being further adapted to encode the value and the position in the upmix matrix of each element in the selected subset of elements.

[0046] 〈IV. Overview - Sparse Matrix Decoder〉 According to a fourth aspect, the exemplary embodiments propose a decoding method, a decoder, and a computer program product for decoding. The proposed method, decoder, and computer program product may generally have the same features and advantages.

[0047] The features and advantages regarding the setup presented in the overview of the sparse matrix encoder above may generally also be valid for the corresponding features and setup of the decoder.

[0048] According to an exemplary embodiment, a method for reconstructing time / frequency tiles of an audio object in an audio decoding system is provided. The method includes: receiving a downmix signal including M channels; receiving at least one encoded element representing a subset of M elements of a row in an upmix matrix, each encoded element including a value and a position in that row of the upmix matrix, the position indicating one of the M channels of the downmix signal corresponding to that encoded element; and reconstructing the time / frequency tile of the audio object from the downmix signal by forming a linear combination of the downmix channels corresponding to the at least one encoded element. In the linear combination, each downmix channel is multiplied by the value of its corresponding encoded element.

[0049] Thus, according to this method, the time / frequency tiles of the audio object are reconstructed by forming a linear combination of a subset of the downmix channels. The subset of downmix channels corresponds to the channels where the upmix coefficients encoded for them were received. Thus, the method allows for reconstructing the audio object despite the fact that only a subset of the upmix matrix, for example a sparse subset, is received. By forming a linear combination of only the downmix channels corresponding to the at least one encoded element, the complexity of the decoding process can be reduced. An alternative would be to form a linear combination of all the downmix signals and then multiply some of them (those not corresponding to the at least one encoded element) by the value 0.

[0050] According to embodiments, the position of the at least one encoded element changes across a plurality of frequency bands and / or across a plurality of time frames. Thus, in other words, different elements of the upmix matrix may be encoded for different time / frequency tiles.

[0051] According to embodiments, the number of elements of the at least one encoded element is equal to 1. That is, the audio object is reconstructed from one downmix channel in each time / frequency tile. However, the one downmix channel used for reconstructing the audio object may vary between different time / frequency tiles.

[0052] According to various embodiments, for a plurality of frequency bands or a plurality of time frames, the value of the at least one encoded element forms one or more vectors, each value is represented by an entropy-encoded symbol, each symbol in each vector of the entropy-encoded symbols corresponds to one of the plurality of frequency bands or one of the plurality of time frames, and the one or more vectors of the entropy-encoded symbols are decoded using a method based on a second aspect. In this way, the values of the elements of the upmix matrix can be efficiently encoded.

[0053] According to various embodiments, for a plurality of frequency bands or a plurality of time frames, the position of the at least one encoded element forms one or more vectors, each position is represented by an entropy-encoded symbol, each symbol in each vector of the entropy-encoded symbols corresponds to one of the plurality of frequency bands or one of the plurality of time frames, and the one or more vectors of the entropy-encoded symbols are decoded using a method based on a second aspect. In this way, the positions of the elements of the upmix matrix can be efficiently encoded.

[0054] According to an exemplary embodiment, there is provided a computer-readable medium having computer code instructions adapted to execute any method of a third aspect when executed on an apparatus having a processing function.

[0055] According to an exemplary embodiment, a decoder is provided that reconstructs time / frequency tiles of an audio object. The decoder is a receiving component configured to receive a downmix signal including M channels and at least one encoded element representing a subset of M elements of a row in an upmix matrix, each encoded element including a value and a position in that row of the upmix matrix, the position indicating one of the M channels of the downmix signal to which the encoded element corresponds; and a reconstruction component configured to reconstruct the time / frequency tile of the audio object from the downmix signal by forming a linear combination of the downmix channels corresponding to the at least one encoded element. In the linear combination, each downmix channel is multiplied by the value of its corresponding encoded element.

Example

[0056] 〈V. Exemplary Embodiments〉 FIG. 1 shows a generalized block diagram of an audio encoding system 100 for encoding an audio object 104. The audio encoding system has a downmix component 106 that generates a downmix signal 110 from a plurality of audio objects 104. The downmix signal 110 may be, for example, a 5.1 or 7.1 surround signal that is backward compatible with an established sound decoding system such as Dolby Digital Plus or an MPEG standard such as AAC, USAC, or MP3. In a further embodiment, the downmix signal is not backward compatible.

[0057] In order to be able to reconstruct the audio object 104 from the downmix signal 110, the upmix parameters are determined in the upmix parameter analysis component 112 from the downmix signal 110 and the audio object 104. For example, the upmix parameters may correspond to the elements of an upmix matrix that allows the reconstruction of the audio object 104 from the downmix signal 110. The upmix parameter analysis component 112 processes the downmix signal 110 and the audio object 104 with respect to individual time / frequency tiles. Thus, the upmix parameters are determined for each time / frequency tile. For example, an upmix matrix may be determined for each time / frequency tile. For example, the upmix parameter analysis component 112 may operate in a frequency domain such as a quadrature mirror filter (QMF) region that allows frequency-selective processing. For this reason, by applying the downmix signal 110 and the audio object 104 to the filter bank 108, the downmix signal 110 and the audio object 104 may be converted to the frequency domain. This may be done, for example, by applying a QMF transform or any other suitable transform.

[0058] The upmix parameter 114 may be organized in a vector format. The vector may represent upmix parameters for reconstructing a particular audio object from the audio objects 104 in various frequency bands at a particular time frame. For example, the vector may correspond to a certain matrix element in the upmix matrix. Here, the vector includes the values of the certain matrix element for a series of frequency bands. In a further embodiment, the vector may represent upmix parameters for reconstructing a particular audio object from the audio objects 104 in various time frames at a particular frequency band. For example, the vector may correspond to a certain matrix element of the upmix matrix, and the vector includes the values of the certain matrix element for a series of time frames, but in the same frequency band.

[0059] Each parameter in the vector corresponds to an aperiodic quantity, for example, a quantity that takes values between -9.6 and 9.4. An aperiodic quantity generally means a quantity that has no periodicity in the values it can take. This is in contrast to a periodic quantity such as an angle where there is a clear periodic correspondence between the values the quantity can take. For example, for an angle, there is a periodicity of 2π, for example, angle 0 corresponds to angle 2π.

[0060] Next, the upmix parameter 114 is received by the upmix matrix encoder 102 in vector format. The upmix matrix encoder will be described in detail here in connection with FIG. 2. The vector is received by the receiving component 202 and has a first element and at least one second element. The number of elements depends, for example, on the number of frequency bands in the audio signal. The number of elements may also depend on the number of time frames of the audio signal encoded in one encoding operation.

[0061] Next, the vector is indexed by the indexing component 204. The indexing component is adapted to represent each parameter in the vector by an index value that can take a predefined number of values. This representation can be done in two steps. First, the parameter is quantized, and then the quantized value is indexed by the index value. As an example, if each parameter in the vector can take values between -9.6 and 9.4, this can be done by using a quantization step of 0.2. The quantized value may then be indexed by index values 0 to 95, i.e., 96 different values. In the following example, the index values are in the range 0 to 95, but this is of course just an example, and other ranges of index values, such as 0 to 191 or 0 to 63, are equally possible. A smaller quantization step may result in a less distorted decoded audio signal at the decoder side, but may also result in a higher required bit rate for the transmission of data between the audio encoding system 100 and the decoder.

[0062] The indexed values are then sent to the association component 206. The association component 206 uses a modulo difference encoding strategy to associate each of the at least one second element with a symbol. The association component 206 is adapted to calculate the difference between the index value of the second element and the index value of the immediately preceding element in the vector. By simply using a normal difference encoding strategy, the difference could be anywhere within the range of -95 to 95. That is, there are 191 possible values. This means that when the difference is encoded using entropy encoding, a probability table containing 191 probabilities is required. That is, one probability for each of the 191 possible values of the difference. Further, for each difference, approximately half of the 191 probabilities are impossible, so the encoding efficiency will decrease. For example, if the second element to be difference encoded has an index value of 90, the possible differences are within the range of -5 to +90. Typically, having an entropy encoding strategy where some of the probabilities are impossible for each value to be encoded reduces the encoding efficiency. The difference encoding strategy in the present disclosure overcomes this problem by applying a modulo 96 operation to the difference and at the same time reducing the number of codes required to 96. Thus, the association algorithm can be expressed as follows.

[0063] Δ idx (b)=(idx(b)-idx(b - 1)) mod N Q (Equation 1) where b is an element in the vector being difference encoded, and N Q is the number of possible index values, and Δ idx (b) is the symbol associated with element b.

[0064] According to some embodiments, the probability table is converted into a Huffman codebook. In this case, the symbol associated with an element in the vector is used as the codebook index. Then, the encoding component 208 can encode each of the at least one second element by representing the second element with the codeword in the Huffman codebook indexed by the codebook index associated with the second element.

[0065] Any other suitable entropy encoding strategy may be implemented by the encoding component 208. For example, such an encoding strategy may be a range coding strategy or an arithmetic coding strategy.

[0066] The following shows that the entropy of the modulo approach is always less than or equal to the entropy of the normal difference approach. The entropy E of the normal difference approach p is

Number

[0067] The entropy E of the modulo approach q is

Number

[0068] Therefore, it becomes as follows.

Number

[0069] [Number] Comparing the sums term by term [Number] so, E p ≥ E q is obtained

[0070] As shown above, the entropy for the modulo approach is always less than or equal to the entropy for the normal difference approach. The case where the entropies are equal is a rare case where the data to be encoded is pathological data, i.e., data with bad behavior, and in most cases, for example, it does not apply to the upmix matrix.

[0071] Since the entropy for the modulo approach is always less than or equal to the entropy for the normal difference approach, the entropy coding of symbols calculated by the modulo approach results in a lower or at least the same bit rate compared to the entropy coding of symbols calculated by the normal difference approach. In other words, the entropy coding of symbols calculated by the modulo approach is usually more efficient than the entropy coding of symbols calculated by the normal difference approach.

[0072] A further advantage is that, as described above, the number of probabilities required in the probability table in the modulo approach is approximately half of the number of probabilities required in the normal non - modulo approach.

[0073] In the above, a modulo approach for encoding the at least one second element in the vector of parameters has been described. The first element may be encoded using an index value representing the first element. Since the probability distributions of the index value of the first element and the modulo difference value of the at least one second element may be very different (see FIG. 3 for the probability distribution of the indexed first element and FIG. 4 for the probability distribution of the modulo difference value, i.e., the symbol for the at least one second element), a dedicated probability table for the first element may be required. This requires that both the audio encoding system 100 and the corresponding decoder have such a dedicated probability table in memory.

[0074] However, the inventors have observed that in some cases, the shapes of the probability distributions may be very similar while shifted relative to each other. This observation can be used to approximate the probability distribution of the indexed first element by a shifted version of the probability distribution of the symbol for the at least one second element. Such a shift may be implemented by the association component 206 associating the first element in the vector with a symbol by shifting the index value representing the first element in the vector by an offset value, and then adapting to apply modulo 96 (or a corresponding value) to the shifted index value.

[0075] Thus, the calculation of the symbol associated with the first element may be idx shifted (1)=(idx(1)-abs_offset) mod N Q (Equation 11) represented as follows.

[0076] The symbol thus achieved is used by the encoding component 208. The encoding component 208 encodes the first element by performing entropy encoding of the symbol associated with the first element using the same probability table that is used to encode the at least one second element. The offset value may be equal to or at least close to the difference between the most likely index value for the first element and the most likely symbol for the at least one second element in the probability table. In FIG. 3, the most likely index value for the first element is represented by arrow 302. If the most likely symbol for the at least one second element is 0, the value represented by arrow 302 becomes the offset value used. By using the offset approach, the peaks of the distributions in FIGS. 3 and 4 are aligned. This approach avoids the need for a dedicated probability table for the first element, thus saving memory in the audio encoding system 100 and the corresponding decoder. On the other hand, it often maintains nearly the same encoding efficiency as that provided by a dedicated probability table.

[0077] When the entropy encoding of the at least one second element is performed using a Huffman codebook, the encoding component 208 may encode the first element in the vector using the same Huffman codebook that is used to encode the at least one second element. This is by representing the first element with the codeword in the Huffman codebook indexed by the codebook index associated with the first element.

[0078] Since search speed can be important when encoding parameters in an audio decoding system, the memory in which the codebook is stored is preferably high-speed memory and thus expensive. Thus, by using only one probability table, the encoder can be less expensive than when two probability tables are used.

[0079] It may be noted that the probability distributions shown in FIGS. 3 and 4 are often pre-computed for the training data set and thus not computed while encoding the vectors. However, of course, it is also possible to compute the distribution "on the fly" while encoding.

[0080] It may be noted that the above description of the audio encoding system 100 using the vector of parameters encoded from the upmix matrix as the vector of parameters to be encoded is merely an exemplary use. The method of encoding a vector of parameters according to the present disclosure may be used in other applications in an audio encoding system. For example, when encoding other internal parameters in a downmix encoding system, such as parameters used in a parametric bandwidth extension system such as spectral band replication (SBR).

[0081] FIG. 5 is a generalized block diagram of an audio decoding system 500 for regenerating an encoded audio object from an encoded downmix signal 510 and an encoded upmix matrix 512. The encoded downmix signal 510 is received by a downmix receiving component 506, where the signal is decoded and converted to a suitable frequency domain if it is not already in a suitable frequency domain. The decoded downmix signal 516 is then sent to an upmix component 508. In the upmix component 508, the encoded audio object is regenerated using the decoded downmix signal 516 and the decoded upmix matrix 504. More specifically, the upmix component 508 may perform a matrix operation in which the decoded upmix matrix 504 is multiplied by a vector containing the decoded downmix signal 516. The decoding process of the upmix matrix is described below. The audio decoding system 500 further has a rendering component 514 that outputs an audio signal based on the reconstructed audio object 518, depending on the type of playback unit connected to the audio decoding system 500.

[0082] The symbolized upmix matrix 512 is received by the upmix matrix decoder 502. This upmix matrix decoder 502 will be described in detail here in connection with FIG. 6. The upmix matrix decoder 502 is configured to decode a vector of entropy-coded symbols into a vector of parameters related to an aperiodic quantity in an audio decoding system. The vector of entropy-coded symbols includes a first entropy-coded symbol and at least one second entropy-coded symbol, and the vector of parameters includes a first element and at least a second element. Thus, the encoded upmix matrix 512 is received by the receiving component 602 in vector format. The decoder 502 further has an indexing component 604 configured to represent each entropy-coded symbol in the vector by a symbol that can take N values using a probability table. N may be, for example, 96. The associating component 606 is configured to associate the first entropy-coded symbol with an index value by any suitable means depending on the encoding method used to encode the first element in the vector of parameters. Then, the symbol for each of the second symbols and the index value for the first symbol are used by the associating component 606. The associating component 606 associates each of the at least one second entropy-coded symbol with an index value. The index value of the at least one entropy-coded symbol is calculated by first calculating the sum of the index value associated with the entropy-coded symbol preceding the second entropy-coded symbol in the vector of entropy-coded symbols and the symbol representing the second entropy-coded symbol. Then, modulo N is applied to the sum. Without loss of generality, assume that the minimum index value is 0 and the maximum index value is N - 1, for example, 95. Then, the association algorithm is: idx(b)=(idx(b - 1)+Δ idx (b)) mod N Q (Equation 12) and may be expressed as follows. Here, b is an element in the decoded vector, and N Q is the number of possible index values.

[0083] The upmix matrix decoder 502 further has a decoding component 608 configured to represent the at least one second element of the parameter vector by a parameter value corresponding to an index value associated with the at least one second entropy - encoded symbol. Thus, this representation is, for example, a decoded version of the parameters encoded by the audio encoding system shown in FIG. 1. In other words, this representation is equal to the quantized parameters encoded by the audio encoding system shown in FIG. 1.

[0084] According to an embodiment of the present invention, each entropy-encoded symbol in a vector of entropy-encoded symbols is represented by a symbol using the same probability table for all entropy-encoded symbols in the vector of entropy-encoded symbols. The advantage of this is that only one probability table needs to be stored in the decoder's memory. In an audio decoding system, when decoding entropy-encoded symbols, the search speed can be important, so the memory in which the probability table is stored is preferably high-speed memory and thus expensive. Therefore, by using only one probability table, the decoder can be less expensive than when two probability tables are used. According to this embodiment, the association component 606 may be configured to associate the first entropy-encoded symbol in the vector of entropy-encoded symbols with an index value by first shifting the symbol representing the first entropy-encoded symbol in the vector of entropy-encoded symbols by a certain offset value. Then modulo N is applied to the shifted symbol. Thus, the association algorithm may be represented as idx(1)=(idx shifted (1)+abs_offset) mod N Q (Equation 13) as shown.

[0085] The decoding component 608 is configured to represent the first element of the vector of parameters by a parameter value corresponding to the index value associated with the first entropy-encoded symbol. Thus, this representation is, for example, the decoded version of the parameters encoded by the audio encoding system 100 shown in FIG. 1.

[0086] A method for differentially encoding an aperiodic quantity will be further described in connection with FIGS. 7-10.

[0087] Figures 7 and 9 describe an encoding method for four second elements in the vector of parameters. Thus, the input vector 902 contains five parameters. These parameters can take any value between a certain minimum value and a certain maximum value. In this example, the minimum value is -9.6 and the maximum value is 9.4. The first step S702 of the encoding method represents each parameter in the vector 902 by an index value that can take N values. In this case, N is chosen to be 96. That is, the quantization step size is 0.2. This gives the vector 904. The next step S704 calculates the difference between each of the second elements, i.e., the four upper parameters in the vector 904, and its preceding element. Thus, the resulting vector 906 contains four difference values - the four upper values in the vector 906. As can be seen in Figure 9, these difference values can be negative, zero, or positive. As described above, it is advantageous to have difference values that can take N values, in this case 96 values. To achieve this, in the next step S706 of this method, modulo 96 is applied to the second element in the vector 906. The resulting vector 908 does not contain any negative values. The symbol thus achieved, shown in the vector 908, is then used to encode the second element of the vector in the final step S708 of the method shown in Figure 7. It is by entropy encoding the symbol associated with the at least one second element based on a probability table containing the probabilities of the symbols shown in the vector 908.

[0088] As can be seen in Figure 9, the first element is not processed after the indexing step S702. In Figures 8 and 10, a method for encoding the first element in the input vector is described. The same assumptions made in the above description of Figures 7 and 9 regarding the minimum and maximum values of the parameters and the number of possible index values are valid when explaining Figures 8 and 10. The first element 1002 is received by the encoder. In the first step S802 of the encoding method, the parameter of the first element is represented by the index value 1004. In the next step S804, the indexed value 1004 is shifted by an offset value. In this example, the offset value is 49. This value is calculated as described above. In the next step S806, modulo 96 is applied to the shifted index value 1006. The resulting value 1008 is then used to encode the first element by performing entropy encoding of symbol 1008 using the same probability table used to encode the at least one element in Figure 7.

[0089] Figure 11 shows an embodiment 102' of the upmix matrix encoder component 102 in Figure 1. The upmix matrix encoder 102' may be used to encode the upmix matrix in an audio encoding system, such as the audio encoding system 100 shown in Figure 1. As described above, each row of the upmix matrix includes M elements that allow the reconstruction of an audio object from a downmix signal including M channels.

[0090] At a low overall target bitrate, encoding and sending each of the M upmix matrix elements for each object and T / F tile, one by one for each downmix channel, can require an undesirably high bitrate. This can be reduced by "sparsening" the upmix matrix, i.e., attempting to reduce the number of non-zero elements. In some cases, four out of five elements are zero, and a single downmix channel is used as the basis for reconstructing the audio object. Sparse matrices have a probability distribution of encoded indices (absolute or differential) that is different from that of non-sparse matrices. When the upmix matrix contains a large proportion of zeros, the probability of the value zero becomes more likely than 0.5, and Huffman coding is used, the coding efficiency decreases. This is because the Huffman coding algorithm is inefficient when a particular value, e.g., zero, has a probability greater than 0.5. Furthermore, since many of the elements in the upmix matrix have the value zero, those elements contain no information at all. Thus, one strategy could be to select a subset of the upmix matrix elements and only encode and transmit those to the decoder. This can reduce the required bitrate of the audio encoding / decoding system since less data is transmitted.

[0091] To increase the efficiency of encoding the upmix matrix, a dedicated coding mode for sparse matrices may be used. This will be described in detail below.

[0092] Encoder 102′ has a receiving component 1102 adapted to receive each row in the upmix matrix. Encoder 102′ further has a selection component 1104 adapted to select a subset of elements from the M elements of a row in the upmix matrix. In most cases, the subset includes all elements that do not have a value of zero. However, in certain embodiments, the selection component may choose not to select elements with non-zero values, for example, elements with values close to zero. According to various embodiments, the selected subset of elements may include the same number of elements for each row of the upmix matrix. To further reduce the required bit rate, the number of elements selected may be one.

[0093] Encoder 102′ further has an encoding component 1106 adapted to represent each element in the selected subset of elements by its value and its position in the upmix matrix. Encoding component 1106 is further adapted to encode the value and the position in the upmix matrix of each element in the selected subset of elements. Encoding component 1106 may be adapted to encode the value, for example, using modulo differential encoding as described above. In this case, for each row in the upmix matrix and for a plurality of frequency bands or a plurality of time frames, the values of the elements of the selected subset of elements form one or more vectors of parameters. Each parameter in the parameter vector corresponds to one of the plurality of frequency bands or the plurality of time frames. The parameter vector may be encoded using the modulo differential encoding described above. In further embodiments, the parameter vector may be encoded using ordinary differential encoding. In yet another embodiment, encoding component 1106 is adapted to encode each value separately using fixed-rate encoding of the true quantization value of each value, i.e., the quantization value not differentially encoded.

[0094] The following examples of average bitrates were observed for typical content. These bitrates were measured for M = 5, with 11 audio objects to be reconstructed at the decoder side, 12 frequency bands, a quantization step size of the parameter quantizer of 0.1, and 192 levels. For the case where all five elements in each row of the upmix matrix were encoded, the following average bitrates were observed.

[0095] Fixed-rate coding: 165 kb / sec Differential coding: 51 kb / sec Modulo differential coding: 51 kb / sec, provided that the size of the probability table or codebook is halved as described above.

[0096] For the case of sparse encoding, where only one element is selected by the selection component 1104 for each row in the upmix matrix, the following average bitrates were observed.

[0097] Fixed-rate coding (using 8 bits for the value and 3 bits for the position): 45 kb / sec Modulo differential coding for both the value and the position of the element: 20 kb / sec.

[0098] The encoding component 1106 may be adapted to encode the position of each element in the upmix matrix of a subset of elements in the same way as the value. The encoding component 1106 may be adapted to encode the position of each element in the upmix matrix of a subset of elements in a different way compared to the encoding of the value. When encoding the position using differential encoding or modulo differential encoding, for each row in the upmix matrix and for a plurality of frequency bands or a plurality of time frames, the positions of the elements of the selected subset of elements form one or more vectors of parameters. Each parameter in the parameter vector corresponds to one of the plurality of frequency bands or a plurality of time frames. The parameter vector is encoded using the differential encoding or modulo differential encoding described above.

[0099] It may be noted that the encoder 102' may be combined with the encoder 102 of FIG. 2 to achieve the modulo differential encoding of the sparse upmix matrix described above.

[0100] Furthermore, it may be noted that the method of encoding rows in a sparse matrix, while illustrated above for encoding rows in a sparse upmix matrix, may be used to encode other types of sparse matrices well known to those skilled in the art.

[0101] The method of encoding a sparse upmix matrix will be further described hereinafter in connection with FIGS. 13 to 15.

[0102] The upmix matrix is received, for example, by the reception component 1102 of FIG. 11. For each row 1402, 1502 in the upmix matrix, the method includes selecting a subset from among M, for example 5, elements of that row of the upmix matrix (S1302). Next, each element in the selected subset of elements is represented by a value and a position in the upmix matrix (S1304). In FIG. 14, one element is selected as the above subset (S1302). For example, it is element number 3 with a value of 2.34. Thus, the representation may be a vector 1404 having two fields. The first field in the vector 1404 represents a value, for example 2.34, and the second field in the vector 1404 represents a position, for example 3. In FIG. 15, two elements are selected as the above subset (S1302). For example, element number 3 with a value of 2.34 and element number 5 with a value of -1.81. Therefore, the representation may be a vector 1504 having four fields. The first field in the vector 1504 represents the value of the first element, for example 2.34, and the second field in the vector 1504 represents the position of the first element, for example 3. The third field in the vector 1504 represents the value of the second element, for example -1.81, and the fourth field in the vector 1504 represents the position of the second element, for example 5. Next, the representations 1404, 1504 are encoded according to the above (S1306).

[0103] FIG. 12 is a generalized block diagram of an audio decoding system 1200 according to an exemplary embodiment. The decoder 1200 has a receiving component 1206 configured to receive a downmix signal 1210 including M channels and at least one encoded element 1204 representing a subset of M elements of a row in an upmix matrix. Each of the encoded elements includes a value and a position in that row of the upmix matrix. The position indicates which of the M channels of the downmix signal 1210 the encoded element corresponds to. The at least one encoded element 1204 is decoded by an upmix matrix element decoding component 1202. The upmix matrix element decoding component 1202 is configured to decode the at least one encoded element 1204 according to an encoding strategy used to encode the at least one encoded element 1204. Examples of such encoding strategies are disclosed above. The at least one decoded element 1214 is then sent to a reconstruction component 1208. The reconstruction component 1208 is configured to reconstruct a time / frequency tile of an audio object from the downmix signal 1210 by forming a linear combination of the downmix channels corresponding to the at least one encoded element 1204. When forming the linear combination, each downmix channel is multiplied by its corresponding encoded element 1204.

[0104] For example, if the decoded element 1214 includes the value 1.1 and the position 2, the time / frequency tile of the second downmix channel is multiplied by 1.1, which is then used to reconstruct the audio object.

[0105] The audio decoding system 500 further has a rendering component 1216 that outputs an audio signal based on the reconstructed audio object 1218. The type of the audio signal depends on what type of playback unit is connected to the audio decoding system 1200. For example, if a pair of headphones is connected to the audio decoding system 1200, a stereo signal may be output by the rendering component 1216.

[0106] 〈Equivalents, Extensions, Alternatives, etc.〉 Upon reviewing the above description, additional embodiments of the present disclosure will be apparent to those skilled in the art. Although this document and the drawings disclose embodiments and examples, the present disclosure is not limited to these individual examples. Numerous modifications and variations can be made without departing from the scope of the present disclosure as defined by the appended claims. Even if reference numerals appear in the claims, they are not to be construed as limiting the scope.

[0107] Furthermore, from a review of the drawings, the present disclosure, and the appended claims, those skilled in the art of implementing the present disclosure can understand and implement variations to the disclosed embodiments. In the claims, the term "comprising" does not exclude other elements or steps, and a singular representation does not exclude a plurality. The mere fact that certain measures are described in mutually different dependent claims does not indicate that a combination of these measures cannot be used advantageously.

[0108] The systems and methods disclosed above may be implemented as software, firmware, hardware or any combination thereof. In a hardware implementation, the partitioning of tasks among the functional units mentioned in the above description does not necessarily correspond to the partitioning into physical units. Conversely, one physical component may have multiple functions, or one task may be executed by several cooperating physical components. Some or all components of a certain type may be implemented as software executed by a digital signal processor or a microprocessor, or may be implemented as hardware or as an application specific integrated circuit. Such software may be distributed on a computer-readable medium that may include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. Further, it is well known to those skilled in the art that a communication medium typically embodies computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and includes any information delivery medium.

[0109] Some aspects are described below. [Aspect 1] A method for encoding a vector of parameters in an audio encoding system, wherein each parameter corresponds to a non-periodic quantity, the vector having a first element and at least one second element, the method comprising: representing each parameter in the vector by an index value that can take N values; associating each of the at least one second element with a symbol, the symbol being: calculating the difference between the index value of the second element and the index value of its preceding element in the vector; calculated by applying modulo N to the difference; encoding each of the at least one second element by entropy encoding the symbol associated with the at least one second element based on a probability table including symbol probabilities; method. [Aspect 2] associating the first element in the vector with a symbol, the symbol being: shifting the index value representing the first element in the vector by an offset value; calculated by applying modulo N to the shifted index value; further comprising encoding the first element by entropy encoding the symbol associated with the first element using the same probability table used to encode the at least one second element; The method according to aspect 1. [Aspect 3] The method according to aspect 2, wherein the offset value is equal to the difference between the most likely index value for the first element and the most likely symbol for the at least one second element in the probability table. [Aspect 4] The method according to any one of aspects 1 to 3, wherein the first element and the at least one second element of the vector of the parameters correspond to different frequency bands used in the audio encoding system in a specific time frame. [Aspect 5] The method according to any one of aspects 1 to 3, wherein the first element and the at least one second element of the vector of the parameters correspond to different time frames used in the audio encoding system in a specific frequency band. [Aspect 6] The method according to any one of aspects 1 to 5, wherein the probability table is converted into a Huffman codebook, the symbol associated with an element in the vector is used as a codebook index, and the encoding step includes representing each of the at least one second element by a codeword in the codebook indexed by the codebook index associated with the second element. [Aspect 7] The method according to aspect 6 when citing aspect 2, wherein the encoding step includes encoding the first element in the vector using the same Huffman codebook used to encode the at least one second element by representing the first element by a codeword in the Huffman codebook indexed by the codebook index associated with the first element. [Aspect 8] The method according to any one of aspects 1 to 7, wherein the vector of the parameters corresponds to an element in an upmix matrix determined by the audio encoding system. [Aspect 9] A computer-readable storage medium having computer code instructions adapted to execute the method according to any one of aspects 1 to 8 when executed on a device having a processing function. [Aspect 10] An encoder for encoding a vector of parameters in an audio encoding system, where each parameter corresponds to an aperiodic quantity, the vector having a first element and at least one second element, the encoder comprising: a receiving component adapted to receive the vector; an indexing component adapted to represent each parameter in the vector by an index value that can take N values; an associating component adapted to associate each of the at least one second element with a symbol, the symbol being: calculating the difference between the index value of the second element and the index value of its preceding element in the vector; an associating component calculated by applying modulo N to the difference; an encoding component that encodes each of the at least one second element by entropy encoding the symbol associated with the at least one second element based on a probability table including the probability of the symbol. Encoder. 〔Aspect 11〕 A method for decoding a vector of entropy-encoded symbols in an audio decoding system into a vector of parameters related to an aperiodic quantity, the vector of entropy-encoded symbols having a first entropy-encoded symbol and at least one second entropy-encoded symbol, the vector of parameters having a first element and at least one second element, the method comprising: representing each entropy-encoded symbol in the vector of entropy-encoded symbols by a symbol that can take N integer values by using a probability table; associating the first entropy-encoded symbol with an index value; associating each of the at least one second entropy-encoded symbol with an index value, wherein the index value of the at least one second entropy-encoded symbol is: calculating a sum of an index value associated with an entropy-encoded symbol that precedes the second entropy-encoded symbol in the vector of entropy-encoded symbols and a symbol representing the second entropy-encoded symbol; the step calculated by applying modulo N to the sum; and expressing at least one second element of the vector of the parameters by a parameter value corresponding to the index value associated with the at least one second entropy-encoded symbol. Method. 〔Aspect 12〕 By a symbol, the step of representing each entropy-encoded symbol in the vector of entropy-encoded symbols is performed using the same probability table for all entropy-encoded symbols in the vector of entropy-encoded symbols, and the index value associated with the first entropy-encoded symbol is: shifting a symbol representing the first entropy-encoded symbol in the vector of entropy-encoded symbols by an offset value; calculated by applying modulo N to the shifted symbol, The method further comprises: expressing the first element of the vector of the parameters by a parameter value corresponding to the index value associated with the first entropy-encoded symbol. The method according to aspect 11. 〔Aspect 13〕 The method according to aspect 11 or 12, wherein the probability table is converted into a Huffman codebook, and each entropy-encoded symbol corresponds to a codeword in the Huffman codebook. [Aspect 14] In the Huffman codebook, each codeword is associated with a codebook index. By a symbol, the step of representing each entropy-encoded symbol in the vector of entropy-encoded symbols includes representing the entropy-encoded symbol by a codebook index associated with the codeword corresponding to the entropy-encoded symbol. The method according to aspect 13. [Aspect 15] Each entropy-encoded symbol in the vector of entropy-encoded symbols corresponds to a different frequency band used in the audio decoding system in a specific time frame. The method according to any one of aspects 11 to 14. [Aspect 16] Each entropy-encoded symbol in the vector of entropy-encoded symbols corresponds to a different time frame used in the audio decoding system in a specific frequency band. The method according to any one of aspects 11 to 14. [Aspect 17] The vector of the parameters corresponds to an element in an upmix matrix used by the audio decoding system. The method according to any one of aspects 11 to 16. [Aspect 18] A computer-readable storage medium having computer code instructions adapted to execute the method according to any one of aspects 11 to 17 when executed on a device having a processing function. [Aspect 19] A decoder that decodes a vector of entropy-encoded symbols in an audio decoding system into a vector of parameters related to an aperiodic quantity. The vector of entropy-encoded symbols has a first entropy-encoded symbol and at least one second entropy-encoded symbol. The vector of the parameters has a first element and at least a second element. The decoder: A receiving component configured to receive the vector of entropy-encoded symbols; An indexing component configured to represent each entropy-encoded symbol in the vector of entropy-encoded symbols by a symbol that can take N integer values by using a probability table; An associating component configured to associate the first entropy-encoded symbol with an index value, wherein the associating component is further configured to associate each of the at least one second entropy-encoded symbol with an index value, and the index value of the at least one second entropy-encoded symbol is: calculated by adding the index value of the entropy-encoded symbol preceding the second entropy-encoded symbol in the vector of entropy-encoded symbols and the symbol representing the second entropy-encoded symbol; calculated by applying modulo N to the sum, the associating component; and a decoding component configured to represent the at least one second element of the vector of parameters by a parameter value corresponding to the index value associated with the at least one second entropy-encoded symbol. A decoder. 〔Aspect 20〕 A method for encoding an upmix matrix in an audio encoding system, wherein each row of the upmix matrix includes M elements that allow reconstruction of a time / frequency tile of an audio object from a downmix signal including M channels, the method comprising: For each row in the upmix matrix: selecting a subset of elements from the M elements of that row in the upmix matrix; Represent each element in the selected subset of elements by a value and a position in the upmix matrix; Including encoding the value and the position in the upmix matrix of each element in the selected subset of elements, Method. 〔Aspect 21〕 For each row in the upmix matrix, the position in the upmix matrix of the elements of the selected subset crosses multiple frequency bands and / or crosses multiple time frames, the method according to Aspect 20. 〔Aspect 22〕 The selected subset of elements includes the same number of elements for each row of the upmix matrix, the method according to Aspect 20 or 21. 〔Aspect 23〕 For each row of the upmix matrix, the selected subset of elements includes exactly one element from among the M elements of that row in the upmix matrix, the method according to any one of Aspects 20 to 22. 〔Aspect 24〕 For each row in the upmix matrix and for multiple frequency bands or multiple time frames, the values of the elements of the selected subset of elements form one or more vectors of parameters, each parameter in the vector of parameters corresponds to one of the multiple frequency bands or the multiple time frames, and the one or more vectors of parameters are encoded using the method according to any one of Aspects 1 to 8, the method according to any one of Aspects 20 to 23. 〔Aspect 25〕 For each row in the upmix matrix and for a plurality of frequency bands or a plurality of time frames, the positions of the elements of a selected subset of the elements form one or more vectors of parameters, each parameter in the vector of parameters corresponding to one of the plurality of frequency bands or the plurality of time frames, and the one or more vectors of parameters are encoded using the method according to any one of Aspects 1 to 8, the method according to any one of Aspects 20 to 24. [Aspect 26] A computer-readable storage medium having computer code instructions adapted to execute the method according to any one of Aspects 20 to 25 when executed on a device having a processing function. [Aspect 27] An encoder for encoding an upmix matrix in an audio encoding system, each row of the upmix matrix including M elements that allow reconstruction of a time / frequency tile of an audio object from a downmix signal including M channels, the encoder comprising: A receiving component adapted to receive each row in the upmix matrix; A selection component adapted to select a subset of elements from the M elements of the row in the upmix matrix; An encoding component adapted to represent each element in the selected subset of elements by a value and a position in the upmix matrix, the encoding component further being adapted to encode the value and the position in the upmix matrix of each element in the selected subset of elements. Encoder. [Aspect 28] A method for reconstructing a time / frequency tile of an audio object in an audio decoding system, comprising: Receiving a downmix signal including M channels; Receiving at least one encoded element representing a subset of M elements of a row in an upmix matrix, each encoded element including a value and a position in its row in the upmix matrix, the position indicating one of the M channels of the downmix signal corresponding to the encoded element; Reconstructing the time / frequency tile of the audio object from the downmix signal by forming a linear combination of the downmix channels corresponding to the at least one encoded element, wherein in the linear combination each downmix channel is multiplied by the value of its corresponding encoded element; Method. 〔Aspect 29〕 The method according to aspect 28, wherein the position of the at least one encoded element changes across a plurality of frequency bands and / or across a plurality of time frames. 〔Aspect 30〕 The method according to aspect 28 or 29, wherein the number of elements of the at least one encoded element is equal to 1. 〔Aspect 31〕 For a plurality of frequency bands or a plurality of time frames, the values of the at least one encoded element form one or more vectors, each value being represented by an entropy-coded symbol, each entropy-coded symbol in each vector of entropy-coded symbols corresponding to one of the plurality of frequency bands or one of the plurality of time frames, and the one or more vectors of entropy-coded symbols being decoded using the method according to any one of aspects 11 to 17. The method according to any one of aspects 28 to 30. 〔Aspect 32〕 For a plurality of frequency bands or a plurality of time frames, the positions of the at least one encoded element form one or more vectors, each position being represented by an entropy-encoded symbol, each symbol in each vector of entropy-encoded symbols corresponding to one of the plurality of frequency bands or one of the plurality of time frames, and the one or more vectors of entropy-encoded symbols being decoded using the method according to any one of aspects 11 to 17, a method according to any one of aspects 28 to 31. [Aspect 33] A computer-readable storage medium having computer code instructions adapted to execute the method according to any one of aspects 28 to 32 when executed on an apparatus having a processing function. [Aspect 34] A decoder for reconstructing a time / frequency tile of an audio object, comprising: A receiving component configured to receive at least one encoded element representing a subset of M elements of a row in a downmix signal and an upmix matrix including M channels, each encoded element including a value and a position in that row of the upmix matrix, the position indicating one of the M channels of the downmix signal to which the encoded element corresponds; A reconstruction component configured to reconstruct the time / frequency tile of the audio object from the downmix signal by forming a linear combination of the downmix channels corresponding to the at least one encoded element, wherein in the linear combination, each downmix channel is multiplied by the value of its corresponding encoded element; Decoder.

Claims

1. A method for encoding a vector of parameters in an audio encoding system, wherein each parameter corresponds to an aperiodic quantity, the vector having a first element and at least one second element, the method comprising: representing each parameter in the vector by an index value that can take N values; associating each of the at least one second element with a symbol, the symbol being: calculating the difference between the index value of the second element and the index value of the element preceding it in the vector; calculated by applying modulo N to the difference; encoding each of the at least one second element by entropy encoding the symbol associated with the at least one second element based on a probability table including symbol probabilities, wherein the first element and the at least one second element of the parameter vector correspond to different frequency bands used in the audio encoding system in a particular time frame; associating the first element in the vector with a symbol, the symbol being: shifting the index value representing the first element in the vector by subtracting an offset value from the index value; calculated by applying modulo N to the shifted index value; encoding the first element by entropy encoding the symbol associated with the first element based on a probability table including symbol probabilities, method.

2. The method according to claim 1, wherein the offset value is equal to the difference between the most likely index value for the first element and the most likely symbol for the at least one second element in the probability table.

3. The probability table is converted into a Huffman codebook, a symbol associated with an element in the vector is used as a codebook index, and the encoding step includes encoding each of the at least one second element by representing the second element with a codeword in a codebook indexed by a codebook index associated with the second element. The method according to claim 1 or 2.

4. The encoding step includes encoding the first element in the vector using the same Huffman codebook used to encode the at least one second element by representing the first element with a codeword in the Huffman codebook indexed by a codebook index associated with the first element. The method according to claim 3.

5. A computer-readable storage medium having computer code instructions adapted to execute the method according to any one of claims 1 to 4 when executed on a device having a processing function.

Citation Information

Patent Citations

  • Encoder and decoder

    JP2002328699A

  • Coding method and device, decoding method and device, transmission method and device, and storage medium

    JP2003110429A

  • Image encoding method, image decoding method, image encoding device, image decoding device and program

    JP2007520948A

  • Method for binary encoding quantization index of signal envelope, method for decoding signal envelope and corresponding encoding and decoding modules

    JP2009527785A

  • Device and method for execution of huffman coding

    WO2012144127A1