Encoding and decoding of a part of an audio signal
The proposed method addresses the challenges of handling large L1-norms and vector lengths in audio encoding and decoding by using a noiseless split and half hyper-pyramid-based segmentation, resulting in efficient noiseless compression and transmission with low complexity.
Patent Information
- Application Number
- PCT/EP2023/087566
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-26
AI Technical Summary
Existing audio encoding and decoding techniques face challenges in efficiently handling large L1-norms and vector lengths, leading to increased complexity and overhead in index codeword handling.
The proposed method involves a noiseless split of a signed integer vector into smaller segments, using a half hyper-pyramid-based segmentation and recursive decomposition, to achieve efficient noiseless compression and transmission of audio signals. This approach allows for low-complexity encoding and decoding, even with large L1-norms and vector lengths.
The method achieves efficient noiseless transmission of integer input signals with low memory and computational requirements, enabling the handling of very large L1-norms and vector lengths, while minimizing overhead and maintaining high coding efficiency.
Smart Images

Figure EP2023087566_26062025_PF_FP_ABST
Abstract
Description
[0001]ENCODING AND DECODING OF A PART OF AN AUDIO SIGNAL TECHNICAL FIELDEmbodiments presented herein relate to a method, an audio encoder, a computerprogram, and a computer program product for encoding a part of an audio signal.Embodiments presented herein further relate to a method, an audio decoder, acomputer program, and a computer program product for decoding the part of theaudio signal. BACKGROUND When audio or video signals are to be transmitted or stored, the signals are typically encoded. In an encoder, vectors representing audio / video signal samples are encoded to be represented by a number of coefficients or parameters. These coefficients or parameters can then efficiently be transmitted or stored. When the coefficients or parameters are received or retrieved, a decoding of the coefficients or parameters into audio / video signal samples is performed to retrieve the original audio / video signal. Many different kinds of encoding techniques have been used for audio / video signals. One approach is based on vector quantization (VQ). It is known that unconstrained vector quantization (VQ) is the optimal quantization method for grouped samples (vectors) of a certain length. However, the memory and search complexity constraints have led to the development of structured vector quantizers. Different structures gives different trade-offs in terms of search complexity and memory requirements.One such method is the gain-shape vector quantization, where the target vector ^ isrepresented using a shape vector ^ and a scalar gain G: The concept is, instead of quantizing directly the target vector V, to quantize pairs of {v, G}. These Gain (in scalar G) and shape (in vector v) components are thenencoded into ^^ using a shape vector quantizer which is tuned for a normalized shapeinput and a scalar gain quantizer into ^^, which handles the dynamics of the signal.This structure is often applied in e.g., audio coding since the division into dynamics and shape (or fine structure) fits well with the perceptual auditory model. A valid entry in the selected structured vector quantizer, is first searched using theknowledge of the structure (e.g., L1 (absolute amplitude)-normalization or L2(energy)-normalization). After a valid vector has been found one needs to efficiently create an index (or codeword) that represents that specific vector and then transmit that index to the receiver. The index creation (also known as indexing or enumeration) will use the properties of the selected structure and create a unique index (codeword) for the found vector in the structured VQ.On the receiving side the decoder needs to efficiently decompose the index into thesame vector that was determined on the encoder side. This decomposition can be made very low complex in terms of operations by using a large table lookup, but then at the cost of huge stored Read-Only Memory (ROM) tables. Alternatively one can design the decomposition (also known as de-indexing) so that it uses knowledge of the selected structure and potentially also the available target hardware numerical operations to algorithmically decompose the index into the unique vector, in an efficient way. A well designed structured VQ, has a good balance between encoder search complexity, encoder indexing complexity and decoder de-indexing complexity in terms of Million Operations Per Second (MOPS) and in terms of Program ROM and dynamic Random Access Memory (RAM) required, and in terms of Table ROM. Many audio codecs such as CELT, IETF / Opus-Audio and ITU-T G.719 use an envelope and shape VQ and an envelope mixed gain-shape VQ to encode the spectral coefficients of the target audio signal (in the Modified Discrete Cosine Transform (MDCT) domain). CELT / IETF OPUS-Audio use a PVQ-Vector Quantizer, while G.719 uses and slightly extended RE8 Algebraic Vector Quantizer for R=1 bit / dimension coding and a very low complexity D8 Lattice quantizer for VQ rates higher than 1bit / dimension. PVQ stands for Pyramid Vector Quantizer, and is a VQ that uses theL1-norm(i.e., for a vector, this is the sum of the absolute values of the vector elements) to enable a fast search. It has also been found that PVQ may provide quite efficient indexing. The PVQ has been around for some time, but the initial conceptwas developed in the time period 1983-86 by Fischer.PVQ-quantizers have also been used for encoding of time domain and Linear Prediction (LP) residual domain samples in speech coders, and for encoding of frequency domain Discrete Cosine Transform (DCT) coefficients. An advantage with the PVQ compared to other structured VQs is that it naturally can handle any vector dimension, while other structured VQs often are limited to the dimension being multiples, e.g. multiples of 4 or multiples of 8. The IETF / OPUS Codec in Audio mode is employing a recursive PVQ-indexing and de-index scheme that has a maximum size of the PVQ-indices / (short) codewords set to 32 bits. If the target vector to be quantized requires more than 32 bits, the original target vector is recursively split in halves into lower dimensions, until all sub-vectors fit into the 32 bit short codeword indexing domain. In the course of the recursivebinary dimension splitting there is an added cost of adding a codeword for encodingthe energy relation (the relative energies, which can be represented by a quantized angle) between the two split sub target vectors. In OPUS-Audio the structured PVQ- search is made on the resulting split smaller dimension target sub-vectors. The original CELT codec (developed by Valin, Terribery and Maxwell in 2009), is employing a similar PVQ-indexing / deindexing scheme, (with a 32 bit codeword limit)but the binary dimension split in CELT is made in the indexing domain aftersearching and after establishing the initial PVQ-structured vector. The integer PVQ- vector to index is then recursively reduced to smaller than or equal to 32 bit PVQ-vector sub-units in the integer domain. This is again achieved by adding an additionalcodeword for the split, this time for the integer relation between the ‘left’ integer sub- vector and the ‘right’ integer sub-vector, so that one can know the L1-norm of each of the sub PVQ-vectors in the decoder. The CELT post-search integer indexing split approach leads to a variable rate (variable total size index), which can be a disadvantage if the media-transport requires fixed rate encoding. One general issue with structured vector quantization is to find a suitable overall compromise including the methods for efficient search, efficient codeword indexing and efficient codeword de-indexing.Long index codewords (e.g., a 400 bit integer codeword) gives larger complexityoverhead in indexing and deindexing calculations (and special software routines will be needed for multiplying and dividing these large integers in the long codeword composition and decomposition).Short index code words can use efficient hardware operators (e.g., Single InstructionMultiple Data (SIMD) instructions in a 32 bit Digital Signal Processor (DSP)), however at the cost of requiring pre-splitting of the target VQ-vectors (like in IETF / OPUS-Audio), or post-search-splitting the integer PVQ-search result-vector (like in original-CELT). These dimension splitting methods adds a transmission cost for the splitting information codeword (splitting overhead), and the shorter thepossible index-codewords are, the higher number of splits are required, and the resultis an increased overhead for the long index codeword splitting. E.g., 16 bit short PVQ-codewords will result in many more splits than 32 bit short codewords, and thus a higher overhead for the splitting. The PVQ (Pyramid Vector Quantizer) readily allows for a very efficient search,through L1-normalization. Typically, the absolute sum normalized target vector iscreated, followed by vector value truncation (or rounding) and then a limited set of corrective iterations are run to reach the target L1-norm (K) for the PVQ-vector (PVQ-vec). The problem of the previously mentioned CELT / OPUS prior art short codeword indexing schemes is that they are limited to a 32-bit integer range (unsigned 32-bitintegers), each CELT binary split also results in a set of at least three codewords,where the additional codeword always is a quite small and thus the CELT-binary split may result in a quite large number of total final short codewords. Further they cannot be efficiently implemented in a DSP-architecture that only supports fast instructions only for signed 32-bit integers, and not for unsigned 32 bit integers. SUMMARY An object of embodiments herein is to address the above noted shortcomings in relation to encoding of audio or video signals using different VQ approaches. In some aspects, a noiseless split (into segments) of a signed integer vectorrepresenting a part of an audio signal is made into smaller signed sub-vectors, wherethe number of small sub-vectors may be larger than two to achieve a high efficiency in terms of reducing the total number of short codewords. The alternative to separate the input vector into short codeword is using very long codewords with huge mantissas, which are non-realizable in most common digital signal processors. The splitting can be very compact in terms of table space (i.e., memory space), use a new type of header construction, and handle a very large range of L1-norms. The low-complexity splitting can be further enhanced by an adaptive parameterized split-rule selection, i.e., an adaptive segmentation method selection, being used to adapt the splitting to the input source signal and to efficiently optimize the finalnoiseless compression. The split-rule selection and the segmentation selection can bemade efficient in terms of cycle complexity by employing a low-complexity approximation of each leaf’s bit-rate consumption. For excessively large L1-norms the method might be preceded by an amplitude-plane (or bit-plane) slicing approach with an optimized ordering of the bit-plane data. Bit- plane norm reduction can be used to maintain a sufficiently low complexity, andfurther make the final signed integer vector presented to the encoder more Laplacianin term of source distribution. The proposed ordering enables loss of the least significant bits (LSBs) bits without losing decoding capability of the most significant bits (MSBs). Conventional clustering techniques, (e.g., binary recursive split and sign-amplitude separation (CPCE) into shorter codewords) may result in that different parts of the input vector have very different short codeword sizes, which could make the positional encoding very inefficient and results in an increased number of total shortcodewords. Thus, there is a need to provide an improved segmentation method tosplit the integer input vector into segments that can then be handled by a highlyoptimized core for the sub-segment short codeword encoding.In combination, the proposed techniques enable low-complexity noiseless transmission of any integer input signal. In some aspects, a half hyper-pyramid-based segmentation and recursive decomposition of large integer vectors corresponding to very large PVQ / PVC L1- norm based positional codes is applied for efficient noiseless compression and transport of audio / video signals. The proposed tree-parsing method can easily handle pyramids with L1-norms as high as 32767 and lengths L as large as 128, resulting in total information sizes up to at least about 800 bits or more, while legacy methods have been restricted about an eighth of that size. Higher L1-norms than 32767, andhigher lengths L than 128 are also possible For example, when the arithmetic logicunit (ALU) of the DSP is optimized for processing words larger than 16 bits, e.g. a 24bit data and 48-bit accumulator, the ALU may handle L1 norms of up to 223-1,employing the herein disclosed embodiments. The efficiency can be characterized by very low table memory usage and a very compact program code, and further the proposed tree-segmentation approach allows for lower computational operations in cycles in terms of parallelization of theenumeration / de-enumeration processes on both encoder and decoder side.A noiseless L1-norm based method that efficiently handles very large L1-norms and large vector lengths (L) is disclosed. Without this, the resulting overly long indexcodewords would be cumbersome to handle for a normal digital signal processor.Further, the method can minimize the overhead caused by vector splits by using a selectively guided segmentation approach to the integer vector splitting / segmentation methodology. The disclosed methods yield a variable rate for the input L1-norm integer vector,without leading to any larger bit-rate excesses. Essentially, the disclosed methods willwork perfectly hand-in-hand with lossy audio compression, as the excess cases may occur mainly for strongly white noise-like input signals, when the ear is less sensitive to degradation and Perceptual Noise Substitution (PNS) can be employed for thetruncated part of the signal. PNS works even better if the excess bit-rate is at higheraudio frequencies. The choice of the final (and best) split rule and segmentation method can beevaluated in a closed total bit-rate estimation loop before the actual enumeration ofthe final leaf indices, saving the cycle complexity and increasing the efficiency. These and other objects are met by embodiments of the proposed technology. According to a first aspect there is presented a method for encoding a part of an audiosignal. The method is performed by an audio encoder. The method comprisesreceiving an input vector comprising integer coefficients derived from said part of the audio signal. The method comprises building an encoder tree by noiselessly splittingthe input vector into segments. Each of the segments is represented by a respectivenode in the encoder tree. Each of the nodes being a header to another node comprisesa vector of sign information of segments represented by nodes being leaves of othernodes. The segments represented by nodes being leaves have a sign set to a fixed signvalue. The method comprises encoding said part of the audio signal by enumeratingthe segments into a sequence of indices, with a unique index per segment. According to a second aspect there is presented an audio encoder for encoding a partof an audio signal. The audio encoder comprises processing circuitry. The processingcircuitry is configured to cause the audio encoder to receive an input vectorcomprising integer coefficients derived from said part of the audio signal. The processing circuitry is configured to cause the audio encoder to build an encoder treeby noiselessly splitting the input vector into segments. Each of the segments isrepresented by a respective node in the encoder tree. Each of the nodes being aheader to another node comprises a vector of sign information of segmentsrepresented by nodes being leaves of other nodes. The segments represented bynodes being leaves have a sign set to a fixed sign value. The processing circuitry isconfigured to cause the audio encoder to encode said part of the audio signal by enumerating the segments into a sequence of indices, with a unique index per segment.According to a third aspect there is presented a computer program for encoding apart of an audio signal. The computer program comprises computer code which,when run on processing circuitry of an audio encoder, causes the audio encoder toperform actions. One action comprises the audio encoder to receive an input vector comprising integer coefficients derived from said part of the audio signal. One action comprises the audio encoder to build an encoder tree by noiselessly splitting theinput vector into segments. Each of the segments is represented by a respective nodein the encoder tree. Each of the nodes being a header to another node comprises a vector of sign information of segments represented by nodes being leaves of other nodes. The segments represented by nodes being leaves have a sign set to a fixed sign value. One action comprises the audio encoder to encode said part of the audio signal by enumerating the segments into a sequence of indices, with a unique index per segment.According to a fourth aspect there is presented a method for decoding a part of anaudio signal. The method is performed by an audio decoder. The method comprisesreceiving a bitstream. The bitstream is composed of a sequence of indices, with aunique index per segment, and represents said part of the audio signal. The method comprises extracting the indices from the bitstream. The method comprises buildinga decoder tree from the indices. There is one unique index per segment. Each of thesegments is represented by a respective node in the decoder tree. Each of the nodes being a header to another node comprises a vector of sign information of segments represented by nodes being leaves of other nodes. The segments represented bynodes being leaves have a sign set to a fixed sign value. The method comprisesrecreating an input vector comprising integer coefficients by moving the sign information from the vector of sign information to the segments and thenconcatenating all the segment in an order given by the decoder tree. The input vectorrepresents the decoded part of said audio signal.According to a fifth aspect there is presented an audio decoder for decoding a part ofan audio signal. The audio decoder comprises processing circuitry. The processingcircuitry is configured to cause the audio decoder to receive a bitstream. Thebitstream is composed of a sequence of indices, with a unique index per segment, and represents said part of the audio signal. The processing circuitry is configured to cause the audio decoder to extract the indices from the bitstream. The processingcircuitry is configured to cause the audio decoder to build a decoder tree from theindices. There is one unique index per segment. Each of the segments is representedby a respective node in the decoder tree. Each of the nodes being a header to anothernode comprises a vector of sign information of segments represented by nodes beingleaves of other nodes. The segments represented by nodes being leaves have a sign setto a fixed sign value. The processing circuitry is configured to cause the audio decoderto recreate an input vector comprising integer coefficients by moving the sign information from the vector of sign information to the segments and thenconcatenating all the segment in an order given by the decoder tree. The input vectorrepresents the decoded part of said audio signal.According to a sixth aspect there is presented a computer program for decoding apart of an audio signal. The computer program comprises computer code which,when run on processing circuitry of an audio decoder, causes the audio decoder to perform actions. One action comprises the audio decoder to receive a bitstream. The bitstream is composed of a sequence of indices, with a unique index per segment, and represents said part of the audio signal. One action comprises the audio decoder to extract the indices from the bitstream. One action comprises the audio decoder tobuild a decoder tree from the indices. There is one unique index per segment. Each ofthe segments is represented by a respective node in the decoder tree. Each of the nodes being a header to another node comprises a vector of sign information of segments represented by nodes being leaves of other nodes. The segments represented by nodes being leaves have a sign set to a fixed sign value. One action comprises the audio decoder to recreate an input vector comprising integer coefficients by moving the sign information from the vector of sign information to the segments and then concatenating all the segment in an order given by the decodertree. The input vector represents the decoded part of said audio signal.According to a seventh aspect there is presented a computer program productcomprising a computer program according to at least one of the third aspect and thesixth aspect and a computer readable storage medium on which the computerprogram is stored. The computer readable storage medium could be a non-transitory computer readable storage medium. Advantageously, these aspects overcome the above noted shortcomings in relation to encoding of audio or video signals using different VQ approaches.Advantageously, these aspects enable the use of very large L1-norm pyramids in anyquantization scheme.Advantageously, these aspects require low memory requirements and lowcomputational requirements. Advantageously, the proposed encoding and decoding can be recursive, resulting in low memory requirements and low computational requirements, even though the vector length and vectors dynamics may be large. Advantageously, the same splitting approach and enumeration features can be usedfor headers as well as for leaves, thereby reducing the algorithmic footprint, andenabling larger utilization of the short codewords.Advantageously, the header and leaf short codeword enumeration can be fullyparallelized on the encoder side.Advantageously, the leaf short codeword de-enumeration can be fully parallelized onthe decoder side. Advantageously, the encoding (and decoding) can be adaptively tailored to other distributions than the natural PVQ-Laplacian source, without requiring large sets of quantization tables, by using an adaptive segmentation approach. Other objectives, features and advantages of the enclosed embodiments will be apparent from the following detailed disclosure, from the attached dependent claims as well as from the drawings. Generally, all terms used in the claims are to be interpreted according to theirordinary meaning in the technical field, unless explicitly defined otherwise herein. Allreferences to "a / an / the element, apparatus, component, means, module, step, etc." are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, module, step, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not have to be performed in the exact order disclosed, unless explicitly stated. BRIEF DESCRIPTION OF THE DRAWINGS The inventive concept is now described, by way of example, with reference to the accompanying drawings, in which:Fig. 1 is a schematic diagram illustrating a system according to embodiments;Figs. 2 and 3 are flowcharts of methods according to embodiments;Fig.4, 5, and 6 are illustration if trees according to embodiments;Fig. 7 is a visualization of different segmentations according to embodiments;Fig. 8, 9, and 10 are block diagrams according to embodiments;Fig. 11 is a schematic diagram showing structural units of an audio encoder / decoderaccording to an embodiment; andFig. 12 shows one example of a computer program product comprising computerreadable means according to an embodiment. DETAILED DESCRIPTION The inventive concept will now be described more fully hereinafter with reference tothe accompanying drawings, in which certain embodiments of the inventive conceptare shown. This inventive concept may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete, and will fully convey the scope of the inventive concept to those skilled in the art. Like numbers refer to like elements throughout the description. Any step or feature illustrated by dashed lines should be regarded as optional. To avoid confusion with traditional lossy pyramid vector quantization (PVQ), the term Half-Pyramid Vector Coding (HPVC) is used to describe the procedure of enumerating an L dimensional vector with a fixed leading sign and fixed L1-norm K. In this case, the coding space subsequent to a fixed leading sign forms a hyper-half pyramid. The embodiments disclosed herein relate to techniques for encoding a part of anaudio signal and decoding a part of an audio signal. In order to obtain suchtechniques, there is provided an audio encoder, a method performed by the audioencoder, a computer program product comprising code, for example in the form of a computer program, that when run on processing circuitry of the audio encoder, causes the audio encoder to perform the method. In order to obtain such techniques,there is further provided an audio decoder, a method performed by the audiodecoder, and a computer program product comprising code, for example in the form of a computer program, that when run on processing circuitry of the audio decoder, causes the audio decoder to perform the method.In Fig. 1 is schematically illustrated a system 100 comprising an audio encoder 200a,an audio decoder 200b, and a storage 110. The audio encoder 200a is configured to encode an audio signal into a bitstream. The bitstream can be send either to the storage 110 or directly to the audio decoder 200b. The audio decoder 200b isconfigured to reproduce the audio signal from the bitstream. The audio decoder 200bcan obtain the bitstream either directly from the audio encoder 200a or from thestorage 110. The audio encoder 200a and the audio decoder 200b could be part ofanother device, such as a user equipment a piece of consumer electronics, etc., where audio encoding and / or decoding is to be performed.Reference is now made to Fig. 2 illustrating a method for encoding a part of an audiosignal as performed by the audio encoder 200a according to an embodiment.S102: The audio encoder 200a receives an input vector. The input vector comprisesinteger coefficients derived from the part of the audio signal.S104: The audio encoder 200a builds an encoder tree by noiselessly splitting theinput vector into segments. Each of the segments is represented by a respective nodein the encoder tree. Each of the nodes being a header to another node comprises a vector of sign information of segments represented by nodes being leaves of other nodes. The segments represented by nodes being leaves have a sign set to a fixed sign value.S106: The audio encoder 200a encodes the part of the audio signal by enumeratingthe segments into a sequence of indices, with a unique index per segment.Embodiments relating to further details of encoding a part of an audio signal asperformed by the audio encoder 200a will now be disclosed with continued referenceto Fig.2.Aspects of pre-processing will be disclosed next.In the context of audio coding, and input audio signal might be segmented intoframes and most likely processed using some form of time-to-frequency transform according to state-of-the-art audio coding techniques. One common technique is to use windowing of overlapping input frames and a transform into a frequency domain,such as to the modified discrete cosine transform (MDCT) domain. This yields a datavector of fixed size for each frame. This data represents vectors (with one vector perframe) that is sent to the (pyramid vector) encoder. The length of each vector is Lin.The L1-norm of the vector is also calculated to get Kin. The following processing is then performed per vector. It is determined if the vector can be encoded into a unique index of limited number ofbits. Here, the required number of bits needed can be determined from Lin and Kin.If this is lower than a certain limit BITLIM, the vector can be encoded directly, andotherwise it needs to be split. This is done by building an encoder tree (such as an HPVC-tree). Aspects of the tree building will be disclosed next.A tree structure is bult when a split is needed to keep track of how the vector is splitinto sub-vectors. The elements in a tree are referred to as nodes and they can beeither headers or leaves, depending of the type of data they contain. The top node isalways a header and it contains information on how the vector is split and the parameters of its sub-vectors. The header only includes the leading signs and the L1norms of its sub-vectors. That is, it does not contain information about other signs inthe sub-vectors, or the individual amplitudes in the sub-vectors. Nodes under theheader can be either headers (if further split is needed) or leaves (if no further split isneeded). The top header itself can also be split, in the case the information in the topheader exceeded the limit BITLIM. That is, in some embodiments, one of the nodes isa top header with respect to all other nodes, and the top header is split into at least two sub-headers when the top header has a bit-rate consumption that exceeds a bit- rate consumption limit (i.e., the limit BITLIM).In general terms, a full pyramid vector input can be separated into a list with aleading sign LSglob and a remaining half-pyramid vector yh. Vector yh is a half-pyramid vector and is created by vector multiplication of the vector input with theleading sign value LSglob. A half-pyramid requires that the first non-zero element ofyh is of a given polarity. The half-pyramids can be selected to have a positive initialsign, for the first non-zero element in yh. In some aspects, during the treeconstruction step, two or more L1-norms of the sub-vectors are moved into the header vector. That is, in some embodiments, the vector of sign information is constructed by two or more L1 norms of the segments being moved from thesegments into the vector of sign information. In some aspects, the headers areconstructed by, during the tree construction step, shifting a leading sign of a sub-segment to the header vector. Therefore, in some embodiments, the vector of signinformation is constructed by leading signs of the segments being shifted to thevector of sign information from the segments. In some aspects, all sign information isshifted to the header. That is, in some embodiments, the leading signs of all the segments with a non-zero L1-norm are shifted to the vector of sign information in the header.The actual split is based on the limit BITLIM, or a value Klim that corresponds to thelimit BITLIM, where Klim (1,L) is the maximum number of unit pulses in a half- pyramid H(L, K) that can fit within a bit limit of BITLIM, for a vector of length L. In some embodiments, how many nodes the encoder tree comprises is determined based on the received signed integer input vector, an approximation of a bit-rate consumption of each segment, and a total bit-rate consumption for the received input vector. In some embodiments, the encoder tree is built according to a splitting rule, where the splitting rule is selected from a set of available splitting rules where the splitting rule yielding lowest total bit-rate consumption is selected for building theencoder tree. In some embodiments, the nodes of the encoder tree are built inparallel.Aspects of segmentation will be disclosed next.Segmentation of a half-pyramid is performed into a list of Ns sub-segment leaves andan initial header, where Ns is determined by a segmentation rule, denoted Sr. In particular, in some embodiments, the signed integer input vector comprises data points, and how the data points are distributed over the segments when noiselesslysplitting the signed integer input vector into the segments is defined by at least onesegmentation rule. The splitting rule Sr can be defined using a lookup table lookuprwith a list of limiting K values for each split possibility 1:NsMax, for each possible length L of yh.Table 1 and Table 2 show realizations of two different splitting rules as tabulated Klimvalues for each vector length L and number of segments Ns. The example tables weregenerated for a short codeword limit BITLIM of 31 bits, a maximum L1 norm of32767, a vector dimension L ranging from 1 to 20, and the possible number ofsegments ranges in the row dimension ranges from 1 to Ns,max, where Ns,max =6was used. The first column with Ns=1 correspond to the limiting Klim values for Ls inthe range of 1 to 20, and may be employed to decide if a subvector may beenumerated into a single short codeword of a size less than 2BITLIM. / *Ns=1, Ns=2, Ns=3, Ns=4, Ns=5, Ns,max=6 * / 32767, 32767, 32767, 32767, 32767, 32767, / *L==1 row* / 32767, 32767, 32767, 32767, 32767, 32767, / *L==2 row* / 32767, 32767, 32767, 32767, 32767, 32767, / *L==3 row* / 1172, 32767, 32767, 32767, 32767, 32767, 238, 32767, 32767, 32767, 32767, 32767, 95, 12257, 32767, 32767, 32767, 32767, 53, 3065, 32767, 32767, 32767, 32767,36, 1164, 30571, 32767, 32767, 32767, / / e.g.K=36 Ns^1,K=322 Ns^2,K=30571 Ns^3 27, 572, 9997, 32767, 32767, 32767, 22, 334, 4246, 32767, 32767, 32767, 18, 219, 2163, 21307, 32767, 32767, 16, 157, 1256, 10054, 32767, 32767, 15, 119, 805, 5415, 32767, 32767, 13, 95, 555, 3228, 18756, 32767, 12, 79, 406, 2083, 10672, 32767, 12, 67, 311, 1431, 6577, 31652, 11, 58, 247, 1035, 4323, 18861, 11, 52, 203, 780, 2996, 11987, 10, 47, 170, 609, 2170, 8038, 10, 43, 146, 489, 1630, 5637, / * L=20 * / Table 1: Example splitting rule r=1 for generating a deeper HPVC-tree; design ofsplitting rule S1. Late split assuming 7% compression gain of the HPVC tree / *Ns=1, Ns=2, Ns=3, Ns=4, Ns=5, Ns,max=6 * / 32767, 32767, 32767, 32767, 32767, 32767, / *L==1 row* / 32767, 32767, 32767, 32767, 32767, 32767, 32767, 32767, 32767, 32767, 32767, 32767, 1172, 32767, 32767, 32767, 32767, 32767, 238, 10781, 32767, 32767, 32767, 32767, 95, 2021, 32767, 32767, 32767, 32767, 53, 682, 13764, 32767, 32767, 32767 36, 321, 4219, 32767, 32767, 32767, / *e.g. K=322 Ns^3, K=30571 Ns^4* / 27, 185, 1767, 18336, 32767, 32767, 22, 122, 910, 7280, 32767, 32767, 18, 89, 540, 3514, 22837, 32767, 16, 69, 356, 1953, 10708, 32767, 15, 56, 253, 1206, 5737, 25761, 13, 47, 191, 807, 3405, 13620, 12, 41, 150, 574, 2188, 7929, 12, 36, 123, 430, 1499, 4984, 11, 33, 103, 335, 1080, 3334, 11, 30, 89, 270, 812, 2346, 10, 28, 78, 223, 632, 1722, 10, 26, 70, 189, 507, 1310, / * L=20 * / Table 2: Example splitting rule lookupr r=2, targeting a flatter HPVC-tree; design ofsplitting rule S2. Early split assuming 14% compression cost of the HPVC-treeUsing the r=1 splitting rule S1 of Table 1, an example sub-vector of length L=8, withan L1-norm K of 30571, would result in an Ns of 3. That is, the sub-vector would besplit into three segments. On the other hand using the r=2 splitting rule S2 of Table 2for the very same sub-vector would result in an Ns of 4 and thus create a flatter treethan the resulting tree of Table 1.Some example segmentation methods, denoted Mt, are: {uniform, log, reverse-log,bell.}. In particular, in some embodiments, the at least one segmentation rule is selected from a set of segmentation rules where a first segmentation rule defines a uniform distribution, a second segmentation rule defines a logarithmic distribution, a third segmentation rule defines a reversed logarithmic distribution, and a fourth segmentation rule defines a bell-shaped distribution.The Sr:s and Mt:s are determined offline, and are the same on the encoder and thedecoder side.During the segmentation, each leaf is treated as a half-pyramid and every header isalso treated as a half-pyramid. That is, if the initial header is too large to fit within BITLIM bits it is treated as a half-pyramid. This means that the herein disclosed segmentation allows for recursion in both the header and in the leaves.Segmentation can be repeated until termination criterion is fulfilled. That is, in someembodiments, at least some of the nodes represent sub-segments resulting from atleast one of the segments having been recursively split into sub-segments, where therecursively split is applied until a termination criterion is met. Further, in someembodiments, the termination criterion is met when the bit-rate consumption of eachsegment no longer exceeds a bit-rate consumption limit. The approximation of thebit-rate consumption of a given segment can be evaluated when determining whetheror not to split this given segment using a dynamic selection of either L-directionequations per branch or K-direction equations per branch (where for L-directionequations, the dimension, or width, is the most contributing factor, and for K-direction equations the L1-norm is the most contributing factor).Aspects of tree candidate evaluation will be disclosed next.The tree building and segmentation describe how one tree is built – and this resultsin a required bit-rate to for the tree to be enumerated and encoded into a bitstream.It is also possible to build different trees using different segmentation rules and segmentations methods and then select the tree candidate that requires the lowesttotal bit-rate for the original input vector (before proceeding to the enumeration).Therefore, in some embodiments, the at least one segmentation rule is selected according to the splitting rule. That is, which segmentation rule to use can be defined by the splitting rule, thereby making it possible to build different trees using different segmentation rules and segmentations methods and then select the tree candidate that requires the lowest total bit-rate for the original input vector.Aspects of enumeration will be disclosed next.Enumeration can be implemented of each half-pyramid vector v into a unique indexthat uniquely identifies the corresponding input vector. In some examples, the half-pyramid enumeration is realized in the index following the leading sign. In generalterms, for a half-pyramid the leading sign is fixed, so there is no index for the leadingsign itself.Aspects of encoding will be disclosed next.As the resulting size of the in the indices often is represented by a fractional bitnumber, it is possible to further reduce the required bit-rate through the use ofarithmetic, or range, coding with a uniform cumulative distribution function (CDF) toconcatenate the fractional indices before the resulting bitstream is sent to thechannel. Further, as there can still be some statistical correlation to be utilized in theencoded indices, it might be possible to further reduce the required bit-rate throughthe use of arithmetic / range coding with a trained (or a nearly uniform) CDF beforethe resulting bitstream is sent to the channel. Therefore, in some embodiments, theaudio encoder 200a is configured to perform (optional) step S106-4 as part ofencoding the part of the audio signal.S106-4: The audio encoder 200a applies arithmetic encoding or range encoding to the enumerated segments.Before transmission, the order of the list of indices in the tree is reversed, so that thedecoder receives the tree structural information (i.e., the headers) before the leaves.Therefore, in some embodiments, the audio encoder 200a is configured to perform(optional) step S106-2 as part of encoding the part of the audio signal.S106-2: The audio encoder 200a sorts the indices into an order of indices in thesequence of indices. Indices of segments representing headers in the encoder tree appear before indices of segments representing leaves in the encoder tree. This enables sequential decoding of the tree at the receiver side. However, tree reversal is generally only needed if the tree construction was a breadth first traversal. In some examples, encoder side breadth first traversal is preferred as it simplifies the tree construction / growth-handling. However it is also possible to use dynamic programming and create the tree in a depth first fashion, and then no reversal of the tree is needed before transmission. The bitstream is then either transmitted to the receiver or stored for later use.Therefore, in some embodiments, the audio encoder 200a is configured to perform(optional) step S108. S108: The audio encoder 200a provides the sequence of indices towards an audio decoder 200b or to a data storage 110.Reference is now made to Fig. 3 illustrating a method for decoding a part of an audiosignal as performed by the audio decoder 200b according to an embodiment.S202: The audio decoder 200b receives a bitstream. The bitstream is composed of asequence of indices, with a unique index per segment, and represents the part of theaudio signal. S204: The audio decoder 200b extracts the indices from the bitstream. S206: The audio decoder 200b builds a decoder tree from the indices. There is one unique index per segment. Each of the segments is represented by a respective node in the decoder tree. Each of the nodes being a header to another node comprises a vector of sign information of segments represented by nodes being leaves of other nodes. The segments represented by nodes being leaves have a sign set to a fixed sign value. S208: The audio decoder 200b recreates an input vector comprising integer coefficients by moving the sign information from the vector of sign information to the segments and then concatenating all the segment in an order given by the decodertree. The input vector represents the decoded part of the audio signal.Embodiments relating to further details of decoding a part of an audio signal asperformed by the audio decoder 200b will now be disclosed with continued referenceto Fig.3. Further aspects of the decoding will be disclosed next. In some embodiments, the nodes of the decoder tree are built in parallel. In the decoder the process in the encoder basically reversed and starts by extractionof the encoded indices by feeding the bitstream to an arithmetic / range decoder,(together with each index size). Hence, in some embodiments, the audio decoder 200b is configured to perform (optional) step S204-2 as part of extracting the indices from the bitstream. S204-2: The audio decoder 200b applies arithmetic decoding or range decoding to the bitstream.The indices contain information that makes it possible for the audio decoder 200b toextract the tree structure and this can then be used by the audio decoder 200b to recreate the vector segments and concatenate them in the correct order to recreate the vector (or at least an approximation in the case of lossy coding) initially sent as input to the encoder.Aspects of post-processing will be disclosed next.In the audio coding context, the decoder vector can then be used as input to theinverse transform to get back to the time domain signal. In continuation of the aboveexample, the inverse modified discrete cosine transform (IMDCT) would be used andits associated overlap add processing to recreate the input signal frame of audio. Further aspects of method for encoding and decoding a part of an audio signal will be disclosed next. The procedure of noiseless encoding into an index is equivalent to the enumeration of the L-dimensional vector into a one-dimensional index. An encoding of the half pyramid vector ^^to the coded index ^^^^^is denoted by: ^^^^^ : = ^^^^(^^, ^, ^),while decoding (that is, de-enumeration of an index into an L-dimensional vector) is denoted by: It should be noted that in the normal case (where all sub-indices are received) it isexpected that vector ^ on the decoding (and receiving side) is equivalent to vector ^^.The L1-norm based tree-splitting conditions using a unsigned input vector will first be described. Then the example of having a signed input vector will be described.Assume the sequence ^^ to be an unsigned (e.g., all positive) integer vector of length^, the all positive vector is an half pyramid vector, ^ ∈ ^(^, ^). For ease of notationthe notation ^ will be used instead of ^^ in the subsequent description. That is^ ≔ (^^, … , ^^) ≝ ^^ .Denote the ^^-norm of the source sequence as ^ ≔ ^|^|^^. Starting from the root node,let ^^ ≔ ^. Then, if ^^ in conjunction with ^ yields a larger number of the HPVCenumeration index size, ^^^^^ (^, ^^). Further aspects of the HPVC enumeration will be disclosed below. For now, it is noted that the HPVC enumeration index sizecorresponds to the total number of points in a half pyramid of dimension L and witha L1-norm of K1), than allowed for by the short codeword length, that is, if^^^^^(^, ^^) > (2^^^^^^ − 1), the method splits the source sequence ^ into ^^^parts, where ^^^ is computed by a known (fixed for a specific tree configuration) splittingrule Sr, which only depends on ^^ and ^. The number ^^^^^(^, ^^) does not need tobe computed at runtime as for a given word length BITLIM a limiting ^^^^value foreach length L may be precomputed and stored in a table of length Lmax. Denote thesequence of vectors obtained by this split … ^^^ . Then this yields a new set ofabsolute sums ^^,^, … , ^^,^^ and lengths ^^,^ … ^^,^^ . For each of these pairs, a newlitting rule S can be applied, yielding a se^,^,^^ sp^ … ^ ^r quence of number of splits, ^^ ^.These again yield new values ^^,^,^, … , and corresponding sequencelengths, and the procedure can be recursively repeated. If at any point ^^,^,…,^ ^= 1, thesub-sequence ^^,^,…,^will not be split further, which results in a leaf of the encoding tree. Once all leaves are created, each sub-sequence can be encoded with a HPVC enumeration under the (short) codeword length constraint given by the systems wordlength. Fig. 4 illustrates an example of a tree 400 built according to these principles.The presented naming convention (with series, sequences, list and vector numbering starting from index 1), results in the following relationships between the tree parameters: To produce a decodable codeword, some requirements should be fulfilled. Accordingto a first requirement, the splitting rule Sr should be uniquely retrievable from onlythe ^ and ^ values of the parent node. If more than a single split rule is in use, extra side information may transmitted to indicate the applied splitting rule Sr, r ∈[0 … maxNbRules] at an efficiency cost. According to a second requirement, each sourcevector split should yield a header containing the ^-values of the created subvector,which needs to be transmitted. These headers are themselves integer sequences andcan therefore be encoded in the same way as for the leaf encoding.A split-rule lookup-table ^^^^^^^, can be used to store the maximum ^^-norm ^^^^for a given maximum number of bits that a short codeword HPVC enumerationadmits, depending on the length L of the sequence to code and the number of splits^^ that is used to create sub-sequences form the input sequence. Each length ^ in^^^^^^^, has an ordered sequence of limiting ^^-norm Klimvalues, and each column in lookupr corresponds to a given number of required splits. Lookup table ^^^^^^^row L and column 1 indicates the maximum K that can be supported by 1 split (where “1”corresponds to no split) and Lookup table ^^^^^^^ row L and column 6 indicates themaximum K that can be supported by 6 splits, etc. These optimal split limiting K-values in ^^^^^^^for columns 2 to Ns,max are source distribution dependent, and alookup table lookupr can be constructed to create early splits or late splits, and thusaffect the structure of the final HPVC-tree.In the case the original source sequence ^ would yield an optimal HPVC-index that islarger than the number of available bits per codeword as allowed for by the targetsystem, row L in the table ^^^^^^^ is used together with an iterative rule for splittingthe source sequence into smaller sub-sequences. Using the splitting rule Sr, the encoding tree can be created recursively. The recursiveencoding determines the node (leaf or header) and leaf order (enumeration). That is,it eventually determines the final order of the HPVC short codewords (in terms of astream of indices) as they appear in the transmitted bitstream. The encodingalgorithm presented here is using a breadth first tree-traversal approach and results in a post-order traversal, and therefore yields a list of short codewords in post-order. With the additional information supplied by the split rule, the full tree can be sequentially reconstructed from only it’s post-order enumeration at the decoder.Fig. 5 illustrates an example tree configuration 500 with post-order numbering,using zero-indexing of the codewords (from 1 to 8). The resulting codeword-sequence would therefore be the HPVC encoded version of the following sequence, where eachelement of the list ^ is encoded separately with the HPVC-short-codeword-encodingengine to obtain the codeword-sequence ^. ^^^,^, ^^,^, ^^,^^ ^That is, the above sequence shows the sub-sequences that will be separately encoded by the HPVC-encoding engine in-order. Together with the final codeword-list ^, theencoder will first transmit the initial values ^^ and ^ (or deduce the top level K1 andtop level L values from the common encoder and decoder configuration). With thisinformation and knowledge about the split rule r, the segmentation method t thesource sequence can be decoded at the decoder. The number of elements in the list, ^̃, for a given tree with number treeNb is denotedas Nidx,treeNb. For example, c^ in the current example has Nidx,treeNb = 8 elements.The set of indices ^^^ for the sub-sequences (i.e., the list of short codewords) ^ aretransmitted over transport channel in the reverse order. That is, ^^^= where Ij(y) indicates the short codeword enumeration of an all-positive integersequence y, and use subindex j to indicate its transmission order. If y would havebeen a signed vector, Ij(y) would indicate the enumeration of a signed vector. Thebenefit of this first transmission order is that the decoder can initiate decoding right away from the top-level index, which is now first, and also handle a truncated series of indices and still provide a partially meaningful result.When received on the decoder side, the list of received indices are denoted Irx, andperfect transmission over any transport channel can be represented by: ^^^: = ^^^ For each source sequence ^, the decoder receives a codeword-string of (post-order) HPVC-indices in the list of short codewords ^, and the values ^, ^, corresponding to the length of ^, and ^^-norm respectively. The decoding procedure then traverses thetree in reverse-encoding order, starting from the root node (corresponding to the topheader), which was transmitted first. To decode any HPVC-index, the decoder needsto possess knowledge about the ^^^^and ^^^^values at the current position in the tree. To decode the root node index, that is, the first element of Irx (corresponding tothe last element at the encoder side ^ and the values ^ and ^ of the complete sourcesequence are used to get: a) the values ^^,^, … , ^^,^^^ , if the original codeword cannot be HPVC-coded as awhole, due to the ^^^^^^ constraint, orb) the source-sequence ^, if ^ ≤ ^^^^, where ^^^^ = lookup^(1, ^ = 8).In case of b), the decoding process is then finished. In case of a), the tree can berecursively traversed further by using the obtained ^ values to decode the next nodes.The result of the HPVC-decoding engine should then be appended in front of the decoded sequence as present so far. To keep track of the nodes as they are traversed by the recursion, the decodermaintains a state variable ^^^^^^^- The variable ^^^^^^^ is initialized to 1 (pointing toI1 ), and is increased by one after each time any node of the decoding tree is processed.For calling the decoding function from the outside, the decoder is provided with theinitial indexes I1 in Itx, and further provided with the expected total length of Lsamples. Through the recursive calls of the decoder tree traversing function, in eachwhen an index Iidx_base of the next node is required, it is obtained by reading it fromthe arithmetic / range decoder using the so far decoded L1-norm and the length value for the node.Next follows an illustration of the decoding procedure using the example trees 400and 500 in Fig. 4 and Fig. 5. In this example the first received (e.g., obtained from thearithmetic / range decoder) index is ^^,^, ^^,^^^, and the last received index atthe decoder side is ^^(^^,^).1. Use ^: = ^^, ^ to compute ^^^ = 3, ^ = (^^, ^^, ^^) with getSplitRule().2. Set ^^^^^^^to 1,3. Use ^, ≔ ^^, ^^^ and demultiplex I1 as ^^^(^^^^^^^) from the A / R decoder usinga uniform CDF over the range [ 1 … ^^^^^(^^, ^^^ )], decode ^^^(^^^^^^^), yields 4. Use ^^,^, ^^ to compute ^ ^,^^ = 2, ^ = ^^^,^, ^^,^, ^ with getSplitRule().5. Use ^ , ^^^^ to demultiplex and decode ^^^(^^^^^^^), yields ^ ^^^^^^^+= 16. Use ^ , ^ ^,^,^^,^,^ ^,^ to compute ^^ = 1, ^ = ^^,^,^ = ^^,^ with getSplitRule().7. Use ^^,^,^, ^^,^ to demultiplex and decode ^^^(^^^^^^^), yields ^ =^^^,^^, ^^^^^^^+= 18. Use ^ ^,^,^^,^,^, ^^,^ to compute ^^ = 1, ^ = ^^,^,^ = ^^,^ with getSplitRule().9. Use ^^,^,^, ^^,^ to demultiplex and decode^^^(^^^^^^^), yields ^ = ^^^,^, ^^,^^,^^^^^^^+= 110. Use ^^,^, ^^ to compute ^ ^,^^ = 1, ^ = ^^,^ = ^^ with getSplitRule()11. Use ^^,^, ^^ to demultiplex and decode ^^^(^^^^^^^), yields ^ = 12. ...After decoding the last codeword, the total L1norm has been reached andresulting integer vector has a dimension of L samples, and the recursive procedure isfinished and yields ^^^^^^ = ^.If the decoder traversing function tries to decode more than ^ samples or cannotobtain a new short codeword index from the arithmetic / range decoder, then a bit error has occurred or the received bitstream was truncated. If the decoder identifies one of these occurrences, it may trigger an appropriate packet loss concealment method, such as zeroing the L samples, or zeroing the non-decoded truncated samples.To reduce rate variability and increase coding efficiency, it can be beneficial toimplement a leading sign coding algorithm into the structured HPVC-tree traversal.Such an algorithm might ensure that the leading sign of any sequence encoded by theHPVC engine is fixed. A fixed leading sign significantly reduces the de-enumeration search space of the coder by a factor of two; instead of full hyper-pyramids, only half(hyper-) pyramids are enumerated during short codeword index encoding / decoding.To indicate the global leading sign (LS) of the full source sequence, a single initial extra bit needs to be allocated in the transmission scheme. Typically the globalleading sign is transmitted before the sequence of half pyramid indices Itx. However,it may also be transmitted after the last element of Itx.When applying the tree splitting, information about the leading sign of each sub- sequence needs to be available in the header and if required propagated through-out the tree for the sequence to be reconstructable at the decoder. One way of ensuring a fixed positive sign for any sequence in the half pyramid coding steps is to invert the signs of any sequence with a negative leading sign. This inversion is then indicated by a negative-signed ^-value in the corresponding header sequence. If the negative ^-value appears at the beginning of a header sequence, the sequence is then again inverted, and the negative leading sign is propagated up one level of the tree. Before the complete encoding procedure, the global leading sign of the full source-sequence is checked, and the source-sequence is inverted if the leadingsign is negative. The overall leading sign ^^ is then indicated by a separate initialtransmission bit. With this procedure, the leading ^ in the final (top-level) headersequence is ensured to be positive. Subsequent ^ values and ^-values for lower-levelheaders therefore can become negative to indicate negative leading signs in the sub- sequences without loss of the reduced half-pyramid search space.Fig. 6(a) shows an example of a tree 600a with the sign propagation for thepreviously introduced exemplary coding tree. Here, Y is the input vector. It is shownwhere the tree-traversing procedure is required to make the sub vector into a true half-pyramid with a positive leading sign, to become possible to enumerate within thebit-rate allocated to a half pyramid node. For every sequence (both Header and Leaf),if there is a negative leading sign, the sequence is inverted (and thus converted tohalf-pyramid with a known positive leading sign), and the ^-value passed to theparent node is assigned a negative sign. Evidently, the resulting sub-sequences can all be encoded with HPVC in a series of short codewords. For large values of ^^,^^^, that is, if the number of sequence splits at every level is high, there is a chance that the constructed header sequences cannot be encoded bythe HPVC engine under the given ^^^^^^ constraint. It is then possible to apply thesame splitting procedure used for the source-sequence to the header-sequence to allow for coding under the short codeword bit-size constraint. For this, the decoder needs to be able to compute the same rule for the received sequence to reconstructthe Header. The decision to split the header should therefore be made with onlyinformation about the ^ and ^^ values of any node. Upon receiving the (previouslysent) values of ^ and ^, the decoder can compute ^^^, and get the same variables to exactly reconstruct the encoder’s decision.Fig. 6(b) shows an example of a tree 600b with preferred encoder analysis traversalorder explicitly indicated for the introduced exemplary coding tree. The encoder willbuild the tree in the breadth-first direction starting in the node using the leftmost samples [1,-2, 6] in vector y, and proceed in the order indicate by the widearrows, and prepare indices Iana,1 to Iana,8 for encoding or actually encode indices Iana,1to Iana,8, and finally end up in the top node K1. Here in Fig 6(c), every node’s final sub-vector below each node has been converted to its half pyramid representation, with apositive leading sign, while the top vector y above the tree shows the incoming half-pyramid vectors with its original signs and the segmentation. It should further benoted that if the content of vector y in Fig 6(b) is to be decoded and synthesized fromleft to right rather than the example right-to-left decoding, the encoder may simply flip the vector y before encoding.Fig. 6(c) shows an example of a tree 600c with the corresponding preferred decodertraversal order indicated for the introduced exemplary coding tree. The decoder willstart in the depth-first direction from the node K1with index Irx,1, and proceed the fractional demultiplexing in the depth-first order as indicated by the wide arrows inFig.6(c), and end up in the end node K . At the receive side, node K1 index Irx,1arrives first and is processed first, and the tree is reconstructed in the order of thewide arrows. However, in an implementation optimized for parallel processing, theexample tree might still be received in the order as indicated by the wide arrows, butthe decoder may also in an initial sequence de-enumerate indices Irx,1, Irx,2, Irx,6corresponding to nodes K1, K1,3, K1,1. This will establish the width and L1-norms ofall the remaining tree nodes, (which are all leaves with a size less than 2BITLIM ). Thatis, during the decoder tree traversal of indices Irx,1 to Irx,8 the decoder may, based onthe local subvector length and the node local L1-norm, determine if an index is a leafand save that index for a subsequent parallel processing de-numeration of all leaves. In case the source sequences are not independent and identically distributed, but follow some kind of pre-determined distribution, an adapted segmentation strategy for the initial source sequence segmentation can lead to decreased overhead caused by Header sequences. Depending on the present structure of the source sequence, thesegmentation method Mt should be chosen in order minimize the overall tree- bit-rate and to decrease the number of overall splits necessary for arriving at valid^^^^^^-bit codewords (that is, reduce tree depth). Consider, e.g., a scenario inwhich the source sequence consists of MDCT coefficients (or a similar frequency representation). Then, if the energy of the signal is concentrated in a specific frequency band or area, it is beneficial to choose the initial segmentation rule such that the band in question is assigned a relatively shorter segment, while theremaining low energy regions are using larger, or wider, segments. Since theseremaining regions will carry only a small part of the signal’s energy (L2-norm, andcorrespondingly a lower L1-norm), they will be codable with a relatively shallow sub-tree, therefore decreasing the necessary Header overhead. Essentially hitting a shortsegment with a large L1-norm, will reduce the enumeration cost for positions in theHPVC-tree. In Fig. 7 is provided a visualization of segmentations for uniformdistribution 700a, logarithmic distribution 700b, reversed logarithmic distribution 700c, and bell-shaped distribution 700d, respectively.The block diagram of Fig. 8 shows how an input vector y of length Lin is split intosub-vectors and encoded. First the L1 norm of the vector is calculated to get Kin. Theinformation of Lin and Kin is then used to determine if the input vector can beencoded into an index or if it needs to be split. The encoder can handle vectors where L<=64 and K<=32767 that can be encoded using an index with at most 32 bits, i.e., B <= 32, without being split. For vector lengths outside these ranges segmentation and segmentation rules are used to determine the number of segments and the length of each segment. The number of splits needed is determined from L and K through a lookup table of limiting Klim values for each L and Ns combination. The input vector (a full pyramid) is initially split into an initial leading sign and ahalf-pyramid, the half-pyramid is then split first into six sub-vectors. The header isfurther split into two sub-sub-vectors to fit the allowed index length. The created tree is then traversed to collect how the tree was split with regards to information of leading signs, number of splits, and L and K for each node (header, sub-header and leaf). This tree information is then feed to the HPVC indexing core, that for each node establishes a unique index to encode the remaining information in that sub-vector. This can to some extent be parallelized for two or more sub-vectors. The indices are ordered and transmitted with headers first and leaves last, such that the structure and vectors can be transmitted and gradually be restored by the decoder.To utilize that the code space of the generated indices might have a fractional bit size,(i.e., the size NHPVC(L,K) of a node is not exactly corresponding to an integer power oftwo and thus log2(NHPVC(L,K)) is not an integer), the encoded information is feed through an arithmetic / range coder before the bitstream is sent through the channel to the decoder, when used with a uniform or trained CDF, the arithmetic / rangeencoder will act as a fractional bit multiplexing unit on the encoder side and as afractional de-multiplexing unit on the decoder side. If an arithmetic / range encoder isnot available, for example due to complexity considerations, the generated nodeindices may be sent as integer codewords of the bit size ceil(log2(NHPVC(L,K)) where“ceil” is the mathematical ceiling operation and log2 is the logarithm to the base 2.The tree is described using data structures for the header and leaves where the header keeps track of the leading signs, lengths L, L1norms K and number of splits, and the sub-vector elements which can be half-pyramid headers or leaves. In text notation this could be represented as S(L,K) which in turn is initially separated in to a leadingsign LS and a half-pyramid H (L,K) as the top node. A split half-pyramid is denotedHPVC (L, K, Sr, Ns, <Ns nodes>) for the half-pyramid headers and each leaf Lf(L,K,x) only keep track of the L and the K, and its sub-vector. The text representation of aleaf could thus be Lf(L, K, x), and the text representation of a header could be Hdr(L,K, x). The leading signs of sub-nodes are handled such that they are positioned in theheader, thus maintaining the remainders as true half-pyramid leaves. In the decoder the information is arriving in depth-first order so that the header information (withL1 norms and sub-heading signs) is arriving prior to its corresponding sub-leafdecoding. On the encoding side it can be favorable to create the tree breadth first, as then the list of information grows from left to right which makes memory handlingsignificantly faster and more practical, as no dynamic resizing of the node datastructures is needed.For the encoding this order (i.e., breadth first) can preferably be used to reduce theamount of dynamically created information that needs to be handled. This structure can be used when the half pyramid sub-vectors to encode and also when the encoded indices are generated by the encoder.The block diagram of Fig. 9 shows how iterations are run over all possible Nr*Ntconfigurations, starting with the same initial target y for every iteration. However, the described HPVC tree configuration analysis could also be made using less complexity at some limited performance decrease, for example by first determining the best segmentation method tbest, and then in a sequential step evaluate the best segmentation rule rbestamong the Nravailable splitting rules, for the now locked segmentation method t = tbest. The figure illustrates that it is possible to evaluate several different segmentation methods and segmentation rules, that will result in different tree structures and then select to use only the combination that results in the lowest bit-rate estimate before one performs the enumeration and indexing of that particular tree.Aspects of low complexity encoder side bit-rate estimation will be disclosed next.For evaluating the best tree-configuration of several possible tree configurations, theencoder tree-traversing can be run for every available split rule r and every availablesegmentation method t using the current input signal yh, but not run the node(header or leaf) enumeration steps. If r∈ [0 … rmax-1], and t∈ [0 … tmax-1], then themaximum number of tree configuration is cfgmax ^ rmax*tmax and cfg ∈ [0, [cfgmax-1], and every configuration should be evaluated in an efficient manner.Instead of running a full tree encoding (including enumeration) for each of thepossible cfgMAX options, the length Lsub and L1-norms Ksub for each segmentedvector may be store for every configuration and later be used to compute a bit-rate foreach node. Then these node costs may be accumulated per configuration to find thebit consumption of the total HVPC-tree HVPCcfg.Essentially, with the availability of a bit-rate for every node in a tree candidateconfiguration cfg the encoder can sum up the estimated bit costs for all nodes in each hypothetical HPVCtree and decide for the best and most efficient combinations of segmentation method(s) and split rule(s), without activating the enumeration procedures. When the best tree HVPCcfg,bestamong all cfgmaxcandidates have been established, the enumeration of each node in that specific list of indices can commence. Thisapproach reduces the closed loop analysis by a factor of close to cfgMAX as theenumeration is the heaviest part of the complete HVPC-tree encoding.The total bit-rate of a tree configuration HVPCcfg with Nidx nodes in the index finallist ^, and with stored subsegment L-values in Lnodes (idx), and subsegment L1-norms in Knodes(idx) can be expressed as: The best configuration is then determined by: ^^^^^^^ ← argmin^^^^^^^ ^^^ ∈ [^… ^^^^^^^^]The HPVC enumeration index size, ^^^^^ (^, ^), above corresponds to the totalnumber of points in a half pyramid of dimension L (and with a L1-norm of K) can be calculated using conditional product code enumeration as where d corresponds to the number of non-zero elements in a vector. In an L lengthvector this can be achieved in ^(^, ^) ways, as where ^(^ − 1, ^ − 1), represents the number of ways ^ non-zero vectors can beachieved using K unit pulses, as Finally, the last factor 2^ represents the number combinations due to signs of ^ non-zero vectors in a full pyramid. The outermost division by 2 reduces the so farcalculated full pyramid cardinality to the desired ^^(^,^) cardinality of the halfpyramid H(L,K).Aspects of different bit-rate estimation approaches will be disclosed next.Each node (header or leaf) may have an NH(L,K) cardinality up to 2BITLIM, and L-values up to Lmax and K values up to Ktree-max, resulting in a very large number ofcombination that might be too many to put in a single table, without using a hugememory space. In the bit-rate estimation equations below, let H(L,K) to denote a halfpyramid defined by L and K, and the cardinality NH(L,K) is used as an equivalentnotation for the NHPVC(L,K) cardinality, to increase the readability of the followingequations.One way of estimating the bit-rate is to let the decoder enumeration routine computeNH(L,K) using combinatorial logic as indicated in the exact ^^^^^(^, ^) = equation above, and then use that actual node bit-rate as e.g. bitsnode =ceil(512∙log2(NH(L,K) )) / 512, this results in a bit-rate quantized into an integer with aprecision of 9 fractional bits (512 equals 29). It is also possible to use single or doubleprecision floating point representations of the bits required for each node, e.g. bitsnode= log2(NH(L,K)). The exact equation approach requires costly combinatorialcalculations and the application of a base-2 logarithm for every node (leaf or header). The number of nodes in each tree may be up to roughly 1.5 times LTREE_max, so in theencoder tree structure analysis it is beneficial if a low complexity node bit-rateestimate may be used instead of the actual exact bit-rate for each node.For part of the combinations ^^(^,^)it is possible to find a closed form expression tocalculate an approximation of the required bit-rate, at least accurate enough toevaluate if the combination of ^ and ^ would fit in the allowed bit-budget. While thisapproximation is not accurate enough for all ^ and ^ combinations it can be used incombination with a small supporting bit-rate lookup table in which theapproximation accuracy is accurate enough.For values along the ^ = 2 axis there is a closed form expression , for ^ = 1 … 32767.Also along the ^ = 1 and ^ = 2 axis there is a closed form expression: When 3≤ ^ ≤ 15 and 3 ≤ ^ ≤ 14 the proposed approximation might not be goodenough and a table lookup from precalculated values of ^^(^,^)can be used.When 3 ≤ ^ ≤ 14 and ^ ≥ 16 the closed form approximation in this K-directionequation is a low complex expression of the form Similarly, when 3≤ ^ ≤ 15 and ^ ≥ 15, the L-direction equation is^^(^,^) = (^ − 1) log^(^) + ^^In the above equations ^^ and ^^ are constants defined in Table 3 and Table 4.L ^^3 +1.000184 +0.415765 -0.5831466 -1.903267 -3.48558 -5.289059 -7.2839910 -9.4474111 -11.761212 -14.210713 -16.7838Table 3: Example values of ^^, as used in the L-direction equation K^^3 -0.5847864 -1.584265 -2.905136 -4.488337 -6.293058 -8.289379 -10.454410 -12.7711 -15.221612 -17.7969Table 4: Example values of ^^ as used in the K-direction equationAspects of handling of very large input vectors or very large bit depths will bedisclosed next.The block diagram of Fig. 10 shows a method that can be used to handle vectors thatare much longer and / or have an L1-norm larger than the HPVC-core can handle. This is handled with the initial step LSB-reduce, where the long vectors are divided into a number of sub-vectors of the size that HPVC can handle. For sub-vectors with a too large L1-norm the LSB-reduce step allows for separate encoding of LSB in successivesteps and for each step the L1-norm is thereby reduced by a factor 2. This can berepeated until the L1-norm is reduced to a level that can be handled by the HPVC- core. The sub-vector division and LSB-reduce information need to be encoded and sent to the decoder so that it can restore the division into sub-vectors and LSBstreams from each reduction step. The LSB_reduce functionality is handled by theEncode Nlev_lsb, Encode Ktree and Encode_LSB reduced signs. The bit-rateestimate for the Ytree is then compared to the limit BITLIM to determine if theremaining tree (in Ytree) needs to be split and handled using the previously described HPVC-tree algorithm or if a single step (no tree) PVQ encoding method, that only uses short code words right away, can be used without any tree recursion. The data from the HPVC and the LSBs and number of reduction steps needs to be packed into the bitstream. This can be done so that the target vector can be restored in thedecoder. If this is done correctly it can be done in a way that even if the actual LSBsare lost (or corrupted or stripped from the data stream) the closest approximation to the input vector can be restored. By transmitting the NLEV-LSBprior to the HPVC-tree information, and the actual LSB-bits after the tree information, it is possible to handle loss of LSBs in transmission and also introduce a form of lossy compression using this specific multiplexing of LSB-reduce and HPVQ tree information. With theavailability of the number of sliced bit-planes, NLEV-LSB the HPVC-tree informationcan be reconstructed at the correct dynamic levels, even without the actual LSB-bits.A summary of at least some of the herein disclosed embodiments will be providednext. According to at least some of the herein disclosed embodiments, an integer vector y of length L, with L1-norm of K is split into more than two parts using a noiselessly coded header. According to at least some of the herein disclosed embodiments, the integer vector y is a signed vector. According to at least some of the herein disclosed embodiments, polarity information (sign) is moved from the target sub-vectors to the split-headers during the encoder side noiseless split. According to at least some of the herein disclosed embodiments the header codeword is noiselessly encoded the sameway / fashion as the leaves. Thus, the headers may correspond to informationotherwise requiring a long codeword. This applies to both signed and non-signed vectors. According to at least some of the herein disclosed embodiments, the initial top level non-binary split / segmentation, is adaptively conditioned to be quasi-linear or quasi-logarithmic (or reversed quasi-logarithmic) or centered, (any deterministicsegmentation method for length L and Ns number of splits. According to at leastsome of the herein disclosed embodiments, subsets of the headers and leaves arecomposed into a tree and then enumerated in a parallel fashion. According to at leastsome of the herein disclosed embodiments, the best tree structure (in a bit-ratesense) is evaluated using low complexity approximations of the node sizes. This allows efficient closed loop optimizations of segmentation method and applied split- rule. According to at least some of the herein disclosed embodiments, for the decoder, subsets of the leaves are decomposed / de-enumerated in a parallel fashion. The headers are decomposed / de-enumerated in serial fashion, while the leaves can beparallelized. According to at least some of the herein disclosed embodiments, for thedecoder, truncated bitstreams with a sequence of short codewords may be decoded up to the last received short codeword index.Fig. 11 schematically illustrates, in terms of a number of functional units, thecomponents of an audio encoder / decoder 200a, 200b according to an embodiment.Processing circuitry 210 is provided using any combination of one or more of a suitable central processing unit (CPU), multiprocessor, microcontroller, digital signal processor (DSP), etc., capable of executing software instructions stored in a computer program product 1210a (as in Fig.12), e.g. in the form of a storage medium 230. The processing circuitry 210 may further be provided as at least one application specific integrated circuit (ASIC), or field programmable gate array (FPGA). Particularly, the processing circuitry 210 is configured to cause the audioencoder / decoder 200a, 200b to perform a set of operations, or steps, as disclosedabove. For example, the storage medium 230 may store the set of operations, and the processing circuitry 210 may be configured to retrieve the set of operations from thestorage medium 230 to cause the audio encoder / decoder 200a, 200b to perform theset of operations. The set of operations may be provided as a set of executableinstructions. Thus, the processing circuitry 210 is thereby arranged to executemethods as herein disclosed. The storage medium 230 may also comprise persistent storage, which, for example, can be any single one or combination of magnetic memory, optical memory, solid state memory or even remotely mounted memory.The audio encoder / decoder 200a, 200b may further comprise a communications(comm.) interface 220 for communications with other entities, functions, nodes, anddevices, such as another audio encoder / decoder 200a, 200b. As such thecommunications interface 220 may comprise one or more transmitters and receivers, comprising analogue and digital components. The processing circuitry 210 controls the general operation of the audioencoder / decoder 200a, 200b e.g. by sending data and control signals to thecommunications interface 220 and the storage medium 230, by receiving data andreports from the communications interface 220, and by retrieving data and instructions from the storage medium 230. Other components, as well as the relatedfunctionality, of the audio encoder / decoder 200a, 200b are omitted in order not toobscure the concepts presented herein.The audio encoder / decoder 200a, 200b may be provided as a standalone device or asa part of at least one further device. A first portion of the instructions performed bythe audio encoder / decoder 200a, 200b (such as the audio encoder 200a or partthereof) may be executed in a first device, and a second portion of the instructionsperformed by the audio encoder / decoder 200a, 200b (such as the audio decoder200b or part thereof) may be executed in a second device; the herein disclosed embodiments are not limited to any particular number of devices on which theinstructions performed by the audio encoder / decoder 200a, 200b may be executed.Hence, the methods according to the herein disclosed embodiments are suitable to beperformed by an audio encoder / decoder 200a, 200b residing in a cloudcomputational environment. Therefore, although a single processing circuitry 210 isillustrated in Fig. 11 the processing circuitry 210 may be distributed among a pluralityof devices, or nodes. The same applies to the computer programs 1220a, 1220b of Fig. 12.Fig. 12 shows one example of a computer program product 1210a, 1210b comprisingcomputer readable means 1230. On this computer readable means 1230, a computer program 1220a can be stored, which computer program 1220a can cause the processing circuitry 210 and thereto operatively coupled entities and devices, such as the communications interface 220 and the storage medium 230, to execute methods according to embodiments described herein. The computer program 1220a and / or computer program product 1210a may thus provide means for performing any stepsof the audio encoder 200a as herein disclosed. On this computer readable means1230, a computer program 1220b can be stored, which computer program 1220b can cause the processing circuitry 310 and thereto operatively coupled entities and devices, such as the communications interface 320 and the storage medium 330, to execute methods according to embodiments described herein. The computer program 1220b and / or computer program product 1210b may thus provide means forperforming any steps of the audio decoder 200b as herein disclosed.In the example of Fig.12, the computer program product 1210a, 1210b is illustrated as an optical disc, such as a CD (compact disc) or a DVD (digital versatile disc) or a Blu-Ray disc. The computer program product 1210a, 1210b could also be embodied as a memory, such as a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or an electrically erasable programmable read-only memory (EEPROM) and more particularly as a non-volatile storage medium of a device in an external memory such as a USB (Universal Serial Bus) memory or a Flash memory, such as a compact Flash memory. Thus, while the computer program 1220a, 1220b is here schematically shown as a track on the depicted optical disk, the computer program 1220a, 1220b can be stored in any way which is suitable for the computer program product 1210a, 1210b. The inventive concept has mainly been described above with reference to a few embodiments. However, as is readily appreciated by a person skilled in the art, other embodiments than the ones disclosed above are equally possible within the scope of the inventive concept, as defined by the appended patent claims.
Claims
CLAIMS1. A method for encoding a part of an audio signal, the method being performed byan audio encoder (200a), the method comprising: receiving (S102) an input vector comprising integer coefficients derived fromsaid part of the audio signal; building (S104) an encoder tree by noiselessly splitting the input vector into segments, wherein each of the segments is represented by a respective node in theencoder tree, wherein each of the nodes being a header to another node comprises avector of sign information of segments represented by nodes being leaves of othernodes, and wherein the segments represented by nodes being leaves have a sign set toa fixed sign value; and encoding (S106) said part of the audio signal by enumerating the segments into a sequence of indices, with a unique index per segment.
2. The method according to claim 1, wherein the vector of sign information isconstructed by two or more L1 norms of the segments being moved from the segments into the vector of sign information.
3. The method according to claim 1, wherein the vector of sign information isconstructed by leading signs of the segments being shifted to the vector of signinformation from the segments.
4. The method according to claim 3, wherein one of the nodes is a top header, andwherein the leading signs of all the segments with a non-zero L1-norm are shifted to the vector of sign information in the top header.
5. The method according to any preceding claim, wherein how many nodes theencoder tree comprises is determined based on the received signed integer inputvector, an approximation of a bit-rate consumption of each segment, and a total bit-rate consumption for the received input vector.
6. The method according to claim 5, wherein one of the nodes is a top header withrespect to all other nodes, and wherein the top header is split into at least two sub-headers when the top header has a bit-rate consumption that exceeds a bit-rateconsumption limit.
7. The method according to claim 5, wherein at least some of the nodes representsub-segments resulting from at least one of the segments having been recursivelysplit into sub-segments, where the recursively split is applied until a terminationcriterion is met.
8. The method according to claim 7, wherein the nodes of the encoder tree arebuilt in parallel.
9. The method according to claim 7 or 8, wherein the termination criterion is metwhen the bit-rate consumption of each segment no longer exceeds a bit-rateconsumption limit.
10. The method according to any of claims 7 to 9, wherein the approximation of thebit-rate consumption of a given segment is evaluated when determining whether ornot to split said given segment using a dynamic selection of either L-direction equations per branch or K-direction equations per branch.
11. The method according to claim 5, wherein the encoder tree is built according toa splitting rule, wherein the splitting rule is selected from a set of available splittingrules where the splitting rule yielding lowest total bit-rate consumption is selected forbuilding the encoder tree.
12. The method according to claim 5, wherein the signed integer input vectorcomprises data points, and wherein how the data points are distributed over the segments when noiselessly splitting the signed integer input vector into the segmentsis defined by at least one segmentation rule.
13. The method according to claim 12, wherein the at least one segmentation rule isselected from a set of segmentation rules where a first segmentation rule defines a uniform distribution, a second segmentation rule defines a logarithmic distribution, a third segmentation rule defines a reversed logarithmic distribution, and a fourth segmentation rule defines a bell-shaped distribution.
14. The method according to a combination of claim 11 and claim 12 or 13, whereinthe at least one segmentation rule is selected according to the splitting rule.
15. The method according to claim 1, wherein encoding said part of the audio signalfurther comprises: sorting (S106-2) the indices into an order of indices in the sequence of indices,wherein indices of segments representing headers in the encoder tree appear before indices of segments representing leaves in the encoder tree.
16. The method according to claim 1, wherein encoding said part of the audio signalfurther comprises: applying (S106-4) arithmetic encoding or range encoding to the enumerated segments.
17. The method according to claim 1, wherein the method further comprises:providing (S108) the sequence of indices towards an audio decoder (200b) or toa data storage (110).
18. A method for decoding a part of an audio signal, the method being performed byan audio decoder (200b), the method comprising: receiving (S202) a bitstream, wherein the bitstream is composed of a sequenceof indices, with a unique index per segment, and represents said part of the audiosignal; extracting (S204) the indices from the bitstream; building (S206) a decoder tree from the indices, wherein there is one unique index per segment, wherein each of the segments is represented by a respective node in the decoder tree, wherein each of the nodes being a header to another node comprises a vector of sign information of segments represented by nodes being leaves of other nodes, and wherein the segments represented by nodes being leaves have a sign set to a fixed sign value; andrecreating (S208) an input vector comprising integer coefficients by moving the sign information from the vector of sign information to the segments and then concatenating all the segment in an order given by the decoder tree, wherein theinput vector represents the decoded part of said audio signal.
19. The method according to claim 18, wherein extracting the indices from thebitstream further comprises: applying (S204-2) arithmetic decoding or range decoding to the bitstream.
20. The method according to claim 18, wherein the nodes of the decoder tree arebuilt in parallel.
21. An audio encoder (200a) for encoding a part of an audio signal, the audioencoder (200a) comprising processing circuitry (210), the processing circuitry beingconfigured to cause the audio encoder (200a) to:receive an input vector comprising integer coefficients derived from said part of the audio signal; build an encoder tree by noiselessly splitting the input vector into segments, wherein each of the segments is represented by a respective node in the encoder tree, wherein each of the nodes being a header to another node comprises a vector of signinformation of segments represented by nodes being leaves of other nodes, andwherein the segments represented by nodes being leaves have a sign set to a fixed sign value; and encode said part of the audio signal by enumerating the segments into a sequence of indices, with a unique index per segment.
22. The audio encoder (200a) according to claim 21, further being configured toperform the method according to any of claims 2 to 17.
23. An audio decoder (200b) for decoding a part of an audio signal, the audiodecoder (200b) comprising processing circuitry (310), the processing circuitry beingconfigured to cause the audio decoder (200b) to:receive a bitstream, wherein the bitstream is composed of a sequence of indices, with a unique index per segment, and represents said part of the audio signal; extract the indices from the bitstream; build a decoder tree from the indices, wherein there is one unique index persegment, wherein each of the segments is represented by a respective node in the decoder tree, wherein each of the nodes being a header to another node comprises a vector of sign information of segments represented by nodes being leaves of other nodes, and wherein the segments represented by nodes being leaves have a sign set to a fixed sign value; and recreate an input vector comprising integer coefficients by moving the sign information from the vector of sign information to the segments and then concatenating all the segment in an order given by the decoder tree, wherein theinput vector represents the decoded part of said audio signal.
24. The audio decoder (200b) according to claim 23, further being configured toperform the method according to any of claims 19 to 20.
25. A computer program (1220a) for encoding a part of an audio signal, thecomputer program comprising computer code which, when run on processingcircuitry (210) of an audio encoder (200a), causes the audio encoder (200a) to:receive (S102) an input vector comprising integer coefficients derived from said part of the audio signal; build (S104) an encoder tree by noiselessly splitting the input vector intosegments, wherein each of the segments is represented by a respective node in the encoder tree, wherein each of the nodes being a header to another node comprises a vector of sign information of segments represented by nodes being leaves of other nodes, and wherein the segments represented by nodes being leaves have a sign set to a fixed sign value; and encode (S106) said part of the audio signal by enumerating the segments into a sequence of indices, with a unique index per segment.
26. A computer program (1220b) for decoding a part of an audio signal, thecomputer program comprising computer code which, when run on processingcircuitry (310) of an audio decoder (200b), causes the audio decoder (200b) to:receive (S202) a bitstream, wherein the bitstream is composed of a sequence of indices, with a unique index per segment, and represents said part of the audio signal; extract (S204) the indices from the bitstream; build (S206) a decoder tree from the indices, wherein there is one unique index per segment, wherein each of the segments is represented by a respective node in the decoder tree, wherein each of the nodes being a header to another node comprises a vector of sign information of segments represented by nodes being leaves of other nodes, and wherein the segments represented by nodes being leaves have a sign set to a fixed sign value; and recreate (S208) an input vector comprising integer coefficients by moving the sign information from the vector of sign information to the segments and then concatenating all the segment in an order given by the decoder tree, wherein theinput vector represents the decoded part of said audio signal.
27. A computer program product (1210a, 1210b) comprising a computer program(1220a, 1220b) according to at least one of claims 25 and 26, and a computer readable storage medium (1230) on which the computer program is stored.
Citation Information
Patent Citations
Method and apparatus for pyramid vector quantization indexing and de-indexing of audio / video sample vectors
US20160088297A1