Encoding of spatial audio direction parameters
Through flexible boundary codebook and entropy coding technology, the directional parameters of spatial audio signals are multi-level quantized and encoded, which solves the problem of low efficiency of spatial metadata compression at low bit rates and realizes efficient audio signal encoding and decoding.
Patent Information
- Application Number
- CN202080092858.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-13
- Filing Date
- 2020-12-07
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2040-12-07
AI Technical Summary
When encoding spatial audio signals, existing technologies have difficulty in effectively compressing and transmitting spatial metadata captured by microphone arrays, especially under low bit rate conditions, resulting in poor codec performance.
Flexible boundary codebook and entropy coding technology are used to quantize and encode the direction parameter values. Multi-level quantization and entropy coding are used to optimize bit rate utilization. The number of coding bits is determined in combination with the energy ratio value to achieve flexible encoding and decoding.
It improves the coding efficiency and decoding quality under low bit rate conditions, ensures the effective reconstruction of spatial audio signals, and is suitable for audio signal processing in various input formats.
Smart Images

Figure CN114945982B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to an apparatus and method for encoding parameters related to a sound field, but is not limited to encoding parameters related to the time-frequency domain of an audio encoder and decoder. Background Art
[0002] Parametric spatial audio processing is an area of audio signal processing that uses a set of parameters to describe the spatial aspects of sound. For example, when performing parametric spatial audio capture from a microphone array, estimating a set of directional metadata parameters from the microphone array signal is a typical and effective choice. This set of parameters, such as the direction of the sound in a frequency band and the ratio of the directional to non-directional portion of the captured sound in the frequency band, is a typical and effective choice. It is well known that these parameters describe the perceived spatial characteristics of the captured sound at the position of the microphone array. These parameters can be used accordingly in the synthesis of spatial sound for binaural headphones, speakers, or other formats such as panoramic surround sound (Ambisonics).
[0003] Therefore, directional metadata such as direction in frequency bands and direct-to-total energy ratio are particularly effective parameterizations for spatial audio capture.
[0004] A directional metadata parameter set comprising one or more directional values for each frequency band and an energy ratio parameter associated with each directional value may also be used as spatial metadata for an audio codec (which may also include other parameters such as spread coherence, number of directions, distance, etc.). The directional metadata parameter set may also include other parameters, or may be associated with other parameters that are considered non-directional (such as surround coherence, diffuse-to-total energy ratio, remainder-to-total energy ratio). For example, these parameters may be estimated from an audio signal captured by a microphone array, and, for example, a stereo signal may be generated from the microphone array signal to be transmitted together with the spatial metadata.
[0005] Since some codecs are expected to operate at a variety of bit rates, ranging from very low to relatively high bit rates, various strategies are needed to compress spatial metadata to optimize codec performance for each operating point. The raw bit rate of the coding parameters (metadata) is relatively high, so, especially at lower bit rates, it is expected that only the most important parts of the metadata can be transmitted from the encoder to the decoder.
[0006] The decoder may decode the audio signal into a PCM signal and process the sound in the frequency bands (using the spatial metadata) to obtain a spatial output, eg, a binaural output.
[0007] The aforementioned solution is particularly suitable for encoding spatial sound captured from a microphone array (e.g., a mobile phone, a video camera, a VR camera, a stand-alone microphone array). However, it may be desirable for such an encoder to have other input types besides the signal captured by the microphone array, such as loudspeaker signals, audio object signals, or ambisonic signals. Summary of the Invention
[0008] According to a first aspect, an apparatus is provided, comprising components configured to: obtain directional parameter values associated with at least two time-frequency portions of at least one audio signal; and encode the obtained directional parameter values based on a codebook, wherein the codebook comprises two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level comprises a first set of quantization values, and a second or subsequent quantization level comprises a second or another set of quantization values and the quantization values of the previous quantization level.
[0009] The component configured to encode the obtained directional parameter values based on a codebook can be further configured to: determine, relative to each of the obtained directional parameter values, a closest quantization value from a set of quantization values of the determined quantization level and a previous quantization level quantization value; and generate a codeword for each of the obtained directional parameter values based on the associated closest quantization value.
[0010] The component configured to encode the obtained directional parameter value based on a codebook may be further configured to entropy encode the codeword generated for the subband within the frame including the subband and the time block.
[0011] The component configured to encode the obtained direction parameter value based on the codebook can be further configured to iteratively: compare the number of bits required for entropy encoding the codeword of the direction parameter value for the subband within the frame including the subband and the time block based on the selected quantization level with the allocated number of bits; depending on the number of bits required for entropy encoding the codeword for the direction parameter value being greater than the allocated number of bits, select a previous quantization level and re-encode the obtained direction parameter value based on the previous quantization level until the number of bits required for entropy encoding the codeword for the direction parameter value is equal to or less than the allocated number of bits.
[0012] The component configured to encode the directional parameter value based on the codebook can be further configured to: sort the closest quantization values determined for the directional parameter values of the subbands within the frame including the subbands and the time block based on the determined angular quantization distortion; iteratively select one of the closest quantization values in the determined order, and: determine whether the selected closest quantization value is a member of the previous quantization level quantization values; and re-encode the obtained directional parameter value using the previous quantization level quantization value until the number of bits required for entropy encoding the codeword for the directional parameter value is equal to or less than the allocated number of bits.
[0013] The component configured to encode the obtained direction parameter value based on the codebook can be further configured to: encode the azimuth direction parameter value based on the codebook; and encode the elevation direction parameter value based on at least one average elevation direction parameter value of the subband within the frame including the subband and the time block.
[0014] The component may be further configured to determine the number of allocated bits for encoding each subband within a frame comprising the subband and the time block based on the value of the energy ratio value associated with the obtained directional parameter value.
[0015] According to a second aspect, an apparatus is provided, comprising components configured to: obtain at least one coded bitstream, the at least one coded bitstream comprising at least one codebook-encoded directional parameter value, wherein the codebook comprises two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level comprises a first set of quantization values, a second or subsequent quantization level comprises a second or another set of quantization values and a previous quantization level quantization value; and decode the at least one codebook-encoded directional parameter value.
[0016] According to a third aspect, a method is provided, comprising: obtaining directional parameter values associated with at least two time-frequency parts of at least one audio signal; and encoding the obtained directional parameter values based on a codebook, wherein the codebook comprises two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level comprises a first set of quantization values, and a second or subsequent quantization level comprises a second or another set of quantization values and the quantization values of the previous quantization level.
[0017] Encoding the obtained directional parameter values based on the codebook may further include: determining, relative to each of the obtained directional parameter values, a closest quantization value from a set of quantization values of the determined quantization level and a previous quantization level quantization value; and generating a codeword for each of the obtained directional parameter values based on the associated closest quantization value.
[0018] Encoding the obtained directional parameter value based on the codebook may further include entropy encoding a codeword generated for a subband within a frame including the subband and the time block.
[0019] Encoding the obtained direction parameter value based on the codebook may further include iteratively: comparing the number of bits required for entropy encoding the codeword of the direction parameter value for the subband within the frame including the subband and the time block based on the selected quantization level with the allocated number of bits; depending on the number of bits required for entropy encoding the codeword for the direction parameter value being greater than the allocated number of bits, selecting a previous quantization level and re-encoding the obtained direction parameter value based on the previous quantization level until the number of bits required for entropy encoding the codeword for the direction parameter value is equal to or less than the allocated number of bits.
[0020] Encoding the directional parameter value based on the codebook may further include: sorting the closest quantization values determined for the directional parameter values of the subbands within the frame including the subbands and the time block based on the determined angular quantization distortion; iteratively selecting one of the closest quantization values in the determined order, determining whether the selected closest quantization value is a member of the previous quantization level quantization values, and re-encoding the obtained directional parameter value using the previous quantization level quantization value until the number of bits required for entropy encoding the codeword for the directional parameter value is equal to or less than the allocated number of bits.
[0021] Encoding the obtained direction parameter value based on the codebook may further include: encoding the azimuth direction parameter value based on the codebook; and encoding the elevation direction parameter value based on at least one average elevation direction parameter value of the subband within the frame including the subband and the time block.
[0022] The method may further include determining an allocated number of bits for encoding each subband within a frame including the subband and the time block based on a value of the energy ratio value associated with the obtained directional parameter value.
[0023] According to a fourth aspect, a method is provided, comprising: obtaining at least one coded bitstream, the at least one coded bitstream comprising at least one codebook-encoded directional parameter value, wherein the codebook comprises two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level comprises a first set of quantization values, a second or subsequent quantization level comprises a second or another set of quantization values and a previous quantization level quantization value; and decoding the at least one codebook-encoded directional parameter value.
[0024] According to a fifth aspect, a device is provided, comprising at least one processor and at least one memory comprising computer program code, wherein the at least one memory and the computer program code are configured to, together with the at least one processor, enable the device to at least: obtain directional parameter values associated with at least two time-frequency parts of at least one audio signal; and encode the obtained directional parameter values based on a codebook, wherein the codebook comprises two or more quantization levels, and the two or more quantization levels are arranged such that a first quantization level comprises a first set of quantization values, and a second or subsequent quantization level comprises a second or another set of quantization values and the quantization values of the previous quantization level.
[0025] The device configured to encode the obtained directional parameter values based on a codebook may be further configured to: determine, relative to each of the obtained directional parameter values, a closest quantization value from a set of quantization values of the determined quantization level and a previous quantization level quantization value; and generate a codeword for each of the obtained directional parameter values based on the associated closest quantization value.
[0026] The apparatus, being caused to encode the obtained directional parameter value based on a codebook, may be further caused to entropy encode a codeword generated for a subband within a frame comprising the subband and the time block.
[0027] The device configured to encode the obtained directional parameter value based on a codebook may be further configured to iteratively: compare the number of bits required for entropy encoding a codeword for the directional parameter value of a subband within a frame including the subband and the time block based on the selected quantization level with the allocated number of bits; select a previous quantization level depending on whether the number of bits required for entropy encoding the codeword for the directional parameter value is greater than the allocated number of bits, and re-encode the obtained directional parameter value based on the previous quantization level until the number of bits required for entropy encoding the codeword for the directional parameter value is equal to or less than the allocated number of bits.
[0028] The device configured to encode the directional parameter value based on a codebook can be further configured to: sort the closest quantization values determined for the directional parameter values of the subbands within the frame including the subbands and the time blocks based on the determined angular quantization distortion; iteratively select one of the closest quantization values in the determined order, and: determine whether the selected closest quantization value is a member of the previous quantization level quantization values; and re-encode the obtained directional parameter value using the previous quantization level quantization value until the number of bits required for entropy encoding the codeword for the directional parameter value is equal to or less than the allocated number of bits.
[0029] The device that is caused to encode the obtained direction parameter value based on the codebook can be further caused to: encode the azimuth direction parameter value based on the codebook; and encode the elevation direction parameter value based on at least one average elevation direction parameter value of the subband within the frame including the subband and the time block.
[0030] The apparatus may be further caused to determine, based on a value of the energy ratio value associated with the obtained directional parameter value, an allocated number of bits for encoding each subband within a frame comprising the subband and the time block.
[0031] According to a sixth aspect, a device is provided, comprising at least one processor and at least one memory comprising computer program code, wherein the at least one memory and the computer program code are configured to, together with the at least one processor, enable the device to at least: obtain at least one encoded bit stream, the at least one encoded bit stream comprising at least one codebook-encoded directional parameter value, wherein the codebook comprises two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level comprises a first set of quantization values, a second or subsequent quantization level comprises a second or another set of quantization values and a previous quantization level quantization value; and decode at least one codebook-encoded directional parameter value.
[0032] According to a seventh aspect, a device is provided, comprising: a component for obtaining directional parameter values associated with at least two time-frequency parts of at least one audio signal; and a component for encoding the obtained directional parameter values based on a codebook, wherein the codebook comprises two or more quantization levels, and the two or more quantization levels are arranged so that a first quantization level comprises a first set of quantization values, and a second or subsequent quantization level comprises a second or another set of quantization values and the quantization values of the previous quantization level.
[0033] According to an eighth aspect, a device is provided, comprising: a component for obtaining at least one encoded bit stream, wherein the at least one encoded bit stream includes at least one codebook-encoded directional parameter value, the codebook including two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level includes a first set of quantization values, a second or subsequent quantization level includes a second or another set of quantization values and a previous quantization level quantization value; and a component for decoding the at least one codebook-encoded directional parameter value.
[0034] According to a ninth aspect, there is provided a computer program comprising instructions [or a computer-readable medium comprising program instructions], which instructions / program instructions are used to cause an apparatus to at least perform the following operations: obtain directional parameter values associated with at least two time-frequency parts of at least one audio signal; and encode the obtained directional parameter values based on a codebook, wherein the codebook comprises two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level comprises a first set of quantization values, and a second or subsequent quantization level comprises a second or another set of quantization values and the quantization values of the previous quantization level.
[0035] According to the tenth aspect, a computer program comprising instructions [or a computer-readable medium comprising program instructions] is provided, wherein the instructions / program instructions are used to cause an apparatus to at least perform the following operations: obtain at least one encoded bit stream, the at least one encoded bit stream comprising at least one codebook-encoded directional parameter value, wherein the codebook comprises two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level comprises a first set of quantization values, a second or subsequent quantization level comprises a second or another set of quantization values and a previous quantization level quantization value; and decode at least one codebook-encoded directional parameter value.
[0036] According to an eleventh aspect, a non-transitory computer-readable medium comprising program instructions is provided, wherein the program instructions are used to cause an apparatus to perform at least the following operations: obtain directional parameter values associated with at least two time-frequency parts of at least one audio signal; and encode the obtained directional parameter values based on a codebook, wherein the codebook comprises two or more quantization levels, and the two or more quantization levels are arranged such that a first quantization level comprises a first group of quantization values, and a second or subsequent quantization level comprises a second or another group of quantization values and the quantization values of the previous quantization level.
[0037] According to the twelfth aspect, a non-transitory computer-readable medium comprising program instructions is provided, which are used to cause an apparatus to perform at least the following operations: obtain at least one encoded bit stream, the at least one encoded bit stream comprising at least one codebook-encoded directional parameter value, wherein the codebook comprises two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level comprises a first set of quantization values, a second or subsequent quantization level comprises a second or another set of quantization values and a previous quantization level quantization value; and decode at least one codebook-encoded directional parameter value.
[0038] According to the thirteenth aspect, a device is provided, comprising: an obtaining circuit configured to obtain directional parameter values associated with at least two time-frequency parts of at least one audio signal; and an encoding circuit configured to encode the obtained directional parameter values based on a codebook, wherein the codebook comprises two or more quantization levels, and the two or more quantization levels are arranged such that a first quantization level comprises a first group of quantization values, and a second or subsequent quantization level comprises a second or another group of quantization values and the quantization values of the previous quantization level.
[0039] According to the fourteenth aspect, a device is provided, comprising: an acquisition circuit configured to obtain at least one encoded bit stream, the at least one encoded bit stream including at least one codebook-encoded directional parameter value, wherein the codebook includes two or more quantization levels, the two or more quantization levels are set so that a first quantization level includes a first group of quantization values, and a second or subsequent quantization level includes a second or another group of quantization values and a previous quantization level quantization value; and a decoding circuit configured to decode the at least one codebook-encoded directional parameter value.
[0040] According to the fifteenth aspect, a computer-readable medium comprising program instructions is provided, which are used to cause an apparatus to perform at least the following operations: obtain directional parameter values associated with at least two time-frequency parts of at least one audio signal; and encode the obtained directional parameter values based on a codebook, wherein the codebook comprises two or more quantization levels, and the two or more quantization levels are arranged such that a first quantization level comprises a first group of quantization values, and a second or subsequent quantization level comprises a second or another group of quantization values and the quantization values of the previous quantization level.
[0041] According to the sixteenth aspect, a computer-readable medium comprising program instructions is provided, which are used to cause an apparatus to perform at least the following operations: obtain at least one encoded bit stream, the at least one encoded bit stream comprising at least one codebook-encoded directional parameter value, wherein the codebook comprises two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level comprises a first set of quantization values, a second or subsequent quantization level comprises a second or another set of quantization values and a previous quantization level quantization value; and decode at least one codebook-encoded directional parameter value.
[0042] An apparatus comprises means for performing the actions of the method as described above.
[0043] An apparatus is configured to perform the actions of the method described above.
[0044] A computer program includes program instructions for causing a computer to execute the method described above.
[0045] A computer program product stored on a medium may cause an apparatus to perform the method described herein.
[0046] An electronic device may include an apparatus as described herein.
[0047] A chipset may include the apparatus as described herein.
[0048] The embodiments of the present application are intended to solve the problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings, in which:
[0050] Figure 1 Schematically illustrates an apparatus system suitable for implementing some embodiments;
[0051] Figure 2 schematically illustrates an encoder according to some embodiments;
[0052] Figure 3 Showing how according to some embodiments Figure 2 A flowchart of the operation of the encoder shown in ;
[0053] Figure 4 Schematically illustrating how Figure 2 The direction encoder shown in ;
[0054] Figure 5 Showing how according to some embodiments Figure 4 A flowchart of the operation of the direction encoder shown in;
[0055] Figure 6 Schematically illustrating how Figure 2 and Figure 4 The flexible boundary codebook encoder portion of the direction encoder shown in FIG;
[0056] Figure 7 Showing how according to some embodiments Figure 6 A flowchart of the operation of the flexible boundary codebook encoder portion shown in;
[0057] Figure 8 shows an example codebook angle allocation according to some embodiments;
[0058] Figure 9 Shown according to some embodiments for Figure 6 A further flow chart of the operation of the flexible boundary codebook selection of the encoder shown in;
[0059] Figure 10 An example device suitable for implementing the illustrated means is schematically illustrated. DETAILED DESCRIPTION
[0060] Suitable means and possible mechanisms for providing combined and encoding metadata parameters derived from spatial analysis are described in more detail below. In the following discussion, multi-channel systems are discussed with respect to multi-channel microphone implementations. However, as discussed above, the input format can be any suitable input format, such as multi-channel loudspeakers, Ambisonic (FOA / HOA), etc. It should be understood that in some embodiments, the channel positions are based on the positions of the microphones or virtual positions or directions.
[0061] Furthermore, in the following examples, the output of the example system is a multi-channel speaker arrangement. In other embodiments, the output may be rendered to the user via means other than speakers. The multi-channel speaker signal may also be summarized as two or more playback audio signals.
[0062] As discussed above, directional metadata associated with an audio signal may include multiple parameters per time-frequency tile (such as multiple directions, and a direct-to-total energy ratio, distance, etc. associated with each direction). Directional metadata may also include other parameters, or may be associated with other parameters that are considered non-directional but can be used to define characteristics of an audio scene when combined with directional parameters (such as surround coherence, diffuse-to-total energy ratio, residual-to-total energy ratio). For example, a reasonable design choice that can produce high-quality output is to determine that the directional metadata includes two directions for each time-frequency subframe (and a direct-to-total energy ratio, distance value, etc. associated with each direction). However, as also discussed above, bandwidth and / or storage limitations may require that the codec not send directional metadata parameter values for every frequency band and time subframe.
[0063] Current proposals include those disclosed in GB patent application 1811071.8, which has considered lossy compression of metadata, and PCT / FI2019 / 050675, which has discussed vector quantization methods when the number of bits available for a given subband is very small. Even with only a 9-bit codebook, the vector quantizer method increases the table ROM of the codec, with approximately 4 kB of memory being used for 4-dimensional codebooks of 2, 3, 4, ..., and 9 bits.
[0064] The concepts discussed in the following embodiments relate to the encoding of a spatial audio stream having a transmitted audio signal and (spatial) directional metadata, wherein apparatus and methods are described for implementing an embedded codebook with flexible boundaries rather than using a vector quantizer. To reduce memory consumption, in some embodiments the codebook is one-dimensional and the codewords are indexed so that those codewords constituting the 1-bit codebook appear first, followed by those codewords belonging to the 2-bit codebook but not to the 1-bit codebook, then those codewords belonging to the 3-bit codebook but not to the 1-bit and 2-bit codebooks, and so on. In such embodiments, one subband is encoded at a time. Based on the number of available bits for the current subband and the distance to the unquantized audio directional data, a codeword is selected for each time-frequency tile and its index is encoded using a suitable entropy coding method (e.g., Golomb Rice coding).
[0065] about Figure 1 , shows an example device and system for implementing embodiments of the present application. System 100 is shown as having an "analysis" portion 121 and a "synthesis" portion 131. The "analysis" portion 121 is the portion that receives the multi-channel signal and encodes the directional metadata and the transmission signal, while the "synthesis" portion 131 is the portion that decodes the encoded directional metadata and the transmission signal and presents the regenerated signal (e.g., in the form of a multi-channel speaker).
[0066] In the following description, the "analysis" section 121 is described as a series of sections, however, in some embodiments, the section may be implemented as a function within the same functional device or section. In other words, in some embodiments, the "analysis" section 121 is an encoder that includes at least one of a transmission signal generator or an analysis processor as described below.
[0067] The input to the system 100 and the "analysis" section 121 is the multi-channel signal 102. The "analysis" section 121 may include a transmission signal generator 103, an analysis processor 105, and an encoder 107. In the following example, microphone channel signal input is described; however, in other embodiments, any suitable input (or synthesized multi-channel) format may be implemented. In such embodiments, directional metadata associated with the audio signal may be provided to the encoder as a separate bitstream. The multi-channel signal is passed to the transmission signal generator 103 and the analysis processor 105.
[0068] In some embodiments, the transmission signal generator 103 is configured to receive a multi-channel signal and generate an appropriate audio signal format for encoding. The transmission signal generator 103 can, for example, generate a stereo or mono audio signal. The transmission audio signal generated by the transmission signal generator can be in any known format. For example, when the input is a mobile phone microphone array audio signal, the transmission signal generator 103 can be configured to select a left and right microphone pair and apply any appropriate processing to the audio signal pair, such as automatic gain control, microphone noise removal, wind noise removal, and equalization. In some embodiments, when the input is a first-order ambisonic / high-order ambisonic (FOA / HOA) signal, the transmission signal generator can be configured to formulate a directional beam signal oriented in the left and right directions, such as two opposing cardioid signals. In addition, in some embodiments, when the input is a speaker surround mix and / or object, the transmission signal generator 103 can be configured to generate a lower condensation signal that combines the left channel into a lower left condensation channel, the right channel into a lower right condensation channel, and adds the center channel to the two transmission channels with an appropriate gain.
[0069] In some embodiments, the transmit signal generator is bypassed (or, in other words, optional). For example, in some cases where analysis and synthesis occur in a single processing step on the same device, no transmit signal is generated without intermediate processing, and the input audio signal is passed through unprocessed. The number of transmit channels generated can be any suitable number, not necessarily one or two channels, for example.
[0070] The output of the transmission signal generator 103 may be passed to the encoder 107 .
[0071] In some embodiments, the analysis processor 105 is also configured to receive the multi-channel signals and analyze these signals to produce directional metadata 106 associated with the multi-channel signals and, therefore, the transmission signal 104 .
[0072] The analysis processor 105 may be configured to generate directional metadata parameters that may include at least one directional parameter 108 and at least one energy ratio parameter 110 for each time-frequency analysis interval (and in some embodiments, other parameters, a non-exhaustive list of which includes number of directions, surround coherence, diffuse to total energy ratio, residual to total energy ratio, extended coherence parameter, and distance parameter). The directional parameters may be represented in any suitable manner, for example, as azimuths. and the spherical coordinates of the elevation angle θ(k,n).
[0073] In some embodiments, the number of directional metadata parameters may vary between time-frequency tiles. Thus, for example, in frequency band X, all directional metadata parameters are obtained (generated) and transmitted, while in frequency band Y, only one of the directional metadata parameters is obtained and transmitted, and furthermore, in frequency band Z, no parameters are obtained or transmitted. A practical example of this could be that for some time-frequency tiles corresponding to the highest frequency band, some directional metadata parameters are not required for perceptual reasons. The directional metadata 106 can be passed to the encoder 107.
[0074] In some embodiments, analysis processor 105 is configured to apply a time-frequency transform to the input signal. Furthermore, for example, in a time-frequency plot, when the input is a mobile phone microphone array, the analysis processor can be configured to estimate delay values between pairs of microphones that maximize the correlation between the microphones. Furthermore, based on these delay values, the analysis processor can be configured to formulate corresponding direction values for directional metadata. Furthermore, the analysis processor can be configured to formulate a direct-to-total ratio parameter based on the correlation value.
[0075] In some embodiments, for example, if the input is a FOA signal, the analysis processor 105 can be configured to determine an intensity vector. Furthermore, the analysis processor can be configured to determine a direction parameter value for the directional metadata based on the intensity vector. Furthermore, a diffuse-to-total ratio can be determined, from which a direct-to-total ratio parameter value for the directional metadata can be determined. This analysis method is referred to in the literature as Directional Audio Coding (DirAC).
[0076] In some examples, for example, if the input is an HOA signal, the analysis processor 105 can be configured to divide the HOA signal into multiple sectors and apply the above method to each sector. This sector-based method is referred to in the literature as high-order DirAC (HO-DirAC). In these examples, there is more than one simultaneous directional parameter value per time-frequency tile (corresponding to multiple sectors).
[0077] Additionally, in some embodiments, if the input is a loudspeaker surround mix and / or an audio object-based signal, the analysis processor may be configured to convert the signal into a FOA / HOA signal format and obtain a direct-to-total ratio parameter value, as described above.
[0078] The encoder 107 may include an audio encoder core 109 configured to receive the transmission audio signals 104 and generate suitable encodings of these audio signals. In some embodiments, the encoder 107 may be a computer (running suitable software stored in memory and on at least one processor), or alternatively may be a specific device, such as an FPGA or ASIC. Any suitable approach may be used to implement audio encoding.
[0079] The encoder 107 may also include a directional metadata encoder / quantizer 111 configured to receive directional metadata and output an encoded or compressed form of the information. Figure 1 Prior to transmission or storage as indicated by the dashed line in FIG, further interleaving, multiplexing into a single data stream, or embedding directional metadata into the coded downmix signal may be performed using any suitable scheme.
[0080] In some embodiments, the transmission signal generator 103 and / or the analysis processor 105 may be located on a device separate (or otherwise decoupled) from the encoder 107. For example, in such an embodiment, the directional metadata (and associated non-directional metadata) associated with the audio signal may be provided to the encoder as a separate bitstream.
[0081] In some embodiments, the transmission signal generator 103 and / or the analysis processor 105 may be part of the encoder 107 , ie, may be located inside the encoder and on the same device.
[0082] In the following description, the "composite" portion 131 is described as a series of portions, however, in some embodiments, the portion may be implemented as a function within the same functional device or portion.
[0083] On the decoder side, the received or retrieved data (stream) may be received by a decoder / demultiplexer 133. The decoder / demultiplexer 133 may demultiplex the encoded stream and pass the audio encoded stream to a transmission signal decoder 135, which is configured to decode the audio signal to obtain a transmission audio signal. Similarly, the decoder / demultiplexer 133 may include a metadata extractor 137, which is configured to receive the encoded directional metadata (e.g., a directional index representing a directional parameter value) and generate directional metadata.
[0084] In some embodiments, the decoder / demultiplexer 133 may be a computer (running suitable software stored on memory and at least one processor), or alternatively may be a dedicated device using, for example, an FPGA or ASIC.
[0085] The decoded metadata and transport audio signal may be passed to a synthesis processor 139 .
[0086] The “Synthesis” portion 131 of the system 100 further illustrates a synthesis processor 139 configured to receive the transmitted audio signal and the directional metadata and, based on the transmitted signal and the directional metadata, recreate synthesized spatial audio in the form of the multi-channel signal 110 in any suitable format (which may be a multi-channel loudspeaker format, or in some embodiments any suitable output format such as a binaural or Ambisonics signal, depending on the use case).
[0087] Therefore, the synthesis processor 139 creates an output audio signal based on any suitable known method, such as a multi-channel speaker signal or a binaural signal. This will not be described in further detail here. However, as a simplified example, the speaker output can be rendered according to any of the following methods. For example, the transmitted audio signal can be divided into a directional stream and an ambient stream based on the direct to total energy ratio and the diffuse to total energy ratio. Furthermore, the directional stream can be rendered based on the directional parameter using amplitude shifting. In addition, the ambient stream can be rendered using decorrelation. Furthermore, the directional stream and the ambient stream can be combined.
[0088] The output signal may be reproduced using a multi-channel loudspeaker setup or headphones capable of head tracking.
[0089] It should be noted that Figure 1 The processing blocks can be located in the same or different processing entities. For example, in some embodiments, the microphone signal from the mobile device is processed by a spatial audio capture system (including an analysis processor and a transmission signal generator), and the resulting spatial metadata and transmission audio signal (e.g., in the form of a MASA stream) are forwarded to an encoder (e.g., an IVAS encoder) including an encoder. In other embodiments, the input signal (e.g., a 5.1-channel audio signal) is forwarded directly to an encoder (e.g., an IVAS encoder) including an analysis processor, a transmission signal generator, and an encoder.
[0090] In some embodiments, there may be two (or more) input audio signals, wherein the first audio signal is Figure 1 , the first audio signal is processed by the apparatus shown in (generating data as input to the encoder), while the second audio signal is directly forwarded to an encoder (e.g., an IVAS encoder) comprising an analysis processor, a transmission signal generator, and an encoder. Furthermore, the audio input signals can be encoded independently in the encoder, or they can be combined in the parameter domain, for example, according to a method that can be referred to as MASA mixing.
[0091] In some embodiments, there may be a composition section that includes separate decoder and composition processor entities or devices, or the composition section may include a single entity that includes both the decoder and composition processor. In some embodiments, the decoder block can process more than one input data stream in parallel. In applications, the term "composition processor" may be interpreted as an internal or external renderer.
[0092] Thus, in summary, first, the system (analysis portion) is configured to receive a multi-channel audio signal. Furthermore, the system (analysis portion) is configured to generate a suitable transmission audio signal (e.g., by selecting some of the audio signal channels). Furthermore, the system is configured to encode the transmission audio signal for storage / transmission. Thereafter, the system can store / send the encoded transmission audio signal and metadata. The system can retrieve / receive the encoded transmission audio signal and metadata. Furthermore, the system is configured to extract the transmission audio signal and metadata from the encoded transmission audio signal and metadata parameters, e.g., demultiplex and decode the encoded transmission audio signal and metadata parameters.
[0093] The system (synthesis part) is configured to synthesize an output multi-channel audio signal based on the extracted transmission audio signal and metadata.
[0094] about Figure 2 , describes in more detail an example analysis processor 105 and metadata encoder / quantizer 111 (e.g., Figure 1 ).
[0095] In some embodiments, the analysis processor 105 includes a time-frequency domain converter 201 .
[0096] In some embodiments, the time-frequency domain converter 201 is configured to receive the multi-channel signal 102 and apply a suitable time-frequency domain transform, such as a short-time Fourier transform (STFT), to convert the input time-domain signal into suitable time-frequency signals. These time-frequency signals can be passed to the spatial analyzer 203 and the signal analyzer 205.
[0097] Thus, for example, the time-frequency signal 202 may be represented in the time-frequency domain as:
[0098] s i (b,n)
[0099] Where b is the frequency bin index, n is the time-frequency block (frame) index, and i is the channel index. In another expression, n can be considered as a time index with a sampling rate lower than the sampling rate of the original time domain signal. These frequency bins can be grouped into multiple subbands, which group one or more bins into subbands with frequency band index k=0,...,K-1. Each subband k has a minimum bin b k,low and the highest position b k,high , and the subband contains k,low to b k,high The widths of the sub-bands can be approximated by any suitable distribution, for example, the Equivalent Rectangular Bandwidth (ERB) scale or the Bark scale.
[0100] In some embodiments, the analysis processor 105 comprises a spatial analyzer 203. The spatial analyzer 203 may be configured to receive the time-frequency signals 202 and, based on these signals, estimate the directional parameters 108. The directional parameters may be determined based on any audio-based determination of "direction."
[0101] For example, in some embodiments, spatial analyzer 203 is configured to estimate direction using two or more signal inputs. This represents the simplest configuration for estimating "direction," and more complex processing can be performed using even more signals.
[0102] Therefore, the spatial analyzer 203 may be configured to provide at least one azimuth and elevation angle for each frequency band and temporal time-frequency block within a frame of the audio signal, denoted as azimuth and the elevation angle θ(k,n). The direction parameter 108 may also be passed to the direction index generator 205.
[0103] The spatial analyzer 203 may also be configured to determine an energy ratio parameter 110. The energy ratio may be considered to determine the energy of the audio signal that may be considered to arrive from a direction. The direct-to-total energy ratio r(k,n) may be estimated, for example, using a stability metric of the directional estimate, or using any correlation metric, or any other suitable method for obtaining a ratio parameter. The energy ratio may be passed to an energy ratio encoder 207.
[0104] The spatial analyzer 203 may also be configured to determine a plurality of coherence parameters 112 , which may include surround coherence (γ(k,n)) and extended coherence (ζ(k,n)), both of which are analyzed in the time-frequency domain.
[0105] Thus, in summary, the analysis processor is configured to receive time domain multi-channel or other formats such as microphone or Ambisonic audio signals.
[0106] Thereafter, the analysis processor may apply a time-domain to frequency domain transform (eg, STFT) to generate a suitable time-frequency domain signal for analysis, and then apply directional analysis to determine direction and energy ratio parameters.
[0107] Furthermore, the analysis processor may be configured to output the determined parameters.
[0108] Although the direction, energy ratio, and coherence parameters are described for each time index n, in some embodiments, these parameters can be combined across several time indices. The same applies to the frequency axis; as already described, the direction of several frequency bins b can be expressed by a direction parameter in a frequency band k composed of several frequency bins b. The same applies to all spatial parameters discussed herein.
[0109] In some embodiments, 16 bits may be used to represent the directional data, such that each azimuth parameter is represented on approximately 9 bits and the elevation angle on 7 bits. In such an embodiment, the energy ratio parameter may be represented on 8 bits. For each frame, there may be N subbands (where N may be between 1 and 24 and may be fixed to 5) and M time-frequency (TF) blocks (where the value of M may be M=4). Thus, in this example, (16+8) x M x N bits are required to store the uncompressed directional and energy ratio metadata for each frame.
[0110] Also like Figure 2 As shown in , an example metadata encoder / quantizer 111 is shown according to some embodiments.
[0111] The metadata encoder / quantizer 111 may include a direction encoder 205. The direction encoder 205 is configured to receive direction parameters (such as azimuth and elevation angles θ(k,n) 108) (in some embodiments, also receiving an expected bit allocation), and generating a suitable encoded output therefrom. In some embodiments, the encoding is based on an arrangement of spheres (forming a spherical grid arranged in a ring on a "surface" sphere), which are defined by a lookup table defined by the determined quantization resolution. In other words, the spherical grid uses the idea of covering a sphere with smaller spheres, and considering the centers of the smaller spheres as points defining a grid of nearly equidistant directions. The smaller spheres thus define cones or solid angles about the center point, which can be indexed according to any suitable indexing algorithm. Although spherical quantization is described herein, any other suitable quantization (linear or non-linear) may also be used.
[0112] Furthermore, the quantized values may be further combined by determining whether the corresponding directional parameter elevation angle values are similar enough to use an embedded flexible boundary codebook.
[0113] In turn, the encoded direction parameters 206 may be passed to a combiner 211 .
[0114] The metadata encoder / quantizer 111 may include an energy ratio encoder 207. The energy ratio encoder 207 is configured to receive the energy ratio and determine the appropriate encoding for compressing the energy ratio of the subbands and time-frequency blocks. For example, in some embodiments, the energy ratio encoder 207 is configured to encode each energy ratio parameter value using 3 bits.
[0115] Furthermore, in some embodiments, instead of sending or storing all energy ratio values for all TF blocks, only a weighted average value per subband is sent or stored. The average value can be determined by considering the total energy of each time block, thereby favoring the values of subbands with more energy.
[0116] In such an embodiment, the quantized energy ratio value 208 is the same for all TF blocks of a given sub-band.
[0117] In some embodiments, the energy ratio encoder 207 is further configured to pass the quantized (encoded) energy ratio value 208 to the combiner 211 .
[0118] The metadata encoder / quantizer 111 may include a combiner 211. The combiner is configured to receive the encoded (or quantized / compressed) directional parameters and energy ratio parameters and combine these parameters to generate a suitable output (e.g., a metadata bitstream, which may be combined with a transmission signal or sent or stored separately from the transmission signal).
[0119] about Figure 3 , showing how according to some embodiments Figure 2 Example operation of the metadata encoder / quantizer shown in .
[0120] like Figure 3 As shown in step 301, the initial operation is to obtain metadata (such as azimuth angle value, elevation angle value, energy ratio, etc.).
[0121] Furthermore, if Figure 3 As shown in step 303, the orientation values (elevation angle, azimuth angle) may be compressed or encoded (eg, by applying spherical quantization or any suitable compression).
[0122] like Figure 3As shown in step 305, the energy ratio values are compressed or encoded (eg, by generating a weighted average for each subband and then quantizing it to a 3-bit value).
[0123] Furthermore, if Figure 3 As shown in step 307, the encoded orientation value, energy ratio, and coherence value are combined to generate encoded metadata.
[0124] about Figure 4 The direction encoder 205 is shown in more detail.
[0125] In some embodiments, the direction encoder may include a quantization determiner 401. The quantization determiner is configured to receive the encoded / quantized energy ratio 208 for each subband and determine from this value the quantization resolution of the azimuth and elevation angles for all time blocks of the current subband. The quantization resolution is set by allowing a predefined number of bits, bits_dir0[0:N-1][0:M-1], given by the value of the energy ratio. This can be output to the bit allocation manager 403.
[0126] The direction encoder 205 may further include a bit allocation manager 403 configured to receive the quantization resolution bits_dir0[0:N-1][0:M-1] of the azimuth and elevation angles of all time blocks of the current subband determined based on the energy ratio value and the allocated bits for the frame, modify the quantization resolution so that the number of allocated bits is reduced to bits_dir1[0:N-1][0:M-1] so that the total amount of allocated bits is equal to the number of available bits remaining after encoding the energy ratio. In turn, the reduced bit allocation bits_dir1[0:N-1][0:M-1] may be passed to the subband-based direction encoder 403.
[0127] The direction encoder 205 may further include a subband-based direction encoder 405 configured to receive the (reduced) bit allocation from the bit allocation manager 403. The subband-based direction encoder 405 is further configured to receive the direction parameters 108 and encode them based on the bit allocation on a subband-by-subband basis.
[0128] about Figure 5 , shows that Figure 4 A flow chart of the operation of the direction encoder 205 is shown in FIG.
[0129] like Figure 5 As shown in step 501, the initial operation is to obtain direction metadata (such as azimuth angle value, elevation angle value, etc.), encoded energy ratio value and bit allocation.
[0130] Furthermore, if Figure 5As shown in step 503 , the quantization resolution is initially determined based on the energy ratio value.
[0131] Furthermore, if Figure 5 As shown in step 505, the quantization resolution may be modified based on the bits allocated for the frame.
[0132] Furthermore, if Figure 5 As shown in step 507, the directional parameters may be compressed / encoded on a sub-band by sub-band basis based on the modified quantization resolution.
[0133] about Figure 6 , which shows in more detail Figure 4 The subband-based direction encoder 405 shown in FIG.
[0134] In some embodiments, the subband-based direction encoder 405 includes a subband bit allocator 601. The subband bit allocator 601 is configured to receive the (reduced) bit allocation bits_dir1[0:N-1][0:M-1] and determine the number of bits allowed for the subband. For example, bits_allowed = sum(bits_dir1[i][0:M-1]).
[0135] In some embodiments, the subband-based direction encoder 405 includes a bit limit determiner 603. The bit limit determiner 603 is configured to find the maximum number of bits max_b=max(bits_dir1[i][0:M-1]) allocated for each TF block of the current subband, and further determine whether the maximum number of bits allocated for each TF block of the current subband is less than or equal to the determined limited number of bits. In other words, whether max_b<=LIMIT_BITS_PER_TF. The value of the determined limited number of bits LIMIT_BITS_PER_TF is the limited number of bits per time-frequency (TF) tile, and a decision is made for the time-frequency tile as to whether to use an embedded quantizer instead of joint coding. For example, for an example TF tile (block) array in a subband, where M=4, when LIMIT_BITS_PER_TF=3, the number of bits in these time-frequency tiles can be (32 1 3) or (3 3 3 3) or (2 1 1 2), respectively, and the method can then start to check whether an embedded quantizer can be used.
[0136] In some embodiments, the subband-based direction encoder 405 includes a distance determiner 605. The distance determiner 605 can be controlled by the bit limit determiner 603 so that the distance d1d2 is determined when the maximum number of bits allocated for each TF block of the current subband is less than or equal to the allowed bits. Wherein, the angular distance is calculated as:
[0137]
[0138] Among them, θ av is the average elevation angle. The d1 distance is an estimate of the quantization distortion when using joint coding, and the d2 distance is an estimate of the quantization distortion when using a flexible embedded codebook.
[0139] In some embodiments, the estimation is performed based on the unquantized angles and the actual values in each codebook, without computing quantized values.
[0140] In some embodiments, the variance of the elevation angle is taken into account, in that if its variance is greater than a determined value, more than one elevation angle value is encoded for the subband. This is further detailed in PCT / FI2019 / 050675.
[0141] The distance determiner 605 is configured to determine whether the distance d2 is smaller than the distance d1 .
[0142] In some embodiments, the subband-based direction encoder 405 includes a joint elevation / azimuth encoder 607. The joint elevation / azimuth encoder 607 may be controlled by the determination of the bit limit determiner 603 to jointly encode the elevation value and the azimuth value of each TF block within the number of bits allocated for the current subband when the maximum number of bits allocated for each TF block of the current subband is greater than the allowed bits.
[0143] In addition, the joint elevation / azimuth encoder 607 can be controlled by the distance determiner 605 to jointly encode the elevation and azimuth values of each TF block within the number of bits allocated for the current subband when the distance d2 is greater than the distance d1. In other words, joint encoding is performed when the estimate of quantization distortion when using joint encoding is less than the estimate of quantization distortion when using the flexible embedded codebook.
[0144] In some embodiments, the subband-based direction encoder 405 includes an average elevation encoder / flexible boundary embedded codebook encoder 609. The average elevation encoder / flexible boundary embedded codebook encoder 609 may receive the distance and direction parameters and may be operable to determine, based on the distance determiner 605, that the distance d2 is greater than the distance d1 (when the maximum number of bits allocated for each TF block of the current subband is less than or equal to the allowed number of bits).
[0145] In some embodiments, the average elevation encoder / flexible boundary embedded codebook encoder 609 is configured to encode the average elevation value with 1 or 2 bits (1 bit for the value 0 degrees and 2 bits for + / -36 degrees). Additionally, the average elevation encoder / flexible boundary embedded codebook encoder 609 is configured to use a flexible boundary embedded codebook for the azimuth value of each considered TF tile.
[0146] Regarding Figure 7 The flowchart shown in
[0147] As Figure 7 shown in step 701 therein, the initial operation may be to allocate subband bits based on a modified quantization resolution.
[0148] Furthermore, as Figure 7 shown in step 703 therein, determine the maximum bits per quantization resolution based on the energy ratio value and perform a check on whether the maximum number of bits is less than the bit limit per time-frequency block.
[0149] As Figure 7 shown in step 710 therein, if the check determines that the maximum number of bits is greater than the limit, jointly encode the elevation and azimuth values of each time-frequency block within the number of bits allocated for the subband.
[0150] If the check determines that the maximum number of bits is less than or equal to the limit, determine distances d1 and d2 for the subframe of the current subband.
[0151] Furthermore, as Figure 7 shown in step 707 therein, check whether d2 < d1.
[0152] As Figure 7 shown in step 710 therein, if the distance d2 >= d1, jointly encode the elevation and azimuth values of each time-frequency block within the number of bits allocated for the subband.
[0153] As Figure 7 shown in step 709 therein, if the distance d2 < d1, encode the average elevation value with 1 or 2 bits (1 bit for the value 0 degrees and 2 bits for + / -36 degrees), and encode the azimuth value using a flexible boundary embedded codebook for each considered TF tile.
[0154] Furthermore, as Figure 7 shown in step 711 therein, output the subband-encoded values.
[0155] An example pseudo-code form of the energy ratio / direction encoding operation may be as follows:
[0156]
[0157]
[0158] about Figure 8 An example of an embedded codebook for azimuth is shown. This example shows a 1-bit codebook (indexes 0 and 1) covering the front 801 and rear 803 directions. The 2-bit codebook further adds left and right directions (indexes 2 and 3). The 3-bit codebook also adds center-left and center-right for the front-back directions (indexes 4, 5, 6, 7).
[0159] The selection of which codebook to encode the azimuth value based on the embedded codebook can be determined by Figure 9 The method shown in the example is shown.
[0160] like Figure 9 As shown in step 901, the first operation is to use the B-bit code to calculate the azimuth angle φ i Quantify
[0161] The next operation is to determine the number of bits (nbits) required to encode the index using an entropy encoder (e.g., a zero-order Golomb Rice encoder). Figure 9 As shown in step 903, the number of bits (nbits) is checked against the allowed number of bits (allowed_bits).
[0162] like Figure 9 As shown in step 904, if nbits<=allowed_bits, entropy coding is performed using the index.
[0163] like Figure 9 As shown in step 905, if nbits>allowed_bits, the angular quantization distortion is calculated for each quantized value (considering the quantized elevation angle value).
[0164] Furthermore, if Figure 9 As shown in step 907, the time-frequency image blocks are sorted in ascending order of their angular quantization distortion.
[0165] Furthermore, if Figure 9 As shown in step 909 , a loop is started for each time-frequency block in ascending order of the angular quantization distortion of the time-frequency block.
[0166] The loop determines if the quantized azimuth angle belongs only to the B-bit codebook and not the B-1-bit codebook, requantizes it in the B-1-bit codebook, and then recalculates the number of bits required to encode the (requantized) index using an entropy encoder (e.g., a Golomb Rice encoder).
[0167] Furthermore, if Figure 9 As shown in step 913, a check is performed to determine whether the number of bits (nbits) is less than or equal to the allowed bits.
[0168] like Figure 9 As shown in step 904, if the check determines that the number of bits (nbits) is less than or equal to the allowed bits, entropy coding is performed using the index.
[0169] like Figure 9 If the check determines that the number of bits (nbits) is greater than the allowed bits, then the loop is checked, as shown in step 915 .
[0170] If there are more TF tiles to be tested, the loop returns to the next increasing angular quantization distortion value.
[0171] like Figure 9 As shown in step 917 , if there are no more TF tiles to be processed, a further check is performed to determine whether the number of bits (nbits) is less than or equal to the allowed number of bits.
[0172] like Figure 9 As shown in step 904, if the check determines that the number of bits (nbits) is less than or equal to the allowed bits, entropy coding is performed using the index.
[0173] like Figure 9 As shown in step 919, if the check determines that the number of bits (nbits) is greater than the allowed bits, the codebook level is reduced, B=B-1 is set, and the angular quantization distortion is estimated for each of the new quantized values.
[0174] This can be summarized in pseudocode as:
[0175]
[0176] The C code implementation of the example can be:
[0177]
[0178]
[0179]
[0180]
[0181]
[0182]
[0183]
[0184] Regarding the decoder, the metadata decoder can be configured to determine whether the average elevation angle is signaled and to read the average elevation angle on 1 or 2 bits, with "0" representing the value 0, "10" representing the value +36, and "11" representing the value -36. Other values can also be used instead of 36 degrees. In addition, in some embodiments, more than one bit can be used to encode the average elevation angle, and thus it is possible to select between 5 values (0, + / - +theta_1, + / - \theta_2). The index of the azimuth angle can then be read using a 0th-order Golomb Rice decoder, and it is no longer necessary to signal which codebook they belong to.
[0185] Thus, the embodiments as discussed herein enable a significant reduction in table ROM for direction encoding, furthermore achieving a 30% reduction in the resulting angular quantization distortion due to the fact that the optimization is done in the angular distortion space and each component is examined individually.
[0186] about Figure 10 , illustrates an example electronic device that can be used as an analysis or synthesis device. The device can be any suitable electronic device or apparatus. For example, in some embodiments, the device 1400 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.
[0187] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 may be configured to execute various program codes, such as the methods described herein.
[0188] In some embodiments, the device 1400 includes a memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any suitable storage component. In some embodiments, the memory 1411 includes a program code portion for storing program codes that can be implemented on the processor 1407. In addition, in some embodiments, the memory 1411 can also include a storage data portion for storing data (e.g., data that has been processed or is to be processed according to the embodiments described herein). Whenever needed, the processor 1407 can obtain the implementation program code stored in the program code portion and the data stored in the storage data portion via the memory-processor coupling.
[0189] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, user interface 1405 can be coupled to processor 1407. In some embodiments, processor 1407 can control the operation of user interface 1405 and receive input from user interface 1405. In some embodiments, user interface 1405 can enable a user to enter commands to device 1400, for example, via a keyboard. In some embodiments, user interface 1405 can enable a user to obtain information from device 1400. For example, user interface 1405 can include a display configured to display information from device 1400 to a user. In some embodiments, user interface 1405 can include a touch screen or touch interface that enables information to be entered into device 1400 and also displays information to a user of device 1400. In some embodiments, user interface 1405 can be a user interface for communicating with a location determiner as described herein.
[0190] In some embodiments, device 1400 includes input / output port 1409. In some embodiments, input / output port 1409 includes a transceiver. In such embodiments, the transceiver can be coupled to processor 1407 and configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver device can be configured to communicate with other electronic devices or devices via a wired or wired coupling.
[0191] The transceiver can communicate with other devices via any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a wireless local area network (WLAN) protocol such as, for example, IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an infrared data communication path (IRDA).
[0192] The transceiver input / output port 1409 may be configured to receive signals and, in some embodiments, determine parameters as described herein using the processor 1407 executing appropriate code.
[0193] In general, various embodiments of the present invention may be implemented using hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented using hardware, while other aspects may be implemented using firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the present invention may be illustrated and described as block diagrams, flow charts, or using some other graphical representation, it is well known that the blocks, devices, systems, techniques, or methods described herein may be implemented using hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof, as non-limiting examples.
[0194] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any block of the logic flow as in the accompanying drawings may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on a physical medium such as a memory chip or a memory block implemented within a processor, on a magnetic medium such as a hard disk or floppy disk, and on an optical medium such as a DVD and its data variant CD.
[0195] The memory may be of any type suitable for the local technical environment and may be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technical environment and may include, by way of non-limiting example, one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit based on a multi-core processor architecture, and a processor.
[0196] Embodiments of the present invention may be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0197] Programs, such as those offered by Synopsys, Inc. of Mountain View, Calif., and Cadence Design, Inc. of San Jose, Calif., can automatically route conductors and position components on a semiconductor chip using well-established design rules and a library of pre-stored design blocks. Once the design of a semiconductor circuit is complete, the resulting design, in a standardized electronic format (e.g., Opus, GDSII, etc.), can be transferred to a semiconductor fabrication facility, or "fab," for fabrication.
[0198] The foregoing description has provided by way of exemplary and non-limiting examples a complete and informative description of the exemplary embodiments of the present invention. However, various modifications and adaptations will become apparent to those skilled in the relevant arts in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. Nevertheless, all such and similar modifications of the teachings of this invention will still fall within the scope of the invention as defined by the appended claims.
Claims
1. An apparatus for spatial audio coding, comprising: obtaining a directional parameter value associated with each of at least two time-frequency map blocks of at least one audio signal; and Based on the codebook, each direction parameter value is encoded, where The codebook includes two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level includes a first set of quantization values and a second or subsequent quantization level includes another set of quantization values and the quantization values of the previous quantization level, The component configured to encode each direction parameter value based on the codebook is further configured to: for each direction parameter value, determining a closest quantization value from said second or subsequent quantization level comprising said further set of quantization values and the quantization value of the preceding quantization level; and generating a codeword index for each direction parameter value based on the closest quantized value; determining a number of bits required to entropy encode a codeword index associated with the at least two time-frequency tiles; Comparing the number of bits required for entropy encoding the codeword index associated with the at least two time-frequency tiles with the allocated number of bits, and when the number of bits required for entropy encoding the codeword index is greater than the allocated number of bits, the apparatus comprises means configured to: sorting the at least two time-frequency tiles based on the determined angular quantization distortion of the closest quantization values determined for the directional parameter values associated with the at least two time-frequency tiles; and Iteratively, in the order of at least two ordered time-frequency tiles and until the number of bits required to entropy encode the codeword index associated with the at least two time-frequency tiles is equal to or less than the allocated number of bits, the component is configured to: select a closest quantization value of the directional parameter value associated with the ordered time-frequency tile; determine whether the selected closest quantization value is a member of the other set of quantization values but not a member of the first set of quantization values at the first quantization level; and when the determination is affirmative, use the first set of quantization values at the first quantization level to regenerate the codeword index for the directional parameter value associated with the ordered time-frequency tile and determine the number of bits required to entropy encode the codeword index associated with the at least two time-frequency tiles.
2. The device according to claim 1, wherein The component configured to encode the direction parameter value based on a codebook is further configured to: Encoding the azimuth direction parameter value based on the codebook; as well as The elevation direction parameter value is encoded based on at least one average elevation direction parameter value of the at least two time-frequency tiles.
3. The device according to claim 1, wherein The component is further configured to: Based on a value of the energy ratio value associated with the directional parameter value, an allocated number of bits for encoding the at least two time-frequency tiles is determined.
4. A method for spatial audio coding, comprising: Obtaining a directional parameter value associated with each of at least two time-frequency map blocks of at least one audio signal; as well as encoding each directional parameter value based on a codebook, wherein the codebook includes two or more quantization levels, the two or more quantization levels being arranged such that a first quantization level includes a first set of quantization values and a second or subsequent quantization level includes another set of quantization values and the quantization values of the previous quantization level, The step of encoding each direction parameter value based on the codebook further includes: for each direction parameter value, determining a closest quantization value from said second or subsequent quantization level comprising said further set of quantization values and the quantization value of the preceding quantization level; and generating a codeword index for each directional parameter value based on the associated closest quantized value; determining a number of bits required to entropy encode a codeword index associated with the at least two time-frequency tiles; Comparing a number of bits required for entropy encoding a codeword index associated with the at least two time-frequency tiles with the allocated number of bits, and when the number of bits required for entropy encoding the codeword index is greater than the allocated number of bits, the method comprises: sorting the at least two time-frequency tiles based on the determined angular quantization distortion of the closest quantization values determined for the directional parameter values associated with the at least two time-frequency tiles; and Iteratively, in the order of at least two ordered time-frequency tiles and until the number of bits required to entropy encode the codeword indices associated with the at least two time-frequency tiles is equal to or less than the allocated number of bits: selecting a closest quantization value of the directional parameter value associated with the ordered time-frequency tiles; determining whether the selected closest quantization value is a member of the other set of quantization values but not a member of the first set of quantization values at the first quantization level; and when the determination is affirmative, using the first set of quantization values at the first quantization level to regenerate the codeword indices for the directional parameter values associated with the ordered time-frequency tiles, and determining the number of bits required to entropy encode the codeword indices associated with the at least two time-frequency tiles.
5. The method according to claim 4, wherein Encoding the direction parameter value based on the codebook further includes: Encoding the azimuth direction parameter value based on the codebook; and The elevation direction parameter value is encoded based on at least one average elevation direction parameter value of the at least two time-frequency tiles.
6. The method according to claim 4, further comprising: Based on the value of the energy ratio value associated with the obtained directional parameter value, the allocated number of bits for encoding the at least two time-frequency tiles is determined.
Citation Information
Patent Citations
Determination of spatial audio parameter encoding and associated decoding
GB2575305A
Determination of spatial audio parameter encoding and associated decoding
WO2020008105A1