Quantization of Spatial Audio Direction Parameters

By deriving, rotating, quantizing and indexing audio direction parameters, the problem of difficult to effectively encode direction-related parameters of multiple types of audio objects and speaker signals in the prior art is solved, and efficient encoding and decoding of complex sound fields and multi-channel inputs is achieved.

CN114207713BActive Publication Date: 2025-06-13NOKIA TECHNOLOGIES OY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080055578.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-31
Filing Date
2020-06-15
Publication Date
2025-06-13
Estimated Expiration
2040-06-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively encode direction-related parameters in multiple types of audio objects and speaker signals, especially when dealing with multi-channel inputs and complex sound fields.

Method used

By exporting audio direction parameters, rotating these parameters, quantizing and indexing according to the differences, encoding audio direction parameters can be achieved. The specific steps include: exporting the audio direction parameters, rotating the parameters, changing the position, calculating the differences and quantizing them, and finally indexing them.

Benefits of technology

It realizes efficient encoding of direction-related parameters in a variety of audio objects and speaker signals, can handle complex sound fields and multi-channel inputs, and improves the encoding and decoding effects of spatial audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114207713B_ABST
    Figure CN114207713B_ABST
Patent Text Reader

Abstract

In particular, a device for spatial audio signal encoding is disclosed, which is configured to: for each of a plurality of audio direction parameters, derive a corresponding derived audio direction parameter including an elevation value and an azimuth value. Each derived audio direction parameter is rotated by the azimuth value of the audio direction parameter at a first position among the plurality of audio direction parameters. The positions of some of the audio direction parameters are changed, and then, for each of the plurality of audio direction parameters, the difference between each audio direction parameter and the corresponding rotated derived audio direction parameter is determined. Furthermore, the difference for each of the plurality of audio direction parameters is quantized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to apparatus and methods for encoding parameters related to sound fields, and more particularly, but not exclusively, to apparatus and methods for encoding direction-related parameters for audio encoders and decoders. Background Art

[0002] Parametric spatial audio processing is an area of audio signal processing that uses a set of parameters to describe the spatial aspects of sound. For example, in parametric spatial audio capture from a microphone array, it is a typical and effective option to estimate a set of parameters from the microphone array signals, such as the direction of sound in a frequency band and the ratio of the directional to non-directional parts of the sound captured in the frequency band. It is well known that these parameters well describe the perceived spatial characteristics of the captured sound at the location of the microphone array. These parameters can accordingly be used in the synthesis of spatial sound for use in, for example, binaural headphones, loudspeakers, or other formats such as Ambisonics.

[0003] Therefore, direction in a frequency band and direct-to-total energy ratio are particularly effective parameterizations for spatial audio capture.

[0004] A set of parameters including direction parameters in a frequency band and energy ratio parameters in a frequency band (indicating the directionality of sound) can also be used as spatial metadata for an audio codec. For example, these parameters can be estimated from an audio signal captured by a microphone array, and for example, a stereo signal can be generated from the microphone array signals to be transmitted together with the spatial metadata. The stereo signal can be encoded, for example, with an AAC encoder. The decoder can decode the audio signal into a PCM signal and process the sound in the frequency band (using the spatial metadata) to obtain a spatial output, for example, a binaural output.

[0005] The foregoing solutions are particularly suitable for encoding captured spatial sound from a microphone array (e.g., in a mobile phone, VR camera, stand-alone microphone array). However, it is desirable for such an encoder to have other input types in addition to the signals captured by the microphone array, such as speaker signals, audio object signals, or Ambisonic signals.

[0006] Analysis of first-order Ambisonics (FOA) inputs for spatial metadata extraction has been well documented in the scientific literature related to Directional Audio Coding (DirAC) and Harmonic Plane Wave Expansion (Harpex). This is because there are microphone arrays that directly provide FOA signals (more precisely: their variant, B-format signals), and thus analysis of such inputs has become a research focus in the field.

[0007] Another input to the encoder can also be a multi-channel loudspeaker input, such as a 5.1 or 7.1 channel surround sound input, or a metadata-assisted spatial audio (MASA) format input.

[0008] However, with respect to the input audio object type to the encoder, there may be accompanying metadata that includes the directional components of each audio object within the physical space. These directional components can include the elevation angle and the azimuth angle of the position of the audio object within the space. SUMMARY OF THE INVENTION

[0009] According to a first aspect, there is provided a method for encoding a spatial audio signal, comprising: for each of a plurality of audio direction parameters, wherein each parameter includes an elevation angle value and an azimuth angle value and each parameter has an ordered position, deriving a corresponding derived audio direction parameter, the corresponding derived audio direction parameter including an elevation angle value and an azimuth angle value; rotating each derived audio direction parameter by the azimuth angle value of the audio direction parameter at a first position among the plurality of audio direction parameters; when the azimuth angle value of an audio direction parameter is closest to the azimuth angle value of another rotated derived audio direction parameter compared to the azimuth angle values of other rotated derived audio direction parameters, changing the ordered position of the audio direction parameter to another position that is consistent with the position of the rotated derived audio direction parameter, and then for each of the plurality of audio direction parameters, determining the difference between each audio direction parameter and the corresponding rotated derived audio direction parameter; and for each of the plurality of audio direction parameters, quantifying the difference.

[0010] The azimuth angle value of each derived audio direction parameter can correspond to a position among a plurality of positions around the circumference of a circle.

[0011] The plurality of positions around the circumference of the circle can be evenly distributed along 360 degrees of the circle, and wherein the number of positions around the circumference of the circle is determined by the number of audio direction parameters.

[0012] Rotating each derived audio direction parameter by the azimuth angle value of a first audio direction parameter among the plurality of audio direction parameters can include: adding the azimuth angle value of the first audio direction parameter to the azimuth angle value of each derived audio direction parameter, wherein the elevation angle value of each derived audio direction parameter is set to zero.

[0013] The method can further include: scalar quantifying the azimuth angle value of the first audio direction parameter; and indexing the position of the audio direction parameter after the change by assigning an index that represents the order of the indices of the arrangement of the positions of the audio direction parameters.

[0014] Determining the difference between each audio direction parameter among a plurality of audio direction parameters and the corresponding rotated derived audio direction parameter may include, for each audio direction parameter among the plurality of audio direction parameters, determining a differential audio direction parameter based at least on the following operations: determining the difference between the audio direction parameter located at a first position and the rotated derived audio direction parameter located at the first position, and / or determining the difference between another audio direction parameter and the rotated derived audio direction parameter, where the position of the another audio direction parameter has not changed, and / or determining the difference between yet another audio direction parameter and the rotated derived audio direction parameter, where the position of the yet another audio direction parameter has been changed to the position of the rotated derived audio direction parameter.

[0015] Determining the difference between an audio direction parameter and the corresponding rotated derived audio direction parameter may include: determining the difference between the azimuth value of the audio direction parameter and the azimuth value of the corresponding rotated derived audio direction parameter; and determining the difference between the elevation value of the audio direction parameter and the elevation value of the corresponding rotated derived audio direction parameter.

[0016] Changing the position of an audio direction parameter to another position may be applicable to any audio direction parameter other than the audio direction parameter located at a first position.

[0017] Quantifying the differential audio direction parameter for each audio direction parameter among a plurality of audio direction parameters may include: for each audio direction parameter among the plurality of audio direction parameters, quantifying the differential audio direction parameter as a vector, where the vector is indexed to a codebook that includes a plurality of indexed elevation values and indexed azimuth values.

[0018] The plurality of indexed elevation values and indexed azimuth values may be points on a grid arranged in the form of a sphere, where the spherical grid may be formed by covering the sphere with smaller spheres, and the smaller spheres define the points of the spherical grid.

[0019] According to a second aspect, there is provided a method for decoding a spatial audio signal, which includes: decoding an index to provide a quantized azimuth value of an audio direction parameter at a first position among a plurality of ordered audio direction parameters, wherein each parameter includes an elevation value and an azimuth value; for each of the plurality of audio direction parameters, deriving a corresponding derived audio direction parameter, the corresponding derived audio direction parameter including an elevation value and an azimuth value; rotating each derived audio direction parameter by the azimuth value of the audio direction parameter at the first position among the plurality of audio direction parameters; decoding the index to provide a quantized difference between the audio direction parameter and its corresponding derived audio direction parameter for each audio direction parameter; for each audio direction parameter, forming a quantized audio direction parameter by adding the quantized difference to its corresponding derived audio direction parameter; and decoding an index representing the order of the plurality of quantized audio direction parameters and reordering the positions of the plurality of quantized audio direction parameters according to the order.

[0020] The azimuth value of each derived audio direction parameter can correspond to a position among a plurality of positions around the circumference of a circle.

[0021] The plurality of positions around the circumference of the circle can be evenly distributed along the 360 degrees of the circle, and the number of positions around the circumference of the circle can be determined by the number of audio direction parameters.

[0022] Rotating each derived audio direction parameter by the azimuth value of the first audio direction parameter among the plurality of audio direction parameters can include: adding the quantized azimuth value of the first audio direction parameter to the azimuth value of each derived audio direction parameter, wherein the elevation value of each derived audio direction parameter is set to zero.

[0023] The index for providing the quantized difference between the audio direction parameter and its corresponding derived audio direction parameter for each audio direction parameter can be an index of a codebook, the codebook including a plurality of indexed elevation values and indexed azimuth values.

[0024] The plurality of indexed elevation values and indexed azimuth values can be points on a grid arranged in the form of a sphere, and the spherical grid can be formed by covering the sphere with smaller spheres, and the smaller spheres can define the points of the spherical grid.

[0025] According to a third aspect, there is provided an apparatus for encoding a spatial audio signal, comprising: for each of a plurality of audio direction parameters, where each parameter includes an elevation value and an azimuth value and each parameter has an ordered position, deriving a corresponding derived audio direction parameter, the corresponding derived audio direction parameter including an elevation value and an azimuth value; rotating each derived audio direction parameter by the azimuth value of the audio direction parameter at a first position among the plurality of audio direction parameters; when the azimuth value of an audio direction parameter is closest to the azimuth value of another rotated derived audio direction parameter compared to the azimuth values of other rotated derived audio direction parameters, changing the ordered position of the audio direction parameter to another position that is consistent with the position of the rotated derived audio direction parameter, and then for each of the plurality of audio direction parameters, determining the difference between each audio direction parameter and the corresponding rotated derived audio direction parameter; and for each of the plurality of audio direction parameters, quantifying the difference.

[0026] The azimuth value of each derived audio direction parameter can correspond to a position among a plurality of positions around the circumference of a circle.

[0027] The plurality of positions around the circumference of the circle can be evenly distributed along 360 degrees of the circle, and wherein the number of positions around the circumference of the circle is determined by the number of audio direction parameters. The apparatus configured to rotate each derived audio direction parameter by the azimuth value of a first audio direction parameter among the plurality of audio direction parameters can be configured to: add the azimuth value of the first audio direction parameter to the azimuth value of each derived audio direction parameter, where the elevation value of each derived audio direction parameter is set to zero.

[0028] The apparatus can be further configured to: scalar quantize the azimuth value of the first audio direction parameter; and index the position of the audio direction parameter after the change by assigning an index that represents the order of the indices arranged for the positions of the audio direction parameters.

[0029] The apparatus configured to determine the difference between each audio direction parameter and the corresponding rotated derived audio direction parameter for each of the plurality of audio direction parameters can be configured to determine a differential audio direction parameter for each of the plurality of audio direction parameters based on at least the following operations: determining the difference between the first located audio direction parameter and the first located rotated derived audio direction parameter, and / or determining the difference between another audio direction parameter and the rotated derived audio direction parameter, where the position of the another audio direction parameter has not changed, and / or determining the difference between yet another audio direction parameter and the rotated derived audio direction parameter, where the position of the yet another audio direction parameter has been changed to the position of the rotated derived audio direction parameter.

[0030] The apparatus configured to determine a difference between an audio direction parameter and a corresponding rotated derived audio direction parameter may be configured to: determine a difference between an azimuth value of the audio direction parameter and an azimuth value of the corresponding rotated derived audio direction parameter; and determine a difference between an elevation value of the audio direction parameter and an elevation value of the corresponding rotated derived audio direction parameter.

[0031] The apparatus configured to change a position of an audio direction parameter to another position may be applicable to any audio direction parameter other than the first-positioned audio direction parameter.

[0032] The apparatus configured to quantify a differential audio direction parameter for each of a plurality of audio direction parameters may be configured to: for each of the plurality of audio direction parameters, quantify the differential audio direction parameter as a vector, where the vector is indexed to a codebook that includes a plurality of indexed elevation values and indexed azimuth values.

[0033] The plurality of indexed elevation values and indexed azimuth values may be points on a grid arranged in the form of a sphere, where the spherical grid may be formed by covering the sphere with smaller spheres, and where the smaller spheres define the points of the spherical grid.

[0034] According to a fourth aspect, there is provided an apparatus for spatial audio signal decoding, configured to: decode an index to provide a quantized azimuth value of an audio direction parameter at a first position among a plurality of ordered audio direction parameters, where each parameter includes an elevation value and an azimuth value; for each of the plurality of audio direction parameters, derive a corresponding derived audio direction parameter, the corresponding derived audio direction parameter including an elevation value and an azimuth value; rotate an azimuth value of each derived audio direction parameter by the azimuth value of the audio direction parameter at the first position among the plurality of audio direction parameters; decode an index to provide a quantized difference between an audio direction parameter and its corresponding derived audio direction parameter for each audio direction parameter; for each audio direction parameter, form a quantized audio direction parameter by adding the quantized difference to its corresponding derived audio direction parameter; and decode an index representing an order of a plurality of quantized audio direction parameters and reorder positions of the plurality of quantized audio direction parameters according to the order.

[0035] The azimuth value of each derived audio direction parameter may correspond to a position among a plurality of positions around a circumference of a circle.

[0036] Multiple positions around the circumference of a circle can be evenly distributed along the 360 degrees of the circle, and, among them, the number of positions around the circumference of the circle can be determined by the number of audio direction parameters.

[0037] The apparatus configured to rotate the azimuth value of the first audio direction parameter among the plurality of audio direction parameters for each derived audio direction parameter can be configured to: add the quantized azimuth value of the first audio direction parameter to the azimuth value of each derived audio direction parameter, wherein the elevation value of each derived audio direction parameter is set to zero.

[0038] The index providing the quantized difference between an audio direction parameter and its corresponding derived audio direction parameter for each audio direction parameter can be an index of a codebook that includes a plurality of indexed elevation values and indexed azimuth values.

[0039] The plurality of indexed elevation values and indexed azimuth values can be points on a grid arranged in the form of a sphere, wherein the spherical grid can be formed by covering the sphere with smaller spheres, and wherein the smaller spheres define the points of the spherical grid.

[0040] According to a fifth aspect, there is provided an apparatus for spatial audio coding, which includes at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, together with the at least one processor, cause the apparatus to: for each audio direction parameter among a plurality of audio direction parameters, wherein each parameter includes an elevation value and an azimuth value and each parameter has an ordered position, derive a corresponding derived audio direction parameter, the corresponding derived audio direction parameter including an elevation value and an azimuth value; rotate the azimuth value of the audio direction parameter at the first position among the plurality of audio direction parameters for each derived audio direction parameter; when the azimuth value of an audio direction parameter is closest to the azimuth value of another rotated derived audio direction parameter compared with the azimuth values of other rotated derived audio direction parameters, change the ordered position of the audio direction parameter to another position consistent with the position of the rotated derived audio direction parameter, and then, for each audio direction parameter among the plurality of audio direction parameters, determine the difference between each audio direction parameter and its corresponding rotated derived audio direction parameter; and for each audio direction parameter among the plurality of audio direction parameters, quantize the difference.

[0041] According to a sixth aspect, there is provided an apparatus for spatial audio decoding, comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to: decode an index to provide a quantized azimuth value of an audio direction parameter at a first position among a plurality of ordered audio direction parameters, wherein each parameter includes an elevation value and an azimuth value; for each audio direction parameter among the plurality of audio direction parameters, derive a corresponding derived audio direction parameter, the corresponding derived audio direction parameter including an elevation value and an azimuth value; rotate each derived audio direction parameter by the azimuth value of the audio direction parameter at the first position among the plurality of audio direction parameters; decode the index to provide a quantized difference between an audio direction parameter and its corresponding derived audio direction parameter for each audio direction parameter; for each audio direction parameter, form a quantized audio direction parameter by adding the quantized difference to its corresponding derived audio direction parameter; and decode an index representing the order of a plurality of quantized audio direction parameters and reorder the positions of the plurality of quantized audio direction parameters according to the order.

[0042] A computer program product stored on a medium can cause an apparatus to perform the methods described herein.

[0043] An electronic device may include an apparatus as described herein.

[0044] A chipset may include an apparatus as described herein.

[0045] Embodiments of the present application are intended to solve problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] To better understand the present application, reference will now be made, by way of example, to the accompanying drawings, in which:

[0047] Figure 1 Schematically shows a system of an apparatus suitable for implementing some embodiments;

[0048] Figure 2 Schematically shows, according to some embodiments, an audio object encoder as shown in Figure 1 ;

[0049] Figure 3a Schematically shows, according to some embodiments, an implemented spherical quantizer & indexer as shown in Figure 2 ;

[0050] Figure 3b Schematically shows, according to some embodiments, a spherical de-indexer as shown in Figure 5 ;

[0051] Figure 3c Schematically shows an example sphere position configuration used in a spherical quantizer & indexer and a spherical deindexer as shown in Figure 3a and Figure 3b ;

[0052] Figure 4 Shows a flowchart of the operation of an audio object encoder as shown in Figure 2 ;

[0053] Figure 5 Schematically shows an audio object decoder as shown in Figure 1 according to some embodiments;

[0054] Figure 6 Shows in more detail a flowchart for generating a direction index based on input direction parameters;

[0055] Figure 7 Shows a flowchart of an example operation of quantizing a direction parameter to obtain a direction index;

[0056] Figure 8 Shows a flowchart of the operation of an audio object decoder as shown in Figure 5 according to some embodiments; and

[0057] Figure 9 Schematically shows an example device adapted to implement the shown apparatus. DETAILED DESCRIPTION

[0058] The following describes in more detail suitable apparatuses and possible mechanisms for providing metadata parameters for effective spatial analysis derivation for multi-channel input format audio signals and input audio objects. In the following discussion, multi-channel systems will be discussed with respect to multi-channel microphone implementations. However, as noted above, the input format can be any suitable input format, such as multi-channel speakers, Ambisonic (FOA / HOA), etc. It should be understood that in some embodiments, the channel positions are based on the positions of the microphones, or on virtual positions or directions. Additionally, the output of the example system is a multi-channel speaker arrangement. However, it should be understood that the output can be rendered to the user by means other than speakers. Further, the multi-channel speaker signals can be generalized to two or more playback audio signals.

[0059] As discussed previously, spatial metadata parameters in a frequency band, such as the direct-to-total energy ratio (or diffuseness-ratio, absolute energies, or any suitable representation indicating the directivity / non-directivity of the sound in a given time-frequency interval) parameter, are particularly suitable for representing the perceptual characteristics of a natural sound field. Synthetic sound scenes such as 5.1 speaker mixes typically utilize audio effects and amplitude panning methods, which provide spatial sounds different from those occurring in a natural sound field. In particular, a 5.1 or 7.1 mix can be configured such that it contains coherent sounds played from multiple directions. For example, some of the sounds in a 5.1 mix that are typically perceived directly in the front are not generated by the center (channel) speaker, but rather, for example, from the left front and right front (channel) speakers and also possibly coherently from the center (channel) speaker. Spatial metadata parameters such as direction and energy ratios do not accurately represent such spatial coherence characteristics. Thus, other metadata parameters such as coherence parameters can be determined from the analysis of the audio signal to represent the audio signal relationship between channels.

[0060] In addition to multi-channel input format audio signals, it may also be necessary for an encoding system to encode audio objects representing various sound sources within a physical space. Whether it is in the form of metadata or some other mechanism, each audio object can be accompanied by directional data in the form of azimuth and elevation values, which indicate the position of the audio object within the physical space.

[0061] As mentioned above, an example of incorporating direction information for an audio object into metadata is to use the determined azimuth and elevation values.

[0062] Thus, this concept attempts to determine direction parameters for audio objects and index the parameters based on an actual sphere-covered direction distribution in order to define a more uniform direction distribution.

[0063] Furthermore, the proposed direction indexing for audio objects can be used with a downmixed signal ("channel") to define, for example, a parameterized immersive format that can be used in an Immersive Voice and Audio Service (IVAS) codec. Alternatively and additionally, a spherical grid format can be used in the codec to quantize the direction.

[0064] In addition, this concept discusses the decoding of such indexed direction parameters to produce quantized direction parameters that can be used in spatial audio synthesis based on audio object sound field-related parameters.

[0065] Regarding Figure 1, which shows an example apparatus and system for implementing embodiments of the present application. System 100 is shown as having an "analysis" section 121 and a "synthesis" section 131. The "analysis" section 121 is the part that encodes from receiving a multi-channel speaker signal to metadata and a downmixed signal, while the "synthesis" section 131 is the part that decodes from the encoded metadata and downmixed signal to the presentation of the regenerated signal (e.g., in the form of multi-channel speakers).

[0066] The input to system 100 and the "analysis" section 121 is the multi-channel signal 102. In the following examples, a microphone channel signal input is described. However, in other embodiments, any suitable input (or synthesized multi-channel) format can be implemented.

[0067] The multi-channel signal is passed to a downmixer 103 and an analysis processor 105.

[0068] In some embodiments, the downmixer 103 is configured to receive the multi-channel signal, downmix these signals to a determined number of channels, and output a downmixed signal 104. For example, the downmixer 103 can be configured to generate a 2-audio-channel downmix of the multi-channel signal. The determined number of channels can be any suitable number of channels. In some embodiments, the downmixer 103 is optional, and the multi-channel signal is passed to the encoder 107 without being processed in the same manner as the downmixed signal in this example.

[0069] In some embodiments, the analysis processor 105 is also configured to receive the multi-channel signal and analyze these signals to generate metadata 106 associated with the multi-channel signal and thus with the downmixed signal 104. The analysis processor 105 can be configured to generate metadata that, for each time-frequency analysis interval, can include a direction parameter 108, an energy ratio parameter 110, a coherence parameter 112, and a diffuseness parameter 114. In some embodiments, the direction, energy ratio, and diffuseness parameters can be considered spatial audio parameters. In other words, spatial audio parameters include parameters intended to characterize the sound field created by the multi-channel signal (or generally two or more played audio signals).

[0070] In some embodiments, the generated parameters can vary from band to band. Thus, for example, in band X, all parameters are generated and sent, while in band Y, only one of the parameters is generated and sent. Additionally, in band Z, no parameters are generated or sent. A practical example in this regard can be that for some bands such as the highest frequency band, certain parameters are not required for perceptual reasons. The downmixed signal 104 and the metadata 106 can be passed to the encoder 107.

[0071] The encoder 107 may include an IVAS stereo core 109 configured to receive downmixed (or other) signals 104 and generate a suitable encoding of these audio signals. In some embodiments, the encoder 107 may be a computer (running suitable software stored on a memory and on at least one processor), or alternatively may be a specific device such as using an FPGA or ASIC. The encoding may be implemented using any suitable scheme. Additionally, the encoder 107 may include a metadata encoder or quantizer 109 configured to receive metadata and output the information in an encoded or compressed form. Additionally, there may also be an audio object encoder 121 within the encoder 107 which, in an embodiment, may be arranged to encode data (or metadata) associated with a plurality of audio objects along input 120. The data associated with the plurality of audio objects may include at least a portion of the directional data.

[0072] In some embodiments, prior to transmission or storage, as indicated by the dashed line in Figure 1 , the encoder 107 may further interleave, multiplex to a single data stream, or embed the metadata within the encoded downmixed signal. The multiplexing may be implemented using any suitable scheme.

[0073] On the decoder side, the received or acquired data (stream) may be received by the decoder / demultiplexer 133. The decoder / demultiplexer 133 may demultiplex the encoded stream and pass the audio encoded stream to a downmix extractor 135 configured to decode the audio signals to obtain a downmixed signal. Similarly, the decoder / demultiplexer 133 may include a metadata extractor 137 configured to receive the encoded metadata and generate the metadata. Additionally, the decoder / demultiplexer 133 may also include an audio object decoder 141 which may be configured to receive the encoded data associated with a plurality of audio objects and decode such data accordingly to produce corresponding decoded data 140. In some embodiments, the decoder / demultiplexer 133 may be a computer (running suitable software stored on a memory and on at least one processor), or alternatively may be a specific device such as using an FPGA or ASIC.

[0074] The decoded metadata and the downmixed audio signal may be passed to a synthesis processor 139.

[0075] The "synthesis" portion 131 of the system 100 further shows a synthesis processor 139 configured to receive the downmix and the metadata and recreate synthetic spatial audio in the form of a multichannel signal 110 in any suitable format based on the downmixed signal and the metadata (which may be a multichannel speaker format according to the usage scenario, or in some embodiments may be any suitable output format such as a binaural or Ambisonics signal).

[0076] The additional input 120 may specifically include the directional data associated with a plurality of audio objects. A specific example of such a use case is a conference call scenario, in which the participants are positioned around a table. Each audio object may represent the audio data associated with each participant. In particular, the audio object may have the position data associated with each participant. The data associated with the audio object is Figure 1 described as being passed to the audio object encoder 121 in

[0077] Return Figure 1 , note that the system 100 may be configured to accept a plurality of audio objects along the input 120, and each audio object may have the associated directional data. Further, the audio object including the associated directional data may be passed to the audio object encoder 121 for encoding and quantization. In this regard, the directional data associated with each audio object may also be represented according to the azimuth angle φ and the elevation angle θ, where the azimuth angle value and the elevation angle value of each audio object indicate the position of the object in space at any point in time. The azimuth angle and the elevation angle values may be updated on a per time frame basis, which does not have to be consistent with the time frame resolution of the directional metadata parameters associated with the multi-channel audio signal.

[0078] Generally, the directional information of N active input audio objects for the audio object encoder 121 may be in the form of P q =(θ q ,φ q ), q = 0:N-1, where P q is the directional information of the audio object with index q, which has a two-dimensional vector including the elevation angle θ and the azimuth angle φ.

[0079] The concept of this article is to find the vector difference between the directional information of the audio object and the "template" audio direction parameters derived for the audio object, and then use a spherical quantization scheme to quantize the vector difference. In this regard, Figure 2 some functions of the audio object encoder 121 are depicted in more detail.

[0080] The audio object encoder 121 may include an audio object direction exporter 201, which is configured to export the appropriate "template" audio direction parameters for each audio object. In an embodiment, this may be exported as an N-dimensional vector, which has N exported audio direction parameters corresponding to the N audio objects as elements. These exported audio direction parameters may be derived from the perspective of considering the circumferential distribution of the audio objects around a circle. In particular, the exported audio direction parameters may be considered from the angle that the audio object directions are evenly distributed as N equally spaced points around the unit circle.

[0081] In the following description, N derived audio direction parameters are disclosed as being formed into a vector structure (referred to as a vector SP), where each element corresponds to a derived audio direction parameter for one of the N audio objects. However, it should be understood that the following disclosure can be applied by treating the derived audio direction parameters as a set of index parameters that do not need to be constructed in the form of a vector.

[0082] The audio object direction exporter 201 can be configured to export a "template" derived audio direction vector SP having N two-dimensional elements, whereby each element represents an azimuth angle and an elevation angle associated with an audio object. Further, the vector SP can be initialized by setting the azimuth angle and elevation angle values of each element such that the N audio objects are evenly distributed around the unit circle. This can be achieved by initializing each audio object direction element within the vector to have an elevation angle value of "zero" and an azimuth angle value where q is the index of the associated audio object. Thus, for N audio objects, the vector SP can be written as:

[0083]

[0084] In other words, the SP vector can be initialized such that the orientation information (derived audio direction parameter) of each audio object is assumed to be evenly distributed along the unit circle starting from an azimuth angle value of 0 0 and proceeding.

[0085] In other embodiments, the SP vector can be initialized with a non-zero elevation angle value. For example, the same elevation angle value can be used for each derived audio direction element of the SP vector. In this case, the SP vector will no longer lie in the horizontal plane but on an inclined plane.

[0086] This processing step of initializing the derived audio direction parameters associated with each audio object is shown as Figure 4 processing step 401 in

[0087] Further, the derived audio direction SP vector having elements including the derived audio direction parameters corresponding to the audio objects can be passed to the audio direction rotator 203 in the audio object encoder 121. The audio direction rotator 203 is also depicted as receiving the audio object 120. In particular, by rotating each derived direction within the SP vector by a first component φ 0 in the azimuth angle value of the first received audio object P 0 the audio direction rotator 203 can then use the audio direction parameter of the first audio object in subsequent processing. That is, by adding the first azimuth angle component φ 0The value can rotate each azimuth component of each derived audio direction parameter within the derived vector SP. In terms of the SP vector, this operation results in each element having the following form:

[0088]

[0089] For embodiments where the derived audio direction vector SP is deployed with an elevation angle of zero, the vector SP can be represented solely in terms of the azimuth angle wherein, is the rotated azimuth component given by and is the rotated derived audio direction vector SP.

[0090] For embodiments where the elevation direction component of the derived vector SP is initialized to some initial elevation value, there can also be a rotation applied to the derived direction elevation value of the derived vector SP. For example, in these embodiments, each element of the derived vector SP can be rotated by a first direction component from the first received audio object θ o As a result of this step, the rotated derived audio direction vector

[0091] is now aligned with the direction of the first audio object on the unit circle.

[0092] Returning to Figure 4 the flowchart, this step can be represented as processing step 403.

[0093] Furthermore, the audio object encoder 121 can be set to quantize and encode the above-mentioned rotated derived audio direction vector In an embodiment, this can simply include quantizing the rotation angle φ 0 to a specific resolution by the quantizer 211. For example, a linear quantizer with a 2.5-degree resolution (i.e., 5 degrees between consecutive points on a linear scale) results in 72 linear quantization levels. Note that the (unrotated) derived audio direction vector SP depends on the number N of active audio objects, and this factor can be passed to the decoder or otherwise agreed upon with the encoder.

[0094] For embodiments where the elevation direction component of the derived vector SP is initialized to some initial elevation value, the first received audio object θ o can also be scalar quantized.

[0095] Figure 4 The step of quantizing the rotation angle is shown as processing step 405 in

[0096] ​​The audio object encoder 121 may also include an audio direction relocator & indexer 205 configured to reorder the positions of the received audio objects to more closely align with the rotated exported audio direction vectors of the rotated exported audio directions of the elements. This can be achieved by reordering the positions of the audio objects such that the azimuth value of each reordered audio object is aligned with the position of the element in the vector having the closest azimuth value. Further, the reordered positions of each audio object can be encoded as a permutation index. This process may include the following algorithmic steps:

[0097] 1. Assign an index as a vector to each active audio object in the order they are received, which can be represented as I = (i 0 , i 1 , i 2 ... i N-1 ).

[0098] 2. Rearrange all indices except the first index i 0 such that if the azimuth associated with the audio object φ i is closest to the azimuth at position j among all azimuths in the rotated exported vector then the index i currently at position i is moved to position j. i

[0099] For example, an example including four active audio objects. The SP code vector can be initialized uniformly along the unit circle as SP = (0, 0; 0, 90; 0, 180; 0, 270). The orientation data ((θ 0 , φ 0 ); (θ 1 , φ 1 );.... (θ N-1, φ N-1 )) can be received as ((0, 130); (0, 210); (0, 39); (0, 310), where the first φ 0 is given as 130 degrees. In this particular example, the rotated azimuths in the vector are given by (0 + 130, 90 + 130, 180 + 130, 270 + 130) = (130; 220; 310; 400) = (130, 220, 310, 40). In this example, the second audio object with azimuth 210 is closest to the second azimuth in the vector , and the third audio object with azimuth 30 is closest to the vector In the fourth azimuth angle, the fourth audio object with azimuth angle 310 is closest to the vector In the third azimuth angle. Thus, in this case, the re-ordered audio object index vector is

[0100] 3. Further, the re-ordered audio object indices can be indexed according to a specific permutation of the indices. Each specific permutation of the indices of the re-ordered audio objects can be assigned an index value. However, it should be understood that the first index position of the re-ordered audio objects is not part of the index permutation, since the index of the first element in the vector has not changed. That is, the first audio object always remains in the first position, since this is the audio object towards which the elements of the derived audio direction vector SP are rotated. Thus, there can be (N - 1)! index permutations of the re-ordered audio objects, which can be represented within the range of log 2 ((N - 1)!) bits.

[0101] Returning to the example of the system with 4 active audio objects above, only the indices of i 3 , i 1 , i 2 need to be indexed. The indices of the possible index permutations of the re-ordered audio objects for the above exemplary example can take the following form:

[0102]

[0103] Thus, in order to represent the re-ordered audio objects, it may be necessary to send the azimuth angle φ 0 of the first object, in order to represent the rotated derived audio parameters and the index indicating the relative order of the positions of the re-ordered audio objects.

[0104] The above processing steps of arranging the positions of the audio objects in an order such that the azimuth angle values of the arranged audio objects correspond to the azimuth angle values closest to the derived direction and indexing the positions of all audio objects except the first audio object are respectively shown as steps 407 and 409 in Figure 4 .

[0105] The K bits (which may be referred to as I 0 ) for scalar quantizing the azimuth angle φ φ0 of the first object and the index I ro representing the index order of the audio direction parameters of audio objects 1 to N - 1 can form part of an encoded bitstream such as from encoder 100.

[0106] In some embodiments, the scalar quantized elevation angle θ o of the first object may also form part of the encoded bitstream.

[0107] As described above, the rotated derived audio direction vector can be a "template" from which an audio orientation difference vector can be derived for the audio orientation parameters of each audio object. This can be performed, for example, by Figure 2 the difference determiner 207 in. In an embodiment, the audio orientation difference vector can be a two-dimensional vector having an elevation difference value and an azimuth difference value.

[0108] It should be understood that the difference determiner can formulate the rotated derived audio direction vector 0 ′ based on the quantized azimuth angle φ of the first object in order to determine the audio difference vector.

[0109] For example, the audio orientation difference vector of an audio object P q with orientation components (θ q ) q can be found as:[[]]

[0110]

[0111] However, in practice, in some embodiments, Δθ q can be θ q because the elevation component of the above SP code vector can be zero. However, it should be understood that other embodiments can derive a vector SP in which the elevation component is non-zero, and in these embodiments, an equivalent rotational transformation can be applied to the elevation component of each element of the derived vector SP. That is, the elevation component of each element of the derived vector SP can be rotated (or aligned to) the elevation of the first audio object.

[0112] It should be understood that the orientation difference of the audio object P q is formed based on the difference between each element of the rotated derived audio direction vector and the corresponding re-ordered (or re-positioned) audio object.

[0113] It should also be understood that the above description is based on the order of re-positioning (or re-arranging) the audio objects, but the above description is equally valid for re-positioning only the audio orientation parameters rather than the entire audio object.

[0114] The step of forming the orientation difference between each re-positioned audio orientation parameter and the corresponding rotated derived direction parameter is shown as processing step 411 in Figure 4 .

[0115] Furthermore, the orientation difference vectors (Δθ q , Δφq ) is quantized.

[0116] In Figure 3a the spherical quantizer & indexer 209 is shown in more detail, where the orientation difference vector 210 is shown to be passed to the spherical quantizer 300 via the input 308.

[0117] The following section describes a suitable spherical quantization scheme for indexing the orientation difference vectors (Δθ q , Δφ q ) of each audio object.

[0118] Hereinafter, the input of the quantizer is generally referred to as (θ, φ) to simplify the nomenclature, and because this method can be used for any pair of elevation and azimuth angles.

[0119] In some embodiments, the directional spherical quantizer 300 includes a quantization input 302. This quantization input (which can also be referred to as an encoding input) is configured to define the granularity of a sphere arranged around a reference position or location, and the direction parameters are determined based on this reference position or location. In some embodiments, the quantization input is a predefined or fixed value. Additionally, in some embodiments, the quantization input 302 can define other aspects or inputs of the configuration enabling the spherical quantization operation. For example, in some embodiments, the quantization input 302 includes a reference direction (e.g., relative to an absolute direction such as magnetic north). In some embodiments, the reference direction is determined or defined based on the analysis of the input signal.

[0120] In some embodiments, the directional spherical quantizer 300 includes a sphere locator 303. This sphere locator is configured to configure the arrangement of the sphere based on the quantization input value. The proposed spherical grid uses the idea of covering a sphere with smaller spheres and treating the centers of the smaller spheres as points of a grid defining approximately equidistant directions.

[0121] The concept shown herein defines the sphere relative to a reference position and a reference direction. The sphere can be visualized as a series of circles (or intersection points), and for each circle intersection, there is a defined number of (smaller) spheres at the circumference of the circle. This is shown, for example, with respect to Figure 3c as shown. For example, Figure 3c shows an example "polarity" reference direction configuration, which shows a first main sphere 370 having a radius defined as the radius of the main sphere. In Figure 3c smaller spheres (shown as circles) 381, 391, 393, 395, 397, and 399 are also shown, which are positioned such that the circumference of each smaller sphere touches the main sphere circumference at one point and touches the circumference of at least another smaller sphere at at least one other point. Thus, as Figure 3cAs shown, the smaller sphere 381 contacts the main sphere 370 as well as the smaller spheres 391, 393, 395, 397, and 399. Additionally, the smaller sphere 381 is positioned such that the center of the smaller sphere lies on the + / −90 degree elevation line (z-axis) extending through the center of the main sphere 370.

[0122] The smaller spheres 391, 393, 395, 397, and 399 are positioned such that each of them contacts the main sphere 370, the smaller sphere 381, and another pair of adjacent smaller spheres. For example, the smaller sphere 391 additionally contacts the adjacent smaller spheres 399 and 393, the smaller sphere 393 additionally contacts the adjacent smaller spheres 391 and 395, the smaller sphere 395 additionally contacts the adjacent smaller spheres 393 and 397, the smaller sphere 397 additionally contacts the adjacent smaller spheres 399 and 391, and the smaller sphere 399 additionally contacts the adjacent smaller spheres 397 and 391.

[0123] Thus, the smaller sphere 381 defines a cone 380 or solid angle with respect to the +90 degree elevation line, and the smaller spheres 391, 393, 395, 397, and 399 define another cone 390 or solid angle with respect to the +90 degree elevation line, where the solid angle of the other cone is larger than that of the cone.

[0124] In other words, the smaller sphere 381 (which defines the first sphere circle) can be considered to be located at a first elevation angle (with the center of the smaller sphere at +90 degrees), and the smaller spheres 391, 393, 395, 397, and 399 (which define the second sphere circle) can be considered to be located at a second elevation angle with respect to the main sphere (with the center of the smaller sphere at <90 degrees) and the elevation angle is lower than that of the previous circle.

[0125] Furthermore, this arrangement can be further repeated with other circles that contact spheres located at other elevation angles with respect to the main sphere and having an elevation angle lower than that of the previous circle.

[0126] Thus, in some embodiments, the sphere locator 303 is configured to perform the following operations to define directions corresponding to the covering spheres:

[0127]

[0128] Thus, according to the above, the elevation angle of each point on circle i is given by the value in θ(i). For each circle above the equator, there is a corresponding circle below the equator (the plane defined by the X-Y axes).

[0129] In addition, as discussed above, each direction point on a circle can be indexed in ascending order with respect to the azimuth value. The index of the first point in each circle is given by an offset that can be derived from the number of points n(i) on each circle. To obtain these offsets, for the order of the circles under consideration, these offsets are calculated as the cumulative number of points on the circles for the given order, starting from the value 0 as the first offset.

[0130] In other words, the circles are arranged downward starting from the "North Pole".

[0131] In another embodiment, the number of points along a circle parallel to the equator can also be obtained by where λ i ≥1, λ i ≤λ i+1 .

[0132] In other words, the spheres along the circles parallel to the equator have a larger radius because they are farther from the North Pole, i.e., they are farther from the North Pole of the main direction.

[0133] Having determined a plurality of circles and the number of circles Nc, the number of points n(i) on each circle, i = 0, Nc - 1, and the sphere locator of the indexing order can be configured to pass this information to the EA to DI converter 305.

[0134] The conversion process from (elevation angle / azimuth angle) (EA) to direction index (DI) and vice versa is provided in the following paragraphs. Alternative orderings of the circles are considered herein.

[0135] The direction metadata encoder 300 includes an elevation angle - azimuth angle to direction index (EA - DI) converter 305. In some embodiments, the elevation angle - azimuth angle to direction index converter 305 is configured to receive a direction parameter input 108 and sphere locator information and convert the elevation angle - azimuth angle value from the direction parameter input 108 into a direction index by quantizing the elevation angle - azimuth angle value.

[0136] Regarding Figure 6 , an example method for generating a direction index according to some embodiments is shown.

[0137] The reception of the quantized input is shown by step 601 in Figure 6 .

[0138] Furthermore, as shown by step 603 in Figure 6 , the method can determine the sphere location based on this quantized input.

[0139] In addition, as shown by step 602 in Figure 6 , the method can include receiving direction parameters.

[0140] As Figure 6 shown in step 605, after the direction parameter and the sphere positioning information have been received, the method may include converting the direction parameter into a direction index based on the sphere positioning information.

[0141] Furthermore, as Figure 6 shown in step 607, the method may output a direction index.

[0142] In some embodiments, the elevation-azimuth to direction index (EA-DI) converter 305 is configured to perform this conversion according to the following algorithm.

[0143] Input: (θ, φ),

[0144] Output: I d

[0145] In some embodiments, S θ may take the form of an index codebook having N discrete entries, each entry θ l corresponding to an elevation value, l = 0:N-1. Additionally, for each discrete elevation value θ l , the codebook further includes a set of discrete azimuth values φ m , where the number of azimuth values in the set depends on the elevation θ l . In other words, for each elevation entry θ l , there may be a different number of discrete azimuth values φ m , j = 0:f(θ l ), where f(θ l ) represents that the number of azimuth values in the set associated with the elevation value θ l is a function of the elevation value θ l .

[0146] Regarding Figure 7 , shows an example method for converting elevation-azimuth to direction index (EA-DI) as shown by step 605 in Figure 6 .

[0147] First step: Quantifying the elevation-azimuth value may include scalar quantifying the elevation value θ by finding the closest codebook entry θ l to give a first quantized elevation value . The elevation value θ may be scalar quantized again by finding the next closest codebook entry. This may be given as either codebook entry θ l+1 or θ l-1 , depending on which is closer to θ, resulting in a second quantized elevation value

[0148] Processing step: The elevation angle θ value is scalar quantized to the closest indexed elevation angle value θ i and additionally scalar quantized to the next closest indexed elevation angle value θ l+1 or θ l-1 are respectively shown as processing steps 701 and 703. For each quantized elevation angle value and the corresponding scalar quantized azimuth angle value can be found. In other words, by finding the closest azimuth angle value from a set of azimuth angle values associated with the indexed elevation angle value θ of the first quantized elevation angle value l the first scalar quantized azimuth angle value corresponding to can be determined. The first scalar quantized azimuth angle value corresponding to the first quantized elevation angle value can be represented as Similarly, the second scalar quantized azimuth angle value corresponding to can also be determined and represented as This can be performed by re - quantizing the azimuth angle value φ, but this time using a set of azimuth angle values associated with the index of the second scalar quantized elevation angle value Processing step: Scalar quantize the azimuth angle value φ corresponding to the closest indexed elevation angle value θ l and additionally scalar quantize the azimuth angle value corresponding to the next closest indexed elevation angle value θ l+1 or θ l-1 are respectively shown as processing steps 705 and 707. Once the first elevation - azimuth scalar quantized value pair and the second elevation - azimuth scalar quantized value pair have been determined, a distance metric on the unit sphere can be calculated for each pair. The distance metric can be considered by taking the L2 - norm distance between two points on the unit sphere. Thus, for the first scalar quantized elevation - azimuth pair the distance d is calculated as the distance between the first scalar quantized elevation - azimuth pair on the unit sphere and the unquantized elevation - azimuth pair (θ, φ). Similarly, for the second scalar quantized elevation - azimuth pair the distance d′ is calculated as the distance between the second scalar quantized elevation - azimuth pair on the unit sphere and the unquantized elevation - azimuth pair (θ, φ).

[0149] It should be understood that in the embodiment, it can be based on ‖x - y‖ 2Consider the L2-norm distance between two points x and y on the unit sphere, where x and y are spherical coordinates in three-dimensional space. In terms of the elevation-azimuth pair (θ, φ), the spherical coordinates can be represented as x = (rcos(θ)cos(φ), rcos(θ)sin(φ), rsin(θ)), and for the elevation-azimuth pair, the spherical coordinates correspond to By considering the unit sphere, the radius r = 1, and the distance d can be simplified to calculate where it can be seen that the distance d depends only on the values of the angles.

[0150] Similarly, the distance d′ between the second scalar-quantized elevation-azimuth pair on the unit sphere and the unquantized elevation-azimuth pair (θ, φ) can be expressed as Processing step: Find the distance between the first scalar-quantized elevation-azimuth pair and the unquantized elevation-azimuth pair (θ, φ) is shown as Figure 7 709 in. Processing step: Find the distance between the second scalar-quantized elevation-azimuth pair and the unquantized elevation-azimuth pair (θ, φ) is shown as Figure 7 711 in.

[0151] Finally, select the scalar-quantized elevation-azimuth pair with the minimum distance metric as the quantized elevation-azimuth value of the elevation-azimuth (θ, φ). Further, the corresponding index associated with the selected quantized elevation and azimuth pair continues to form the direction index I d . Processing step: Find the minimum distance is shown as Figure 7 713 in.

[0152] Processing step: Select between the index with the quantized elevation-azimuth θ, φ according to the minimum distance in and the index of is shown as Figure 7 715 in.

[0153] It should be understood that even though the above spherical quantization scheme has been defined based on the unit sphere, other embodiments may deploy the above quantization scheme based on a general sphere whose radius is not equal to 1. In such embodiments, the above step of finding the minimum distance still holds because the minimum distance calculations corresponding to both the first scalar-quantized elevation-azimuth pair and the second scalar-quantized elevation-azimuth pair are independent of or the radius r.

[0154] In another embodiment, a scalar quantizer may also be used to quantize the elevation and azimuth angles. Regardless of whether the spherical grid or the scalar quantizer is used to quantize the azimuth and elevation angles, the indices generated by the azimuth and elevation angles can be used for encoding and transmission in place of the direction index. This can be particularly useful in instances where bit consumption is to be saved.

[0155] The direction index I can be output d 306.

[0156] Return Figure 4 , and the entire step of quantizing the audio directional difference is shown depicted as processing step 413.

[0157] Regarding Figure 5 , an audio object decoder according to Figure 1 141 is shown. It can be seen that the audio object decoder 141 can be configured to receive the direction index I from the encoded bitstream d , K bits for scalar quantizing the azimuth angle φ of the first object 0 (referred to as ), and the index I representing the index order of the audio direction parameters of audio objects 1 to N - 1 ro . Within the audio object decoder 141, the direction index I d can be passed to the spherical deindexer 511. In this regard, an example spherical deindexer 511 is shown in Figure 3b , which can be used to decode the direction data index I associated with the audio object d and generate a quantized directional difference vector.

[0158] Associated with Figure 5 is Figure 8 , Figure 8 depicting the processing steps of the audio object decoder 141.

[0159] Here, the naming of the audio difference direction vector is restored to (Δθ q , Δφ q ), and the quantized audio difference direction vector will be referred to as (Δθ′ q , Δφ′ q ).

[0160] The spherical deindexer 511 may include a quantization input 352. In some embodiments, this is passed from the encoder or otherwise agreed upon with the encoder. The quantization input is configured to define the granularity of a sphere arranged around a reference position or location. Additionally, in some embodiments, the quantization input defines the configuration of the sphere, e.g., the orientation of a reference direction (relative to an absolute direction such as magnetic north).

[0161] The spherical deindexer 511 may include a direction index input 351. This may be received from an encoder or obtained by any suitable means.

[0162] In some embodiments, the spherical deindexer 511 includes a sphere locator 353. The sphere locator 353 is configured to receive a quantization input as an input and generate a sphere arrangement in the same manner as generated in the encoder. Further, this sphere arrangement is used to generate a codebook as previously described for generating dequantized elevation and azimuth values.

[0163] The spherical deindexer 511 includes a direction index to elevation-azimuth (DI-EA) converter 355. The direction index to elevation-azimuth converter 355 is configured to receive a direction index and, in addition, a spherical codebook as generated by the sphere arrangement. Further, by referring to the index of the spherical codebook and obtaining the corresponding quantized elevation and azimuth values, the index to elevation-azimuth converter 355 converts the direction index into quantized elevation and azimuth values.

[0164] As previously discussed, the above spherical quantization scheme may be set to quantize and index the directional difference vectors corresponding to the audio object P q As previously mentioned, the quantization and dequantization of the directional difference vectors are performed sequentially for each audio object on a per-direction vector basis. Thus, the final result of the spherical dequantization process may be N quantized directional difference vectors (Δθ′ q , Δφ′ q ), each corresponding to the audio object P q q = 0:N-1.

[0165] The step of dequantizing the audio directional difference between each repositioned audio direction parameter and the corresponding rotated derived audio direction parameter is depicted as processing step 801 in Figure 8 .

[0166] Additionally, Figure 5 shows the index for the K bits used to scalar quantize the azimuth φ of the first object 0 which are passed to the dequantizer 505 to produce the dequantized first object azimuth φ′ 0 . The step of dequantizing the azimuth value of the first audio object is shown as Figure 8 processing step 803 in

[0167] The audio object decoder 141 may include an audio direction extractor 501, which has the same functions as the audio direction extractor 201 at the encoder 121. In other words, the audio direction extractor 501 may be configured to form and initialize the SP vectors in the same manner as performed at the encoder. That is to say, each derived audio direction component of the SP vector is formed on the premise that the orientation information of the audio object can be initialized to a series of points evenly distributed along the circumference of the unit circle, starting from the azimuth value 0 0 and then. Furthermore, the SP vector containing the derived audio directions may be passed to the audio direction rotator 503.

[0168] Reference Figure 8 , the step of initializing the derived directions associated with each audio object is shown as processing step 807.

[0169] The audio direction rotator 503 may also be configured to accept the dequantized azimuth φ′ of the first object 0 as another input. The audio direction rotator 503 may use the dequantized azimuth value of the first object to reform the rotated derived audio direction "template" vector for the N-1 audio object directions through the following calculations

[0170]

[0171] That is to say, by adding the dequantized azimuth φ′ of the first object to each derived audio direction component of the SP vector to form the rotated derived vector 0 Reference

[0172] , processing step 807 represents rotating each derived direction by the dequantized azimuth value of the first audio object. Figure 8

[0173] Furthermore, the rotated derived audio directions may be passed to the adder 507.

[0174] After the N directional difference vector indices associated with the N audio objects have been decoded in processing step 801, the quantized audio direction difference vectors (Δθ′ q ,Δφ′ q ,Δφ′ q ) corresponding to q = 0:N-1 for the audio object P q

[0175] The adder 507 may be configured to, for each audio object P q q = 0:N-1, add the quantized orientation vectors (Δθ′ q ,Δφ′ q) with the corresponding rotated derived audio direction (from the rotated derived audio direction “template” vector from dequantization ) are added to form a quantized orientation vector for each audio object. This can be expressed as:

[0176]

[0177] For those embodiments where the rotation is based only on the azimuth value, i.e., for each element of the “template” code vector SP, the elevation component is 0, the above equation simplifies to:

[0178]

[0179] Processing step: for each audio object P q q = 0:N-1 the quantized orientation vector (Δθ′ q , Δφ′ q ) is added to the corresponding rotated derived audio direction as shown in Figure 8 as step 809.

[0180] Return Figure 5 , the index I representing the index order of the audio direction parameters for audio objects 1 to N-1 ro is shown to be received by the audio direction deindexer and relocator 509. Additionally, the audio direction deindexer and relocator 509 can also be configured to receive N quantized audio direction vectors from the adder 507.

[0181] In an embodiment, the audio direction deindexer and relocator 509 can be configured to decode the index I ro in order to find a specific index permutation of the reordered audio directions. Further, the audio direction deindexer and relocator 509 can use this index permutation to reorder the audio direction parameters back to their original order as presented to the audio object encoder 121 for the first time. Thus, the output from the audio direction deindexer and relocator 509 can be the ordered quantized audio directions associated with the N audio objects. Further, these ordered quantized audio parameters can form part of the decoded plurality of audio object streams 140.

[0182] The step of deindexing the positions of all audio object direction parameters except for the first audio object direction parameter is shown as Figure 8 processing step 811 in.

[0183] The step of arranging the positions of the audio object direction parameters to have their original order as received at the encoder is shown as Figure 8 processing step 813 in.

[0184] Regarding FIG. 10, an example electronic device that can be used as an analysis or synthesis device is shown. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.

[0185] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. The processor 1407 can be configured to execute various program codes such as the methods described herein.

[0186] In some embodiments, device 1400 includes a memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any suitable storage component. In some embodiments, the memory 1411 includes a program code portion for storing program codes that can be implemented on the processor 1407. Additionally, in some embodiments, the memory 1411 may also include a stored data portion for storing data (e.g., data that has been processed or is to be processed according to the embodiments described herein). Whenever needed, the processor 1407 can obtain the implementation program codes stored in the program code portion and the data stored in the stored data portion via the memory-processor coupling.

[0187] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, the user interface 1405 can be coupled to the processor 1407. In some embodiments, the processor 1407 can control the operation of the user interface 1405 and receive inputs from the user interface 1405. In some embodiments, the user interface 1405 can enable a user to input commands into the device 1400, for example, via a keyboard. In some embodiments, the user interface 1405 can enable a user to obtain information from the device 1400. For example, the user interface 1405 can include a display configured to display information from the device 1400 to the user. In some embodiments, the user interface 1405 can include a touch screen or a touch interface, which can enable information to be input into the device 1400 and also display information to the user of the device 1400. In some embodiments, the user interface 1405 can be a user interface for communicating with a position determiner as described herein.

[0188] In some embodiments, device 1400 includes an input / output port 1409. In some embodiments, the input / output port 1409 includes a transceiver. In such an embodiment, the transceiver can be coupled to the processor 1407 and is configured to enable communication with other devices or electronic apparatuses, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver device can be configured to communicate with other electronic devices or apparatuses via a wired or wired coupling.

[0189] The transceiver can communicate with other devices via any suitable known communication protocol. For example, in some embodiments, the transceiver or transceiver components can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an infrared data communication path (IRDA).

[0190] The transceiver input / output port 1409 can be configured to receive signals and, in some embodiments, determine the parameters as described herein by using a processor 1407 that executes suitable code. Additionally, the device can generate a suitable down-converted signal and parameter output for transmission to a synthesis device.

[0191] In some embodiments, the device 1400 can be part of at least a synthesis device. Thus, the input / output port 1409 can be configured to receive a down-converted signal and, in some embodiments, receive the parameters determined at a capture device or a processing device as described herein, and generate a suitable audio signal format output by using a processor 1407 that executes suitable code. The input / output port 1409 can be coupled to any suitable audio output, such as being coupled to a multi-channel speaker system and / or headphones or the like.

[0192] Generally, various embodiments of the present invention can be implemented using hardware or a dedicated circuit, software, logic, or any combination thereof. For example, some aspects can be implemented using hardware, while other aspects can be implemented using firmware or software executable by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although the various aspects of the present invention can be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it is well known that the blocks, devices, systems, techniques, or methods described herein can be implemented as non-limiting examples using hardware, software, firmware, a dedicated circuit or logic, general hardware or a controller or other computing devices, or some combination thereof.

[0193] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. Additionally, in this regard, it should be noted that any block of the logical flow in the drawings can represent a program step, or interconnected logical circuits, blocks, and functions, or a combination of program steps and logical circuits, blocks, and functions. The software can be stored on a physical medium such as a memory chip or a memory block implemented within a processor, a magnetic medium such as a hard disk or a floppy disk, and an optical medium such as a DVD and its data variant CD.

[0194] The memory can be of any type suitable for the local technical environment and can be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and, by way of non-limiting example, can include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.

[0195] Embodiments of the present invention can be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Sophisticated and powerful software tools can be used to transform a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0196] The program can automatically route conductors and position components on a semiconductor chip using well-established design rules and a library of pre-stored design modules. Once the design of the semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be transferred to a semiconductor manufacturing facility or “fab” for fabrication.

[0197] The foregoing description has provided a complete and beneficial description of exemplary embodiments of the present invention by way of example and not limitation. However, various modifications and adaptations will become apparent to those skilled in the relevant art in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of the present invention will still fall within the scope of the present invention as defined by the appended claims.

Claims

1. A method for encoding spatial audio signals, comprising: For each of a plurality of audio direction parameters, where each parameter includes an elevation value and an azimuth value and each parameter has an ordered position, deriving a corresponding derived audio direction parameter, the corresponding derived audio direction parameter including an elevation value and an azimuth value; Rotating each derived audio direction parameter by the azimuth value of the audio direction parameter at the first position among the plurality of audio direction parameters; When the azimuth value of an audio direction parameter is closest to the azimuth value of another rotated derived audio direction parameter compared to the azimuth values of other rotated derived audio direction parameters, changing the ordered position of the audio direction parameter to another position that is consistent with the position of the rotated derived audio direction parameter, and then for each of the plurality of audio direction parameters, determining the difference between each audio direction parameter and the corresponding rotated derived audio direction parameter; and For each of the plurality of audio direction parameters, quantifying the difference.

2. The method according to claim 1, wherein, The azimuth value of each derived audio direction parameter corresponds to a position among a plurality of positions around the circumference of a circle.

3. The method according to claim 2, wherein, The plurality of positions around the circumference of the circle are evenly distributed along the 360 degrees of the circle, and wherein the number of positions around the circumference of the circle is determined by the number of audio direction parameters.

4. The method according to any one of claims 1 to 3, wherein, Rotating each derived audio direction parameter by the azimuth value of the first audio direction parameter among the plurality of audio direction parameters includes: Adding the azimuth value of the first audio direction parameter to the azimuth value of each derived audio direction parameter, where the elevation value of each derived audio direction parameter is set to zero.

5. The method according to claim 4, further comprising: Scalar quantifying the azimuth value of the first audio direction parameter; and Indexing the position of the audio direction parameter after the change by assigning an index arranged in the order representing the position of the audio direction parameter.

6. The method according to any one of claims 1 to 3, wherein, Determining the difference between each audio direction parameter and the corresponding rotated derived audio direction parameter for each of the plurality of audio direction parameters includes: for each of the plurality of audio direction parameters, determining a differential audio direction parameter based on at least the following operations: Determining the difference between the first located audio direction parameter and the first located rotated derived audio direction parameter, and / or Determining the difference between another audio direction parameter and the rotated derived audio direction parameter, where the position of the another audio direction parameter has not changed, and / or Determining the difference between yet another audio direction parameter and the rotated derived audio direction parameter, where the position of the yet another audio direction parameter has been changed to the position of the rotated derived audio direction parameter.

7. The method according to any one of claims 1 to 3, wherein, determining the difference between the audio direction parameter and the corresponding rotated derived audio direction parameter includes: determining the difference between the azimuth value of the audio direction parameter and the azimuth value of the corresponding rotated derived audio direction parameter; and determining the difference between the elevation value of the audio direction parameter and the elevation value of the corresponding rotated derived audio direction parameter.

8. The method according to any one of claims 1 to 3, wherein, changing the position of the audio direction parameter to another position is applicable to any audio direction parameter other than the audio direction parameter at the first position.

9. The method according to claim 6, wherein, quantifying the differential audio direction parameter for each of the plurality of audio direction parameters includes: for each of the plurality of audio direction parameters, quantifying the differential audio direction parameter as a vector, wherein the vector is indexed to a codebook, the codebook including a plurality of indexed elevation values and indexed azimuth values.

10. The method according to claim 9, wherein, the plurality of indexed elevation values and indexed azimuth values are points on a spherical grid arranged in the form of a sphere, wherein the spherical grid is formed by covering the sphere with smaller spheres, and wherein the smaller spheres define the points of the spherical grid.

11. A method for decoding a spatial audio signal, comprising: decoding an index to provide a quantized azimuth value of an audio direction parameter at a first position among a plurality of ordered audio direction parameters, wherein each parameter includes an elevation value and an azimuth value; for each of the plurality of audio direction parameters, deriving a corresponding derived audio direction parameter, the corresponding derived audio direction parameter including an elevation value and an azimuth value; rotating each derived audio direction parameter by the azimuth value of the audio direction parameter at the first position among the plurality of audio direction parameters; decoding an index to provide a quantized difference between the audio direction parameter and its corresponding derived audio direction parameter for each audio direction parameter; for each audio direction parameter, forming a quantized audio direction parameter by adding the quantized difference to its corresponding derived audio direction parameter; and decoding an index representing the order of a plurality of quantized audio direction parameters and reordering the positions of the plurality of quantized audio direction parameters according to the order.

12. The method according to claim 11, wherein, the azimuth value of each derived audio direction parameter corresponds to a position among a plurality of positions around the circumference of a circle.

13. The method according to claim 12, wherein, the plurality of positions around the circumference of the circle are evenly distributed along the 360 degrees of the circle, and wherein the number of positions around the circumference of the circle is determined by the number of audio direction parameters.

14. The method according to any one of claims 11 to 13, wherein, Rotating each of the derived audio direction parameters by the azimuth value of the first audio direction parameter among the plurality of audio direction parameters includes: Adding the quantized azimuth value of the first audio direction parameter to the azimuth value of each of the derived audio direction parameters, wherein the elevation value of each of the derived audio direction parameters is set to zero.

15. The method according to any one of claims 11 to 13, wherein, The index providing the quantization difference between the audio direction parameter and its corresponding derived audio direction parameter for each audio direction parameter is an index of a codebook, the codebook including a plurality of indexed elevation values and indexed azimuth values.

16. The method according to claim 15, wherein, The plurality of indexed elevation values and indexed azimuth values are points on a spherical grid arranged in the form of a sphere, wherein the spherical grid is formed by covering the sphere with smaller spheres, and wherein the smaller spheres define the points of the spherical grid.

17. An apparatus for spatial audio signal coding, comprising: For each audio direction parameter among a plurality of audio direction parameters, wherein each parameter includes an elevation value and an azimuth value and each parameter has an ordered position, deriving a corresponding derived audio direction parameter, the corresponding derived audio direction parameter including an elevation value and an azimuth value; Rotating each of the derived audio direction parameters by the azimuth value of the audio direction parameter in the first position among the plurality of audio direction parameters; When the azimuth value of an audio direction parameter is closest to the azimuth value of another rotated derived audio direction parameter compared to the azimuth values of other rotated derived audio direction parameters, changing the ordered position of the audio direction parameter to another position consistent with the position of the rotated derived audio direction parameter, and then the apparatus is configured to: for each audio direction parameter among the plurality of audio direction parameters, determine the difference between each audio direction parameter and the corresponding rotated derived audio direction parameter; and For each audio direction parameter among the plurality of audio direction parameters, quantizing the difference.

18. The apparatus according to claim 17, wherein, The azimuth value of each derived audio direction parameter corresponds to a position among a plurality of positions around the circumference of a circle.

19. According to the apparatus of claim 18, the plurality of positions around the circumference of the circle are evenly distributed along the 360 degrees of the circle, and wherein, The number of positions around the circumference of the circle is determined by the number of audio direction parameters.

20. The apparatus according to any one of claims 17 to 19, wherein, The apparatus configured to rotate each of the derived audio direction parameters by the azimuth value of the first audio direction parameter among the plurality of audio direction parameters is configured to: Adding the azimuth value of the first audio direction parameter to the azimuth value of each of the derived audio direction parameters, wherein the elevation value of each of the derived audio direction parameters is set to zero.

21. The apparatus according to claim 20, further configured to: Scalar quantize the azimuth value of the first audio direction parameter; and Index the positions of the audio direction parameters after the change by assigning indices that order the permutation of the indices representing the positions of the audio direction parameters.

22. The apparatus according to any one of claims 17 to 19, wherein, the apparatus configured to determine the difference between each audio direction parameter of the plurality of audio direction parameters and a corresponding rotated derived audio direction parameter is configured to, for each audio direction parameter of the plurality of audio direction parameters, determine a differential audio direction parameter based at least on the following operations: determine the difference between a first-positioned audio direction parameter and a first-positioned rotated derived audio direction parameter, and / or determine the difference between another audio direction parameter and a rotated derived audio direction parameter, wherein the position of the another audio direction parameter has not changed, and / or determine the difference between yet another audio direction parameter and a rotated derived audio direction parameter, wherein the position of the yet another audio direction parameter has been changed to the position of the rotated derived audio direction parameter.

23. The apparatus according to any one of claims 17 to 19, wherein, the apparatus configured to determine the difference between an audio direction parameter and a corresponding rotated derived audio direction parameter is configured to: determine the difference between the azimuth value of the audio direction parameter and the azimuth value of the corresponding rotated derived audio direction parameter; and determine the difference between the elevation value of the audio direction parameter and the elevation value of the corresponding rotated derived audio direction parameter.

24. The apparatus according to any one of claims 17 to 19, wherein, the apparatus configured to change the position of an audio direction parameter to another position is applicable to any audio direction parameter other than the first-positioned audio direction parameter.

25. The apparatus according to claim 22, wherein, the apparatus configured to quantify the differential audio direction parameter for each audio direction parameter of the plurality of audio direction parameters is configured to: for each audio direction parameter of the plurality of audio direction parameters, quantify the differential audio direction parameter as a vector, wherein the vector is indexed to a codebook that includes a plurality of indexed elevation values and indexed azimuth values.

26. The apparatus according to claim 25, wherein, the plurality of indexed elevation values and indexed azimuth values are points on a spherical grid arranged in the form of a sphere, wherein the spherical grid is formed by covering the sphere with smaller spheres, and wherein the smaller spheres define the points of the spherical grid.

27. An apparatus for spatial audio signal decoding, configured to: decode an index to provide a quantified azimuth value of an audio direction parameter at a first position of a plurality of ordered audio direction parameters, wherein, each parameter includes an elevation value and an azimuth value; for each audio direction parameter of the plurality of audio direction parameters, derive a corresponding derived audio direction parameter, the corresponding derived audio direction parameter including an elevation value and an azimuth value; Rotate the azimuth value of the audio direction parameter at the first position among the plurality of audio direction parameters for each exported audio direction parameter; Decode an index to provide a quantization difference between an audio direction parameter and its corresponding exported audio direction parameter for each audio direction parameter; For each audio direction parameter, form a quantized audio direction parameter by adding the quantization difference to its corresponding exported audio direction parameter; and Decode an index representing the order of the plurality of quantized audio direction parameters and reorder the positions of the plurality of quantized audio direction parameters according to the order.

28. The apparatus according to claim 27, wherein, The azimuth value of each exported audio direction parameter corresponds to a position among a plurality of positions around the circumference of a circle.

29. The apparatus according to claim 28, wherein, The plurality of positions around the circumference of the circle are evenly distributed along 360 degrees of the circle, and wherein the number of positions around the circumference of the circle is determined by the number of audio direction parameters.

30. The apparatus according to any one of claims 27 to 29, wherein, The apparatus configured to rotate the azimuth value of the first audio direction parameter among the plurality of audio direction parameters for each exported audio direction parameter is configured to: Add the quantized azimuth value of the first audio direction parameter to the azimuth value of each exported audio direction parameter, wherein the elevation value of each exported audio direction parameter is set to zero.

31. The apparatus according to any one of claims 27 to 29, wherein, The index that provides the quantization difference between an audio direction parameter and its corresponding exported audio direction parameter for each audio direction parameter is an index of a codebook, and the codebook includes a plurality of indexed elevation values and indexed azimuth values.

32. The apparatus according to claim 31, wherein, The plurality of indexed elevation values and indexed azimuth values are points on a spherical grid arranged in the form of a sphere, wherein the spherical grid is formed by covering the sphere with smaller spheres, and wherein the smaller spheres define the points of the spherical grid.

Citation Information

Patent Citations

  • Quantization of spatial audio direction parameters

    CN114207713A