Determining frequency sub-bands for spatial audio parameters
By adjusting the frequency subband and merging/removing unnecessary sets of spatial audio parameters, the bandwidth mismatch problem in the prior art when transmitting audio streams and analyzing MASA metadata streams is solved, and more efficient encoding and transmission efficiency is achieved.
Patent Information
- Application Number
- CN202280101964.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art leads to unnecessary encoding and bandwidth mismatch when transmitting audio streams and analyzing MASA metadata streams, especially when the subbands of the audio stream are minimal to the overall synthesis.
By determining the set of spatial audio parameters and adjusting the number and boundaries of the frequency subbands based on the encoding rate and bandwidth values, unwanted sets of spatial audio parameters are merged or removed to match the bandwidth of the audio signal.
It effectively reduces unnecessary encoding bit overhead, improves bandwidth utilization, and optimizes the transmission efficiency of audio signals.
Smart Images

Figure CN120226076A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to apparatus and methods for altering the bandwidth of a spatial audio signal. Background Art
[0002] Immersive audio codecs are being implemented to support a wide range of operating points from low bitrate operation to transparency. An example of such a codec is the Immersive Voice and Audio Service (IVAS) codec, which is designed to be suitable for use on communication networks such as 3GPP 4G / 5G networks, including use in immersive services such as immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of voice, music, and general audio. In addition, it is expected to support channel-based audio and scene-based audio inputs, including spatial information about the sound field and sound sources. It is also expected that the codec operate with low latency to enable session services, and support high error robustness under various transmission conditions.
[0003] Metadata-Assisted Spatial Audio (MASA) is an input format for IVAS. It uses an audio signal along with corresponding spatial metadata. This spatial metadata includes parameters that define the spatial aspects of the audio signal, and it can include, for example, direction and direct-to-total energy ratio in a frequency band. A MASA stream can be obtained, for example, by capturing spatial audio using the microphones of a suitable capture device. For example, a mobile device including multiple microphones can be configured to capture microphone signals, where a set of spatial metadata can be estimated based on the captured microphone signals. A MASA stream can also be obtained from other sources (such as specific spatial audio microphones (such as Ambisonics or array-microphones), studio mixes (e.g., 5.1 audio channel mixes)) or other content through appropriate format conversion.
[0004] An audio signal input to an immersive voice codec (such as IVAS) can be simultaneously encoded as 1 - N audio signals to give a transmission audio stream, and analyzed to give a MASA metadata stream. In such a setup, the analysis and encoding for the MASA metadata stream can be performed separately from the encoding for the transmission audio stream. This can lead to unnecessary encoding of some sets of MASA metadata. In particular, for subbands of the transmission audio stream that contribute minimally to the overall synthesized spatial audio signal. In addition, this can result in a mismatch between the bandwidth of the input audio signal encoded for the transmission audio stream and the bandwidth of the input audio signal analyzed for the MASA metadata stream. Summary of the Invention
[0005] According to a first aspect, an apparatus for spatial audio coding, the apparatus comprising components configured to perform the following operations: for each frequency subband of a plurality of frequency subbands of one or more audio signals, determine a set of spatial audio parameters; receive a coding rate associated with the one or more audio signals; based on the coding rate, map at least two consecutive subbands of the plurality of frequency subbands to an extended frequency subband to give a plurality of frequency subbands adjusted by the coding rate; receive a bandwidth value associated with the one or more audio signals; starting from the highest frequency subband of the plurality of frequency subbands adjusted by the coding rate, remove a certain number of frequency subbands to give a plurality of frequency subbands adjusted by the bandwidth, wherein the number of removed frequency subbands is based on the bandwidth value associated with the one or more audio signals; under the condition that the highest frequency subband of the plurality of frequency subbands adjusted by the bandwidth exceeds the bandwidth value associated with the one or more audio signals, reduce the highest frequency subband of the plurality of frequency subbands adjusted by the bandwidth to the bandwidth value or below the bandwidth value; merge the set of spatial audio parameters associated with the first frequency subband of the at least two consecutive frequency subbands and the set of spatial audio parameters associated with the second frequency subband of the at least two consecutive frequency subbands to give a merged set of spatial audio parameters for the extended frequency subband; remove the set of spatial audio parameters corresponding to each removed frequency subband; and under the condition that the highest frequency subband of the plurality of frequency subbands adjusted by the bandwidth exceeds the bandwidth value, remove the set of spatial audio parameters associated with the plurality of frequency subbands adjusted by the bandwidth that exceed the bandwidth value.
[0006] The highest frequency subband of the plurality of frequency subbands adjusted by the bandwidth may include the upper subband boundary value and the lower subband boundary value of a plurality of frequency subbands of the plurality of frequency subbands including one or more audio signals, and the component configured to reduce the highest frequency subband of the plurality of frequency subbands adjusted by the bandwidth to the bandwidth value or below the bandwidth value may include a component configured to perform the following operations: adjust the upper subband boundary value within the bandwidth value; and wherein the component configured to remove the set of spatial audio parameters associated with the plurality of frequency subbands adjusted by the bandwidth that exceed the bandwidth value includes a component configured to perform the following operations: remove the set of spatial audio parameters associated with the plurality of frequency subbands of the one or more audio signals that are higher than the adjusted upper subband boundary value.
[0007] The component configured to map at least two consecutive subbands of the plurality of frequency subbands to an extended frequency subband based on the coding rate to give a plurality of frequency subbands adjusted by the coding rate may include a component configured to perform the following operations: map the high band boundary value and the low band boundary value of the at least two consecutive frequency subbands of the plurality of frequency subbands to the low band boundary value and the high band boundary value of the extended frequency subband.
[0008] The low-frequency sub-band boundary value and the high-frequency sub-band boundary value of the extended frequency sub-bands can be given by the low-frequency sub-band boundary value and the high-frequency sub-band boundary value of a frequency sub-band reduction array, which includes a plurality of frequency sub-band boundaries in ascending order of frequency sub-bands. Among them, a sub-band boundary value in ascending order and the next higher sub-band boundary value in the frequency sub-band reduction array are respectively the low-frequency sub-band boundary value and the high-frequency sub-band boundary value of the extended frequency sub-bands.
[0009] The plurality of frequency sub-band boundaries in the frequency sub-band reduction array can form fewer frequency sub-bands than the plurality of frequency sub-bands of one or more audio signals, and the plurality of frequency sub-bands with adjusted coding rates can be given by the frequency sub-band reduction array. The frequency sub-band reduction array can be selected from a plurality of frequency sub-band reduction arrays. The selection can be based on the coding rate associated with the one or more audio signals. Each frequency sub-band reduction array in the plurality of frequency sub-band reduction arrays can include a different number of frequency sub-bands, and each frequency sub-band reduction array in the plurality of frequency sub-band reduction arrays can be associated with a different coding rate associated with the one or more audio signals.
[0010] A certain number of frequency sub-bands to be removed can be selected from a plurality of certain numbers of frequency sub-bands to be removed. Among them, the selection can be based on the bandwidth value, and each of the plurality of certain numbers of frequency sub-bands to be removed can be associated with a different bandwidth value.
[0011] The plurality of frequency sub-bands with adjusted sampling frequencies can be in the form of an array, which includes a plurality of frequency sub-band boundary values in ascending order of frequency sub-bands.
[0012] The device can include a first encoder and a second encoder for encoding one or more audio signals at a coding rate. The coding rate can include the sum of the coding rate for the first encoder and the coding rate for the second encoder. The first encoder can encode an audio transmission signal associated with the one or more audio signals, and the second encoder can encode a plurality of sets of spatial audio parameters associated with the frequency sub-bands of the one or more audio signals.
[0013] According to a second aspect, an apparatus for spatially audio encoding one or more audio signals, wherein the apparatus comprises means for performing the following operations: for each frequency sub-band of a plurality of frequency sub-bands of the one or more audio signals, determining a set of spatial audio parameters; receiving a coding rate associated with the one or more audio signals; based on the coding rate, mapping at least two consecutive sub-bands of the plurality of frequency sub-bands to an extended frequency sub-band to give a plurality of frequency sub-bands adjusted by the coding rate; merging the set of spatial audio parameters associated with the first frequency sub-band of the at least two consecutive frequency sub-bands and the set of spatial audio parameters associated with the second frequency sub-band of the at least two consecutive frequency sub-bands to give a merged set of spatial audio parameters for the extended frequency sub-band; for each frequency bin of the one or more audio signals, determining an energy level; determining a cut-off frequency sub-band of the one or more audio signals by: determining the highest frequency bin having an energy level greater than a predetermined energy level, and designating the cut-off frequency sub-band as the frequency sub-band containing the highest frequency bin; comparing the cut-off frequency sub-band of the one or more audio signals with a bandwidth value of the one or more audio signals; under the condition that the cut-off frequency sub-band is less than the bandwidth value of the one or more audio signals, starting from the highest frequency sub-band of the plurality of frequency sub-bands adjusted by the coding rate, removing a certain number of frequency sub-bands to give a plurality of frequency sub-bands adjusted by the bandwidth, and removing the set of spatial audio parameters corresponding to each removed frequency sub-band, wherein the number of removed frequency sub-bands is based on the cut-off frequency sub-band; under the condition that the highest frequency sub-band of the plurality of frequency sub-bands adjusted by the bandwidth exceeds the cut-off frequency sub-band, reducing the highest frequency sub-band of the plurality of frequency sub-bands adjusted by the bandwidth to the cut-off frequency sub-band value or below the cut-off frequency sub-band value, and removing the set of spatial audio parameters associated with the plurality of frequency sub-bands adjusted by the bandwidth that exceed the cut-off frequency sub-band; and encoding an index of the cut-off frequency sub-band.
[0014] The means configured to encode an index of the cut-off frequency sub-band may be further configured to: encode each set of spatial audio parameters associated with a frequency sub-band lower than the cut-off frequency sub-band.
[0015] A component configured to encode each set of spatial audio parameters associated with a frequency subband below a cut-off frequency subband may further include a component configured to perform the following operations: for each frequency subband among a plurality of frequency subbands of one or more audio signals, determine an energy ratio parameter; quantize the energy ratio for each frequency subband among the plurality of frequency subbands that is greater than or equal to the cut-off frequency subband to a minimum quantization level; quantize the energy ratio for each frequency subband among the plurality of frequency subbands that is less than the cut-off frequency subband; and encode an indication that the number of encoded sets of spatial audio parameters is less than the number of frequency subbands of the one or more audio signals; and encode the plurality of unencoded sets of spatial audio parameters using a Golomb Rice code.
[0016] The highest frequency subband among the plurality of frequency subbands with bandwidth adjustment may include the upper subband boundary value and the lower subband boundary value of the plurality of frequency subbands among the plurality of frequency subbands including one or more audio signals, and the component configured to reduce the highest frequency subband among the plurality of frequency subbands with bandwidth adjustment to the cut-off frequency subband value or below the cut-off frequency subband value, and remove the set of spatial audio parameters associated with the plurality of frequency subbands with bandwidth adjustment that exceed the cut-off frequency subband may include a component configured to perform the following operations: adjust the upper subband boundary value within the cut-off frequency subband value; and remove the set of spatial audio parameters associated with the plurality of frequency subbands among the plurality of frequency subbands of the one or more audio signals that are higher than the adjusted upper subband boundary value.
[0017] The apparatus including a component configured to map at least two consecutive subbands among the plurality of frequency subbands to an extended frequency subband based on a coding rate to give a plurality of frequency subbands with coding rate adjustment may include a component configured to perform the following operations: map the high band boundary value and the low band boundary value for at least two consecutive frequency subbands among the plurality of frequency subbands to the low band boundary value and the high band boundary value of the extended frequency subband.
[0018] The low subband boundary value and the high subband boundary value of the extended frequency subband are given by the low subband boundary value and the high subband boundary value of a frequency subband reduction array, the frequency subband reduction array including a plurality of frequency subband boundaries in ascending order of frequency subbands, wherein one subband boundary value and the next higher subband boundary value in ascending order in the frequency subband reduction array are the low subband boundary value and the high subband boundary value of the extended frequency subband respectively.
[0019] Multiple frequency sub-band boundaries in a frequency sub-band reduction array may constitute fewer frequency sub-bands than the multiple frequency sub-bands of one or more audio signals, and a plurality of frequency sub-bands with adjusted coding rates may be given by the frequency sub-band reduction array, which may be selected from a plurality of frequency sub-band reduction arrays, wherein the selection may be based on the coding rate associated with the one or more audio signals, and wherein each frequency sub-band reduction array in the plurality of frequency sub-band reduction arrays may include a different number of frequency sub-bands, and each frequency sub-band reduction array in the plurality of frequency sub-band reduction arrays may be associated with a different coding rate associated with the one or more audio signals.
[0020] Multiple frequency sub-bands with adjusted sampling frequencies may be in the form of an array that includes a plurality of frequency sub-band boundary values in ascending order of frequency sub-bands.
[0021] The apparatus may include a first encoder and a second encoder for encoding one or more audio signals at a coding rate, which may include the sum of the coding rate for the first encoder and the coding rate for the second encoder. The first encoder may encode an audio transmission signal associated with the one or more audio signals, and the second encoder may encode a plurality of sets of spatial audio parameters associated with the frequency sub-bands of the one or more audio signals.
[0022] According to a third aspect, a method for spatially audio encoding one or more audio signals, the method comprising: for each of a plurality of frequency subbands of the one or more audio signals, determining a set of spatial audio parameters; receiving a coding rate associated with the one or more audio signals; based on the coding rate, mapping at least two consecutive subbands of the plurality of frequency subbands to an extended frequency subband to give a plurality of frequency subbands adjusted by the coding rate; receiving a bandwidth value associated with the one or more audio signals; starting from the highest frequency subband of the plurality of frequency subbands adjusted by the coding rate, removing a number of frequency subbands to give a plurality of frequency subbands adjusted by the bandwidth, wherein the number of removed frequency subbands is based on the bandwidth value associated with the one or more audio signals; in the condition that the highest frequency subband of the plurality of frequency subbands adjusted by the bandwidth exceeds the bandwidth value associated with the one or more audio signals, reducing the highest frequency subband of the plurality of frequency subbands adjusted by the bandwidth to the bandwidth value or below the bandwidth value; combining the set of spatial audio parameters associated with the first frequency subband of the at least two consecutive frequency subbands and the set of spatial audio parameters associated with the second frequency subband of the at least two consecutive frequency subbands to give a combined set of spatial audio parameters for the extended frequency subband; removing the set of spatial audio parameters corresponding to each removed frequency subband; and in the condition that the highest frequency subband of the plurality of frequency subbands adjusted by the bandwidth exceeds the bandwidth value, removing the set of spatial audio parameters associated with the plurality of frequency subbands adjusted by the bandwidth that exceed the bandwidth value.
[0023] According to a fourth aspect, a method for spatially audio encoding one or more audio signals, the method comprising: for each frequency subband of a plurality of frequency subbands of the one or more audio signals, determining a set of spatial audio parameters; receiving a coding rate associated with the one or more audio signals; based on the coding rate, mapping at least two consecutive subbands of the plurality of frequency subbands to an extended frequency subband to give a plurality of frequency subbands adjusted by the coding rate; combining the set of spatial audio parameters associated with the first frequency subband of the at least two consecutive frequency subbands and the set of spatial audio parameters associated with the second frequency subband of the at least two consecutive frequency subbands to give a combined set of spatial audio parameters for the extended frequency subband; for each frequency bin of the one or more audio signals, determining an energy level; determining a cut-off frequency subband of the one or more audio signals by: determining the highest frequency bin having an energy level greater than a predetermined energy level, and designating the cut-off frequency subband as the frequency subband containing the highest frequency bin; comparing the cut-off frequency subband of the one or more audio signals with a bandwidth value of the one or more audio signals; under the condition that the cut-off frequency subband is less than the bandwidth value of the one or more audio signals, starting from the highest frequency subband of the plurality of frequency subbands adjusted by the coding rate, removing a certain number of frequency subbands to give a plurality of frequency subbands adjusted by the bandwidth, and removing the set of spatial audio parameters corresponding to each removed frequency subband, wherein the number of removed frequency subbands is based on the cut-off frequency subband; under the condition that the highest frequency subband of the plurality of frequency subbands adjusted by the bandwidth exceeds the cut-off frequency subband, reducing the highest frequency subband of the plurality of frequency subbands adjusted by the bandwidth to the cut-off frequency subband value or a value lower than the cut-off frequency subband value, and removing the set of spatial audio parameters associated with the plurality of frequency subbands adjusted by the bandwidth that exceed the cut-off frequency subband; and encoding an index of the cut-off frequency subband.
[0024] According to a fifth aspect, an apparatus for spatial audio coding, the apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least: for each frequency subband among a plurality of frequency subbands of one or more audio signals, determine a set of spatial audio parameters; receive a coding rate associated with the one or more audio signals; based on the coding rate, map at least two consecutive subbands among the plurality of frequency subbands to an extended frequency subband to give a plurality of frequency subbands adjusted by the coding rate; receive a bandwidth value associated with the one or more audio signals; starting from the highest frequency subband among the plurality of frequency subbands adjusted by the coding rate, remove a certain number of frequency subbands to give a plurality of frequency subbands adjusted by the bandwidth, wherein the number of the removed frequency subbands is based on the bandwidth value associated with the one or more audio signals; under the condition that the highest frequency subband among the plurality of frequency subbands adjusted by the bandwidth exceeds the bandwidth value associated with the one or more audio signals, reduce the highest frequency subband among the plurality of frequency subbands adjusted by the bandwidth to the bandwidth value or below the bandwidth value; merge the set of spatial audio parameters associated with the first frequency subband among the at least two consecutive frequency subbands and the set of spatial audio parameters associated with the second frequency subband among the at least two consecutive frequency subbands to give a merged set of spatial audio parameters for the extended frequency subband; remove the set of spatial audio parameters corresponding to each of the removed frequency subbands; and under the condition that the highest frequency subband among the plurality of frequency subbands adjusted by the bandwidth exceeds the bandwidth value, remove the set of spatial audio parameters associated with the plurality of frequency subbands adjusted by the bandwidth that exceed the bandwidth value.
[0025] According to a sixth aspect, an apparatus for spatial audio coding, the apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least: for each frequency subband among a plurality of frequency subbands of one or more audio signals, determine a set of spatial audio parameters; receive a coding rate associated with the one or more audio signals; based on the coding rate, map at least two consecutive subbands among the plurality of frequency subbands to an extended frequency subband to give a plurality of frequency subbands adjusted by the coding rate; merge the set of spatial audio parameters associated with the first frequency subband among the at least two consecutive frequency subbands and the set of spatial audio parameters associated with the second frequency subband among the at least two consecutive frequency subbands to give a merged set of spatial audio parameters for the extended frequency subband; for each frequency bin of the one or more audio signals, determine an energy level; determine a cut-off frequency subband of the one or more audio signals by: determining the highest frequency bin having an energy level greater than a predetermined energy level, and designating the cut-off frequency subband as the frequency subband containing the highest frequency bin; compare the cut-off frequency subband of the one or more audio signals with a bandwidth value of the one or more audio signals; under the condition that the cut-off frequency subband is less than the bandwidth value of the one or more audio signals, starting from the highest frequency subband among the plurality of frequency subbands adjusted by the coding rate, remove a certain number of frequency subbands to give a plurality of frequency subbands adjusted by the bandwidth, and remove the set of spatial audio parameters corresponding to each removed frequency subband, wherein the number of removed frequency subbands is based on the cut-off frequency subband; under the condition that the highest frequency subband among the plurality of frequency subbands adjusted by the bandwidth exceeds the cut-off frequency subband, reduce the highest frequency subband among the plurality of frequency subbands adjusted by the bandwidth to the cut-off frequency subband value or a value lower than the cut-off frequency subband value, and remove the set of spatial audio parameters associated with the plurality of frequency subbands adjusted by the bandwidth that exceed the cut-off frequency subband; and encode an index of the cut-off frequency subband.
[0026] A computer program product stored on a medium can cause an apparatus to perform the method as described herein.
[0027] An electronic device can include the apparatus as described herein.
[0028] A chipset can include the apparatus as described herein.
[0029] Embodiments of the present application are intended to solve problems associated with the prior art. Description of the Drawings
[0030] To better understand the present application, reference will now be made, by way of example, to the accompanying drawings, in which:
[0031] Figure 1 Schematically shows a device system suitable for implementing some embodiments;
[0032] Figure 2 Schematically shows an analysis processor according to some embodiments;
[0033] Figure 3 Schematically shows a spatial analyzer according to some embodiments;
[0034] Figure 4 Shows a band adjuster according to some embodiments;
[0035] Figure 5 Shows, as in Figure 4 a flowchart of the operation of the band adjuster shown;
[0036] Figure 6 Shows a band adjuster according to another embodiment;
[0037] Figure 7 Shows according to Figure 6 a flowchart of the operation of the band adjuster according to another embodiment shown in;
[0038] Figure 8 Shows a flowchart of the operation of the band adjuster of an embodiment combining Figure 5 and Figure 7 ; and
[0039] Figure 9 Schematically shows an example device suitable for implementing the devices shown. Detailed Description
[0040] Suitable devices and possible mechanisms for providing metadata parameters (MASA parameters) for effective spatial analysis derivation are described in more detail below.
[0041] As described above, Metadata-Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation suitable as an input format for IVAS.
[0042] It can be regarded as an audio representation consisting of "N channels + spatial metadata". It is a scene-based audio format, particularly suitable for spatial audio capture on consumer devices such as smart phones. The idea is to describe the sound scene in terms of the direction of the time-frequency varying sound source and, for example, the energy ratio. The sound energy not defined (described) by the direction is described as diffuse (from all directions).
[0043] As discussed above, the spatial metadata associated with an audio signal can include multiple parameters per time-frequency bin (such as multiple directions and the direct-to-total ratio, spread coherence, distance, etc. associated with each direction). The spatial metadata can also include other parameters or can be associated with other parameters that are considered non-directional (such as surround coherence, diffuse-to-total energy ratio, remainder-to-total energy ratio), but when combined with the directional parameters, can be used to define the characteristics of an audio scene. For example, a reasonable design choice that can result in a good quality output is to determine that the spatial metadata includes one or more directions (and the direct-to-total energy ratio, spread coherence, distance values, etc. associated with each direction) for each time-frequency sub-frame.
[0044] As described above, a parametric spatial metadata representation can use multiple concurrent spatial directions. With MASA, the proposed maximum number of concurrent directions is two. For each concurrent direction, there can be associated parameters such as: a direction index; a direct-to-total energy ratio; spread coherence; and distance. In some embodiments, other parameters are defined such as a diffuse-to-total energy ratio; surround coherence; and a remainder-to-total energy ratio.
[0045] In the following discussion, multi-channel systems will be discussed with respect to multi-channel microphone implementations. However, as described above, the input format can be any suitable input format such as multi-channel speakers, Ambisonic (FOA / HOA), etc. Additionally, the output of the example system is a multi-channel speaker device. However, it will be understood that the output can be rendered to the user via means / components other than speakers, such as binaural channel output. Additionally, the multi-channel speaker signal can be generalized to two or more playback audio signals.
[0046] Additionally, the IVAS codec as an extension to EVS can be used for store-and-forward applications where audio and voice content is encoded and stored in a file for playback.
[0047] For each time-frequency (TF) block or tile (also referred to as a time / frequency sub-band) under consideration, the MASA metadata can include at least spherical direction (elevation, azimuth), at least one direct-to-total energy ratio of the resulting direction, extended coherence, and surround coherence independent of direction. Overall, MASA can have multiple different types of metadata parameters for each time-frequency (TF) tile. Table 1 below shows the types of spatial audio parameters that constitute the metadata for MASA.
[0048] This data can be encoded and sent (or stored) by the encoder so that the spatial signal can be reconstructed at the decoder.
[0049] In addition, in some cases, metadata-assisted spatial audio (MASA) can support up to 2 directions per TF tile, which will require encoding and sending the above parameters for each direction on a per-TF-tile basis. Thus, according to Table 1 below, the required bitrate may double.
[0050]
[0051] The bitrate allocated for metadata in a practical immersive audio communication codec can vary widely. The typical overall operating bitrate of the codec may only leave 2 to 10 kbps for the transmission / storage of spatial metadata. However, some other implementations may allow a bitrate of up to 60 kbps or higher for the transmission / storage of spatial metadata. The encoding of the direction parameters and energy ratio components and the encoding of the coherence data have been examined previously. However, regardless of the bitrate specified for the transmission / storage of spatial metadata, it will always be necessary to represent these parameters using as few bits as possible, especially when a TF tile can support multiple directions corresponding to different sound sources in a spatial audio scene.
[0052] Figure 1 An example apparatus and system for implementing embodiments of the present application are shown. System 100 is shown to have an "analysis" section 121 and a "synthesis" section 131. The "analysis" section 121 is the part that encodes from the received multi-channel signal to the metadata and the transmission signal, while the "synthesis" section 131 is the part that decodes from the encoded metadata and the transmission signal to the presentation of the regenerated signal (e.g., in the form of multi-channel speakers).
[0053] The input to system 100 and the "analysis" section 121 is the input audio signal 102.
[0054] In the following example, the audio input signal 102 may be from a microphone array. However, it will be understood that the audio input can be any suitable audio input format, and the differences in processing when different input formats are employed will be described in detail below.
[0055] The audio input signal 102 can be from any suitable source, such as: two or more microphones mounted on a mobile phone, other microphone arrays, such as, a B-format microphone or an Eigenmike. In some embodiments, as described above, the input can be any suitable audio signal input, such as an Ambisonic signal, e.g., first-order Ambisonics (FOA), higher-order Ambisonics (HOA) or speaker surround mix and / or object or any combination of the above.
[0056] In an embodiment, the microphone array audio input signal 102 can be provided to an analysis processor 105, which is configured to generate or determine suitable (spatial) metadata associated with the audio input signal 102. Additionally, the (microphone array) audio input signal 102 can also be provided to a suitable transmission signal generator 103 to generate an audio transmission signal 104.
[0057] Thus, the analysis processor 105 is configured to perform a spatial analysis on the audio input signal 102 to generate suitable spatial audio (MASA) metadata 106 in a frequency band. For all the above input types, there are known methods to generate suitable spatial metadata, e.g., direction and direct-to-total energy ratio (or similar parameters, such as diffuseness, i.e., ambient-to-total ratio) in a frequency band. These methods are not described in detail herein. However, some examples may include performing a suitable time-frequency transform on the input signal, and then estimating the delay value between microphone pairs that maximizes the inter-microphone correlation in a frequency band when the input is a mobile phone microphone array, formulating a corresponding direction value for the delay, and formulating a ratio parameter based on the correlation value.
[0058] In some embodiments, when the audio input is a FOA signal or a B-format microphone, the analysis processor 105 can be configured to determine parameters such as an intensity vector (from which a direction parameter is obtained), and compare the intensity vector length with an estimate of the total sound field energy to determine a ratio parameter. This method is known in the literature as Directional Audio Coding (DirAC).
[0059] In some embodiments, when the audio input signal 102 is a HOA signal, the analysis processor 105 may take the FOA subset of the signal and use the methods described above, or divide the HOA signal into multiple parts / sectors and use the methods described above in each part / sector. This part / sector-based method is known as higher-order DirAC (HO-DirAC) in the literature. In this case, there are more than one simultaneous direction parameters per frequency band.
[0060] In some embodiments, when the audio input signal 102 is a loudspeaker surround mix and / or object, the analysis processor 105 may be configured to convert the signal into a FOA signal (by using spherical harmonic coding gain) and analyze the direction and ratio parameters as described above.
[0061] Thus, the output of the analysis processor 105 is the spatial audio (MASA) metadata 106 in the frequency band. The spatial audio (MASA) metadata 106 may relate to the direction and energy ratio in the frequency band, but may also have any of the metadata types listed previously. The spatial audio (MASA) metadata 106 may vary over time and over frequency.
[0062] In some embodiments, the analysis processor functionality is implemented external to the system 100. For example, in some embodiments, the spatial audio (MASA) metadata 106 associated with the audio input signal 102 may be provided to the encoder 107 as a separate bitstream. In some embodiments, the spatial audio (MASA) metadata 106 may be provided as a set of spatial (direction) index values.
[0063] The system 100 described above is further configured to implement a transmission signal generator 103 to generate a suitable audio transmission signal 104. The transmission signal generator 103 is configured to receive the audio input signal 102 (which may be, for example, a microphone array audio signal) and generate the audio transmission signal 104. The audio transmission signal 104 may be a multi-channel, stereo, binaural, or monaural audio signal. The generation of the audio transmission signal 104 may be implemented using any suitable method, as outlined below.
[0064] When the audio input signal 102 is a microphone array audio signal, the transmission signal generator 103 may select the left and right microphone pair and apply suitable processing to the signal pair, such as automatic gain control, microphone noise removal, wind noise removal, and equalization.
[0065] When the input is a FOA / HOA signal or a B-format microphone, the audio transmission signal 104 may be a directional beam signal directed towards the left and right directions, such as two opposite cardioid signals.
[0066] When the input is a loudspeaker surround mix and / or an object, the audio transport signal 104 may be a downmix signal that combines the left channels into a left downmix channel, and likewise for the right, with the center channel added to both transport channels with an appropriate gain.
[0067] In some embodiments, the audio transmission signal 104 is an audio input signal 102, such as a microphone array audio signal. For example, in some cases, analysis and synthesis occur in a single processing step at the same device without intermediate encoding. The number of audio transmission channels can also be any suitable number (rather than one or two channels as discussed in the example).
[0068] In some embodiments, the transmission signal generator 103 and the analysis processor 105 may be computers (running appropriate software stored on a memory and on at least one processor), or alternatively, may be specific devices using, for example, an FPGA or ASIC.
[0069] The transport signal 104 and the spatial audio (MASA) metadata 106 may be passed to an encoder 107 .
[0070] The encoder 107 may include an audio encoder core 109 configured to receive the audio transmission (e.g., downmix) signals 104 and generate suitable encodings of these audio signals. In some embodiments, the encoder 107 may be a computer (running suitable software stored on a memory and on at least one processor), or alternatively, may be a specific device using, for example, an FPGA or ASIC. The encoding may be implemented using any suitable scheme. The encoder 107 may also include a metadata encoder / quantizer 111 configured to receive the spatial audio (MASA) metadata 106 and output an encoded or compressed form of the information. In some embodiments, the encoder 107 may further interleave, multiplex into a single data stream, or transmit or store (in Figure 1 The metadata is embedded into the encoded downmix signal before (indicated by the dotted line). The multiplexing can be achieved using any suitable scheme.
[0071] At the decoder side, the received or retrieved data (stream) can be received by the decoder / demultiplexer 133. The decoder / demultiplexer 133 can demultiplex the encoded stream and pass the audio encoded stream to the transmission extractor 135, which is configured to decode the audio signal to obtain the transmission signal. Similarly, the decoder / demultiplexer 133 can include a metadata extractor 137, which is configured to receive the encoded metadata and decode the metadata. In some embodiments, the decoder / demultiplexer 133 can be a computer (running suitable software stored on a memory and at least one processor), or alternatively, can be a specific device using, for example, FPGA or ASIC.
[0072] The decoded metadata and the transmitted audio signal can be passed to the synthesis processor 139.
[0073] The "synthesis" part 131 of the system 100 further shows a synthesis processor 139, which is configured to receive the encoded audio transmission signal and the encoded spatial audio (MASA) metadata, and recreate, in any suitable format, the synthesized spatial audio in the form of multi-channel spatial audio signals 110 (these signals can be in a multi-channel speaker format, or, in some embodiments, can be any suitable output format, such as binaural or Ambisonics signals, depending on the use case) based on the encoded audio transmission signal and the encoded spatial audio (MASA) metadata.
[0074] Thus, in summary, first, the system (analysis part) is configured to receive multi-channel audio signals.
[0075] Then, the system (analysis part) is configured to generate suitable transmitted audio signals (e.g., by selecting or downmixing certain audio signal channels) and spatial audio parameters as metadata.
[0076] Furthermore, the system is configured to encode the audio transmission signal and the spatial audio (MASA) metadata for storage / transmission.
[0077] Thereafter, the system can store / send the encoded audio transmission signal and the encoded spatial audio (MASA) metadata.
[0078] The system can retrieve / receive the encoded audio transmission signal and the encoded spatial audio (MASA) metadata.
[0079] Then, the system is configured to extract the audio transmission signal and the spatial audio (MASA) metadata from the encoded audio transmission signal and the encoded spatial audio (MASA) metadata parameters, e.g., by demultiplexing and decoding the encoded audio transmission signal and the encoded spatial audio (MASA) metadata parameters.
[0080] The system (synthesis section) is configured to synthesize an output multi-channel spatial audio signal based on the extracted audio transmission signal and the spatial audio (MASA) metadata.
[0081] Figure 2 is an example analysis processor 105 and metadata encoder / quantizer 111 according to some embodiments (as Figure 1 shown), and is described in further detail below.
[0082] Figure 1 and Figure 2 shows that the metadata encoder / quantizer 111 and the analysis processor 105 are coupled together. However, it will be understood that in some embodiments, these two separate processing entities may not be coupled so closely, such that the analysis processor 105 may be located on a different device from the metadata encoder / quantizer 111. Thus, a device including the metadata encoder / quantizer 111 may be provided with the audio transmission signal 104 and the metadata stream for processing and encoding, independent of the capture and analysis process.
[0083] In some embodiments, the analysis processor 105 includes a time-frequency domain transformer 201.
[0084] In some embodiments, the time-frequency domain transformer 201 is configured to receive the audio input signal 102 and apply a suitable time-domain to frequency-domain transform, such as a short-time Fourier transform (STFT), in order to convert the audio input time-domain signal into a suitable time-frequency audio signal 202. These time-frequency audio signals 202 may be passed to the spatial analyzer 203.
[0085] Thus, for example, the time-frequency audio signal 202 may be represented in the time-frequency domain by the following equation as:
[0086] s i (b,n)
[0087] where b is the frequency bin index, n is the time-frequency block (frame) index, and i is the channel index. In other words, n may be considered a time index with a sampling rate lower than that of the original time-domain signal. These frequency bins may be grouped into subbands, which group one or more bins into subbands with band indices k = 0,..., K-1. Each subband k has a lowest bin b k,low and a highest bin b k,high , and the subband contains all bins from b k,low to b k,high . The width of the subbands may approximate any suitable distribution, for example, the equivalent rectangular bandwidth (ERB) scale or the Bark scale.
[0088] Thus, a time-frequency (TF) tile (or block) is a specific sub-band within a sub-frame of a frame.
[0089] It can be understood that the number of bits required to represent spatial audio parameters can depend at least in part on the TF (time-frequency) tile resolution (i.e., the number of TF sub-frames or tiles). For example, a 20 ms audio frame can be divided into 4 time-domain sub-frames of 5 ms each, and each time-domain sub-frame can have up to 24 frequency sub-bands, approximations thereof, or any other suitable division in the frequency domain according to the Bark scale. In this particular example, the audio frame can be divided into 96 TF sub-frames / tiles, in other words, 4 time-domain sub-frames with 24 frequency sub-bands. Thus, the number of bits required to represent the spatial audio parameters of an audio frame can depend on the resolution of the TF tiles. For example, if each TF tile is encoded according to the assignment in Table 1 above, each TF tile will require 64 bits (for one sound source direction per TF tile) and 104 bits (for two sound source directions per TF tile, considering parameters independent of the sound source direction).
[0090] In an embodiment, the analysis processor 105 can include a spatial analyzer 203. The spatial analyzer 203 can be configured to receive the time-frequency audio signal 202 and estimate a set of spatial audio parameters for each TF tile based on these signals. This is shown overall as spatial audio (MASA) metadata 106 in Figure 1 The spatial audio (MASA) metadata 106 can include direction parameters. The direction parameter can be determined based on any audio-based "direction" determination.
[0091] For example, in some embodiments, the spatial analyzer 203 is configured to estimate the direction of a sound source having two or more signal inputs.
[0092] Thus, the spatial analyzer 203 can be configured to provide, for each frequency band and time-frequency block within a frame of the audio signal, at least one azimuth angle and elevation angle (spatial audio direction parameters), labeled as azimuth angle φ(k,n) and elevation angle θ(k,n). The spatial audio direction parameters for the time sub-frame can be passed to the spatial parameter set encoder 207.
[0093] The spatial analyzer 203 can also be configured to determine an energy ratio parameter. This energy ratio can be considered as the determination of the energy of the audio signal that can be considered to arrive from a certain direction. For example, the direct-to-total energy ratio r(k,n) can be estimated using a stability measure of the direction estimation, or using any correlation measure, or any other suitable method for obtaining the ratio parameter (such as those described in the patent publication EP3542546). Each direct-to-total energy ratio corresponds to a specific spatial direction and describes how much of the total energy is from this specific spatial direction. This value can also be represented separately for each time-frequency bin. The spatial direction parameter and the direct-to-total energy ratio describe how much of the total energy for each time-frequency bin is from this specific direction. Generally, the spatial direction parameter can also be considered as the direction of arrival (DOA).
[0094] The direct-to-total energy ratio parameter of the multi-channel captured microphone array signal can be estimated based on the normalized cross-correlation parameter cor′(k,n) between microphone pairs in the frequency band k, and the value of this cross-correlation parameter is between -1 and 1. The direct-to-total energy ratio parameter r(k,n) can be determined by comparing the normalized cross-correlation parameter with the diffuse-field normalized cross-correlation parameter cor′ D (k,n) as The direct-to-total energy ratio is further illustrated in the PCT publication WO2017 / 005978, which is incorporated herein by reference. This energy ratio can be passed to the spatial parameter set encoder 207.
[0095] The spatial analyzer 203 can also be configured to determine a plurality of coherence parameters 112, which can include the ambient coherence (γ(k,n)) and the spread coherence (ζ(k,n)), both of which are analyzed in the time-frequency domain.
[0096] The term "audio source" can be related to the dominant direction of the propagating sound wave, and this dominant direction can include the actual direction of the sound source.
[0097] Therefore, for each sub-band k, there will be a set of spatial audio parameters (or a group of spatial audio parameters) associated with the sub-band k and the sub-frame n. In this example, each sub-band k and sub-frame n (in other words, the TF bin) can have the following spatial audio parameters associated with it based on each audio source direction: at least one azimuth and elevation angle, labeled as the azimuth φ(k,n) and the elevation angle θ(k,n), the spread coherence (ζ(k,n)), and the direct-to-total energy ratio parameter r(k,n). If there is more than one direction per TF bin, the TF bin can have each of the above-listed parameters associated with each sound source direction. In addition, the set of spatial audio parameters can also include the ambient coherence (γ(k,n)). The parameter can also include the diffuse-to-total energy ratio r diff(k,n).
[0098] In an embodiment, the diffusion-to-total energy ratio r diff (k,n) is the energy ratio of non-directional sound in the surrounding direction, and there is typically a single diffusion-to-total energy ratio (and surrounding coherence (γ(k,n))) per TF tile. The diffusion-to-total energy ratio can be considered as the energy ratio remaining when the direct-to-total energy ratio (for each direction) has been subtracted from one. Next, the above parameters can be referred to as a set of spatial audio parameters (or a collection of spatial audio parameters) for a specific TF tile. The collection of sets of spatial audio parameters associated with TF tiles is called the spatial audio (MASA) metadata signal 106.
[0099] Furthermore, the set of spatial parameter data is passed to the metadata encoder / quantizer 111 for encoding and quantization. In Figure 2 this, this is illustrated by the spatial parameter set encoder 207, which can be arranged to receive the set of spatial parameter data (represented as the spatial audio MASA metadata stream 106) and quantize and encode the set of spatial parameters associated with each TF tile.
[0100] The audio input signal 102 can be processed by both the transmission signal generator 103 and the analysis processor 105 in the frequency domain for the same frequency subband resolution. However, some of the resulting frequency subbands in the audio transmission signal 104 (from the processing of the transmission signal generator 103) can contain very small (or even zero) signal energy levels, resulting in a negligible contribution, or at most a very small contribution, of the signals associated with these subbands to the overall synthesized multi-channel (spatial) audio signal 110. This would indicate that the audio signals within the "low energy" frequency subbands can be ignored and may not be encoded (by the audio encoder core 109) for subsequent transmission and storage.
[0101] However, for the current system, the analysis processor 105 generates a set of spatial audio parameters for each subband of the subframes of the processed audio input signal 102. Thus, there will be a mismatch between the number of subbands of the audio transmission signal 104 on which there are so-called active audio signals and the number of subbands of the audio input signal 102 on which the spatial audio (MASA) metadata signal 106 is analyzed. In other words, the analysis processor 105 can generate a set of spatial parameter data for each subband of the subframes of the audio input signal 102, regardless of whether the corresponding frequency subbands of the audio transmission stream / signal 104 contain active audio signals.
[0102] Accordingly, a set of spatial audio parameters corresponding to frequency sub-bands of the audio transmission signal / stream 10 with inactive audio signals (on a per sub-frame basis) can be considered to be encoded unnecessarily. Encoding these sets of spatial audio parameters in turn results in an unnecessary bit overhead.
[0103] Note that as applied above, the term "active audio signal" refers to the case where the audio signal of a sub-band of a sub-frame of the audio transmission signal 104 has a high enough energy level such that the audio signal is considered to contribute to the synthesized multi-channel spatial audio signal 110. In contrast, the term "inactive audio signal" can refer to the case where a sub-band of a sub-frame of the audio transmission signal 104 has a very low audio signal energy level such that the sub-band can be considered to make no significant contribution to the synthesized multi-channel spatial audio signal 110.
[0104] Accordingly, the embodiments are based on the consideration that if the contribution of the energy in the frequency sub-bands of the audio transmission signal 104 to the output multi-channel spatial audio signal 110 is negligible, the number of sets of spatial audio parameters of the spatial audio (MASA) meta data stream 106 can be reduced.
[0105] In addition, as previously described, the IVAS codec can operate at a series of different coding rates and different bandwidths, and this can lead to a mismatch between the number of sub-bands on which the audio input signal 102 is processed for the audio transmission stream 104 and the number of sub-bands on which the audio input signal 102 is analyzed for the spatial (MASA) meta data stream 106. The mismatch in the number of sub-bands (the number of sub-bands on which the audio transmission stream 104 and the spatial (MASA) meta data stream 106 are processed) can be at least partially attributed to the coding rate assigned for the coding of the audio transmission stream 104 and the separate coding rate assigned for the spatial (MASA) meta data stream 106. For example, the audio transmission stream 104 and the spatial (MASA) meta data stream 106 can each be encoded according to any one of a plurality of different coding rates. The coding rate assigned for each stream in turn affects the number of sub-bands on which the audio transmission stream 104 and the spatial (MASA) meta data stream 106 are generated. For example, the coding rate assigned for the coding of the audio transmission stream 104 can result in a smaller number of generated sub-bands than the number of sub-bands on which the spatial audio parameters of the spatial (MASA) meta data stream 106 are generated.
[0106] Accordingly, the frequency bands of the spatial (MASA) meta data stream 106 can extend beyond the frequency bands of the audio transmission stream 104. This results in unnecessary encoding of the spatial audio parameters associated with the sub-bands of the spatial (MASA) meta data stream 106 that extend beyond the sub-bands of the audio transmission stream 104, which in turn results in an unnecessary coding bit overhead during the encoding of the spatial (MASA) meta data stream 106.
[0107] In this regard, Figure 3 the spatial analyzer 203 is shown in more detail. Among them, the time-frequency audio signal 202 is received by the spatial parameter set determiner 301. The spatial parameter set determiner 301 can be configured to determine a set of spatial parameters for each sub-band of the time-frequency audio signal 202. The components of each parameter set can be at least some of the spatial audio parameters discussed above and listed in Table 1.
[0108] It should be noted that in some other embodiments, the spatial parameter set determiner 301 can be implemented in the spatial analyzer 203 in the analysis processor 105, and the frequency sub-band adjuster 303 and the parameter set combiner / reducer 305 can form part of the metadata encoder / quantizer 111, and the analysis processor 105 can be present on a device different from the metadata encoder / quantizer 111.
[0109] In addition, the frequency sub-band adjuster 303 is also shown in Figure 3 . The frequency sub-band adjuster 303 can be configured to receive input configuration information, such as the total (IVAS) coding rate 206 (selected) and the audio signal bandwidth 208 (selected). In addition, the frequency sub-band adjuster 303 can also be configured to receive the audio transmission signal 104.
[0110] Furthermore, the frequency sub-band adjuster 303 can generate another sub-band arrangement in response to the received input configuration information, the total (IVAS) coding rate 206, and the audio signal bandwidth 208. This another sub-band arrangement can be based on the original sub-band arrangement of the time-frequency audio signal 202. However, it has some changes in the distribution and width of certain frequency sub-bands, and thus changes the number of sub-bands on the bandwidth of the signal. For example, when compared with the pattern of the sub-bands of the time-frequency audio signal 202, this another sub-band arrangement can include fewer and wider sub-bands.
[0111] In an embodiment where the input of the frequency sub-band adjuster 303 includes the audio transmission signal 104, the frequency sub-band arrangement of the original time-frequency audio signal 102 can be reduced in response to the energy of each corresponding sub-band of the audio transmission signal 104. In other words, the frequency sub-band arrangement generated by the frequency sub-band adjuster 303 can be made fewer by removing frequency sub-bands from the original pattern of the sub-bands of the time-frequency audio signal 202. Therefore, in response to the energy level of the frequency sub-bands of the audio transmission signal 104, the resulting frequency sub-band arrangement can include fewer sub-bands of the original width.
[0112] The output from the frequency sub-band adjuster 303 is in Figure 3is shown as the adjusted subband configuration array 302. This parameter can reflect changes (or removals of frequency subbands) to the boundaries of the frequency subbands of the original time-frequency audio signal 202 in the form of an array of subband boundary values. In other words, the adjusted subband configuration array 302 can represent the subband boundary pattern after the encoder operating conditions of the selected coding rate (total IVAS coding rate 206) and the selected bandwidth (audio signal bandwidth 208) have been considered. Note that the total (IVAS) coding rate 206 is only used as an example of how the coding rate can be parameterized. This does not exclude any other parameter that can indicate the coding rate for the encoder. For example, the coding rate parameter (such as input 206) can be set according to the coding rate of the audio encoder 109, or the coding rate associated with the metadata encoder and quantizer 111.
[0113] Furthermore, the adjusted subband configuration parameter 302 can be passed to the parameter set combiner / reducer 305.
[0114] In addition to the adjusted subband configuration parameter 302, the parameter set combiner / reducer 305 also receives a set of spatial audio parameters 304 for each frequency subband of the time-frequency audio signal 202.
[0115] When the input to the spatial analyzer 203 includes the total (IVAS) coding rate 206 and the audio signal sampling frequency 208, the parameter set combiner / reducer 305 can be set to perform a merging operation between some of the spatial parameter sets 304. This merging operation can be performed according to the subband configuration of the subband configuration parameter / array 302. Essentially, some of the spatial parameter sets (for the time-frequency audio signal 202) can be merged with adjacent spatial parameter sets so that the resulting spatial parameter set distribution reflects / mirrors the subband distribution as indicated by the adjusted subband configuration parameter 302. A description of the merging process can be found in patent application publication WO2021 / 130404. It teaches that the sets of spatial audio parameters on adjacent subbands can be merged to give fewer sets of spatial audio parameters on a fewer number of merged frequency bands.
[0116] When the input to the spatial analyzer 203 includes the audio transmission signal 104, the parameter set combiner / reducer 305 can be set to reduce a plurality of spatial audio parameter sets from the signal 304 as indicated by the subband cut-off signal 306. In this case, the adjusted subband cut-off signal 306 can contain information indicating the spatial parameter sets that are to be removed from the spatial audio parameter set signal 304.
[0117] Furthermore, the output from the parameter set combiner / reducer 305 (i.e., the spatial audio metadata 106) may include spatial audio parameter sets in the signal 304 that have been combined into a smaller number of spatial audio parameter sets, and / or spatial audio parameter sets in the signal 304 that have been reduced to a smaller number of spatial audio parameter sets.
[0118] Embodiments for generating the spatial audio metadata 106 in response to the adjusted subband configuration array 302 will be described below. That is, the output from the parameter set combiner / reducer 305 is the spatial audio metadata 106, which includes a smaller number of combined spatial audio parameter sets in the spatial audio parameter set signal 304 in response to the audio signal bandwidth (parameter) 208 and the total (IVAS / coding system) coding rate 206.
[0119] To this end, the spatial audio (MASA) metadata 106 may be encoded (by the encoder 207) at various coding rates between 2.5 kbps and 65 kbps. The particular coding rate selected may be related to the total (IVAS or system) coding rate 206, which for IVAS may be one of the following:
[0120] / *IVAS_13k2,IVAS_16k4,IVAS_24k4,IVAS_32k,IVAS_48k,IVAS_64k,IVAS_80k,IVAS_96k,IVAS_128k,IVAS_160k,IVAS_192k,IVAS_256k,IVAS_384k,IVAS_512k.* /
[0121] Where, for example, IVAS_13k2 represents an IVAS coding rate of 13.2 kbps. The total (IVAS) coding rate 206 may be used by the frequency subband adjuster 303 to partially determine the subband boundaries for the adjusted subband configuration parameters 302.
[0122] In this regard, Figure 4 The frequency subband adjuster 303 is shown in more detail for the case of generating the adjusted subband configuration array 302 in response to a combination of inputs including the total (IVAS) coding rate 206 and the audio signal bandwidth (parameter) 208.
[0123] The total system coding rate (IVAS coding rate) 206 is shown as being received by the coding rate subband adjuster 401. The output from the coding rate subband adjuster 401 is shown as the coded rate adjusted subband array 402.
[0124] In an embodiment, the coding rate sub-band adjuster 401 may be configured to perform a mapping function between the total coding rate 206 and a particular frequency sub-band distribution regarding the distribution of sub-bands in the time-frequency audio signal 202. In other words, the result of this mapping function is the coded rate-adjusted sub-band array 402.
[0125] The mapping may be performed such that the distribution of the coded rate-adjusted sub-bands is more closely aligned with the width and number of the frequency sub-bands of the transmitted audio signal 104. The time-frequency audio signal 202 may include 24 frequency sub-bands over its bandwidth. Further, the mapping function in 401 may be configured to take the total (IVAS) coding rate 206 and map that coding rate to a frequency sub-band distribution different from the distribution of the frequency sub-bands of the time-frequency audio signal 202. The coded rate-adjusted sub-band array 402 may have a smaller number of sub-bands, where some sub-bands are wider than their corresponding sub-bands in the time-frequency audio signal 202. Thus, the resulting coded rate-adjusted sub-bands continue to extend across the equivalent bandwidth of the original time-frequency audio signal 202 (of 24 frequency bands), but with fewer sub-bands.
[0126] In an embodiment, the mapping function may be implemented by initially mapping the received total (IVAS) coding rate 206 to a parameter indicating the number of sub-bands in the coded rate-adjusted sub-band array. There is a one-to-one mapping between each total (IVAS) coding rate 206 and the parameter indicating the reduced number of sub-bands. An example of the one-to-one mapping for IVAS is shown in Table 2 below.
[0127]
[0128] Table 2
[0129] For example, a total IVAS coding rate of 160 kbps will result in the number of sub-bands in the coded rate-adjusted sub-band array 402 being reduced from 24 to 12.
[0130] It should be noted that each number of sub-bands in Table 2 above refers to the continuously operating sub-bands starting from the lowest sub-band, and the coded rate-adjusted sub-band array 402 extends across the entire bandwidth occupied by the 24 sub-bands of the time-frequency audio signal 202.
[0131] Each parameter in the table above indicating the reduced number of sub-bands is mapped to the IVAS coding rate in ascending order of bit rate. As another example, for an IVAS coding rate of 32 kbps, the parameter indicating the reduced number of sub-bands is 5 sub-bands. In this example, the coded rate-adjusted sub-band array will include elements marking the sub-band boundaries of those 5 sub-bands.
[0132] It should be clear that any reduction in the number of sub - bands related to the time - frequency audio signal 202 that can be performed is based on the maximum number of sub - bands, which is given as 24 in the above example. Thus, any adjustment to the number of sub - bands is made while preserving the full bandwidth of the signal. In essence, the width of some frequency sub - bands is expanded to occupy a wider range of frequency bins while preserving the full bandwidth associated with the time - frequency audio signal 202 (for IVAS, it is 24 sub - bands or 60 frequency bins, where each frequency bin has a width of 400 Hz).
[0133] Once the parameter indicating the reduced number of sub - bands has been found in Table 2 above, the width of some of the remaining sub - bands can be changed so as to preserve the full bandwidth of the signal, as described above. The redistribution of multiple frequency bins for some of the sub - bands with a reduced number can be found by using the following mapping array.
[0134] The frequency - bin distribution for the 24 sub - bands of the time - frequency audio signal 202 can be given by the following array MASA_band_grouping_24. In other words, this is the frequency - bin distribution for each sub - band in the 24 - sub - band grouping, where the maximum number of frequency bins is 60 and each bin has a width of 400 Hz.
[0135]
[0136] Each member in MASA_band_grouping is an indication of the frequency - bin index indicating the lower / upper boundaries of the sub - band. The frequency - bin indices are grouped uniformly in ascending order in the above - mentioned grouping array. For example, in the above MASA band grouping, the 24th sub - band is assigned the frequency - bin range from 40 to 60, the 23rd sub - band is assigned the frequency - bin range from 30 to 40, and the first sub - band is assigned the frequency bins from 0 to 1. It should be noted that the assignment / specification of frequency bins for a sub - band generally does not include the last value of the frequency - bin range. Thus, in reality, the frequency bins assigned to the 24th sub - band will be from 40 to 59, and similarly, the frequency - bin range assigned to the 23rd sub - band will be from 30 to 39.
[0137] The redistribution of frequency sub - bands for each value of the parameter indicating the reduced number of sub - bands in Table 2 can be given by the following MASA_band_mapping array.
[0138] For example, if the total (IVAS) coding rate is 192 kbps, the coded - rate - adjusted sub - band array 402 reduced from 24 sb to 18 sb can be given by the following array:
[0139]
[0140] In this example, the 18th sub-band in the coded rate adjusted sub-band array is assigned to the frequency bins covering the 23rd to 24th sub-bands related to the original sub-bands of the time-frequency audio signal 202, where the frequency bins assigned to each sub-band are given by the above MASA_band_grouping array. In other words, the 18th sub-band occupies the frequency bin range from 40 to 60. The 17th sub-band in the coded rate adjusted sub-band array 402 is assigned to cover the frequency bin range of the 22nd and 23rd sub-bands in the above MASA_band_grouping array, that is, the frequency bin range is from 30 to 40, and so on.
[0141] Similarly, in this example, since the IVAS coding rate is 160 kbps, the reduced sub-band count is 12 sb. The coded rate adjusted sub-band array 402 is given by the following array.
[0142]
[0143] In a similar manner, for the reduced sub-band counts from 24 sb to 8 sb and from 24 sb to 5 sb, the coded rate adjusted sub-band array 402 can be given by the following arrays respectively.
[0144]
[0145] It should be noted that the coded rate adjusted sub-band array can be any one from MASA_band_mapping_24_to_18 to MASA_band_mapping_24_to_5. Therefore, the coded rate adjusted sub-band array contains the "pattern" of sub-bands (sub-bands according to MASA_band_grouping_24) in response to the coding rate 206 determined by the total (IVAS) system.
[0146] Returning to Figure 4 , it can be seen that the output from the coded rate frequency sub-band adjuster 401, that is, the coded rate adjusted sub-band array 402, can be passed to the coded bandwidth sub-band adjuster 403. The coded bandwidth sub-band adjuster 403 is configured to receive the audio signal bandwidth 208, and this audio signal bandwidth 208 can be used to reduce the sampling frequency / bandwidth associated with the coded rate adjusted sub-band array 402 regarding the time-frequency audio signal 202. In an embodiment, this process generally requires removing the higher sub-bands in the coded rate adjusted sub-band array 402 so that the full bandwidth of the audio signal associated with the coded rate adjusted sub-band array 402 is reduced according to the bandwidth indicated by the audio signal bandwidth input 208.
[0147] The output from the sampling frequency sub-band adjuster can be referred to as the bandwidth adjusted sub-band array 404.
[0148] The reduction in the bandwidth of the coded rate-adjusted subband array 402 (due to the audio signal bandwidth 208) can be performed using a table in which, for each possible input audio signal bandwidth 208, the reduction in the number of subbands from the (full-band) coded rate-adjusted subband array is given.
[0149] Note that the encoder 121 is capable of operating at one of a plurality of different predetermined bandwidths indicated by the audio signal bandwidth signal line 208. For example, the IVAS encoder can be configured to operate at any one of the audio signal bandwidths specified in Table 3.
[0150] In this regard, the following bandwidth adjustment Table 3 shows the relationship between the input audio signal bandwidth 208 and the coded rate-adjusted subband array 402. The various allowed audio signal sampling frequencies / bandwidths are listed along the columns in Table 3, and the coded rate-adjusted subband arrays are listed along the rows of the mapping Table 3. Table 3 provides, for each value of the audio signal bandwidth 208, the number of subbands that need to be removed from the coded rate-adjusted subband array 402 in order to achieve the bandwidth associated with the specified audio signal bandwidth 208. This mapping is given for each combination of the audio signal bandwidth 208 and the coded rate-adjusted subband count array 402. The values specified by Table 3 are based on the number of subbands removed starting from the highest subband in the coded rate-adjusted subband array 402.
[0151]
[0152] Table 3
[0153] The operation mechanism of the bandwidth adjustment Table 3 can be further understood by the above example (wherein, since the IVAS coding rate is 160 kbps, the number of subbands is adjusted from 24 sb to 12 sb). In other words, the coded rate-adjusted subband array for a coding rate of 160 kbps includes 12 frequency subbands. The coded rate-adjusted subband array with 12 sb in turn constitutes one input to the table, and the other input is the bandwidth specified by the audio signal bandwidth 208. Thus, for an example input audio signal bandwidth 208 in the wideband (WB) mode, the table will produce an adjustment factor of 2 sb. That is, two highest frequency subbands are removed from the coded rate-adjusted subband array, thereby giving a bandwidth-adjusted subband array 404 of 10 sb for the combination of the 160 kbps IVAS coding rate and the selected WB bandwidth. Thus, overall, the combination of the total (IVAS) coding rate of 160 kbps (206) and the WB audio signal bandwidth (208) will produce a bandwidth-adjusted subband array 404, the first 10 subbands of which are spread over the 8 kHz bandwidth (16 kHz sampling frequency) of the wideband audio signal.
[0154] In addition to the adjustment obtained from the total coding rate 206 and the audio signal bandwidth 208, the width of the final frequency band for some combinations of the subbands 402 with adjusted coding rate and the audio signal bandwidth 208 can also be considered for further adjustment. For those cases where the highest remaining frequency subband in the bandwidth-adjusted subbands (as indicated by the bandwidth-adjusted subband array 404) is found to exceed the bandwidth associated with the audio signal bandwidth 208, this further adjustment can be applied.
[0155] In Figure 4 FIG. shows this final adjustment process as performed by the highest subband limiter 405, where the highest subband limiter 405 receives the bandwidth-adjusted subband array 404 and the audio signal bandwidth 208, and produces as output the adjusted subband configuration array 302.
[0156] In this regard, Table 3 above discloses the bandwidth according to the number of frequency bins for each possible value of the audio signal bandwidth 208. In addition, in Table 3, the bandwidth according to the subband numbers of the 24 subbands of the original time-frequency audio signal 202 is also shown in the same column. Further, this column can be used to determine whether the last subband of the bandwidth-adjusted subband array 404 exceeds the actual bandwidth allowed by the audio signal bandwidth 208.
[0157] For example, as can be seen from the above table, a narrowband signal (NB) can have a maximum signal bandwidth of 10 frequency bins, and a full-band signal (FB) can have a maximum signal bandwidth of 60 frequency bins. As described above, in some cases, the highest sub-band of the bandwidth-adjusted sub-band array 404 may exceed the actual bandwidth of the audio signal bandwidth (parameter) 208. This situation may be particularly common for NB (narrowband) signals with an actual bandwidth of only 10 frequency bins. For example, if MASA_band_grouping_24_to_12 (which is an adjustment to the number of sub-bands for a total coding rate of 160 kbps) is examined, the highest sub-band has been assigned the frequency bins (frequency bins 40 to 60) corresponding to sub-bands 22 to 24 of the original time-frequency audio signal 202. If this coding rate is further adjusted for a narrowband signal (NB), it can be seen from Table 3 that the four highest sub-bands are removed, leaving the following sub-bands {0, 1, 2, 3, 4, 5, 7, 9, 12}. The last sub-band occupies the frequency sub-bands 9 to 12 relative to the sub-bands of the MASA_band_grouping_24 array. Thus, the last sub-band in this example will exceed the bandwidth of the NB signal of the 10th frequency bin. Clearly, in such a case, it would be advantageous to perform a further adjustment where the last sub-band is trimmed to fall within the actual bandwidth of the audio signal bandwidth 208. In this regard, Table 4 below lists the corresponding sub-band boundaries for each combination of the audio signal bandwidth 208 and the coding rate-adjusted sub-band array 402. It can be seen that some of the entries in Table 4 have trimmed the frequency bins of the highest sub-band to fall within the bandwidth of the audio signal bandwidth (parameter) 208. For clarity, these entries have been marked with an asterisk (*).
[0158] In an embodiment, once the "mode" of the sub-band boundaries has been determined in response to the total (IVAS) coding rate 206 (given by Table 2) and the bandwidth of the audio signal bandwidth input 208 (given by Table 3), which is shown as the bandwidth-adjusted sub-band array 404 in Figure 4 it is possible to further examine the sub-band boundaries of the bandwidth-adjusted sub-band array 404 against Table 4 to determine whether the highest sub-band is to be capped (or limited) to align it with the actual bandwidth of the audio signal bandwidth 208.
[0159]
[0160] Table 4
[0161] In Table 4, the "number of sub-bands" is the initial number of sub-bands before adjustment according to the total (IVAS) coding rate 206 and the audio signal bandwidth 208. Note that the sub-band boundaries are given according to the sub-band count of 24 sub-bands of the time-frequency audio signal 202. In other words, the sub-band boundaries are relative to the original MASA_band_grouping_24. For example, a full-band signal reduced to 5 sub-bands has a mapping according to MASA_band_mapping 24_to_5, where it can be seen that the last sub-band occupies sub-bands 15 to 24 of the original audio signal (of 24 sub-bands), which is equivalent to the highest sub-band occupying frequency bins from 15 to 60. Clearly, for a 32 kHz SWB signal, the bandwidth is 16 kHz (or, according to the time-frequency audio signal 202 of 24 sub-bands, the sub-band count is 23), which is equivalent to the frequency bin width from 30 to 40 of the MASA_band_grouping_24 array (i.e., 12 kHz to 16 kHz). Thus, for the SWB signal, the mapping from 24 sub-bands to 5 sub-bands is restricted to sub-band count 23 (which, according to the array MASA_band_grouping_24, is equivalent to a frequency bin count of 40) to ensure that the signal does not exceed the bandwidth of the SWB signal (16 kHz).
[0162] The output from the highest sub-band limiter 405 is the adjusted sub-band configuration array 302. In an instance where the highest sub-band of the bandwidth sub-band array 404 falls within the bandwidth of the audio bandwidth 208, the adjusted sub-band configuration array 302 will be the bandwidth-adjusted sub-band array 404. In other words, no limiting / capping operation is applied to the highest sub-band. However, in an instance where the highest sub-band of the bandwidth sub-band array 404 extends beyond the bandwidth of the audio signal bandwidth 208, the adjusted sub-band configuration array 302 will be the bandwidth-adjusted sub-band array 404, where the highest sub-band is limited in its width.
[0163] Furthermore, the adjusted sub-band configuration array 302 can be passed to the parameter set combiner / reducer 305, as Figure 3 shown.
[0164] Figure 5Illustrates a computer software or hardware-implemented process of the frequency subband adjuster 303 for determining the adjusted subband configuration array 302. The adjusted subband configuration array 302 is shown to be determined based on the total (IVAS) coding rate 206 and the audio signal bandwidth 208. For clarity, in an embodiment, the adjusted subband configuration array 302 (or vector) may include member values that specify the boundaries of subbands for the parameter set combiner / reducer 305. In fact, the adjusted subband configuration array 302 may be one of the subband boundary arrays from Table 4 above. Further, the parameter set combiner / reducer 305 uses the adjusted subband configuration array 302 to combine adjacent spatial audio parameter sets from adjacent subbands. The parameter set combiner / reducer 305 may also be set to remove spatial audio parameter sets corresponding to frequency subbands greater than those in the adjusted subband configuration array 302. The result of the combining and reducing process is a spatial audio parameter set for the subbands that reflects / mirrors the subband pattern as given by the adjusted subband configuration array 302. In some embodiments, the adjusted subband configuration array 302 may be set to an index or pointer that points to one of the subband boundary arrays in Table 4.
[0165] Return to Figure 5 , the process of determining the adjusted subband configuration array 302 by the frequency subband adjuster 303 is shown to receive an input 206 that includes an indication of the total coding rate (for the IVAS encoder). Processing step 501 describes a mapping step between the received total (IVAS) coding rate 206 and the number of frequency subbands allowed in the coding rate-adjusted subband array 402. This can be performed by using Table 2.
[0166] Figure 5 Processing step 503 therein describes the selection of the MASA_band_mapping array as determined by the number of frequency subbands from step 501. Note that a higher coding rate from Table 2 does not require reducing the number of subbands. The selected MASA_band_mapping array constitutes the coding rate-adjusted subband array 402.
[0167] Figure 5 Processing step 505 therein describes the step of removing a plurality / a certain number of high-frequency subbands from the coding rate subband array 402 in response to the audio signal bandwidth 208. For example, this step can be implemented by using Table 3.
[0168] Processing step 507 describes the process of checking Table 4 to determine whether the highest sub-band of the sub-band array 404 with bandwidth adjustment exceeds the bandwidth of the audio signal sampling frequency 208. If the highest sub-band exceeds the bandwidth, the width of the highest sub-band is adjusted to be within the bandwidth. This step can be performed by using Table 4. The output of this step can be one of the arrays from Table 4, which specifies the sub-band boundaries of the adjusted sub-band configuration array 302.
[0169] According to Figure 5 The processing step has the following advantages: no additional signaling bits need to be sent from the encoder to the decoder. The reason is that the decoder can know both the coding rate and the bandwidth at the encoder through system-level configuration information, and both the encoder and the decoder can access the above tables.
[0170] It will be understood that Figure 3 In combination with Figure 4 It is shown that the set of spatial parameters associated with the original sub-band pattern of the time-frequency audio signal 202 is merged and reduced (as a final stage) according to the sub-band pattern given by the adjusted sub-band configuration array 302. In other words, the merging of the set of spatial audio parameters indicated by the sub-band array 402 adjusted by the coding rate, the reduction of the set of spatial parameters indicated by the sub-band array 404 adjusted by the bandwidth, and the conditional pruning of the highest sub-band indicated by 405 can be performed in the parameter set combiner / reducer 305 as a single processing stage according to the "final" adjusted sub-band configuration array 302.
[0171] However, it will also be understood that in other embodiments, the process of merging and reducing the set of spatial parameters can also be performed sequentially when determining the corresponding sub-band pattern. Therefore, in these embodiments, the merging of the set of sub-band parameters can be performed when determining the sub-band array 402 adjusted by the coding rate. Furthermore, the reduction of the set of spatial parameters can be performed when determining the sub-band array 404 adjusted by the bandwidth after this step. Finally, the spatial parameter set of the highest frequency sub-band can be conditionally pruned by the highest sub-band limiter 405.
[0172] Figure 6 The frequency sub-band adjuster 303 is shown for an embodiment in which the sub-band cut-off signal 306 is deployed according to the energy level of the sub-bands of the audio transmission signal 104.
[0173] In Figure 6In it, the sub-band adjuster 303 is shown as receiving the audio transmission signal 104 by the bin energy determiner 601. The bin energy determiner 601 is configured to measure / determine the energy of the audio signal in each bin of the audio transmission signal 104, in other words, the bin energy 605. Considering that the audio transmission signal 104 can include up to two transmission signals, the bin energy determiner 601 is set to determine the energy in each bin for all transmission signals. The energy calculation can be performed on a per-audio-frame basis.
[0174] Furthermore, the output of the bin energy determiner 601, that is, the bin energy 605 (for each transmission signal) is passed to the frequency sub-band reducer 603 for further processing.
[0175] The frequency sub-band reducer 603 can be set to determine whether any bin energy is below a predetermined energy. This can be performed by scanning the energy of each bin in descending order of the bin index of the bin energy signal 605 and checking the case where the energy of the bin first exceeds the minimum energy level. After determining this index b m the energy cut-off bin index b e can be determined as b m +1. Furthermore, the frequency sub-band reducer 603 can be configured to determine the frequency sub-band k e in which the bin index b e is located. This is determined as the cut-off frequency sub-band above which the contribution of the audio transmission signal 104 to the final multi-channel spatial audio signal 110 is considered negligible. In other words, any sub-band above index k e and above is considered to have insufficient energy level, and thus, the set of spatial parameters associated with these sub-bands can be effectively removed by not encoding them.
[0176] Regarding the case where the audio transmission signal 104 has more than one channel, the above process can be performed sequentially for each channel, so that multiple bin indices (b e1 , b e2 ....) (one for each channel) can be found. The highest bin index is selected, and the frequency sub-band associated with the highest bin index can be determined as the cut-off frequency sub-band index k e for all channels of the audio transmission signal 104.
[0177] The cut-off frequency sub-band index k e can be transmitted as the signal 306 to the parameter set combiner / reducer 305. k e >B W (k)
[0178] k e <BW (k)
[0179] When the parameter set combiner / reducer 305 receives the signal 306, it can be set to remove all sets of spatial parameters associated with all frequency sub-bands k e and above. That is, all sets of parameters associated with the frequency sub-bands k e to K-1 (where K-1 is the highest sub-band index associated with the audio transmission signal 104) are set to zero (or removed) and will thus not form part of the spatial metadata signal 106 that is passed to the metadata encoder / quantizer 111.
[0180] Furthermore, the remaining sets of spatial parameters of the spatial audio metadata 106 can be encoded by the spatial parameter set encoder 207 using the techniques described in patent application EP3818525. In addition, the spatial parameter set encoder 207 can also be set to use a zero-order Golomb Rice code to encode the number of sub-bands that do not contain the encoded set of spatial parameters (the number of sub-bands from k e to K-1).
[0181] It should be noted that for the case where the energy levels of all frequency bins are higher than a predetermined energy level, no set of spatial parameters will be removed from the spatial metadata signal 106. This case can be signaled using a single bit.
[0182] Thus, in this embodiment, the encoded spatial metadata information can include the encoded set of spatial parameters and an additional signaling bit. Wherein, one state of the signaling bit indicates that the encoded spatial metadata 106 includes the encoded set of spatial audio parameters for all frequency bands, and the other state of the signaling bit indicates that a partial number of the frequency band spatial parameter sets of the spatial metadata 106 have been encoded.
[0183] In another embodiment, the spatial parameter set encoder 207 can be set to cancel the single bit indicating that no set of spatial parameters has been removed from the spatial metadata signal 106. Instead, in the special case where the number of sub-band spatial parameter sets is less than the total number of sub-bands and the direct-to-total energy ratio of the spatial audio parameter sets associated with the sub-bands k e to K-1 (the remaining sub-bands) is quantized to the minimum quantization level, a single bit is added to the encoded stream.
[0184] It should be noted that the quantization and coding of the energy ratio values can be performed separately from the quantization and coding of the other spatial audio parameters in the set of subband spatial audio parameters. Thus, each subband can have at least an associated quantized energy ratio. And the other parameters in the set of spatial audio parameters associated with that subband do not need to be quantized and coded (and thus, do not form part of the coded bitstream).
[0185] For example, the energy ratio values (for each subband) can be quantized with a 3-bit scalar quantizer, and the other spatial audio parameters in the set of spatial audio parameters for the subband can be quantized and coded according to the disclosure EP3818525.
[0186] Regarding other embodiments, Figure 7 a further process for quantizing the set of subband spatial parameters is shown when the subbands of the audio transmission signal 104 are considered to have sufficiently low energy so as not to contribute to the synthesized multi-channel spatial audio signal 110.
[0187] As Figure 7 shown, the process starts with receiving a k e value associated with K-1 frequency subbands of a subframe.
[0188] First, the subband cut-off value k e is checked to determine whether k e <K-1. As described above, this indicates that the set of spatial audio parameters for frequency subbands k e to K-1 can be removed from the metadata encoding process performed by 111. This is shown in Figure 7 by processing step 701.
[0189] If at step 701 it is determined that k e <K-1, then the processing path 702 is taken according to Figure 7 .
[0190] Furthermore, the processing path 702 sets the energy ratio associated with the frequency subband k e <K-1 to have the minimum quantization level. This is shown in Figure 7 as processing step 703.
[0191] Then, the energy ratios associated with the frequency subbands 0 to k e -1 are quantized according to their values. As described above, this can be performed with a scalar quantizer, thereby generating a quantization index (or codeword) for each energy ratio value. This is shown in Figure 7 as processing step 705.
[0192] Furthermore, a single bit can be appended to the bitstream (for a subframe) to indicate that the number of sets of coded spatial audio parameters encoded in the bitstream for that subframe is not a full complement for subband K-1. This is shown in Figure 7 as processing step 707.
[0193] The number of frequency subbands without any associated set of spatial parameters, that is, subband k e < K-1, is encoded using a Golomb Rice code of order 0. This is shown in Figure 7 as processing step 709. Clearly, this encoded number of frequency subbands also forms part of the encoded bitstream for the frame.
[0194] Finally, the "other" spatial audio parameters in the sets of spatial audio parameters for subbands 0 to k e -1 can be quantized and encoded according to publications WO2022 / 129672, WO2021 / 048468, WO2020 / 070377, WO2020 / 008105 and WO2021 / 144498. As described above, these quantized sets of spatial audio parameters can also form part of the encoded bitstream for the frame. This is shown in Figure 7 as processing step 711.
[0195] It should be clear that, in this context, the term "other" spatial audio parameters refers to the spatial audio parameters in the set of spatial audio parameters (for a subband) that do not include the above-mentioned energy ratio.
[0196] Returning to Figure 7 the determination step 701 in, it can be seen that the result of this step determines whether all frequency subbands are to have their corresponding sets of spatial parameters encoded. In other words, the frequency subband reducer 603 effectively determines that all subbands are above the minimum energy level by returning a value k e = K-1 in at least one way. It should be understood that those skilled in the art will recognize that other ways / means can be used to signal this condition. Furthermore, the process can be set to take processing path 704.
[0197] Once it is decided to take processing path 704, the parameter set combiner / reducer 305 can be set to quantize and encode the energy ratios corresponding to all frequency subbands 0 to K-1. This is shown in Figure 7 as processing step 713.
[0198] Furthermore, the process determines whether the energy ratio associated with the last subband (K-1) has been quantized to the minimum quantization level. This decision step is shown as Figure 7Processing step 715 therein. When it is determined that the result of step 715 indicates that the energy ratio associated with the last subband (K - 1) has not been quantized to the minimum quantization level, the process is set to continue to processing step 717, where the set of spatial parameters associated with all subbands 0 to K - 1 is quantized and encoded.
[0199] However, when it is determined that step 715 indicates that the energy ratio associated with the last subband (K - 1) has been quantized to the minimum quantization level, the process is set to continue to processing step 719. At processing step 719, a single bit is appended to the encoded stream (for the frame). The state of this bit (shown as being set to 0 in Figure 7 indicates that although the energy ratio associated with the last subband K - 1 has been encoded to the minimum quantization level, all K - 1 sets of spatial parameters have been encoded.
[0200] Finally, after step 719, Figure 7 the process of moving to processing step 717 is shown, but before the set of spatial parameters associated with all subbands 0 to K - 1 is quantized and encoded.
[0201] It should be clear that the bitstream of processing path 702 can include, for each frame, at least the encoded and quantized energy ratios associated with bands 0 to k e - 1, the energy ratios associated with subbands k e to K - 1 that are quantized and encoded to the minimum quantization level, a bit signaling the number of encoded sets of spatial parameters < K - 1, a zero - order GR code indicating the number of subbands given by the values k e to K - 1, and the quantized and encoded sets of spatial data associated with subbands 0 to k e - 1 (each set including other spatial parameters for the energy ratio).
[0202] Continuing from the above, the bitstream of processing path 704 can include, for each frame, at least the encoded and quantized energy ratios associated with bands 0 to K - 1, and the quantized and encoded sets of spatial parameters associated with subbands 0 to K - 1 (each set including other spatial parameters for the energy ratio). In addition, the bitstream of processing path 704 can also include a bit signaling the number of encoded sets of spatial parameters corresponding to the subbands from 0 to K - 1 in the case where the energy ratio associated with the last band K - 1 is quantized to the minimum level.
[0203] The above - mentioned embodiments can be executed on a per - frame basis. In addition, a second embodiment can be deployed in combination with the first embodiment frame - by - frame. For example, it can be decided at the start of a new frame whether to use the first embodiment or the second embodiment.
[0204] It will be understood that in some other embodiments, the above energy-based embodiments may be combined with embodiments that previously employed processing steps according to Figure 5 and performed. In other words, the above energy-based embodiments may be integrated into embodiments in which a set of spatial parameters associated with subbands of the time-frequency audio signal 202 are combined and reduced in response to the total coding rate 206 and the audio signal bandwidth 208.
[0205] In this regard, Figure 8 shows how the above energy-based embodiments can be implemented in a system that deploys a previous embodiment according to Figure 5 .
[0206] The processing steps 801 and 803 may be set up according to Figure 5 , where the selected total (IVAS) coding rate 206 is received, and on this basis, the coded rate-adjusted subband array 402 can be obtained by determining the MASA band mapping array.
[0207] The processing step 803 is set to receive the bandwidth B W , which is the specified audio signal bandwidth 208. For example, this may be given by Table 3, where various allowed sampling frequencies are listed as a function of the number of subbands k.
[0208] Furthermore, the processing step 805 compares the audio signal bandwidth B W 208 with the cut-off frequency subband index k e 306.
[0209] Then, Figure 8 continuing to show that when the cut-off frequency subband index k e is greater than (or equal to) the bandwidth B W , the process proceeds to step 807, where a processing step similar to step 505 is performed. In other words, the processing step 807 performs the process of removing high-frequency subbands from the coded rate subband array 402 in response to the audio signal bandwidth 208B W . Thus, the processing step 807 is shown to also receive the coded rate-adjusted subband array 402 from the processing step 803, as well as the audio signal bandwidth 208. Therefore, the result of this processing step is the bandwidth-adjusted subband array 404.
[0210] Alternatively, the comparison step 803 may determine that the energy-based cut-off frequency subband index k e is less than the bandwidth B W . When this condition is met, the process may be set to transition to step 809. At step 809, the process is set to remove those higher than the index k eThe frequency sub-bands of the cut-off frequency sub-bands. In other words, processing step 809 takes the coded rate-adjusted sub-band array 402 and removes those sub-bands whose indices are higher than the cut-off index k e of the sub-bands. Thus, the result of this processing step can be regarded as a version of the bandwidth-adjusted sub-band array 404, where the higher sub-bands are limited according to the cut-off index 306. Thus, processing step 809 is shown as also receiving the coded rate-adjusted sub-band array 402 from processing step 803, as well as the cut-off frequency sub-band index k e 306, thus allowing the formation of the above-mentioned variant of the bandwidth-adjusted sub-band array 404.
[0211] Figure 8 It is shown that the output from step 807 (bandwidth-adjusted sub-band array 404) is passed to processing step 811. Step 811 performs a processing function similar to that of Figure 5 step 505 in. In other words, step 811 performs the process of determining whether the highest sub-band of the bandwidth-adjusted sub-band array 404 exceeds the audio signal bandwidth 208, and if it is found that the highest sub-band exceeds the audio signal bandwidth 208, the width of the highest sub-band is adjusted to be within this bandwidth. The output from step 811 is the adjusted sub-band configuration array 302.
[0212] Processing step 813 is shown as receiving the bandwidth-adjusted sub-band array 404 from step 809. In a similar manner to step 811, processing step 813 can be set to perform the process of determining whether the highest sub-band of the bandwidth-adjusted sub-band array 404 exceeds the cut-off frequency sub-band 306 with an index of k e . If it is determined that this is the case, the width of the highest sub-band is adjusted to be within the bandwidth of the cut-off frequency sub-band index k e 306. The output from step 813 is also another variant of the adjusted sub-band configuration array 302.
[0213] In addition, Figure 8 it is also shown that when following the processing path 805, the cut-off frequency sub-band index k e 306 can be encoded. This is shown as processing step 815, where the encoding of the cut-off frequency sub-band index k Figure 7 can be performed according to the processing steps of e 306.
[0214] Regarding Figure 9 , an example electronic device that can be used as an analysis or synthesis device is shown. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.
[0215] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. The processor 1407 may be configured to execute various program codes, such as the methods described herein.
[0216] In some embodiments, device 1400 includes a memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 may be any suitable storage component. In some embodiments, the memory 1411 includes a program code portion for storing program code that can be implemented on the processor 1407. Additionally, in some embodiments, the memory 1411 may further include a stored data portion for storing data (e.g., data that has been processed or will be processed according to the embodiments described herein). The implemented program code stored within the program code portion and the data stored within the stored data portion may be retrieved by the processor 1407 via the memory-processor coupling when needed.
[0217] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, the user interface 1405 may be coupled to the processor 1407. In some embodiments, the processor 1407 may control the operation of the user interface 1405 and receive input from the user interface 1405. In some embodiments, the user interface 1405 may enable a user to input commands to the device 1400, for example, via a keyboard. In some embodiments, the user interface 1405 may enable a user to obtain information from the device 1400. For example, the user interface 1405 may include a display configured to display information from the device 1400 to the user. In some embodiments, the user interface 1405 may include a touch screen or touch interface that enables information to be input into the device 1400 and also displays information to the user of the device 1400. In some embodiments, the user interface 1405 may be a user interface for communicating with the position determiner described herein.
[0218] In some embodiments, device 1400 includes an input / output port 1409. In some embodiments, the input / output port 1409 includes a transceiver. In such embodiments, the transceiver may be coupled to the processor 1407 and configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver components may be configured to communicate with other electronic devices or apparatuses via a wired or wired coupling.
[0219] The transceiver can communicate with other devices via any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol (such as, for example, IEEE 802.X), a suitable short-range radio frequency communication protocol (such as Bluetooth), or an Infrared Data Association (IrDA) communication path.
[0220] The transceiver input / output port 1409 can be configured to receive signals and, in some embodiments, determine the parameters as described herein by executing suitable code using the processor 1407. Additionally, the device can generate suitable down-converted signals and parameter outputs to be sent to a synthesizing device.
[0221] In some embodiments, the device 1400 can be used as at least a part of a synthesizing device. Thus, the input / output port 1409 can be configured to receive down-converted signals and, in some embodiments, receive the parameters determined at a capture device or a processing device as described herein, and generate a suitable audio signal format output by executing suitable code using the processor 1407. The input / output port 1409 can be coupled to any suitable audio output, for example, coupled to a multi-channel speaker system and / or headphones or similar devices.
[0222] In general, various embodiments of the present invention can be implemented using hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects can be implemented using hardware, while other aspects can be implemented using firmware or software executable by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although the various aspects of the present invention can be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it is well known that the blocks, devices, systems, techniques, or methods described herein can be implemented as non-limiting examples using hardware, software, firmware, dedicated circuits or logic, general hardware or a controller or other computing devices, or some combination thereof.
[0223] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. Additionally, in this regard, it should be noted that any block of the logical flow in the figures can represent a program step, or an interconnected logical circuit, block, and function, or a combination of program steps and logical circuits, blocks, and functions. The software can be stored on a physical medium such as a memory chip or a memory block implemented within a processor, a magnetic medium such as a hard disk or a floppy disk, and an optical medium such as a DVD and its data variants, a CD.
[0224] The memory can be of any type suitable for the local technical environment and can be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and, by way of non-limiting example, can include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), gate-level circuitry, and a processor based on a multi-core processor architecture.
[0225] Embodiments of the present invention may be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Sophisticated and powerful software tools are available to transform a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0226] The program can automatically route conductors and place components on a semiconductor chip using well-established design rules and a library of pre-stored design modules. Once the semiconductor circuit design is complete, the resulting design in a standardized electronic format can be transferred to a semiconductor manufacturing facility or "fab" for fabrication.
[0227] The foregoing description has provided a complete and beneficial description of exemplary embodiments of the present invention by way of example and not limitation. However, various modifications and adaptations will become apparent to those skilled in the relevant art in view of the above description when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of the present invention will still fall within the scope of the present invention as defined by the appended claims.
Claims
1. An apparatus for spatially audio encoding one or more audio signals, wherein, The apparatus includes components configured to perform the following operations: For each of a plurality of frequency sub-bands of the one or more audio signals, determine a set of spatial audio parameters; Receive a coding rate associated with the one or more audio signals; Based on the coding rate, map at least two consecutive sub-bands of the plurality of frequency sub-bands to an extended frequency sub-band to give a plurality of frequency sub-bands adjusted by the coding rate; Receive a bandwidth value associated with the one or more audio signals; Starting from the highest frequency sub-band of the plurality of frequency sub-bands adjusted by the coding rate, remove a certain number of frequency sub-bands to give a plurality of frequency sub-bands adjusted by the bandwidth, wherein the number of removed frequency sub-bands is based on the bandwidth value associated with the one or more audio signals; Under the condition that the highest frequency sub-band of the plurality of frequency sub-bands adjusted by the bandwidth exceeds the bandwidth value associated with the one or more audio signals, reduce the highest frequency sub-band of the plurality of frequency sub-bands adjusted by the bandwidth to the bandwidth value or below the bandwidth value; Combine the set of spatial audio parameters associated with the first frequency sub-band of the at least two consecutive frequency sub-bands and the set of spatial audio parameters associated with the second frequency sub-band of the at least two consecutive frequency sub-bands to give a combined set of spatial audio parameters for the extended frequency sub-band; Remove the set of spatial audio parameters corresponding to each removed frequency sub-band; and Under the condition that the highest frequency sub-band of the plurality of frequency sub-bands adjusted by the bandwidth exceeds the bandwidth value, remove the set of spatial audio parameters associated with the plurality of frequency sub-bands adjusted by the bandwidth that exceed the bandwidth value.
2. The apparatus for spatial audio coding according to claim 1, wherein, The highest frequency sub-band of the plurality of frequency sub-bands adjusted by the bandwidth includes the upper sub-band boundary value and the lower sub-band boundary value of a plurality of frequency sub-bands in the plurality of frequency sub-bands including the one or more audio signals, and wherein the component configured to reduce the highest frequency sub-band of the plurality of frequency sub-bands adjusted by the bandwidth to the bandwidth value or below the bandwidth value includes a component configured to perform the following operations: Adjust the upper sub-band boundary value within the bandwidth value; and wherein the component configured to remove the set of spatial audio parameters associated with the plurality of frequency sub-bands adjusted by the bandwidth that exceed the bandwidth value includes a component configured to perform the following operations: Remove the set of spatial audio parameters associated with the plurality of frequency sub-bands in the plurality of frequency sub-bands of the one or more audio signals that are higher than the adjusted upper sub-band boundary value.
3. The apparatus for spatial audio coding according to claim 1 and 2, wherein, The apparatus including the component configured to map at least two consecutive sub-bands of the plurality of frequency sub-bands to an extended frequency sub-band based on the coding rate to give a plurality of frequency sub-bands adjusted by the coding rate includes a component configured to perform the following operations: Map the high-band boundary value and the low-band boundary value for the at least two consecutive frequency sub-bands of the plurality of frequency sub-bands to the low-band boundary value and the high-band boundary value of the extended frequency sub-band.
4. The apparatus for spatial audio coding according to claim 3, wherein, The low-frequency sub-band boundary value and the high-frequency sub-band boundary value of the extended frequency sub-bands are given by the low-frequency sub-band boundary value and the high-frequency sub-band boundary value of a frequency sub-band reduction array, the frequency sub-band reduction array including a plurality of frequency sub-band boundaries in ascending order of frequency sub-bands, wherein a sub-band boundary value and the next higher sub-band boundary value in ascending order in the frequency sub-band reduction array are the low-frequency sub-band boundary value and the high-frequency sub-band boundary value of the extended frequency sub-bands, respectively.
5. The apparatus for spatial audio coding according to claim 4, wherein, The plurality of frequency sub-band boundaries in the frequency sub-band reduction array constitute fewer frequency sub-bands than the plurality of frequency sub-bands of the one or more audio signals, and wherein the plurality of frequency sub-bands with adjusted coding rate are given by the frequency sub-band reduction array, wherein the frequency sub-band reduction array is selected from a plurality of frequency sub-band reduction arrays, wherein the selection is based on the coding rate associated with the one or more audio signals, and wherein each frequency sub-band reduction array in the plurality of frequency sub-band reduction arrays includes a different number of frequency sub-bands, and wherein each frequency sub-band reduction array in the plurality of frequency sub-band reduction arrays is associated with a different coding rate associated with the one or more audio signals.
6. The apparatus for spatial audio coding according to claims 1 to 5, wherein, A certain number of frequency sub-bands to be removed are selected from a plurality of certain numbers of frequency sub-bands to be removed, wherein the selection is based on the bandwidth value, and wherein each of the plurality of certain numbers of frequency sub-bands to be removed is associated with a different bandwidth value.
7. The apparatus for spatial audio coding according to any one of claims 1 to 6, wherein, The plurality of frequency sub-bands with adjusted sampling frequency are in the form of an array including a plurality of frequency sub-band boundary values in ascending order of frequency sub-bands.
8. The apparatus for spatial audio coding according to any one of claims 1 to 7, wherein, The apparatus includes a first encoder and a second encoder for encoding the one or more audio signals at the coding rate, wherein the coding rate includes the sum of the coding rate for the first encoder and the coding rate for the second encoder, wherein the first encoder encodes an audio transmission signal associated with the one or more audio signals, and the second encoder encodes the plurality of sets of spatial audio parameters associated with the frequency sub-bands of the one or more audio signals.
9. An apparatus for spatially audio encoding one or more audio signals, wherein, The apparatus includes components for performing the following operations: Determining, for each frequency sub-band of the plurality of frequency sub-bands of the one or more audio signals, a set of spatial audio parameters; Receiving a coding rate associated with the one or more audio signals; Based on the coding rate, mapping at least two consecutive sub-bands of the plurality of frequency sub-bands to an extended frequency sub-band to give a plurality of frequency sub-bands with adjusted coding rate; Combining the set of spatial audio parameters associated with the first frequency sub-band of the at least two consecutive frequency sub-bands and the set of spatial audio parameters associated with the second frequency sub-band of the at least two consecutive frequency sub-bands to give a combined set of spatial audio parameters for the extended frequency sub-band; Determining, for each frequency bin of the one or more audio signals, an energy level; The cut-off frequency sub-band of the one or more audio signals is determined by: determining the highest frequency bin having an energy level greater than a predetermined energy level, and designating the cut-off frequency sub-band as the frequency sub-band containing the highest frequency bin; Comparing the cut-off frequency sub-band of the one or more audio signals with the bandwidth value of the one or more audio signals; Under the condition that the cut-off frequency sub-band is less than the bandwidth value of the one or more audio signals, starting from the highest frequency sub-band among the plurality of frequency sub-bands adjusted by the coding rate, a certain number of frequency sub-bands are removed to give a plurality of frequency sub-bands adjusted by the bandwidth, and a set of spatial audio parameters corresponding to each removed frequency sub-band is removed, wherein the number of removed frequency sub-bands is based on the cut-off frequency sub-band; Under the condition that the highest frequency sub-band among the plurality of frequency sub-bands adjusted by the bandwidth exceeds the cut-off frequency sub-band, reducing the highest frequency sub-band among the plurality of frequency sub-bands adjusted by the bandwidth to the cut-off frequency sub-band value or a value lower than the cut-off frequency sub-band value, and removing the set of spatial audio parameters associated with the plurality of frequency sub-bands adjusted by the bandwidth that exceed the cut-off frequency sub-band; and Encoding the index of the cut-off frequency sub-band.
10. The apparatus for spatial audio coding according to claim 9, wherein, The component configured to encode the index of the cut-off frequency sub-band is further configured to: encode each set of spatial audio parameters associated with frequency sub-bands lower than the cut-off frequency sub-band.
11. The device according to claim 10, wherein The component configured to encode each set of spatial audio parameters associated with frequency sub-bands lower than the cut-off frequency sub-band further includes a component configured to perform the following operations: Determining an energy ratio parameter for each frequency sub-band among the plurality of frequency sub-bands of the one or more audio signals; Quantizing the energy ratio for each frequency sub-band among the plurality of frequency sub-bands that is greater than or equal to the cut-off frequency sub-band to a minimum quantization level; Quantizing the energy ratio for each frequency sub-band among the plurality of frequency sub-bands that is less than the cut-off frequency sub-band; And Encoding an indication that the number of encoded sets of spatial audio parameters is less than the number of frequency sub-bands of the one or more audio signals; And Encoding the plurality of unencoded sets of spatial audio parameters using a Golomb-Rice code.
12. The apparatus for spatial audio coding according to claims 9 to 11, wherein, The highest frequency sub-band among the plurality of frequency sub-bands adjusted by the bandwidth includes the upper sub-band boundary value and the lower sub-band boundary value of a plurality of frequency sub-bands among the plurality of frequency sub-bands of the one or more audio signals, and wherein the component configured to reduce the highest frequency sub-band among the plurality of frequency sub-bands adjusted by the bandwidth to the cut-off frequency sub-band value or a value lower than the cut-off frequency sub-band value, and remove the set of spatial audio parameters associated with the plurality of frequency sub-bands adjusted by the bandwidth that exceed the cut-off frequency sub-band includes a component configured to perform the following operations: Adjusting the upper sub-band boundary value within the cut-off frequency sub-band value; and Remove a set of spatial audio parameters associated with a plurality of frequency subbands in the one or more audio signals that are higher than an adjusted upper subband boundary value.
13. The apparatus for spatial audio coding according to claims 9 to 12, wherein, The apparatus including components configured to map at least two consecutive subbands in the plurality of frequency subbands to extended frequency subbands based on the coding rate to give a plurality of coding rate adjusted frequency subbands includes components configured to perform the following operations: Map a high band boundary value and a low band boundary value for the at least two consecutive frequency subbands in the plurality of frequency subbands to a low band boundary value and a high band boundary value of the extended frequency subbands.
14. The apparatus for spatial audio coding according to claim 13, wherein, The low subband boundary value and the high subband boundary value of the extended frequency subbands are given by a low subband boundary value and a high subband boundary value of a frequency subband reduction array, the frequency subband reduction array including a plurality of frequency subband boundaries in increasing order of frequency subbands, wherein a subband boundary value and a next higher subband boundary value in increasing order in the frequency subband reduction array are the low subband boundary value and the high subband boundary value of the extended frequency subbands, respectively.
15. The apparatus for spatial audio coding according to claim 14, wherein, The plurality of frequency subband boundaries in the frequency subband reduction array constitute fewer frequency subbands than the plurality of frequency subbands of the one or more audio signals, and wherein the plurality of coding rate adjusted frequency subbands are given by the frequency subband reduction array, wherein the frequency subband reduction array is selected from a plurality of frequency subband reduction arrays, wherein the selection is based on the coding rate associated with the one or more audio signals, and wherein each frequency subband reduction array in the plurality of frequency subband reduction arrays includes a different number of frequency subbands, and wherein each frequency subband reduction array in the plurality of frequency subband reduction arrays is associated with a different coding rate associated with the one or more audio signals.
16. The apparatus for spatial audio coding according to claims 9 to 15, wherein, The plurality of frequency subbands adjusted by a sampling frequency are in the form of an array including a plurality of frequency subband boundary values in increasing order of frequency subbands.
17. The spatial audio encoding apparatus according to claims 9 to 16, wherein, The apparatus includes a first encoder and a second encoder for encoding the one or more audio signals at the coding rate, wherein the coding rate includes the sum of a coding rate for the first encoder and a coding rate for the second encoder, wherein the first encoder encodes an audio transmission signal associated with the one or more audio signals, and the second encoder encodes the plurality of sets of spatial audio parameters associated with the frequency subbands of the one or more audio signals.
18. A method for spatially audio encoding one or more audio signals, wherein, The method includes: Determine a set of spatial audio parameters for each frequency subband in a plurality of frequency subbands of the one or more audio signals; Receive a coding rate associated with the one or more audio signals; Based on the coding rate, map at least two consecutive subbands in the plurality of frequency subbands to extended frequency subbands to give a plurality of coding rate adjusted frequency subbands; Receive a bandwidth value associated with the one or more audio signals; Starting from the highest frequency sub-band among the plurality of frequency sub-bands with adjusted coding rate, a certain number of frequency sub-bands are removed to give a plurality of frequency sub-bands with adjusted bandwidth, wherein the number of removed frequency sub-bands is based on the bandwidth value associated with the one or more audio signals; Under the condition that the highest frequency sub-band among the plurality of frequency sub-bands with adjusted bandwidth exceeds the bandwidth value associated with the one or more audio signals, the highest frequency sub-band among the plurality of frequency sub-bands with adjusted bandwidth is reduced to the bandwidth value or below the bandwidth value; The set of spatial audio parameters associated with the first frequency sub-band among the at least two consecutive frequency sub-bands and the set of spatial audio parameters associated with the second frequency sub-band among the at least two consecutive frequency sub-bands are combined to give a combined set of spatial audio parameters for the extended frequency sub-band; The set of spatial audio parameters corresponding to each removed frequency sub-band is removed; and Under the condition that the highest frequency sub-band among the plurality of frequency sub-bands with adjusted bandwidth exceeds the bandwidth value, the set of spatial audio parameters associated with the plurality of frequency sub-bands with adjusted bandwidth that exceed the bandwidth value is removed.
19. A method for spatially audio encoding one or more audio signals, wherein, The method includes: For each frequency sub-band among the plurality of frequency sub-bands of the one or more audio signals, determining a set of spatial audio parameters; Receiving a coding rate associated with the one or more audio signals; Based on the coding rate, mapping at least two consecutive sub-bands among the plurality of frequency sub-bands to extended frequency sub-bands to give a plurality of frequency sub-bands with adjusted coding rate; The set of spatial audio parameters associated with the first frequency sub-band among the at least two consecutive frequency sub-bands and the set of spatial audio parameters associated with the second frequency sub-band among the at least two consecutive frequency sub-bands are combined to give a combined set of spatial audio parameters for the extended frequency sub-band; For each frequency bin of the one or more audio signals, determining an energy level; Determining the cut-off frequency sub-band of the one or more audio signals by: determining the highest frequency bin having an energy level greater than a predetermined energy level, and designating the cut-off frequency sub-band as the frequency sub-band containing the highest frequency bin; Comparing the cut-off frequency sub-band of the one or more audio signals with the bandwidth value of the one or more audio signals; Under the condition that the cut-off frequency sub-band is less than the bandwidth value of the one or more audio signals, starting from the highest frequency sub-band among the plurality of frequency sub-bands with adjusted coding rate, a certain number of frequency sub-bands are removed to give a plurality of frequency sub-bands with adjusted bandwidth, and the set of spatial audio parameters corresponding to each removed frequency sub-band is removed, wherein the number of removed frequency sub-bands is based on the cut-off frequency sub-band; Under the condition that the highest frequency sub-band among the bandwidth-adjusted multiple frequency sub-bands exceeds the cut-off frequency sub-band, reducing the highest frequency sub-band among the bandwidth-adjusted multiple frequency sub-bands to the cut-off frequency sub-band value or lower than the cut-off frequency sub-band value, and removing the set of spatial audio parameters associated with the bandwidth-adjusted multiple frequency sub-bands that exceed the cut-off frequency sub-band; and Encoding the index of the cut-off frequency sub-band.
Citation Information
Patent Citations
Analysis of spatial metadata from multi-microphones having asymmetric geometry in devices
EP3542546A1
Determination of spatial audio parameter encoding and associated decoding
EP3818525A1
Spatial audio processing apparatus
WO2017005978A1
Determination of spatial audio parameter encoding and associated decoding
WO2020008105A1
Selection of quantisation schemes for spatial audio parameter encoding
WO2020070377A1