Determining frequency subbands for spatial acoustic parameters
By adjusting and combining spatial audio parameter sets across frequency subbands to align bandwidths, the apparatus optimizes immersive audio encoding, addressing inefficiencies in existing codecs and reducing bitrate requirements for spatial metadata.
Patent Information
- Application Number
- JP2025529763
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2026-02-04
AI Technical Summary
Existing immersive audio codecs like IVAS face issues with mismatched bandwidths between transport audio and spatial metadata streams, leading to unnecessary encoding of metadata sets and inefficiencies, particularly in subbands with minimal contribution to the overall spatial audio signal.
The apparatus and method involve adjusting frequency subbands by combining spatial audio parameter sets, removing excess subbands, and encoding these adjustments to align with a specified bandwidth, thereby optimizing the encoding process for immersive audio signals.
This approach reduces unnecessary encoding, aligns bandwidths, and minimizes bitrate requirements for spatial metadata, enhancing the efficiency and effectiveness of immersive audio encoding and decoding processes.
Smart Images

Figure 2026504248000001_ABST
Abstract
Description
[Technical Field]
[0001] This application relates to an apparatus and method for modifying the bandwidth of a spatial audio signal. [Background technology]
[0002] Immersive audio codecs have been implemented that support multiple operating points ranging from low bitrate operation to transparency. One example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is designed for use over communication networks such as 3GPP 4G / 5G networks, including for use in immersive services such as immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general audio. It is also expected to support channel-based and scene-based audio inputs, including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services and support high error resilience under various transmission conditions.
[0003] Metadata-assisted spatial audio (MASA) is one input format for IVAS. It uses an audio signal along with corresponding spatial metadata. The spatial metadata defines the spatial aspects of the audio signal and includes parameters that may include, for example, direction within a frequency band and the direct-to-total energy ratio. The MASA stream can be obtained, for example, by capturing spatial audio using the microphones of a suitable capture device. For example, a mobile device with multiple microphones can be configured to capture microphone signals, and a set of spatial metadata can be estimated based on the captured microphone signals. The MASA stream can also be obtained from other sources, such as a specific spatial audio microphone (such as an Ambisonics or array microphone), a studio mix (e.g., a 5.1 audio channel mix), or other content using a suitable format conversion.
[0004] An audio signal input to an immersive audio codec (such as IVAS) may be simultaneously encoded as 1 to N audio signals to provide a transport audio stream and analyzed to provide a MASA metadata stream. In such a configuration, the analysis and encoding for the MASA metadata stream may be performed separately from the encoding for the transport audio stream. This may result in unnecessary encoding of some MASA metadata sets, particularly for subbands of the transport audio stream that have minimal contribution to the overall composite spatial audio signal. Furthermore, this may result in a mismatch between the bandwidth at which the input audio signal is encoded for the transport audio stream and the bandwidth at which the input audio signal is analyzed for the MASA metadata stream. Summary of the Invention
[0005] According to a first aspect, there is provided an apparatus for spatial audio coding, the apparatus comprising: means for determining a spatial audio parameter set for each of a plurality of frequency subbands of one or more audio signals; receiving a coding rate associated with the one or more audio signals; mapping at least two consecutive subbands of the plurality of frequency subbands to an extended frequency subband based on the coding rate to provide a plurality of bandwidth-adjusted frequency subbands; receiving a bandwidth value associated with the one or more audio signals; removing a number of frequency subbands starting from a highest frequency subband of the plurality of bandwidth-adjusted frequency subbands to provide a plurality of bandwidth-adjusted frequency subbands, the number of frequency subbands removed being based on the bandwidth value associated with the one or more audio signals; and removing a number of frequency subbands starting from a highest frequency subband of the plurality of bandwidth-adjusted frequency subbands to provide a plurality of bandwidth-adjusted frequency subbands, the number of frequency subbands removed being based on the bandwidth value associated with the one or more audio signals. an apparatus for spatial audio coding, the apparatus comprising: means configured to: reduce a highest frequency subband of the at least two consecutive frequency subbands so that it lies below a bandwidth value; combine a spatial audio parameter set associated with a first of the at least two consecutive frequency subbands with a spatial audio parameter set associated with a second of the at least two consecutive frequency subbands to provide an combined spatial audio parameter set for the expanded frequency subband; remove a spatial audio parameter set corresponding to each removed frequency subband; and remove spatial audio parameter sets associated with a plurality of bandwidth-adjusted frequency subbands that extend beyond the bandwidth value, provided that a highest frequency subband of the bandwidth-adjusted plurality of frequency subbands extends beyond the bandwidth value.
[0006] The highest frequency subband of the bandwidth-adjusted plurality of frequency subbands may include upper and lower subband boundary values that encompass more than one of the plurality of frequency subbands of the one or more acoustic signals, and the means configured to reduce the highest frequency subband of the bandwidth-adjusted plurality of frequency subbands to be below the bandwidth value may include means configured to adjust the upper subband boundary value to be within the bandwidth value, and the means configured to remove spatial acoustic parameter sets associated with the bandwidth-adjusted plurality of frequency subbands that extend beyond the bandwidth value includes means configured to remove spatial acoustic parameter sets associated with more than one of the plurality of frequency subbands of the one or more acoustic signals that are above the adjusted upper subband boundary value.
[0007] The means configured to map at least two consecutive subbands of the plurality of frequency subbands to the expanded frequency subband to provide a plurality of code rate adjusted frequency subbands based on the code rate may include means configured to map upper and lower frequency band boundary values for the at least two consecutive frequency subbands of the plurality of frequency subbands to lower and upper frequency band boundary values of the expanded frequency subband.
[0008] The lower frequency subband boundary value and upper frequency subband boundary value of the expanded frequency subband may be given by the lower frequency band boundary value and upper frequency band boundary value of a frequency subband reduction array that includes a plurality of frequency subband boundaries in ascending order of the frequency subbands, and the subband boundary value and the next upper subband boundary value in ascending order of the frequency subband reduction array are the lower frequency subband boundary and upper frequency subband boundary, respectively, of the expanded frequency subband.
[0009] The multiple frequency subband boundaries in the frequency subband reduction array may constitute fewer frequency subbands than the multiple frequency subbands of the one or more acoustic signals, the coding rate adjusted multiple frequency subbands may be provided by the frequency subband reduction array, the frequency subband reduction array may be selected from the multiple frequency subband reduction arrays, the selection may be based on the coding rate associated with the one or more acoustic signals, each of the multiple frequency subband reduction arrays may include a different number of frequency subbands, and each of the multiple frequency subband reduction arrays may be associated with a different coding rate associated with the one or more acoustic signals.
[0010] The number of frequency subbands to be removed may be selected from a plurality of numbers of frequency subbands to be removed, and the selection may be based on a bandwidth value, and each of the plurality of numbers of frequency subbands to be removed may be associated with a different bandwidth value.
[0011] The sampling frequency adjusted frequency subbands may be in the form of an array including a plurality of frequency subband boundary values in ascending order of frequency subbands.
[0012] The apparatus may include a first encoder and a second encoder for encoding one or more acoustic signals at a coding rate, where the coding rate may include a sum of an encoding rate for the first encoder and an encoding rate for the second encoder, and where the first encoder may encode an acoustic transport signal associated with the one or more acoustic signals, and the second encoder may encode a plurality of spatial acoustic parameter sets associated with frequency subbands of the one or more acoustic signals.
[0013] According to a second aspect, there is provided an apparatus for spatial audio coding of one or more audio signals, the apparatus comprising means for determining a spatial audio parameter set for each of a plurality of frequency subbands of the one or more audio signals, receiving a coding rate associated with the one or more audio signals, mapping at least two consecutive subbands of the plurality of frequency subbands to an expanded frequency subband based on the coding rate to provide a code-rate adjusted plurality of frequency subbands, integrating the spatial audio parameter set associated with a first of the at least two consecutive frequency subbands with the spatial audio parameter set associated with a second of the at least two consecutive frequency subbands to provide an integrated spatial audio parameter set for the expanded frequency subband, determining an energy level for each frequency bin of the one or more audio signals, and determining a highest frequency bin having an energy level greater than a predetermined energy level and allocating a cutoff frequency subband as the frequency subband incorporating the highest frequency bin. determining a cutoff frequency subband for one or more acoustic signals; comparing the cutoff frequency subband for the one or more acoustic signals with a bandwidth value for the one or more acoustic signals; removing some frequency subbands, starting from a highest frequency subband among the plurality of code rate adjusted frequency subbands, to obtain a plurality of bandwidth adjusted frequency subbands, provided that the cutoff frequency subband is smaller than the bandwidth value for the one or more acoustic signals, wherein the number of frequency subbands removed is based on the cutoff frequency subband, and removing a spatial acoustic parameter set corresponding to each removed frequency subband; reducing a highest frequency subband among the plurality of bandwidth adjusted frequency subbands to be below the cutoff frequency subband value, provided that the highest frequency subband among the plurality of bandwidth adjusted frequency subbands extends beyond the cutoff frequency subband, and removing the spatial acoustic parameter sets associated with the plurality of bandwidth adjusted frequency subbands that extend beyond the cutoff frequency subband;An apparatus for spatial audio coding, comprising means adapted to:
[0014] The means configured to encode the exponent of the cut-off frequency sub-band may further be configured to encode each spatial acoustic parameter set associated with a frequency sub-band below the cut-off frequency sub-band.
[0015] The means configured to encode each spatial acoustic parameter set associated with a frequency subband below a cut-off frequency subband may further include means configured to determine an energy ratio parameter for each of a plurality of frequency subbands of the one or more acoustic signals, quantize the per-frequency-subband energy ratios for a plurality of frequency subbands equal to or greater than the cut-off frequency band to a minimum quantization level, quantize the per-frequency-subband energy ratios for a plurality of frequency subbands below the cut-off frequency band, as well as encode an indication that the number of encoded spatial acoustic parameter sets is less than the number of frequency subbands of the one or more acoustic signals, and encode the number of unencoded spatial acoustic parameter sets using a Golomb-Rice code.
[0016] The highest frequency subband of the plurality of bandwidth-adjusted frequency subbands may include upper and lower subband boundary values that encompass more than one of the plurality of frequency subbands of the one or more acoustic signals, and the means configured to reduce the highest frequency subband of the plurality of bandwidth-adjusted frequency subbands to be below a cutoff frequency subband value and to remove spatial acoustic parameter sets associated with the plurality of bandwidth-adjusted frequency subbands that extend beyond the cutoff frequency subband may include means configured to adjust the upper subband boundary value to be within the cutoff frequency subband value, and to remove spatial acoustic parameter sets associated with more than one of the plurality of frequency subbands of the one or more acoustic signals that are above the adjusted upper subband boundary value.
[0017] An apparatus comprising means configured to map at least two consecutive subbands of a plurality of frequency subbands to an expanded frequency subband to provide a plurality of code rate adjusted frequency subbands based on a code rate may comprise means configured to map upper and lower frequency band boundary values for the at least two consecutive frequency subbands of the plurality of frequency subbands to lower and upper frequency band boundary values of the expanded frequency subband.
[0018] The lower frequency subband boundary value and upper frequency subband boundary value of the expanded frequency subband are given by the lower frequency band boundary value and upper frequency band boundary value of a frequency subband reduction array that includes a plurality of frequency subband boundaries in ascending order of the frequency subbands, and the subband boundary value and the next upper subband boundary value in ascending order of the frequency subband reduction array are the lower frequency subband boundary and upper frequency subband boundary, respectively, of the expanded frequency subband.
[0019] The multiple frequency subband boundaries in the frequency subband reduction array may constitute fewer frequency subbands than the multiple frequency subbands of the one or more acoustic signals, the coding rate adjusted multiple frequency subbands may be provided by the frequency subband reduction array, the frequency subband reduction array may be selected from the multiple frequency subband reduction arrays, the selection may be based on the coding rate associated with the one or more acoustic signals, each of the multiple frequency subband reduction arrays may include a different number of frequency subbands, and each of the multiple frequency subband reduction arrays may be associated with a different coding rate associated with the one or more acoustic signals.
[0020] The sampling frequency adjusted frequency subbands may be in the form of an array including a plurality of frequency subband boundary values in ascending order of frequency subbands.
[0021] The apparatus may include a first encoder and a second encoder for encoding one or more acoustic signals at a coding rate, where the coding rate may include a sum of an encoding rate for the first encoder and an encoding rate for the second encoder, and where the first encoder may encode an acoustic transport signal associated with the one or more acoustic signals, and the second encoder may encode a plurality of spatial acoustic parameter sets associated with frequency subbands of the one or more acoustic signals.
[0022] According to a third aspect, there is provided a method for spatial audio coding of one or more audio signals, the method comprising: determining a spatial audio parameter set for each of a plurality of frequency subbands of the one or more audio signals; receiving a coding rate associated with the one or more audio signals; mapping at least two consecutive subbands of the plurality of frequency subbands to an expanded frequency subband based on the coding rate to provide a plurality of bandwidth-adjusted frequency subbands; receiving a bandwidth value associated with the one or more audio signals; removing a number of frequency subbands starting from a highest frequency subband of the plurality of bandwidth-adjusted frequency subbands to provide a plurality of bandwidth-adjusted frequency subbands, wherein the number of frequency subbands removed is based on the bandwidth value associated with the one or more audio signals; a method for reducing a highest frequency subband of a plurality of bandwidth-adjusted frequency subbands to reside below a bandwidth value, on condition that the highest frequency subband of the plurality of subbands extends beyond the bandwidth value associated with the one or more acoustic signals; aggregating a spatial acoustic parameter set associated with a first of at least two consecutive frequency subbands with a spatial acoustic parameter set associated with a second of the at least two consecutive frequency subbands to provide an aggregating spatial acoustic parameter set for the expanded frequency subband; removing a spatial acoustic parameter set corresponding to each removed frequency subband; and aggregating a spatial acoustic parameter set associated with a plurality of bandwidth-adjusted frequency subbands that extends beyond the bandwidth value, on condition that the highest frequency subband of the plurality of bandwidth-adjusted frequency subbands extends beyond the bandwidth value.
[0023] According to a fourth aspect, there is provided a method for spatial audio coding of one or more audio signals, the method comprising: determining a spatial audio parameter set for each of a plurality of frequency subbands of the one or more audio signals; receiving a coding rate associated with the one or more audio signals; mapping at least two consecutive subbands of the plurality of frequency subbands to an extended frequency subband based on the coding rate to provide a code-rate adjusted plurality of frequency subbands; integrating the spatial audio parameter set associated with a first of the at least two consecutive frequency subbands with the spatial audio parameter set associated with a second of the at least two consecutive frequency subbands to provide an integrated spatial audio parameter set for the extended frequency subband; determining an energy level for each frequency bin of the one or more audio signals; determining a highest frequency bin having an energy level greater than a predetermined energy level and allocating a cutoff frequency subband as the frequency subband incorporating the highest frequency bin. or determining a cutoff frequency subband for a plurality of acoustic signals; comparing the cutoff frequency subband for one or more acoustic signals with a bandwidth value for the one or more acoustic signals; removing some frequency subbands, starting from a highest frequency subband among the plurality of code rate adjusted frequency subbands, to obtain a plurality of bandwidth adjusted frequency subbands, provided that the cutoff frequency subband is smaller than the bandwidth value for the one or more acoustic signals, wherein the number of frequency subbands to be removed is based on the cutoff frequency subband; and removing a spatial acoustic parameter set corresponding to each removed frequency subband; and reducing the highest frequency subband among the plurality of bandwidth adjusted frequency subbands to be below the cutoff frequency subband value, provided that the highest frequency subband among the plurality of bandwidth adjusted frequency subbands extends beyond the cutoff frequency subband, and removing the spatial acoustic parameter set associated with the plurality of bandwidth adjusted frequency subbands that extend beyond the cutoff frequency subband;and encoding the exponents of the cutoff frequency sub-bands.
[0024] According to a fifth aspect, there is provided an apparatus for spatial audio coding, comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code causing the apparatus, using the at least one processor, to at least: determine a spatial audio parameter set for each of a plurality of frequency subbands of one or more audio signals; receive a coding rate associated with the one or more audio signals; map at least two consecutive subbands of the plurality of frequency subbands to an extended frequency subband based on the coding rate to provide a code-rate-adjusted plurality of frequency subbands; receive a bandwidth value associated with the one or more audio signals; and remove a number of frequency subbands, starting from a highest frequency subband of the code-rate-adjusted plurality of frequency subbands, to provide a bandwidth-adjusted plurality of frequency subbands, wherein the number of removed frequency subbands is equal to the number of bandwidths associated with the one or more audio signals. an apparatus for spatial audio coding configured to perform filtering based on a bandwidth value; reducing a highest frequency subband of the bandwidth-adjusted plurality of frequency subbands to be equal to or less than the bandwidth value, provided that the highest frequency subband of the bandwidth-adjusted plurality of frequency subbands extends beyond the bandwidth value associated with the one or more audio signals; integrating a spatial audio parameter set associated with a first of at least two consecutive frequency subbands with a spatial audio parameter set associated with a second of the at least two consecutive frequency subbands to obtain an integrated spatial audio parameter set for the expanded frequency subband; removing a spatial audio parameter set corresponding to each removed frequency subband; and removing a spatial audio parameter set associated with a plurality of bandwidth-adjusted frequency subbands that extends beyond the bandwidth value, provided that the highest frequency subband of the bandwidth-adjusted plurality of frequency subbands extends beyond the bandwidth value.
[0025] According to a sixth aspect, there is provided an apparatus for spatial audio coding, comprising at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code cause the apparatus, using the at least one processor, to at least: determine a spatial audio parameter set for each of a plurality of frequency subbands of one or more audio signals; receive a coding rate associated with the one or more audio signals; map at least two consecutive subbands of the plurality of frequency subbands to an expanded frequency subband based on the coding rate to provide a code-rate adjusted plurality of frequency subbands; integrate the spatial audio parameter set associated with a first of the at least two consecutive frequency subbands with the spatial audio parameter set associated with a second of the at least two consecutive frequency subbands to provide an integrated spatial audio parameter set for the expanded frequency subband; and determine an energy level for each frequency bin of the one or more audio signals. determining a cutoff frequency subband for the one or more acoustic signals by determining a highest frequency bin having an energy level greater than a predetermined energy level and allocating the cutoff frequency subband as a frequency subband incorporating the highest frequency bin; comparing the cutoff frequency subband for the one or more acoustic signals with a bandwidth value for the one or more acoustic signals; removing several frequency subbands, starting from a highest frequency subband among the code rate adjusted frequency subbands, to provide a plurality of bandwidth adjusted frequency subbands, on condition that the cutoff frequency subband is smaller than the bandwidth value for the one or more acoustic signals, wherein the number of frequency subbands removed is based on the cutoff frequency subband; and removing a spatial acoustic parameter set corresponding to each removed frequency subband; and comparing the highest frequency subband among the bandwidth adjusted frequency subbands with a condition that the highest frequency subband among the bandwidth adjusted frequency subbands extends beyond the cutoff frequency subband.An apparatus for spatial audio coding is provided, the apparatus being configured to: remove spatial audio parameter sets associated with a plurality of bandwidth-adjusted frequency subbands that are reduced to lie below a cutoff frequency subband value and extend beyond the cutoff frequency subband; and encode exponents of the cutoff frequency subbands.
[0026] A computer program product stored on the medium can cause an apparatus to perform the methods as described herein.
[0027] The electronic device may comprise an apparatus as described herein.
[0028] A chipset may comprise an apparatus as described herein.
[0029] SUMMARY OF THE INVENTION Embodiments of the present application aim to address problems associated with the state of the art.
[0030] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings in which: [Brief explanation of the drawings]
[0031] [Figure 1] FIG. 1 is a schematic diagram of a system of apparatus suitable for implementing some embodiments. [Figure 2] FIG. 1 is a schematic diagram illustrating an analysis processor according to some embodiments. [Figure 3] FIG. 1 is a schematic diagram illustrating a spatial analyzer according to some embodiments. [Figure 4] FIG. 2 illustrates a frequency band adjuster according to an embodiment. [Figure 5] 5 illustrates a flow diagram of the operation of a frequency band adjuster as shown in FIG. 4, according to some embodiments. [Figure 6] FIG. 10 illustrates a frequency band adjuster according to a further embodiment. [Figure 7]7 shows a flow diagram of the operation of the frequency band adjuster according to the further embodiment shown in FIG. 6; [Figure 8] 8 shows a flow diagram of the operation of a frequency band adjuster incorporating the embodiments of FIGS. 5 and 7. [Figure 9] FIG. 1 shows a schematic diagram of an exemplary device suitable for implementing the depicted apparatus. DETAILED DESCRIPTION OF THE INVENTION
[0032] The following describes in more detail suitable apparatus and possible mechanisms for providing effective spatial analysis derived metadata parameters (MASA parameters).
[0033] As mentioned above, Metadata-Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation that is suitable as an input format for IVAS.
[0034] It can be thought of as an audio representation consisting of "N channels + spatial metadata". It is a scene-based audio format particularly suited for spatial audio capture on practical devices such as smartphones. The idea is to describe the audio scene in terms of time- and frequency-varying sound source directions and, for example, energy ratios. Acoustic energy that is not defined (described) by direction is described as diffuse (coming from all directions).
[0035] As described above, spatial metadata associated with an acoustic signal may include multiple parameters per time-frequency tile (such as multiple directions and a direct-to-total ratio, diffuse coherence, distance, etc. associated with each direction). The spatial metadata may also include other parameters, or may be considered non-directional but associated with other parameters (such as surround coherence, diffuse-to-total energy ratio, remainder-to-total energy ratio, etc.) that, when combined with the directional parameters, may be used to define characteristics of the acoustic scene. For example, a reasonable design choice that can produce good quality output is determined to be that the spatial metadata includes one or more directions per time-frequency subframe (and a direct-to-total energy ratio, diffuse coherence, distance value, etc. associated with each direction).
[0036] As mentioned above, the parametric spatial metadata representation can use multiple simultaneous spatial directions. With MASA, the suggested maximum number of simultaneous directions is two. For each simultaneous direction, there may be associated parameters, such as a direction index, a direct-to-total energy ratio, a diffuse coherence, and a distance. In some embodiments, other parameters are also defined, such as a diffuse-to-total energy ratio, a surround coherence, and a residual-to-total energy ratio.
[0037] In the following description, the multi-channel system is described with respect to a multi-channel microphone implementation. However, as noted above, the input format can be any suitable input format, such as multi-channel loudspeaker, Ambisonic (FOA / HOA), etc. Furthermore, the output of the exemplary system is a multi-channel loudspeaker configuration. However, it will be understood that the output may be rendered to the user via means other than loudspeakers, such as binaural channel output. Furthermore, the multi-channel loudspeaker signal may be generalized to two or more reproduced sound signals.
[0038] Additionally, the IVAS codec, as an extension to EVS, can be used in store-and-forward applications where audio and speech content is encoded and stored in files for playback.
[0039] The MASA metadata may consist of at least the spherical direction (elevation, azimuth), at least one direct-to-global energy ratio of the resulting direction, the diffuse coherence, and the direction-independent surround coherence for each considered time-frequency (TF) block or tile, otherwise known as a time / frequency subband. In total, the MASA may have many different types of metadata parameters per time-frequency (TF) tile. The types of spatial acoustic parameters that can constitute the metadata for MASA are shown in Table 1 below.
[0040] This data may be encoded and transmitted (or stored) by an encoder to enable the spatial signal to be reconstructed at the decoder.
[0041] Furthermore, in some cases, metadata assisted spatial audio (MASA) may support up to two directions per TF tile, which would require the above parameters to be coded and transmitted per direction per TF tile, potentially doubling the bitrate required according to Table 1 below.
[0042] [Table 1]
[0043] The bitrate allocated for metadata in practical immersive audio communication codecs can vary widely. A typical overall operating bitrate of a codec may leave only 2-10 kbps for the transmission / storage of spatial metadata. However, some further implementations may allow up to 60 kbps or more for the transmission / storage of spatial metadata. The coding of directional parameters and energy ratio components has been previously examined along with the coding of coherence data. However, whatever the transmission / storage bitrate allocated for spatial metadata, it will always be necessary to represent these parameters using as few bits as possible, especially when TF tiles can support multiple directions corresponding to different sound sources within a spatial audio scene.
[0044] 1 illustrates an exemplary apparatus and system for implementing embodiments of the present application. System 100 is shown having an "analysis" portion 121 and a "synthesis" portion 131. "Analysis" portion 121 is the portion from receiving a multi-channel signal to encoding metadata and a transport signal, and "synthesis" portion 131 is the portion from decoding the encoded metadata and transport signal to presenting a reproduced signal (e.g., in the form of a multi-channel loudspeaker).
[0045] The input to the system 100 and the “analysis” portion 121 is the input acoustic signal 102 .
[0046] In the following example, the acoustic input signal 102 may be from a microphone array, however, it is understood that the acoustic input may be in any suitable acoustic input format, as detailed in the following description, where differences in processing arise when different input formats are employed.
[0047] The acoustic input signal 102 may be from any suitable source, for example: two or more microphones on a mobile phone, other microphone arrays such as B-format microphones, or Eigenmike. In some embodiments, as described above, the input may be any suitable acoustic signal input, such as an Ambisonic signal, for example, first-order Ambisonics (FOA), higher-order Ambisonics (HOA), or loudspeaker surround mix and / or objects, or any combination of the above.
[0048] In various embodiments, the microphone array acoustic input signal 102 may be provided to an analysis processor 105 configured to generate or determine suitable (spatial) metadata associated with the acoustic input signal 102. In addition, the (microphone array) acoustic input signal 102 may also be provided to a suitable transport signal generator 103 to generate an acoustic transport signal 104.
[0049] The analysis processor 105 is therefore configured to perform a spatial analysis on the acoustic input signal 102, resulting in suitable spatial acoustic (MASA) metadata in frequency bands 106. For all of the input types mentioned above, there are known methods for generating suitable spatial metadata in frequency bands, such as direction and direct-to-total energy ratio (or similar parameters such as diffuseness, i.e., ambient-to-total ratio). These methods will not be detailed herein, but some examples may include performing a suitable time-to-frequency transform for the input signal, and then, when the input is a cell phone microphone array, estimating a delay value between pairs of microphones that maximizes the correlation between the microphones in the frequency band, and constructing a corresponding direction value for that delay, and constructing a ratio parameter based on the correlation value.
[0050] In some embodiments, when the acoustic input is an FOA signal or a B-format microphone, the analysis processor 105 may be configured to determine parameters, such as intensity vectors, from which directional parameters are derived, and compare the intensity vector lengths with the full sound field energy estimates to determine ratio parameters, a method known in the literature as Directional Audio Coding (DirAC).
[0051] In some embodiments, when the acoustic input signal 102 is an HOA signal, the analysis processor 105 may take an FOA subset of the signal and use the methods described above, or may divide the HOA signal into multiple sectors, each of which utilizes the methods described above. This sector-based method is known in the literature as higher order directional acoustic coding (HO-DirAC). In this case, there is more than one simultaneous directional parameter per frequency band.
[0052] In some embodiments, when the acoustic input signal 102 is a loudspeaker surround mix and / or object, the analysis processor 105 may be configured to convert the signal into a FOA signal (through the use of spherical harmonics coding gain) and analyze the direction and ratio parameters as described above.
[0053] The output of the analysis processor 105 is therefore spatial acoustic (MASA) metadata 106 determined in frequency bands. The spatial acoustic (MASA) metadata 106 may include direction and energy ratios within frequency bands, but may also include any of the metadata types previously listed. The spatial acoustic (MASA) metadata 106 can vary over time and across frequency.
[0054] In some embodiments, the analysis processor functionality is implemented external to the system 100. For example, in some embodiments, the spatial acoustic (MASA) metadata 106 associated with the acoustic input signal 102 may be provided to the encoder 107 as a separate bitstream. In some embodiments, the spatial acoustic (MASA) metadata 106 may be provided as a set of spatial (directional) index values.
[0055] The above-described system 100 is further configured to implement a transport signal generator 103 to generate a suitable acoustic transport signal 104. The transport signal generator 103 is configured to receive an acoustic input signal 102, which may be, for example, a microphone array acoustic signal, and generate the acoustic transport signal 104. The acoustic transport signal 104 may be a multi-channel, stereo, binaural, or mono acoustic signal. The generation of the acoustic transport signal 104 may be implemented using any suitable method, such as those summarized below.
[0056] When the acoustic input signal 102 is a microphone array acoustic signal, the functionality of the transport signal generator 103 may select a left and right microphone pair and apply suitable processing to the signal pair, such as automatic gain control, microphone noise cancellation, wind noise cancellation, and equalization.
[0057] When the input is a FOA / HOA signal or a B-format microphone, the acoustic transport signal 104 can be a directional beam signal pointing in the left and right directions, such as two opposing cardioid signals.
[0058] When the input is a loudspeaker surround mix and / or objects, the audio transport signal 104 can be a downmix signal that combines the left channel into a left downmix channel, does the same for the right, and adds the center channel with a suitable gain to both transport channels.
[0059] In some embodiments, the acoustic transport signal 104 is an acoustic input signal 102, e.g., a microphone array acoustic signal. For example, in some situations, analysis and synthesis are performed in the same device in a single processing step, without intermediate encoding. The number of acoustic transport channels can also be any suitable number (rather than one or two channels as described in the examples).
[0060] The transport signal generator 103 and analysis processor 105 may, in some embodiments, be a computer (executing suitable software stored in memory and on at least one processor) or, alternatively, a specific device utilizing, for example, an FPGA or an ASIC.
[0061] The transport signal 104 and spatial audio (MASA) metadata 106 may be passed to an encoder 107 .
[0062] The encoder 107 may include an audio encoder core 109 configured to receive the audio transport (e.g., downmix) signal 104 and generate a suitable encoding of these audio signals. The encoder 107, in some embodiments, may be a computer (executing suitable software stored in memory and on at least one processor) or, alternatively, a specific device utilizing, for example, an FPGA or ASIC. The encoding may be performed using any suitable scheme. The encoder 107 may further include a metadata encoder / quantizer 111 configured to receive the spatial audio (MASA) metadata 106 and output an encoded or compressed form of the information. In some embodiments, the encoder 107 may further interleave, multiplex into a single data stream, or embed the metadata within the encoded downmix signal before transmission or storage, as indicated by the dashed lines in FIG. 1 . The multiplexing may be performed using any suitable scheme.
[0063] On the decoder side, the received or acquired data (stream) may be received by a decoder / demultiplexer 133. The decoder / demultiplexer 133 may demultiplex the encoded stream and pass the audio encoded stream to a transport extractor 135 configured to decode the audio signal and obtain a transport signal. Similarly, the decoder / demultiplexer 133 may include a metadata extractor 137 configured to receive the encoded metadata and decode the metadata. The decoder / demultiplexer 133 may, in some embodiments, be a computer (executing suitable software stored on a memory and on at least one processor) or, alternatively, a specific device utilizing, for example, an FPGA or an ASIC.
[0064] The decoded metadata and transport audio signal may be passed to a synthesis processor 139 .
[0065] The "synthesis" part 131 of the system 100 further illustrates a synthesis processor 139 configured to receive the encoded acoustic transport signal and the encoded spatial audio (MASA) metadata and to reproduce, based on the encoded acoustic transport signal and the encoded spatial audio (MASA) metadata, synthesized spatial audio in any suitable format in the form of a multi-channel spatial audio signal 110 (which may be in multi-channel loudspeaker format or in some embodiments any suitable output format such as a binaural or Ambisonics signal depending on the use case).
[0066] So, in summary, first the system (analysis part) is arranged to receive a multi-channel acoustic signal.
[0067] The system (analysis part) is then configured to generate a suitable transport audio signal (for example by selecting or downmixing some of the audio signal channels) and spatial audio parameters as metadata.
[0068] The system is then configured to encode the audio transport signal and spatial audio (MASA) metadata for storage / transmission.
[0069] After this, the system may store / transmit the encoded audio transport signal and the encoded spatial audio (MASA) metadata.
[0070] The system may acquire / receive an encoded audio transport signal and encoded spatial audio (MASA) metadata.
[0071] The system is then configured to extract the acoustic transport signal and the spatial acoustic (MASA) metadata from the encoded acoustic transport signal and the encoded spatial acoustic (MASA) metadata parameters, for example by demultiplexing and decoding the encoded acoustic transport signal and the encoded spatial acoustic (MASA) metadata parameters.
[0072] The system (synthesis part) is configured to synthesize an output multi-channel spatial audio signal based on the extracted acoustic transport audio signal and spatial audio (MASA) metadata.
[0073] FIG. 2 is an exemplary analysis processor 105 and metadata encoder / quantizer 111 (as shown in FIG. 1) according to some embodiments, which will be described in further detail.
[0074] 1 and 2 depict the metadata encoder / quantizer 111 and the analysis processor 105 as being coupled to one another. However, it should be understood that some embodiments may not tightly couple these two respective processing entities, such that the analysis processor 105 may reside on a different device than the metadata encoder / quantizer 111. Thus, a device including the metadata encoder / quantizer 111 may be provided with the acoustic transport signal 104 and metadata stream for processing and encoding independently of the capture and analysis processes.
[0075] The analysis processor 105 includes, in some embodiments, a time-frequency domain transformer 201 .
[0076] In some embodiments, the time-frequency domain transformer 201 is configured to receive the acoustic input signal 102 and apply a suitable time-to-frequency domain transform, such as a Short Time Fourier Transform (STFT), to transform the acoustic input time domain signal into suitable time-frequency acoustic signals 202. These time-frequency acoustic signals 202 may be passed to a spatial analyzer 203.
[0077] Thus, for example, the time-frequency audio signal 202 may be represented in a time-frequency domain representation by: s i (b,n), where b is the frequency bin index, n is the time-frequency block (frame) index, and i is the channel index. In another expression, n can be considered as a time index with a lower sampling rate than that of the original time-domain signal. These frequency bins can be grouped into subbands, which group one or more of the bins into subbands with band index k=0,...,K-1. Each subband k is assigned the lowest bin b k,low and the highest bin b k,high and the sub-bands are b k,low ~b k,high The widths of the subbands can approximate any suitable distribution, for example, the equivalent rectangular bandwidth (ERB) measure or the Bark measure.
[0078] Hence, a time-frequency (TF) tile (or block) is a particular sub-band within a sub-frame of a frame.
[0079] It can be appreciated that the number of bits required to represent spatial audio parameters may depend at least in part on the TF (time-frequency) tile resolution (i.e., the number of TF subframes or tiles). For example, a 20 ms audio frame may be divided into four 5 ms time-domain subframes, each having up to 24 frequency subbands divided in the frequency domain according to the Bark scale, an approximation thereof, or any other suitable division. In this particular example, the audio frame may be divided into 96 TF subframes / tiles, or in other words, four time-domain subframes each having 24 frequency subbands. Therefore, the number of bits required to represent spatial audio parameters for an audio frame may depend on the TF tile resolution. For example, if each TF tile were coded according to the allocation in Table 1 above, each TF tile would require 64 bits (for one audio source direction per TF tile) and 104 bits (for two audio source directions per TF tile, taking into account parameters unrelated to the audio source direction).
[0080] In various embodiments, the analysis processor 105 may include a spatial analyzer 203. The spatial analyzer 203 may be configured to receive time-frequency acoustic signals 202 and estimate a set of spatial acoustic parameters for each TF tile based on these signals, which are collectively shown in FIG. 1 as spatial acoustic (MASA) metadata 106. The spatial acoustic (MASA) metadata 106 may include directional parameters, which may be determined based on any acoustic-based "directional" determination.
[0081] For example, in some embodiments, the spatial analyzer 203 is configured to estimate the direction of a sound source using two or more signal inputs.
[0082] Therefore, the spatial analyzer 203 may be configured to provide at least one azimuth angle and elevation angle (spatial acoustic direction parameters) for each frequency band and temporal time-frequency block within a frame of the acoustic signal, denoted as azimuth angle φ(k,n) and elevation angle θ(k,n). The spatial acoustic direction parameters for a temporal subframe may be passed to the spatial parameter set encoder 207.
[0083] The spatial analyzer 203 may also be configured to determine an energy ratio parameter. The energy ratio may be considered to be a determination of the energy of an acoustic signal that can be considered to come from a certain direction. For example, the direct-to-total energy ratio r(k,n) may be estimated using a stability measure of the directivity estimation, or any correlation measure, or any other suitable method for obtaining a ratio parameter, such as that described in patent publication EP 3542546. Each direct-to-total energy ratio corresponds to a specific spatial direction and describes how much of the energy compared to the total energy comes from that specific spatial direction. This value may also be expressed separately for each time-frequency tile. The spatial direction parameter and the direct-to-total energy ratio describe how much of the total energy per time-frequency tile comes from a specific direction. Generally, the spatial direction parameter may also be considered as a direction of arrival (DOA).
[0084] The direct-to-total energy ratio parameter for a multi-channel captured microphone array signal can be estimated based on the normalized cross-correlation parameter cor'(k,n) between a pair of microphones in band k, where the value of the cross-correlation parameter is between -1 and 1. The direct-to-total energy ratio parameter r(k,n) is calculated by multiplying the normalized cross-correlation parameter by the diffuse field normalized cross-correlation parameter cor' D By comparing with (k,n),
[0085]
number
[0086] The spatial analyzer 203 may be further configured to determine a number of coherence parameters 112, which may include ambient coherence (γ(k,n)) and diffuse coherence (ζ(k,n)), both analyzed in the time-frequency domain.
[0087] The term audio source may refer to the dominant direction of a propagating sound wave, which may include the actual direction of the sound source.
[0088] Thus, for each subband k, there will be a collection (or set) of spatial acoustic parameters associated with the subband k and subframe n. In this case, each subband k and subframe n (in other words, a TF tile) may have the following spatial acoustic parameters associated with it for each sound source direction: at least one azimuth and elevation angle, denoted as azimuth angle φ(k,n) and elevation angle θ(k,n), as well as diffuse coherence (ζ(k,n)), and a direct-to-global energy ratio parameter r(k,n). If there is more than one direction per TF tile, then the TF tile may have each of the above-listed parameters associated with each sound source direction. In addition, the collection of spatial acoustic parameters may also include ambient coherence (γ(k,n)). The parameters may also include the diffuse-to-global energy ratio r(k,n). diff It may also include (k,n).
[0089] In embodiments, the diffuse-to-total energy ratio r diffwhere (k,n) is the energy ratio of non-directional sound across the surrounding directions, and typically there is a single diffuse-to-global energy ratio (and ambient coherence (γ(k,n))) per TF tile. The diffuse-to-global energy ratio can be thought of as the energy ratio remaining when the direct-to-global energy ratio (per direction) is subtracted from 1. Hereafter, the above parameters can be referred to as the set of spatial acoustic parameters for a particular TF tile (spatial acoustic parameter set). The collection of spatial acoustic parameter sets associated with a TF tile is known as the spatial acoustic (MASA) metadata signal 106.
[0090] The spatial parameter data set is then passed to a metadata encoder / quantizer 111 for encoding and quantization. In Figure 2, this is illustrated by a spatial parameter set encoder 207, which may be configured to receive the spatial parameter data set (shown as spatial audio MASA metadata stream 106) and quantize and encode the associated spatial parameter set, each TF tile.
[0091] The acoustic input signal 102 may be processed in the frequency domain by both the transport signal generator 103 and the analysis processor 105 at the same frequency subband resolution. However, some of the resulting frequency subbands in the acoustic transport signal 104 (from processing in the transport signal generator 103) may contain negligible (or even zero) levels of signal energy, the effect of which is that the signals associated with these subbands contribute negligibly, or at best, only a small amount, to the overall synthesized multi-channel (spatial) acoustic signal 110. This would indicate that the acoustic signals in the "low-energy" frequency subbands can be ignored and do not need to be coded (by the acoustic core encoder 109) for subsequent transmission and storage.
[0092] However, in the current system context, the analysis processor 105 generates a spatial acoustic parameter set for each subband of a subframe of the processed acoustic input signal 102. Hence, a mismatch may arise between the number of subbands in which the acoustic transport signal 104 contains so-called active acoustic signals and the number of subbands in which the acoustic input signal 102 is analyzed for spatial acoustics (MASA) metadata signal 106. In other words, the analysis processor 105 can generate a spatial parameter data set for each subband of a subframe of the acoustic input signal 102, regardless of whether the corresponding frequency subband of the acoustic transport stream / signal 104 contains an active acoustic signal or not.
[0093] Therefore, spatial acoustic parameter sets corresponding to frequency subbands (per subframe) of the audio transport signal / stream 10 having inactive audio signals may be considered to be coded unnecessarily, and therefore coding of these spatial acoustic parameter sets may result in unnecessary bit consumption.
[0094] It should be noted that the term active acoustic signal, as applied above, refers to a situation in which the acoustic signals of the subbands of a subframe of the acoustic transport signal 104 have a sufficiently high level of energy that the acoustic signals are considered to contribute to the synthesized multi-channel spatial audio signal 110. Conversely, the term inactive acoustic signal may refer to a situation in which the subbands of a subframe of the acoustic transport signal 104 have such a low acoustic signal energy level that the subbands are considered not to contribute significantly to the synthesized multi-channel spatial audio signal 110.
[0095] Therefore, embodiments start from the consideration that the number of spatial acoustic parameter sets in the spatial acoustic (MASA) metadata stream 106 can be reduced if the contribution of energy in frequency sub-bands of the acoustic transport signal 104 to the output multi-channel spatial audio signal 110 is negligible.
[0096] Additionally, as mentioned above, the IVAS codec can operate at a wide range of different encoding rates and different bandwidths, which can result in a mismatch between the number of subbands into which the acoustic input signal 102 is processed for the audio transport stream 104 and the number of subbands into which the audio input signal 102 is analyzed for the spatial (MASA) metadata stream 106. Some of the reasons for the mismatch between the number of subbands (in which the audio transport stream 104 and the spatial (MASA) metadata stream 106 are processed) can be due, at least in part, to the encoding rate assigned for encoding the audio transport stream 104 and the separate encoding rate assigned for the spatial (MASA) metadata stream 106. For example, the audio transport stream 104 and the spatial (MASA) metadata stream 106 can each be encoded according to any of a number of different encoding rates. The encoding rate assigned per stream can consequently affect the number of subbands into which the audio transport 104 and spatial (MASA) metadata 106 streams are created. For example, the code rate allocated for encoding the audio transport stream 104 may result in fewer subbands being generated than the number of subbands from which the spatial audio parameters of the spatial (MASA) metadata stream 106 are generated.
[0097] Therefore, the frequency band of the spatial (MASA) metadata stream 106 may extend beyond the frequency band of the audio transport stream 104. This may result in unnecessary encoding of spatial audio parameters associated with sub-bands of the spatial (MASA) metadata stream 106 that extend beyond the sub-bands of the audio transport stream 104, which in turn results in unnecessary consumption of coding bits during encoding of the spatial (MASA) metadata stream 106.
[0098] In this regard, Figure 3 shows the spatial analyzer in more detail, where the time-frequency acoustic signal 202 is received by a spatial parameter set determiner 301. The spatial parameter set determiner 301 may be configured to determine a spatial parameter set for each sub-band of the time-frequency acoustic signal 202. The components of each parameter set may be at least some of the spatial acoustic parameters as described above and listed in Table 1.
[0099] It should be noted that in some other embodiments, the spatial parameter set determiner 301 may be implemented within the spatial analyzer 203 within the analysis processor 105, the frequency sub-band adjuster 303 and the parameter set combiner / reducer 305 may form part of the metadata encoder / quantizer 111, and the analysis processor 105 may reside on a different device metadata than the encoder / quantizer 111.
[0100] 3 is a frequency sub-band adjuster 303. The frequency sub-band adjuster 303 may be configured to receive input configuration information, such as a (selected) overall (IVAS) coding rate 206 and a (selected) audio signal bandwidth 208. In addition, the frequency sub-band adjuster 303 may also be configured to receive the audio transport signal 104.
[0101] The frequency subband adjuster 303 may then create a further configuration of subbands depending on the received input configuration information, the overall (IVAS) coding rate 206, and the audio signal bandwidth 208. This further configuration of subbands is based on the original configuration of subbands of the time-frequency audio signal 202, but may involve some changes in the distribution and width of some of the frequency subbands, and therefore a change in the number of subbands across the bandwidth of the signal. For example, the further configuration of subbands may include fewer and wider subbands compared to the pattern of subbands for the time-frequency audio signal 202.
[0102] In various embodiments, the input to the frequency subband adjuster 303 now includes the acoustic transport signal 104. The frequency subband configuration of the original time-frequency acoustic signal 102 may be reduced depending on the energy of each corresponding subband of the acoustic transport signal 104. In other words, the frequency subband configuration as produced by the frequency subband adjuster 303 may be made fewer by removing frequency subbands from the original pattern of subbands of the time-frequency acoustic signal 202. Thereby, the resulting frequency subband configuration may include fewer subbands of the original width depending on the energy levels of the frequency subbands of the acoustic transport signal 104.
[0103] The output from the frequency subband adjuster 303 is shown in FIG. 3 as an adjusted subband configuration array 302. This parameter may reflect the modification of frequency subband boundaries (or removal of frequency subbands) of the original time-frequency audio signal 202 in the form of an array of subband boundary values. In other words, the adjusted subband configuration array 302 may represent the pattern of subband boundaries after the encoder operating conditions of the selected encoding rate (total IVAS encoding rate 206) and the selected bandwidth (audio signal bandwidth 208) are taken into account. Note that the total (IVAS) encoding rate 206 merely serves as an example of how the encoding rate may be parameterized. This does not exclude any other parameters that may indicate the encoding rate for an encoder. For example, the encoding rate parameter (such as input 206) may be set to the encoding rate of the audio encoder 109 or according to the encoding rate associated with the metadata encoder and quantizer 111.
[0104] The adjusted subband configuration parameters 302 may then be passed to a parameter set combiner / reducer 305 .
[0105] In addition to the adjusted subband configuration parameters 302 , the parameter set integrator / reducer 305 also receives a spatial acoustic parameter set for each frequency subband 304 of the time-frequency acoustic signal 202 .
[0106] When the input to the spatial analyzer includes the overall (IVAS) coding rate 206 and the audio signal sampling frequency 208, the parameter set integrator / reducer 305 may be configured to perform an integrating operation between some of the spatial parameter sets 304. The integrating operation may be performed according to the subband configuration of the subband configuration parameter / array 302. Essentially, some of the spatial parameter sets (for the time-frequency audio signal 202) may be integrated with adjacent spatial parameter sets such that the resulting distribution of spatial parameter sets reflects the distribution of subbands as indicated by the adjusted subband configuration parameter 302. A description of the integrating process may be found in patent application publication WO2021 / 130404, which teaches that spatial audio parameter sets across adjacent subbands may be integrated to provide a smaller number of spatial audio parameter sets across a smaller number of integrated frequency bands.
[0107] When the input to the spatial analyzer 203 includes an acoustic transport signal 104, the parameter set combiner / reducer 305 may be configured to reduce the number of spatial acoustic parameter sets from the signal 304 as indicated by the sub-band cutoff signal 306. In this case, the adjusted sub-band cutoff signal 306 may contain information indicating the spatial parameter sets that should be removed from the spatial acoustic parameter set signal 304.
[0108] The output from the parameter set combiner / reducer 305 (i.e., the spatial acoustic metadata 106) may then include the spatial acoustic parameter sets of the signal 304 combined into a smaller number of spatial acoustic parameter sets and / or the spatial acoustic parameter sets of the signal 304 reduced to a smaller number of spatial acoustic parameter sets.
[0109] An embodiment for generating the spatial acoustic metadata 106 in response to the adjusted subband configuration array 302 is described below: the output from the parameter set combiner / reducer 305 is the spatial acoustic metadata 106 comprising a smaller number of combined spatial acoustic parameter sets of the spatial acoustic parameter set signal 304 in response to the audio signal bandwidth (parameters) 208 and the overall (IVAS / coding system) coding rate 206.
[0110] To that end, the spatial audio (MASA) metadata 106 may be encoded (by the encoder 207) at various encoding rates ranging from 2.5 kbps to 65 kbps. The particular rate chosen may be tied to an overall (IVAS or system) encoding rate 206, which for IVAS may be one of the following: / * IVAS_13k2,IVAS_16k4,IVAS_24k4,IVAS_32k,IVAS_48k,IVAS_64k,IVAS_80k,IVAS_96k,IVAS_128k,IVAS_160k,IVAS_192k,IVAS_256k,IVAS_384k,IVAS_512k.* /
[0111] Here, for example, IVAS_13k2 means an IVAS encoding rate of 13.2 kbps. The overall (IVAS) coding rate 206 may be used by the frequency sub-band adjuster 303, in part, to determine the sub-band boundaries for the adjusted sub-band configuration parameters 302.
[0112] In this regard, FIG. 4 illustrates in further detail the frequency sub-band adjuster 303 where an adjusted sub-band configuration array 302 is generated in response to a combination of inputs including the overall (IVAS) coding rate 206 and the audio signal bandwidth (parameter) 208.
[0113] The overall system code rate (IVAS code rate) 206 is shown as being received by code rate sub-band adjuster 401. The output from code rate sub-band adjuster 401 is shown as rate adjusted sub-band array 402.
[0114] In various embodiments, the code rate sub-band adjuster 401 may be configured to perform a mapping function between the overall code rate 206 and a particular distribution of frequency sub-bands relative to the distribution of sub-bands in the time-frequency audio signal 202. In other words, the result of the mapping function is the code rate adjusted sub-band array 402.
[0115] The mapping may be performed so that the distribution of the code rate-adjusted subbands more closely aligns with the width and number of frequency subbands of the transmitted acoustic signal 104. The time-frequency acoustic signal 202 may include 24 frequency subbands across its bandwidth. The mapping functionality in 401 may then be configured to take the overall (IVAS) code rate 206 and map the code rate to a distribution of frequency subbands that differs from the distribution of frequency subbands of the time-frequency acoustic signal 202. The code rate-adjusted subband array 402 may have fewer subbands, some of which are wider than their counterparts in the time-frequency acoustic signal 202. Thus, the resulting code rate-adjusted subbands still span the equivalent bandwidth (of 24 bands) of the original time-frequency acoustic signal 202, but have fewer subbands.
[0116] In various embodiments, the mapping function may be performed by first mapping the received overall (IVAS) code rate 206 to a parameter indicating the number of subbands in the code rate adjusted subband array. There may be a one-to-one mapping between each overall (IVAS) code rate 206 and the parameter indicating the reduced number of subbands. An example of a one-to-one mapping for IVAS is shown by Table 2 below.
[0117] [Table 2]
[0118] For example, a total IVAS encoding rate of 160 kbps would result in a reduction in the number of subbands in the code rate adjusted subband array 402 from 24 to 12.
[0119] Note that each subband number in Table 2 above refers to a contiguous series of subbands starting from the lowest subband, and that the rate-adjusted subband array 402 spans the entire bandwidth occupied by the 24 subbands of the time-frequency acoustic signal 202.
[0120] Each parameter indicating the reduced number of subbands in the table above corresponds to an IVAS coding rate in ascending order of bit rate. Taking another example, the parameter indicating the reduced number of subbands for an IVAS encoding rate of 32 kbps is 5 subbands. In this example, the code rate adjusted subband array would include elements defining the subband boundaries of the 5 subbands.
[0121] For clarity, any reduction in the number of subbands for the time-frequency acoustic signal 202 that may be performed may be performed with respect to the maximum number of subbands, given as 24 in the example above. Therefore, any adjustment made to the number of subbands is performed on the basis that the full bandwidth of the signal is preserved. Essentially, the width of some of the frequency subbands is expanded to occupy a wider range of frequency bins while preserving the full bandwidth associated with the time-frequency acoustic signal 202 (for IVAS, this is 24 subbands, or 60 frequency bins with each frequency bin having a width of 400 Hz).
[0122] Once the parameters indicating the reduced number of sub-bands are found from Table 2 above, the widths of some of the remaining sub-bands can be modified to preserve the full bandwidth of the signal as explained above. The reallocation of the number of frequency bins for some of the reduced number of sub-bands can be found by using the following mapping array:
[0123] The distribution of frequency bins for the 24 sub-bands of the time-frequency acoustic signal 202 may be given by the following array MASA_band_grouping_24: In other words, this is the distribution of frequency bins per sub-band of the 24 sub-band groups, where the maximum number of frequency bins is 60 and each bin has a width of 400 Hz. int16 MASA_band_grouping_24[24+1]= { 0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,25,30,40,60 };
[0124] Each element of MASA_band_grouping is an indication of the frequency bin index of the lower / upper boundary of a subband. The frequency bin indexes are collectively grouped in ascending order within the group array described above. For example, within the above MASA band group, the 24th subband is allocated a frequency bin range of 40 to 60, the 23rd subband is allocated a frequency bin range of 30 to 40, and the 1st subband is allocated a frequency bin range of 0 to 1. Note that the allocation of frequency bins to subbands typically does not include the last value of the frequency bin range; therefore, in practice, the frequency bins allocated to the 24th subband would be 40 to 59, and similarly, the frequency bin range allocated to the 23rd subband would be 30 to 39.
[0125] The reallocation of frequency subbands for each reduced number of subbands in Table 2, and for each value of the parameter indicating the reduction, can be given by the MASA_band_mapping array below.
[0126] For example, the resulting rate adjusted subband array 402 for a reduction from 24sb to 18sb when the overall (IVAS) code rate is 192kbps may be given by the following array: int16 MASA_band_mapping_24_to_18[18+1]= { 0,1,2,3,4,5,6,7,8,9,10,12,14,17,20,21,22,23,24 };
[0127] In this example, the 18th subband of the rate adjusted subband array is allocated to the frequency bin spanning the 23rd to 24th subbands with respect to the original subbands of the time-frequency acoustic signal 202, where the frequency bins allocated per subband are given by the MASA_band_grouping array described above. In other words, the 18th subband occupies the range of frequency bins from 40 to 60. The 17th subband of the rate adjusted subband array 402 is allocated the range of frequency bins spanning the 22nd and 23rd subbands from the MASA_band_grouping array described above, i.e., the range of frequency bins from 30 to 40, and so on.
[0128] Similarly, in this example, due to an IVAS coding rate of 160 kbps, the total number of reduced subbands is 12sb. The code rate adjusted subband array 402 is given by the following array: int16 MASA_band_mapping_24_to_12[12+1]= { 0,1,2,3,4,5,7,9,12,15,20,22,24 };
[0129] In a similar manner, the code rate adjusted subband array 402 for the total number of subbands reduced from 24sb to 8sb and from 24sb to 5sb may be given by the following arrays, respectively: int16 MASA_band_mapping_24_to_8[8+1]= { 0,1,2,3,5,8,12,20,24 }; int16 MASA_band_mapping_24_to_5[5+1]= { 0,1,3,7,15,24 };
[0130] Note that the rate-adjusted subband array can be any of the arrays MASA_band_mapping_24_to_18 through MASA_band_mapping_24_to_5. Thus, the rate-adjusted subband array encompasses a "pattern" of subbands (for the subbands in MASA_band_grouping_24) according to the overall (IVAS) system-determined code rate 206.
[0131] 4, it can be seen that the output from the code rate frequency sub-band adjuster 401, the coding adjusted sub-band array 402, can be passed to the coding bandwidth sub-band adjuster 403. The coding bandwidth sub-band adjuster 403 is configured to receive the audio signal bandwidth 208, which can be used to reduce the sampling frequency / bandwidth associated with the code rate adjusted sub-band array 402 relative to the time-frequency audio signal 202. In various embodiments, this process typically involves removing higher sub-bands of the code rate adjusted sub-band array 402 so that the complete bandwidth of the audio signal associated with the code rate adjusted sub-band array 402 is reduced to match the bandwidth indicated by the audio signal bandwidth input 208.
[0132] The output from the sampling frequency sub-band adjuster may be referred to as a bandwidth adjusted sub-band array 404 .
[0133] The reduction in bandwidth of the code rate adjusted subband array 402 (due to the audio signal bandwidth 208) may be accomplished using a table that gives the reduction in the number of subbands from the (full-band) code rate adjusted subband array for each possible input audio signal bandwidth 208.
[0134] Note that encoder 121 has the capability to operate at one of several different pre-specified bandwidths as indicated by audio signal bandwidth signal line 208. For example, an IVAS encoder can be configured to operate at any one of the audio signal bandwidths specified in Table 3.
[0135] In this regard, Bandwidth Adjustment Table 3 below illustrates the relationship between input audio signal bandwidth 208 and code rate adjusted subband array 402. Various allowable audio signal sampling frequencies / bandwidths are listed along the columns of Table 3, and code rate adjusted subband arrays are listed along the rows of the mapping Table 3. For each value of audio signal bandwidth 208, Table 3 provides the number of subbands that need to be removed from code rate adjusted subband array 402 to achieve the bandwidth associated with the specified audio signal bandwidth 208. This mapping is provided for each combination of audio signal bandwidth 208 and code rate adjusted subband total array 402. The values specified by Table 3 relate to the number of subbands to be removed, starting with the highest subband in code rate adjusted subband array 402.
[0136] [Table 3]
[0137] The operating mechanism of bandwidth adjustment Table 3 can be further understood by taking the example above in which the number of subbands was adjusted from 24 sb to 12 sb for an IVAS code rate of 160 kbps. In other words, the code rate adjusted subband array includes 12 frequency subbands for the 160 kbps code rate. In this case, the code rate adjusted subband array with 12 sb constitutes one entry into the table, and the other entry is the bandwidth as specified by the audio signal bandwidth 208. Thus, for the exemplary input audio signal bandwidth 208 in wide band (WB) mode, the table results in an adjustment factor of 2 sb. That is, the two highest frequency subbands are removed from the code rate adjusted subband array, resulting in a bandwidth adjusted subband array 404 of 10 sb for the combination of the 160 kbps IVAS code rate and the selected bandwidth of WB. Overall, therefore, a total (IVAS) coding rate (206) of 160 kbps with an audio signal bandwidth (208) of WB will result in a bandwidth-adjusted subband array 404 having first 10 subbands spanning the 8 kHz bandwidth (16 kHz sampling frequency) of the wideband audio signal.
[0138] In addition to the adjustments provided by the overall code rate 206 and the audio signal bandwidth 208, the width of the final frequency band for some combination of code rate adjusted subbands 402 and audio signal bandwidth 208 may also be considered for further adjustment. Further adjustments may be applied in those cases of bandwidth adjusted subbands (as indicated by the bandwidth adjusted subband array 404) where the highest remaining frequency subband is found to be wider than the bandwidth associated with the audio signal bandwidth 208.
[0139] This final adjustment process is shown in FIG. 4 as being performed by a highest sub-band limiter 405, which receives the bandwidth-adjusted sub-band array 404 along with the acoustic signal bandwidth 208 and produces an adjusted sub-band configuration array 302 as an output.
[0140] In this regard, Table 3 above discloses bandwidths in terms of the number of frequency bins for each possible value of the acoustic signal bandwidth 208. Similarly, within the same column in Table 3, bandwidths are shown in terms of the number of subbands out of the 24 subbands of the original time-frequency acoustic signal 202. This column can then be used to determine whether the final subbands of the bandwidth-adjusted subband array 404 are wider than the actual bandwidth allowed by the acoustic signal bandwidth 208.
[0141] For example, from the table above, a narrow band (NB) signal may have a maximum signal bandwidth of 10 frequency bins, while a full band (FB) signal may have a maximum signal bandwidth of 60 frequency bins. As explained above, in some situations, the highest subband of the bandwidth-adjusted subband array 404 may extend beyond the actual bandwidth of the acoustic signal bandwidth (parameter) 208. This situation may be particularly prevalent for narrow band (NB) signals, which have an actual bandwidth of only 10 frequency bins. For example, when examining MASA_band_grouping_24_to_12, which is the adjustment made to the number of subbands for a full coding rate of 160 kbps, the highest subband is assigned the frequency bin corresponding to subbands 22-24 (frequency bins 40-60) of the original time-frequency acoustic signal 202. Next, if this code rate is further adjusted for a narrowband signal (NB), it can be seen from Table 3 that the four highest subbands are removed, leaving the following subbands: {0, 1, 2, 3, 4, 5, 7, 9, 12}. The final subband occupies frequency subbands 9 through 12 relative to the subbands in the MASA_band_grouping_24 array. Thus, the final subband in this example extends beyond the bandwidth of the NB signal at the 10th frequency bin. Naturally, in these situations, it would be advantageous to perform a further adjustment in which the final subband is clipped to fit within the actual bandwidth of the audio signal bandwidth 208. In this regard, Table 4 below lists the respective subband boundaries for each combination of audio signal bandwidth 208 and code rate adjusted subband array 402. It can be seen that some of the entries in Table 4 had the frequency bin of the highest subband clipped to fit within the bandwidth of the audio signal bandwidth (parameter) 208. These entries are marked with an asterisk * for clarity.
[0142] In embodiments, once a "pattern" of subband boundaries (shown in FIG. 4 as bandwidth-adjusted subband array 404) is determined as a function of the overall (IVAS) code rate 206 (given by Table 2) and the bandwidth of the audio signal bandwidth input 208 (given by Table 3), the subband boundaries of bandwidth-adjusted subband array 404 may be further checked against Table 4 to determine whether the highest subband should be capped (or limited) to match the actual bandwidth of audio signal bandwidth 208.
[0143] [Table 4]
[0144] In Table 4, "Number of Subbands" is the initial number of subbands before adjustments are made for the overall (IVAS) coding rate 206 and the audio signal bandwidth 208. Note that the subband boundaries are given in terms of the total number of subbands of 24 subbands in the time-frequency audio signal 202; in other words, the subband boundaries are relative to the original MASA_band_grouping_24. For example, a full band signal reduced to five subbands would have a mapping according to MASA_band_mapping24_to_5, in which case it can be seen that the final subband occupies subbands 15 through 24 (of the original audio signal of 24 subbands), which is equivalent to the highest subband occupying frequency bins 15 through 60. Obviously, the bandwidth for the 32 kHz SWB signal is 16 kHz (or 23 subbands total for the 24 subband time-frequency acoustic signal 202), which is equal to the 30-40 (i.e., 12 kHz-16 kHz) frequency bin width from the MASA_band_grouping_24 array. Therefore, the mapping of 24 subbands to 5 subbands for the SWB signal is capped at 23 subbands total (which is equal to the 40 frequency bins total for the array MASA_band_grouping_24) to ensure that the signal does not extend beyond the bandwidth of the SWB signal (16 kHz).
[0145] The output from the highest sub-band limiter 405 is the adjusted sub-band configuration array 302. In the case where the highest sub-band of the bandwidth sub-band array 404 falls within the bandwidth of the acoustic bandwidth 208, the adjusted sub-band configuration array 302 will be the bandwidth adjusted sub-band array 404. In other words, there is no limiting / capping action applied to the highest sub-band. However, in the case where the highest sub-band of the bandwidth sub-band array 404 is wider than the bandwidth of the acoustic signal bandwidth 208, the adjusted sub-band configuration array 302 will be the bandwidth adjusted sub-band array 404 where the highest sub-band is limited in terms of its width.
[0146] The adjusted subband configuration array 302 may then be passed to a parameter set combiner / reducer 305 as shown in FIG.
[0147] 5 illustrates a computer software or hardware implementable process of the frequency subband adjuster 303 for determining the adjusted subband configuration array 302. The adjusted subband configuration array 302 is shown as being determined from the overall (IVAS) coding rate 206 and the audio signal bandwidth 208. For clarity, in various embodiments, the adjusted subband configuration array 302 (or vector) may include element values that specify subband boundaries for the parameter set combiner / reducer 305. In effect, the adjusted subband configuration array 302 may be one of the subband boundary arrays from Table 4 above. The parameter set combiner / reducer 305 then combines adjacent sets of spatial acoustic parameters from adjacent subbands using the adjusted subband configuration array 302. The parameter set combiner / reducer 305 may also be configured to remove spatial acoustic parameter sets that correspond to frequency subbands that are larger than those of the adjusted subband configuration array 302. The result of the integration and reduction process is a set of spatial acoustic parameters for the subbands that reflect the pattern of the subbands as given by the adjusted subband configuration array 302. In some embodiments, the adjusted subband configuration array 302 may be configured as an index or pointer to one of the subband boundary arrays of Table 4.
[0148] 5, the process of determining the adjusted subband configuration array 302 by the frequency subband adjuster 303 is shown as receiving an input 206 containing an indication of the overall code rate (for the IVAS encoder). Processing step 501 shows a mapping step between the received overall (IVAS) code rate 206 and the number of frequency subbands allowed in the code rate adjusted subband array 402. This can be accomplished by using Table 2.
[0149] 5 illustrates the selection of the MASA_band_mapping array as determined by the number of frequency subbands from step 501. Note that higher code rates from Table 2 do not require a reduction in the number of subbands. The selected MASA_band_mapping array forms the code rate adjusted subband array 402.
[0150] 5 shows a step of removing some high frequency sub-bands from the code rate sub-band array 402 according to the audio signal bandwidth 208. This step can be implemented, for example, by using Table 3.
[0151] Processing step 507 shows the process of checking Table 4 to determine whether the highest subband of the bandwidth adjusted subband array 404 is wider than the bandwidth of the acoustic signal sampling frequency 208. If the highest subband is wider than the bandwidth, then the width of the highest subband is adjusted to be within the bandwidth. This step can be accomplished by using Table 4. The output from this step can be one of the arrays from Table 4 that specify the subband boundaries of the adjusted subband configuration array 302.
[0152] The processing steps according to Figure 5 have the advantage that no extra signaling bits are required to be sent from the encoder to the decoder, since both the encoder and decoder have access to the table above, and the decoder can be made aware of both the code rate and bandwidth at the encoder through system level configuration information.
[0153] It should be appreciated that Figure 3, in conjunction with Figure 4, illustrates that the spatial parameter sets associated with the original pattern of sub-bands of the time-frequency acoustic signal 202 are combined and reduced as a final stage according to the sub-band pattern provided by the adjusted sub-band configuration array 302. In other words, the combination of spatial acoustic parameter sets as indicated by the code rate adjusted sub-band array 402, the reduction of spatial parameter sets as indicated by the bandwidth adjusted sub-band array 404, and the conditional trimming of the highest sub-band as indicated by 405 may be performed as a single processing stage in the parameter set combiner / reducer 305 according to the "final" adjusted sub-band configuration array 302.
[0154] However, it should also be understood that in other embodiments, the spatial parameter set merging and reduction process may be performed sequentially once the patterns for each subband are determined. Thus, in these embodiments, subband parameter set merging may be performed when the code rate-adjusted subband array 402 is determined. This step may then be followed by spatial parameter set reduction when the bandwidth-adjusted subband array 404 is determined. Finally, the spatial parameter set of the highest frequency subband may then be conditionally trimmed by the highest subband limiter 405.
[0155] FIG. 6 shows a frequency sub-band adjuster 303 for an embodiment that develops a sub-band cutoff signal 306 derived from the energy level of the sub-band of the acoustic transport signal 104 .
[0156] 6, the sub-band adjuster 303 is shown receiving the acoustic transport signal 104 via a frequency bin energy determiner 601. The frequency bin energy determiner 601 is configured to measure / determine the energy of the acoustic signal in each frequency bin of the acoustic transport signal 104, in other words, the frequency bin energy 605. Bearing in mind that the acoustic transport signal 104 may include up to two transmitted signals, the frequency bin energy determiner 601 is configured to determine the energy in each frequency bin for all transport signals. The energy calculation may be performed for each acoustic frame.
[0157] The output of the frequency bin energy determiner 601, the frequency bin energy (per transport signal) 605, is then passed to the frequency sub-band reducer 603 for further processing.
[0158] The frequency sub-band reducer 603 may be configured to determine whether any of the frequency bin energies are below a predetermined energy. This may be accomplished by scanning the energy of each frequency bin in descending order of frequency bin index in the frequency bin energy signal 605 and checking the first instance when the energy of the frequency bin is above a minimum energy level. m Once the energy cutoff frequency bin index b is determined, e A b m +1. The frequency sub-band reducer 603 then determines the frequency bin index b e There is a frequency subband k e This is determined to be a cutoff frequency sub-band above which the acoustic transport signal 104 is considered to make little contribution to the final multi-channel spatial audio signal 110. In other words, k e Any sub-bands mentioned above with indices above or equal to this are considered to have insufficient energy levels, and therefore the spatial parameter sets associated with these sub-bands may be effectively removed by not being coded.
[0159] For cases where the acoustic transport signal 104 has more than one channel, the above process may be performed using multiple frequency bin indices (b e1 ,b e2 ....) can be found (one for each channel). The highest frequency bin index is selected and the frequency sub-band associated with the highest frequency bin index is the cutoff frequency sub-band index k for all channels of the acoustic transport signal 104. e can be determined as:
[0160] Cutoff frequency subband index k e may be communicated as signal 306 to the parameter set combiner / reducer 305. e >B w (k) k e w (k)
[0161] When the parameter set integrator / reducer 305 receives the signal 306, it combines all the frequency sub-bands k e may be configured to remove all spatial parameter sets associated with frequency sub-band k e All parameter sets associated with ∼K-1 (where K-1 is the highest subband index associated with the acoustic transport signal 104) are set to 0 (or removed) and therefore will not form part of the spatial metadata signal 106 passed to the metadata encoder / quantizer 111.
[0162] The remaining spatial parameter sets of the spatial audio metadata 106 may then be coded by the spatial parameter set coder 207 according to the technique described in patent application EP 3818525. Furthermore, the spatial parameter set coder 207 also uses a Golomb-Rice code of order 0 to code the number of sub-bands (k) that do not contain any coded spatial parameter sets. e ∼K-1 number of sub-bands).
[0163] Note that for cases where all frequency bins have energy levels above a predetermined energy level, then no spatial parameter sets are removed from the spatial metadata signal 106. This case can be signaled using a single bit.
[0164] Thus, in this embodiment, the encoded spatial metadata information may include encoded spatial parameter sets and an additional signaling bit, where one state of the signaling bit indicates that the encoded spatial metadata 106 includes encoded spatial parameter sets for all frequency bands, and another state of the signaling bit indicates that only a partial number of frequency band spatial parameter sets have been encoded in the spatial metadata 106.
[0165] In another embodiment, the spatial parameter set encoder 207 may be configured to eliminate a single bit indicating the absence of a spatial parameter set removed from the spatial metadata signal 106. Alternatively, if the number of sub-band spatial parameter sets is less than the number of complete sub-bands, and the number of sub-band k spatial parameter sets is less than the number of complete sub-bands, the spatial parameter set encoder 207 may be configured to eliminate a single bit indicating the absence of a spatial parameter set removed from the spatial metadata signal 106. e Only in the specific case when the direct-to-global energy ratio of the spatial acoustic parameter set associated with ∼K-1 (the remaining sub-bands) is quantized to the minimum quantization level, a single bit is added to the coded stream.
[0166] It should be noted that the quantization and encoding of the energy ratio value may be performed separately from the quantization and encoding of the other spatial acoustic parameters of the sub-band spatial acoustic parameter set. Thus, each sub-band may have at least a quantized energy ratio associated with it, whereas the other parameters of the spatial acoustic parameter set associated with the sub-band may not be quantized and coded (and thus do not form part of the coded bitstream).
[0167] For example, the energy ratio value (for each sub-band) can be quantized using a 3-bit scalar quantizer, and the quantization and encoding of other spatial acoustic parameters of the spatial acoustic parameter set for the sub-band can be quantized and encoded according to the published EP3818525.
[0168] Regarding other embodiments, FIG. 7 shows a further process of quantizing the sub-band spatial parameter set when the sub-bands of the acoustic transfer signal 104 are considered to have energy low enough not to contribute to the synthesized multi-channel spatial acoustic signal 110.
[0169] As seen in FIG. 7, the process starts by receiving the value of k associated with K-1 frequency sub-bands of the sub-frame. e and starts by receiving the value of k for the K-1 frequency sub-bands of the sub-frame.
[0170] First, the cut-off sub-band value of k is examined to determine whether k < K-1. As described above, this indicates that the spatial acoustic parameter set for frequency sub-bands k to K-1 can be removed from the metadata encoding process performed by 111. This is shown by the processing step 701 in FIG. 7. e <K-1 to determine whether k < K-1, the cut-off sub-band value of k is examined. As described above, this indicates that the spatial acoustic parameter set for frequency sub-bands k to K-1 can be removed from the metadata encoding process performed by 111. This is shown by the processing step 701 in FIG. 7. e As described above, this indicates that the spatial acoustic parameter set for frequency sub-bands k to K-1 can be removed from the metadata encoding process performed by 111. This is shown by the processing step 701 in FIG. 7. e ~K-1 can be removed from the metadata encoding process performed by 111. This is shown by the processing step 701 in FIG. 7.
[0171] In step 701, if it is determined that k < K-1, then, according to FIG. 7, the processing path 702 is taken. e <K-1, then, according to FIG. 7, the processing path 702 is taken.
[0172] The process path 702 then sets the energy ratio associated with the frequency sub-band k < K-1 to have the minimum quantization level. This is shown as the processing step 703 in FIG. 7. e <K-1 to have the minimum quantization level. This is shown as the processing step 703 in FIG. 7.
[0173] Next, for frequency sub-bands 0 to k e eThe energy ratios associated with -1 are quantized according to those values. As described above, this can be accomplished using a scalar quantizer, which creates a quantization index (or codeword) for each energy ratio value. This is shown, in FIG. 7, as processing step 705.
[0174] Next, a single bit can be added to the bitstream for the subframe to indicate that the number of encoded spatial acoustic parameter sets encoded within the bitstream for the subframe is not the total number for subband K-1. This is shown, in FIG. 7, as processing step 707.
[0175] The frequency subbands having no associated spatial parameter sets, i.e., subbands k e <K-1, the number of which is encoded using a Golomb-Rice code of order zero. This is shown, in FIG. 7, as processing step 709. Of course, this encoded number of frequency subbands also forms part of the encoded bitstream for the frame.
[0176] Finally, the "other" spatial acoustic parameters of the spatial acoustic parameter sets for subbands 0 to k e -1 can be quantized and encoded according to WO 2022 / 129672, WO 2021 / 048468, WO 2020 / 070377, WO 2020 / 008105, and WO 2021 / 144498. As described above, these quantized spatial acoustic parameter sets can also form part of the encoded bitstream for the frame. This step is shown, in the figure, as processing step 711.
[0177] Note that, for clarity, the term "other" spatial acoustic parameters in this context refers to the spatial acoustic parameters of the spatial acoustic parameter sets (for the subbands) that do not include the energy ratios described above.
[0178] Returning to decision step 701 in Figure 7, it can be checked whether the result of this step determines that all frequency subbands should have their respective spatial parameter sets coded. e =K-1, effectively determining that all sub-bands are above a minimum energy level. It should be understood that those skilled in the art will recognize that other means may be used to signal this condition. The process may then be configured to take processing path 704.
[0179] Once the decision is made to take processing path 704, parameter set combiner / reducer 305 may be configured to quantize and encode the energy rations corresponding to all frequency subbands 0 to K−1, which is shown as processing step 713 in FIG.
[0180] The process then determines whether the energy ratio associated with the last subband (K-1) has been quantized to the minimum quantization level. This determination step is shown as processing step 715 in Figure 7. When the result of determination step 715 indicates that the energy ratio associated with the last subband (K-1) has not been quantized to the minimum quantization level, the process is configured to proceed to processing step 717, where the spatial parameter sets associated with all subbands 0 to K-1 are quantized and coded.
[0181] However, when the decision step 715 indicates that the energy ratio associated with the last sub-band (K-1) has been quantized to the minimum quantization level, the process is configured to proceed to processing step 719. In processing step 719, a single bit is added to the encoded stream (for the frame). The state of the bit (shown as set to 0 in FIG. 7) indicates that while the energy ratio associated with the final sub-band K-1 was encoded to the minimum quantization level, all K-1 spatial parameter sets were encoded.
[0182] Finally, after step 719, FIG. 7 shows that the process moves to processing step 717. Prior to this, the spatial parameter sets associated with all sub-bands 0 to K-1 are quantized and encoded.
[0183] For clarity, the bitstream for processing route 702 includes, at least per frame, the encoded and quantized energy ratios associated with frequency bands 0 to k e -1, the sub-band k quantized and encoded to the minimum quantization level, the energy ratios associated with sub-bands k e ~K-1, a bit to notify that the number of encoded spatial parameter sets is <K-1, the 0th order GR code indicating the number of sub-bands given by the values of k e ~K-1, and the quantized and encoded spatial parameter sets (each including other spatial parameters for the energy ratio) associated with sub-bands 0 to k e -1.
[0184] Following the above, the bitstream for processing route 704 may include, for each frame, at least, coded, quantized energy ratios associated with frequency bands 0 to K-1 and quantized, coded spatial parameter sets (each including other spatial parameters for the energy ratios) associated with sub-bands 0 to K-1. In addition, the bitstream for processing route 704 may also include a bit for indicating that the number of coded spatial parameter sets corresponds to sub-bands 0 to K-1 for the situation when the energy ratio associated with the last frequency band K-1 is quantized to the minimum level.
[0185] The above-described embodiment may be performed on a frame-by-frame basis. Furthermore, the second embodiment may be deployed in conjunction with the first embodiment on a frame-by-frame basis. For example, a decision on whether to use the first or second embodiment may be made at the start of a new frame.
[0186] It should be appreciated that in some further embodiments, the energy-based embodiments described above may be performed in conjunction with previous embodiments employing processing steps according to Figure 5. In other words, the energy-based embodiments described above may be integrated into embodiments in which spatial parameter sets associated with sub-bands of the time-frequency audio signal 202 are aggregated and reduced according to the overall code rate 206 and audio signal bandwidth 208.
[0187] In this regard, FIG. 8 illustrates how the energy-based embodiment described above may be implemented within a system that deploys the previous embodiment of FIG.
[0188] Processing steps 801 and 803 may be configured as in FIG. 5, where a selected overall (IVAS) coding rate 206 is received and, based on this, the coding rate adjusted subband array 402 may be determined by determining the MASA band mapping array.
[0189] Processing step 803 and 804 determine the specified acoustic signal bandwidth 208, bandwidth B w This can be given, for example, by Table 3, in which various allowable sampling frequencies are listed as a function of the number of sub-bands k.
[0190] Processing step 805 then determines the audio signal bandwidth B W 208 is the cutoff frequency subband index k e Compare with 306.
[0191] FIG. 8 then shows the cutoff frequency subband index k e is the bandwidth B W If the audio signal bandwidth 208 B is found to be greater than (or equal to), the process proceeds to step 807 where processing steps similar to those of step 505 are performed. In other words, processing step 807 W 8. Accordingly, processing step 807 is shown to receive the rate-adjusted subband array 402 from processing step 803, and also the audio signal bandwidth 208. The result of this processing step is therefore the bandwidth-adjusted subband array 404.
[0192] Alternatively, the comparison step 803 may be performed using an energy-based cutoff frequency sub-band index k e is the bandwidth B W When this condition is met, the process may be configured to move to step 809. In step 809, the process may determine that the exponent k e In other words, processing step 809 is configured to remove frequency subbands above a cutoff frequency subband having an index equal to the cutoff index k eThe result of this processing step can therefore be seen as a version of the bandwidth-adjusted subband array 404, with the upper subbands limited according to the cutoff index 306. Thus, processing step 809 takes the rate-adjusted subband array 402 from processing step 803 and the cutoff frequency subband index k e 306, which is shown to allow the above-described variants of the bandwidth-adjusted sub-band array 404 to be formed.
[0193] 8 shows how the output from step 807 (bandwidth-adjusted subband array 404) is passed to processing step 811. Step 811 performs the same processing function as step 505 in FIG. 5. In other words, step 811 performs a process to determine whether the highest subband of bandwidth-adjusted subband array 404 is wider than acoustic signal bandwidth 208, and if the highest subband is found to be wider than acoustic signal bandwidth 208, then the width of the highest subband is adjusted to be within this bandwidth. The output from step 811 is adjusted subband configuration array 302.
[0194] Process step 813 is shown accepting the bandwidth-adjusted subband array 404 from step 809. In a manner similar to that of step 811, process step 813 determines whether the highest subband of the bandwidth-adjusted subband array 404 is further widened by an index k e 306. If it is determined that this is indeed the case, then the width of the highest sub-band is determined by the cutoff frequency sub-band index k e 306. The output of step 813 is also a further variant of the adjusted subband configuration array 302.
[0195] Furthermore, FIG. 8 also illustrates that when processing path 805 is followed, the cutoff frequency subband index ke 306 indicates that the cutoff frequency subband index k e The encoding of 306 is shown as process step 815, which may be performed according to the process steps of FIG.
[0196] 9, an exemplary electronic device that may be used as an analysis or synthesis device is shown. The device may be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, sound reproduction device, etc.
[0197] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 can be configured to execute various program code, such as methods, such as those described herein.
[0198] In some embodiments, device 1400 comprises memory 1411. In some embodiments, at least one processor 1407 is coupled to memory 1411. Memory 1411 may be any suitable storage means. In some embodiments, memory 1411 includes program code sections for storing program code executable on processor 1407. Additionally, in some embodiments, memory 1411 may further include a storage data section for storing data, e.g., data that has been processed or is to be processed according to embodiments as described herein. The implementing program code stored in the program code sections and the data stored in the storage data section may be retrieved by processor 1407 via the memory-processor coupling whenever needed.
[0199] In some embodiments, device 1400 comprises a user interface 1405. User interface 1405, in some embodiments, may be coupled to processor 1407. In some embodiments, processor 1407 may control the operation of user interface 1405 and receive input from user interface 1405. In some embodiments, user interface 1405 may allow a user to input commands into device 1400, for example, via a keypad. In some embodiments, user interface 1405 may allow a user to obtain information from device 1400. For example, user interface 1405 may include a display configured to display information from device 1400 to a user. User interface 1405, in some embodiments, may include a touchscreen or touch interface capable of both allowing information to be input into device 1400 and further displaying information to a user of device 1400. In some embodiments, user interface 1405 may be a user interface for communicating with a position determiner as described herein.
[0200] In some embodiments, device 1400 comprises an input / output port 1409. Input / output port 1409, in some embodiments, includes a transceiver. The transceiver in such embodiments may be coupled to processor 1407 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver, or any suitable transceiver or transmitter and / or receiver means, in some embodiments, may be configured to communicate with other electronic devices or apparatuses via wiring or a wired coupling.
[0201] The transceiver may communicate with the further device by any suitable known communication protocol, for example, in some embodiments the transceiver may use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an infrared data communication pathway (IRDA).
[0202] The transceiver input / output port 1409 may be configured to receive the signal and, in some embodiments, determine parameters as described herein by using a processor 1407 executing suitable code. Additionally, the device may generate a suitable downmix signal and parameter output to be transmitted to a combining device.
[0203] In some embodiments, device 1400 may be employed as at least part of a synthesis device. Thus, input / output port 1409 may be configured to receive the downmix signal determined in a capture device or processing device as described herein, and in some embodiments, parameters, and generate a suitable audio signal format output by using processor 1407 to execute suitable code. Input / output port 1409 may be coupled to any suitable audio output, for example, to a multi-channel speaker system and / or headphones or the like.
[0204] In general, various embodiments of the present invention may be implemented in the form of hardware or special purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in the form of hardware, while other aspects may be implemented in the form of firmware or software that may be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. While various aspects of the present invention may be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it is appreciated that these blocks, apparatus, systems, techniques, or methods described herein may be implemented in the form of, by way of non-limiting example, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller, or other computing device, or any combination thereof.
[0205] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device, for example, in a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any blocks of logic flow as depicted in the figures may represent program steps, or interconnected logic circuits, blocks, and functions, or a combination of program steps and logic circuits, blocks, and functions. Software may be stored on physical media such as memory chips, or memory blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs, and the like.
[0206] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memories, etc. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting example, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.
[0207] Embodiments of the present invention may be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available for converting logic-level designs into semiconductor circuit designs ready to be etched and formed on semiconductor substrates.
[0208] The program can use well-established design rules and a library of pre-stored design modules to route the conductors and place the components on the semiconductor chip. Once the design for a semiconductor circuit is complete, the resulting design, in a standardized electronic format, can be transmitted to a semiconductor manufacturing facility or "fab" for fabrication.
[0209] The foregoing description has provided a complete and informative description of exemplary embodiments of the present invention, by way of illustrative and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the art in light of the above description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of the present invention as defined in the appended claims.
Claims
1. 1. An apparatus for spatial audio coding of one or more audio signals, said apparatus comprising means: determining a spatial acoustic parameter set for each of a plurality of frequency subbands of the one or more acoustic signals; receiving a coding rate associated with the one or more acoustic signals; mapping at least two consecutive subbands of the plurality of frequency subbands to an expanded frequency subband based on the coding rate to provide a plurality of coding rate adjusted frequency subbands; receiving a bandwidth value associated with the one or more acoustic signals; removing a number of frequency subbands, starting with a highest frequency subband from the rate-adjusted plurality of frequency subbands, to provide a plurality of bandwidth-adjusted frequency subbands, the number of frequency subbands removed being based on the bandwidth value associated with the one or more acoustic signals; reducing a highest frequency subband of the bandwidth-adjusted plurality of frequency subbands to reside at or below a bandwidth value associated with the one or more acoustic signals, provided that the highest frequency subband of the bandwidth-adjusted plurality of frequency subbands extends beyond the bandwidth value associated with the one or more acoustic signals; aggregating a spatial acoustic parameter set associated with a first of the at least two consecutive frequency subbands with a spatial acoustic parameter set associated with a second of the at least two consecutive frequency subbands to provide an aggregated spatial acoustic parameter set for the expanded frequency subband; removing a spatial acoustic parameter set corresponding to each removed frequency subband; removing spatial acoustic parameter sets associated with the plurality of bandwidth-adjusted frequency subbands that extend beyond the bandwidth value, with the condition that the highest frequency subband of the plurality of bandwidth-adjusted frequency subbands extends beyond the bandwidth value; 1. An apparatus for spatial audio coding, comprising means configured to:
2. the highest frequency subband of the bandwidth-adjusted plurality of frequency subbands includes upper and lower subband boundary values that encompass more than one of the plurality of frequency subbands of the one or more acoustic signals, and the means configured to reduce the highest frequency subband of the bandwidth-adjusted plurality of frequency subbands to reside below the bandwidth values; 2. The apparatus for spatial audio coding of claim 1, further comprising: means configured to adjust the upper sub-band boundary value to lie within the bandwidth value; and means configured to remove spatial audio parameter sets associated with the bandwidth-adjusted plurality of frequency sub-bands that extend beyond the bandwidth value, said means configured to remove spatial audio parameter sets associated with the plurality of frequency sub-bands of the one or more audio signals that are above the adjusted upper sub-band boundary value.
3. the apparatus comprising means configured to map at least two consecutive subbands of the plurality of frequency subbands to an expanded frequency subband based on the coding rate to provide the plurality of code rate adjusted frequency subbands, 3. An apparatus for spatial audio coding according to claim 1, comprising means configured to map upper and lower frequency band boundary values for the at least two consecutive frequency subbands of the plurality of frequency subbands to lower and upper frequency band boundary values of the expanded frequency subband.
4. 4. The apparatus for spatial audio coding of claim 3, wherein the lower frequency subband boundary value and the upper frequency subband boundary value of the expanded frequency subband are given by lower frequency band boundary value and upper frequency band boundary value of a frequency subband reduction array including a plurality of frequency subband boundaries in ascending order of frequency subbands, and the subband boundary value and the next upper subband boundary value in ascending order of the frequency subband reduction array are the lower frequency subband boundary and the upper frequency subband boundary, respectively, of the expanded frequency subband.
5. 5. The apparatus for spatial audio coding of claim 4, wherein the plurality of frequency subband boundaries in the frequency subband reduction array constitute fewer frequency subbands than the plurality of frequency subbands of the one or more acoustic signals, the code rate adjusted plurality of frequency subbands are provided by the frequency subband reduction array, the frequency subband reduction array is selected from a plurality of frequency subband reduction arrays, the selection being based on the code rate associated with the one or more acoustic signals, each of the plurality of frequency subband reduction arrays comprising a different number of frequency subbands, and each of the plurality of frequency subband reduction arrays is associated with a different code rate associated with the one or more acoustic signals.
6. 6. An apparatus for spatial audio coding according to claim 1, wherein the number of frequency subbands to be removed is selected from a plurality of numbers of frequency subbands to be removed, the selection being based on the bandwidth value, each of the plurality of numbers of frequency subbands to be removed being associated with a different bandwidth value.
7. 7. An apparatus for spatial audio coding according to claim 1, wherein the plurality of sampling frequency adjusted frequency sub-bands is in the form of an array comprising a plurality of frequency sub-band boundary values in ascending order of frequency sub-bands.
8. 8. The apparatus for spatial audio coding of claims 1 to 7, comprising a first encoder and a second encoder for encoding the one or more audio signals at said coding rate, said coding rate comprising the sum of an encoding rate for the first encoder and an encoding rate for the second encoder, wherein the first encoder encodes an audio transport signal associated with the one or more audio signals, and wherein the second encoder encodes the plurality of spatial audio parameter sets associated with the frequency subbands of the one or more audio signals.
9. 1. An apparatus for spatial audio coding of one or more audio signals, said apparatus comprising means: determining a spatial acoustic parameter set for each of a plurality of frequency subbands of the one or more acoustic signals; receiving a coding rate associated with the one or more acoustic signals; mapping at least two consecutive subbands of the plurality of frequency subbands to an expanded frequency subband based on the coding rate to provide a plurality of coding rate adjusted frequency subbands; aggregating a spatial acoustic parameter set associated with a first of the at least two consecutive frequency subbands with a spatial acoustic parameter set associated with a second of the at least two consecutive frequency subbands to provide an aggregated spatial acoustic parameter set for the expanded frequency subband; determining an energy level for each frequency bin of the one or more acoustic signals; determining a cutoff frequency sub-band for the one or more acoustic signals by determining a highest frequency bin having an energy level greater than a predetermined energy level and allocating the cutoff frequency sub-band as a frequency sub-band incorporating the highest frequency bin; comparing the cutoff frequency subbands for the one or more acoustic signals to a bandwidth value for the one or more acoustic signals; removing a number of frequency subbands to provide a plurality of bandwidth-adjusted frequency subbands, starting from the highest frequency subband of the code rate adjusted frequency subbands, on the condition that the cutoff frequency subband is smaller than the bandwidth value for the one or more acoustic signals, wherein the number of removed frequency subbands is based on the cutoff frequency subband, and removing a spatial acoustic parameter set corresponding to each removed frequency subband; reducing a highest frequency subband of the bandwidth-adjusted plurality of frequency subbands to lie below the cutoff frequency subband value, provided that the highest frequency subband of the bandwidth-adjusted plurality of frequency subbands extends beyond the cutoff frequency subband, and removing spatial acoustic parameter sets associated with the bandwidth-adjusted plurality of frequency subbands that extend beyond the cutoff frequency subband; encoding the exponents of the cutoff frequency sub-bands; 1. An apparatus for spatial audio coding, comprising means configured to:
10. 10. The apparatus for spatial audio coding according to claim 9, wherein the means configured to encode the exponent of the cut-off frequency sub-band is further configured to encode each spatial audio parameter set associated with the frequency sub-bands below the cut-off frequency sub-band.
11. the means adapted to encode each spatial acoustic parameter set associated with the frequency sub-bands below the cut-off frequency sub-band, determining an energy ratio parameter for each of the plurality of frequency subbands of the one or more acoustic signals; quantizing the energy ratio for each of the plurality of frequency subbands equal to or greater than the cutoff frequency band to a minimum quantization level; quantizing the energy ratio for each frequency subband of the plurality of frequency subbands smaller than the cutoff frequency band; and encoding an indication that the number of encoded spatial acoustic parameter sets is less than the number of frequency subbands of the one or more acoustic signals; and encoding the number of uncoded spatial acoustic parameter sets using Golomb-Rice coding; The apparatus of claim 10 , further comprising means configured to:
12. the highest frequency subband of the bandwidth-adjusted plurality of frequency subbands includes upper and lower subband boundary values that encompass more than one of the plurality of frequency subbands of the one or more acoustic signals, the means configured to reduce the highest frequency subband of the bandwidth-adjusted plurality of frequency subbands to be below the cut-off frequency subband value and to remove spatial acoustic parameter sets associated with the bandwidth-adjusted plurality of frequency subbands that extend beyond the cut-off frequency subband, adjusting the upper sub-band boundary value to be within the cutoff frequency sub-band value; and removing spatial acoustic parameter sets associated with the plurality of frequency subbands of the one or more acoustic signals that are above the adjusted upper subband boundary value; An apparatus for spatial audio coding according to claims 9 to 11, comprising means adapted to:
13. the apparatus comprising means configured to map at least two consecutive subbands of the plurality of frequency subbands to an expanded frequency subband based on the coding rate to provide the plurality of code rate adjusted frequency subbands, 13. An apparatus for spatial audio coding according to claim 9, comprising means configured to map upper and lower frequency band boundary values for the at least two consecutive frequency subbands of the plurality of frequency subbands to lower and upper frequency band boundary values of the extended frequency subband.
14. 14. The apparatus for spatial audio coding of claim 13, wherein the upper frequency subband boundary value and the upper frequency subband boundary value of the expanded frequency subband are given by a lower frequency band boundary value and an upper frequency band boundary value of a frequency subband reduction array comprising a plurality of frequency subband boundaries in ascending order of frequency subbands, and the subband boundary value and the next upper subband boundary value in ascending order of the frequency subband reduction array are the lower frequency subband boundary and the upper frequency subband boundary, respectively, of the expanded frequency subband.
15. 15. The apparatus for spatial audio coding of claim 14, wherein the plurality of frequency subband boundaries in the frequency subband reduction array constitute fewer frequency subbands than the plurality of frequency subbands of the one or more acoustic signals, the code rate adjusted plurality of frequency subbands are provided by the frequency subband reduction array, the frequency subband reduction array is selected from a plurality of frequency subband reduction arrays, the selection being based on the code rate associated with the one or more acoustic signals, each of the plurality of frequency subband reduction arrays comprising a different number of frequency subbands, and each of the plurality of frequency subband reduction arrays is associated with a different code rate associated with the one or more acoustic signals.
16. 16. An apparatus for spatial audio coding according to claim 9, wherein the plurality of sampling frequency adjusted frequency sub-bands is in the form of an array comprising a plurality of frequency sub-band boundary values in ascending order of frequency sub-bands.
17. 17. An apparatus for spatial audio coding according to claim 9, comprising a first encoder and a second encoder for encoding the one or more audio signals at said coding rate, said coding rate comprising the sum of an encoding rate for the first encoder and an encoding rate for the second encoder, wherein the first encoder encodes an audio transport signal associated with the one or more audio signals, and wherein the second encoder encodes the plurality of spatial audio parameter sets associated with the frequency subbands of the one or more audio signals.
18. 1. A method for spatial audio coding of one or more audio signals, said method comprising: determining a spatial acoustic parameter set for each of a plurality of frequency subbands of the one or more acoustic signals; receiving a coding rate associated with the one or more acoustic signals; mapping at least two consecutive subbands of the plurality of frequency subbands to an expanded frequency subband based on the coding rate to provide a plurality of coding rate adjusted frequency subbands; receiving a bandwidth value associated with the one or more acoustic signals; removing a number of frequency subbands, starting with a highest frequency subband from the rate-adjusted plurality of frequency subbands, to provide a plurality of bandwidth-adjusted frequency subbands, wherein the number of frequency subbands removed is based on the bandwidth value associated with the one or more acoustic signals; reducing a highest frequency subband of the bandwidth-adjusted plurality of frequency subbands to reside at or below a bandwidth value associated with the one or more acoustic signals, provided that the highest frequency subband of the bandwidth-adjusted plurality of frequency subbands extends beyond the bandwidth value associated with the one or more acoustic signals; aggregating a spatial acoustic parameter set associated with a first of the at least two consecutive frequency subbands with a spatial acoustic parameter set associated with a second of the at least two consecutive frequency subbands to provide an aggregated spatial acoustic parameter set for the expanded frequency subband; removing a spatial acoustic parameter set corresponding to each removed frequency subband; removing spatial acoustic parameter sets associated with the plurality of bandwidth-adjusted frequency subbands that extend beyond the bandwidth value, with the condition that the highest frequency subband of the plurality of bandwidth-adjusted frequency subbands extends beyond the bandwidth value; A method comprising:
19. 1. A method for spatial audio coding of one or more audio signals, said method comprising: determining a spatial acoustic parameter set for each of a plurality of frequency subbands of the one or more acoustic signals; receiving a coding rate associated with the one or more acoustic signals; mapping at least two consecutive subbands of the plurality of frequency subbands to an expanded frequency subband based on the coding rate to provide a plurality of coding rate adjusted frequency subbands; aggregating a spatial acoustic parameter set associated with a first of the at least two consecutive frequency subbands with a spatial acoustic parameter set associated with a second of the at least two consecutive frequency subbands to provide an aggregated spatial acoustic parameter set for the expanded frequency subband; determining an energy level for each frequency bin of the one or more acoustic signals; determining a cutoff frequency sub-band for the one or more acoustic signals by determining a highest frequency bin having an energy level greater than a predetermined energy level and allocating the cutoff frequency sub-band as a frequency sub-band incorporating the highest frequency bin; comparing the cutoff frequency subbands for the one or more acoustic signals to a bandwidth value for the one or more acoustic signals; removing a number of frequency subbands to provide a plurality of bandwidth-adjusted frequency subbands, starting from the highest frequency subband of the code rate adjusted frequency subbands, on the condition that the cutoff frequency subband is smaller than the bandwidth value for the one or more acoustic signals, wherein the number of removed frequency subbands is based on the cutoff frequency subband, and removing a spatial acoustic parameter set corresponding to each removed frequency subband; reducing a highest frequency subband of the bandwidth-adjusted plurality of frequency subbands to lie below the cutoff frequency subband value, provided that the highest frequency subband of the bandwidth-adjusted plurality of frequency subbands extends beyond the cutoff frequency subband, and removing spatial acoustic parameter sets associated with the bandwidth-adjusted plurality of frequency subbands that extend beyond the cutoff frequency subband; encoding the exponents of the cutoff frequency sub-bands; A method comprising:
Citation Information
Patent Citations
A method for parametric multichannel encoding
JP2016509260A
Determining the importance of spatial audio parameters and related coding
JP2022528660A
The reduction of spatial audio parameters
WO2021250312A1