Audio Signal Representation Decoding Unit, Apparatus, Audio Signal Representation Encoding Unit, Audio Encoder, Methods, Non-Transient Storage Unit and Compressed Ambisonic Audio Signal Representation

BR112025017189A2Pending Publication Date: 2026-08-25
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
BR112025017189
Authority / Receiving Office
BR · BR
Patent Type
Applications
Publication Date
2026-08-25

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

1 / 74 “AUDIO SIGNAL REPRESENTATION DECODING UNIT, APPARATUS, AUDIO SIGNAL REPRESENTATION CODING UNIT, AUDIO ENCODER, METHODS, NON-TRANSIENT STORAGE UNIT AND COMPRESSED AMBISONIC AUDIO SIGNAL REPRESENTATION”

[001] This document refers to an audio signal representation decoding unit, an audio signal representation encoding unit, apparatus comprising them, and non-transient storage methods and units.

[002] This document details a new framework for higher-order directional audio coding (HO-DirAC) for higher-order Ambisonics (HOA) input-to-output transmission. This document is also directed towards a Sector-based DirAC system with combined first-order and higher-order DoA and diffusion estimation.

[003] According to conventional representations of standard ambisonic audio signals, there is a single direction of arrival (DoA) and a single diffusion throughout the space. However, it is understood that it is possible to have multiple DoAs and diffusion in multiple spatial sectors. Therefore, a different and more accurate sound field parameter model is proposed here.

[004] In contrast to the commonly employed first-order parameter estimators, use is made of the additional information available in the higher-order channels of the HOA. Specifically, the sound field can be characterized by more than one dominant direction of arrival (DoA), which allows resolving more than one sound source per critical band in the encoder.

[005] In the decoder, these additional DoAs control the synthesis of multiple directional HOA streams that can originate from multiple sound sources.

[006] Despite the additional information, the proposed technique can implement the structure of the current encoder, preserving the robustness of the current first-order ambisonic DirAC (FOA), allowing seamless switching between the two designs for Petition 870250071890, dated 08 / 14 / 2025, pages 222 / 295 2 / 74 different coding scenarios.

[007] In essence, the proposed technique improves upon previously known HO-DirAC methods by making use of local and global sectoral diffusion information. Specifically, the new technique is able to more accurately model realistic soundscapes, correctly and robustly reproducing the overall diffuse energy ratio of the input signal, resulting in improved perceptual quality during HOA audio coding and spatial enhancement compared to previous designs.

[008] This occurs because the global diffusion path resembles that of a first-order system, thus maintaining its relative robustness and stability. At the same time, greater spatial image accuracy can be obtained by measuring a DoA for each sector of an arbitrary number of sectors and reconstructing multiple directional components of the sound field.

[009] Furthermore, the integration of multiple DoAs into existing first-order DirAC systems is greatly simplified with the present invention.

[010] Directional audio coding parameterizes a spatial audio scene as perceptually relevant parameters. These parameters comprise, for each time-frequency mosaic, the direction of arrival (DoA)Ω of the incident sound field and a measure of sound field diffusionψ, indicating the relationship between the directional and diffuse sound field components. Both parameters are extracted from the active intensity vector, estimated from first-order Ambisonics (FOA) (see [US20100169103A1]). The active intensity vectori is conveniently derived from FOA, according to the well-known formula (as per [Pulkki2007]), i = p V .

[011] The direction of the sound provides an estimate of DoA, while the length compared to the acoustic energy provides a measure of diffusion.

[012] The decoder can restore certain higher-order signal components of transmitted FOA signals, as detailed in a Patent of Petition 870250071890, dated 08 / 14 / 2025, pp. 223 / 295 3 / 74 HO-DirAC Encoder [WO 2020 / 115311 A1], According to WO 2020 / 115311 A1, the input FOA signals are split into two rendering paths based on the estimated diffusion ψ; to perform directional (i.e., 1-ψ) and fuzzy (ψ) rendering. The directional components are assumed to be plane waves and therefore decoded as HOA signals in the direction of Ω, by a plane wave continuation of the omnidirectional pressure signal. The latter is extracted from the transmitted subset signal. Fuzzy components result in FOA signals, scaled by a function that depends on ψ.

[013] HOA signals allow segmenting the input sound field, for example, by multiple spatial weights, i.e., beamformer(s), such as those in Figure 3. The HOA input, therefore, allows formulating weighted FOA signals accordingly, as shown in Figure 4. Therefore, these segmented FOA channels, i.e., sound field sectors, allow the simultaneous estimation of multiple Ω and ψ in sound field sectors, (as in HO-DirAC Sector Processing [US 10,313,815 B2]).

[014] Sectoral parameters were proposed in [Politis2015], however, not for spatial audio encoding and compression, but for speaker-based rendering and spatial sharpness.

[015] A state-of-the-art device transmits a single DoA and diffusion (first-order estimates Ψ'Ω), or partially recovers these estimates in the decoder.

[016] The current conventional sound field model assumes a mix of a single directional source and a diffuse sound field, per time-frequency block. However, this conventional model is frequently violated in practice, for example, by multiple directional sources in the same time-frequency block or by specular reflections. A multi-DoA model, such as the proposed sector model, can resolve these scenarios for multiple directional sound sources, thus increasing the perceived audio quality.

[017] In addition, the sector model can stabilize parameter estimation. Petition 870250071890, dated 08 / 14 / 2025, pages 224 / 295 4 / 74 in situations with competing directions; sector weighting distorts the DoA estimator, leading to less directional fluctuation, stabilizing and improving performance. In general, this technique improves rendering situations that consist of highly spatial and directional sound events.

[018] Combining the use of first-order (global) and higher-order (directionally local) sound field diffusion estimation during rendering can increase performance in an organized coding framework. This is because the diffusion level is critical for rendering impression, as it distributes signal energy between the directional and diffuse rendering stream, see Figure 5a (ψ block). Spatially averaged global diffusion accurately captures this feature of the sound scene with good stability and can therefore provide better perceptual quality in practice.

[019] Lower bit rate scenarios allow transmitting only a single set of (first-order) estimates, therefore switching to higher bit rates and enabling the proposed architecture should not rebalance the direct-to-diffuse ratio of the rendered HOA signals. This is avoided by using global diffusion ψ to balance the global direct-to-diffuse ratio. Directionally local sector diffusion is then used to balance the directional local sector recoding.

[020] The combination of global diffusion and sector diffusion also allows diffusion-dependent bit economy in metadata, for example, by limiting the quantization steps of directional parameters to sectors that have predominantly diffuse content. In sound scenarios with high ψ, only little energy is distributed to the directional stream, thus requiring only coarse quantization of the directional parameterization.

[021] Furthermore, it can be assumed that the FOA was sufficiently restored in the decoder for a sufficient bit rate, which allows recovering first-order estimates in the decoder. This includes, in particular, ψ, therefore, no transmission is necessary.

[022] Figure 2 shows an example related to the previous technique. It can be seen Petition 870250071890, dated 08 / 14 / 2025, pp. 225 / 295 5 / 74 that a first-order ambisonic (FOA) signal 202 is split in the signal splitter 204 between a single directional path 221 and a global diffusion path 205. The signal splitter 204 is conditioned by a global diffusion Ψ of the FOA signal 202 (or, alternatively, by a complement to the global diffusion Ψ of the FOA signal 202, which may be 1 - Ψ). In the signal splitter 204, the FOA signal 202 is scaled by a weight that is in accordance with the directionality of the signal (1 - Ψ). In block 224, an omnidirectional pressure is applied to the FOA, to obtain a resulting directional signal 226. The directional signal 226 is also transformed in block 228, applying a spherical harmonic function of a DoA (Ω). In the splitter 204, the signal splitter 204 also emits a global broadcast signal 210, which is routed to a second path 205, for example, by weighting the FOA signal 202 by a weight conditioned by Ψ.In a power compensator (208) (as per document WO 2020 / 115311 A1), the global diffusion signal 210 is obtained. In block 260, the global diffusion signal 210 and the direction signal 222 are added together to obtain a HOA signal 262. The Ω DoA and the global diffusion Ψ are obtained from the bit stream.

[023] The aim is to model realistic sound scenarios more accurately by resolving simultaneous scenarios from multiple sources, resulting in improved perceptual quality during HOA audio coding and spatial enhancement compared to the current design. SUMMARY

[024] According to one aspect, an audio signal representation decoding unit is provided to generate an uncompressed ambisonic spatial audio signal representation from a compressed ambisonic spatial audio signal representation representing an audio signal, wherein the compressed ambisonic spatial audio signal representation includes at least one transport channel and side information, the side information includes sound field parameters, the sound field parameters include, for each spatial sector of a plurality of spatial sectors, the directional parameter(s) that provide(s) information about a direction of arrival, DoA, in Petition 870250071890, dated 08 / 14 / 2025, pages 226 / 295 6 / 74 spatial sector, the sound field parameters include, for at least one spatial sector, the sector diffusion parameter(s) that provide(s) information about the sector diffusion of the audio signal in at least one spatial sector, the audio signal representation decoding unit, includes a plurality of sector decoding paths, each sector decoding path being configured to decode a directional sector signal from the uncompressed ambisonic spatial audio signal representation in each spatial sector, applying to at least one transport channel or to a sector signal derived from at least one transport channel, the directional parameter(s) and the spatial sector diffusion parameter(s), the audio signal representation decoding unit, includes a global diffusion signal decoding path configured to derive a global diffusion signal applying to at least one transport channel,A global diffusion parameter or other information about the global diffusion of the audio signal, the audio signal representation decoding unit, includes a global diffusion signal inserter to combine the plurality of decoded directional sector signals and the global diffusion signal, to produce the uncompressed ambisonic spatial audio signal representation.

[025] In some examples, at least one transport channel may actually include (or at least may be processed, for example, by upmixing, to obtain) a plurality of transport channels. For example, at least one transport channel may actually include a plurality of transport channels mixed from a first number of transport channels (which may be 1 or a plural number) to a second number of transport channels (the second number of transport channels being greater than the first number of transport channels and therefore always being a plural number). Therefore, even if the bitstream includes a single transport channel (or a given number of transport channels), in some examples, the audio signal representation decoding unit may process the single transport channel (or a given number of transport channels). Petition 870250071890, dated 08 / 14 / 2025, pages 227 / 295 7 / 74 of transport channels) to obtain a mixed plural number (greater than the determined number of transport channels). Subsequently, the directional paths and the global broadcast signal decoding path are applied to the mixed plural number transport channels.

[026] According to one aspect, the audio signal representation decoding unit is configured to apply, to at least one transport channel or to a sector signal derived from the transport channel, the sector diffusion parameter(s) by weighting the transport channel, in at least one sector decoding path, using a mixing weight derived from the sector diffusion parameter(s), in order to derive the sector directional signal.

[027] According to one aspect, the audio signal representation decoding unit is configured to weight at least one transport channel or sector signal derived from the transport channel using the mixing weight which is, or is derived from, a positive coefficient received or processed from the sector diffusion parameter(s).

[028] According to one aspect, the audio signal representation decoding unit is configured to weight at least one transport channel, or sector signal derived from the transport channel, using the mixing weight, for at least one spatial sector, the mixing weight is, or is derived from, a coefficient indicative of a sector directionality in the specific spatial sector.

[029] According to one aspect, the audio signal representation decoding unit is configured to weight at least one transport channel or sector signal derived from the transport channel using the mixing weight, for each spatial sector, the mixing weight which is, or is derived from, a coefficient indicative of the relative directionality of the signal in the specific spatial sector over the relative directionalities of the total spatial sectors.

[030] According to one aspect, the representation decoding unit Petition 870250071890, dated 08 / 14 / 2025, pages 228 / 295 The 8 / 74 audio signal is configured to weight at least one transport channel or sector signal derived from the transport channel for at least one first spatial sector using a first mixing weight that is, or is derived from, a coefficient indicative of the sector directionality in the first spatial sector, and configured to weight at least one transport channel or sector signal derived from the transport channel for at least one second spatial sector using a second mixing weight, the audio signal representation decoding unit being configured to recover the second mixing weight being recovered, complementing itself, to a predetermined fixed value, the coefficient indicative of the sector directionality in the first spatial sector.

[031] According to one aspect, the audio signal representation decoding unit is configured to derive each of the (N1)-th mixing weights from parameters written in the side information, and to derive an N-th mixing weight by complementing the other (N1)-th mixing weights to a constant positive value, where N is the number of spatial sectors.

[032] According to one aspect, the decoding unit is configured, in each sector decoding path, to apply, to at least one sector signal, the directional parameter(s) by multiplying the signal of at least one sector by a vector of spherical harmonic functions evaluated along the DoA^s) in the spatial sector, so as to extend the directional signal to the spatial sector by a higher ambisonic order.

[033] According to one aspect, the decoding unit is configured to apply a spatial filter to at least one transport channel or processed version of at least one transport channel, to limit at least one transport channel to a spatial sector for each sector decoding path.

[034] According to one aspect, the decoding unit is configured to calculate at least one directional sector signal using xs = * Y(íls) = [xs * Yoo(íl), xE * Yi- * Yio(íl), xs * Yn(íl), ], Petition 870250071890, dated 08 / 14 / 2025, pages 229 / 295 9 / 74 where s indicates the space sector, s = the transport channel, or processed version thereof, in the specific space sector s, the directional parameter for the specific space sector s, and Y, which is a function of Ωε, is the vector of spherical harmonic functions given by (^b),Yi-i(^b),Yio(^b),Xli(^b), -^01(¾)] θ Ynm(^s) is a spherical harmonic of order ne degree m.

[035] According to one aspect, the decoding unit of any of the previous aspects, configured to calculate at least one directional sector signal for at least the specific spatial sector using = (1 - Ψ)» a, » Y(ílB) * xsem where ψ is the global diffusion parameter, Mo is the sector diffusion parameter expressed as the relative directionality of the sector in at least one sector signal, γ(Ώ) is a vector of spherical harmonic functions evaluated along the DoA in the specific spatial sector.

[036] According to one aspect, the decoding unit is configured to read the global diffusion parameter from the side information.

[037] According to one aspect, the decoding unit is configured to estimate the global diffusion parameter of at least one transport channel.

[038] According to one aspect, the decoding unit is configured to apply a global diffusion weight obtained from the global diffusion parameter, or information about the global diffusion of the audio signal, to weight at least one transport channel, thus obtaining a version of the global diffusion signal to be used in the decoding path of the global diffusion signal, and to apply a second weight, complementary to the global diffusion weight, to weight at least one transport channel, thus obtaining at least one globally non-diffuse signal to be processed in the sector's plurality of decoding paths.

[039] According to one aspect, the decoding unit is configured to derive the mixing weight(s) of the global broadcast signal and the directional sector signals from the global broadcast parameter, or the global broadcast information. Petition 870250071890, dated 08 / 14 / 2025, pages 230 / 295 10 / 74 of the audio signal.

[040] According to one aspect, the decoding unit is configured to apply, to at least one transport channel, a weighting parameter complementary to the global broadcast parameter used to derive the global broadcast signal, so that, for each sector decoding path, at least the transport channel is weighted using the weighting parameter.

[041] According to one aspect, the global diffusion signal decoding path is configured to weight at least one transport channel by a global diffusion gain, which is, or is derived from, the global diffusion parameter, or from other information about the global diffusion of the audio signal. According to one aspect, each of the multiple sector decoding paths is configured to weight at least one transport channel by a global directionality gain, which is, or is derived from, the global diffusion parameter, or from other information about the global diffusion of the audio signal.

[042] According to one aspect, the global diffusion gain is 1 + e is in accordance with , I ZH + 1\ 1+ σ(ψ) = 1+Ψ*16J XL+ 1 / where ψ is, or is derived from, the global diffusion parameter, or from other information about the global diffusion of the audio signal, L is an ambisonic input order and H is an ambisonic output order.

[043] According to one aspect, the global diffusion gain is1 +e is in accordance with + gW = , where ψ is either derived from the global diffusion parameter, or from other information about the global diffusion of the audio signal, and &°™ρ is a diffuse compensation factor.

[044] According to one aspect, the fuzzy compensation factor is given by Petition 870250071890, dated 08 / 14 / 2025, pages 231 / 295 11 / 74f(Σί=οΣ'=~ί 2 * ί + 1) 7cc.mp, , , 1 ί=ο ™=-ί 2 * I + 1 / where i is the degree of a spherical harmonic and L is the ambisonic order of the input signal and H is a higher ambisonic order, or a signal comprising the transport channels and channels generated through the use of decorrelators, where is the index of a spherical harmonic and takes values ​​of -iai.

[045] According to one aspect, the range of values ​​of the global diffusion gain is limited to a certain range of values ​​to avoid very strong deviations from the global diffusion signal.

[046] According to one aspect, the global diffusion signal decoding path includes a power compensation unit to apply gain to the global diffusion signal to adjust the power distribution in order to obtain a more physically realistic ambisonic output signal.

[047] According to one aspect, the audio signal representation decoding unit is configured to switch between: a low-order mode of operation, in which, among the plurality of sector decoding paths, at least one of the sector decoding paths is deactivated, while only one of the sector decoding paths is activated, wherein the side information does not contain the sound field parameter(s) for the deactivated at least one of the sector decoding paths; and a high-order mode of operation, in which, among the plurality of sector decoding paths, all sector decoding paths of the plurality are activated, or at least fewer sector decoding paths are deactivated in relation to the low-order mode of operation, wherein the side information also contains the sound field parameter(s) for the entire plurality of sector decoding paths, as well as the global diffusion parameter.

[048] According to one aspect, the representation of the audio signal is Petition 870250071890, dated 08 / 14 / 2025, pp. 232 / 295 12 / 74 configured to convert the spatial audio signal representation of at least one encoded transport channel into a decoded version of the encoded transport channel.

[049] According to one aspect, the audio signal representation comprises an EVS decoder for decoding at least one encoded transport channel into a decoded version of the encoded transport channel.

[050] According to one aspect, the audio signal representation decoding unit is configured to convert the decoded ambisonic spatial audio signal representation from the filter bank domain to the time domain.

[051] According to one aspect, the audio signal representation decoding unit is configured to perform upmixing of at least one transport channel from a first transport channel number to a second transport channel number that is larger than the first number.

[052] According to one aspect, the audio signal representation decoding unit comprises a mixing matrix estimator configured to process the sound field parameters, to derive a covariance matrix, or other covariance information, between different transport channels, wherein the mixing matrix estimator is configured to reconstruct a mixing matrix, or other mixing information, from the covariance matrix, or other covariance information, and apply the mixing matrix, or other mixing information, to the transport channels.

[053] According to one aspect, the mixing matrix estimator is configured to process the sound field parameter(s), including the DoA parameter(s) and the spatial sector plurality diffusion parameter(s) and the global diffusion parameter, or other global diffusion information, to derive the covariance matrix, or other covariance information, between different transport channels, wherein the mixing matrix estimator is configured to reconstruct a mixing matrix from the covariance matrix, Petition 870250071890, dated 08 / 14 / 2025, pages 233 / 295 13 / 74 or other covariance information, so as to employ the sound field parameter(s) to derive the covariance matrix, or other covariance information, for at least one frequency band, wherein the audio signal representation decoding unit is configured to derive the covariance matrix, or other covariance information, for at least one other frequency band without using the sound field parameters.

[054] According to one aspect, the audio signal representation decoding unit is configured to derive, for at least one other frequency band, the mixing matrix or other mixing information, from covariance information received from the side information.

[055] According to one aspect, the sound field parameters are modified to obtain a rotation of the sound field represented by the output ambisonic signal.

[056] According to one aspect, an apparatus method is provided, which comprises: the unit for decoding the representation of the sound signal from any of the previous aspects; A bitstream reader and dequantizer, configured to read a bitstream in which the low-order spatial audio signal representation is encoded, and to provide the high-order spatial audio signal representation to the audio signal representation decoding unit.

[057] According to one aspect, the apparatus further comprises a renderer, for rendering the audio signal of the ambisonic spatial audio signal representation.

[058] According to one aspect, the apparatus further comprises an encoding unit for encoding the representation of the high-order spatial audio signal into a second spatial audio signal representation.

[059] According to one aspect, an audio signal representation encoding unit is provided for encoding an audio signal representation. Petition 870250071890, dated 08 / 14 / 2025, pages 234 / 295 14 / 74 spatial input, representing an audio signal, in a compressed ambisonic spatial audio signal representation representing the audio signal, the audio signal representation encoding unit being configured to perform downmixing of the spatial input audio signal representation to derive at least one transport channel; The audio signal representation coding unit is configured to derive lateral information, wherein the lateral information includes sound field parameters; the sound field parameters include, for each spatial sector of a plurality of spatial sectors, the directional parameter(s) that provide(s) information about a direction of arrival (DoA) in the spatial sector; the sound field parameters include the sector diffusion parameter(s) that provide(s) information about the diffusion of the audio signal in at least one spatial sector; the audio signal representation coding unit includes a plurality of sector parameter estimators, wherein each sector parameter estimator is configured to process a specific sector signal of the input spatial sound signal representation in a specific spatial sector of the plurality of spatial sectors.In order to derive the directional parameter(s) and information about the diffusion of the audio signal in at least one spatial sector, the audio signal representation encoding unit includes a bitstream recorder to encode at least one transport channel and sideband information.

[060] According to one aspect, the audio signal representation coding unit includes a global diffusion parameter estimator to estimate a global diffusion parameter to be inserted into the side information.

[061] According to one aspect, the audio signal representation encoding unit is configured not to write a global diffusion parameter to the bitstream.

[062] According to one aspect, the coding unit of representation of Petition 870250071890, dated 08 / 14 / 2025, pages 235 / 295 The 15 / 74 audio signal is configured to estimate the relative directionality of each specific spatial sector with respect to the directionalities of all spatial sectors, and to write (e.g., in the sidebars) the coefficient (e.g., indicated as ai, a2, etc. below, and which may be mixing weights), or information indicative of the relative directionality, such as a sector diffusion parameter (e.g., a sector diffusion parameter for each spatial sector of the plurality of spatial sectors, or a sector diffusion parameter for at least one of the plurality of spatial sectors, or a sector diffusion parameter for each of all but one of the plurality of spatial sectors).

[063] In some examples, it is possible to write (for example, in the side information) each coefficient (ai, a2, etc.), or information indicating relative directionality, for each sector diffusion parameter. In some examples, all coefficients except one, or information indicating relative directionality, are written (for example, in the side information), while a single coefficient, or information indicating relative directionality, is ignored: since the sum of the coefficients (or information indicating relative directionality) can be known (for example, 1), it is possible to skip one of the coefficients, while the decoder can reconstruct it. For example, it may be that = 1 and therefore it is possible to simply write, as side information, \ while the audio signal representation decoding unit can derive a2 = 1 - ai.Therefore, while in some examples the secondary information may include all the coefficients, in other examples there may be at least one coefficient written (for example, just ai), θ ΡθΙο minus one coefficient may be obtained from the first (for example, as a2 = 1 _ ai).

[064] According to one aspect, the audio signal representation coding unit is configured to estimate relative directionality as including at least one of a first and a second spatial sector, respectively indicated by ai and a2, and satisfies: Petition 870250071890, dated 08 / 14 / 2025, pages 236 / 295 16 / 74 1-Ψιai(1 - Ψ,) + (1- Ψ2) and a2— 1—where is, or is obtained from the diffusion information of the sector for the first space sector and is, or is obtained from the diffusion information of the sector for the second space sector.

[065] According to one aspect, the audio signal representation coding unit is configured to estimate the relative directionality to include two or more sectors according to -Ψί A — -----------1Σ / ι with Sj3j = 1, where i indicates the i-th specific spatial sector, ej indicates a j-th generic spatial sector of the plurality of spatial sectors, and * indicates the sector diffusion information for the i-th, specific, spatial sector and each j-th generic spatial sector.

[066] According to one aspect, the audio signal representation encoding unit is configured to perform an active downmix of the audio signal, or a processed version thereof, using a downmix matrix, or other downmix information, calculated by a downmix information calculator, wherein the downmix information calculator is configured to process the sound field parameter(s) to derive the downmix matrix, or other downmix information, based on the global diffusion parameter and the sector diffusion parameters and the directional parameters for each spatial sector of the spatial sector plurality.

[067] According to one aspect, the information matrix calculator is configured to perform a cross-channel prediction to derive the mixing matrix. Petition 870250071890, dated 08 / 14 / 2025, pages 237 / 295 17 / 74 descending, or other descending mixing information, based on an inter-channel covariance matrix, or other inter-channel covariance information, the inter-channel covariance matrix or other inter-channel covariance information that is derived from the directional parameter(s) and sector diffusion parameter(s) for each spatial sector of the sector plurality and a global diffusion.

[068] According to one aspect, the interchannel covariance matrix C is defined as having the element between the ambisonic channel with degree and index 1el*, respectively, and the ambisonic channel with degree and index era, respectively, and is computed according to (1 - Ψ) * Es* a2* * Yjv(nj + (1 - Ψ)* Ex*(1- a)2* Ylm(n2) * ΥI%ί(Ω2) + +Ψ*σ- * Es* 8lmJv where E* is the signal energy, θ0ç|e|taqeKronecker being 1 on the diagonal of the interchannel covariance matrix and 0 off the diagonal of the interchannel covariance matrix, and are the first and second directional parameters, respectively, and “a” is a relative directionality, or other parameter indicative of a ratio, or other information about the relationship, between the directionality in the spatial sector. Regarding the total directionalities of all spatial sectors, ψ is indicative of the global diffusion parameter, and σ is an energy scaling factor.

[069] According to one aspect, the covariance matrix between channels or other covariance information between channels is based on an energy weighted by the spherical harmonics evaluated in the DoAs (^1. ^2j' and mixing weights ( alra2,..,aN^ for spatial cac|a.

[070] According to one aspect, the audio signal representation encoding unit is additionally configured to convert the input spatial audio signal representation into the filter bank domain to derive a filter bank version of the input spatial audio signal representation, additionally configured to perform downmixing of the version Petition 870250071890, dated 08 / 14 / 2025, pages 238 / 295 18 / 74 of the filter bank domain of the input spatial audio signal representation to derive at least one transport channel in the filter bank domain, and configured to perform a filter bank synthesis of at least one transport channel from the filter bank domain in the time domain.

[071] According to one aspect, the audio signal representation encoding unit is configured to perform down-mixing of the input spatial audio signal representation using a channel selector to derive at least one transport channel, selecting lower-order channels from higher-order channels of the input spatial audio signal representation.

[072] According to one aspect, the audio signal representation coding unit is configured to perform enhanced voice services coding, EVS, in order to provide an EVS-encoded version of at least one transport channel.

[073] In one respect, the audio signal representation encoding unit is provided configured to switch between: A low-order mode of operation, in which, among a plurality of sector paths (e.g., sector encoding paths), at least one of the sector paths is deactivated, while only one of the sector paths is activated, wherein the side information does not contain the sound field parameter(s) for the deactivated at least one of the sector paths; and a high-order mode of operation, in which, among the plurality of sector paths, all sector paths of the plurality are activated, or at least fewer sector paths are deactivated compared to the low-order mode of operation, wherein the side information also contains the sound field parameter(s) for the entire plurality of activated sector paths, as well as a global diffusion parameter.

[074] According to one aspect, the audio signal representation encoding unit can be configured to select between low-order and high-order operating modes based on the bit rate, so Petition 870250071890, dated 08 / 14 / 2025, pages 239 / 295 19 / 74 selects the low-order operating mode in case of low bit rate and the high-order operating mode in case of a bit rate higher than the low bit rate.

[075] In one aspect, the audio signal representation encoding unit can be configured to select between low-order and high-order operating modes based on measurements related to network connection quality (e.g., latency-related measurements and / or error rate measurements and / or connection bandwidth measurements, etc.), such that: If measurements related to network connection quality indicate poor quality (e.g., high latency and / or high error rate and / or low connection bandwidth, etc.), the audio signal representation encoding unit selects the low-order operating mode and, If measurements related to network connection quality indicate high quality, with high quality being superior to low quality (e.g., low latency and / or low error rate and / or connection bandwidth, etc.), the audio signal representation encoding unit selects the high-order operating mode.

[076] (To perform the selection, measurements related to network connection quality can be evaluated against a predetermined quality threshold, in order to classify the network connection quality. For example, to determine whether the quality is high or low, measurements related to network connection quality can be evaluated against at least one quality threshold, in order to distinguish between high quality and low quality. For example, latencies can be evaluated against a latency threshold, in order to classify the quality as low if the latencies, for example, the average latencies, are above the latency threshold, and to classify the quality as high if the latencies, for example, the average latencies, are below the latency threshold. Or the error rate can be evaluated against an error rate threshold, in order to Petition 870250071890, dated 08 / 14 / 2025, pages 240 / 295 20 / 74 to classify the quality as low if the error rate, for example, the average error rate, is higher than the error rate threshold, and to classify the quality as high if the error rate, for example, the average error rate, is lower than the error rate threshold. Or the connection bandwidth can be evaluated relative to a connection bandwidth threshold, so as to classify the quality as low if the connection bandwidth, for example, the average connection bandwidth, is below the connection bandwidth threshold, and to classify the quality as high if the connection bandwidth, for example, the average connection bandwidth, is above the connection bandwidth threshold.

[077] In one aspect, the audio signal representation encoding unit can be configured to select between low-order and high-order operating modes based on battery supply-related measurements, such that: If battery power measurements indicate low battery power for a battery powering the audio signal representation encoding unit, the audio signal representation encoding unit selects low-order operating mode and, If the battery power measurements indicate a high battery power level, greater than the low battery power level, the audio signal representation encoding unit selects the high-order operating mode.

[078] (To perform the selection, the battery power can be evaluated against a predetermined battery power threshold (charge threshold) in order to classify the battery power and perform the selection based on the classification. For example, to determine whether the battery power is high or low, battery power measurements can be evaluated against at least one battery power threshold (charge threshold) in order to distinguish between high battery power and low battery power.) Petition 870250071890, dated 08 / 14 / 2025, pages 241 / 295 21 / 74

[079] According to one aspect, the audio signal representation encoding unit can be configured to select between low-order operating mode and high-order operating mode based on a feedback signal from a receiver (e.g., decoding unit), so as to select the operating mode requested in the feedback signal.

[080] In one aspect, an audio encoder is provided, which comprises: the encoding unit for representing the sound signal of a previous aspect; A bitstream quantizer and recorder for writing, in a bitstream, the low-order spatial audio signal representation.

[081] According to one aspect, a method is provided for decompressing an ambisonic spatial audio signal representation representing an audio signal, wherein the compressed ambisonic spatial audio signal representation includes at least one transport channel and side information, wherein the side information includes sound field parameters, the sound field parameters include, for each spatial sector of a plurality of spatial sectors, the directional parameter(s) that provide(s) information about a direction of arrival, DoA, in the spatial sector, the sound field parameters include, for at least one spatial sector, the sector diffusion parameter(s) that provide(s) information about the sector diffusion of the audio signal in at least one spatial sector, the method includes the use of a plurality of sector decoding paths,whereby each sector decoding path decodes a directional sector signal from the ambisonic spatial audio signal representation in each spatial sector, applying to at least one transport channel or to a sector signal derived from the transport channel, the directional parameter(s) and the spatial sector diffusion parameter(s), the method includes the use of a diffusion signal decoding path, Petition 870250071890, dated 08 / 14 / 2025, pages 242 / 295 22 / 74 global to derive a global diffusion signal by applying, at least, to a transport channel, a global diffusion parameter or other information about the global diffusion of the audio signal, the method includes combining, through a global diffusion signal inserter, the plurality of decoded directional sector signals and the global diffusion signal, to produce the uncompressed ambisonic spatial audio signal representation.

[082] According to one aspect, a method is provided for encoding an input spatial audio signal representation, representing an audio signal, into a compressed ambisonic spatial audio signal representation representing the audio signal, the method includes deriving at least one transport channel and side information, wherein the side information includes sound field parameters, the sound field parameters include, for each spatial sector of a plurality of spatial sectors, the directional parameter(s) that provide(s) information about a direction of arrival, DoA, in the specific spatial sector, the sound field parameters include the sector diffusion parameter(s) that provide(s) information about the diffusion of the audio signal in at least one spatial sector, the method includes the use of a plurality of sector parameter estimators,whereby each sectorial parameter estimator processes a specific sectorial signal from the representation of the input spatial audio signal in a specific spatial sector of the plurality of spatial sectors, in order to derive the directional parameter(s) and information about the diffusion of the audio signal in at least one spatial sector, the method including the use of at least one transport channel encoding and secondary information in a bitstream.

[083] According to one aspect, a non-transient storage unit is provided that stores instructions which, when executed by a processor, cause the processor to execute the method of the aspects. Petition 870250071890, dated 08 / 14 / 2025, pp. 243 / 295 23 / 74 previous.

[084] According to one aspect, a representation of a compressed ambisonic audio signal is provided that includes at least one transport channel and side information, wherein the side information includes sound field parameters, the sound field parameters include, for each spatial sector of a plurality of spatial sectors, the directional parameter(s) that provide(s) information about a direction of arrival, DoA, in the spatial sector, the sound field parameters include, for at least one spatial sector, the sector diffusion parameter(s) that provide(s) information about the sector diffusion of the audio signal in at least one spatial sector, and a global diffusion parameter.

[085] In one aspect, a compressed ambisonic audio signal representation is provided, for example, generated according to the encoding method of an input spatial audio signal representation described above. FIGURES Figure 1 shows examples of first-order basis functions of an ambisonic representation; Figure 2 shows one embodiment of the previous technique; Figure 3 shows an example of how to divide the sphere into two sectors; Figure 4 shows the basic functions of Figure 1 after filtering with one of the sectors of Figure 3; Figures 5a-5c show examples of audio signal representation decoding units according to the present disclosure; Figure 6 shows the results of a subjective hearing test comparing the invention with the prior art; Figure 7 shows an example of an audio signal representation encoding unit according to the present disclosure; Figure 8 shows an example of a device that includes an audio signal representation decoding unit from and of Figures 5a-5c; Petition 870250071890, dated 08 / 14 / 2025, pages 244 / 295 24 / 74 Figures 9a and 9b show examples of current techniques; Figures 9c and 9d show the results obtained with the current techniques; Figure 10 shows an example of an audio signal representation encoding unit according to the present disclosure. EXAMPLES

[086] Reference is made to Figures 7 and 10. Each shows an audio signal representation coding unit (700, 700b) (also called an encoder) for encoding an input spatial audio signal representation (702), representing an audio signal (e.g., in higher-order ambisonics), into a compressed ambisonic spatial audio signal representation (502, 802) representing the audio signal (702). The audio signal representation coding unit 700, 700b can be configured to perform downmixing (e.g., in downmixing step 1700a or 1700b) of the input spatial audio signal representation (702) to derive at least one transport channel (736). The audio signal representation coding unit can be configured to derive side information (503). The side information (503) may include sound field parameters (e.g., 714, 718, 549, 529).The sound field parameters (e.g., 714, 718, 549, 529) may include, for each spatial sector of a plurality of spatial sectors, directional parameter(s) that provide information about a direction of arrival, DoA, in the spatial sector. The sound field parameters may include sector diffusion parameter(s) that provide information about the diffusion of the audio signal (702) in at least one spatial sector (e.g., the diffusion parameters may be written in the 502, 802 audio signal representation for at least one spatial sector, but may provide diffusion information for all spatial sectors). The audio signal representation coding unit (700, 700b) may include a plurality of sector parameter estimators (712, 7211, 7212, 721n). Each sector parameter estimator (712, 7211, 7212, 721n) can be configured for. Petition 870250071890, dated 08 / 14 / 2025, pages 245 / 295 25 / 74 process a specific sector signal (710, 7101, 7102, 710n) from the input spatial audio signal representation (702) in a specific spatial sector from the plurality of spatial sectors, so as to derive the directional parameter(s) and diffusion information of the audio signal (702) in at least one spatial sector. The audio signal representation encoding unit may include a bitstream recorder (750) to encode at least one transport channel (736, 501) and side information (503), which may be understood as incorporating the compressed ambisonic spatial audio signal representation (502, 802).

[087] Figures 5a-5c show examples of audio signal representation decoding units (500, 500b, 500c) (also called decoders) for generating an uncompressed ambisonic spatial audio signal representation (562) from a compressed ambisonic spatial audio signal representation (502) representing an audio signal, the compressed ambisonic spatial audio signal representation (502, 802) which may be, for example, the compressed ambisonic spatial audio signal representation (502, 802) generated by the audio signal representation encoding unit (700, 700b). As explained above, the ambisonic spatial audio signal representation (502, 802) may include at least one transport channel (501) and side information (503). The side information (503) may include sound field parameters (e.g., 529, 549, 718).The sound field parameters may include, for each spatial sector of the plurality of spatial sectors, directional parameter(s) (e.g., 529, 549, 718) that provide information about a direction of arrival, DoA, in the spatial sector. The sound field parameters may include, for at least one spatial sector, sector diffusion parameter(s) (529, 549) that provide information about the sector diffusion of the audio signal in at least one spatial sector (as explained above, the diffusion parameters may be in the 502, 802 audio signal representation for at least one spatial sector, but may provide diffusion information for all spatial sectors). The audio signal representation decoding unit may... Petition 870250071890, dated 08 / 14 / 2025, pages 246 / 295 26 / 74 include a plurality of sector decoding paths (521, 541). Each sector decoding path (521, 541) can be configured to decode a directional sector signal (532, 552) from the uncompressed ambisonic spatial audio signal representation (562) in each spatial sector, applying to at least one transport channel (501) or to a sector signal (528, 548), derived from at least one transport channel, the directional parameter(s) (529, 549) and the spatial sector diffusion parameter(s) (529, 549). The audio signal representation decoding unit may include a global diffusion signal decoding path (505) for deriving a global diffusion signal (510) applying at least to a transport channel (501), a global diffusion parameter (507, 507', Ψ) or other information about the global diffusion of the audio signal.The audio signal representation decoding unit may include a global diffusion signal inserter (560) to combine the plurality of decoded directional sector signals (532, 552) and the global diffusion signal (510), to produce the uncompressed ambisonic spatial audio signal representation (562).

[088] Below, the units discussed above are exemplified in detail.

[089] Figure 5a shows the audio signal representation decoding unit 500. The audio signal representation decoding unit (audio signal representation decoding unit) 500 can receive, at the input, a compressed ambisonic spatial audio signal representation (e.g., FOA signal) 502 and provide, at the output, an uncompressed ambisonic spatial audio signal representation 562 (e.g., HOA, or high-order uncompressed ambisonic spatial audio signal representation). (The FOA signal 502 can be replaced by a low-order ambisonic signal, and the HOA signal 562 can be a higher-order ambisonic signal than the low-order ambisonic signal 502).

[090] The ambisonic compressed spatial audio signal representation 502 may include at least one 501 transport channel (also represented in some cases mathematically as xl). The 501 transport channel may include, for example, a downmixed version of a signal representation of Petition 870250071890, dated 08 / 14 / 2025, pages 247 / 295 27 / 74 original audio (from an audio signal). In general terms, at least one 501 transport channel can be understood as having downmix channels relative to the original 702 audio signal channels. Each channel can be an ambisonic component (ambisonic components are represented in Figure 1). For example, there can be a plurality of channels, e.g., four channels, in the case of the compressed ambisonic spatial audio signal representation 502 being an FOA signal. Each 501 transport channel can be provided in, or converted from, a filter bank domain. Although here below it is often referred to as at least one transport channel, this is also valid for a plurality of channels (e.g., four or more channels). Notably, the 501 transport channel(s) can be processed, through the elements of Figure 5a, to become the uncompressed 562 spatial audio signal representation.

[091] If the input compressed ambisonic spatial audio signal representation 502 is not in the filter bank domain, there may be a filter bank (not shown in Figure 5a) that would convert the compressed ambisonic spatial audio signal representation into the filter bank domain, upstream of the elements shown in Figure 5a. In some examples, downstream of the elements shown in Figure 5a, there may be another filter bank synthesizer (also not shown in Figure 5a), to provide the uncompressed ambisonic spatial audio signal representation 562 in a time domain, for example.

[092] At least one (or more) transport channel 501 (or the compressed ambisonic spatial audio signal representation 502) may be an FOA signal or, more generally, a low-order ambisonic signal. A task of the audio signal representation decoding unit 500 (audio signal representation decoding unit) may be to obtain the uncompressed ambisonic spatial audio signal representation 562 as the HOA signal, or at least a higher-order ambisonic signal, corresponding to and providing possibly reliable audio information from the HOA signal 702 inserted into the encoder 700 or 700b.

[093] The compressed ambisonic spatial audio signal representation 502 Petition 870250071890, dated 08 / 14 / 2025, pages 248 / 295 28 / 74 may include 503 side information. 503 side information may include sound field parameters. In examples, all time frequency blocks within the same spatial sector may or may not share at least some of the sound field parameters. For example, some sound field parameters may be the same for all bands. In some examples, bands may be grouped to save on metadata. Furthermore, for some signals, the parameters may be the same for some of all bands (generally, they may be different for each band). In some cases, different bands may have different sound field parameters.

[094] The sound field parameters 529, 549 may include, for specific spatial sectors of the plurality of spatial sectors, at least one directional parameter for each specific spatial sector. For example, if there are two spatial sectors, there may be two directional parameters (one for each sector) for each frequency band. The sector index is denoted as s. The total number of sectors is N. Here it is often exemplified with N=2 (or sometimes, for generality, with N>2). Space may be divided into spatial sectors. The position of the spatial sectors may be fixed (i.e., it may be known as priority by either the encoding unit 700, 700b or the decoding unit 500). The directional parameter may be, or provide information about, the direction of arrival, DoA, in the specific spatial sector. The DoA may be represented, for a specific spatial sector s, with the symbol where s indicates the specific sector.For example, in the case of having two space sectors s=1 and s=2, we can have s=3, etc. In the case of more than two space sectors, there will also be s=2, s=4, etc. Therefore, at least one directional parameter s=2, s=4 can be provided. Thus, for each space sector, a specific directional parameter can be defined (for example, for each time-frequency block). Unlike the state of the art, as shown in the example in Figure 2 (where there is only a single DoA in the entire space), here there is a directional parameter for each space sector, and the space sectors are more than one. Petition 870250071890, dated 08 / 14 / 2025, pages 249 / 295 29 / 74

[095] The sound field parameters 529, 549 may also include at least one sector diffusion parameter (529, 549 using the same reference numerals used for the directional parameter), which may provide information on sector (local) diffusion or, complementarily, on sector (local) directionality. The sector diffusion parameter will often be denoted as Ψδ where s indicates the sector and, in the case of two sectors, may be denoted as Ψι and Ψ2 (or as 1-Ψι and 1-Ψ2 when referring here to sector directionality). Another name for diffusion could be, for example, diffuse energy ratio Ψ = (diffuse energy) / (total energy) (where / means mathematical division). Global diffusion (or global diffusion energy ratio) can be T^(diffuse energy in space) / (total energy in space), while sector diffusion (or sector diffusion energy ratio) can be ML^(diffuse energy in sector s) / (total energy in sector s).

[096] It should be noted here that directionality must be understood as a concept complementary to diffusion (global diffusion and sectoral diffusion). Whether reference is made to Ψ1 or Ψ2 (in terms of diffusion) or to 1-Ψι and 1-Ψ2 (in terms of “directionality”), the indicative information of “sectoral diffusion” is nevertheless present, since 1-Ψι and 1-Ψ2 are also indicative of Ψ1 and Ψ2, and vice versa (the same applies to global diffusion Ψ and its complementary information 1-Ψ). Another name for directionality could be, for example, directional energy ratio and diffusion which is the complement of the diffuse energy ratio (e.g., 1-Ψ=1-(diffuse energy) / (total energy)).

[097] In more detail, a distinction is made here between “sector directionality” (1Ψ1, 1-Ψ2) and “directional information” (e.g., given in terms of ^eΩ-): “directional information” and DoA provide geometric information about the direction of the signal (e.g., intensity vector), without specifically indicating any weight or energy or intensity or pressure information; while “directionality” (1-Ψ1, 1-Ψ2) refers to concepts such as weight, intensity, energy, pressure, etc., that characterize sound, without providing information about DoA. In general terms, the more the audio signal is Petition 870250071890, dated 08 / 14 / 2025, pages 250 / 295 30 / 74 is locally diffuse in the space sector, except the audio signal is locally directional in the same space sector.

[098] Secondary information 503 may also include, for example, as a sound field parameter, a global diffusion parameter or other global diffusion parameters (507, 507', Ψ), or other information about the global diffusion of the sound signal. The global diffusion parameter is usually indicated by Ψ without subscripts, and is a global characteristic that describes the input signal. The global diffusion parameter Ψ can therefore provide information for weighting the FOA transport channels 501 (e.g., in a splitter 504, see below) to derive a diffusion component 506 in the path 505 (also shown in Figures 1 and 4 as a diffusion component). Another way to provide information about the global diffusion of the audio signal could be to indicate 1-Ψ (or B-Ψ with B>0, e.g., fixed, e.g., known a priori): even though 1-Ψ is complementary information to global diffusion, it is still information about global diffusion.

[099] As such, the global diffusion parameters (507, 507', Ψ), or other information about the global diffusion of the audio signal, can in some cases also be estimated, for example, when the compressed ambisonic spatial audio signal representation 502 has four or more transport channels 501. In other situations, the global diffusion parameters (507, 507', Ψ), or other information about the global diffusion of the audio signal, can be encoded in the side information 503.

[0100] At least one transport channel can be divided between a globally diffused signal (weighted by Ψ, for example) and a non-globally diffused signal (weighted by 1-Ψ, for example). However, the inventors understood that the non-globally diffused signal is not necessarily a “fully directional signal” (i.e., it is not necessarily fully directional) and not necessarily uniquely distributed in a single DoA, but can also be distributed, across multiple sectors, between a local directional component (sector directional component) and a local diffused component (sector diffused component). Petition 870250071890, dated 08 / 14 / 2025, pages 251 / 295 31 / 74

[0101] It will be shown that the inventors also understood that it is not strictly necessary to calculate (or have written in the sound field parameters), for each sector, both the sector's directional component and the sector's diffusion component. In contrast, it is more easily possible to derive, for each sector, a relative directionality by measuring the directionality of each spatial sector over the total amount of directionalities of all spatial sectors. By weighting at least one 501 transport channel, for each spatial sector, through a mixed weight derived from the relative directionality of that spatial sector, it is simply possible to derive a sectoral directional signal that takes into account, in itself, both its DoA and its sectoral diffusion. (It will be shown that the relative directionality can be the ratio between the sector directionality in a spatial sector and the sum of the sector directionalities in the totality of spatial sectors).

[0102] Reference is frequently made below to the sector diffusion parameters. Relative directionalities can be examples of diffusion parameters.

[0103] As can be seen in Figure 5a, at least one transport channel 501 can be subjected to a splitter 504, or other element in which weights conditioned by the global diffusion parameters (507, 507', Ψ), or other information about the global diffusion of the audio signal, are applied. The splitter (or other element) 504 can divide the compressed ambisonic spatial audio signal representation 502 into two signals, so as to emit a globally diffused signal 506 (in the compressed version, FOA) and a remaining globally non-diffuse signal 520 (in the compressed version, FOA).The global diffusion signal 506 can be understood as being weighted by a weight (e.g., Ψ) that increases with increasing diffusion (e.g., high diffusion will cause a high Ψ, total diffusion will cause Ψ=1 or another maximum value, and low diffusion will cause a low Ψ, and no diffusion will cause Ψ=0; at Ψ>0.5 the diffusion signal 506 tends to dominate the globally non-diffuse signal, while at Ψ<0.5 the remaining, globally non-diffuse signal 520 tends to dominate over the diffusion signal 506). The globally non-diffuse signal 520 can be understood as being obtained from the transport channels 501. Petition 870250071890, dated 08 / 14 / 2025, pages 252 / 295 32 / 74 weighted by a weight (e.g., 1-Ψ) that increases with decreasing global diffusion (i.e., increases with increasing global directionality). The globally non-diffuse 520 signal may have the remaining energy. It may be, however, that the energy of the globally non-diffuse 520 signal is, in turn, locally diffuse within a single sector. Furthermore, the globally non-diffuse 520 signal will therefore be filtered into multi-sector signals (528, 548), and each sector signal (528, 548) will, in turn, be continued to arbitrary higher ambisonic orders using the spherical harmonics of those orders into a directional sector signal (532, 552).

[0104] In the global broadcast signal decoding path 505, in block 508 (energy compensator), a gain1 + can be provided, to weight the global broadcast signal (component) 506 so that its energy matches the energy of a physically correct HOA signal (see WO document). 2020 / 115311 A1). In some cases, the gain l+g^) may be I +^(ψ) = where ψ is either obtained from the global diffusion parameter (507, 509), or from other information about the global diffusion of the audio signal; L is an ambisonic input order and H is an ambisonic output order. (Other formulas are possible).

[0105] Alternatively, the gain can be chosen as + g(T) = Jl+T*UCPínp-l) , where the diffuse compensation factor can be / VH π f 1\ f _1=0m=-l 2*7 + 1 / / comp / _ .1 xj π L- V1il=02jm=-[ 2*7 + 1 / where í is the degree of a spherical harmonic and L is the ambisonic order of the input signal, or a signal comprising the transport channels and, optionally, channels generated through the use of decorrelators, where is the index of Petition 870250071890, dated 08 / 14 / 2025, pages 253 / 295 33 / 74 is a spherical harmonic and assumes values ​​of -!a\ and H is a higher ambisonic order.

[0106] The range of values ​​for the global diffusion gain can be limited to a certain range of values ​​to avoid very strong deviations from the global diffusion signal (506).

[0107] The 508 power compensator unit can apply gain to the global diffusion signal 506 to adjust the power distribution in order to obtain a more physically realistic ambisonic output signal.

[0108] Note that it is generally considered °-ψ-1 (where ψ = 0 when the signal is completely directional and ψ = 1 when the signal is completely diffused; in some examples, 1 may be replaced by a value B>0). Through the gain, the global diffusion signal FOA (component) 506 is amplified through the gain1 +5(ψ). Notably, a higher diffusion (e.g., ψ close to 1) implies a higher gain than in the case of lower diffusion (e.g., ψ close to 0).

[0109] Figure 5a shows that the global diffusion parameters (507, 507', Ψ) or other information about the global diffusion of the audio signal can be obtained, as parameter 507, from the side information 503 or, alternatively, can be estimated in an optional diffusion estimator 570 estimated as 507' and / or 509' (for example, when the signal is ambisonic or multichannel), for example, using a pseudo-intensity vector or covariance-related techniques (507' and 509' can be the same). The diffusion 507' (509') can be obtained, by the optional diffusion estimator 570, from the intensity vector and the average energy (this is the conformance form used for DirAC).

[0110] The output of the 508 power compensator block, here indicated as 510, is the global diffusion signal 510 of the HOA signal, wherein the diffuse component of the HOA output signal. Typically, it contributes only to the first-order channels of the higher-order output signals.

[0111] In parallel with the processing in the global diffusion path 505, the globally non-diffuse signal 520 (for example, resulting as output 520 of the divider 520, Petition 870250071890, dated 08 / 14 / 2025, pages 254 / 295 34 / 74, for example, after the FOA 501 signal is scaled by ψ, can be processed in the plurality of sector decoding paths 521, 541. For simplicity, Figure 5a shows only two sector decoding paths, 521 and 541. In general, however, there can be an arbitrary number of sector decoding paths. In some implementations, sector decoding paths N=2 may be a reasonable design choice, representing a good trade-off between the need for good results and the need to keep computational effort low.

[0112] The globally non-diffuse signal 520 can be subjected, in the spatial filtering stage 574, to spatial filtering. The inputs of the spatial filtering stage 574 are indicated with 522 and 542 (which can be signals equal to each other, and also equal to the globally non-diffuse signal 520), each of the inputs 522 and 542 entering a respective spatial filtering block 524, 544. The spatial filtering blocks 524, 544 are part of the spatial filtering stage 574, and each spatial filtering block filters the globally non-diffuse signal 520, to limit the globally non-diffuse signal 520, in each sector decoding path, to a specific spatial sector.Therefore, at the output of spatial filtering block 524 on path 521, the transport channel(s) (in their sector-limited version 528) is limited to sector s=1, while at the output of spatial filtering block 526 on sector decoding path 522, the transport channel(s) (in their sector-limited version 548) is limited to sector s=2. To achieve spatial filtering, spatial filtering stage 574 can perform beamforming. Downstream of spatial filtering stage 574, the signal from sector 528 of spatial sector s=1 is different from the signal from sector 548 of spatial sector s=2, since they are limited to different spatial sectors.

[0113] Notably, the spatially filtered signal 528, 548 from each sector decoding path 521, 541 can be understood as still subject to another subdivision between a diffuse sector component and a DoA component: the sector signal simply lacks the global diffusion component (signal 510) that was already Petition 870250071890, dated 08 / 14 / 2025, pp. 255 / 295 35 / 74 polished in the divider 504 (the global diffusion component can be seen as acting as a common mode, which was removed in the divider 504). Therefore, each spatially filtered signal 528, 548 emitted by a relative block 526, 546 can be considered a sector signal, which provides signal information for the specific spatial sector.

[0114] The spatial filtering stage 574 can be instantiated by a plurality of spatial filters in blocks 524, 544 and so on. In each of the blocks 524, 544 and so on, for each spatial sector, a beamforming can be performed. In some examples, these can be obtained by _ f xs-w3* xL where w contains the beamforming weight vector for the spatial sector if Ai is the signal 528, 548 emitted by block 574 and T represents the transpose operator. The beamforming weight vector can be known a priori by the decoder.Notably, 3 can be an operator in the abstract representation T with multiple elements, each of which is a weight (for example, =[w,i.0'wi-xs'hio,3'wii,s,]).

[0115] In the subsequent stage of the 572 sector signal processor (including the 528, 548 sector signal processor blocks), the 528, 548 sector signals (processed transport channels, spatially filtered signals, etc.) are continued (extended) to a higher ambisonic order using a spherical harmonic vector evaluated along the spatial sector DoA (i.e., the local DoA internal to a specific spatial sector, for example, as indicated in the 529, 549 sound field parameters of the 503 side information). This can, for example, be achieved (or, in any case, verified) by the formula: x3 = y(n3) * x3 = x3 * Fi_lrx3 * ...] with the scalar value x* being one of the sector signals 528, 548 (for example, remembering that Xj = * 'ϊζ) indicating the specific spatial sector. A scale by 1-ψ (applied in the divisor 504) is not shown explicitly here, as it enters the formula through the sector signals Xs. is a vector of spherical harmonic functions (for example, calculated or read from tables by the decoder) that allows reconstructing the higher-order directional sector signal 532, 555 along the Petition 870250071890, dated 08 / 14 / 2025, pages 256 / 295 36 / 74 respective space sector DoA^a 'l'3éo vector of the decoded directional component of the space sector s in HOA.

[0116] The vector components are the real spherical harmonics in the ACN (ambisonic channel numbering) order (where Ω can be in terms of p, for example). These are defined as (see document [WO 2020 / 115311 A1]) Mm* P,m* sin / ? * sinÇlml * ω) if m < 0 * P™ * sin Θ * cos (|m| * φ) if m > 0 (where ll indicates the absolute value, i.e., l = +1, l°l = -0 and l+1l = +1) with the associated Legendre polynomials and a normalization term for the Legendre functions and the trigonometric functions that takes the following form for SN3D ([WO 2020 / 115311 A1 ], [Zotter and Frank]): / V; = ----------- Ϊ ------------l'mJ 4ít (i + |m|)!

[0117] For the ambisonic order L, the indices in traverse 1=0,..,L in=-;,..,f, respectively, where $mé 1 for m=0 and 0 otherwise, and ! indicates the factorial.

[0118] The spatially filtered signal 528, 548 can therefore be submitted to stage 572 of the sector signal processor, which can include the plurality of blocks 530, 550 for paths 522, 542, to obtain directional sector signals 532, 552 (in higher-order ambisonic format), respectively. For example, a sector 530 signal processor block can be applied to the filtered signal 528 as emitted by block 524, while the sector 550 signal processor block can be applied to the signal 548 as emitted by block 544 for path 541. This can be repeated, for example, for each spatial sector s.It will be shown that, operating accordingly, each block 530, 550, and so on of the 572 stage of the sector signal processor can make use of a sector diffusivity parameter (e.g., Ψι and Ψ2, or ai, a2, as discussed below) and / or DoA for the spatial sector s, as instantiated in and for the path 521, 541,. Petition 870250071890, dated 08 / 14 / 2025, pages 257 / 295 37 / 74 respectively. In practice, for each spatially filtered signal 528, 548 from each path 521, 541, a directional signal 532, 552 of the spatially filtered signal 528, 548 is recovered. Since, as explained above, the spatially filtered signal 528, 548 from each path 521, 541 is the signal of a specific sector, it is possible to imagine the signal 528, 548 as a directional component (local component) of the audio signal in the specific spatial sector.

[0119] For example, in the case of two spatial sectors (s=2), it is possible to use the coefficients (mixing weights) ai and a2, each expressed as a ratio indicating, for example, the directionality of sector 1 - (respectively 1 - ψ2) over the sum of the directionalities of all sectors 1 - + 1 - (respectively, the same 1 ~ %. + 1 ~ ψ2). An example might be: - Ψχ+ 1 - Ψ2e =1-ψ21-ΨΧ+ 1-Ψ2 which in this case (since s=2) is a2= 1 — ox. ai and a2 are therefore complementary to 1 (or another B>0 in other examples). At least one of the parameters °'11 - Ψχ+ 1 - T2a2— l — Oi can be obtained by processing the diffusion information of sectors 529, 549, for example, as received from the side information 503. In other examples, ai and / or a2 can be recovered directly from the side information 503. Notably, in the case of s=2 sectors, it is possible to encode only ai (respectively a2), so that the audio signal representation decoding unit 500 recovers a2 (respectively ai) by subtracting 1 (or B). Therefore, in some cases, there may be provision, from the secondary information, of parameters of lower sectoral diffusion, although, nevertheless, they Petition 870250071890, dated 08 / 14 / 2025, pages 258 / 295 38 / 74 provide a description of sectoral diffusion (and sectoral directionality) for all spatial sectors. For example, given that the sum of the N relative directionalities aj is 1 (or another B>0), it is possible that n-1 relative directionalities are simply encoded, so as to obtain the Nth relative directionality as 1-(ai + a2 + ... + aN-i).

[0120] The coefficients ai, a2 (usually also indicated as as) can be applied to the spherical harmonic evaluated along the DoA sector (that is, the local DoA inside a specific spatial sector). We can see that the directional signal of the 532, 552 sector can be—(1—Ψ) * * ^s·

[0121] In this expression, 1 _ ψ is applied in the divider 504, Xs in the spatial filtering stage 574 and in the signal processor of the stage sector 572. However, different treatments can be performed.

[0122] Nevertheless, it is important to note that the directional sector signal 532, 552 can be seen as having the following terms (at least in some examples): - the spherical harmonic evaluated along the DoA of the space sector (or other information that allows applying the DoA to the transport channel); - global directionality1 - ψ (or other global diffusion information that allows polishing the global diffusion component of the transport channel); - the coefficient (or other sector diffusion parameter). In this specific case, fljJ can be seen as the relative sectoral directionality of the current sector s over the totality of sectors n.

[0123] Basically, in the decoding paths of sectors 522, 542, etc., the directional signals of sectors 532, 552 can be weighted according to their relative sector directionality after which all of them (520, 522, 542) were weighted by the global directionality1-ψ. In this way, they also take into account the sectoral diffusion of each of them.

[0124] Other types of coefficients, also different from ai and a2, can be used. For example, diffusion parameters can directly indicate Ψι or 1 Petition 870250071890, dated 14 / 08 / 2025, page 259 / 295 39 / 74 Ψι or Ψ2 or 1-Ψ2, for example.

[0125] In the case of more than two sectors, as(with s>2) can be used (in some cases, the condition = 1 or Σ5α5 = B > 0 p0C|e is given).

[0126] The coefficients ai and a2 can be applied to the spatially filtered signals 528, 548 and so on. An example of applying the coefficients could be, for example, = a^ * for cac|sector, thus obtaining the directional signals of the sector 532, 552 and so on.

[0127] Thanks to the application of the coefficients as (applied to different spatial sectors s=1, 2...), it is possible to provide different diffusions for different sectors.

[0128] The coefficient ai, for example, is high where the directionality of the signal 522 in a first spatial sector s=1 is high relative to a second spatial sector s=2. Therefore, if the signal in sector s=1 (path 521) is very directional and the signal in sector s=2 is locally very diffuse, then there will tend to be ai > a2, while if the signal in sector s=1 (path 521) is very diffuse and the signal in sector s=2 is very directional, then there will tend to be ai < a2.

[0129] The as coefficient can therefore be considered as providing diffusion information (and therefore a sector diffusion parameter), because it provides information about sector diffusion (in the specific spatial sector). However, as is a sectoral directionality of sector s over the sum of the sectoral directionalities of all spatial sectors (i.e., a relative directionality). The most directional sectors will have the highest as (e.g., close to 1, particularly if the other sectors are extremely diffuse), and the most locally diffuse sectors will have the lowest as (e.g., close to 0, particularly if the other sectors are extremely directional).

[0130] It can be understood that the coefficient °* will weight the intensity of the 528, 548 audio signal along the DoAJ (at least, in relation to the other directions of the same spatial sector), and the weight will tend to be high if the signal is highly directional in the s spatial sector (and tend to be low if the signal is highly diffuse locally in the s spatial sector), and the weight will tend to be higher than other Petition 870250071890, dated 08 / 14 / 2025, pages 260 / 295 40 / 74 sector S2 if the signal is more directional in spatial sector s than in the other sector S2 (and tending to be smaller than in the other sector S2 if the signal is more diffuse in spatial sector s than in the other sector S2). (Note that the transport channel conversion can be obtained from the relation = 1W * * ^11—1 )

[0131] In general terms, the as coefficient can be an example of a mixing weight derived from the sector diffusion parameter (it can be, for example, the sector diffusion parameter itself). The larger the as, the greater the mixing weight applied to the specific path.

[0132] Basically, it is possible to overcome the problems and issues of the state of the art that go beyond the shortcomings of the standard ambisonic model. Therefore, multiple directional sources in the same time-frequency block and specular reflections must be taken into account.

[0133] It should be noted that the directional signals 522, 542 (are, when in their 532, 552 version, HOA signals. The vector of spherical harmonic functions can be trivially evaluated in arbitrary ambisonic orders. Thus, it allows reconstructing the signal in the originally recorded order or artificially extending it to a higher order to create a better listening experience.

[0134] The audio signal representation decoding unit 500 may include a global diffusion signal inserter 560. The global diffusion signal inserter 560 may combine the plurality of decoded sector signals (532, 552, etc.) with the global diffusion signal 510, so as to insert the global diffusion (510) into the sector signals 532, 552, etc. The output of the global diffusion signal inserter 560 may therefore be the compressed ambisonic spatial audio signal representation 562.

[0135] The signal 559, here conceived as a juxtaposition of these directional signals 532, 552 and so forth, and the global diffusion signal 510 emitted by the power compensator block 508 is therefore indicated with the reference numeral 559.

[0136] In summary, the audio signal representation decoding unit Petition 870250071890, dated 08 / 14 / 2025, pages 261 / 295 41 / 74 500 can generate an uncompressed ambisonic spatial audio signal representation 562 from a compressed ambisonic spatial audio signal representation 502 representing an audio signal, using, as side information 503: Information about the sound field, including, for each specific spatial sector: a. directional parameter(s) (e.g., 529, 549, which provide information about a direction of arrival, DoA, in the specific spatial sector; b. sector diffusion parameter(s) (e.g., 529, 549, ψι, ai, a2, etc.) that provide information about the sector diffusion of the audio signal in the specific spatial sector; a global diffusion parameter (507, 507', 509', Ψ) or other information about the global diffusion of the audio signal (which may or may not be part of the side information 503 and / or may or may not be part of the sound field parameters; or which may be estimated, for example, by the global diffusion estimator 570).

[0137] These parameters can be easily applied to at least one transport channel 501 (in the FOA version) of the audio signal representation 502, to obtain the uncompressed HOA version 562 of the audio signal. In different sector decoding paths (521, 541, etc.), spatially filtered versions (528, 548) of the transport channels 501 can be obtained, each spatially filtered version (528, 548) representing the audio signal within the specific spatial sector.After that, each spatially filtered 528, 548 transport channel is continued into an HOA signal using spherical harmonics evaluated in the DoA for each sector, it is possible to apply a mixing weight to each 528, 548 sector signal that (in some examples) weights the 528, 548 sector signals according to the sector directionality of the audio signal in the specific sector (the mixing weight can be a relative directionality of each spatial sector versus the sum of the directionalities of all spatial sectors).

[0138] Figure 8 shows an example of an 800 device including the Petition 870250071890, dated 08 / 14 / 2025, pages 262 / 295 42 / 74 compressed ambisonic spatial audio signal representation 502, which renders an audio signal (as rendered audio signal 814) or transcodes the audio signal (as transcoded signal 816) from the compressed ambisonic spatial audio signal representation 502. Furthermore, the compressed ambisonic spatial audio signal representation 502 can also be obtained from a bitstream 802 (encoded signal). The device 800 may have a bitstream reader (encoded signal reader) and a dequantizer 804, which can read the bitstream 802 (encoding the compressed ambisonic spatial audio signal representation 502 or 502b) and provide the compressed ambisonic spatial audio signal representation 502 or 502b to the audio signal representation decoding unit 500 (or 500b).The uncompressed ambisonic spatial audio signal representation 562 can be output by the audio signal representation decoding unit 500 or 500b to a renderer 812, to render the audio signal 800 into an audio signal 814 (which should generally be the best possible reproduction of the original audio signal 702) or an encoding unit 813, which can recode the ambisonic spatial audio signal representation 562 into a different spatial audio signal representation 816. The different compressed ambisonic spatial audio signal representation 816 can also be stored and / or transmitted (sent) to other devices or units. In this way, if the renderer 812 is not used to obtain the signal 814, the encoding unit 813 is used to obtain a second spatial audio signal representation 816, then a transcoder is performed by the device 800.In some examples, neither the 812 renderer nor the 813 encoding unit are present, and the output is simply the uncompressed 562 ambisonic spatial audio signal representation.

[0139] Figure 10 shows an audio signal representation encoding unit (e.g., encoder) 700b, which can be used, for example, to provide the bitstream (encoded signal) 802 (generally speaking, only the metadata that must be provided to the audio representation decoding unit 500 Petition 870250071890, dated 08 / 14 / 2025, pages 263 / 295 43 / 74 are the representation of a compressed ambisonic spatial audio signal 502, as well as the plurality of directional parameters and the parameter(s) indicating sector diffusion, which may vary in different spatial sectors. The 700b audio signal representation coding unit can provide the representation of a compressed ambisonic spatial audio signal 502, in this case encapsulated in a bitstream (encoded signal) 802.

[0140] The audio signal representation coding unit 700b of Figure 10 can be inserted with an audio signal (input audio signal representation, which represents an audio signal) 702, which can be, for example, an ambisonic time-domain signal (the audio signal representation coding unit 700b can include a higher-order ambisonic, HOA, time-domain converter, for example, of a non-ambisonic time-domain version, which is not shown in the figure, but is upstream of the HOA signal 702 in Figure 10). Furthermore, the input audio signal representation 702 can be obtained from a purchased version of the microphone(s) or can be synthesized. The input audio signal representation 702 can therefore generally be an uncompressed HOA representation of an audio signal.The 700b audio signal representation coding unit can therefore compress the 702 input audio signal representation into a FOA (or at least a lower-order ambisonic), compressed 502 (802) version of the 702 input audio signal representation, so as to represent the same audio signal in the compressed version. It will be shown that the 502 encoded audio signal representation can include at least one 501 transport channel (in one of its 736 or 739 versions) and 503 side information, such as sound field parameters (e.g., as discussed above and below). In particular, at least one 501 transport channel can represent a downmixed version of the 702 HOA signal (e.g., at least one 501 transport channel, e.g., in its 736 version, can have a selected number of channels relative to the 702 HOA signal).

[0141] The ambisonic audio signal representation coding unit Petition 870250071890, dated 08 / 14 / 2025, pp. 264 / 295 A high-order 44 / 74 (HOA) signal 700b can be fed to an analysis filter bank 704 to obtain a HOA signal version 706 of the input audio signal representation 702 in the filter bank domain, i.e., in the time-frequency domain (so that the audio signal is subdivided into time-frequency blocks). The filter bank domain version 706 of the input audio signal representation 702 can be fed to a spatial filter stage 708. The spatial filtering stage 708 can perform beamforming, for example, by applying beamforming weights to the filter bank domain HOA signal 706. The HOA signal 706 can correspond to the decompressed HOA signal 562 of Figure 5a. The spatial filtering stage 708 can be instantiated by the spatial filters 7071, 7072, ..., 707n, one for each spatial sector (for example, if there are two spatial sectors, there will be two filters, i.e., N=2; in general, N>1). Each spatial filter 7071, 7072, ...707n can cut the audio signal 706 into spatial sectors (for example, spatial sectors N=2 can be two hemispheres or other subdivisions of space can be defined). What is obtained in the spatial filtering stage 708 is a spatially filtered uncompressed ambisonic signal 710, formed by several sectoral directional signals 7101, 7102, ..., 710n (one for each spatial sector s of the spatial sectors N>1). The directional signals of sector 7101, 7102, ..., 710n can correspond to the directional signals of sector 532, 552 of Figure 5a, while the spatially filtered ambisonic signals 710 can correspond to signals 528, 548 of Figure 5a.

[0142] The spatially filtered version of the uncompressed HOA signal 710 can be fed to a sector parameter estimator stage 712, which may include a plurality of sector parameter estimators 7211, 7212, 721n, each of which is configured to derive the sound field parameters 7141, 7142, ..., 714n respectively from signals 7071, 7072, ..., 707n. Basically, each value 7141, 7142, 714n may include some parameters such as directional DoA information (e.g., Ωι, Ω2, ..., Ων) for each specific spatial sector 1, 2, ..., N, and local (sector) diffusion information (e.g., in terms of sector diffusion Ψ1, Petition 870250071890, dated 08 / 14 / 2025, pages 265 / 295 45 / 74 Ψ2, ..., Ψν and / or, complementarily, in terms of local directivity 1 -Ψι, 1-Ψ2, ..., 1-Ψν) for each specific spatial sector 1, 2, ..., N (in some cases, not all sector diffusion information is calculated; for example, from all N spatial sectors, it is possible to calculate the diffusion parameter for each of the N-1 spatial sectors).

[0143] In parallel, a global diffusion estimator 7129 (in the global diffusion path 709) can be provided to supply the global diffusion parameter (e.g., global diffusion information), here indicated by 7149. The parameters 714 (7141, 7142, ..., 714n), i.e., directional parameters and / or diffusion parameters, can then be encoded in the bitstream (encoded signal) 802 (502), directly or in a processed version, as sound field parameters encoded in the side information 503. A parameter converter unit 716 (if present) can supply the directional parameters and the diffusion parameters 718 in processed form. If the parameter converter is provided, then the coefficients a1, a2... etc. can be derived, for example, 1-Ψ! ι-ψ3, £2 —------------ £2-2 — ----------- — 1 — £2 using the formulas ι-Ψι+ι-ψ, and / or ι-ψ1+ι-ψϊ as above (as explained above, in the case of N=2 spatial sectors, the encoding of ai or a2 can be ignored).Parameters 714 (7141, 7142, ..., 714n, 7129) and / or 718 can therefore be processed to obtain secondary information 503 (529, 549, 507).

[0144] The parameter converter unit 716 can therefore convert the sector diffusion parameter(s) from a first representation 714 associating, to each specific sector component, information indicative of the sector diffusion (ψι ,ψ2), to a second representation 718 associating, to the specific sector signal, information (ai, a2) indicative of a relative directionality of the sector signal in relation to the directionalities of the total sector signals. (The parameter converter unit 716 may not be necessary in some cases, for example, when the sector diffusion is written directly to the bit stream 802).

[0145] A 720 parameter quantizer can quantize 718 parameters. The Petition 870250071890, dated 08 / 14 / 2025, pages 266 / 295 46 / 74 quantized parameter 724 can be provided to a parameter encoder 740, which can encode parameters 718 (e.g., in the quantized version 724) in the bitstream 802, as secondary information 503. Therefore, the secondary information 503 can be present in the sound field parameters 729, 549, such as at least some of Ψι, Ψ2, 1-Ψι, 1-Ψ2, Ωι, Ω2, in some cases also the global diffusion parameter Ψ or other information about the global diffusion of the audio signal). To save on metadata bitrate, quantization can be automatically reduced to coarser steps for sectors where diffusion is high.

[0146] The representation of the input audio signal 702 can also be provided to an analysis filter bank unit 704a. A filtered version 729 of the input signal 702 in the filter bank domain (e.g., time-frequency blocks) can therefore be output by the analysis filter bank unit 704a. The version 729 of the representation of the input audio signal 702 can be the same as the version 706 output by the analysis filter bank 704, but in other cases it may be different.

[0147] The 700b encoder (audio signal representation coding unit) of Figure 10 may also include a down-mixing stage 1700b to reduce the 702 audio signal into a compressed (down-mixed) version 736. The down-mixing stage 1700b may be instantiated, for example, by a channel selector. By virtue of the representation of the 702 input audio signal (HOA signal) including a plurality of channels, the channel selector 1700b may simply select the channels corresponding to the FOA version (or at least a lower-order version) of the 702 HOA signal (the selected channels may be, for example, in a plurality of channels; for example, there may be four channels, for example, in the case of FOA, or more channels). This selection operation, of a trivial nature, allows compressing the 702 audio signal, which will therefore require fewer bits.However, most of the audio information will not be lost, as it will be reconstructed by the audio signal representation decoding unit (e.g., 500) using the sidebar information 503. For example, the channel(s) of... Petition 870250071890, dated 08 / 14 / 2025, pages 267 / 295 47 / 74 transport (down-mix, compressed) 736 may include four channels, for example, fewer than the HOA signal 702. The transport channels 736, together with the side information 503, may constitute the representation of the compressed ambisonic spatial audio signal. Notably, Figure 10 shows that an EVS (enhanced voice signal) encoder 738 may be present, to convert the transport channels 736 into an encoded version 739 of the transport channels 736.

[0148] Figure 7 shows another example of an audio signal representation coding unit 700. Here, components 704, 708, 704, 712, 716 and 720 (or at least some of them) may be basically the same as example 700b in Figure 10. However, an analysis filter bank 704a (which may or may not be the same as analysis filter bank 704) may provide a filter bank domain version 729 of the input audio signal representation 702.The 729 version of the filter bank domain can be downloaded into a 1700a drop-down mixing stage. The 1700a drop-down mixing stage may include a 730 drop-down mixing unit (e.g., mix down or drop-down mix) to obtain a 732 drop-down mixing version (in the filter bank domain) of the 702 HOA signal. The 732 drop-down mixing version may have a single transport channel, or multiple transport channels, depending on the specific drop-down mixing performed. The 732 drop-down mixing version of the 702 HOA signal may be submitted to a 734 synthesis filter bank to obtain 501 transport channels (e.g., drop-down mixing, compressed version 736 of the 702 HOA signal in the time domain).The synthesized version 736 (which may or may not be an FOA signal, but in any case in fewer channels than the original signal 702) of the downmixed version 732 (compressed transport channel(s)) of the audio signal 702 can then be provided, as version 736, to one or more instances of an EVS encoder 738 or any other mono audio encoder. The transport channel(s) 501 (736) can, for example, be stored and / or transmitted, together with the side information 503, to form a compressed version of the spatial audio signal representation. Petition 870250071890, dated 08 / 14 / 2025, pages 268 / 295 48 / 74 ambisonic 502. Notably, however, the representation of compressed ambisonic spatial audio signal 502 may include, in addition to the side information 503, any of the compressed (down-mixed) versions 732, 736, 739.

[0149] The filtered 729 input signal can be mixed, and mixing information or covariance information can be provided to the audio signal representation decoding unit (e.g., on at least one transport channel), but this is not the case in all examples. In some cases, which will be discussed below, the audio signal representation decoding unit (e.g., 500b, see below) can reconstruct the mixing information even if the covariance information is not recorded in the 802 bitstream (encoded signal).

[0150] Version 729 of the filter bank domain of the input audio signal representation 702 can be downmixed in a downmix unit 730 to obtain a downmix version 732. The downmix version 732 can be submitted to a synthesis filter bank 734. The synthesized version 736 of the downmix version 732 of the audio signal 702 can then be fed to the EVS encoder 738. Then, the compressed transport channels 736 can be encoded by an instance of the EVS encoder 738 or any other mono audio encoder in block 738 to obtain at least one encoded transport channel 739 (encoded version of the compressed transport channel 736).The compressed transport channel(s) 739 may, for example, be stored and / or transmitted, together with the side information 503, to form a compressed version 502 (in the encoded signal or bitstream 802) of the ambisonic spatial audio signal representation 502.

[0151] It should be noted that the compressed transport channel 736, 501 may or may not be an instantiation of at least one transport channel 501 of the compressed ambisonic spatial audio signal representation 502 (which may be decompressed, for example, in Figure 5a). If at least one transport channel 736 is compressed, it must be decompressed (augmented) by a Petition 870250071890, dated 08 / 14 / 2025, pages 269 / 295 49 / 74 up-mixing unit (550, see below in Figure 5b) to recover the transport channel(s) of the compressed ambisonic spatial audio signal representation 502. The fact that it is further compressed increases efficiency.

[0152] A descending mix matrix calculator 726 can provide a descending mix matrix 728 for the descending mix unit 730. The descending mix matrix calculator 726 can make use of covariance information (e.g., covariance matrix) detailed below to perform a cross-channel prediction. The descending mix matrix calculator 726 could be more generally called a descending mix information calculator, but for simplicity, the descending mix matrix calculator will be preferred.

[0153] Version 732, 736 or 739 of the ambisonic spatial audio signal representation 702 may be supplied to a bitstream recorder (multiplexer, encoded signal encoder) (signal encoder) 750 to supply the compressed ambisonic spatial audio signal representation 502 (bitstream, encoded signal) to an external device (e.g., by transmission) or a storage unit.

[0154] The 730 downmixing unit can apply sound field parameters (sector diffusion parameters, directional parameter(s) (529, 549, 718) providing information about an arrival direction, DoA, for each parameter, etc.) to perform downmixing. For this purpose, the down-mixing unit 730 can make use of the down-mixing matrix 728, which can be output by a down-mixing matrix calculator 726. The down-mixing matrix calculator 726 can obtain the down-mixing matrix 728 from a covariance matrix (or, more generally, covariance information) which is, in turn, estimated from the sound field parameters 718 or their quantized versions 722. As can be seen, in fact, the down-mixing matrix calculator 726 is shown as being entered with an input 722 including the sound field parameters 718 (by Petition 870250071890, dated 08 / 14 / 2025, pages 270 / 295 50 / 74 example, in quantized form, but in other cases they may be in non-quantized form, for example, 718 or 714, for example, including 7141, 7142, ..., 714n). The descending mixing matrix calculator 726 may or may not perform the same operations as a mixing matrix estimator 100 discussed below on the decoder side (and, in particular, the covariance matrix synthesizer 102 and a mixing matrix constructor 106 of Figure 5b). The descending mixing unit 730 can, in principle, be considered as corresponding to the ascending mixing block 110 of Figure 5b (but in the two cases, the matrices do not correspond to each other).

[0155] The operations on the 726 descending mixing matrix calculator are now discussed. The descending mixing matrix (or more general descending mixing information) can be obtained from the covariance matrix (or more general covariance information). Therefore, first, how to obtain the covariance matrix (or covariance information) is discussed. Initially, the covariance matrix between C channels can be defined. The covariance matrix between C channels can be a square matrix (L+1 )2x (L+1 )2 (i.e., with (L+1 )2 rows and (L+1 )2 columns), given that L is the order of the ambisonic signal in version 501 to be inserted into decoder 500 (or version 501c to be inserted into portion 500 of decoder 500b, see below in relation to Figure 5b) (for example, for an FOA signal, L=1 and there will be (L+1 )2=4 rows and (L+1 )2=4 columns).Each of the (L+1)2 columns and each of the (L+1)2 rows of the covariance matrix corresponds to one of the (L+1)2 ambisonic channels according to a predefined order, so that an entry provides the covariance between the ambisonic channel corresponding to the row and the ambisonic channel corresponding to the column. Here, the generic matrix element will be indicated with . With, for example, an FOA signal, this would be I equal to 0 or 1, em=0 for m=0, em=-1, 0, +1 for 1=1, which leads to four combinations and a 4x4 covariance matrix. The covariance matrix can be a symmetric matrix. Each non-diagonal element of the covariance matrix provides covariance information between the two ambisonic channels. In general terms, the more correlated two ambisonic channels are, the greater a will be. Petition 870250071890, dated 08 / 14 / 2025, pages 271 / 295 51 / 74 covariance in the corresponding matrix elements, whereas the more uncorrelated two ambisonic channels are, the smaller the covariance in the corresponding matrix elements. For a non-diagonal matrix element, the element between the ambisonic channel of degree and index I in, respectively, and that with degree and index m' and respectively, can be (1 - Ψ) « Ex« ax2« + (i — ψ) * * a2* rlm(n2) * where is the signal energy, e2 are the first and second directional parameters (e.g., sector DoAs), respectively, and “ai” ea2_ 1-ψ±_ lHj _ é2j_ — ------------- £2-2 — ------------- — 1 — £2j_ can be the coefficient ι-ψ±+ι-ψ3 e as discussed above, for example, indicating the relative directivity of the 702 audio signal in the spatial sector relative to the signal directivity for all sectors, ψ is the global diffusion parameter, and ee are the spherical harmonics evaluated in the DoAs for each sector and for each ambisonic channel.This is valid in the case of two space sectors.

[0156] Since there is no covariance in the elements of the diagonal matrix (I=Γ, m=m'), the generic element of the diagonal matrix can be written as (1 - Ψ) *ar2*ϊ^(Πι) + (1 - Ψ) * Ex *a22*^(Π2) +Ψ * ο2* Ex with the same meanings as the symbols, eυe a predetermined energy scaling factor.

[0157] A more compact representation is (1 - Ψ) . Ex* aL2« + (i - ψ)« ex«a22* y!m(íi2) * riIm,(fi2) + +Ψ » σ2* Ex* Petition 870250071890, dated 08 / 14 / 2025, pages 272 / 295 52 / 74 where is the Kronecker delta being 1 on the diagonal of the covariance matrix between channels and 0 off the diagonal of the covariance matrix between channels.

[0158] The covariance matrix may include: 1) In non-diagonal elements for spatial caqasetor, a component obtained by a product between: a. a *E of non-global diffusion energy of the audio signal b. the spherical harmonic, evaluated at a first DoA° <e dimensionado pela diretividade relativa do setor espacial sobre a soma das diretividades dos outros setores espaciais c. the spherical harmonic, evaluated at the second DoA and dimensioned by the relative directivity of the space sector over the sum of the directivities of the other space sectors 2) in diagonal elements a. for each spatial sector, a component obtained by a product between: ia(1 - Ψ) * Ej, qe energjaç|edifusao não global do sinal de áudio II. a directivity energy aia* (respectively aa2* for each sector) W Φ p b. a global component scaling the global diffuse energy by a predefined scaling factor °.

[0159] The covariance matrix between channels C (with the elements can therefore be estimated from the sound field parameters 722, including nww paratQdos os s, where s=1...neo sector index, in the top-down mixing calculator 726.

[0160] This inter-channel covariance matrix, in turn, allows deriving the descending mix matrix and the ascending mix matrix in the audio signal representation encoding unit and the audio signal representation decoding unit, respectively. Therefore, a covariance matrix encoding step in the 502 bitstream can be advantageously skipped. Specifically, the 726 descending mix matrix calculator can be Petition 870250071890, dated 08 / 14 / 2025, pages 273 / 295 53 / 74 configured to perform an inter-channel prediction (between ambisonic channels). This prediction is based on the inter-channel covariance matrix 732 (the inter-channel covariance matrix 732 being derived from the directional parameters and sector diffusion parameter(s) for each spatial sector and a global diffusion or one or more parameters indicative of a ratio or relationship between the diffusion or diffuse energy of one sector to the diffusion or diffuse energy of all sectors) and can achieve an energy compression between audio channels 729.

[0161] In the audio signal representation coding unit, the descending mixing matrix can, for example, with an FOA input signal, be derived from the covariance matrix, for example, via the formula / 1 0 0 0\ _ í ~ Cl-iOo / 1 0 0 \ I—Çlxoo / Qw>o 0 1 0 I \ ~ Cii,oo / Coo,oo 0 0 1 /

[0162] Additional terms may or may not be included in the matrix to model uncorrelated signals using decorrelators in the decoder.

[0163] Reference is now made to the audio signal representation decoding unit 500b (Figure 5b), which may comprise a portion 500 (the portion 500 which is identical to the audio signal representation decoding unit 500 of Figure 5a and is therefore indicated with the same numeral 500). In Figure 5b, the transport channel(s) 501 is / are converted from the compressed and mixed version 501b to a version 501c (but still compressed, at least in the sense of being an FOA version) and is provided to portion 500, thus providing the transport channels 501. Therefore, in portion 500, the audio signal representation decoding unit 500b of Figure 5b operates identically to the audio signal representation decoding unit 500 of Figure 5a, to provide the uncompressed ambisonic spatial audio signal representation 562.

[0164] In the audio signal representation decoding unit 500b (Figure 5b), the mixing matrix reconstructor 106 (see below) calculates the ascending mixing matrix (mixing matrix) 108 as the inverse of the matrix of Petition 870250071890, dated 08 / 14 / 2025, pages 274 / 295 54 / 74 descending mix, which can, in turn, be calculated from the channel covariance matrix 104. This can also be done by using a predefined formula derived to provide the inverse of the descending mix matrix 728.

[0165] More generally, covariance information can be based on an energy weighted by the spherical harmonics evaluated in the DoAs (e.g., and mixing weights .. °* for each spatial sector.

[0166] More generally, instead of the covariance matrix, covariance information can be used, for example, in conjunction with global diffusion information (e.g., global diffusion parameter Ψ or other information about the global diffusion of the sound signal).

[0167] Figure 5b shows an example of an audio signal representation decoding unit 500b generating an uncompressed ambisonic spatial audio signal representation 562, for example, from the mix transport channel(s) 501 into the compressed ambisonic spatial audio signal representation 502 inserted into the bitstream 802 (encoded signal), for example, by the audio signal representation decoding unit 700 of Figure 7. Here, at least one down-mix transport channel 501 of the compressed ambisonic spatial audio signal representation 502 may be a compressed, down-mixed FOA version of the input audio signal 702.

[0168] Figure 5b shows an example of the audio signal representation decoding unit 500b, which includes, along with block 500 of Figure 5a, also an upmixing unit 550. The upmixing unit 550 can receive, from the bitstream 802 (encoded signal), the side information 503 (for example, at least some of Ψi, Ψ2, ai, a2, Ωi, Ω2 ...) and at least one transport channel 501b. At least one transport channel 501b can be an example of at least one transport channel 501 corresponding to the transport channel(s) 501 upstream of the EVS 738 encoder. However, in this case, at least one transport channel 501b (501) is mixed so as to present Petition 870250071890, dated 08 / 14 / 2025, pages 275 / 295 55 / 74 a greater number of transport channels. At least one 501b (501) transport channel can be received from the 802 bitstream as a 739 encoded transport channel and can be decoded by an EVS 738b decoder (if the 700 encoder does not have the EVS 738 encoder, the EVS 738b decoder can be avoided). The mixed transport channels are therefore indicated with 501c, which is also an example of 501 transport channels. In this case, however, the upmix is ​​performed from the 503 side information (a directional parameter, sector diffusion parameters and global diffusion), which have already been discussed above.

[0169] As can be seen in Figure 5b, a covariance matrix synthesizer 102 can receive side information 503, including directional parameters Ωι, Ω2, ..., second diffusion parameters Ψι, Ψ2 (or in the form of relative directionalities ai, a2, etc.) ..., and global diffusion information for other information, which allows deriving global diffusion. Here, the covariance matrix synthesizer 102 can therefore obtain an inter-channel covariance matrix 104 (see above). The inter-channel covariance matrix 104 can include information about the covariance between the different ambisonic channels. Alternatively, covariance information can be derived. The covariance matrix can be estimated as in the encoder 700 and is therefore not repeated here.

[0170] Basically, the channel covariance matrix 104 can be obtained as the covariance matrix Gm / m_{1-ψ).^*^.^{Ω1).^(ί11)+(1-ψ). ^*3ζ2*^(Ω2)*^(η2) + UJ # fr 2 $ J7 ΐ / i' í í m; as described above.

[0171] The channel covariance matrix 104 can then be provided to a mixing matrix reconstructor 106, which reconstructs the mixing matrix 108 according to the number of transport channels that must be in the 501c version of the transport channel 501b. Once the mixing matrix 108 is obtained by the mixing matrix reconstructor 106, an upmixing block 110 can use the mixing matrix 108 and apply it to the transport channel(s) 501b, to convert them into a 501c transport channel version (501) (e.g., in Petition 870250071890, dated 08 / 14 / 2025, pages 276 / 295 56 / 74 multiple channels, for example, represented as an FOA signal, for example, with four channels) to be supplied to block 500 of Figure 5a. The sound field parameters 503 are also supplied to block 500.

[0172] It should be noted that the 500b technique can be ignored in cases where the inter-channel covariance matrix or the mixing matrix is ​​written in the 802 bitstream or is obtained in another way. The 102 covariance matrix synthesizer and the 106 mixing matrix reconstructor together form a 100 mixing matrix estimator. Notably, other techniques can be used to obtain the mixing matrix.

[0173] The above operations can be performed band by band. Reference is made to Figure 5c (showing a variant 500b' of Figure 5b). The audio signal representation decoding unit 500b' may comprise a band combiner 570. The band combiner 570 may combine the 104 covariance-related information (from the 802 bitstream and / or the 102 covariance matrix synthesizer) so that some of the 104a covariance information comes from the 802 bitstream for some bands, while other 104 covariance information comes from the 102 covariance matrix synthesizer for the other bands. Note, however, that some parameters (e.g., sound field parameters such as DoAs, and / or diffusion parameter(s)) may be the same for groups of bands.Furthermore, it is understood that the 104 covariance information in the 802 bitstream may contain the elements of the covariance matrix directly or any other representation derived from them. For example, prediction coefficients or decorrelator channel weights may be such representations. In general, different representations may be mixed.

[0174] The 108 mixing matrix can be reconstructed for each band (e.g., in example 500b of Figure 5b). However, in some examples (e.g., in variant 500b' of Figure 5c), for some bands, the entries of the 108 mixing matrix, or other parameters that encode covariance or prediction information, may be encoded in side information 503, while for others Petition 870250071890, dated 08 / 14 / 2025, pages 277 / 295 57174 bands they are ignored. This is done (e.g., in variant 500b' of Figure 5c) in the band combiner unit 595, which provides the inputs of the mixing matrix 108 as 104a. Here, it may be that the sector directional parameter(s) and the sector diffusion parameter(s) (e.g., relative directionalities) are used to recover the mixing matrix 108 via the covariance matrix only for some bands (e.g., high-frequency bands), while for other bands (e.g., lower-frequency bands), the inputs of the mixing matrix 108 (or the mixing information) can be written to the side information parameters 503. For example, sound field models (i.e., to recover the covariances of the sound field parameters) can be employed only in high-frequency bands, where the perceptual impact of inaccuracies is smaller.

[0175] It should also be noted that in some examples it is possible (for example, in the audio signal representation decoding unit) to switch between: a low-order mode of operation, in which, among the plurality of sector decoding paths (521, 541), at least one of the sector decoding paths (521, 541) is deactivated, while only one of the sector decoding paths (521, 541) is activated, in which the side information (503) does not contain the sound field parameter(s) (549) for the deactivated at least one of the sector decoding paths (521, 541); and a high-order mode of operation, in which, among the plurality of sector decoding paths (521, 541), all sector decoding paths of the plurality (521, 541) are activated, or at least fewer sector decoding paths are deactivated in relation to the low-order mode of operation, in which the side information (503) also contains the sound field parameter(s) (529, 549) for the entire plurality of sector decoding paths (521, 541), as well as the global diffusion parameter (507, 509).

[0176] In some examples, sidebar information 503 may also include the Petition 870250071890, dated 08 / 14 / 2025, pages 278 / 295 58 / 74 global diffusion parameter in low-order operating mode.

[0177] It should also be noted that in some examples it may be possible (for example, in the audio signal representation coding unit, for example, 700 or 700b) to switch between: A low-order mode of operation, in which, among a plurality of sector paths, at least one of the sector paths is deactivated, while only one of the sector paths is activated, so that the side information does not contain the sound field parameter(s) for the deactivated at least one of the sector paths (the global diffusion parameter may also be encoded in the side information); and a high-order mode of operation, in which, among the plurality of sector paths, all sector paths of the plurality are activated, or at least, fewer sector paths are deactivated compared to the low-order mode of operation, so that the side information also contains the sound field parameter(s) for the entire plurality of activated sector paths, as well as the global diffusion parameter.

[0178] (In Figure 7 or 10, a first sector path may include a series formed by blocks 7071, 7121, providing the parameter set for sector 1, 7141; a second sector may include a series formed by blocks 7072, 7122, providing the parameter set for sector 2, 7142; an nth sector path may include a series formed by blocks 707n, 712n, providing the parameter set for sector n, 714n in the side information).

[0179] For example, in the audio signal representation coding unit (e.g., 700 or 700b), in low-order operating mode, it may be that only the sector 1 parameter set (7141) is provided, while in high-order operating mode the second parameter sets 1 and 2 (and perhaps also n) are provided in the side information, and also the global diffusion parameter 7149 (507) may be provided in the side information (e.g., the global diffusion parameter may be provided in the side information in both Petition 870250071890, dated 08 / 14 / 2025, pages 279 / 295 59 / 74 high-order operating mode as well as low-order operating mode).

[0180] In some examples, the choice between low-order and high-order operating modes can be made by the encoding unit of the audio signal representation (e.g., 700 or 700b) and signaled in sidebar 503 of bitstream 802, and the decoding unit of the audio signal representation (e.g., 500, 500b, 500b'), after retrieving the signaling in sidebar 503, will also change according to the low-order operating mode or high-order operating mode under the control of the (signaling).

[0181] Optionally, the selection between low-order and high-order operating modes can be static or bit-rate dependent.

[0182] In some examples, the switching (for example, selecting between low-order operating mode and high-order operating mode) is only for some bands, while in some other examples the switching is for all bands.

[0183] Consequently, a satisfactory trade-off can be achieved between maintaining low bandwidth overhead (reducing 503 side information) and quality for the more important bands.

[0184] Optionally, the selection between low-order and high-order operating modes can be static or bit-rate dependent.

[0185] In a mobile communications scenario, for example, the available bit rate may depend on the quality of the network connection, which may vary over time. Thus, according to examples, the audio signal representation encoding unit (e.g., 700, 700b) and / or the audio signal representation decoding unit (e.g., 500, 500b, 500b') may dynamically switch between different bit rates. The audio signal representation encoding unit and / or the audio signal representation decoding unit may be configured to select, at high bit rates (e.g., at bit rates above a bit rate threshold), Petition 870250071890, dated 08 / 14 / 2025, pages 280 / 295 60 / 74 predetermined), the high-order operating mode and / or to select, at low bit rates (e.g., the predetermined bit rate limit mentioned above), the low-order operating mode (low bit rates are lower than high bit rates). Quality can be measured, for example, from measurements related to network connection quality. (For example, quality can be measured through latency measurements, so that the higher the average latency of one or more messages, the lower the quality; and the lower the average latency of one or more messages, the higher the quality; in this case, the predetermined limit related to quality is a latency limit, so that higher bit rates are chosen for lower average latencies, and lower bit rates are chosen for higher average latencies.)Quality can be measured through error rate measurements, for example, based on checking the CRC field of messages, so that the higher the number of incorrect messages received, the lower the quality, and the lower the number of incorrect messages received, the higher the quality; in this case, the predetermined limit related to quality can be an incorrect message limit (error rate limit), such as the highest bit rate is chosen for a number of incorrect messages that is below the incorrect message limit, and the lowest bit rate is chosen for a number of incorrect messages that is above the incorrect message limit. Or quality can be measured through connection bandwidth measurements, for example, based on the average bandwidth of the connection.In this case, the predetermined limit related to quality can be a bandwidth limit, so that higher bit rates are chosen for higher average bandwidth and lower bit rates are chosen for lower average bandwidth. Measurements related to network connection quality can be obtained, for example, by the cooperation of the audio signal representation encoding unit (e.g., transmitter) with the audio signal representation decoding unit (e.g., receiver). For example, the error rate can be... Petition 870250071890, dated 08 / 14 / 2025, pages 281 / 295 61 / 74 measured by the audio signal representation decoding unit (receiver) and its value can be encoded and provided as feedback to the audio signal representation encoding unit. Furthermore, latencies can be measured by the audio signal representation encoding unit (e.g., transmitter) upon receiving a response to a specific pilot signal sent at specific time instants, and by measuring the reception time of a specific response signal sent by the audio signal representation decoding unit (e.g., receiver) in response to receiving the specific pilot signal. By subtracting the reception time of the specific response signal from the transmission time of the specific pilot signal, the latency can be calculated.Another way, for the audio signal representation encoding unit (e.g., transmitter), to obtain latencies could be, for example, to read a timestamp in a message from the audio signal representation decoding unit (e.g., receiver) in order to determine the latency of that message. Or, connection bandwidth measurements could be performed. Other quality-related measurements could be taken. Therefore, the selection between high-order and low-order operating modes can be based on network connection quality measurements.

[0186] In examples, in the audio signal representation coding unit (e.g., 700, 700b), the bit rate can be, for example, selected by the user or by a pre-selection (e.g., a default pre-selection) or automatically, depending on the quality of the network connection (e.g., so that the lower the quality, the lower the bit rate; and the higher the quality, the higher the bit rate). The audio signal representation coding unit can then select between high-order and low-order operating modes depending on the bit rate (e.g., a bit rate below a predetermined bit rate limit, indicative of low quality, implying the selection of the low-order operating mode; and a bit rate above the predetermined bit rate limit, indicative of high quality). Petition 870250071890, dated 08 / 14 / 2025, pages 282 / 295 62 / 74 superior to low quality, implying the selection of the high-order operating mode).

[0187] In the audio signal representation coding unit (e.g., 700, 700b), the selection of the operating mode between the low-order operating mode and the high-order operating mode may, for example, depend (partially or totally) on the input audio signal (e.g., totally on the input audio signal or at least on the input audio signal). When a high-order input signal is available, the high-order operating mode can be selected. When only a low-order input signal is present, the audio signal representation coding unit can revert to the low-order operating mode.

[0188] In another example, the audio signal representation coding unit (e.g., 700, 700b) can be configured to select the higher-order operating mode when the battery supplying the audio signal representation coding unit (e.g., the battery of a user device comprising the audio signal representation coding unit) is fully charged (or at least charged above a predetermined charge limit, or battery power limit) and to select the lower-order operating mode when the battery is not fully charged (or at least charged below the predetermined charge limit, or battery power limit), for example, when in a power saving mode.

[0189] The bitrate selected by the audio signal representation encoding unit (e.g., 700, 700b) is detected by the audio signal representation decoding unit. For example, the audio signal representation decoding unit (e.g., 500, 500b, 500b') can select the high-order operating mode when a high bitrate is received (e.g., above a predetermined bitrate threshold), and the audio signal representation decoding unit (e.g., 500, 500b, 500b') can select the low-order operating mode when a lower bitrate is received. Petition 870250071890, dated 08 / 14 / 2025, pages 283 / 295 63 / 74 low (e.g., below the predetermined bit rate limit) is received.

[0190] In another example, the audio signal representation decoding unit (e.g., 500, 500b, 500b') can select the high-order operating mode when it is signaled in the bitstream (e.g., between sidebars 503) that a high-order audio signal has been encoded by the audio signal representation encoder, and the audio signal representation decoding unit can select the low-order operating mode when it is signaled in the bitstream (e.g., at 503) that a low-order audio signal has been encoded by the encoder.

[0191] In some other examples, the audio signal representation decoding unit (e.g., 500, 500b, 500b') may request a high bit rate from a network and select the high-order operating mode, or the audio signal representation decoding unit may request a low bit rate from a network and select the low-order operating mode. The bit rate selection may, for example, depend on a user setting (or pre-selection, such as a default pre-selection) or on the capabilities of user equipment comprising the audio signal representation decoding unit.

[0192] In the examples above, the sound field parameters can be modified to obtain a rotation of the sound field represented by the output ambisonic signal (502). The DoAs can contain the direction from which the sound comes. This is because, if it is necessary to obtain a rotation of the sound field (for example, for head tracking), the audio signal representation decoding unit modifies these parameters and saves the complexity of an extra rotation step. The audio signal representation decoding unit will therefore operate according to the parameter modifications.

[0193] Below are some examples of the evolution of the representation of the 702 audio signal in Figure 7: In the 700 audio signal representation encoding unit: 1) The representation of uncompressed 702 audio signals can be, therefore Petition 870250071890, dated 08 / 14 / 2025, pages 284 / 295 64 / 74 example, a time-domain HOA signal, for example, with more than four channels; 2) In the 1700a down-mix stage: a. In the analysis block of filter bank 704a, the representation of the uncompressed audio signal 702 can be converted to the domain of filter bank version 729; b. In the 732 drop-down mixing unit, the 729 version of the filter bank domain can be converted to the (compressed) 732 version; c) After the synthesis block of the filter bank 734, the downmix (compressed) version 732 can be converted into a time-domain version 736 (notably, the downmix version 732 can have a single transport channel or a plurality of transport channels) 3) In the EVS encoder, the 736 version of the descending mix, compressed in the time domain, can be converted into a 739 encoded version. 4) In the 750 bitstream recorder (e.g., multiplexer), the 802 bitstream is recorded.

[0194] Meanwhile, the HOA signal 702 can be processed to obtain the sound field parameters 718, including the sector directional parameters and the sector diffusion parameters. From the sound field parameters 718, the covariance matrix and the downmixing matrix 728 can be calculated, so as to allow downmixing in 730. In addition, a quantized representation of the sound field parameters is recorded in the bitstream (802) in the bitstream recorder (multiplexer) (750).

[0195] In device 800 of Figure 8 (including audio signal representation decoding unit 500b of Figure 5b or 500b' of Figure 5c): 1) The 802 bitstream is read by the 804 bitstream reader and dequantizer as a representation of a 502 encoded and compressed ambisonic spatial audio signal (502b); 2) From the bit stream 502 (502b), at least one channel is obtained Petition 870250071890, dated 08 / 14 / 2025, pages 285 / 295 65 / 74 transport coded 739 (501); 3) In the EVS 738b decoder, at least one coded transport channel 739 (501) is converted into at least one compressed and mixed transport channel 501b (corresponding to transport channel 736 in Figure 7); 4) In the upmix block 110, through the mix information (e.g., mix matrix) 108, at least one transport channel 501b is mixed to transport channels 501c (501) (e.g., four FOA channels) 5) In portion 500 of the audio signal representation decoding unit 500b (corresponding entirely to the audio signal representation decoding unit 500 of Figure 5a), the four FOA channels 501 are split in the splitter block 504, between the globally diffused FOA signal 506 (in the globally diffused path 505) and the globally non-diffuse FOA signal 520; a. In the global broadcast path 505, the global broadcast signal 506 is subject to the gain provided by the power compensation block 508 to thus obtain the power-compensated global broadcast signal 510; b. In each of the decoding paths of sectors 521, 541, etc., the transport channels first evolve through spatial filtering in the spatial filtering stage 574 and then through the signal processor of sector 572, in order to obtain, for each spatial sector, a directional signal from sector 532; 6) In the global broadcast signal inserter 560, the directional sector signals 532, 552 and the energy-compensated global broadcast signal 510 are added together to derive the uncompressed ambisonic spatial audio signal representation HOA 562; 7) The uncompressed ambisonic spatial audio signal representation HOA 562 can then be recoded as 816, or rendered as 814, or stored or transmitted as is.

[0196] As explained above, the channel covariance matrix 104 and the mixing matrix 108 can be reconstructed from the field parameters. Petition 870250071890, dated 08 / 14 / 2025, pages 286 / 295 66 / 74 sound field in the side information, to perform the upmixing of at least one transport channel 501b into the transport channels 501 (501c) to be fed to portion 500 of the audio signal representation decoding unit 500b. Within portion 500 of the audio signal representation decoding unit 500b, the same sound field parameters are also used to process the transport channels in paths 505, 521 and 541.

[0197] With reference to the example of the audio signal representation encoding unit 700b of Figure 10 and the audio signal representation decoding unit 500 of Figure 5a, it is basically the same, with the difference that the down-mixing stage 1700a is performed by channel selection, no covariance matrix or down-mixing matrix is ​​calculated and / or reconstructed, and the transport channel(s) 736 or 739 (e.g., four transport channels) are provided, as transport channel(s) 501, directly to the audio signal representation decoding unit 500 and, in particular, to the splitter 504.

[0198] In the examples, the audio signal representation encoding unit (e.g., 700, 070b, etc.) may be a transmitter or integrated into a transmitter (e.g., transmitting via wired or wireless or mixed transmission, e.g., via geographic networks and / or local area networks) and / or the audio signal representation decoding unit (e.g., 500, 500b, 500b', etc.) may be a receiver or integrated into a receiver (e.g., receiving via wired or wireless or mixed transmission, e.g., via geographic networks and / or local area networks). DISCUSSION

[0199] The invention uses a combination of first-order estimators and higher-order sector estimators for higher-order directional audio coding (HO-DirAC).

[0200] In particular, it makes use of a combination of global diffusion and sectoral diffusion, improving the state of the art shown in Figure 2. Petition 870250071890, dated 08 / 14 / 2025, pages 287 / 295 67 / 74

[0201] Global diffusion defines the balance of direct to diffuse flow, noting that ψ can be restored in the decoder (audio signal information decoding unit, e.g., 500, 500b, 500b', etc.). The decoder can receive ψ from the bitstream or calculate it from the transport channels (501, 501c).

[0202] The decoder (audio signal information decoding unit, e.g., 500, 500b, 500b', etc.) in the direct stream can extract the sector signals, for example, by beamforming (spatial / directional filtering) of the transmitted FOA signals according to the encoder sector design. The sector signals are obtained by where the beamforming weight vector for the sector is s, and the FOA signal vector is s. From this, the HOA signals are restored by continuing the plane wave coefficients SH in the direction of the sector DoA x3= y(ü5) * x3= K-Yc,. χL�-ι,χLΛο,^ϊΐ!,...].

[0203] The sectors (spatial sectors) are balanced based on the sector diffusion rate °J, therefore, less diffuse sectors contribute more to the directional flow. The sector ratio can be defined, for example, for two sectors, as •| __\pi Oχ=fOj=1 1I> 1-^1+1-^3 where the sector diffusion is estimated in the encoder.

[0204] The proposed project is flexible in terms of the number of sectors, with the only restrictions being ; Σ^- = 1 |ssodecorre direct da preservação da amplitude sobre os setors detalhes em [Hold2021],

[0205] The restored directional part of the HOAXh^ signal vector for a sector s is xH^ = (1 - Ψ) * a3* * x3.

[0206] The diffuse part is represented as a FOA signal vector, amplified by a gain factor dependent on the total diffusion and the in-order1 and the out-order, detailed in the document [WO 2020 / 115311 A1], fI + 1 \ ι+5(ψ)=

[0207] The sum of all HOA signal vectors and FOA signal vector results in Petition 870250071890, dated 08 / 14 / 2025, pages 288 / 295 68 / 74 HOA signal vector output from the decoder.

[0208] The proposed design with two DoAs and higher-order sector processing according to Figure 5a can show significant improvements in a loudspeaker listening test (CICP19), where the result is shown in Figure 6. In particular, items such as 4, which contains a sound scene with wide spatial distribution of surrounding raindrops with distinct location, benefit significantly, as the spatial impression tends to collapse with the last-generation method, which is improved in the proposed method. Other scenes tend to benefit or not deteriorate with the proposed method.

[0209] Furthermore, the more accurate multi-DoA model of the directional signal allows for a more precise estimate of the interchannel covariance matrix Cx, which can be given by + (1 — Ψ) * * a22* (Ω2) + +Ψ v:2* V

[0210] This will lead to more efficient compression of the transport channels in the current IVAS system.

[0211] In Figure 9a, 900 refers to a sector beam of a spatial sector. Panel 902 shows four spherical harmonic functions belonging to the first channels (FOA) of an ambisonic signal (e.g., 501). Panel 904 shows filtered versions of these functions, where the sector beam from the first panel has been applied. Therefore, it can be understood as the contribution of the respective FOA channel to the filtered signal (528) of this specific sector.

[0212] Figure 9b shows the same, but for two spatial sectors (e.g., s=1 and s=2).

[0213] Figure 9c. shows the signal energy as a function of DoA. The input signal (top panel) 910 is the reference and corresponds to the audio signal 702 (or a version thereof) encoded by the audio signal representation coding unit 700. The proposed method provides an output signal 914 (which may correspond to the decoded representation 562 and / or its rendered version 814) where the energy distribution is more similar to the reference signal 702, in Petition 870250071890, dated 08 / 14 / 2025, pages 289 / 295 69 / 74 comparison with a 912 output signal (e.g., corresponding to representation 262 of Figure 2) according to the state of the art with a single DoA. Specifically, the two independent sources in too many directions are resolved much more clearly, unlike 912, where much energy leaks into the region between the actual sources.

[0214] In the figures, ordinate and abscissa refer to the zenith coordinate and azimuthal coordinate. RMS means root mean square of the signal energy.

[0215] Figure 9d displays the direction and diffusion parameters for four spatial sectors and the entire signal. The comparison between the upper and lower plots demonstrates that the inventive sector-based method can resolve different DoAs and diffusion values ​​for different sectors. In contrast, other methods with only one sector can resolve only one DoA and one diffusion.

[0216] The present disclosure also relates to an audio encoder comprising the audio signal representation encoding unit (for example, of Figures 7 or 10) and a bitstream quantizer and recorder (for example, element 40) that can record the compressed audio signal representation 502 in the bitstream. ASPECTS

[0217] Some aspects are summarized here.

[0218] Compared to the state of the art: - Sector processing, i.e., more than one DoA (more than one diffusion measure) acting on the sector beam signals during reconstruction (delta for document [WO 2020 / 115311 A1]). Key aspects of the new feature: - Combination of total diffusion (e.g., first order) and sector diffusion rate (e.g., higher order) the delta for document [US10313815B2], which is the speaker rendering, not the HOA encoding

[0219] Details: Petition 870250071890, dated 08 / 14 / 2025, pages 290 / 295 70 / 74 A device parameterizing a spatial audio scene from higher-order (HO) spherical harmonic domain (SHD) signals, that is, higher-order Ambisonic (HOA) signals. - 1a Transmission of a subset of the HOA input signals, such as, among others, first-order ambisonics (FOA), and spatial parameterization as a set of metadata. 1b. Reconstructing the untransmitted HOA signal components using the transmitted metadata - 1c Metadata that includes more than one direction of arrival (DoA) - 1d The combination of a general estimate of sound field diffusion, estimated from the first-order SHD, together with more than one measure of spatially localized diffusion in the sound field, estimated from the higher-order SHD. - 1e Psychoacoustic frequency weighting in the average of the spatial parameterization group - 2 An apparatus for reconstructing the HOA signal using more than one DoA and both the general diffusion measure and the spatially localized sound field diffusion measures. - 2nd Re-estimation (parts) of first-order parameters in the decoder - 2b Use parameters estimated from HOA signals to improve reconstruction performance based on parameters estimated from FOA, thus employing first-order and higher-order estimators. - 3 Use sector parameterization, such as, but not limited to, more than one DoA to predict the HOA channel covariance (SPAR). ADDITIONAL SPECIFICATIONS

[0220] Possible metadata (e.g., sound field parameters in side information 503). Petition 870250071890, dated 08 / 14 / 2025, pages 291 / 295 71 / 74

[0221] Possible comparative examples: Previous technique: DoA and Diffusion: lA^J (f) Current technique: 2 DoAs Diffusion (but can be estimated in the 500 decoder) and Sector Diffusion Ratio (or more generally sector diffusion information, such as directionality relative to 1): A' A'fíl (f) Notably, it is possible to make use of the currently used infrastructure and signal encoder (802 bitstream recorder), and this is transparent to the 500 decoder.

[0222] Other important aspects: 1) Multi-DoA rendering using HO sectors (only needs to transmit 2 sets of DirAC pairs, but it's possible to use the same audio channels) 2) It was observed that the direct energy of the proposed HO design is equal to the direct energy in the state of the art. 3) The reconstruction of the sector signal in the decoder depends on FOA signals (suitable for higher bit rates). 4) Use existing coders • Multi DoA can improve covariance prediction and decrease residual Cff C = Currently: fF Proposed: + -a) * _rS2* )][...]T ADDITIONAL IMPLEMENTATIONS

[0223] Depending on certain implementation requirements, the examples can be implemented in hardware. Implementation can be carried out using a digital storage medium, for example, a floppy disk, a Digital Versatile Disc (DVD), a Blu-ray Disc, a Compact Disc (CD), a Read-Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), a Memory Petition 870250071890, dated 08 / 14 / 2025, pages 292 / 295 72 / 74 An Electrically Erasable Read-Only Programmable Memory (EEPROM) or flash memory has electronically readable control signals stored within it that cooperate (or are capable of cooperating) with a programmable computer system so that the respective method is performed. Therefore, the digital storage medium can be computer-readable.

[0224] In general, the examples can be implemented as a computer program product with program instructions, where the program instructions can be operated on to execute one of the methods when the computer program product is run on a computer. The program instructions can, for example, be stored on machine-readable media.

[0225] Other examples include the computer program for performing one of the methods described in this document, stored on machine-readable media. In other words, an example of a method is therefore a computer program that has program instructions for executing one of the methods described in this document when the computer program is run on a computer.

[0226] Another example of the methods is, therefore, a data transport medium (or a digital storage medium or a computer-readable medium) comprising, recorded on it, the computer program to execute one of the methods described in this document. The data transport medium, the digital storage medium or the recording medium are tangible and / or non-transitory, as opposed to signals which are intangible and transient.

[0227] A further example comprises a processing unit, for example, a computer or a programmable logic device that performs one of the methods described in this document.

[0228] A further example comprises a computer with the computer program installed to perform one of the methods described in this document.

[0229] A further example comprises a device or system that transfers (for example, electronically or optically) a computer program to Petition 870250071890, dated 08 / 14 / 2025, pages 293 / 295 73 / 74 perform one of the methods described in this document to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device or similar. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0230] In some instances, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functionalities of the methods described in this document. In some instances, a field-programmable gate array can cooperate with a microprocessor in order to perform one of the methods described in this document. In general, the methods can be performed by any suitable hardware device.

[0231] The examples described above are only to illustrate the principles discussed above. It is understood that modifications and variations to the provisions and details described herein will become apparent. Therefore, the intention is to be limited only by the scope of the imminent patent claims and not by the specific details presented by way of description and explanation of the examples contained herein. REFERENCES Document [US20100169103A1] Pulkki, Method and apparatus for enhancement of audio reconstruction [Pulkki2007] Pulkki, V.: Spatial Sound Reproduction with Directional Audio Coding, J. Audio Eng. Soc, 2007, 55, 503-516 [Zotter and Frank] Ambisonics - A Practical 3D Audio Theory for Recording, Studio Production, Sound Reinforcement, and Virtual Reality, Springer, 2019 Documento [WO 2020 / 115311 A1] Fuchs, Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to dirac based spatial audio coding using low-order, mid-order and high-order components generators Documento [US10313815B2] Kuech, Apparatus and method for generating Petição 870250071890, de 14 / 08 / 2025, pág. 294 / 295 74 / 74 a plurality of parametric audio streams and apparatus and method for generating a plurality of loudspeaker signals [Politis2015] A. Politis, J. Vilkamo e V. Pulkki, Sector-Based Parametric Sound Field Reproduction in the Spherical Harmonic Domain, no IEEE Journal of Selected Topics in Signal Processing, vol. 9, n° 5, pp. 852-866, agosto de 2015, doi: 10.1109 / JSTSP.2015.2415762. [Segurar2021] Segure, Christoph, et al. Spatial filter bank design in the spherical harmonic domain. 2021 29aConferência Europeia de Processamento de Sinais (EUSIPCO). IEEE, 2021. Petition 870250071890, dated 08 / 14 / 2025, pp. 295 / 295

Claims

1 / 16 CLAIMS 1. Audio signal representation decoding unit (500) for generating an uncompressed ambisonic spatial audio signal representation (562) from a compressed ambisonic spatial audio signal representation (502) representing an audio signal, characterized in that the compressed ambisonic spatial audio signal representation (502) includes at least one transport channel (501) and side information (503), wherein the side information (503) includes sound field parameters (529, 549, 718), the sound field parameters include, for each spatial sector of a plurality of spatial sectors, the directional parameter(s) (529, 549, 718) that provide(s) information about a direction of arrival, DoA, in the spatial sector, the sound field parameters include, for at least one spatial sector, the sector diffusion parameter(s). (529,549) which provides information about the sectoral diffusion of the audio signal in at least one spatial sector, the audio signal representation decoding unit includes a plurality of sector decoding paths (521, 541), each sector decoding path (521, 541) being configured to decode a directional sector signal (532, 552) from the uncompressed ambisonic spatial audio signal representation (562) in each spatial sector, applying to at least one transport channel (501) or to a sector signal (528, 548) derived from at least one transport channel, the directional parameter(s) (529, 549) and the sector diffusion parameter(s) (529, 549) of the spatial sector, the audio signal representation decoding unit includes a global diffusion signal decoding path (505) configured to to derive a global broadcast signal (510) by applying at least one transport channel (501),a global diffusion parameter (507, 507', Ψ) or other information about the global diffusion of the audio signal, the audio signal representation decoding unit includes a global diffusion signal inserter (560) to combine the plurality of decoded directional sector signals (532, 552) and the global diffusion signal (510), to produce the uncompressed ambisonic spatial audio signal representation (562).

2. Audio signal representation decoding unit, according to claim 1, characterized by being configured to apply, to at least one transport channel (501) or to a sector signal derived from the transport channel, the sector diffusion parameter(s) (529, 549) by weighting the transport channel (501), in at least one sector decoding path, using a mixing weight derived from the sector diffusion parameter(s) (529, 549), to thus derive the sector directional signal (532, 552).

3. Audio signal representation decoding unit, according to claim 2, characterized in that it is configured to weight at least one transport channel or sector signal derived from the transport channel using the mixing weight which is, or is derived from, a positive coefficient received or processed from the sector diffusion parameter(s).

4. Audio signal representation decoding unit, according to claim 2 or 3, characterized by being configured to weight at least one transport channel, or sector signal derived from the transport channel, using the mixing weight, for at least one spatial sector, the mixing weight being, or derived from, a coefficient indicative of a sector directionality in the specific spatial sector.

5. Audio signal representation decoding unit, according to any one of claims 2 to 4, characterized in that it is configured to weight at least one transport channel or sector signal derived from the transport channel using the mixing weight, for each spatial sector, the mixing weight being, or derived from, a coefficient indicative of the relative directionality of the signal in the specific spatial sector over the relative directionalities of the total spatial sectors.

6. Audio signal representation decoding unit, according to any one of claims 2 to 5, characterized by being configured to weigh at least one transport channel or sector signal derived from the transport channel for at least one first spatial sector using a first mixing weight that is, or is derived from, a coefficient indicative of the sector directionality in the first spatial sector, and configured to weight at least one transport channel or sector signal derived from the transport channel for at least one second spatial sector using a second mixing weight, the audio signal representation decoding unit being configured to recover the second mixing weight being recovered by complementing, to a predetermined fixed value, the coefficient indicative of the sector directionality in the first spatial sector.

7. Audio signal representation decoding unit, according to any one of claims 2 to 6, characterized in that it is configured to derive each of the N-1 mixing weights from parameters written in the side information (503), and to derive an N-th mixing weight by complementing the other N-1 mixing weights to a constant positive value, wherein N is the number of spatial sectors.

8. Decoding unit, according to any of the preceding claims, characterized in that it is configured, in each sector decoding path, to apply, to at least one sector signal (528, 548), the directional parameter(s) (529, 549) by multiplying the signal of at least one sector by a vector of spherical harmonic functions evaluated along the DoA ^Ωε) in the space sector, so as to extend the directional signal to the space sector by a higher ambisonic order.

9. Decoding unit, according to any of the preceding claims, characterized by being configured to apply a spatial filter (574, 524, 544) to at least one transport channel (520, 521) or processed version of at least one transport channel, to limit at least one transport channel (520, 521) to a spatial sector for each decoding path of Petition 870250071890, dated 08 / 14 / 2025, page 189 / 295 4 / 16 sector.

10. Decoding unit, according to any of the preceding claims, characterized by being configured to calculate at least one directional sector signal (532, 552) using xs = * Y(íls) = [xs * Yoo(íl), xE * Yi- * Yio(íl),xs * Yu(fi), -], where s indicates the space sector, Xséo transport channel, or processed version thereof, in the specific space sector s, the directional parameter for the specific space sector s, ev, which is a function of Ωε, is the vector of spherical harmonic functions given by ΙΧοϋ(βBλΥι-ι(βB),Υιο(βBλΥιι(ΠB),-Ynm(íU] and Ynm(ílB) is a spherical harmonic of order ne degree m.

11. Decoding unit, according to any of the preceding claims, characterized by being configured to calculate at least one directional sector signal (532, 552) for at least the specific spatial sector using = (1 - Ψ) * aE * Y(ílB) * xE where ψ is the global diffusion parameter, Mo is the sector diffusion parameter expressed as the relative sector directionality in at least one sector signal, and is a vector of spherical harmonic functions evaluated along the DoA in the specific spatial sector.

12. Decoding unit, according to any of the preceding claims, characterized in that it is configured to read the global diffusion parameter from the side information (507, 509).

13. Decoding unit, according to any one of claims 1 to 11, characterized in that it is configured to estimate (570) the global diffusion parameter of at least one transport channel (501).

14. Decoding unit, according to any of the preceding claims, characterized in that it is configured to apply a global diffusion weight obtained from the global diffusion parameter (507, Ψ), or information about the global diffusion of the audio signal, to weight at least one transport channel (501), thus obtaining a version of the global diffusion signal (506) to be used in the decoding path of the global diffusion signal (505), and to apply a second weight, complementary to the global diffusion weight, to weight at least one transport channel (501), thus obtaining at least one globally non-diffuse signal (520) to be processed in the sector plurality of decoding paths (521, 541).

15. Decoding unit, according to any of the preceding claims, characterized in that it is configured to derive the mixing weight(s) of the global diffusion signal (510) and the directional sector signals (532, 552) from the global diffusion parameter (507, 509, Ψ), or information about the global diffusion of the audio signal.

16. Decoding unit, according to any of the preceding claims, characterized in that it is configured to apply, to at least one transport channel (501), a weighing parameter complementary to the global diffusion parameter used to derive the global diffusion signal (510), such that, for each sector decoding path, at least the transport channel is weighed using the weighing parameter.

17. Audio signal representation decoding unit, according to any of the preceding claims, characterized in that the global diffusion signal decoding path (505) is configured to weight at least one transport channel (501) by a global diffusion gain, which is, or is derived from, the global diffusion parameter (507, Ψ), or from other information about the global diffusion of the audio signal, and each of the multiple sector decoding paths (521, 541) is configured to weight at least one transport channel (501) by a global directionality gain, which is, or is derived from, the global diffusion parameter (507), or from other information about the global diffusion of the audio signal.

18. Audio signal representation decoding unit, according to claim 17, characterized in that the global diffusion gain is Petition 870250071890, dated 08 / 14 / 2025, page 191 / 295 6 / 16 1 + §(Ψ) θ esfar ç-ιθ in accordance with , I / Ή+ 1 X 1+ §(ψ) = Jl+Ψ*--1 , BJ XL+ 1 1 where ψ is, or is derived from the global diffusion parameter (507, 509), or from other information about the global diffusion of the audio signal, L is an ambisonic input order and H is an ambisonic output order.

19. Audio signal representation decoding unit, according to claim 17, characterized in that the global diffusion gain is 1 + gÇP) θ esfar ç-ιθ according to I + g(T) = .^ι + Ψ*(Λοίηρ-ΐ) , where ψ is, or is derived from the global diffusion parameter (507, 509), or from other information about the global diffusion of the audio signal, and is a diffuse compensation factor.

20. Audio signal representation decoding unit, according to claim 19, characterized in that the diffuse compensation factor is given by f (Σ ί=0 Σ ™=-i 2 * l + 1) 7cc.mp , , , 1 x (Ã. !=0 L + where i is the degree of a spherical harmonic and L is the ambisonic order of the input signal (501, -Ví) and H is a higher ambisonic order, or a signal comprising the transport channels and channels generated through the use of decorrelators, where is the index of a spherical harmonic and takes values ​​of - ; to L 21. Audio signal representation decoding unit, according to any one of claims 17 to 20, characterized in that the range of values ​​of the overall diffusion gain is limited to a certain range of values ​​to avoid very strong deviations from the overall diffusion signal (506).

22. Audio signal representation decoding unit, according to any one of claims 17 to 21, characterized in that the global diffusion signal decoding path (505) includes an energy compensation unit (508) to apply gain to the global diffusion signal (506) to adjust the energy distribution so as to obtain a more physically realistic ambisonic output signal (502).

23. Audio signal representation decoding unit, according to any of the preceding claims, characterized in that it is configured to switch between: a low-order operating mode, in which, among the plurality of sector decoding paths (521, 541), at least one of the sector decoding paths (521, 541) is deactivated, while only one of the sector decoding paths (521, 541) is activated, in which the side information (503) does not contain the sound field parameter(s) (549) for the deactivated at least one of the sector decoding paths (521, 541);and a high-order mode of operation, in which, among the plurality of sector decoding paths (521, 541), all sector decoding paths of the plurality (521, 541) are activated, or at least fewer sector decoding paths are deactivated in relation to the low-order mode of operation, in which the side information (503) also contains the sound field parameter(s) (529, 549) for the entire plurality of sector decoding paths (521, 541), as well as the global diffusion parameter (507, 509).

24. Audio signal representation decoding unit, according to any of the preceding claims, characterized in that it is configured to convert the spatial audio signal representation (502) of at least one encoded transport channel (739) into a decoded version of the encoded transport channel (739).

25. Audio signal representation decoding unit according to claim 24, characterized by further comprising an EVS decoder (738b) for decoding at least one encoded transport channel (739) into a decoded version of the encoded transport channel (739).

26. Audio signal representation decoding unit, of Petition 870250071890, of 08 / 14 / 2025, p. 193 / 295 8 / 16 according to any of the preceding claims, characterized by being configured to convert the decoded ambisonic spatial audio signal representation (562) from the filter bank domain to the time domain.

27. Audio signal representation decoding unit, according to any of the preceding claims, characterized in that it is further configured to perform upmixing (110) of at least one transport channel (503) from a first transport channel number (501b) to a second transport channel number (501c) greater than the first number.

28. Audio signal representation decoding unit, according to any of the preceding claims, characterized by comprising a mixing matrix estimator (106) configured to process the sound field parameters (503), to derive a covariance matrix (104), or other covariance information, between different transport channels (501), wherein the mixing matrix estimator (106) is configured to reconstruct a mixing matrix (108), or other mixing information, from the covariance matrix (104), or other covariance information, and apply the mixing matrix (108), or other mixing information, to the transport channels (501).

29. Audio signal representation decoding unit, according to claim 28, characterized in that the covariance matrix synthesizer (102) is configured to process the sound field parameter(s), including the DoA parameter(s) and the spatial sector plurality diffusion parameter(s) and the global diffusion parameter, or other global diffusion information, to derive the covariance matrix (104), or other covariance information, between different transport channels, wherein the mixing matrix estimator (106) is configured to reconstruct a mixing matrix (108), or other mixing information, from the covariance matrix (104), or other covariance information, so as to employ the sound field parameter(s) to derive the covariance matrix (104), or other covariance information, for at least one frequency band,where the audio signal representation decoding unit (500b) is configured to derive the covariance matrix (104), or other covariance information, for at least one other frequency band without using the sound field parameters.

30. Audio signal representation decoding unit, according to claim 29, characterized in that it is configured to derive, for at least one other frequency band, the mixing matrix (104) or other mixing information, from covariance information received from the side information (503).

31. Apparatus (100) characterized by comprising: the audio signal representation decoding unit (500), according to any of the preceding claims; a bitstream reader and dequantizer (804), configured to read a bitstream (802), in which the low-order spatial audio signal representation (502) is encoded, and to provide the high-order spatial audio signal representation (502) to the audio signal representation decoding unit (500).

32. Apparatus, according to claim 31, characterized by further comprising a renderer (812), for rendering the audio signal (814) of the ambisonic spatial audio signal representation (562).

33. Apparatus, according to claim 31 or 32, characterized by further comprising an encoding unit (813) for encoding the representation of the high-order spatial audio signal (562) into a second spatial audio signal representation (816).

34. Audio signal representation coding unit for encoding an input spatial audio signal representation (702), which represents an audio signal, into a compressed ambisonic spatial audio signal representation (502, 802) which represents the audio signal, wherein the audio signal representation coding unit is Petition 870250071890, dated 08 / 14 / 2025, page 195 / 295 10 / 16 characterized by being configured to perform down-mixing (1700a, 1700b) of the input spatial audio signal representation (702) to derive at least one transport channel (736, 739, 501); the audio signal representation encoding unit that is configured to derive side information (503), wherein the side information (503) includes sound field parameters (714, 718, 549, 529), the sound field parameters (714, 718, 549, 529) include, for each spatial sector of a plurality of spatial sectors,The directional parameter(s) that provide information about a direction of arrival, DoA, in the spatial sector; the sound field parameters include the sector diffusion parameter(s) that provide information about the diffusion of the audio signal (702) in at least one spatial sector; the coding unit of the sound signal representation includes a plurality of sector parameter estimators (712, 7211, 7212, 721n), wherein each sector parameter estimator (712, 7211, 7212, 721n) is configured to process a specific sector signal (710, 7101, 7102, 710n) of the input spatial sound signal representation (702) in a specific spatial sector of the plurality of spatial sectors, in order to derive the directional parameter(s) and information about the diffusion of the audio signal (702) in at least one sector. space,The audio signal representation encoding unit includes a bitstream recorder (750) for encoding at least one transport channel (736, 501) and side information (503).

35. Audio signal representation coding unit (700), according to claim 34, characterized by additionally including a global diffusion parameter estimator (7129) for estimating a global diffusion parameter (7149, 507, Ψ) to be inserted (716) into the side information (718, 503).

36. Audio signal representation encoding unit (700), according to claim 34, characterized in that it is configured not to write, in the bitstream, a global diffusion parameter (7129). Petition 870250071890, dated 14 / 08 / 2025, pp. 196 / 295 11 / 16 37. Audio signal representation coding unit (700), according to any one of claims 34 to 36, characterized in that it is further configured to estimate a relative directionality of each specific spatial sector with respect to the directionalities of all spatial sectors and to write the coefficient, or information indicative of the relative directionality, as a sector diffusion parameter.

38. Audio signal representation coding unit, according to claim 37, characterized in that it is further configured to estimate the relative directionality as including at least one of a first and a second spatial sector, respectively indicated with ai and a2, and satisfies 1-Ti ai (1 - TJ + (1 - Ψ2) and a2 = 1 — ai where is, or is obtained from the sector diffusion information for the first spatial sector and ^2 is, or is obtained from the sector diffusion information for the second spatial sector.

39. Audio signal representation coding unit, according to claim 37, characterized by being further configured to estimate the relative directionality to include two or more sectors according to g. — ------------ b S / i-Tj) with SjSj = 1, where i indicates the i-th specific spatial sector, ej indicates a j-th generic spatial sector of the plurality of spatial sectors, and > indicate the sector diffusion information for the i-th, specific, spatial sector and each j-th generic spatial sector.

40. Audio signal representation encoding unit, according to any one of claims 34 to 39, characterized in that it is configured Petition 870250071890, dated 08 / 14 / 2025, page 197 / 295 12 / 16 to perform (1700a) an active downmix (730) of the audio signal (702), or a processed version (729) thereof, using a downmix matrix (728), or other downmix information, calculated by a downmix information calculator (726), wherein the downmix information calculator (726) is configured to process the sound field parameter(s) to derive the downmix matrix (728), or other downmix information, based on the global diffusion parameter and the sector diffusion parameters and the directional parameters for each spatial sector of the spatial sector plurality.

41. Audio signal representation coding unit, according to claim 40, wherein the information matrix calculator (726) is characterized by being configured to perform an inter-channel prediction to derive the downmix matrix (728), or other downmix information, based on an inter-channel covariance matrix, or other inter-channel covariance information, the inter-channel covariance matrix or other inter-channel covariance information that is derived from the directional parameter(s) and sector diffusion parameter(s) for each spatial sector of the sector plurality and a global diffusion.

42. Audio signal representation coding unit, according to claim 41, wherein the interchannel covariance matrix C is characterized by being defined as having the element between the ambisonic channel with degree and index 1 and l', respectively, and the ambisonic channel with degree and index e, respectively, and is computed according to Clni,l'n / = (1 - Ψ) » Ex * a2 » Υιη1(Ωθ « Υ / η / (Ωι) + (1 - Ψ) * Εχ « (1 - a)2 * Υ]η1(Ω2) »Y1W(íl2) + +Ψ * σ2 * Εχ * where E is the signal energy, is the Kronecker delta being 1 on the diagonal of the interchannel covariance matrix and 0 off the diagonal of the covariance matrix. 870250071890, dated 08 / 14 / 2025, page.198 / 295 13 / 16 covariance between channels, and are the first and second directional parameters, respectively, and “a” is a relative directionality, or other parameter indicative of a ratio, or other information about the relationship, between the directionality in the spatial sector over the total directionalities of all spatial sectors, ψ is indicative of the global diffusion parameter, and ° is an energy scaling factor.

43. Audio signal representation coding unit, according to any one of claims 37 to 38, characterized in that the inter-channel covariance matrix or other inter-channel covariance information is based on an energy weighted by the spherical harmonics evaluated in the DoAs ( A θ weights ,-ιθ mix (°ΐ'°2,·-'β,ν) for each spatial sector.

44. Audio signal representation encoding unit, according to any one of claims 34 to 42, characterized in that it is further configured to convert the input spatial audio signal representation (702) in the filter bank domain to derive a filter bank version (729) of the input spatial audio signal representation (702), further configured to perform down-mixing of the filter bank domain version (729) of the input spatial audio signal representation (706) to derive at least one transport channel (732) in the filter bank domain, and configured to perform a filter bank synthesis (734) of at least one transport channel (732) from the filter bank domain to the time domain.

45. Audio signal representation encoding unit, according to any one of claims 34 to 44, characterized in that it is configured to perform down-mixing of the input spatial audio signal representation (706) using a channel selector (1700b) to derive at least one transport channel (736, 501), selecting lower-order channels from higher-order channels of the input spatial audio signal representation (706).

46. ​​Audio signal representation coding unit, according to Petition 870250071890, dated 08 / 14 / 2025, pp. 199 / 295 14 / 16 with any of claims 34 to 45, characterized by being further configured to perform enhanced voice service coding, EVS, so as to provide an EVS-encoded version (739) of at least one transport channel (736, 501).

47. Audio encoder (700) characterized by comprising: the encoding unit of the sound signal representation, as defined in any one of claims 34 to 46; a bitstream quantizer and recorder for writing, in a bitstream (802), a low-order spatial audio signal representation (502) and / or the compressed ambisonic spatial audio signal representation (502).

48. Method for decompressing an ambisonic spatial audio signal representation (562) representing an audio signal, characterized in that the compressed ambisonic spatial audio signal representation (502) includes at least one transport channel (501) and side information (503), wherein the side information (503) includes sound field parameters (529, 549, 718), the sound field parameters include, for each spatial sector of a plurality of spatial sectors, the directional parameter(s) (529, 549, 718) that provide information about a direction of arrival, DoA, in the spatial sector, the sound field parameters include, for at least one spatial sector, the sector diffusion parameter(s) (529, 549) that provide information about the sectoral diffusion of the audio signal in at least one spatial sector, the method includes the use of a plurality of sector decoding paths. (521, 541), where each sector decoding path (521,541) decodes a directional sector signal (532, 552) from the ambisonic spatial audio signal representation (562) in each spatial sector, applying to at least one transport channel (501) or to a sector signal (528, 548) derived from the transport channel, the directional parameter(s) (529, 549) and the spatial sector diffusion parameter(s) (529, 549), the method includes the use of a diffusion signal decoding path Petition 870250071890, dated 14 / 08 / 2025, p. 200 / 295 15 / 16 global (505) to derive a global diffusion signal (510) by applying at least one transport channel (501), a global diffusion parameter (507, 507', Ψ) or other information about the global diffusion of the audio signal, the method includes combining, through a global diffusion signal inserter (560), the plurality of decoded directional sector signals (532, 552) and the global diffusion signal (510),to produce the uncompressed ambisonic spatial audio signal representation (562)., 49. Method for encoding an input spatial audio signal representation (706), representing an audio signal, into a compressed ambisonic spatial audio signal representation (502, 802) representing the audio signal, wherein the method is characterized by including deriving at least one transport channel (736, 501) and side information (503), wherein the side information (503) includes sound field parameters (714, 718, 549, 529), the sound field parameters (714, 718, 549, 529) include, for each spatial sector of a plurality of spatial sectors, the directional parameter(s) that provide(s) information about a direction of arrival, DoA, in the specific spatial sector, the sound field parameters include the sector diffusion parameter(s) that provide(s) information about the diffusion of the audio signal (702) in at least one spatial sector, the method includes the use of a plurality of sectoral parameter estimators (712, 7211, 7212, 721n),whereby each sector parameter estimator (712, 7211, 7212, 721n) processes a specific sector signal (710, 7101, 7102, 710n) from the representation of the input spatial audio signal (706) in a specific spatial sector of the plurality of spatial sectors, in order to derive the directional parameter(s) and information about the diffusion of the audio signal (702) in at least one spatial sector, the method includes the use of encoding of at least one transport channel (736, 501) and secondary information (503) in a bitstream. Petition 870250071890, dated 14 / 08 / 2025, p. 201 / 295 16 / 16, 50. A non-transient storage unit characterized by storing instructions that, when executed by the processor, cause the processor to execute the method as defined in claim 48 or 49.

51. Representation of a compressed ambisonic audio signal (802) characterized by including at least one transport channel (501) and side information (503), wherein the side information (503) includes sound field parameters (529, 549, 718), the sound field parameters include, for each spatial sector of a plurality of spatial sectors, the directional parameter(s) (529, 549, 718) that provide information about a direction of arrival, DoA, in the spatial sector, the sound field parameters include, for at least one spatial sector, the sector diffusion parameter(s) (529, 549) that provide information about the sector diffusion of the audio signal in at least one spatial sector, and a global diffusion parameter (509, 7149). Petition 870250071890, dated 14 / 08 / 2025, p. 202 / 295