General spatial audio format transcoder

By using a cost function-based transformation matrix optimization method, the problem of generating the optimal transformation matrix between arbitrary input and output spatial audio formats is solved, achieving audio signal reproduction that meets psychoacoustic requirements and minimizes speaker size.

CN122122925APending Publication Date: 2026-05-29DOLBY INTERNATIONAL AB

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2024-09-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to generate optimal transformation matrices between arbitrary input and output spatial audio formats, especially when decoding to speaker layouts, failing to effectively meet psychoacoustic requirements and the need to minimize the number of speakers.

Method used

By optimizing the transformation matrix based on a cost function, a general transformation matrix is ​​generated. This matrix depends on the encoding and decoding schemes of the input and output formats. Combined with psychoacoustic quantities and speaker layout, an optimal transformation matrix is ​​generated to meet the needs of a specific setup.

Benefits of technology

It achieves optimal conversion between arbitrary input and output spatial audio formats, meets psychoacoustic requirements, minimizes the number of active speakers, and improves the quality and efficiency of audio signal reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122122925A_ABST
    Figure CN122122925A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method for determining an optimized conversion matrix for converting between an input spatial audio format and an output spatial audio format. The input spatial audio format is specified by an encoding scheme for encoding a first set of directions into the input spatial audio format, and the output spatial audio format is specified by a decoding scheme for decoding the output spatial audio format into an output layout. The method comprises optimizing the conversion matrix based on a cost function. The cost function is based on parameters describing the encoding scheme, parameters describing the decoding scheme, and the conversion matrix. The present disclosure also relates to a corresponding apparatus, computer program and computer readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to Spanish Patent Application No. P202330824, filed September 29, 2023; U.S. Provisional Application No. 63 / 606,751, filed December 6, 2023; European Patent Application No. 23214758.7, filed December 6, 2023; and U.S. Provisional Application No. 63 / 557,948, filed February 26, 2024, the entire contents of each of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to techniques for determining optimized transformation matrices (e.g., transcoding matrices) for converting between input spatial audio formats and output spatial audio formats. In particular, this disclosure relates to techniques for determining a universal transformation matrix that is independent of the specific audio signal and depends only on the details of the input and output spatial formats. Background Technology

[0004] Spatial audio signals can take on a wide variety of formats. Some of these formats are layout-independent, such as Ambisonics (e.g., Zotter & Frank, 2019) or spherical wavelet (SWF) formats (e.g., Scaini & Arteaga, 2020), while others are tailored to specific speaker layouts, such as multichannel formats like 5.1, 7.1, or 7.1.4. Additionally, there can be spatial audio formats that are directly associated with the recording setup, such as direct recording with a microphone array.

[0005] Spatial audio formats independent of the recording setup (such as Atmos or SWF) cannot be directly reproduced to a speaker setup; they must first be decoded to a specific speaker setup. Decoding these formats to a speaker setup is often not straightforward, and existing techniques frequently involve decoding to a virtual intermediate speaker setup (e.g., Zotter & Frank 2012; Zotter & Frank, 2018), or optimizing decoding based on psychoacoustic standards (e.g., Scaini & Arteaga, 2014). Similarly, spatial audio formats associated with the recording setup cannot be directly reproduced but must be transformed, either into a multichannel format suitable for direct reproduction or into an intermediate spatial format, such as Atmos. This transformation can be based on knowledge of sound field principles (e.g., Blanco Galindo, Coleman & Jackson, 2020), beamforming techniques (e.g., Blanco Galindo, Coleman & Jackson, 2020), or established stereo mixing principles (e.g., Leal & Elmar, 1991). Finally, even when the intended reproduction layout is readily available, and the multichannel format can be directly reproduced to the speaker setup, it is often the case that the reproduction layout differs from the multichannel format configuration (such as reproducing a 7.1 mix on a 5.1 setup, for example), or the reproduction layout does not conform to the intended layout specifications (such as reproducing a 7.1 mix on a 7.1 layout where the speakers are in irregular positions).

[0006] Therefore, there is a need for improved techniques that, given any input spatial audio format and output spatial audio format, can generate an optimal transformation matrix (e.g., transcoding matrix) from the input format to the output format. Summary of the Invention

[0007] In view of this need, this disclosure provides a method for determining an optimized transformation matrix for converting between an input spatial audio format and an output spatial audio format, as well as corresponding apparatus, computer programs and computer-readable storage media, having the features of the respective independent claims.

[0008] One aspect of this disclosure relates to a method (e.g., a computer-implemented method) for determining an optimized transformation matrix for converting between an input spatial audio format and an output spatial audio format. For example, the format conversion may be related to transcoding. Therefore, in some examples, the transformation matrix may be a transcoding matrix. The input spatial audio format may be specified by an encoding scheme for encoding a first set of directions into the input spatial audio format. The output spatial audio format may be specified by a decoding scheme for decoding the output spatial audio format into an output layout. The method may include optimizing the transformation matrix based on a cost function. The cost function may be based on parameters describing the encoding scheme, parameters describing the decoding scheme, and the transformation matrix. Specifically, the cost function may be based on parameters describing the encoding scheme, parameters describing the decoding scheme, and parameterization of the transformation matrix (e.g., coefficients / elements of the transformation matrix).

[0009] Therefore, given the input and output formats, the proposed method can provide an optimal transformation (e.g., transcoding) matrix tailored to a specific setting. This can be done independently of the actual audio being encoded, transcoded, and / or decoded, to provide a universal transformation matrix for a given (but arbitrary) input and output scheme. Furthermore, adaptation of the cost function, for example by adjusting the (directional) weights, can also satisfy user preferences and / or specific implementation methods.

[0010] In some embodiments, the input spatial audio format can be specified by a first set of directions and an encoding matrix, which encodes input audio signals from various directions in the first set of directions into channels of a spatial input format. The first set of directions can be a set of sampling directions (e.g., the source directions from which the input spatial audio format is obtained). Additionally, the output spatial audio format can be specified by a second set of directions and a decode-to-speaker matrix, which decodes channels of the output spatial audio format into output audio signals for speakers at various directions in the second set of directions. The second set of directions can be a set of speaker directions. Speaker positions can be located on a unit sphere around the origin, with directions in the second set pointing to the corresponding speaker positions. For such input and output spatial formats, the cost function can be based on the encoding matrix, the decode-to-speaker matrix, and the transformation matrix (a parameterization of the transformation matrix).

[0011] In some embodiments, the cost function may be based on a system matrix used to link the input audio signal to the output audio signal. This system matrix may be based on a transformation matrix. The system matrix (or transfer matrix) may be the input-output matrix of an encoder-transcoder-decoder processing chain. The system matrix may be based on a parameterization of the transformation matrix.

[0012] In some embodiments, the system matrix can also be based on an encoding matrix and a decoded loudspeaker matrix.

[0013] In some embodiments, the cost function may also be based on a first set of directions and a second set of directions.

[0014] In some embodiments, at least one term of the cost function may be based on a psychoacoustic quantity. The psychoacoustic quantity may be related to a psychoacoustic quantity determined with respect to the origin of the output layout (e.g., based on or according to a system matrix, assuming the source (e.g., the input audio signal) has unit amplitude).

[0015] This ensures that the generated transformation matrix meets the specific psychoacoustic requirements for the reproduced sound field.

[0016] In some embodiments, at least one term of the cost function may be based on one or more of the following: sound pressure, sound intensity, sound energy, and sound speed. The aforementioned psychoacoustic quantities may be related to (e.g., reconstructed) psychoacoustic quantities determined with respect to the origin of the output layout (e.g., based on or according to a system matrix, assuming the source (e.g., the input audio signal) has unit amplitude).

[0017] In some embodiments, the cost function may include a term penalizing the non-sparseness of the system matrix. This allows for optimization of the transcoding matrix that minimizes the number of active speakers during decoding.

[0018] In some embodiments, the cost function may include a term indicating the mean square deviation of the sound pressure relative to zero phase.

[0019] In some embodiments, the cost function may include a term indicating the mean square deviation of the radial component of the sound velocity relative to zero phase.

[0020] In some embodiments, the cost function may include a term indicating the mean square deviation of the zero phase of the transverse component of the sound velocity.

[0021] In some embodiments, the terms may be based on the system matrix. That is, assuming the source (e.g., the input audio signal) has unit amplitude, the terms of the cost function may be a function of the coefficients of the system matrix.

[0022] In some embodiments, the cost function may include quadratic terms with respect to the coefficients of the system matrix. Assuming the source (e.g., an input audio signal) has unit amplitude, said terms of the cost function may be a function of the coefficients of the system matrix.

[0023] In some embodiments, the cost function may include a fourth-order term with respect to the coefficients of the system matrix. Assuming the source (e.g., an input audio signal) has unit amplitude, the term of the cost function may be a function of the coefficients of the system matrix.

[0024] In some embodiments, the cost function may include a term penalizing the non-sparseness of the system matrix. This allows for optimization of the transcoding matrix that minimizes the number of active speakers during decoding.

[0025] In some embodiments, the cost function may include a term that penalizes larger values ​​of the elements (or coefficients, entries, etc.) of the transformation matrix. This term may be penalized with a measure positively correlated with the magnitude of the values ​​of one or more (e.g., all) elements of the transformation matrix. This avoids excessive gains in the transformation from input to output format.

[0026] In some embodiments, the cost function may include a term that penalizes elements in a transformation matrix associated with a channel whose noise level exceeds a predetermined threshold.

[0027] In some embodiments, the encoding scheme may be a complex-valued scheme. For example, the encoding matrix may be a complex-valued matrix. Alternatively, the encoding scheme may be a real-valued scheme. For example, the encoding matrix may be a real-valued matrix.

[0028] In some embodiments, the output layout may be associated with the physical speaker layout. Alternatively, the output layout may be associated with the virtual speaker layout.

[0029] In some embodiments, the output spatial audio format may be one of Dolby Atmos, multi-channel format, and linear decoding format.

[0030] In some embodiments, the input spatial audio format may be one of Dolby, multichannel, and linear coding formats.

[0031] In some embodiments, encoding the first set of directions into the input spatial audio format may be associated with (or may involve) recording a set of input audio signals via a microphone array.

[0032] In some embodiments, encoding the first set of directions into the input spatial audio format can involve a one-to-one mapping between the input audio signal and the channels of the input spatial audio format. That is, the encoding matrix can be a permutation matrix or an identity matrix. This allows for decoding optimization.

[0033] In some embodiments, decoding the output spatial audio format to the output layout can involve a one-to-one mapping between the channels of the output spatial audio format and the output audio signal. That is, decoding to the speaker matrix can be a permutation matrix or an identity matrix. This allows for encoding optimization.

[0034] In some embodiments, at least one of the encoding and decoding schemes may be frequency-dependent. The transformation matrix can then be optimized individually in each of the multiple frequency bands.

[0035] In some embodiments, the cost function may be frequency-dependent.

[0036] In some embodiments, the transformation matrix can be optimized simultaneously across multiple frequency bands. In this case, the cost function may include a contribution for each of the multiple frequency bands and one or more terms penalizing discontinuities in the transformation matrix between adjacent frequency bands.

[0037] In some embodiments, the method may further include determining an effective cost function for a given frequency band. The transformation matrix can be optimized based on the effective cost function for the given frequency band. Furthermore, the determination of the effective cost function for a given frequency band may be based on multiple cost functions at various frequencies between the lower and upper bound frequencies of the given frequency band. The effective cost function for a given frequency band can be used for individual optimization of the transformation matrix within the given frequency band, or for simultaneous optimization of the transformation matrix across frequency bands. The effective cost function can be an aggregated cost function (e.g., average, weighted average, sum, weighted sum, etc.) of the cost functions at multiple frequencies between the lower and upper bound frequencies. The cost functions at multiple frequencies can be determined, for example, based on a frequency-dependent coding matrix.

[0038] According to another aspect, an apparatus is provided. This apparatus may include a processor and a memory coupled to the processor and storing instructions for the processor. The processor may be configured to perform methods or method steps as outlined throughout this disclosure.

[0039] According to another aspect, a computer program is described. This computer program may include executable instructions, which, when executed by a computing device (e.g., a processor), are used to perform the methods or method steps outlined throughout this disclosure.

[0040] According to another aspect, a computer-readable storage medium is described. The storage medium may store a computer program adapted to execute on a computing device (e.g., a processor) and, when executed on a computing device, to perform the methods or method steps outlined throughout this disclosure.

[0041] It should be noted that the methods and apparatus, and their preferred embodiments thereof, summarized throughout this disclosure, can be used alone or in combination with other methods and apparatus disclosed in this document. Furthermore, all aspects of the methods and apparatus summarized throughout this disclosure can be combined in any manner. In particular, the features of the claims can be combined with each other in any way.

[0042] It will be understood that the features of the apparatus and the steps of the method can be interchanged in many ways. In particular, the details of the disclosed method can be implemented by the corresponding apparatus and vice versa, as will be understood by those skilled in the art. Furthermore, any of the foregoing statements made for a method (e.g., its steps) are considered equally applicable to the corresponding apparatus (e.g., its blocks, stages, units) and vice versa. Attached Figure Description

[0043] The invention will now be described by way of example with reference to the accompanying drawings, wherein...

[0044] Figure 1 An overview of an example system for transformation matrix optimization according to embodiments of the present disclosure is illustrated schematically;

[0045] Figure 2 The illustration schematically shows embodiments of the present disclosure that can be optimized. Figure 1 An example of the system matrix of a system;

[0046] Figure 3 This is a flowchart illustrating an example of a method for determining an optimized transformation matrix according to embodiments of the present disclosure; and

[0047] Figure 4 This is a block diagram schematically illustrating an example of an apparatus suitable for implementing a method according to embodiments of the present disclosure. Detailed Implementation

[0048] In the following description, exemplary embodiments of the present disclosure will be described with reference to the accompanying drawings. The same elements in the drawings may be denoted by the same reference numerals, and repeated descriptions thereof may be omitted.

[0049] introduce

[0050] In general, this disclosure relates to a general audio format transcoding method that, given any input spatial audio format (such as recording with a circular microphone array, third-order immersive, or 5.1 multichannel) and output spatial audio format (such as second-order immersive or 7.1 multichannel), generates an optimal transcoding matrix from the input format to the output format based on the optimization of a cost function. The cost function can be based on psychoacoustic principles. For this purpose, the proposed method requires neither detailed specifications of the input format nor detailed specifications of the output format; for the input format, the only information required is the encoding of the plane wave in a set of sampling directions, and for the output format, the only information required is the additional decoding from the output format to the speaker layout to the speaker matrix.

[0051] As a specific case, the transcoding method applicable to the technology according to this disclosure includes encoding from recording settings to an output spatial audio format, and decoding from an input spatial audio format to a speaker layout. In the latter decoding case, the additional decoding-to-speaker matrix can simply be an identity matrix or a routing matrix (e.g., a permutation matrix) that maps each channel to each speaker.

[0052] In another specific case, the proposed method allows finding the optimal panning matrix from any input location to any speaker layout.

[0053] System Overview

[0054] Figure 1 An overview of an example system for optimizing a transformation matrix (e.g., a transcoding matrix) 110 according to embodiments of the present disclosure is illustrated. This transformation matrix (e.g., transcoding matrix) 110 is used for conversion between an input spatial audio format and an output spatial audio format. While transcoding matrices will be frequently mentioned throughout the disclosure, this should not be taken as biased or intended as limiting, and it should be understood that the present disclosure is equally applicable to any transformation matrix used for conversion between an input spatial audio format and an output spatial audio format.

[0055] Not intended to limit, the input spatial audio format (input format) is - Channel format, and is used to direct the first group of directions (e.g., Each sampling direction) is encoded into the input spatial audio format (its sampling direction). The encoding scheme for each channel is specified. Specifically, the input spatial audio format can be specified by a first set of directions and an encoding matrix (input matrix) 120 for encoding the input audio signals from each direction in the first set of directions into channels of the spatial input format. Spatial orientation and input format One audio channel, 120-bit encoding matrix It will be matrix.

[0056] Not intended to limit, the output spatial audio format (output format) is - Channel format, and is used to output spatial audio format (its The decoding scheme for decoding (each channel) to the output layout is specified. This output spatial audio format can be determined by a second set of directions (e.g., ...). The speaker matrix 130 specifies the decoding of the output spatial audio format channels to the output audio signals for the speakers in each of the second set of directions (speaker directions). A second direction and output space audio format Each channel is decoded to a 130-speaker matrix. It will be matrix.

[0057] As mentioned above, the first set of directions can be A set of 125 sampling directions (e.g., microphone direction). The second set of directions can be... A set of speaker directions, wherein the speaker position may be located on a unit sphere around the origin, and the directions in the second set of directions point to the corresponding speaker positions.

[0058] In practice, it is possible to give Given a set of 125 sampling directions (the first group of directions), generate an input matrix 120. Encoding each sampling direction into an input spatial audio format can be associated with recording a set of input audio signals via a microphone array. For example, input matrix 120 could be a spatial sample polar plot of a set of microphones (e.g., gain and phase in the case of a complex-valued matrix). As another example, input matrix 120 could be the spatial sample-coded gain of a linear coding format (e.g., Dolby, SWF, or VBAP) (in the case of a real-valued matrix). In other examples, input matrix 120 could be an identity matrix (or permutation matrix) assigning each sampling direction to a separate channel.

[0059] Depending on its details and implementation, the encoding scheme (e.g., encoding matrix 120) can be complex-valued or real-valued. Furthermore, the encoding scheme can be frequency-dependent.

[0060] Furthermore, as stated above, without limitation, the input spatial audio format (input format) can be, for example, one of Dolby, multi-channel, and linear coding formats.

[0061] In all cases, it should be understood that the input matrix 120 is a known input to the method according to embodiments of this disclosure. Similarly, the sampling direction (first set of directions) 125 is also considered to be known.

[0062] Similarly, not intended to limit, the output spatial audio format can be one of Dolby, multi-channel, and linear decoding formats.

[0063] In addition, the above output layout may involve The physical speaker layout refers to the physical speaker placement of each physical speaker in its corresponding position / or orientation, or alternatively, to the virtual speaker layout depending on the desired optimization.

[0064] It should be understood that decoding to the speaker matrix 130 is also a known input to the method according to embodiments of this disclosure. Similarly, the speaker orientation (second set of orientations) is also considered known.

[0065] Since the encoding and decoding schemes are known inputs to the process, only the transformation matrix (e.g., the transcoding matrix) 110 undergoes optimization. For input spatial audio formats... Individual channels and output space audio formats One channel, 110 transcoding matrix It will be matrix.

[0066] The optimization process for optimizing the transcoding matrix 110 optimizes (e.g., minimizes) the cost function (described in more detail below) by adjusting the value of the transcoding matrix 110, which can typically be a complex-valued matrix.

[0067] exist Figure 1 In the example, the cost function used for optimization depends on the product (matrix product) of the transcoding matrix 110 and the decoded-to-speaker matrix 130, for example, as a function of that product, which can typically be a real-valued matrix. The product may involve methods for transcoding the input spatial audio format. Each channel is decoded as The decoding matrix for each speaker signal is 140.

[0068] The product of the transcoding matrix 110 and the decoded speaker matrix 130 can be called the decoding matrix 140, as described above. Figure 2 As shown in the example, the product of input matrix 120 and decoding matrix 140 can be called system matrix 150. System matrix 150 describes the entire encoding, transcoding, and decoding process. In other words, it can be said that system matrix (or transfer matrix) 150 is related to the input-output matrix of the encoding-transcoding-decoding processing chain, and therefore it can be said that it will... The input audio signal is linked to One output audio signal (speaker signal). For One input audio signal and One output audio signal, system matrix 150. It will be matrix.

[0069] The cost function described above for optimization can be based on the cost function used to optimize the cost function. The input audio signal is linked to A system matrix 150 outputs an audio signal. The system matrix 150 then depends on a transformation matrix 110 (e.g., its parameterization, such as in terms of its elements / coefficients / entries), and further on an encoding matrix 120 and a decoded-to-speaker matrix 130. In general, the cost function can be said to be based on parameters describing the encoding scheme, parameters describing the decoding scheme, and the transformation matrix (e.g., its parameterization, such as in terms of its elements / coefficients / entries). For example, the cost function can be based on the encoding matrix, the decoding matrix, and the transformation matrix. Furthermore, as described in more detail below, the cost function can be further based on a first set of directions and a second set of directions.

[0070] The cost function's functional dependency on transformation matrix 110 (e.g., through its dependency on system matrix 150 or its elements / coefficients / entries, etc.) can be determined by the desired transcoding type and can be fixed at initialization. For example, this can apply to the multiplication coefficients of the cost function (e.g., as defined below). Weights) and / or (optionally) directional weights (e.g., directional weights defined below) ).

[0071] It should be pointed out that, Figure 1 and Figure 2 The document does not specify where user input might be required, what user input might be needed, or where the user input might be directed. In any case, this is irrelevant to the technology according to this disclosure. Typically, the user can configure all entities indicated by solid boxes in these figures. The only system part that is not directly configurable by the user is the transcoding matrix 110 generated by the optimization process. However, optimization can also depend on user input or preferences through a cost function.

[0072] Symbols and Definitions

[0073] Unless intended to be limiting, throughout this disclosure it will be assumed that M is the number of channels in the input format, N is the number of channels in the output format, and L is the set of directions (first set of directions) for sampling the (unit) sphere. It is a set of P output speaker positions (second group of directions), { It encodes unit amplitude point sources from each of the L sampling directions into an L×M input matrix (encoding matrix) across M channels in the input format. The P×M decoding matrix { Generate L×P output signal set { That is, the signal fed to the loudspeaker p and reproduced from the unit amplitude source in direction l.

[0074] (1)

[0076] Or, more simply, written in matrix notation as,

[0077] (2)

[0079] In a more general case, the complete decoding matrix { From the transformation matrix (e.g., the transcoding matrix) { and additional decoding to speaker matrix composition,

[0080] (3)

[0082] Or, more simply, written in matrix notation as,

[0083] (4)

[0085] N×M transcoding matrix Convert the input format (M channels) to the output format (N channels). Assume an additional P×N decoding to the speaker matrix. It is known that the goal is to optimize only the transcoding matrix. Decoding to speaker matrix Map N output format channels to P speaker channels.

[0086] In some embodiments, the speaker's output layout may correspond to a real (i.e., physical) listening layout. In other cases, the output layout may not correspond to any existing listening layout and may therefore be associated with a virtual listening layout. This distinction may not be known according to the techniques of this disclosure and may be applied equivalently to either case.

[0087] In one embodiment (e.g., for decoding optimization), the output format (output spatial audio format) can be a multi-channel setup corresponding to the listening environment. The input matrix (encoding matrix) can be real-valued, and the number of output speakers... Can be related to the number of output channels Consistent, Furthermore, in this case, decoding the output format to the output layout can involve the output format... Each channel and A one-to-one mapping between each output audio signal. For example, additional decoding to a speaker matrix. It can be an identity matrix. And the transcoding matrix This can correspond to the decoding matrix. In this case, as described above, the speaker output layout corresponds to the actual listening layout. One example could be atmos decoding, where the input matrix is ​​a spatially sampled atmos codec, with dimensions corresponding to the number of sampling directions multiplied by the number of atmos channels. The decoding matrix maps from the atmos channels to the speakers (e.g., Scaini & Arteaga, 2014). In a variation of this embodiment, the decode-to-speaker matrix is ​​not an identity matrix, but could be a permutation matrix mapping each output path to the corresponding speaker.

[0088] Because in this embodiment, the transcoding matrix Corresponding to the decoding matrix Therefore, the optimization according to this disclosure will affect the decoding matrix. Optimize.

[0089] In another embodiment (e.g., microphone encoding optimization), the input format (input spatial audio format) can be a direct recording of the microphone array, evaluated for each frequency band, so the input matrix (encoding matrix) for each frequency band can be complex-valued. The output format (output spatial audio format) can be a layout-independent format (e.g., Dolby Atmos) or a multi-channel format (e.g., 7.1). In this case, typically, the reproduced layout can be a conventional virtual layout that explores all relevant spatial locations but does not necessarily correspond to any actual (i.e., physical) layout. One example could be the optimization of encoding recordings from tetrahedral sound field microphones into a first-order Dolby Atmos signal. For example, the virtual speaker layout could be an octagonal grid of speakers. Other examples are encoding recordings from microphones in a mobile phone into a two-dimensional first-order Dolby Atmos or 5.1 layout. In each of these latter examples, the reproduced layout could be a two-dimensional virtual hexagonal arrangement.

[0090] In this case, the optimization according to this disclosure will be applied to a specific encoding matrix { Optimize the transformation matrix.

[0091] In yet another embodiment (e.g., virtual source encoding optimization), the input format may include A set of channels, with each sampling point corresponding to one channel. Encoding of individual channels to the input format can involve Each input audio signal (channel) and input format A one-to-one mapping between each channel. For example, the input matrix { It can be simply the identity matrix (therefore, The representation layout can be an actual representation layout or a regular virtual layout that explores all relevant spatial locations. This embodiment allows for optimization of the encoding of input sources in each of the sampling directions. In one example, this can be used to find the optimal encoding for a set of sampled point sources in Dolby Atmos. In another example, the method can be used to find the optimal panning function for a set of point sources in a multichannel 5.1 format.

[0092] In yet another embodiment (e.g., transcoding optimization), the input format can be complex or real, multi-channel, or layout-independent, and the output format can similarly be complex or real, layout-dependent, or layout-independent. The speaker layout under consideration can be actual (i.e., physical) or virtual. One example is finding the optimal transcoding matrix from 5.1 to Dolby Atmos, where the speaker layout corresponds to a virtual speaker layout with 8 speakers placed at the vertices of an octagon.

[0093] Regardless of the specific details of the encoding and decoding schemes, this disclosure aims to obtain an optimized transcoding matrix by optimizing (e.g., minimizing) the cost function. Preferably, the cost function is a cost function based on psychoacoustic principles.

[0094] Figure 3 This is a flowchart illustrating a method (e.g., a computer-implemented method) 300 for determining an optimized transformation matrix (e.g., a transcoding matrix) according to embodiments of the present disclosure. Method 300 includes steps S310 to S330, wherein steps S310 and S320 can be performed in any order.

[0095] In step S310, parameters describing the coding scheme are obtained. This may involve obtaining (e.g., receiving as input) a first set of directions and obtaining (e.g., receiving as input, or determining) the coding matrix. For example, if it is not received as input, the coding matrix may be determined based on the first set of directions and the input spatial audio format.

[0096] In step S320, parameters of the decoding scheme are obtained. This may involve obtaining (e.g., receiving as input) a second set of directions, and obtaining (e.g., receiving as input, or determining) the decoded-to-speaker matrix. For example, if there are no inputs received, the decoded-to-speaker matrix may be determined based on the second set of directions and the output spatial audio format.

[0097] In step S330, the transformation matrix is ​​optimized based on the cost function. As described above, this cost function is based on the parameters describing the encoding scheme, the parameters describing the decoding scheme, and the transformation matrix.

[0098] For frequency-dependent settings, optimization can be performed individually in each of the multiple frequency bands, or simultaneously across the frequency bands, by utilizing the respective contribution of each frequency band to the cost function.

[0099] The details of the cost function and possible terms (cost function terms) will be described next.

[0100] Cost function

[0101] The cost function according to this disclosure can be a cost function based on psychoacoustic principles, or in other words, at least one term of the cost function (cost function term) can be based on psychoacoustic quantities. Examples of such cost function terms will be described in more detail below.

[0102] The cost function may contain, for example, any, some, or all of the following terms:

[0103] + + (5)

[0105] in yes The weight of each sub-cost (cost function term). Weights can be selected by the user. For example, weights can be selected based on optimization preferences. Assuming a cost function based on psychoacoustic principles, then weights can be selected... Weights are used to emphasize certain psychoacoustic quantities whose accurate reproduction is of particular concern, while potentially suppressing (or completely eliminating) other psychoacoustic quantities. Therefore, the cost function may contain any, some, or all of the cost function terms listed in equation (5).

[0106] The cost function described above can specifically include two types of terms or sub-cost functions: quadratic terms and quartic terms. The former assumes the coherent behavior of the sound field and can be used in the encoding process (encoding optimization) because they are directly related to the accurate reconstruction of the sound field. The latter assumes the incoherent behavior of the sound field and can be preferably used for decoding optimization because incoherent behavior is typically found in the reconstructed environment. For example, the selection of preferred terms during the optimization phase can also be related to the frequency range of interest in transcoding; this selection can be achieved, for example, through appropriate selection... Weights are used to achieve this. As an example, at low frequencies, the reconstruction of pressure and velocity may be preferred, while at high frequencies, the reconstruction of energy and intensity may be preferred. This preference arises from psychoacoustic considerations (e.g., Gerzon, 1992; Frank, 2013).

[0107] As described above, at least one term of the cost function can be based on a psychoacoustic quantity. This psychoacoustic quantity can be related to a psychoacoustic quantity determined (e.g., reconstructed) with respect to the origin of the output layout, for example, based on or according to the aforementioned system matrix, assuming the source (e.g., the input audio signal) has unit amplitude. For example, at least one term of the cost function can be based on sound pressure level. Additionally or alternatively, at least one term of the cost function can be based on sound intensity. Additionally or alternatively, at least one term of the cost function can be based on sound energy. Additionally or alternatively, at least one term of the cost function can be based on sound speed. Examples of specific cost function terms based on these psychoacoustic quantities will be described in more detail below.

[0108] The quadratic term of the cost function

[0109] As described above, the cost function may include one or more quadratic terms with respect to the coefficients of the system matrix. These terms of the cost function may also be functions of the coefficients of the system matrix, assuming the source (e.g., the input audio signal) has unit amplitude.

[0110] The following section will introduce additional quantities used to describe the aforementioned cost function terms.

[0111] The input signal reconstructed at the origin pressure It is given by the following formula:

[0112] (6)

[0114] as well as

[0115] (7)

[0117] Each has a real part and an imaginary part. and ,

[0118] (8)

[0120] Furthermore, the input signal reconstructed at the origin speed of sound The real and imaginary parts are given by the following equation:

[0121] (9)

[0123] as well as

[0124] (10)

[0126] in It is the unit vector pointing in the direction of the speaker p. The real part of the velocity vector. It can be projected onto both the radial and lateral portions, as shown below.

[0127] (11)

[0129] (12)

[0131] in It refers to the sampling position / direction. The unit vector. The imaginary part of the velocity vector. This can be similarly projected onto the radial and lateral portions, as follows:

[0132] (13)

[0134] (14)

[0136] Among them, the real radial component This represents the desired component of the real intensity vector, while all imaginary parts and This represents the unwanted component. In an ideal system, , , , , ,as well as .

[0137] Based on the above quantities, the following six different cost function terms can be defined.

[0138] (15)

[0140] in

[0141] (16)

[0143] (17)

[0145] in

[0146] (18)

[0148] (19)

[0150] in

[0151] (20)

[0153] (twenty one)

[0155] in

[0156] (twenty two)

[0158] (twenty three)

[0160] in

[0161] (twenty four)

[0163] as well as

[0164] (25)

[0166] in

[0167] (26)

[0169] In the above text, These are (optional) directional weights that allow some directions to be preferred over others.

[0170] The above contributions can be explained as follows: It is the mean square deviation relative to the correct pressure level; It is the mean square deviation relative to zero phase; It is the mean square deviation relative to the optimal directionality, where This means that the apparent size of the source is not optimal; It is the mean square deviation relative to zero phase; and finally, and These are the mean square values ​​of the directional error components and their deviations from zero phase, thus representing the positioning error.

[0171] The cost function according to this disclosure may include cost function terms related to any, some, or all of the foregoing.

[0172] Specifically, the cost function may include a term indicating the mean square deviation of the sound pressure relative to zero phase (e.g., based on...). Or a corresponding term). Additionally or alternatively, the cost function may include a term indicating the mean square deviation of the radial component of the sound velocity relative to zero phase (e.g., based on...). Or a corresponding term). Additionally or alternatively, the cost function may include a term indicating the mean square deviation of the transverse component of the sound velocity relative to zero phase (e.g., based on...). (or their corresponding terms). Any of these terms can be based on the system matrix. That is, assuming the source (e.g., the input audio signal) has unit amplitude, the individual terms of the cost function can be functions of the coefficients of the system matrix.

[0173] The fourth term of the cost function

[0174] As mentioned above, the cost function can also include one or more quartic terms with respect to the coefficients of the system matrix. Assuming the source (e.g., the input audio signal) has unit amplitude, these terms of the cost function can also be functions of the coefficients of the system matrix.

[0175] The following will describe the additional quantities used to describe the cost function term described above. The input signal is reconstructed at the origin (e.g., under the assumption of incoherence). Energy from

[0176] (27)

[0178] Furthermore, the input signal is reconstructed at the origin (e.g., under the assumption of incoherence). The sound intensity is given by the following formula.

[0179] (28)

[0181] vector It can be projected onto both the radial and lateral portions, as follows.

[0182] (29)

[0184] (30)

[0186] radial portion This represents the desired component of the intensity vector and can be related to the size of the visible source. Tangent (i.e., transverse) portion. This represents the unwanted component and can be related to positioning error. In an ideal system, , ,and .

[0187] Based on the above quantities, the following three different cost function terms can be defined:

[0188] (31)

[0190] (32)

[0192] as well as

[0193] (33)

[0195] In the above text, These are (optional) directional weights that allow some directions to be preferred over others.

[0196] These contributions can be explained as follows: C E It is the mean square deviation relative to the correct energy reconstruction; C RI It is the mean square deviation relative to the optimal directionality, where C RI >0 means the source's view size is not optimal; finally, C TI It is the mean square value of the directional error component, and therefore represents the positioning error.

[0197] The cost function according to this disclosure may include cost function terms related to any, some, or all of the foregoing.

[0198] Other cost function terms

[0199] In-phase decoding can be achieved through an additional cost function term to penalize out-of-phase contributions from different speakers. This is based on a quadratic term. and fourth term Two versions were proposed.

[0200] (34)

[0202] (35)

[0204] (36)

[0206] (37)

[0208] Based on the above quantities, the following two different cost function terms can be defined:

[0209] (38)

[0211] (39)

[0213] The above terms of the cost function and This is proportional to the square or fourth power of the system matrix components, and may favor smaller transcoding matrices (e.g., transcoding matrices with fewer entries) during minimization. To avoid this unwanted side effect, in some other examples, it is based on quadratic terms. and fourth term The cost function term can be alternatively defined as:

[0214] (40)

[0216] (41)

[0218] The cost function according to this disclosure may include a cost function term associated with any one or both of the cost function terms in equations (38) and (39), or a cost function term associated with any one or both of the cost function terms in equations (40) and (41).

[0219] Additionally, a symmetry cost can be added to the cost function to force similar reconstruction of psychoacoustic variables in left-right symmetrical directions. Once left-right symmetrical pairs are identified in the destination layout... This allows us to quantify the asymmetric quantities that arise during decoding (again, both exist in the quadratic and linear versions).

[0220] (42)

[0222] (43)

[0224] In some examples, the following two different cost function terms can be defined based on the quantities mentioned above:

[0225] (44)

[0227] (45)

[0229] To avoid favoring a smaller transcoding matrix during the minimization process, based on the aforementioned asymmetric quantities... and The cost function term can be alternatively defined as:

[0230] (46)

[0232] (47)

[0234] Wherein, the weighting function w l It is an optional bias factor that allows for improved decoding performance in some regions of the space (at the expense of other regions). Bias-free decoding is achieved by... Provided.

[0235] The cost function according to this disclosure may include a cost function term associated with any one or both of the cost function terms in equations (44) and (45), or a cost function term associated with any one or both of the cost function terms in equations (46) and (47).

[0236] Additionally, quadratic and / or quartic sparsity enhancement terms can be added to the cost function. These terms penalize non-sparse solutions to the transcoding matrix (or system matrix).

[0237] In a mathematical or computational context, a sparse vector or matrix is ​​a vector or matrix in which most of its elements are zero or close to zero. In spatial audio, a reproduction method that activates a small number of speakers for a given virtual sound source is often preferred (e.g., typically 1 to 4, but fewer the better, depending on the implementation). In this context, this preference can be translated into the requirement that each row of the system matrix, which represents the input source and speakers, be as sparse as possible. A linked matrix.

[0238] Not intended to restrict, one possible way to quantify the non-sparseness of a solution is based on the L1 and L2 norms of the rows of the system matrix. For example, non-sparseness can be quantized based on the difference between the L1 and L2 norms of the rows of the system matrix:

[0239] (48)

[0241] (49)

[0243] Alternatively, non-sparseness measures can be defined based on the ratio of the L1 and L2 norms of the rows of the system matrix; note that substantially equal L1 and L2 norms indicate high sparsity. Further examples of non-sparseness measures that can be used in the context of this disclosure are summarized in Niall Hurley and Scott Rickard, “Comparing Measures of Sparsity,” arXiv:0811.4706v2 [cs.IT], April 27, 2009, the full text of which is incorporated herein by reference.

[0244] In some examples, the following two different cost function terms can be defined based on the quantities mentioned above:

[0245] (50)

[0247] (51)

[0249] The proposed cost function terms measure the non-sparseness of the rows of the system matrix by comparing the L2 and L1 norms. When one or more of these terms are present in the cost function, the optimizer will attempt to minimize these non-sparse cost function terms, resulting in sparser decoding and fewer speaker activations.

[0250] To avoid favoring a smaller transcoding matrix during the minimization process, the cost function term based on the aforementioned measure of non-sparseness (or its equivalent) can alternatively be defined as:

[0251] (52)

[0253] (53)

[0255] The cost function according to this disclosure may include a cost function term associated with any one or both of the cost function terms in equations (50) and (51), or a cost function term associated with any one or both of the cost function terms in equations (52) and (53). Generally, the cost function according to this disclosure may include one or more cost function terms (e.g., quadratic or quartic) that depend on a measure of the sparsity (or non-sparseness) of the system matrix, or in other words, penalize the non-sparseness of the system matrix.

[0256] Depending on the implementation and circumstances, additional cost function terms can be added. For example, a cost function term could be added that penalizes large values ​​in the transcoding matrix elements to avoid excessive gain. One such possible term could be...

[0257] (54)

[0259] For example, the value of the exponent q can be 1 or 2. In yet another example, this term can be activated only when a certain value t is exceeded.

[0260] (55)

[0262] in, It is a heaviside step function, which allows for a value of the exponent in the latter case. .

[0263] In any case, in some embodiments, the cost function may include a term that penalizes large values ​​of the transformation matrix elements. Generally, this term may impose a penalty that is positively correlated with a measure of the magnitude of the values ​​of one or more (e.g., all) elements of the transformation matrix. This avoids excessive gain in the transformation from the input format to the output format. Examples of such terms are given in equations (54) and (55) above.

[0264] In the case of microphone encoding, an additional term can be added to the cost function to limit the gain in the transcoding matrix relative to a metric of the microphone's self-noise. Given a metric of the self-noise for each input microphone (e.g., broadband, or per frequency or band), an example of such a cost function term could be the sum of the differences between the transcoding matrix gain and the product of the self-noise values, and some fixed threshold.

[0265] (56)

[0267] in Each microphone Self-noise (in dB), It is the threshold (in dB), and It is the Heaviside step function. That is, the cost function typically includes a term that penalizes elements in the transformation matrix associated with the channel whose noise level exceeds a predetermined threshold.

[0268] Frequency-related cost function term

[0269] The process described so far can be generalized to the case where any component of the system is frequency-dependent. For example, at least one of the encoding and decoding schemes can be frequency-dependent. In other words, the input matrix or the decoding matrix, or both, can be frequency-dependent. This typically results in a frequency-dependent cost function (e.g., via a frequency-dependent system matrix).

[0270] In one embodiment, in the desired frequency band The input (encoding) and decoding matrices are obtained. Then, the cost function (and thus the transcoding matrix) is optimized separately (independently) in each frequency band.

[0271] , (57)

[0273] turn out An independent (optimized) transcoding matrix This includes cost items. Weights and directional weights It may be frequency-related.

[0274] Independently optimizing the transcoding matrix in each frequency band may result in arbitrarily large variations in the (transcoding) gain across frequencies, which may negatively impact the listening experience.

[0275] To mitigate this, in one embodiment, optimization is performed simultaneously across (multiple) frequency bands, with the total cost function... Including contributions to each of the multiple frequency bands. And one or more terms of the penalty transformation matrix that cause discontinuities between adjacent frequency bands. For example, the cost function It can be the cost at each frequency. The sum is added to the following term, which is obtained by adding T( ) to the adjacent frequency bands. The absolute values ​​of the differences between adjacent frequency bands are added together to penalize discontinuities, for example as follows:

[0276] (58)

[0278] in This represents the Frobenius norm.

[0279] In another embodiment, a cross-frequency averaging method can be utilized. Each frequency band can be sequentially optimized via a cost function that, given a specific averaging strategy between design points, relates to the emergent response at frequencies surrounding the design point (e.g., the center frequency in a frequency-correlated approach). For example, a simple averaging approach might be assumed as follows:

[0280] (59)

[0282] Where N is and The number of frequency points between (the lower and upper limits of frequency band b, respectively). Optionally, the summation in equation (59) may include a frequency-dependent pre-factor to implement, for example, a window function across frequencies.

[0283] In conjunction with the foregoing, the method according to an embodiment of the present invention may include: determining an effective cost function for a given frequency band b. ; and based on the effective cost function for a given frequency band b To optimize the transformation matrix (e.g., the transcoding matrix). This involves determining the effective cost function for a given frequency band b. It can be based on the lower limit frequency of the given frequency band. and the upper limit frequency of the given frequency band Multiple cost functions C(f) at various frequencies f between these frequencies. For example... and The frequency f between them can be in and A given number of N frequencies are equidistantly distributed between them, optionally including and .

[0284] In optimization, the effective cost function for a given frequency band can be used for individual optimization of the transformation matrix within that band, such as in equation (57), or for simultaneous optimization across frequency bands, such as in equation (58). The effective cost function can be an aggregate cost function (e.g., average, weighted average, sum, weighted sum, etc.) of cost functions at multiple frequencies between the lower and upper bound frequencies. The cost functions at multiple frequencies can be determined, for example, based on the frequency-dependent system matrix (e.g., based on the frequency-dependent input matrix), as described above.

[0285] Obtain the input matrix

[0286] Spatial audio coding methods

[0287] Any spatial audio coding method that can encode a unit amplitude (virtual) source from a spatial direction into a spatial audio format can be used to generate the appropriate input matrix required for the purposes disclosed herein. This is not intended to be limiting; for example, it includes the most commonly used image localization methods, such as stereo, vector-based amplitude shift (VBAP), distance-based amplitude shift (DBAP), or tri-balanced Atmos. ® Sound image locators, and the most commonly used sound field coding methods, such as Dolby Atmos or SWF.

[0288] As a first example, the VBAP method encodes each virtual source direction as positive coefficients into a maximum of 3 channels. As a second example, each virtual source can be encoded as real coefficients in immersive sound, which is different for each immersive sound channel. The number of channels depends on the order of immersive sound.

[0289] microphone system

[0290] In the most general case, the signal from direction l will be at each frequency. Each of the input channels results in different amplitudes and phases. This is the case with non-overlapping microphones (e.g., microphone arrays), where the signal is picked up by each microphone with different phases and amplitudes at each frequency and direction.

[0291] To capture this behavior, the input matrix... It will be frequency-dependent (i.e., a different matrix is ​​obtained for each correlated frequency band) and will be complex-valued (or alternatively, each entry must contain an amplitude term and a phase term).

[0292] The input matrix of such a system can be obtained empirically, for example by measuring or simulating the microphone signal at frequency f generated from a source at direction l, or by measuring or simulating the impulse response. Then, its corresponding complex-valued frequency domain representation is obtained through time-frequency transformations such as the discrete Fourier transform. In some embodiments, full-resolution frequency domain It can be converted into a smaller number of frequency band values. , Each frequency band contains a complex value, which is a frequency belonging to the corresponding frequency band. The weighted sum.

[0293] It should also be noted that different input matrix terms By using different weights in the cost function The appropriate multiplication and weighting reflect, for example, the fact that some microphones may have a better signal-to-noise ratio than others.

[0294] Cost function optimization

[0295] The optimization objective according to embodiments of this disclosure is to obtain transcoding matrix coefficients that minimize the value of the cost function C. The optimization problem can be equated with the general nonlinear programming problem.

[0296] Several existing methods and corresponding software packages exist in the literature for finding the optimal values ​​of transcoding matrix coefficients. Examples of these methods include gradient descent, the Broyden-Fletcher-Goldfarb-Shanno (BFGS) method, the Newton method, and the Naider-Mead simplex method. Some of these methods (such as gradient descent or BFGS) may require knowledge of the gradient (first derivative) of the cost function; others (such as the Newton method) may require additional knowledge of the Hessian (second derivative) of the cost function. In contrast, some other methods (such as the Naider-Mead simplex method) do not require knowledge of the derivative at all.

[0297] If first or second derivatives are required, they can be calculated manually from the analytical expression of the cost function terms, or more conveniently, automatically from the software expression of the cost function using automatic differentiation software frameworks (such as Jax, PyTorch, or Tensorflow). As a third alternative, numerical differentiation methods can also be used to numerically estimate the derivatives from finite differences.

[0298] Devices, programs and storage media

[0299] Although methods for optimizing transformation matrices (e.g., transcoding matrices) have been described above, it is to be understood that this disclosure also relates to apparatus (e.g., computer apparatus or apparatus generally having processing capabilities) for implementing these methods (or, in general, techniques).

[0300] An example of this device 400 is in Figure 4 The device 400 is schematically illustrated. It includes a processor 410 and a memory 420 coupled to the processor 410. The memory 420 may store instructions for execution by the processor 410. The processor 410 may be adapted to perform the methods described throughout the disclosure (e.g., a method for determining an optimized transformation matrix). The device 400 may receive inputs (e.g., indications of input and / or output schemes, indications of cost function weights, direction weights, etc.) and generate outputs (e.g., an optimized transformation matrix, etc.), as described throughout the disclosure.

[0301] Furthermore, this disclosure also relates to a program including instructions that, when executed by a processor, cause the processor to perform the methods described throughout the disclosure, and this disclosure also relates to a computer-readable storage medium storing such a program.

[0302] explain

[0303] The aspects of the systems described herein can be implemented in a suitable computer-based processing network environment (e.g., a server or cloud environment). A portion of these systems may include one or more networks containing any desired number of individual machines, including one or more routers (not shown) that act as buffers and routers for data transmission between computers. Such networks can be built on a variety of different network protocols and can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0304] One or more of the components, blocks, processes, or other functional components may be implemented by a computer program executed by a processor-based computing device controlling the system. It should also be noted that the various functions disclosed herein can be described and / or described using any number of combinations of hardware and firmware, and embodied in data and / or instructions in various machine-readable or computer-readable media, in terms of their behavior, register transfers, logical components, and other characteristics. Such computer-readable media embodying the formatted data and / or instructions include, but are not limited to, various forms of physical (non-transitory), non-volatile storage media, such as optical, magnetic, or semiconductor storage media.

[0305] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules, which, for the purposes of discussion, may be shown and described as if most components were implemented solely in hardware. However, those skilled in the art, and based on this detailed description, will recognize that in at least one embodiment, the electronic aspects may be implemented in software (e.g., stored on a non-transitory computer-readable medium) and executed by one or more electronic processors (such as microprocessors and / or application-specific integrated circuits (“ASICs”)). Therefore, it should be noted that embodiments may be implemented using a variety of hardware and software-based devices and a variety of different structural components. For example, the apparatus described herein may include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., system buses) connecting the various components.

[0306] While one or more embodiments have been described by way of example and specific implementation, it is to be understood that the one or more embodiments are not limited to the disclosed embodiments. Rather, it is contemplated to cover various modifications and similar arrangements that will be apparent to those skilled in the art. Therefore, the scope of the appended claims should be given the broadest interpretation in order to cover all such modifications and similar arrangements.

[0307] Furthermore, it should be understood that the terminology and wording used herein are for descriptive purposes and should not be considered restrictive. The use of “including,” “comprising,” or “having,” and variations thereof is intended to cover the items listed thereafter and their equivalents, as well as additional items. Unless otherwise specified or limited, the terms “installation,” “connection,” “support,” and “coupled,” and variations thereof, are used extensively and cover direct and indirect installation, connection, support, and coupling.

[0308] Examples and Implementations

[0309] Various aspects and implementations of the present invention can also be understood from the following enumerated exemplary embodiments (EEE), which are not claims.

[0310] EEE 1. A method for determining an optimized transformation matrix, the transformation matrix being used to transform between an input spatial audio format and an output spatial audio format, wherein the input spatial audio format is specified by an encoding scheme for encoding a first set of directions into the input spatial audio format, and the output spatial audio format is specified by a decoding scheme for decoding the output spatial audio format into an output layout, the method comprising:

[0311] The transformation matrix is ​​optimized based on the cost function.

[0312] The cost function is based on the parameters describing the encoding scheme, the parameters describing the decoding scheme, and the transformation matrix.

[0313] EEE 2, the method according to EEE 1, wherein the input spatial audio format is specified by the first set of directions and the encoding matrix, the encoding matrix being used to encode input audio signals from each of the first set of directions into the channels of the spatial input format;

[0314] The output spatial audio format is specified by a second set of directions and a decode-to-speaker matrix, wherein the decode-to-speaker matrix is ​​used to decode the channels of the output spatial audio format to output audio signals for speakers at each direction in the second set of directions; and

[0315] The cost function is based on the encoding matrix, the decoded-to-speaker matrix, and the transformation matrix.

[0316] EEE 3, the method according to EEE 1 or 2, wherein the cost function is based on a system matrix for linking the input audio signal to the output audio signal; and

[0317] The system matrix is ​​based on the transformation matrix.

[0318] EEE 4, the method according to EEE 3 which depends on EEE 2, wherein the system matrix is ​​further based on the encoding matrix and the decoded-to-speaker matrix.

[0319] EEE 5, the method according to EEE 2 or any EEE depending on EEE 2, wherein the cost function is further based on the first set of directions and the second set of directions.

[0320] EEE 6. The method according to any of the preceding EEEs, wherein at least one term of the cost function is based on a psychoacoustic quantity.

[0321] 7. The method according to any of the preceding EEE methods, wherein at least one term of the cost function is based on one or more of the following:

[0322] Sound pressure;

[0323] Sound intensity;

[0324] Sound energy; and

[0325] Speed ​​of sound.

[0326] EEE 8, the method according to EEE 3 or any EEE depending on EEE 3, wherein the cost function includes terms based on a measure of the sparsity or non-sparseness of the system matrix.

[0327] EEE 9. The method according to any of the preceding EEEs, wherein the cost function includes a term indicating the mean square deviation of the sound pressure relative to zero phase.

[0328] EEE 10, the method according to any of the preceding EEEs, wherein the cost function includes a term indicating the mean square deviation of the radial component of the sound velocity relative to zero phase.

[0329] EEE 11, the method according to any of the preceding EEEs, wherein the cost function includes a term indicating the mean square deviation of the transverse component of the sound velocity relative to zero phase.

[0330] EEE 12, the method according to any one of EEE 9 to 11 which depends on EEE 3, wherein the item is based on the system matrix.

[0331] EEE 13, the method according to EEE 3 or any EEE dependent on EEE 3, wherein the cost function comprises quadratic terms with respect to the coefficients of the system matrix.

[0332] EEE 14, the method according to EEE 3 or any EEE dependent on EEE 3, wherein the cost function includes quartic terms with respect to the coefficients of the system matrix.

[0333] EEE 15, the method according to EEE 3 or any EEE dependent on EEE 3, wherein the cost function includes a term penalizing the non-sparseness of the system matrix.

[0334] EEE 16, the method according to any of the preceding EEEs, wherein the cost function includes terms that penalize the larger values ​​of the elements of the transformation matrix.

[0335] EEE 17, the method according to any one of EEE 1 to 16, wherein the cost function includes a term that penalizes elements in a transformation matrix associated with a channel whose noise level exceeds a predetermined threshold.

[0336] EEE 18, the method according to any one of EEE 1 to 17, wherein the encoding scheme is complex-valued.

[0337] EEE 19, the method according to any one of EEE 1 to 17, wherein the encoding scheme is real-valued.

[0338] EEE 20, the method according to any one of EEE 1 to 19, wherein the output layout is related to the physical speaker layout.

[0339] EEE 21, the method according to any one of EEE 1 to 19, wherein the output layout is related to the virtual speaker layout.

[0340] EEE 22. The method described according to any of the preceding EEEs, wherein the output spatial audio format is one of Dolby Atmos, multi-channel format and linear decoding format.

[0341] EEE 23. The method according to any of the preceding EEEs, wherein the input spatial audio format is one of Dolby Atmos, Multichannel format and Linear Encoding format.

[0342] EEE 24, the method according to any one of EEE 1 to 23, wherein the first set of directions is encoded into an input spatial audio format in relation to a set of input audio signals recorded via a microphone array.

[0343] EEE 25, the method according to any one of EEE 1 to 23, wherein encoding the first set of directions into the input spatial audio format involves a one-to-one mapping between the input audio signal and the channels of the input spatial audio format.

[0344] EEE 26, the method according to any one of EEE 1 to 25, wherein decoding the output spatial audio format to the output layout involves a one-to-one mapping between the channels of the output spatial audio format and the output audio signal.

[0345] EEE 27. The method according to any of the preceding EEEs, wherein at least one of the encoding scheme and the decoding scheme is frequency-dependent; and

[0346] The transformation matrix is ​​optimized individually in each of the multiple frequency bands.

[0347] EEE 28, according to the method described in EEE 27, wherein the cost function is frequency-dependent.

[0348] EEE 29, the method according to EEE 27 or 28, wherein the transformation matrix is ​​optimized simultaneously in each frequency band of multiple frequency bands; and

[0349] The cost function includes contributions for each of the multiple frequency bands and one or more terms penalizing the discontinuities in the transformation matrix between adjacent frequency bands.

[0350] EEE 30, the method according to any one of EEE 27 to 29, further includes:

[0351] Determine the effective cost function for a given frequency band.

[0352] The transformation matrix is ​​optimized based on the effective cost function for a given frequency band; and

[0353] The determination of the effective cost function for a given frequency band is based on multiple cost functions at various frequencies between the lower limit frequency and the upper limit frequency of the given frequency band.

[0354] EE 31. An apparatus comprising a processor and a memory, the memory being coupled to the processor and storing instructions for the processor.

[0355] The processor is configured to perform the method according to any one of EEE 1 to 30.

[0356] EEE 32. A program comprising instructions that, when executed by a processor, cause the processor to perform the method described according to any one of EEE1 to 30.

[0357] EEE 33, a computer-readable storage medium storing a program according to EEE 32.

[0358] literature

[0359] Blanco Galindo, M., Coleman, P., & Jackson, PJ (2020). MicrophoneArray Geometries for Horizontal Spatial Audio Object Capture With Beamforming. J. Audio Eng. Soc., Vol. 68, Issue 5, pp. 324-337, May 2020.

[0360] Frank, M. (2013). Phantom Sources using Multiple Loudspeakers in the Horizontal Plane.

[0361] Gerzon, MA (1992). General Metatheory of Auditory Localization. Audio Engineering Society Convention 92.

[0362] Leal, O., & Elmar, H. (1991). Visual Mixing and Stereo Panning Techniques. Audio Engineering Society Convention 91.

[0363] Scaini, D., & Arteaga, D. (2014). Decoding of Higher Order Ambisonics to Irregular Periphonic Loudspeaker Arrays. Audio Engineering Society Conference: 55th International Conference: Spatial Audio.

[0364] Seefeldt, AJ, Lando, JB, & Arteaga, D. (2021). Patent Application Publication No. WO2021021682A1

[0365] Zotter, F., & Frank, M. (2012). All-round ambisonic panning and decoding. J. Audio Eng. Soc., Vol. 6, Issue 10, pp. 807-820, October 2012.

[0366] Zotter, F., & Frank, M. (2018). Ambisonic Decoding with Panning-Invariant Loudness on Small Layouts (AllRAD2). Audio Engineering Society Convention 144.

[0367] Zotter, F., & Frank, M. (2019). Ambisonics: A Practical 3D Audio Theory for Recording, Studio Production, Sound Reinforcement, and Virtual Reality. Springer.

[0368] Hurley, N. & Rickard, S. (2009). Comparing Measures of Sparsity. arXiv:0811.4706v2 [cs.IT].

Claims

1. A method for converting an input audio signal having an input spatial audio format into an output audio signal having an output spatial audio format via a transformation matrix, wherein the transformation matrix is ​​used to perform conversion between the input spatial audio format and the output spatial audio format, wherein, The input spatial audio format is specified by an encoding scheme for encoding a first set of directions into the input spatial audio format, and the output spatial audio format is specified by a decoding scheme for decoding the output spatial audio format into an output layout, the method comprising: Receive the input audio signal; The transformation matrix is ​​optimized based on the cost function. Wherein, the cost function is based on the parameters describing the encoding scheme, the parameters describing the decoding scheme, and the transformation matrix; and The input audio signal is processed using an optimized transformation matrix to output the output audio signal.

2. The method according to claim 1, wherein, The input spatial audio format is specified by the first set of directions and the encoding matrix, the encoding matrix being used to encode the input audio signals from each of the first set of directions into the channels of the spatial input format; The output spatial audio format is specified by a second set of directions and a decode-to-speaker matrix, wherein the decode-to-speaker matrix is ​​used to decode the channels of the output spatial audio format to output audio signals for speakers at each direction in the second set of directions; and The cost function is based on the encoding matrix, the decoded-to-speaker matrix, and the transformation matrix.

3. The method according to claim 1 or 2, wherein, The cost function is based on a system matrix used to link the input audio signal to the output audio signal; as well as The system matrix is ​​based on the transformation matrix.

4. The method according to claim 3, which depends on claim 2, wherein, The system matrix is ​​also based on the encoding matrix and the decoded-to-speaker matrix.

5. The method according to claim 2 or any claim relying on claim 2, wherein, The cost function is also based on the first set of directions and the second set of directions.

6. The method according to any of the preceding claims, wherein, At least one term of the cost function is based on psychoacoustic quantities.

7. The method according to any of the preceding claims, wherein, At least one term of the cost function is based on one or more of the following: Sound pressure; Sound intensity; Sound energy; and Speed ​​of sound.

8. The method according to claim 3 or any claim relying on claim 3, wherein, The cost function includes terms based on a measure of the sparsity or non-sparseness of the system matrix.

9. The method according to any of the preceding claims, wherein, The cost function includes a term indicating the mean square deviation of the sound pressure relative to zero phase.

10. The method according to any of the preceding claims, wherein, The cost function includes a term indicating the mean square deviation of the radial component of the sound velocity relative to zero phase.

11. The method according to any of the preceding claims, wherein, The cost function includes a term indicating the mean square deviation of the lateral component of the sound velocity relative to zero phase.

12. The method according to any one of claims 9 to 11, depending on claim 3, wherein, The item is based on the system matrix.

13. The method according to claim 3 or any claim relying on claim 3, wherein, The cost function includes quadratic terms with respect to the coefficients of the system matrix.

14. The method according to claim 3 or any claim relying on claim 3, wherein, The cost function includes fourth-order terms with respect to the coefficients of the system matrix.

15. The method according to claim 3 or any claim relying on claim 3, wherein, The cost function includes a term that penalizes the non-sparseness of the system matrix.

16. The method according to any of the preceding claims, wherein, The cost function includes terms that penalize the larger values ​​of the elements of the transformation matrix.

17. The method according to any one of claims 1 to 16, wherein, The cost function includes a term that penalizes elements in the transformation matrix associated with channels whose noise levels exceed a predetermined threshold.

18. The method according to any one of claims 1 to 17, wherein, The encoding scheme is complex.

19. The method according to any one of claims 1 to 17, wherein, The encoding scheme is real-valued.

20. The method according to any one of claims 1 to 19, wherein, The output layout is related to the physical speaker layout.

21. The method according to any one of claims 1 to 19, wherein, The output layout is related to the virtual speaker layout.

22. The method according to any of the preceding claims, wherein, The output spatial audio format is one of Dolby, multi-channel, or linear decoding formats.

23. The method according to any of the preceding claims, wherein, The input spatial audio format is one of Dolby, Multichannel, and Linear Encoding formats.

24. The method according to any one of claims 1 to 23, wherein, The first set of directions is encoded into the input space audio format and associated with a set of input audio signals recorded via a microphone array.

25. The method according to any one of claims 1 to 23, wherein, Encoding the first set of directions into the input space audio format involves a one-to-one mapping between the input audio signal and the channels of the input space audio format.

26. The method according to any one of claims 1 to 25, wherein, Decoding the output spatial audio format to the output layout involves a one-to-one mapping between the channels of the output spatial audio format and the output audio signal.

27. The method according to any of the preceding claims, wherein, At least one of the encoding and decoding schemes is frequency-dependent; and The transformation matrix is ​​optimized individually in each of the multiple frequency bands.

28. The method according to claim 27, wherein, The cost function is frequency-dependent.

29. The method according to claim 27 or 28, wherein, The transformation matrix is ​​optimized simultaneously across multiple frequency bands; and The cost function includes contributions for each of the multiple frequency bands and one or more terms penalizing the discontinuities in the transformation matrix between adjacent frequency bands.

30. The method according to any one of claims 27 to 29, further comprising: Determine the effective cost function for a given frequency band. The transformation matrix is ​​optimized based on the effective cost function for a given frequency band; and The determination of the effective cost function for a given frequency band is based on multiple cost functions at various frequencies between the lower limit frequency and the upper limit frequency of the given frequency band.

31. An apparatus comprising a processor and a memory, the memory being coupled to the processor and storing instructions for the processor. in, The processor is configured to perform the method according to any one of claims 1 to 30.

32. A program comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 30.

33. A computer-readable storage medium storing the program according to claim 32.