Apparatus, method, and computer program for enabling spatial audio rendering

By using spatial metadata to determine directional distribution for indirect audio, the system accurately renders both direct and indirect audio components, addressing misclassification issues and enhancing spatial audio immersion.

JP7829707B2Active Publication Date: 2026-03-13NOKIA TECHNOLOGIES OY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-01-11
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing audio rendering technologies struggle with accurately distinguishing between direct and indirect audio sources, leading to incorrect spatial representation, especially when multiple sound sources with similar intensity levels are present, resulting in diffuse audio being misclassified and improperly rendered.

Method used

The system employs spatial metadata to determine directional distribution information for indirect audio, using techniques like parametric encoding and decoding to separate direct and indirect audio components, and applies rendering information to correctly position sound sources in the original sound scene.

Benefits of technology

This approach enhances spatial audio rendering by accurately positioning sound sources, improving the immersive experience by correctly rendering both direct and indirect audio components, even in complex sound scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007829707000037
    Figure 0007829707000037
  • Figure 0007829707000038
    Figure 0007829707000038
  • Figure 0007829707000039
    Figure 0007829707000039
Patent Text Reader

Abstract

Examples of the disclosure relate to an apparatus, method, and computer program for enabling rendering of spatial audio including both direct audio and indirect audio. The apparatus may be configured to obtain a spatial audio signal comprising one or more audio signals and associated spatial metadata. The associated spatial metadata is configured to enable rendering of spatial audio from the one or more audio signals. The spatial audio includes direct audio and indirect audio. The apparatus is also configured to determine directional distribution information for the indirect audio using at least the associated spatial metadata. The apparatus is also configured to determine rendering information corresponding to the determined directional distribution information, and enable rendering of the spatial audio using the determined rendering information, the one or more audio signals, and the associated spatial metadata.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Examples of the present disclosure relate to apparatuses, methods, and computer programs for enabling the rendering of spatial audio. Some relate to apparatuses, methods, and computer programs for enabling the rendering of spatial audio that includes both direct audio and indirect audio.

Background Art

[0002] Spatial audio enables the reproduction of the spatial characteristics of a sound scene for a user so that the user can perceive the spatial characteristics. This can provide an immersive audio experience for the user or can be used for other applications.

Summary of the Invention

[0003] According to various examples of the present disclosure, but not necessarily all, there may be provided an apparatus comprising means for obtaining a spatial audio signal comprising one or more audio signals and associated spatial metadata, the associated spatial metadata being configured to enable the rendering of spatial audio from the one or more audio signals, the spatial audio including direct audio and indirect audio, means for determining directivity distribution information for the indirect audio using at least the associated spatial metadata, means for determining rendering information corresponding to the determined directivity distribution information, and means for enabling the rendering of the spatial audio using the determined rendering information, the one or more audio signals, and the associated spatial metadata.

[0004] Indirect audio may include omnidirectional audio.

[0005] Indirect audio may include diffuse audio.

[0006] The determined directivity distribution information may indicate one or more directions associated with the indirect audio.

[0007] Rendering information may include the target covariance matrix of the audio signal.

[0008] Rendering information may include diffuse audio gain for the channels in a multi-channel loudspeaker configuration.

[0009] The means could be to use at least relevant spatial metadata to directly determine directional information for audio.

[0010] The relevant spatial metadata may include information that enables the mixing of audio signals to allow for the rendering of spatial audio in the selected audio format.

[0011] The associated spatial metadata may include information indicating at least one of sound direction and sound directivity for one or more frequency subbands.

[0012] The relevant spatial metadata may include one or more prediction coefficients for one or more frequency subbands.

[0013] The associated spatial metadata may include one or more coherence parameters.

[0014] In some but not all cases of this disclosure, electronic devices comprising the devices described herein are provided, the electronic devices being at least one of a telephone, a camera, a computing device, or a video conferencing device.

[0015] A method is provided, which includes, but is not necessarily limited to, obtaining a spatial audio signal comprising one or more audio signals and associated spatial metadata, wherein the associated spatial metadata is configured to enable the rendering of spatial audio from one or more audio signals, and the spatial audio comprises the steps of: obtaining direct audio and indirect audio; determining directional distribution information for the indirect audio using at least the associated spatial metadata; determining rendering information corresponding to the determined directional distribution information; and enabling the rendering of spatial audio using estimated target spatial features, one or more audio signals, and associated spatial metadata.

[0016] A computer program is provided that causes a processing circuit to execute a computer program instruction comprising the steps of: obtaining a spatial audio signal including one or more audio signals and associated spatial metadata, wherein the associated spatial metadata is configured to enable the rendering of spatial audio from one or more audio signals, and the spatial audio includes direct audio and indirect audio; determining directional distribution information for the indirect audio using at least the associated spatial metadata; determining rendering information corresponding to the determined directional distribution information; and enabling the rendering of spatial audio using estimated target spatial features, one or more audio signals, and associated spatial metadata.

[0017] Next, we will explain some examples with reference to the attached diagrams. [Brief explanation of the drawing]

[0018] [Figure 1] This is a diagram illustrating an exemplary system. [Figure 2] This diagram illustrates an example method. [Figure 3]This is a diagram illustrating an exemplary decoder. [Figure 4] This is a diagram illustrating an exemplary spatial synthesizer. [Figure 5] This is a diagram illustrating an exemplary apparatus. [Modes for carrying out the invention]

[0019] Examples of this disclosure enable the rendering of spatial audio that includes both direct and indirect audio. In some situations, the finite temporal and / or frequency resolution of audio processing may lead to some audio being identified as indirect or diffuse audio, even if it is not the case in the original sound scene.

[0020] For example, in Directional Audio Coding (DirAC), the direction and spread of audio are analyzed in the frequency band. This analysis is based on intensity. If the original sound scene contains only one primary source within a time-frequency tile, it is likely to be analyzed as non-spread. However, if the original sound scene contains two or more sound sources with similar dominance levels, there may be several time-frequency tiles where sounds from different sources have similar intensity levels. This can lead to the sound being treated as spread audio or indirect audio, even if it is not true in the original sound scene. This can lead to other sounds being analyzed as spread audio in addition to true spread audio.

[0021] An example of this disclosure addresses this problem by estimating directional distribution information for indirect audio based on spatial metadata associated with an audio signal. The directional distribution information can then be used to enable the rendering of the audio signal and to avoid the analysis of non-spread audio as purely spread sound with surround spatial distribution. This leads to improved spatial audio.

[0022] Figure 1 shows an exemplary system 101 that may be used to implement an example of the present disclosure. This system comprises an encoder 105 and a decoder 109. In some examples, the encoder 105 and decoder 109 may be located in different devices. In some examples, the encoder 105 and decoder 109 may be located in the same device.

[0023] System 101 is configured such that encoder 105 acquires an input containing a spatial audio signal 103. In this example, the spatial audio signal 103 may be a primary ambisonics (FOA) signal. Other types of spatial audio signals 103 may be used in other examples of this disclosure.

[0024] The spatial audio signal 103 may be acquired from two or more microphones configured to capture spatial audio. In examples where the audio signal 103 includes an FOA audio signal, the FOA audio signal may be acquired from a dedicated ambisonic microphone such as an Eigenmike or any other suitable means. The spatial audio signal represents a sound scene. A sound scene may include one or more sound sources. In some examples, the spatial audio signal 103 may be acquired from sources other than microphones, for example, they may consist of multi-channel loudspeaker signals such as 5.1 signals.

[0025] The spatial audio signal 103 may include direct audio and indirect audio. Direct audio may include audio reaching the microphone from a specific direction. Direct audio may include audio from one or more dominant sound sources in the sound scene. Indirect audio may include background noise or ambient noise that appears omnidirectional and / or comes from a wide range of directions. Indirect audio may also include audio that may be presumed to be omnidirectional or diffuse, but is not actually diffuse in the original sound scene. For example, if there are multiple sound sources with similar intensity levels, they may be analyzed as omnidirectional, or there may be reflections from one or more sound sources.

[0026] The encoder 105 may include any means by which it can be configured to encode the audio signal 103 and provide a bitstream 107 as an output. The encoder 105 may be configured to use a parametric method to encode the audio signal 103. The parametric method may include an IVAS (Immersive Voice and Audio Services) method or any other suitable type of method. The encoder 105 may be configured to use the audio signal 103 to determine a transport audio signal and spatial metadata. The transport audio signal and spatial metadata may be multiplexed to provide the bitstream 107.

[0027] In some examples, bitstream 107 may be transmitted from a device with encoder 105 to a device with decoder 109. In some examples, bitstream 107 may be stored in the device with encoder 105 and, as appropriate, retrieved and decoded by decoder 109.

[0028] Decoder 109 is configured to receive bitstream 107 as input. Decoder 109 may include means that can be configured to decode bitstream 107. Decoder 109 can decode the bitstream to generate transport audio and spatial metadata. Decoder 109 may be configured to render spatial audio output 111 using the decoded spatial metadata. Spatial audio output 111 may be provided in any suitable format, such as a binaural audio signal.

[0029] Figures 2-4 illustrate exemplary methods and parts of System 101, which enables the determination of directional distribution information for indirect audio and can be used to render improved spatial audio.

[0030] Figure 2 shows an exemplary method that can be used to enable the rendering of spatial audio in different audio formats. This method can be implemented in a system 101, such as the system 101 shown in Figure 1.

[0031] The method includes acquiring a spatial audio signal 103 in block 201. The spatial audio signal 103 may include an encoded audio signal. The audio signal may include multiple channels of audio or a single channel of audio. In some examples, the spatial audio signal 103 may consist of an unencoded audio signal, such as an audio signal based on a signal from a microphone integrated into the device or any other suitable type of audio.

[0032] The spatial audio signal 103 may include one or more audio signals and associated spatial metadata. The associated spatial metadata is configured to enable the rendering of spatial audio from one or more audio signals.

[0033] Spatial metadata may include information that enables the mixing of audio signals to allow for the rendering of spatial audio in the selected audio format.

[0034] Spatial metadata may be provided in frequency subbands. Spatial metadata may include information indicating the direction and directivity of a sound for one or more frequency subbands. Sound directivity can be an indicator of how directional or omnidirectional a sound is. Sound directivity can indicate whether a sound is ambient or from a point source. Sound directivity can be provided as the direct-to-ambient energy ratio or in any other preferred format. In some examples, spatial metadata may include one or more coherence parameters or any other preferred parameters.

[0035] Spatial metadata can be provided in any suitable format corresponding to the format used for the audio signal and / or rendering. For example, if the format used is an FOA signal, the spatial metadata may include information indicating how to predict the FOA signal from the audio signal for one or more frequency subbands. The audio signal may include multiple channels of audio or a single channel of audio. Such information may include prediction coefficients for predicting the FOA signal from the transport audio signal. The transport audio signal may include multiple channels of audio or a single channel of audio. For example, an omnidirectional FOA signal W can be used as the transport audio signal, and the dipole signals X, Y, and Z can be predicted from the transmit signal W using prediction coefficients.

[0036] Spatial audio includes direct and indirect audio. Indirect audio may include omnidirectional audio. In some examples, indirect audio may include diffuse audio. Indirect audio may also include audio that is not truly indirect in the original sound scene but can be analyzed as indirect due to limitations in processing in frequency and / or temporal and / or spatial resolution.

[0037] Indirect audio can include audio that covers a wider range of angles than direct audio. Direct audio can originate from a single angle or a narrow range of angles, while indirect audio can originate from a wide range of angles. In some examples, indirect audio can originate from all directions, rather than from a specific direction.

[0038] In block 203, the method includes the step of using relevant spatial metadata to determine directional distribution information for indirect audio. In some examples, information in addition to spatial metadata may be used to determine directional distribution information.

[0039] The determined directional distribution information indicates one or more directions related to the indirect audio. For example, if the indirect audio includes audio from multiple sources of similar dominance levels, the directional distribution information may include information about the direction of each of the different sources.

[0040] The determined directional distribution information may include information that can be used to adjust the spatialization of the audio in order to render the indirect audio from the correct direction, rather than rendering the indirect audio as ambient or completely ambient.

[0041] In block 205, the method includes determining rendering information corresponding to the determined directional distribution information. The rendering information may include target spatial features. The target spatial features include features that should enable the listener to perceive spatial audio having sound sources in the correct positions corresponding to the original sound scene.

[0042] Rendering information may include parameters that indicate how directional distribution information corresponds to spatial audio features that can be perceived by the listener.

[0043] Rendering information can be determined in any suitable format. The format used for rendering information may depend on the format used for the spatial audio output 111. For example, rendering information may include the target covariance matrix of the audio signal or the diffuse audio gain of the channels in a multi-channel loudspeaker array or any other suitable type of information. Diffuse audio gain can provide an approximation of how diffused the sound is and how the sound should be distributed in space.

[0044] In block 207, the method includes enabling the rendering of spatial audio using determined rendering information, one or more audio signals, and associated spatial metadata.

[0045] In some examples, the method also includes determining directional distribution information for both direct and indirect audio. Spatial metadata and any other suitable information may be used to determine the directional distribution information for direct audio.

[0046] Figure 3 schematically shows an exemplary decoder 109. The exemplary decoder 109 can be installed in a system 101, such as the system in Figure 1. The exemplary decoder 109 may be configured to implement methods such as the method in Figure 2 to enable improved spatial audio rendering.

[0047] Decoder 109 receives bitstream 107 as input. Bitstream 107 may contain spatial metadata and corresponding audio signals. Audio signals may include direct audio and indirect audio.

[0048] Bitstream 107 is supplied to demultiplexer 301. Demultiplexer 301 is configured to demultiplex bitstream 107 into multiple streams. In the example in Figure 3, demultiplexer 301 demultiplexes bitstream 107 into a first stream and a second stream. The first stream contains encoded spatial metadata 303, and the second stream contains encoded transport audio signal 319.

[0049] The encoded transport audio signal 319 is supplied to the transport audio signal decoder 321. The transport audio signal decoder 321 is configured to decode the encoded transport audio signal 319 and provide the decoded transport audio signal 323 as an output. The process used to decode the encoded transport audio signal 319 may include the corresponding process used by the encoder 105 to encode the audio signal. The transport audio signal decoder 321 may comprise an Enhanced Voice Services (EVS) decoder, an Advanced Audio Coding (AAC) decoder, or any other suitable type of decoder.

[0050] The decoded transport audio signal 323 is fed to a time-frequency conversion block 325. The time-frequency conversion block 325 is configured to modify the domain of the decoded transport audio signal 323. In some examples, the time-frequency conversion block 325 is configured to convert the decoded transport audio signal 323 into a time-frequency representation. The time-frequency conversion block 325 may be configured to use any suitable means for modifying the domain of the decoded transport audio signal 323. For example, the time-frequency conversion block 325 may be configured to use a short-time Fourier transform (STFT), a complex modulated quadrature mirror filter (QMF) bank, their low-latency variations, or any other suitable means.

[0051] The time-frequency conversion block 325 provides a time-frequency transport audio signal 327 as its output.

[0052] The encoded first spatial metadata 303 is provided as input to the metadata decoder 305. The metadata decoder 305 is configured to decode the encoded first spatial metadata 303 and provide the decoded spatial metadata 307 as output. The process used to decode the encoded spatial metadata 303 may include the corresponding process used by the encoder 105 to encode the first spatial metadata. The metadata decoder 305 may include any suitable type of decoder.

[0053] The decoded spatial metadata 307 may be provided in any preferred format. For example, if the audio signal is encoded for FOA rendering, the decoded spatial metadata 307 may be in a format that enables FOA rendering. For example, the decoded spatial metadata 307 may include FOA prediction coefficients. The FOA prediction coefficients include information that can be converted into rendering information, such as a mixture matrix. The rendering information or mixture matrix may include any information indicating how the audio signals should be mixed and / or decorrelated to produce an audio output in ambisonic format. In other examples, different types of audio formats may be used.

[0054] The decoded first spatial metadata 307 is provided as input to the rendering information determiner block 309. The rendering information determiner block 309 may be configured to determine rendering information 311. The rendering information 311 may include one or more mixing matrices and / or any other suitable rendering information. The rendering information determiner block 309 provides rendering information 311, such as a mixing matrix, as an output. The rendering information 311 may differ from the rendering information obtained in block 205 of Figure 2, which corresponds to the directional distribution information.

[0055] The rendering information 311 can be determined based on the decoded spatial metadata 307.

[0056] In an example where an audio signal is encoded for FOA rendering, the rendering information 311 may include a mixture matrix. The mixture matrix can be denoted as A(i,j,k,n), where i is the output channel index, j is the input channel, k is the frequency band, and n is the time frame. The mixture matrices can be used to render the FOA signal by applying them to the transport audio signal and / or a decorrelated version of the transport audio signal.

[0057] In other examples, other types of rendering information can be obtained. In some examples, the decoded spatial metadata 307 may already be a mixture matrix or other rendering information, so in such examples, it is not necessary to use the rendering information determiner block 309.

[0058] A mixing matrix or other rendering information may be used to render the decoded time-frequency transport audio signal 327. For example, the decoded time-frequency transport audio signal 327 may be represented as a column vector s(b,n), where b is the frequency bin index, and the rows of the vector represent the transport audio signal channels. The number of rows may be between 1 and 4, depending on the bitrate applied and any other preferred factors. If the number of channels is less than 4, the number of rows in the vector s(b,n) is also less than 4. In such an example, a new channel can be appended to the column vector to form a vector s'(b,n) with 4 rows. The new channel may be a decorrelated version of the first channel signal of s(b,n).

[0059] The FOA signal then, y(b,n)=A(k,n)s'(b,n) This is rendered by the following: where k is the frequency band in which bin frequency b resides. Spatial metadata for the frequency band can correspond to one or more frequency bins of the filter bank used to transform the audio signal.

[0060] In the above equation, the confusion matrix A can be expressed as follows:

number

[0061] This notation means that the time resolution (i.e., metadata time resolution) of the signal s(b,n) and the mixture matrix A(k,n) are the same. This could be the case for system 101 using a filter bank such as STFT configured to apply a coarse time resolution. A coarse time resolution allows for a time step of about 20 milliseconds for the filter bank. Other filter banks may have finer time resolutions. In such cases, the time resolution of the spatial metadata will be sparser than the resolution of the audio signal. In these examples, the same mixture matrix can be applied to multiple different time indices of the audio signal, or the mixture matrix can be temporally interpolated.

[0062] In the examples of this disclosure, the decoded spatial metadata 307 and / or rendering information obtained from the spatial metadata may be used to determine directional distribution information 331 indicating directions related to at least some of the indirect audio.

[0063] In the example shown in Figure 3, rendering information 311 is provided as input to the rendering metadata determiner 313. The rendering metadata determiner 313 also receives a time-frequency transport audio signal 327 as input.

[0064] The rendering metadata determiner 313 is configured to determine rendering metadata 315 suitable for rendering spatial audio in any appropriate format. The format of the rendering metadata 315 may be different from the format used for the decoded spatial metadata 307. For example, the decoded spatial metadata 307 may be suitable for the FOA format, while the rendering metadata 315 may be suitable for other formats such as the binaural format.

[0065] In the example in Figure 3, the rendering metadata 315 includes the direction (azimuth, elevation) θ(k,n), φ(k,n) parameters and the direct-to-total energy ratio r(k,n) parameter. The parameters are provided in frequency bands. Other types of parameters may be used in other examples of this disclosure. Other parameters may be used instead of, or in combination with, the direction and energy ratio parameters.

[0066] In some examples, the decoded spatial metadata 307 may already contain the directional parameters and energy ratio parameters and / or any other appropriate parameters of the rendering metadata 315. In such examples, it is not necessary to use the rendering metadata determinator block 313.

[0067] Rendering information 311 is also provided as input to the directional distribution determiner block 329. The directional distribution determiner block 329 is configured to use rendering information 311 or spatial metadata or other information obtained from spatial metadata to determine the directional distribution information. The directional distribution information can be determined for both indirect and direct audio. In some examples, the directional distribution information may be determined only for indirect audio, and the directional distribution for direct audio may be provided by the directional parameters of the spatial metadata.

[0068] In the example where rendering information 311 includes a mixture matrix A(i,j,k,n), the mixture matrix A(i,j,k,n) is transferred to the directional distribution determiner block 329. In some examples, there may be only one transport audio signal, and the remainder of the audio signal to be processed by the mixture matrix A(i,j,k,n) may be a decorrelated version of the single transport audio signal. In other examples, there may be multiple transport audio signals. In such cases, the energy of the signals must be taken into consideration. When there are multiple transport audio signals, the directional distribution determiner block 329 can also be configured to receive a time-frequency transport audio signal 327, or the covariance matrix of such a signal, as input.

[0069] In the example where FOA type parameters are used in the decoded spatial metadata 307, the first column of the mixture matrix A(i,j,k,n) (i.e., j=1) provided as rendering information 311 corresponds to direct audio, or primarily direct audio. This portion of the audio is coherently rendered from the same input, e.g., an omnidirectional signal W.

[0070] Conversely, the other columns of the mixture matrix A(i,j,k,n) (i.e., 2≦j≦4) correspond to indirect audio, or primarily indirect audio. This portion of the audio is rendered decorrelated or incoherently.

[0071] In the examples of this disclosure, the directional distribution information 331 can be estimated from the columns of the mixing matrix A(i,j,k,n) corresponding to indirect audio. For example, it can be estimated from the last three columns of the mixing matrix A(i,j,k,n).

[0072] The directional distribution information 331 can be determined using any appropriate process. In some examples, the directional distribution information 331 can be estimated by determining the gain for indirect audio. Indirect audio can be diffuse audio. The gain for indirect audio can be determined in the X, Y, and Z directions. In this example, we assume that the columns of the mixing matrix A(i,j,k,n) are in the order X, Y, and Z. In other examples, they may be in a different order.

[0073] The gain is as follows:

number

number

number

[0074] These gains are then the indirect sound gain g sum,diff (k,n) can be added together to determine the sum, resulting in the following equation.

number

[0075] Next, the ratio of indirect audio is determined for each of the X, Y, and Z directions.

number

number

number

[0076] The directional distribution information 331 is the ratio r of indirect audio. x,diff (k,n), r y,diff (k,n), and r z,diff(k,n) may be included. Indirect audio ratio r x,diff (k,n), r y,diff (k,n), and r z,diff (k,n) can be provided as the output of the directional distribution determiner 329.

[0077] Note that this is only one example of parameterization of the directional distribution information 331. Other parameterizations for the directional distribution information 331 may be used in other examples.

[0078] The rendering metadata 315, time-frequency transport audio signal 327, rendering information 311, and directional distribution information 331 are provided as inputs to the spatial combiner 317. The spatial combiner 317 is configured to render the spatial audio output 111 using the rendering metadata 315, time-frequency transport audio signal 327, rendering information 311, and directional distribution information 331. The spatial audio output may consist of a binaural output, a multi-channel speaker signal, a higher-order ambisonic (HOA) signal, or any other suitable audio format.

[0079] Figure 4 schematically shows an exemplary spatial combiner 317. The exemplary spatial combiner 317 can be installed in a decoder 109, such as the decoder 109 shown in Figure 3.

[0080] The spatial synthesizer 317 receives rendering metadata 315, and the time-frequency transport audio signal 327, directional distribution information 331, and rendering information 311 are used as inputs.

[0081] As shown in Figure 4, the time-frequency transport audio signal 327 and rendering information 311 are provided as inputs to the combining input generator 401. The combining input generator 401 is configured to convert the time-frequency transport audio signal 327 into a format suitable for processing by the remaining blocks in the spatial combiner 317.

[0082] The process performed by the synthesizing input generator 401 may depend on the number of transport channels used. In an example where the time-frequency transport audio signal 327 contains a single channel (mono transport), the synthesizing input generator 401 can allow the time-frequency transport audio signal 327 to pass through without performing any processing on it. In some examples, the single-channel signal can be duplicated to generate a signal containing two or more channels. This can provide a dual-mono or pseudo-stereo signal, and as a result, all subsequent blocks in the spatial combiner 317 can assume a stereo track, even though originally there was only one transport audio signal.

[0083] In an example where the time-frequency transport audio signal 327 has multiple channels, the synthesizing input generator 401 may be configured to generate a stereo track. The stereo track can represent a cardioid pattern pointing in different directions, such as left and right.

[0084] The cardioid pattern can be obtained by using any suitable process. For example, it can be obtained by applying the matrix A'(k,n) to the time-frequency transport audio signal 327. Matrix A'(k,n) includes the first two rows of matrix A(k,n). Thus, this provides the W and Y spherical harmonic signals.

[0085] After applying matrix A'(k,n), a left-right cardioid beamforming matrix can be applied to provide a preprocessed transport audio signal 403.

[0086] The preprocessed transport audio signal x(b,n)403 can be expressed as follows:

number

[0087] The preprocessed transport audio signal x(b,n) 403 is provided as the output of the synthesis input generator 401. The covariance matrix determiner may be configured to determine an input covariance matrix and a target covariance matrix.

[0088] The preprocessed transport audio signal 403, the rendering metadata 315, and the directivity distribution information 331 may be provided as inputs to the covariance matrix determiner 411. The covariance matrix determiner 411 is configured to determine an input covariance matrix and a target covariance matrix. The input covariance matrix represents the preprocessed transport audio signal 403, and the target covariance matrix represents the time-frequency spatial audio signal 407.

[0089] The input covariance matrix may be determined from the preprocessed transport audio signal 403 by the following equation.

Equation

[0090] As described above, in this example, the time resolution of the covariance matrix is the same as the time resolution of the audio signal. In other examples, for example, in an example where a filter bank with high time selectivity is used, the time resolution may be different.

[0091] When there are two or more transport audio signals in of s(b,n), the covariance matrix C x (k,n) can be expressed by the following equation.

Equation

Equation

[0092] Note that s'(b,n) was s(b,n) with the decorrelated version of the first channel of s(b,n) added. Therefore, C s’ (k,n) can also be estimated by first estimating the covariance matrix of s(b,n) and then zero-padding it to a size of 4x4. The zero-padding diagonal entries can then be placed in the first diagonal entries of the estimated covariance matrix.

[0093] The target covariance matrix can be determined based on rendering metadata 315 and the overall signal energy and directional distribution information.

[0094] The total signal energy O(k,n) can be obtained as the average of the diagonal values ​​of C_x(k,n), or it can be determined based on the omnidirectional components of the signal A'(k,n)s'(b,n). Then, in some examples, the rendering metadata 315 includes the directions θ(k,n), φ(k,n), and the direct-to-total ratio parameter r(k,n). Assuming the output is a binaural signal, the target covariance matrix is ​​as follows:

number

[0095] Here, h(k,θ(k,n),φ(k,n)) is the head transfer function sequence vector with bandwidth k and direction θ(k,n),φ(k,n).

[0096] h(k,θ(k,n),φ(k,n)) is a column vector. The vector contains two values, which can be complex. These values ​​correspond to the head-related transfer function (HRTF) amplitude and phase of the left and right ears. At high frequencies, the phase difference is not required for perceptual reasons, so the HRTF values ​​can include real values.

[0097] HRTF can be obtained using any appropriate process. HRTF can be obtained for a given direction and frequency.

[0098] In the above formula, C d (k) is the indirect audio binaural covariance matrix. In systems that do not use the implementation of this disclosure, it can be assumed that the indirect audio is diffuse. It can be assumed that the indirect audio is completely ambient and originates from all directions.

[0099] The indirect audio binaural covariance matrix corresponding to the fully ambient and diffuse audio distributions can be obtained by selecting a uniform spatial distribution of direction DOA_d, where d=1..D, and is expressed as follows:

number

[0100] In the examples of this disclosure, indirect audio is not assumed to be always uniform. Indirect audio is not assumed to be perfectly diffuse or originating from all directions. Heterogeneity of indirect audio must be taken into account by the indirect audio binaural covariance matrix. As an example, heterogeneity of indirect audio can be described as follows:

[0101] First, the indirect binaural covariance matrices are determined for X, Y, and Z. These matrices can be determined as follows:

number

number

number

[0102] Here, x(DOA d ), y(DOA d), and z(DOA d ) is DOA d The corresponding Cartesian coordinates. C bin,diff,x (k) includes the contribution of indirect audio mainly originating from the X direction, C bin,diff,y (k) includes the contribution of indirect audio mainly originating from the Y direction, C (bin,diff,z (k) includes the contribution of indirect audio, mainly originating from the Z direction.

[0103] The indirect audio binaural covariance matrix with a directional distribution is given by the diffuse audio ratio r x,diff (k,n), r y,diff (k,n), and r z,diff (k,n), and the diffusion binaural covariance matrix C in the X, Y, and Z directions. bin,diff,x (k), C bin,diff,y (k), C bin,diff,z This can be determined using (k). For example, the indirect audio binaural covariance matrix with a directional distribution can be determined as follows:

number

[0104] Indirect audio ratio r x,diff (k,n), r y,diff (k,n), and r z,diff In the example where (k,n) are equal (i.e.), each ratio is 1 / 3, and the resulting indirect audio binaural covariance matrix C d (k,n) is the uniform diffusion binaural covariance matrix C bin,diff,uniform (k) is a possible possibility. In some cases, C bin,diff,x (k), C bin,diff,y (k), and C bin,diff,z It may be possible to make slight adjustments to the value of (k), and as a result, their average can be exactly C bin,diff,uniform (k). Therefore, rendering of an audio signal with a uniform diffusion distribution (i.e., r x,diff (k,n), r y,diff (k,n), and r z,diffThe case where (k,n) is all 1 / 3 yields the same result as an implementation that does not use the examples of this disclosure.

[0105] The covariance matrix determination unit 411 provides a covariance matrix 413 as its output. The covariance matrix 413 provided as output is equal to the input covariance matrix C x (k,n) and the target covariance matrix C y (k,n) may be included.

[0106] The above equation implies that the processing is performed uniformly within each bin of the bandwidth k. In some examples, the processing is performed with a higher frequency resolution, which is for each frequency bin b. In such examples, the equation given above is adapted so that the covariance matrix is ​​determined for each bin b, but using the parameters of the rendering metadata 315 for the bandwidth k in which the bins reside.

[0107] In some examples, the input covariance matrix and the target covariance matrix can be time-averaged. Time averaging can be implemented using infinite impulse rest number (IIR), finite impulse response (FIR) averaging, or any other preferred type of time averaging. The covariance matrix decisioner 411 may be configured to perform time averaging such that the time-averaged covariance matrix 413 is provided as the output.

[0108] In this example for obtaining the target covariance matrix, only parameters related to direction and energy ratio are considered. In other examples, other parameters may be taken into account when obtaining the target covariance matrix. For example, in addition to direction and energy ratio, spatial coherence parameters, or any other suitable parameters, may be considered. The use of other types of parameters may allow the spatial audio output to be provided in a format other than binaural format and / or improve the accuracy with which the spatial sound can be reproduced.

[0109] The processing matrix decisioner 415 uses the covariance matrix 413C. x (k,n) and C yIt is configured to accept (k,n) as input. The processing matrix decisioner 415 is configured to accept the covariance matrix C x (k,n) and C y Using (k,n), the processing matrix M(k,n) and M r It is configured to determine (k,n). Using any appropriate process, the processing matrix M(k,n) and M r (k,n) can be determined. In some examples, the process used is the target covariance matrix C of the audio signal. y The measured covariance matrix C is used to achieve (k,n). x The method may involve determining a mixing matrix for processing an audio signal having (k,n). Such a method may be used to generate a binaural audio signal, a surround loudspeaker signal, or other types of audio signals. To formulate the processing matrix, the method may involve using a matrix such as a prototype matrix. The prototype matrix is ​​a matrix that indicates, for the sake of optimization, what kind of signal is intended for each output. This may be within the constraint that the output must achieve a target covariance matrix. In the example where the spatial audio output format is a binaural format, the prototype matrix is ​​as follows:

number

[0110] This prototype matrix indicates that the signal for the left ear is primarily rendered from the left preprocessing transport channel, and the signal for the right ear is primarily rendered from the right preprocessing transport channel. In some examples, the orientation of the user's head can be tracked. If it is determined that the user is now facing towards the posterior sphere, the prototype matrix will be as follows:

number

[0111] The processing matrix decisioner 415 may be configured to determine the processing matrix M(k,n) and M_r(k,n) based on the prototype matrix and the input and target covariance matrices, using the means described in Vilkamo, J., Backstroom, T., & Kuntz, A. (2013) "An optimized covariance domain framework for time-frequency processing of spatial audio" "Journal of the Audio Engineering Society", 61(6), 403-411. The processing matrix decisioner 415 determines the processing matrix M(k,n) and M r It is configured to provide (k,n)417 as an output.

[0112] Processing matrix M(k,n) and M r (k,n)417 is provided as input to the decorrelation and mixing block 405. The decorrelation and mixing block 405 also receives the preprocessed transport audio signal x(b,n)403 as input. The decorrelation and mixing block 405 processes the matrix M(k,n) and M r The system may include any means that can be configured to decorrelate and mix the preprocessed transport audio signal x(b,n)403 based on (k,n)417.

[0113] The preprocessed transport audio signal x(b,n)403 can be decorrelated and mixed using any appropriate process. In some examples, the decorrelation and mixing of the preprocessed transport audio signal x(b,n)403 involves processing the preprocessed transport audio signal x(b,n)403 using the same prototype matrix applied by the processing matrix decisioner 415, and the decorrelated signal x D This may include decorrelation the results to generate (b,t). Decorrelation signal x D (b,t) (preprocessed transport audio signal x(b,n)403) can then be mixed using any suitable mixing procedure to generate a time-frequency audio signal 407.

[0114] In some examples, the following mixing procedure may be used to generate the time-frequency audio signal 407.

number

[0115] As mentioned above, the notation used here refers to the processing matrix M(k,n) and M r This means that (k,n)417 and the preprocessed transport audio signal x(b,n)403 have the same time resolution. In other examples, they may have different time resolutions. For example, the time resolution of the processing matrix 417 may be sparser than the time resolution of the preprocessed transport audio signal 403. In such examples, an interpolation process such as linear interpolation can be applied to the processing matrix 417 to achieve the same time resolution as the preprocessed transport audio signal 403. The interpolation rate may depend on any appropriate factor. For example, the interpolation rate may depend on whether a start has been detected. If a start has been detected, fast interpolation can be used; if a start has not been detected, normal interpolation can be used.

[0116] The decorrelation and mixing block 405 provides a time-frequency spatial audio signal 407 as an output.

[0117] The time-frequency space audio signal 407 is provided as input to the inverse filter bank 409. The inverse filter bank 409 is configured to apply an inverse transform to the time-frequency space audio signal 407. The inverse transform applied to the time-frequency space audio signal 407 may be a transform corresponding to the transform used to convert the decoded transport audio signal 323 to the time-frequency transport audio signal 327 in Figure 3.

[0118] The inverse filter bank 409 is configured to provide a spatial audio output 111 as an output. The spatial audio output 111 is provided in any suitable audio format.

[0119] Different examples may use different methods instead of the covariance matrix-based rendering other than the example used in Figure 4. For example, in other examples, an audio signal may be split into a directional portion and an omnidirectional portion. Ratio parameters from spatial metadata may be used to split the signal into a directional portion and an omnidirectional portion. The directional portion may then be positioned to virtual loudspeakers using amplitude panning or any other suitable means. The omnidirectional portion may be distributed to all speakers and decorrelated. The processed directional and omnidirectional portions can then be added together. Each of the virtual speakers may then be processed with an HRTF to obtain a binaural output.

[0120] Therefore, a system implementing an example of this disclosure will not distribute indirect audio equally, but instead will have an indirect audio ratio r x,diff (k,n), r y,diff (k,n), and r z,diff It is possible to provide indirect audio distribution to virtual speakers based on (k,n) and the direction of the virtual speakers.

[0121] For example, the omnidirectional sound gain of a virtual speaker is the spread audio ratio r multiplied by the squared x-coordinate of the virtual speaker. x,diff Multiply (k,n) by the spread audio ratio r and set the squared y coordinate of the virtual speaker. y,diff The spread audio ratio r can be obtained by multiplying (k,n) by the squared z coordinate of the virtual speaker. z,diff Multiply by (k,n) and sum the results. Then, the omnidirectional sound gains of all virtual speakers can be normalized such that the sum of their squares equals 1.

[0122] A similar approach can be used for multi-channel speaker output, in which case the virtual loudspeaker is replaced by the actual loudspeaker.

[0123] In the example above, bitstream 107 contains only a single encoded transport audio signal 319, and mixing is performed using the single transport audio signal and its decorrelated version. As a result, each input signal to the mix has the same energy, and the spread audio ratio r x,diff (k,n), r y,diff (k,n), and r z,diff (k,n) can be calculated without considering energy. Since the energy of the decorrelated channels should correspond to the energy of the transport audio signal from which they were generated (which is the first omnidirectional transport audio signal), there is no need to calculate the energy. The energy value of the transport audio signal can be used instead.

[0124] In some cases, bitstream 107 may contain multiple transport audio signals. In such cases, the spread audio ratio r x,diff (k,n), r y,diff (k,n), and r z,diff Energy must be considered when calculating (k,n). In such an example, the energy E(k,n,j) of the transport audio signal is calculated in the frequency band.

number

[0125] The diffuse audio gain can then be calculated using the energy E(k,n,j) and the mixture matrix A(i,j,k,n).

number

number

number

[0126] It should be noted that these formulas can be used in scenarios where they are used with a single transport audio signal, even if the calculations are more complex.

[0127] The remaining processing can be done as described above. That is, the spread audio ratio can be calculated using these spread audio gains.

[0128] An example of this disclosure may be implemented in a system that enables head tracking of the orientation of a listener's head. In such an example, the matrix A(i,j,k,n) is the spread audio gain g x,diff (k,n), g y,diff (k,n), and g z,diff Before (k,n) is estimated, it can be rotated using a rotation matrix according to the orientation of the listener's head. For example, when the FOA signal is generated from the transport audio signal by y(b,n)=A(k,n)s'(b,n), the rotated FOA signal is as follows:

number

[0129] Here, R(n) is the FOA rotation matrix corresponding to the orientation of the listener's head. The rotation matrix can mix the X,Y,Z channels into new X,Y,Z channels and align them according to the current listener's head position. Therefore, head orientation can be taken into account by using the following rotation mixing matrix instead of A(k,n) in the equation.

number

[0130] In the example above, spatial metadata in the spatial audio signal was provided in FOA format. In other examples, other formats may be used for spatial metadata. In such examples, the spreading audio ratio r x,diff (k,n), r y,diff (k,n), and r z,diff(k,n) can be estimated in different ways. For example, in some cases, spatial metadata may include the direction (azimuth, elevation) θ(k,n), φ(k,n), and direct-to-total energy ratio r(k,n) parameters. This type of spatial metadata can be obtained from a mobile device to which a microphone array is attached, or from any other suitable type of device, or by using any suitable process.

[0131] In such cases, the diffuse audio ratio can be estimated based on the mean direction. This can be done as follows: The direction is transformed into the following Cartesian coordinates. d x (k,n)=cosθ(k,n)cosφ(k,n) d y (k,n)=sinθ(k,n)cosφ(k,n) d z (k,n)=sinφ(k,n)

[0132] Diffuse audio gain can be determined by averaging the absolute values ​​of how the sound is estimated to "spread" over time, weighted by (e.g., 1-r(k,n)).

number

number

number

[0133] In some examples, spatial metadata may consist of multiple different types of parameters, or different types of parameters may be derived from spatial metadata. For example, spatial metadata may include both SPAR (Spatial Audio Rendering) parameters and DirAC (Directional Audio Coding) parameters, or any other appropriate type of parameter. In such cases, the first type of parameter may be used to calculate the diffuse audio ratio, and the second type of parameter may be used for rendering. For example, the diffuse audio ratio can be calculated using the SPAR parameter, and the DirAC parameter can be used for rendering.

[0134] In some examples, spatial metadata may be available in a first format for a first set of frequencies, and spatial metadata may be available in a second format for a second set of frequencies. For example, spatial metadata may comprise SPAR parameters for the first set of frequencies and DirAC parameters for the second set of frequencies. In these cases, different processes may be used to estimate the spread audio ratio for different frequencies. In some cases, different processes may also be used to determine the rendering metadata 315 at different frequencies.

[0135] In the examples shown in Figures 3 and 4, the diffusion audio ratio r x,diff (k,n), ry,diff (k,n), and r z,diff Directional distribution information such as (k,n) is estimated within the decoder 109. In some examples, the spread audio ratio r x,diff (k,n), r y,diff (k,n), and r z,diff Directional distribution information such as (k,n) can be estimated within the encoder 105 and then transmitted to the decoder 109. This directional distribution information may be transmitted along with spatial metadata and any other appropriate parameters.

[0136] In the example above, directional distribution information was acquired separately for the X, Y, and Z axes. In other examples, directional distribution information can be acquired for other coordinate systems. For example, the diffuse component can be determined in a rotating coordinate system such that rotation adaptively maximizes, or substantially maximizes, the energy of the first axial diffuse component. For example, there may be one sound source at an azimuth of 45 degrees and an elevation of 45 degrees, and another sound source at an azimuth of -135 degrees and an elevation of -45 degrees (i.e., in opposite directions). In this case, one embodiment measures and reproduces the diffuse component primarily on the axes, rather than focusing on fixed X, Y, and Z coordinates.

[0137] In some cases, the composition of the diffusion component does not have to follow an arbitrary set of rotational or non-rotating axes. For example, if the FOA covariance matrix is ​​determined, it is possible to determine the spatial energy spectrum in the spatial distribution in the surrounding directions using the minimum variance-free response (MVDR) beamforming method, based on the measurement of the signal covariance matrix y(b,n)=A(k,n)s'(b,n). To do this, the energy from the d-th direction (i.e., the direction of arrival d(DOA)) is used. d The sound arriving from () can be represented as E'(d,k,n). These energy values ​​can be used to weight the following diffuse binaural covariance matrix in space.

number

[0138] Even when the method for determining the diffuse audio distribution follows a coordinate system such as X, Y, Z or a rotated one, the ratio value g is first used. (x,y,z),diff By mapping (k,n) (or a similar rotated value) to the spatial energy distribution value E'(d,k,n), and then applying the above formula, C d The above formula can be used to determine (k,n).

[0139] In the example above, the decorrelated sound was generated by decorrelating the left and right preprocessed transport audio signals 403 and mixing them to obtain the residual component. In the case of mono, where a single transport audio signal exists, this means that the mono sound is decorrelated into left and right decorrelated sounds. The left and right decorrelated sounds are mixed using a covariance matrix-based rendering scheme. This can be assumed to be a diagonal matrix for the input covariance matrix (of the decorrelated portion).

[0140] In some examples, a monaural transport audio signal may be decorrelated just once, and a two-channel signal may be generated by providing the decorrelated signal to the first channel and the inverted (multiplied by -1) decorrelated signal to the second channel. In such cases, the input covariance matrix of the decorrelation portion is not diagonal, but the cross-terms of the covariance matrix are the same as the diagonal values ​​but have a negative sign. This procedure can be used in situations where a decorrelator itself is lacking when generating a signal that can be assumed to be perfectly incoherent.

[0141] In some examples, the diffusion binaural covariance matrix can be generated based on the estimated FOA covariance matrix. In such examples, the transport signal covariance matrix is as follows.

Number

[0142] Next, C s (k,n) is zero-padded to size 4×4, and the value of the first diagonal value is placed on the zero-padded diagonal values to obtain the padded matrix C s’ (k,n). The FOA covariance matrix is as follows.

Number

[0143] The FOA matrix can be rotated by the following equation.

Number

[0144] Next, the diffusion binaural covariance matrix can be determined as the following equation.

Number

[0145] Here, H FOA (k) is the FOA pair binaural processing matrix. In this method, when the spatial metadata indicates that the audio is indirect audio, the target covariance matrix is mainly based on the FOA-binaural rendering scheme. However, when the spatial metadata indicates a sound that is direct audio, the target covariance matrix is mainly based on the rendering metadata consisting of the direction parameters. This covariance matrix C d (k,n) can also be estimated by first generating the following signals.

Number

[0146] Figure 5 schematically shows an exemplary apparatus 501 that may be used in some examples of the present disclosure. Apparatus 501 may include a controller device, which may be located within an electronic device such as a telephone, camera, computing device, video conferencing equipment, or any other suitable type of device.

[0147] In the example shown in Figure 5, the device 501 comprises at least one processor 503 and at least one memory 505. It should be understood that the device 501 may include additional components not shown in Figure 5.

[0148] In the example shown in Figure 5, the device 501 can be implemented as a processing circuit. In some examples, the device 501 may be implemented as hardware only, or it may have several forms in software including firmware only, or it may be a combination of hardware and software (including firmware).

[0149] As shown in Figure 5, the device 501 may be implemented by using executable instructions of a computer program 507 in a general-purpose or dedicated processor 503, which can be stored in a computer-readable storage medium (disk, memory, etc.) executed by such a processor 503, using instructions that enable hardware functionality.

[0150] The processor 503 is configured to read from and write to memory 505. The processor 503 may also have an output interface to which data and / or commands are output by the processor 503, and an input interface to which data and / or commands are input to the processor 503.

[0151] Memory 505 is configured to store a computer program 507, which includes computer program instructions (computer program code 509) that control the operation of the device 501 when loaded into the processor 503. The computer program instructions of computer program 507 provide logic and routines that enable the device 501 to perform the actions shown in Figures 2 to 4. The processor 503 can load and execute computer program 507 by reading from memory 505.

[0152] Accordingly, the device 501 includes at least one processor 503, at least one memory 505 including computer program code 509, and computer program code 509 configured to cause the device 501 to perform the steps of: obtaining a spatial audio signal including at least one or more audio signals and associated spatial metadata using at least one memory 505 and at least one processor 503, wherein the associated spatial metadata is configured to enable the rendering of spatial audio from one or more audio signals, and the spatial audio includes direct audio and indirect audio; determining directional distribution information for indirect audio using at least the associated spatial metadata; determining rendering information corresponding to the determined directional distribution information; and enabling the rendering of spatial audio using the determined rendering information, one or more audio signals, and associated spatial metadata.

[0153] As shown in Figure 5, the computer program 507 may be delivered to the device 501 via any suitable delivery mechanism 511. The delivery mechanism 511 may be, for example, a machine-readable medium, a computer-readable medium, a non-temporary computer-readable storage medium, a computer program product, a memory device, a recording medium such as a compact disc read-only memory (CD-ROM) or a digital versatile disc (DVD), or a solid-state memory, or a manufactured product containing or tangibly embodying the computer program 507. The delivery mechanism may be a signal configured to reliably transfer the computer program 507. The device 501 may propagate or transmit the computer program 507 as a computer data signal. In some examples, the computer program 507 may be transmitted to the device 501 using a radio protocol such as Bluetooth®, Bluetooth® Low Energy, Bluetooth® Smart, 6LoWPan (IPv6 over Low Power Personal Area Network), ZigBee, ANT+, Near Field Communication (NFC), Radio Frequency Identification, Wireless Local Area Network (Wi-Fi), or any other suitable protocol.

[0154] The computer program 507 includes the steps of: obtaining a spatial audio signal from the device 501, which includes at least one or more audio signals and associated spatial metadata, wherein the associated spatial metadata is configured to enable the rendering of spatial audio from one or more audio signals, and the spatial audio includes direct audio and indirect audio; determining directional distribution information for the indirect audio using at least the associated spatial metadata; determining rendering information corresponding to the determined directional distribution information; and enabling the rendering of spatial audio using the determined rendering information, one or more audio signals, and associated spatial metadata; and computer program instructions to perform the following:

[0155] Computer program instructions may be included in computer programs 507, non-temporary computer-readable media, computer program products, and machine-readable media. In some, but not all, computer program instructions may be distributed across multiple computer programs 507.

[0156] Although memory 505 is shown as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which may be integrated / removable and / or may provide permanent / semi-permanent / dynamic / cache storage.

[0157] Although the processor 503 is illustrated as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which can be integrated / removed. The processor 503 can be a single-core or multi-core processor.

[0158] References to "computer-readable storage media," "computer program products," "tangibly embodied computer programs," or to "controllers," "computers," and "processors" should be understood to encompass not only computers with different architectures such as single / multiprocessor architectures and sequential (Von Neumann) / parallel architectures, but also specialized circuits such as field-programmable gate arrays (FPGAs), application-specific circuits (ASICs), signal processing devices, and other processing circuits. References to computer programs, instructions, code, etc., should be understood to encompass software or firmware for programmable processors, such as programmable content for hardware devices, or configuration settings for fixed-function devices, gate arrays, or programmable logic devices, regardless of whether they are instructions for the processor.

[0159] As used in this application, the term "circuit" may refer to one, more, or all of the following: (a) Hardware-only circuit implementation (such as implementation in analog and / or digital circuits only) (b) A combination of hardware circuitry and software, for example (where applicable) (i) combination of analog and / or digital hardware circuits with software / firmware (ii) Any part of a hardware processor having software (including a digital signal processor), software, and memory(s) that works together to cause a device such as a mobile phone or server to perform various functions. (c) A microprocessor or part of a microprocessor that requires software (e.g., firmware) to operate, but the software may not be present when not needed for operation, hardware circuitry and / or processor

[0160] This definition of circuit applies to all use of the term in this application, including any claim. As a further example, as used in this application, the term circuit also encompasses implementations of hardware circuits or processors and their associated software and / or firmware. The term circuit also encompasses, for example, a baseband integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or network device, as applicable to the elements of a particular claim.

[0161] The blocks shown in Figures 2 to 4 may represent steps of a method and / or sections of code within the computer program 507. The examples of specific orders for the blocks do not necessarily suggest that there is a required or preferred order for the blocks, and the order and arrangement of the blocks can be changed. Furthermore, some blocks may be omitted.

[0162] The term “comprise” is used herein in a comprehensive, not exclusive, sense. That is, any reference to X containing Y indicates that X may contain only one Y or multiple Ys. If “comprise” is intended to be used in an exclusive sense, it will be evident in the context by referring to “comprising only one…” or by using “consisting.”

[0163] This explanation refers to various examples. Descriptions of features or functions relating to an example indicate that those features or functions are present in that example. The use of the terms “example,” “for example,” “can,” or “may” in the text, whether explicitly stated or not, indicates that such features or functions are present in at least the example being described, whether or not they are described as examples, and that they may, but not necessarily, be present in some or all other examples. Thus, “example,” “for example,” “can,” or “may” refers to a specific instance in the example class. Properties of an instance may be properties of that instance only, or properties of the class, or properties of a subclass of the class that includes some, but not all, instances within that class. Thus, it is implicitly disclosed that features described in one example but not in another may, where possible, be used as part of a working combination in other examples, but are not necessarily required to be used in other examples.

[0164] Examples have been described in the preceding paragraphs with reference to various embodiments, but it should be understood that modifications to the given embodiments may be made without departing from the scope of the claims.

[0165] The features described above may be used in combinations other than those explicitly described above.

[0166] While the functions are described by referring to specific features, these functions may be performable by other features, whether or not they are described.

[0167] While we have described the features by referring to specific examples, these features may also exist in other examples, whether or not they are described.

[0168] The terms "a" or "the" are used herein in an inclusive, not exclusive, sense. That is, any reference to X containing "a" / "the" Y indicates that X may contain only one Y or more Y, unless the context explicitly indicates otherwise. If "a" or "the" is intended to be used in an exclusive sense, it will be evident from the context. In some situations, the use of "at least one" or "one or more" may be used to emphasize the inclusive meaning, but the absence of these terms should not be interpreted as inferring any exclusive meaning.

[0169] The presence of a feature (or combination of features) in a claim is a reference to the feature or (combination of features) itself, and to features (equivalent features) that achieve substantially the same technical effect. Equivalent features include, for example, variations of features that achieve substantially the same result in substantially the same way. Equivalent features include, for example, features that perform substantially the same function in substantially the same way to achieve substantially the same result.

[0170] This explanation refers to various examples in which adjectives or adjective phrases are used to describe the characteristics of the examples. Such descriptions of characteristics of the examples indicate that the characteristics are present in some examples exactly as described and in others substantially as described.

[0171] While the foregoing specification attempts to draw attention to these features considered important, it should be understood that the applicant may seek protection through the claims with respect to any patentable feature or combination of features mentioned herein and / or shown in the drawings, whether or not they are emphasized.

Claims

1. A device comprising at least one processor and at least one memory containing computer program code, wherein the at least one memory and the computer program code are connected by the at least one processor. A step of obtaining a spatial audio signal including one or more audio signals and associated spatial metadata, wherein the associated spatial metadata is configured to enable rendering of spatial audio from the one or more audio signals, and the spatial audio includes direct audio and indirect audio, The steps include determining directional distribution information for the indirect audio using at least the relevant spatial metadata, A step of determining rendering information corresponding to the determined directional distribution information for the indirect audio, An apparatus for performing the steps of enabling the rendering of spatial audio using the determined rendering information, the one or more audio signals, and the associated spatial metadata.

2. The apparatus according to claim 1, wherein the indirect audio includes omnidirectional audio.

3. The apparatus according to claim 1, wherein the indirect audio includes diffuse audio.

4. The apparatus according to claim 1, wherein the determined directional distribution information indicates one or more directions related to the indirect audio.

5. The apparatus according to claim 1, wherein the rendering information includes the target covariance matrix of the audio signal.

6. The apparatus according to claim 1, wherein the rendering information includes diffuse audio gain for channels in a multi-channel loudspeaker arrangement.

7. The apparatus according to claim 1, which determines directional information for the direct audio using at least the relevant spatial metadata.

8. The apparatus according to claim 1, wherein the associated spatial metadata includes information that enables mixing of audio signals to enable rendering of the spatial audio in a selected audio format.

9. The apparatus according to claim 1, wherein the associated spatial metadata includes information indicating at least one of sound direction and sound directivity for one or more frequency subbands.

10. The apparatus according to claim 1, wherein the associated spatial metadata includes, for one or more frequency subbands, one or more prediction coefficients and at least one of one or more coherence parameters.

11. A step of obtaining a spatial audio signal including one or more audio signals and associated spatial metadata, wherein the associated spatial metadata is configured to enable rendering of spatial audio from the one or more audio signals, and the spatial audio includes direct audio and indirect audio, The steps include determining directional distribution information for the indirect audio using at least the relevant spatial metadata, A step of determining rendering information corresponding to the determined directional distribution information for the indirect audio, A method comprising the step of enabling the rendering of the spatial audio using the determined rendering information, the one or more audio signals, and the associated spatial metadata.

12. The method according to claim 11, wherein the indirect audio includes at least one of omnidirectional audio and diffuse audio.

13. The method according to claim 11, wherein the determined directional distribution information indicates one or more directions related to the indirect audio.

14. The method according to claim 11, wherein the rendering information includes the target covariance matrix of the audio signal.

15. The method according to claim 11, wherein the rendering information includes diffuse audio gain for channels in a multi-channel loudspeaker arrangement.

16. The method according to claim 11, wherein the step of using at least the relevant spatial metadata includes the step of determining directional information for the direct audio.

17. The method according to claim 11, wherein the associated spatial metadata includes information that enables mixing of audio signals to enable rendering of the spatial audio in a selected audio format.

18. The aforementioned related spatial metadata is Information indicating at least one of sound direction and sound directivity for one or more frequency subbands, and For one or more frequency subbands, one or more prediction coefficients, The method according to claim 11, comprising at least one of the following.

19. The aforementioned related spatial metadata is For one or more frequency subbands, one or more prediction coefficients, and One or more coherence parameters, The method according to claim 11, comprising at least one of the following.

20. A step of obtaining a spatial audio signal including one or more audio signals and associated spatial metadata, wherein the associated spatial metadata is configured to enable rendering of spatial audio from the one or more audio signals, and the spatial audio includes direct audio and indirect audio, The steps include determining directional distribution information for the indirect audio using at least the relevant spatial metadata, A step of determining rendering information corresponding to the determined directional distribution information for the indirect audio, A computer program that causes a processing circuit to execute a computer program instruction, which includes the step of enabling the rendering of the spatial audio using the determined rendering information, the one or more audio signals, and the associated spatial metadata.

Citation Information

Patent Citations

  • Audio signal processor, system and method for distributing an ambient signal among multiple ambient signal channels

    JP2021512570A

  • Spatial audio representation and rendering

    WO2021240053A1