Adjustment of reverberator based on source directivity

The apparatus and method address the challenge of recreating reverberation based on sound source directionality by using directivity data and room parameters to generate directionally affected reverberant audio signals, enhancing spatial audio perception.

JP2025134998APending Publication Date: 2025-09-17NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025113095
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-03
Filing Date
2025-07-03
Publication Date
2025-09-17

AI Technical Summary

Technical Problem

Existing spatial audio reproduction methods fail to accurately recreate reverberation based on sound source directionality, resulting in a diffuse reverberation perception that lacks environmental specificity.

Method used

An apparatus and method that utilize directivity data to determine averaged gain data, incorporating room parameters and sound source directionality to generate a bitstream for rendering spatial audio, employing digital reverberators to create directionally affected reverberant audio signals.

Benefits of technology

Enhances the perception of spatial audio by accurately recreating reverberation based on sound source directionality, improving the environmental specificity of the audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025134998000001_ABST
    Figure 2025134998000001_ABST
Patent Text Reader

Abstract

To aim to address problems associated with prior arts.SOLUTION: An apparatus for assisting spatial rendering for room acoustics comprises means configured to: obtain directivity data having an identifier, the directivity data comprising data for at least two separate directions; obtain at least one room parameter; determine information associated with the directivity data; determine gain data based on the determined information; determine averaged gain data based on the gain data; and generate a bitstream defining a rendering, the bitstream comprising the averaged gain data and the at least one room parameter such that at least one audio signal associated with the identifier is configured to be rendered based on the at least one room parameter and the determined averaged gain data.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus and method for spatial audio reproduction by adjusting a reverberator based on the directional characteristics of a sound source, but is not limited to an apparatus and method for spatial audio reproduction by adjusting a reverberator based on directional positioning of a sound source in an augmented reality and / or virtual reality device. [Background technology]

[0002] Reverberation refers to the persistence of sound in a space after the actual sound source has stopped. Different spaces have different reverberation characteristics. Reproducing reverberation perceptually accurately is important to convey the spatial impression of an environment. Room acoustics is often represented by statistical models that combine discrete early reflections and diffuse late reverberation. Figure 1 shows an example of a room impulse response in which a direct sound 101 is followed by discrete early reflections 103 with a direction of arrival (DOA) and diffuse late reverberation 105, which can be combined without a specific direction of arrival. The delay d1(t) 102 in Figure 1 can be viewed as the direct sound arrival delay from the sound source to the listener, and the delay d2(t) 104 can be viewed as the delay from the sound source to the listener for one of the early reflections (in this case, the first arriving reflection).

[0003] One way to recreate reverberation is to use N loudspeakers (or virtual loudspeakers recreated binaurally using a set of head-related transfer functions (HRTFs)). The loudspeakers are more or less uniformly positioned around the listener. These loudspeakers reproduce mutually incoherent reverberation signals, resulting in the perception of a diffuse reverberation in the surroundings.

[0004] Reverberation generated by different loudspeakers must be mutually incoherent. In the simple case, reverberation can be generated using different channels of the same reverberator, where the output channels are uncorrelated but have the same acoustic characteristics, such as RT60 time and level (in particular, the diffuse-to-direct or reverberant-to-direct ratio). Such uncorrelated outputs sharing the same acoustic characteristics can be obtained, for example, from the output taps of a feedback-delay-network (FDN) reverberator with appropriately adjusted delay line lengths, or from a reverberator based on the use of uncorrelated noise sequences that are attenuated by using different uncorrelated noise sequences in each channel. In this case, the different reverberation signals effectively have the same characteristics, and the reverberation is generally perceived as similar in all directions. Summary of the Invention [Problem to be solved by the invention]

[0005] The embodiments of the present application aim to solve the problems associated with the prior art. [Means for solving the problem]

[0006] According to a first aspect, there is provided an apparatus for assisting spatial rendering for room acoustics, the apparatus comprising means configured to: acquire directivity data having an identifier, the directivity data including data regarding at least two distinct directions; acquire at least one room parameter; determine information related to the directivity data; determine gain data based on the determined information; determine averaged gain data based on the gain data; and generate a bitstream defining the rendering, the bitstream including the averaged gain data and the at least one room parameter such that at least one audio signal associated with the identifier is configured to be rendered based on the at least one room parameter and the determined averaged gain data.

[0007] The means configured to determine information relating to the directional data may be configured to determine a directional model based on the directional data.

[0008] The directional model may be either a two-dimensional directional model in which at least two directions are arranged on a plane, or a three-dimensional directional model in which at least two directions are arranged in space.

[0009] The means configured to determine the averaged gain data may be configured to determine the averaged gain data based on spatial averaging of gain data independent of the direction and / or orientation of the sound source.

[0010] The means configured to determine information relating to the directional data may be configured to estimate a continuous directional model based on the obtained directional data.

[0011] The means configured to determine the averaged gain data may be configured to determine the gain data based on the determined directivity model and further based on spatial averaging of the gains for at least two distinct directions.

[0012] The means configured to obtain at least one room parameter may be configured to obtain at least one digital reverberator parameter.

[0013] The means configured to determine averaged gain data based on the gain data may be configured to determine frequency dependent gain data.

[0014] The frequency dependent gain data may be graphic equalizer coefficients.

[0015] According to a second aspect, there is provided an apparatus for spatial rendering for room acoustics, the apparatus comprising: means configured to: obtain a bitstream, the bitstream including averaged gain data based on averaging of gain data and an identifier associated with at least one audio signal or the at least one audio signal and at least one room parameter; configure at least one reverberator based on the averaged gain data and the at least one room parameter; and apply the at least one reverberator to the at least one audio signal as at least part of the rendering of the at least one audio signal.

[0016] The at least one room parameter may include at least one digital reverberator parameter.

[0017] The averaged gain data may include frequency dependent gain data.

[0018] The frequency dependent gain data may be graphic equalizer coefficients.

[0019] The averaged gain data may be spatially averaged gain data.

[0020] The means configured to apply at least one reverberator to the at least one audio signal as at least part of the rendering of the at least one audio signal may be further configured to: apply the averaged gain data to the at least one audio signal to generate a directionally affected audio signal; and apply a digital reverberator configured based on the at least one room parameter to the directionally affected audio signal to generate a directionally affected reverberant audio signal.

[0021] The averaged gain data may include at least one set of grouped gains, the grouped gains being grouped because of similar directional patterns.

[0022] Similar directional patterns may include differences between the directional patterns that are less than a determined threshold.

[0023] According to a third aspect, there is provided a method for an apparatus for assisting spatial rendering for room acoustics, the method comprising: acquiring directional data having an identifier, the directional data including data relating to at least two distinct directions; acquiring at least one room parameter; determining information related to the directional data; determining gain data based on the determined information; determining averaged gain data based on the gain data; and generating a bitstream specifying the rendering, the bitstream including the averaged gain data and the at least one room parameter such that at least one audio signal associated with the identifier is configured to be rendered based on the at least one room parameter and the determined averaged gain data.

[0024] Determining information related to the directional data may include determining a directional model based on the directional data.

[0025] The directional model may be either a two-dimensional directional model in which at least two directions are arranged on a plane, or a three-dimensional directional model in which at least two directions are arranged in space.

[0026] Determining the averaged gain data may include determining the averaged gain data based on spatial averaging of gain data that is independent of the direction and / or orientation of the sound source.

[0027] Determining information related to the directional data may include estimating a continuous directional model based on the obtained directional data.

[0028] Determining the averaged gain data may include determining the gain data based on spatial averaging of gains for at least two distinct directions further based on the determined directional model.

[0029] Obtaining at least one room parameter may include obtaining at least one digital reverberator parameter.

[0030] Determining averaged gain data based on the gain data may include determining frequency dependent gain data.

[0031] The frequency dependent gain data may be graphic equalizer coefficients.

[0032] According to a fourth aspect, there is provided a method for an apparatus for spatial rendering for room acoustics, the method comprising: obtaining a bitstream, the bitstream including averaged gain data based on averaging of gain data and an identifier associated with at least one audio signal or the at least one audio signal and at least one room parameter; configuring at least one reverberator based on the averaged gain data and the at least one room parameter; and applying the at least one reverberator to the at least one audio signal as at least part of rendering the at least one audio signal.

[0033] The at least one room parameter may include at least one digital reverberator parameter.

[0034] The averaged gain data may include frequency dependent gain data.

[0035] The frequency dependent gain data may be graphic equalizer coefficients.

[0036] The averaged gain data may be spatially averaged gain data.

[0037] Applying at least one reverberator to the at least one audio signal as at least part of rendering the at least one audio signal may include applying averaged gain data to the at least one audio signal to generate a directionally affected audio signal, and applying a digital reverberator configured based on the at least one room parameter to the directionally affected audio signal to generate a directionally affected reverberant audio signal.

[0038] The averaged gain data may include at least one set of grouped gains, the grouped gains being grouped because of similar directional patterns.

[0039] Similar directional patterns may include differences between the directional patterns that are less than a determined threshold.

[0040] According to a fifth aspect, there is provided an apparatus for assisting spatial rendering for room acoustics, the apparatus comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code, together with the at least one processor, configured to cause the apparatus to at least: acquire directional data having an identifier, the directional data including data for at least two distinct directions; acquire at least one room parameter; determine information related to the directional data; determine gain data based on the determined information; determine averaged gain data based on the gain data; and generate a bitstream specifying the rendering, the bitstream including the averaged gain data and the at least one room parameter, such that at least one audio signal associated with the identifier is configured to be rendered based on the at least one room parameter and the determined averaged gain data.

[0041] The device adapted to determine information relating to the directional data may be adapted to determine a directional model based on the directional data.

[0042] The directional model may be either a two-dimensional directional model in which at least two directions are arranged on a plane, or a three-dimensional directional model in which at least two directions are arranged in space.

[0043] The device adapted to determine the averaged gain data may be adapted to determine the averaged gain data based on spatial averaging of gain data independent of the direction and / or orientation of the sound source.

[0044] The device adapted to determine information relating to the directional data may be adapted to estimate a continuous directional model based on the obtained directional data.

[0045] The device adapted to determine the averaged gain data may be adapted to determine the gain data based on the determined directional model and further based on spatial averaging of the gains for at least two distinct directions.

[0046] The device adapted to obtain at least one room parameter may be adapted to obtain at least one digital reverberator parameter.

[0047] The apparatus adapted to determine averaged gain data based on the gain data may be adapted to determine frequency dependent gain data.

[0048] The frequency dependent gain data may be graphic equalizer coefficients.

[0049] According to a sixth aspect, there is provided an apparatus for spatial rendering for room acoustics, comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code configured by the at least one processor to cause the apparatus to at least: obtain a bitstream, the bitstream containing averaged gain data based on averaging of gain data and an identifier associated with the at least one audio signal or the at least one audio signal and at least one room parameter; configure at least one reverberator based on the averaged gain data and the at least one room parameter; and apply the at least one reverberator to the at least one audio signal as at least part of the rendering of the at least one audio signal.

[0050] The at least one room parameter may include at least one digital reverberator parameter.

[0051] The averaged gain data may include frequency dependent gain data.

[0052] The frequency dependent gain data may be graphic equalizer coefficients.

[0053] The averaged gain data may be spatially averaged gain data.

[0054] The apparatus may be configured to apply at least one reverberator to at least one audio signal as at least part of rendering the at least one audio signal, further applying the averaged gain data to the at least one audio signal to generate a directionally affected audio signal, and applying a digital reverberator configured based on the at least one room parameter to the directionally affected audio signal to generate a directionally affected reverberant audio signal.

[0055] The averaged gain data may include at least one set of grouped gains, where the grouped gains are grouped because of similar directional patterns.

[0056] Similar directional patterns may include differences between the directional patterns that are less than a determined threshold.

[0057] According to a seventh aspect, there is provided an apparatus including: an acquisition circuit configured to acquire directional data having an identifier, the directional data including data relating to at least two distinct directions; an acquisition circuit configured to acquire at least one room parameter; a determination circuit configured to determine information related to the directional data; a determination circuit configured to determine gain data based on the determined information; a determination circuit configured to determine averaged gain data based on the gain data; and a generation circuit configured to generate a bitstream specifying rendering, the bitstream including the averaged gain data and the at least one room parameter, wherein at least one audio signal associated with the identifier is configured to be rendered based on the at least one room parameter and the determined averaged gain data.

[0058] According to an eighth aspect, there is provided an apparatus comprising: an acquisition circuit configured to acquire a bitstream, the bitstream including averaged gain data based on averaging gain data, an identifier associated with at least one audio signal, or the at least one audio signal, and at least one room parameter; a configuration circuit configured to configure at least one reverberator based on the averaged gain data and the at least one room parameter; and an application circuit configured to apply the at least one reverberator to the at least one audio signal as at least part of rendering the at least one audio signal.

[0059] According to a ninth aspect, there is provided a computer program comprising instructions (or a computer-readable medium comprising program instructions) to cause an apparatus to at least: acquire directional data having an identifier, the directional data comprising data relating to at least two distinct directions; acquire at least one room parameter; determine information relating to the directional data; determine gain data based on the determined information; determine averaged gain data based on the gain data; and generate a bitstream specifying rendering, the bitstream comprising the averaged gain data and the at least one room parameter such that at least one audio signal associated with the identifier is configured to be rendered based on the at least one room parameter and the determined averaged gain.

[0060] According to a tenth aspect, there is provided a computer program comprising instructions (or a computer-readable medium comprising program instructions) to cause an apparatus to at least: obtain a bitstream, the bitstream comprising averaged gain data based on averaging of gain data and an identifier associated with at least one audio signal or the at least one audio signal and at least one room parameter; configure at least one reverberator based on the averaged gain data and the at least one room parameter; and apply the at least one reverberator to the at least one audio signal as at least part of rendering the at least one audio signal.

[0061] According to an eleventh aspect, there is provided a non-transitory computer-readable medium comprising program instructions to cause an apparatus to at least: acquire directional data having an identifier, the directional data including data regarding at least two distinct directions; acquire at least one room parameter; determine information related to the directional data; determine gain data based on the determined information; determine averaged gain data based on the gain data; and generate a bitstream specifying rendering, the bitstream including the averaged gain data and the at least one room parameter, such that at least one audio signal associated with the identifier is configured to be rendered based on the at least one room parameter and the determined averaged gain data.

[0062] According to a twelfth aspect, there is provided a non-transitory computer-readable medium comprising program instructions for causing an apparatus to at least: obtain a bitstream, the bitstream including averaged gain data based on averaging of gain data and an identifier associated with at least one audio signal or the at least one audio signal and at least one room parameter; configure at least one reverberator based on the averaged gain data and the at least one room parameter; and apply the at least one reverberator to the at least one audio signal as at least part of rendering the at least one audio signal.

[0063] According to a thirteenth aspect, there is provided an apparatus comprising: means for acquiring directional data having an identifier, the directional data including data for at least two distinct directions; means for acquiring at least one room parameter; means for determining information related to the directional data; means for determining gain data based on the determined information; means for determining averaged gain data based on the gain data; and means for generating a bitstream specifying rendering, the bitstream including the averaged gain data and the at least one room parameter such that at least one audio signal associated with the identifier is configured to be rendered based on the at least one room parameter and the determined averaged gain data.

[0064] According to a fourteenth aspect, there is provided an apparatus comprising: means for obtaining a bitstream, the bitstream including averaged gain data based on averaging of gain data, an identifier associated with at least one audio signal, or the at least one audio signal, and at least one room parameter; means for configuring at least one reverberator based on the averaged gain data and the at least one room parameter; and means for applying the at least one reverberator to the at least one audio signal as at least part of rendering the at least one audio signal.

[0065] According to a fifteenth aspect, there is provided a computer-readable medium comprising program instructions to cause an apparatus to perform at least the following: acquire directional data having an identifier, the directional data including data regarding at least two distinct directions; acquire at least one room parameter; determine information related to the directional data; determine gain data based on the determined information; determine averaged gain data based on the gain data; and generate a bitstream specifying rendering, the bitstream including the averaged gain data and the at least one room parameter such that at least one audio signal associated with the identifier is configured to be rendered based on the at least one room parameter and the determined averaged gain data.

[0066] According to a sixteenth aspect, there is provided a computer-readable medium comprising program instructions to cause an apparatus to at least: obtain a bitstream, the bitstream including averaged gain data based on averaging of gain data and an identifier associated with at least one audio signal or the at least one audio signal and at least one room parameter; configure at least one reverberator based on the averaged gain data and the at least one room parameter; and apply the at least one reverberator to the at least one audio signal as at least part of rendering the at least one audio signal.

[0067] An apparatus comprising means for performing the operations of the methods described above.

[0068] An apparatus configured to perform the operations of the method described above.

[0069] A computer program comprising program instructions for causing a computer to carry out the method described above.

[0070] A computer program product stored on the medium can cause an apparatus to perform the methods described herein.

[0071] The electronic device may comprise an apparatus as described herein.

[0072] The chipset may comprise an apparatus as described herein. [Brief explanation of the drawings]

[0073] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings in which: [Figure 1] Figure 1 shows a model of the room acoustics and the room impulse response. [Figure 2] FIG. 2 illustrates schematically an exemplary device in which some embodiments may be implemented. [Figure 3] FIG. 3 illustrates a flow diagram of the operation of an exemplary device such as that shown in FIG. [Figure 4] FIG. 4 illustrates schematically an exemplary directivity-affected reverberation gain determiner as shown in FIG. 2, according to some embodiments. [Figure 5] FIG. 5 illustrates a flow diagram of the operation of an exemplary directivity affected reverberation gain determiner as shown in FIG. [Figure 6] FIG. 6 illustrates a schematic diagram of an exemplary reverberator such as that shown in FIG. 2 according to some embodiments. [Figure 7] FIG. 7 is a diagram that schematically illustrates an exemplary FDN reverberator such as that shown in FIGS. 2 and 6 in accordance with some embodiments. [Figure 8] FIG. 8 is a schematic diagram of an exemplary FDN reverberator with a directional bus having a directional gain filter as shown in FIG. 6 in accordance with some embodiments. [Figure 9] FIG. 9 shows a flow diagram of the operation of an FDN reverberator with a directional bus with a directional gain filter, as shown in FIG. [Figure 10]FIG. 10 illustrates a schematic diagram of an exemplary device having transmission and / or storage in which some embodiments may be implemented. [Figure 11] FIG. 11 illustrates a schematic diagram of an exemplary system arrangement in which some embodiments may be implemented. [Figure 12] FIG. 12 shows an example of an apparatus suitable for implementing the apparatus shown in the previous figures. DETAILED DESCRIPTION OF THE INVENTION

[0074] In the following, preferred apparatus and possible mechanisms for parameterizing and rendering reverberant audio scenes are described in more detail.

[0075] As mentioned above, reverberation can be represented, for example, using a feedback-delay-network (FDN) reverberator with appropriately adjusted delay line lengths. FDNs allow for independent control of the reverberation time (RT60) and the energy at various frequencies. This allows them to be used to render reverberation based on the characteristics of a room or modeled space. The reverberation time and the energy at various frequencies are affected by the frequency-dependent absorption characteristics of the room.

[0076] Furthermore, the directionality of a sound source affects energy at various frequencies. For example, a human speaker is affected because the human head and body cast an acoustic shadow. This can result in the phenomenon of direct sound being attenuated when listening from behind the speaker more than when listening from directly in front of the speaker. This attenuation is frequency-dependent, since the shadow cast by the head or body is wavelength-dependent. Simply put, a human speaker can be thought of as being nearly omnidirectional at low frequencies (long wavelengths), but highly directional at high frequencies (short wavelengths).

[0077] Directivity also affects late reverberation (also using the human speech example below): At low frequencies, the sound source (human speech) is effectively omnidirectional, so reverberation can be applied directly, with known frequency-dependent energy and reverberation time (typically determined for an omnidirectional sound source).

[0078] However, at high frequencies, considering all sound sources omnidirectional does not result in optimal reverberation quality. While a human talker radiates sound with a "normal" frequency response to the front (audio signals are typically captured with microphones at the front, and therefore the captured audio signal contains this frequency response), the sound is significantly attenuated at the rear at high frequencies. Thus, in practice, the contribution of late reverberation is attenuated at high frequencies due to the directionality of the human talker. In general, the more directional a sound source is at a certain frequency, the less energy it will contribute to the reverberation in the room at that frequency.

[0079] Therefore, when rendering late reverberation, it is necessary to take into account the frequency-dependent reverberation time and energy of the room, as well as the frequency-dependent directivity of the sound source.

[0080] The directivity of a sound source can be obtained in a number of ways. One example is to measure (or model) the amplitude frequency response of the sound source for various directions around the source, and calculate the ratio between these amplitude frequency responses and the amplitude frequency response in the frontal direction. It is therefore possible to describe how much sound is attenuated at various frequencies for these directions.

[0081] If this directional data is available for every direction with infinite resolution, we can choose a spatially uniform (or indeed pseudo-uniform) distribution for all 3D directions (e.g., 100 data points evenly distributed in 3D), calculate their average, and process the reverberation signal with the resulting amplitude-frequency response (or filter it with a corresponding filter).

[0082] However, directional data are rarely available with infinite resolution. They are usually only available for a limited number of directions. Furthermore, the distribution of data points may not be (spatially) uniform. For example, directional data may be available for many directions in front of the sound source, but only for a few directions behind it. In such cases, calculating a simple average value can result in a large bias in the amplitude-frequency response.

[0083] In some cases, an acoustic engineer could manually adjust the appropriate amplitude-frequency response based on available directional information, but this is not possible for an automated system or without some manual (possibly artistic) work by the acoustic engineer or similar.

[0084] Thus, there is a need to be able to determine and render the effect of sound source directionality on the amplitude-frequency response of late reverberation in an efficient manner and without significant or any user input (or interaction).

[0085] The concepts in the embodiments described in further detail herein relate to the reproduction of late reverberant components of a sound source. In these embodiments, an apparatus and method are proposed that allow rendering late reverberation based on directional data of the sound source to take into account spectral effects caused by directivity. In some embodiments, this is achieved by obtaining directional data in multiple different directions, determining whether the directional data is two-dimensional or three-dimensional from a directional model, estimating a spatial domain of the directional data based on the directional model, estimating frequency-dependent directivity-affected reverberation gain data based on the spatial domain, and rendering late reverberation based on the directivity-affected reverberation gain data (and room-related parameters such as audio signal(s) and frequency-dependent reverberation time and energy).

[0086] Furthermore, in some embodiments, sound sources whose determined directivity-affected reverberation gain data are close to each other (the difference is below a threshold) are pooled, an average directivity-affected reverberation gain data is determined for each pool, and the average directivity-affected reverberation gain data can be applied only once to the total audio signals of the pool, thereby improving computational efficiency.

[0087] In some embodiments, the directivity-affected reverberation gain data may be determined at the encoder, which may then transmit them to the decoder (e.g., as graphical equalizer coefficients), which may then apply them when rendering the late reverberation. Furthermore, in some embodiments in situations where there is pooled directivity-affected reverberation gain data, only the average directivity-affected reverberation gain data may be transmitted, and an index to the various average directivity-affected reverberation gain data may be transmitted for each sound source, minimizing the required bit rate for transmission.

[0088] In some embodiments, the directivity-affected reverberation gain data of various audio elements can be collated into a common source directivity pattern with a unique identifier. For example, each audio element (audio object or channel) can have a source directivity pattern with a unique identifier, and all audio sources with the same source directivity pattern can be pooled by the renderer. In other words, the same source directivity pattern identifier can be pooled by the renderer in some embodiments.

[0089] MPEG-I Audio Phase 2 normatively standardizes the bitstream and renderer process. It also provides a reference implementation of the encoder, but allows for changes to the encoder implementation as long as the output bitstream conforms to the normative specification. This allows new encoder implementations to improve the quality of the codec after standardization.

[0090] In some embodiments, the encoder reference implementation is configured to receive an encoder input format description having one or more sound sources with directivity and room-related parameters. In further embodiments, the encoder is configured to estimate frequency-dependent reverberation gain data based on the directivity of the sound sources. Embodiments may then be configured to estimate reverberator parameters based on the room-related parameters. Further, embodiments may be configured to write a bitstream description including the reverberator parameters and the frequency-dependent reverberation gain data.

[0091] Further, in some embodiments, the example bitstream is configured to include frequency-dependent reverberation gain data and reverberator parameters written using the syntax described herein. The example renderer in some embodiments is configured to decode the bitstream to obtain the frequency-dependent reverberation gain data and reverberator parameters, initialize processing components for reverberation rendering using the parameters, and perform reverberation rendering using the initialized processing components using the presented methods.

[0092] Thus, in some embodiments, for VR rendering, the reverberator parameters are derived in the encoder and transmitted in the bitstream. For AR rendering, the reverberator parameters are derived in the renderer based on a Listening Space Description Format (LSDF) file or a corresponding representation. In some embodiments, the source directivity data is available at the encoder. The embodiments described herein do not exclude implementations in which new sound sources are provided directly to the renderer, which would also mean that the sound source directivity data arrives directly at the renderer.

[0093] With reference to FIG. 2, an exemplary implementation of a directionally sensitive reverberator 299 according to some embodiments is shown.

[0094] In some embodiments, the directionally influenced reverberator 299 is configured to receive the directivity data 200, the audio signal 204, and the room parameters 206. The directionally sensitive reverberator 299 is further configured to impart a directionally influenced reverberation to the audio signal 204 based on the room parameters 206 and the directivity data 200, and to output a directionally influenced reverberant audio signal, or a reverberant audio signal incorporating the influence of source directionality (or generally, a reverberant audio signal) 208. These reverberant audio signals 208 may be in any suitable output format. For example, the multi-channel output format may be a 7.1+4 channel system format, a binaural audio signal, or a mono audio signal.

[0095] The directivity data 200 is forwarded to a directivity affected reverberation gain determiner 201. In some embodiments, the directivity data 200 is calculated as gain values ​​g for a number of directions θ(i), φ(i). dir It is of the form (i,k), where i is the index of the data point, k is the frequency, θ is the azimuth angle, and φ is the elevation angle. Optimally, the directions should evenly or uniformly cover the entire sphere around the sound source, but the distribution in some embodiments may be uneven or non-uniform, or may include only a small number of data points.

[0096] In some embodiments, the directivity affected reverberator 299 includes a directivity affected reverberation gain determiner 201. The directivity affected reverberation gain determiner acquires or receives directivity data 200 and determines a directivity affected reverberation gain 202 g that describes how the directivity of the sound source affects the amplitude frequency response of the late reverberation. dir,rev (k). The operation of the directivity affected reverberation gain determiner 201 is presented in more detail later.

[0097] As a result, the directivity-affected reverberation gain 202 is transferred to the reverberator 203 .

[0098] In some embodiments, the directionally affected reverberator 299 comprises a reverberator 203. The reverberator receives the directionally affected reverberation gain 202 and also receives the audio signal 204 s in (t) (where t is time) and room parameters 206 .

[0099] The room parameters can take various forms, for example, in some embodiments, the room parameters 206 include the energy in frequency band k (typically as a diffuse-to-whole ratio DDR or a reverberant-to-direct ratio RDR) and the reverberation time (typically as RT60).

[0100] The reverberator 203 is configured to reverberate the audio signal 204 based on the room parameters 206 and the directionality-affected reverberation gain 202. For example, the reverberator may include an FDN reverberator implementation configured in a manner described in more detail below.

[0101] The resulting directivity-affected reverberant audio signal 208 s rev (j,t), where j is the output audio channel index, is output. The output reverberant audio signal may, in some embodiments, be rendered for a multi-channel loudspeaker setup (such as 7.1+4). This reverberation may be based not only on the directivity data 200 but also on the room parameters 206.

[0102] With reference to FIG. 3, a flow diagram illustrating the operation of the exemplary directionally affected reverberator 299 shown in FIG. 2 is shown.

[0103] In FIG. 3, as indicated by step 301, the first action may be to obtain the audio signal, directivity data, and room (reverberation) parameters.

[0104] Then, as shown by step 303 in FIG. 3, the directivity-affected reverberation gain can be determined.

[0105] In FIG. 3 , as indicated by step 305, after determining the directionally affected reverberation gain, a directionally affected reverberation audio signal is generated from the audio signal based on the directionally affected reverberation gain and room parameters.

[0106] Next, as indicated by step 307 in FIG. 3, the directionally affected reverberant audio signal is output.

[0107] 4 illustrates in more detail the directivity affected reverberation gain determiner 201 according to some embodiments. In some embodiments, the directivity affected reverberation gain determiner 201 is configured to receive directivity data 200. The directivity data 200 in some embodiments includes gains g for directions θ(i), φ(i). dir Contains (i,k).

[0108] In some embodiments, the directivity affected reverberation gain determiner 201 comprises a directivity model determiner 401. The directivity model determiner 401 is configured to analyze the input directional data 200 to determine whether the data is three-dimensional or two-dimensional. This is done by determining the gain p(i)=[x(i),y(i),z(i)] in all directions i provided by the directional data 200. TThis can be implemented by analyzing the axes of an array of orthogonal points. If none of the three axes are all zero, it means that the directional data 200 has three dimensions (3D), and therefore the directional model is three-dimensional. If one of the axes is all zero, the directional data includes data provided in two dimensions (2D), and the directional model is two-dimensional. The resulting directional model 402 information in some embodiments is forwarded to a (spatial) area-weighted gain determiner 403. The area-weighted gain may also be referred to as gain data.

[0109] In some embodiments, the directivity affected reverberation gain determiner 201 comprises a region weighting gain determiner 403. The region weighting gain determiner 403 is configured to receive the directivity model 402 information and divide the total region, such as a sphere (for a 3D model) or a plane (for a 2D model), into sub-regions. The region weighting gain determiner 403 is further configured to receive the directivity data 200 and assign the directivity data to the sub-regions corresponding to the provided directivity values.

[0110] In some embodiments for a three-dimensional (3D) directional model, a spherical Voronoi cover is formed from the orthogonal points p(i) of the gain in all directions. The Voronoi cover divides the sphere into regions near each of the points p(i).

[0111] For each squared gain, the magnitude of the area weight is calculated as follows:

number

[0112] In some embodiments, the total area is calculated by summing the areas of all Voronoi cells.

number

[0113] In some embodiments for a two-dimensional (2D) directional model, the Cartesian coordinates of the centroid of the planar region are calculated as follows:

number

[0114] In the above example, the non-zero Cartesian elements are shown as x and y, indicating that z was all zero. However, this need not be the case, and the method presented here will always get you two axes of non-zero elements, whether x, y, or z.

[0115]

number

[0116] The area of ​​the triangle can then be calculated in some embodiments as follows:

number

[0117] For each squared gain, the magnitude of the region weight is calculated as follows:

number

[0118] Since two-dimensional directivity models are not necessarily circular, this embodiment shows a method that can be used as a general-purpose technique for obtaining an estimated area from directivity data 200 of any two-dimensional shape.

[0119] An alternative to obtaining an area estimate for a 2D shape is to use the arc length, i.e., the angle in radians, from each midpoint (on the circle) between two directional samples to the next midpoint instead of the area of ​​a triangle.

[0120] The resulting spatial domain weighted gain (or domain weighted gain) is 404 g dir,area (i,k) 2 may be forwarded to the average gain determiner 405.

[0121] In some embodiments, the directivity affected reverberation gain determiner 201 includes an average gain determiner 405. The average gain determiner 405 is configured to receive the spatial domain weighted gains and determine the directivity affected reverberation gain 202. In some embodiments, the directivity affected reverberation gain may be determined by calculating an average of the spatial domain weighted gains. For example, in some embodiments, the directivity affected reverberation gain 202 is determined as follows:

number

[0122] In some embodiments, the directivity-affected reverberation gain g dir,rev (k) 202 is the output. The reverberation gain g affected by directivity dir,rev (k) is the original gain data g dir Since (i, k) is spatially averaged over the provided (spatial) direction θ(i)φ(i), it can also be called the averaged gain. The averaged gain data g dir,rev Note that (k) no longer depends on direction, but on frequency k.

[0123] Also, the averaged gain data g dir,rev (k) can be directly represented in the bitstream using the original frequency k and encoded. Alternatively, 20*log10(g dir,rev(k)) can be calculated and converted to decibels. Alternatively or additionally, the averaged gain data can be expressed in other frequency resolutions, such as octave or triple-octave frequencies. As yet another alternative, the averaged gain data can be expressed in terms of the coefficients of a graphic equalizer filter that includes the coefficients of a cascaded filter bank of second-order section IIR filters. Such a filter bank can be designed such that its magnitude response resembles the input command gain in decibels, which can be expressed as 20*log10(g dir,rev (b)), where g dir,rev (b) shows the averaged gain evaluated at the filter bank center frequency (eg, octave center frequency).

[0124] 5, there is a flow diagram illustrating the operation of the directivity affected reverberation gain determiner 201 of the example shown in FIG.

[0125] In FIG. 5, the first action is to obtain directional data, as indicated by step 501.

[0126] In FIG. 5, as indicated by step 503, once the directional data is obtained, it is used to determine a directional model.

[0127] Then, as indicated by step 505 in FIG. 5, spatial domain weighted gains are determined based on the determined directional model and directional data.

[0128] In FIG. 5, as indicated by step 507, after determining the spatial domain weighted gains, the average gain is determined.

[0129] Then, as shown in step 509 in FIG. 5, the reverberation gain (average gain) affected by the directivity is output.

[0130] With reference to FIG. 6, an exemplary reverberator 203 according to some embodiments is shown. The reverberator 203 may be implemented as any suitable, directionally influenced digital reverberator 600 enabled or configured to generate reverberation characteristics consistent with room parameters. The exemplary reverberator implementation includes a feedback delay network (FDN) reverberator and directionally influenced filters that enable reproduction of reverberation with a desired frequency-dependent RT60 time and level. Room parameters 206 are used to adjust the FDN reverberator parameters to generate the desired RT60 time and level. An example of a level parameter may be the direct-to-diffuse ratio (DDR) (or the diffuse-to-total energy ratio, as used in MPEG-I). A directionally influenced reverberation gain 202 is input to the reverberator and applied to the input or output of the reverberator so that the reverberation spectrum (level) is appropriately adjusted according to the source directionality. The input to the directionally influenced FDN reverberator 600 is an audio signal 204, which may be a monophonic, multi-channel, or Ambisonics input. The output from the directionally affected FDN reverberator 600 is a directionally affected reverberated audio signal 208, which in the case of binaural headphone playback is then rendered into two output signals, while a loudspeaker output typically means two or more output audio signals. Regenerating some outputs, such as the FDN delay line output 15, into a binaural output can be done, for example, via HRTF filtering.

[0131] 7 shows in more detail an exemplary directionally influenced FDN reverberator 600 that can be used to generate D uncorrelated output audio signals, in this example each of which can be rendered at a specific spatial location around the listener for an enveloping reverberation effect.

[0132] The exemplary implementation of the directivity-based FDN reverberator 600 processes the reverberation parameters to calculate the coefficients of each attenuation filter 761, GEQ d (GEQ1, GEQ2,... GEQ D), coefficients A and D of the feedback matrix 757, and the length m of the delay line 759 d (m1,m2,···m D ) and a 753-coefficient GEQ based on directivity dir 7. The example FDN reverberator 601 thus provides the D channel output by providing the output from each FDN delay line as a separate output. The example directivity affected FDN reverberator 600 of FIG. 7 is configured to generate a single directivity affected filter GEQ dir 753, although in some embodiments there are multiple such directionally affected filters.

[0133] In some embodiments, each attenuation filter GEQ d 761 is implemented as a graphic EQ filter using M biquad IIR band filters. Thus, for octave bands M=10, the parameters of each graphic EQ include the feedforward and feedback coefficients of the biquad IIR filters, the gains of the biquad band filters, and the overall gain. In some embodiments, any suitable method can be implemented to determine the FDN reverberator parameters; for example, the method described in GB patent application GB2101657.1 can be implemented to derive FDN reverberator parameters to reproduce the desired RT60 time of a virtual / physical scene.

[0134] The reverberator generates a very dense impulse response for the late portion using a network of delays 759, feedback elements (shown as damping filters 761, feedback matrix 757 and combiner 755, and output gain 763). Input samples 751 are input to the reverberator to generate reverberant audio signal components, which can then be output.

[0135] The FDN reverberator includes multiple recirculation delay lines. A unitary matrix A757 is used to control recirculation within the network. In some embodiments, an attenuation filter 761, which may be implemented as a graphic EQ filter implemented as a cascade of second-order section IIR filters, can facilitate control of the rate of energy decay at various frequencies. The filter 761 is designed to attenuate each pulse passing through the delay line by a desired amount in decibels to obtain the desired RT60 time. Thus, the input to the encoder can provide the desired RT60 time for each specified frequency f, denoted as RT60(f). For a frequency f, the desired attenuation per signal sample is calculated as attenuationPerSample(f)=-60 / (samplingRate*RT60(f)). The length m d The attenuation in decibels for a delay line of is attenuationDb(f)=m d *attenuationPerSample(f).

[0136] The attenuation filters are designed as cascade graphic equalizer filters for each delay line, as described in V. Valimaki and J. Liski, “Accurate cascade graphic equalizer,” IEEE Signal Process. Lett., vol. 24, no. 2, pp. 176–180, February 2017. The design procedure outlined in the above-referenced paper takes as input a set of command gains in octave bands. There are also methods for similar graphic EQ structures that support a third octave band, increasing the number of biquad filters to 31 and providing better matching to detailed target responses, as described in “Third-Octave and Bark Graphic-Equalizer Design with Symmetric Band Filters,” https: / / www.mdpi.com / 2076-3417 / 10 / 4 / 1222 / pdf.

[0137] Further, in some embodiments, the design procedure of V. Valimaki and J. Liski, “Accurate cascade graphic equalizer,” IEEE Signal Process. Lett., vol. 24, no. 2, pp. 176-180, February 2017, is a reverberation directional filter GEQ. dir The input to the design procedure is the directivity-affected reverberation gain 202 in decibels.

[0138] The parameters of the FDN reverberator 601 can be adjusted to generate a reverberation whose characteristics match the input room parameters. For this reverberator 601, the parameters are adjusted for each attenuation filter GEQ d The coefficient 761 of the feedback matrix, the coefficient A757 of the feedback matrix, and the length m of the D delay lines 759 d , and the spatial location of the delay line d.

[0139] Also, the directional gain filter 753GEQ dir The coefficients of are obtained based on the directivity-affected reverberation gain 202. In the present invention, each attenuation filter GEQ d and directional gain filter GEQ dir is a graphic EQ filter using M biquad IIR bandpass filters. There are as many directional gain filters GEQ as there are unique directional patterns for the input signal. dir Note that in an embodiment, the number of biquad filters in the various graphic EQ filters may differ, and need not be the same in the delay line attenuation filters and the directionally affected reverberation gain filters.

[0140] The number of delay lines D can be adjusted depending on the quality requirements and the desired tradeoff between reverberation quality and computational complexity. In the present embodiment, an efficient implementation with D=15 delay lines is used. This allows the feedback matrix coefficients A proposed by Rocchesso: "Maximally Diffusive Yet Efficient Feedback Delay Networks for Artificial Reverberation," IEEE Signal Processing Letters, Vol. 4, No. 9, September 1997, to be defined in terms of Galois sequences, which facilitates an efficient implementation.

[0141] The length m of the delay line d d can be determined based on the size of the virtual room. For example, a shoebox (or cuboid) shaped room can be defined with dimensions xDim, yDim, and zDim. If the room is not cubic (or shoebox shaped), a shoebox or cube can be fitted into the room and the size of the fitted shoebox can be used for the delay line length. Alternatively, the dimensions can be taken as the three longest dimensions of a non-shoebox shaped room, or any other suitable method can be used.

[0142] The delay can be set in some embodiments proportional to the standing wave resonant frequency in the virtual or physical room. The delay line length m d may in some embodiments be further configured to be primes of each other.

[0143] FIG. 8 illustrates a schematic diagram of the directivity-affected filter 753 in accordance with some embodiments in more detail. The goal of this example is to group sources with identical or similar directivity patterns so that there can be B directional buses, fewer than the number of sources, S. A simple grouping would combine sources that share the same directivity pattern because they have the same directivity-affected reverberation gain 202. Furthermore, in some embodiments, B can be less than the number of distinct directivity patterns of the S sources. In this case, the grouping method combines sources that have directivity-affected reverberation gains that are close to each other. In some embodiments, closeness can be defined as the average absolute difference in decibels of the directivity-affected reverberation gains or by other suitable metrics, such as the log-spectral distortion of the average directivity pattern.

[0144] The closeness criterion may depend on the available computing power and number of sources. Thus, as computing power decreases, the threshold for combining two sources with similar directional patterns can be increased. As the number of sources increases, the threshold for combining two sources with similar directional patterns can be increased.

[0145] 8, a first set of combiners are shown that receive inputs of audio sources. For example, a first set of sources including audio source 1 8001, audio source 2 8002, and audio source 3 8003 are shown input to a first combiner 8011 (because sources 1, 2, and 3 have similar or the same directivity-affected reverberation gains) and a second set of sources including audio source 4 8004 and audio source 5 8005 are shown input to a second combiner 8012 (because sources 4 and 5 have similar or the same directivity-affected reverberation gains). Furthermore, a Bth combiner 801 B Audio source input to the S-1 800 S-1and audio source S 800 S (Sources S-1 and S have similar or the same directivity-affected reverberation gains)

[0146] The output of each combiner 801 then forms the input of a directional affected filter GEQ. dir ,1 8031 ​​has input 1 8021 and is a second group of directional-influenced filters GEQ dir ,2 8032 has input 2 8022 and Bth group of directional filter GEQ dir ,B 803 B is input B 802 B It has.

[0147] The outputs of the filters 803, which have been influenced by the directivity of each group, are then passed to a combiner 805.

[0148] The directivity affected filter 753 may further comprise a combiner 805 that receives the outputs of the group of directivity affected filters and then combines them to generate the input to the FDN reverberator 601.

[0149] With reference to FIG. 9, a flow diagram illustrating the operation of the directivity affected filter 753 / FDN 601 configuration shown in FIGS. 7 and 8 is shown.

[0150] For example, in FIG. 9, the first action is to obtain the directivity data of the sound source, as indicated by step 901.

[0151] Next, as shown by step 903 in FIG. 9, the directivity-affected reverberation gain for the sound source is calculated or determined.

[0152] Then, as shown by step 905 in FIG. 9, the directivity-affected reverberation gains for the sound sources are compared with each other.

[0153] Then, as shown in step 907 in FIG. 9, if the reverberation gain data affected by directivity is close to the reverberation gain data affected by directivity of another sound source, this sound source and the other sound source are grouped together to configure them to use the reverberation gain data affected by the same directivity.

[0154] 10 is a schematic diagram illustrating an example implementation in which an encoder device is configured to implement some of the functionality of a reverberator. For example, as shown in FIG. 10, the encoder is configured to generate a directivity-affected reverberation gain and write this information together with the audio signal and room parameters into a bitstream to send to a renderer (and / or store this information for later consumption).

[0155] In this embodiment, there are three sound sources as inputs: a first sound source having directivity data 2001 and an audio signal 2041; a second sound source having directivity data 2002 and an audio signal 2042; and a third sound source having directivity data 2003 and an audio signal 2044. q and audio signal 204 q However, there may be any number of input sound sources. The directivity data of each sound source is input to an associated directivity influenced reverberation gain determiner (e.g., a first directivity influenced reverberation gain determiner 2011 associated with the first audio source (directivity data 2001), a second directivity influenced reverberation gain determiner 2012 associated with the second audio source (directivity data 2002), a third directivity influenced reverberation gain determiner 2013 associated with the qth audio source (directivity data 2004), a fourth directivity influenced reverberation gain determiner 2014 associated with the qth audio source (directivity data 2005), a fifth directivity influenced reverberation gain determiner 2016 associated with the qth audio source (directivity data 2006), a sixth directivity influenced reverberation gain determiner 2018 associated with the qth audio source (directivity data 2007), a sixth directivity influenced reverberation gain determiner 2019 associated with the qth audio source (directivity data 2008), a sixth directivity influenced reverberation gain determiner 2020 associated with the qth audio source (directivity data 2009), a sixth directivity influenced reverberation gain determiner 2021 associated with the qth audio source (directivity data 2009), a sixth directivity influenced reverberation gain determiner 2022 associated with the qth audio source (directivity data 2009), a sixth directivity influenced reverberation gain determiner 2023 associated with the qth audio source (directivity data 2009), a sixth directivity influenced reverberation gain determiner 2024 associated with the qth audio source (directivity data 2009), a sixth directivity influenced reverberation gain determiner 2025 associated with the qth audio source (directivity data 2009), a sixth directivity influenced reverberation gain determiner 2026 associated with q ) related to the q-th directivity-affected reverberation gain determiner 201 q )

[0156] Reverberation gain determiners 2011, 2012, and 201 affected by each directivity q are the associated audio signals 2041, 2042, and 204 qand a set of directionally influenced reverberation gains 2021, 2022, and 2023 that can be coded / quantized and combined into a bitstream with room parameters 206 that can then be passed to the reverberator 203. q The apparatus is configured to output:

[0157] In some embodiments, the conversion from room parameters to reverberator parameters is performed by the encoder device, in which case the reverberator parameters are signaled from the encoder to the renderer.

[0158] In some embodiments, a directional influenced filter GEQ dir,j The "room parameters" mapped to the digital reverberator parameters with the filter parameters are described in the bitstream definition below.

[0159] [Table 1]

[0160] The AudioObjectsStruct() example above can be summarized as follows: numberOfAudioObjects specifies the number of audio objects in the audio scene. The id uniquely identifies the audio object within the audio scene. directivityPresentFlag equal to 1 indicates that there is directionality associated with the audio object. If directivityPresentFlag is equal to 1, then directivityId is the directionality profile description identifier for each of the audio objects. Each directionality file present in the audio scene description has a unique directivityId. LocationStruct() provides information about the location of an audio object in the audio scene. This can be provided in any appropriate coordinate system (e.g., Cartesian, Polar, etc.). This data structure can also provide directional information for the audio object, which may be of greater relevance for audio objects that are not omnidirectional point sources.

[0161] [Table 2]

[0162] The AudioChannelSourcesStruct() example above can be summarized as follows: numberOfAudioChannelSources specifies the number of channel sources in the audio scene. numberOfLoudspeakers defines the number of loudspeakers for a particular channel source. id defines the channel source by a unique identifier. channel_index specifies the index of each channel in a given channel source. When directivityPresentFlag is equal to 1, it indicates that the channel has directivity associated with it. directivityId is the directionality profile description identifier for the channels of the channel source whose directivityPresentFlag is equal to 1. This identifier is unique for each of the directionality files present in the audio scene description.

[0163] [Table 3]

[0164] The reverbPayloadStruct() example above can be summarized as follows: numberOfSpatialPositions specifies the number of output delay line positions for the late reverberation payload. This value is specified using an index that corresponds to a specific number of delay lines. A bit string value of "0b00" informs the renderer of 15 spatial direction values ​​for the delay lines. The other three values ​​"0b01", "0b10", and "0b11" are reserved. azimuth specifies the azimuth angle of the delay line relative to the listener, ranging from -180 degrees to 180 degrees. Elevation specifies the elevation angle of the delay line relative to the listener. Range is -90 degrees to 90 degrees. numberOfAcousticEnvironments specifies the number of acoustic environments in the audio scene. reverbPayloadStruct() carries information about one or more acoustic environments currently present in the audio scene. An acoustic environment has specific "room parameters" such as the RT60 time used to obtain the FDN reverberation parameters. environmentId This value specifies a unique identifier for the acoustic environment. delayLineLength Specifies the length, in samples, of the graphic equalizer (GEQ) filter used to set the delay line attenuation filter. The lengths of different delay lines corresponding to the same acoustic environment are prime of each other. filterParamsStruct() This configuration describes the graphic equalizer cascade filters for configuring the delay line attenuation filters. The same configuration is used later to configure the filters for the diffuse to direct reverberation ratio and the reverberation source directional gain. The details of this configuration are explained in the following table.

[0165] The above source-oriented processing example can be summarized as follows. If directivitiesPresent is equal to 1, it indicates the presence of source-directive audio elements in the acoustic environment. If the value is equal to 0, source-directive processing metadata may not be present in the late reverberation metadata. numberOfDirectivities indicates the number of source directivities present in a particular acoustic environment. A directionality can be applied to one or more audio elements within the acoustic environment.

[0166] In some embodiments, the directivitiesPresent flag and associated checks can be skipped: the directionality processing metadata can be present directly.

[0167] The above filterParamsStruct() example can be summarized as follows: SOSLength is the length of each of the second order section filter coefficients. The b1,b2,a1,a2 filter is made up of coefficients b1,b2,a1,a2. These are the feedforward and feedback IIR filter coefficients of a second order section IIR filter. globalGain specifies the gain factor of the GEQ in decibels. levelDB specifies the sound level offset in decibels for each delay line.

[0168] The association between the source directivity profile and the reverberation payload directivity processing metadata in some embodiments is performed by the renderer / decoder. In some embodiments, this may be performed by first checking the associated audio source for a particular acoustic environment (e.g., contained within an acoustic environment range). The associated audio source feeding the reverberation is checked for the presence of source directivity information (e.g., numberOfDirectivities is greater than 0 and directivitiesPresentFlag is equal to 1). The reverberation metadata is then checked for the presence of a corresponding reverbDirectivityGainFilterId. The associated source-directivity filtering is then applied before the audio is fed for late-reverberation rendering.

[0169] As noted above, it can be seen that AudioChannelSourcesStruct and AudioObjectsStruct have a directivityId, and the reverberation metadata payload has a reverbDirectivityGainFilterId. In some embodiments, the directivityId and reverbDirectivityGainFilterId may be the same. In such a scenario, the number of directivityIds in an audio scene corresponding to audio elements in a particular acoustic environment shall be equal to the numberOfDirectivities in the reverberation payload metadata. In other embodiments, if several directivities in an audio scene description (represented by unique directivityIds) are determined to be similar or exceed a threshold and can therefore be clustered or combined using the method shown in FIG. 9, there may be a fewer number of reverbDirectivityGainFilterIds corresponding to fewer GEQs for performing source directivity-related filtering. Such clustering of multiple directivityIds of an audio scene into fewer reverbDirectivityGainFilterIds can be exploited by the renderer to gain greater computational efficiency, as illustrated in the embodiment of FIG. 8.

[0170] In cases where multiple directivityIds are mapped to a single reverbDirectivityGainFilterId, the bitstream may contain additional data structures. In such situations, the clustering to obtain fewer reverbDirectivityGainFilterIds may be performed by the encoder, taking into account the information contained in the bitstream. In another implementation, such remapping may also be implemented by the renderer after performing its own analysis to combine multiple directivityIds into a fewer number of reverbDirectivityGainFilterIds.

[0171]

number

[0172] The above reverbDirectivityGainFilterMappingStruct() example can be summarized as follows: numReverbDirectivityGainFilters is the number of directional gain filters included in the reverberation metadata. reverbDirectivityGainFilterId specifies the GEQ filter identifier for the reverberation source directional gain control. numDirectivityIds specifies the number of source directivityIds that are mapped to one reverbDirectivityGainFilterId.

[0173] Below is an example of metadata that an encoder can provide to allow a renderer to choose, based on the audio scene and computational load, whether to render each directivityId or combine multiple directivities with a single reverberation directivity gain filter (GEQ) specified by a single reverbDirectivityGainFilterId. Thus, this feature gives the renderer flexibility to make runtime decisions.

[0174]

number

[0175] The above reverbDirectivitySimilarityStruct() example can be summarized as follows: numDirectivityIds is the number of directions for which similarity data exists. directivityId is an identifier of a source directivity specified in the audio scene description for one or more audio elements. The similarity_index specifies a numerical value that describes the characteristics of the directivity profile specified by the corresponding directivityId. It can be derived based on the value derived for an appropriate similarity index. An example of such a value is that the gain difference for various frequency bins is less than a preset threshold. The smaller the threshold, the greater the similarity index. That is, if the directivities are identical, the similarity_index will be 255, and if they are most different, the similarity_index will be 0. Other similarity_index measures can be derived based on the requirements of the application. In an embodiment, the similarity_index may be the result of step 905 of FIG. 9.

[0176] In some embodiments, a determination or check of whether a sound source has a directionally affected filter is performed at the initialization of the reverberator instance, and a sound source that has a directionally affected filter may receive a valid pointer to the directionally affected filter instance, and a sound source that is not directionally affected may receive a null pointer as the pointer to the directionally affected filter instance.

[0177] In some embodiments, all filterParamsStruct() are deserialized into GEQ objects in the renderer, forming the association between the directionality and the GEQ. The renderer associates the directionality model of each audio object with the corresponding GEQ that is used to apply filtering to each audio item.

[0178] In some embodiments, the implementation of reverberation directional gain filtering in the renderer can be performed as follows: The input signal has a pointer to the directional filter that initializes the directional filter on bus B. Each directional gain filter has an input bus. The digital reverberator also has an input bus.

[0179] In each rendering loop through the input audio signal to the digital reverberator, the method first sets the input buffers of all directional gain filters to zero, and resets a status flag for each directional filter that indicates whether a signal has been added to the input bus of that directional gain filter.

[0180] When an audio signal is selected to be input to the reverberator, the method first checks whether the audio signal is associated with a directional gain filter. This can be done by checking whether the directional gain filter pointer associated with the input audio signal has a valid value. If so, the method adds the input audio signal to the input bus of the corresponding directional gain filter. A status flag is set to indicate that the audio signal has been added to the input bus of this directional filter. If the pointer is NULL, the input audio signal is added directly to the reverberator input bus.

[0181] When all input audio signals have been added to either the reverberation input bus (without any directivity affected gain filters) or the directivity affected gain filter input bus, the method performs filtering with the directivity affected filters that have at least one audio signal added to their input bus. The method loops through the directivity affected filters, and for each directivity affected gain filter, determines from a status flag whether at least one audio signal has been added to the input bus of that directivity affected gain filter. If at least one audio signal has been added to the input bus, the method performs filtering with that directivity affected gain filter and adds the output of that directivity affected gain filter to the reverberator input bus. Directivity affected filters that have no audio signals added to their input bus may remain unprocessed.

[0182] Finally, a digital reverberator is used to process the reverberator input bus signal and generate an output signal.

[0183] 11 shows an example implementation of the system of the above-described embodiment. The encoder unit 1101 may be implemented, for example, on a suitable content creator's computer and / or on a network server computer.

[0184] The encoder 1101 is configured to receive a virtual scene description 1100 and an audio signal 1102. The virtual scene description may be provided in MPEG-I Encoder Input Format (EIF) or another suitable format. Generally, the virtual scene description includes an acoustically relevant description of the contents of the virtual scene, such as scene geometry as a mesh, acoustic materials, an acoustic environment with reverberation parameters, sound source locations, and other audio element-related parameters, such as whether reverberation is rendered for the audio element. In some embodiments, the encoder 1101 includes a reverberation parameter obtainer 1103 configured to receive the virtual scene description 1100 and obtain reverberation parameters. The reverberation parameters may, in embodiments, be obtained from RT60, DDR, and pre-delay from the acoustic environment.

[0185] The encoder 1101, in some embodiments, further includes a directionality-affected reverberation gain determiner 1105. The directionality-affected reverberation gain determiner 1105 is configured to receive the virtual scene description 1100, more specifically the directionality data of the sound sources it contains, and to generate directionality-affected reverberation gains that can be passed to the directionality-affected reverberation gain combiner 1107 and the reverberation parameter encoder 1108.

[0186] Encoder 1101, in some embodiments, further comprises a directionally affected reverberation gain combiner 1107. The directionally affected reverberation gain combiner 1107 obtains the directionally affected reverberation gains and determines whether any gain grouping should be applied. This information can be passed to reverberation parameter encoder 1108. Combiner 1107 is optional.

[0187] The encoder 1101, in some embodiments, further comprises a directionally affected reverberation parameter encoder 1108. The directionally affected reverberation parameter encoder 1108 in some embodiments is configured to obtain the directionally affected reverberation gains and optionally combiner information and write a bitstream description including reverberator parameters and frequency-dependent reverberation gain data, which can then be output to a bitstream encoder 1109.

[0188] The encoder 1101, in some embodiments, further includes a bitstream encoder 1109 configured to receive the output of the reverberation parameter encoder 1109 and the audio signal and generate a bitstream 1111 that can be passed to the bitstream decoder 1123. In other words, an exemplary bitstream can be configured to include frequency-dependent reverberation gain data and reverberator parameters described using the syntax described herein. The bitstream 1111 in some embodiments can be streamed to an end user device, made available for download, or stored.

[0189] The output of the encoder is a bitstream 1111 that is made available for download or streaming. The decoder / renderer 1121 functionality runs on an end user device which may be a mobile terminal, a personal computer, a sound bar, a tablet computer, a car media system, a home HiFi or theater system, a head mounted display for AR or VR, a smart watch, or any suitable system that uses audio.

[0190] The decoder 1121 in some embodiments includes a bitstream decoder 1123 configured to decode the bitstream to obtain frequency-dependent reverberation gain data and reverberator parameters.

[0191] The decoder 1121 may further comprise a reverberation parameter decoder 1127 configured to obtain the encoded frequency-dependent reverberation gain data and reverberator parameters from the bitstream decoder 1123 and decode them in an opposite or inverse operation to the reverberation parameter encoder 1108.

[0192] In some embodiments, the decoder 1121 comprises a reverberation directivity gain filter generator 1125 that receives the output of the reverberation parameter decoder 1127, generates a gain filter influenced by reverberation directivity, and passes it to the reverberation directivity gain filter 1131.

[0193] In some embodiments, the decoder 1121 includes a reverberation directivity affected gain filter 1131 configured to filter the reverberation affected directional gain and provide input to an FDN reverberator 1133. The FDN reverberator 1133 may be initialized with reverberator parameters provided by the reverberation parameter decoder 1127.

[0194] The decoder 1121 is configured to apply an FDN reverberator 1133 to generate a late-reverberant audio signal that is passed to a head-related transfer function (HRTF) processor 1135 .

[0195] In some embodiments, the decoder 1121 comprises an HRTF processor 1135 configured to apply HRTF processing to the late reverberant audio signal to generate a binaural audio signal, which is output to the binaural signal combiner 1139.

[0196] Further, the decoder / renderer 1121 is configured to receive the decoded audio signal from the bitstream decoder 1123 and comprises a direct sound processor 1129 configured to perform any direct sound processing such as air absorption and distance-gain attenuation, which may be passed to an HRTF processor 1137 which may generate a direct sound component together with a head direction determination, which may be passed to a binaural signal combiner 1139 together with a reverberant component from the HRTF processor 1135. The binaural signal combiner 1139 is configured to combine the direct sound portion and the reverberant sound portion to generate a suitable output (e.g., for headphone playback).

[0197] Furthermore, in some embodiments, the decoder comprises a head direction determiner 1141 that passes head direction information to the HRTF processor 1137 .

[0198] The decoder further comprises a binaural signal combiner configured to take inputs from HRTF processor 1135 and HRTF processor 1137 and generate a binaural audio signal that can be output to a suitable set of transducers, such as a headphone / speaker set. Although not shown, various other audio processing methods may be applied, such as early reflection rendering in combination with the proposed method.

[0199] The MPEG-I Audio Phase 2 standard is designed to normatively standardize the bitstream and renderer process. While some encoders have implementations based on it, they can be changed later as long as the resulting bitstream conforms to the normative specifications. This allows new encoder implementations to improve the quality of the codec even after the standard has been finalized.

[0200] In the main embodiment of the present invention, the parts directed to the different parts of the MPEG-I standard are as follows, with reference to FIG.

[0201] The reference implementation of the encoder includes: Receiving an encoder input format description including a virtual scene description having one or more sound sources with directivity and room-related parameters. Obtaining reverberator parameters from room-related parameters. Determining directivity-influenced reverberation gain. Optionally combine directivity-influenced reverberation gain. o Writing a bitstream description containing reverberator parameters and frequency dependent reverberation gain data.

[0202] The normative bitstream shall contain frequency-dependent reverberation gain data and reverberator parameters described using the syntax described herein. The bitstream shall be streamed to a terminal device or made available for download or storage.

[0203] The reference renderer shall decode the bitstream to obtain frequency-dependent reverberation gain data and reverberator parameters, initialize the processing components for reverberation rendering using the parameters, and perform reverberation rendering using the initialized processing components in the proposed manner. For VR rendering, the reverberator parameters are derived in the encoder and transmitted in the bitstream, as shown in Figure 11. o In AR rendering, the reverberator parameters are derived in the renderer based on a Listening Space Description Format (LSDF) file or a corresponding representation (not shown in Figure 11). The source directivity data is available in the encoder, and there is currently no use case to directly provide new sources to the renderer, however, such a use case may emerge in the future, where directivity-affected reverberation gain decisions may be made in the renderer.

[0204] The full canonical renderer also retrieves other parameters related to room acoustics and source characteristics from the bitstream and uses them to render the diffuse late reverberation as well as the direct sound, early reflections, diffraction, source spatial extent or width, and other acoustic effects. The invention presented here focuses on the rendering of the diffuse late reverberation part, and in particular on how to adjust the diffuse late reverberation spectrum based on the directional characteristics of the source.

[0205] 12 is an exemplary electronic device that can be used as any of the device portions of the systems described above. The device may be any suitable electronic device or device. For example, in some embodiments, device 2000 is a mobile terminal, user equipment, tablet computer, computer, audio playback device, etc. The device may be configured to implement, for example, an encoder or renderer or any of the functional blocks described above.

[0206] In some embodiments, device 2000 includes at least one processor or central processing unit (CPU) 2007. Processor 2007 may be configured to execute various program code, such as the methods described herein.

[0207] In some embodiments, device 2000 comprises a memory (MEM) 2011. In some embodiments, at least one processor 2007 is coupled to memory 2011. Memory 2011 can be any suitable storage means. In some embodiments, memory 2011 includes a program code section for storing program code executable by processor 2007. Additionally, in some embodiments, memory 2011 can further comprise a storage data section for storing data, e.g., data that has been processed or is to be processed according to embodiments described herein. The implemented program code stored in the program code section and the data stored in the storage data section can be retrieved by processor 2007 whenever needed via the memory-processor coupling.

[0208] In some embodiments, device 2000 comprises a user interface (UI) 2005. User interface 2005, in some embodiments, may be coupled to processor 2007. In some embodiments, processor 2007 may control the operation of user interface 2005 and receive input from user interface 2005. In some embodiments, user interface 2005 may allow a user to input commands into device 2000, for example, via a keypad. In some embodiments, user interface 2005 may allow a user to obtain information from device 2000. For example, user interface 2005 may comprise a display configured to display information from device 2000 to a user. User interface 2005, in some embodiments, may comprise a touchscreen or touch interface that can both allow information to be input into device 2000 and also display information to a user of device 2000. In some embodiments, user interface 2005 may be a user interface for communication.

[0209] In some embodiments, device 2000 comprises an input / output port 2009. The input / output port 2009 in some embodiments comprises a transceiver. The transceiver in such embodiments may be coupled to processor 2007 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitting and / or receiving means may, in some embodiments, be configured to communicate with other electronic devices or apparatuses via a wire or a wired connection.

[0210] The transceiver may communicate with the further device by any suitable known communication protocol, for example, in some embodiments the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth®, or an Infrared Data Path (IRDA).

[0211] The input / output port 2009 may be configured to receive a signal.

[0212] In some embodiments, device 2000 may be employed as at least part of a renderer. Input / output port 2009 may be connected to headphones (which may be head-tracked or non-tracked headphones) or the like.

[0213] In general, various embodiments of the present invention may be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Although various aspects of the present invention may be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that these blocks, devices, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controller or other computing device, or any combination thereof.

[0214]

[0013] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile terminal, such as within a processor entity, or by hardware, or by a combination of software and hardware. Further, in this regard, it should be noted that any blocks of logic flows, such as those shown, may represent program steps, or interconnected logic circuits, blocks and functions, or combinations of program steps and logic circuits, blocks and functions. Software may be stored on physical media, such as memory chips or memory blocks implemented within a processor, magnetic media, such as hard disks or floppy disks, and optical media, e.g., DVDs and their data variants, CDs.

[0215] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memory, etc. The data processing device may be of any type suitable for the local technology environment and may include, by way of non-limiting examples, one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), gate-level circuits, and processors based on multi-core processing architectures.

[0216] Embodiments of the present invention can be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert logic-level designs into semiconductor circuit designs ready to be etched onto semiconductor substrates.

[0217] Programs offered by Synopsys, Inc. of Mountain View, California, and Cadence Design, Inc. of San Jose, California, use established design rules and pre-stored libraries of design modules to automatically route conductors and place components on a semiconductor chip. Once the design of a semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII) may be sent to a semiconductor manufacturing facility or "fab" for fabrication.

[0218] The foregoing description provides a complete and informative description of exemplary embodiments of the present invention, by way of illustrative and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the relevant art in light of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of the present invention, as defined by the appended claims.

Claims

1. 1. An apparatus for spatial rendering for room acoustics, the apparatus comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code, together with the at least one processor, causing the apparatus to perform at least: obtaining a bitstream, the bitstream including averaged gain data based on averaging the gain data, an identifier associated with the at least one audio signal, or the at least one audio signal, and at least one room parameter; configuring at least one reverberator based on the averaged gain data and the at least one room parameter; processing the at least one audio signal through the at least one reverberator as at least part of the rendering of the at least one audio signal; configured to cause the the averaged gain data includes frequency dependent gain data; The apparatus is adapted to perform, as at least part of the rendering of the at least one audio signal, processing the at least one audio signal with the at least one reverberator, the apparatus further comprising: applying the averaged gain data to the at least one audio signal to generate a directionally affected audio signal; applying the configured at least one reverberator to the directionally affected audio signal based on the at least one room parameter to generate a directionally affected reverberant audio signal; the averaged gain data includes at least one set of gains grouped based on a common directional pattern; Device.

2. The apparatus of claim 1 , wherein the at least one room parameter includes at least one digital reverberator parameter.

3. The apparatus of claim 1 , wherein the frequency-dependent gain data are graphic equalizer coefficients.

4. The apparatus of claim 1 , wherein the averaged gain data is spatially averaged gain data.

5. The apparatus of claim 1 , wherein the common directional pattern comprises a difference between directional patterns that is less than a determined threshold.

6. The apparatus of claim 1 , wherein the at least one reverberator comprises a feedback-delay-network reverberator.