Adjustment of the reverberator based on the directivity of the sound source
The apparatus and method address the issue of uniform reverberation perception by using directivity data to adjust reverberation gain, ensuring spatially accurate and directionally varied reverberation reproduction.
Patent Information
- Application Number
- JP2022192647
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-03
- Filing Date
- 2022-12-01
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-01
AI Technical Summary
Existing spatial audio reproduction systems fail to accurately reproduce reverberation characteristics based on the directivity of sound sources, leading to a uniform perception of reverberation in all directions, which does not reflect the actual spatial impression of a room.
An apparatus and method that utilize directivity data to determine averaged gain data, incorporating room parameters and directivity models to adjust reverberation gain, ensuring that the reverberation is rendered based on the directivity characteristics of sound sources, using techniques such as spatial averaging and frequency-dependent gain adjustments.
Enhances the spatial accuracy of reverberation reproduction by accounting for the directivity of sound sources, resulting in a more realistic and directionally varied reverberation effect.
Smart Images

Figure 0007708733000013 
Figure 0007708733000014 
Figure 0007708733000015
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus and method for spatial audio reproduction by adjusting a reverberator based on the directivity characteristics of a sound source, and is not limited to a spatial audio reproduction apparatus and method by adjusting a reverberator based on the directivity positioning of a sound source in an augmented reality and / or virtual reality apparatus.
Background Art
[0002] Reverberation refers to the persistence of sound in a space after the actual sound source has stopped. The reverberation characteristics vary depending on the space. In order to convey the spatial impression of the environment, it is important to accurately reproduce the reverberation perceptually. Room acoustics often synthesize the early reflection part individually and represent the diffused late reverberation with a statistical model. FIG. 1 shows an example of the impulse response of a room in which a discrete early reflection 103 having an arrival direction (DOA) and a diffused late reverberation 105 that can be synthesized without a specific arrival direction are synthesized after the direct sound 101. The delay d1(t)102 in FIG. 1 can be regarded as indicating the direct sound arrival delay from the sound source to the listener, and the delay d2(t)104 can be regarded as indicating the delay from the sound source to the listener for one of the early reflections (in this case, the first arriving reflection).
[0003] One method of reproducing reverberation is to use N loudspeakers (or virtual loudspeakers reproduced binaurally using a set of head-related transfer functions (HRTFs)). The loudspeakers are arranged around the listener to some extent evenly. Incoherent reverberation signals are reproduced from these loudspeakers, and the reverberation diffused around is perceived.
[0004] The reverberations generated by different speakers must be mutually incoherent. In a simple case, the reverberations can be generated using different channels of the same reverberator, and the output channels are uncorrelated, but the acoustic characteristics such as RT60 time and level (especially the diffusion-to-direct ratio or reverberation-to-direct ratio) are the same. Such uncorrelated outputs sharing the same acoustic characteristics can be obtained, for example, from the output taps of a feedback-delay-network (FDN) reverberator with appropriately adjusted delay lengths, or from a reverberator based on the use of uncorrelated noise sequences that are attenuated by using different uncorrelated noise sequences for each channel. In this case, the different reverberation signals effectively have the same characteristics, and the reverberation is generally perceived to be similar in all directions. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0005] Embodiments of the present application aim to solve problems related to the prior art. MEANS FOR SOLVING THE PROBLEMS
[0006] According to a first aspect, there is provided an apparatus for assisting spatial rendering for room acoustics, the apparatus obtaining directional data having an identifier, the directional data including data regarding at least two distinct directions, obtaining at least one room parameter, determining information related to the directional data, determining gain data based on the determined information, determining averaged gain data based on the gain data, generating a bitstream defining the rendering, the bitstream being configured such that at least one audio signal related to the identifier is rendered based on at least one room parameter and the determined averaged gain data, comprising means configured to include the averaged gain data and at least one room parameter.
[0007] Means configured to determine information related to directivity data may be configured to determine a directivity model based on the directivity data.
[0008] The directivity model may be either a two-dimensional directivity model in which at least two directions are arranged on a plane or a three-dimensional directivity model in which at least two directions are arranged in space.
[0009] Means configured to determine averaged gain data may be configured to determine the averaged gain data based on spatial averaging of gain data that is independent of the direction and / or azimuth of the sound source.
[0010] Means configured to determine information related to directivity data may be configured to estimate a continuous directivity model based on the acquired directivity data.
[0011] Means configured to determine averaged gain data may be further configured to determine the gain data based on spatial averaging of gains for at least two separate directions based on the determined directivity model.
[0012] Means configured to acquire at least one room parameter may be configured to acquire at least one digital reverberator parameter.
[0013] Means configured to determine averaged gain data based on the gain data may be configured to determine frequency-dependent gain data.
[0014] The frequency-dependent gain data may be graphic equalizer coefficients.
[0015] According to a second aspect, an apparatus for spatial rendering for room acoustics is provided, the apparatus obtaining a bitstream, the bitstream including gain data averaged based on averaging of gain data, an identifier associated with at least one audio signal, or at least one audio signal, and at least one room parameter, and configuring at least one reverberator based on the averaged gain data and the at least one room parameter, and applying at least one reverberator to at least one audio signal as at least a part of rendering of the at least one audio signal, comprising means configured to do so.
[0016] The at least one room parameter may include at least one digital reverberator parameter.
[0017] The averaged gain data may include frequency-dependent gain data.
[0018] The frequency-dependent gain data may be graphic equalizer coefficients.
[0019] The averaged gain data may be spatially averaged gain data.
[0020] The means configured to apply at least one reverberator to at least one audio signal as at least a part of rendering of the at least one audio signal may further apply the averaged gain data to at least one audio signal to generate an audio signal affected by directivity, and apply a digital reverberator configured based on the at least one room parameter to the audio signal affected by directivity to generate a reverberant audio signal affected by directivity, and may be configured to do so.
[0021] The averaged gain data may include at least one set of grouped gains, the grouped gains being grouped for reasons of similar directivity patterns.
[0022] A similar directivity pattern may include differences between directivity patterns that are smaller than a determined threshold value.
[0023] According to a third aspect, a method for an apparatus for assisting spatial rendering for room acoustics is provided, the method comprising: obtaining directivity data having an identifier, the directivity data including data regarding at least two distinct directions; obtaining at least one room parameter; determining information related to the directivity data; determining gain data based on the determined information; determining averaged gain data based on the gain data; and generating a bitstream defining the rendering, the bitstream being configured such that at least one audio signal related to the identifier is rendered based on at least one room parameter and the determined averaged gain data, including the averaged gain data and the at least one room parameter.
[0024] Determining information related to the directivity data may include determining a directivity model based on the directivity data.
[0025] The directivity model may be either a two-dimensional directivity model in which at least two directions are arranged on a plane, or a three-dimensional directivity model in which at least two directions are arranged in space.
[0026] Determining the averaged gain data may include determining the averaged gain data based on a spatial averaging of gain data that is independent of the direction and / or azimuth of the sound source.
[0027] Determining information related to the directivity data may include estimating a continuous directivity model based on the obtained directivity data.
[0028] Determining the averaged gain data may include determining the gain data based on the spatial averaging of the gain for at least two distinct directions, further based on the determined directivity model.
[0029] Obtaining at least one room parameter may include obtaining at least one digital reverberator parameter.
[0030] Determining the averaged gain data based on the gain data may include determining frequency-dependent gain data.
[0031] The frequency-dependent gain data may be graphic equalizer coefficients.
[0032] According to a fourth aspect, a method for an apparatus for spatial rendering for room acoustics is provided, the method comprising obtaining a bitstream, the bitstream including averaged gain data based on the averaging of gain data, an identifier associated with at least one audio signal, or at least one audio signal, and at least one room parameter, obtaining, configuring at least one reverberator based on the averaged gain data and at least one room parameter, and applying at least one reverberator to at least one audio signal as at least a part of the rendering of the at least one audio signal.
[0033] The at least one room parameter may include at least one digital reverberator parameter.
[0034] The averaged gain data may include frequency-dependent gain data.
[0035] The frequency-dependent gain data may be graphic equalizer coefficients.
[0036] The averaged gain data may be spatially averaged gain data.
[0037] Applying at least one reverberator to at least one audio signal as at least a part of the rendering of the at least one audio signal may include applying the averaged gain data to the at least one audio signal to generate an audio signal affected by directivity, and applying a digital reverberator configured based on at least one room parameter to the audio signal affected by directivity to generate a reverberant audio signal affected by directivity.
[0038] The averaged gain data may include at least one set of grouped gains, and the grouped gains are grouped due to similar directivity patterns.
[0039] The similar directivity patterns may include differences between directivity patterns that are smaller than a determined threshold.
[0040] According to a fifth aspect, an apparatus for assisting spatial rendering for room acoustics is provided, the apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, together with the at least one processor, cause the apparatus to at least: acquire directivity data having an identifier, the directivity data including data in at least two distinct directions; acquire at least one room parameter; determine information related to the directivity data; determine gain data based on the determined information; determine averaged gain data based on the gain data; and generate a bitstream defining the rendering, the bitstream being configured such that at least one audio signal related to the identifier is rendered based on at least one room parameter and the determined averaged gain data, the generated bitstream including the averaged gain data and the at least one room parameter.
[0041] The apparatus configured to determine information related to the directivity data may be configured to determine a directivity model based on the directivity data.
[0042] The directivity model may be either a two-dimensional directivity model in which at least two directions are arranged on a plane or a three-dimensional directivity model in which at least two directions are arranged in space.
[0043] The apparatus configured to determine the averaged gain data may be configured to determine the averaged gain data based on a spatial averaging of gain data independent of the direction and / or azimuth of the sound source.
[0044] The apparatus configured to determine information related to the directivity data may be configured to estimate a continuous directivity model based on the acquired directivity data.
[0045] An apparatus adapted to determine averaged gain data may be adapted to determine the gain data based on the determined directivity model and further based on the spatial averaging of the gain for at least two distinct directions.
[0046] An apparatus adapted to obtain at least one room parameter may be adapted to obtain at least one digital reverberator parameter.
[0047] An apparatus adapted to determine averaged gain data based on the gain data may be adapted to determine frequency-dependent gain data.
[0048] The frequency-dependent gain data may be graphic equalizer coefficients.
[0049] According to a sixth aspect, an apparatus for spatial rendering for indoor acoustics, comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to cause the apparatus, at least by the at least one processor, to obtain a bitstream, the bitstream including averaged gain data based on averaging of gain data, an identifier associated with at least one audio signal, or at least one audio signal, at least one room parameter, to configure at least one reverberator based on the averaged gain data and the at least one room parameter, and to apply the at least one reverberator to the at least one audio signal as at least part of the rendering of the at least one audio signal.
[0050] The at least one room parameter may include at least one digital reverberator parameter.
[0051] The averaged gain data may include frequency-dependent gain data.
[0052] The frequency-dependent gain data may be graphic equalizer coefficients.
[0053] The averaged gain data may be spatially averaged gain data.
[0054] An apparatus adapted to apply at least one reverberator to at least one audio signal as at least part of the rendering of the at least one audio signal may further apply the averaged gain data to the at least one audio signal to generate an audio signal affected by directivity, and apply a digital reverberator configured based on at least one room parameter to the audio signal affected by directivity to generate a reverberant audio signal affected by directivity.
[0055] The averaged gain data may include at least one set of grouped gains, and the grouped gains are grouped due to similar directivity patterns.
[0056] The similar directivity patterns may include differences between directivity patterns that are less than a determined threshold.
[0057] According to a seventh aspect, there is provided an acquisition circuit configured to acquire directivity data having an identifier, the directivity data including data regarding at least two distinct directions; an acquisition circuit configured to acquire at least one room parameter; a determination circuit configured to determine information related to the directivity data; a determination circuit configured to determine gain data based on the determined information; a determination circuit configured to determine averaged gain data based on the gain data; and a generation circuit configured to generate a bitstream defining rendering, the bitstream being configured such that at least one audio signal related to the identifier is rendered based on at least one room parameter and the determined averaged gain data, the generation circuit including the averaged gain data and the at least one room parameter.
[0058] According to an eighth aspect, there is provided an acquisition circuit configured to acquire a bitstream, the bitstream including averaged gain data based on averaging of gain data, and an identifier related to at least one audio signal, or the at least one audio signal and at least one room parameter; a configuration circuit configured to configure at least one reverberator based on the averaged gain data and the at least one room parameter; and an application circuit configured to apply the at least one reverberator to the at least one audio signal as at least part of rendering of the at least one audio signal.
[0059] According to a ninth aspect, there is at least obtaining directivity data having an identifier, the directivity data including data regarding at least two distinct directions; obtaining at least one room parameter; determining information related to the directivity data; determining gain data based on the determined information; determining averaged gain data based on the gain data; and generating a bitstream defining rendering, the bitstream including at least one audio signal related to the identifier, averaged gain data, and at least one room parameter, configured to render based on the at least one room parameter and the determined averaged gain. A computer program including instructions [or a computer-readable medium including program instructions] for causing the apparatus to execute the above is provided.
[0060] According to a tenth aspect, there is at least obtaining, by an apparatus, a bitstream including averaged gain data based on averaging of gain data, an identifier related to at least one audio signal, or at least one audio signal, and at least one room parameter; configuring at least one reverberator based on the averaged gain data and the at least one room parameter; and applying the at least one reverberator to the at least one audio signal as at least a part of rendering of the at least one audio signal. A computer program including instructions [or a computer-readable medium including program instructions] for causing the apparatus to execute the above is provided.
[0061] According to an eleventh aspect, there is at least obtaining directivity data having an identifier, the directivity data including data regarding at least two distinct directions; obtaining at least one room parameter; determining information related to the directivity data; determining gain data based on the determined information; determining averaged gain data based on the gain data; and generating a bitstream defining rendering, the bitstream including at least one audio signal related to the identifier being configured to be rendered based on at least one room parameter and the determined averaged gain data, including the averaged gain data and at least one room parameter. There is provided a non-transitory computer-readable medium including program instructions for causing the apparatus to perform the above operations.
[0062] According to a twelfth aspect, there is at least obtaining, by an apparatus, a bitstream including averaged gain data based on averaging of gain data, and an identifier related to at least one audio signal, or at least one audio signal, and at least one room parameter; configuring at least one reverberator based on the averaged gain data and the at least one room parameter; and applying the at least one reverberator to the at least one audio signal as at least part of rendering of the at least one audio signal. There is provided a non-transitory computer-readable medium including program instructions for causing the apparatus to perform the above operations.
[0063] According to a 13th aspect, there is provided a means for acquiring directional data having an identifier, the directional data including data in at least two distinct directions, a means for acquiring at least one room parameter, a means for determining information related to the directional data, a means for determining gain data based on the determined information, a means for determining averaged gain data based on the gain data, and a means for generating a bitstream that defines rendering, the bitstream being configured such that at least one audio signal related to the identifier is rendered based on at least one room parameter and the determined averaged gain data, the generating means including the averaged gain data and the at least one room parameter. An apparatus is provided that includes these means.
[0064] According to a 14th aspect, there is provided a means for acquiring a bitstream, the bitstream including averaged gain data based on averaging of gain data, an identifier related to at least one audio signal, or the at least one audio signal and at least one room parameter, a means for configuring at least one reverberator based on the averaged gain data and the at least one room parameter, and a means for applying the at least one reverberator to the at least one audio signal as at least a part of rendering of the at least one audio signal. An apparatus is provided that includes these means.
[0065] According to a fifteenth aspect, there is at least obtaining directivity data having an identifier, the obtaining including directivity data including data regarding at least two distinct directions, obtaining at least one room parameter, determining information related to the directivity data, determining gain data based on the determined information, determining averaged gain data based on the gain data, and generating a bitstream that defines rendering, the bitstream being configured such that at least one audio signal related to the identifier is rendered based on at least one room parameter and the determined averaged gain data, the generating including the averaged gain data and at least one room parameter. A computer-readable medium including program instructions for causing a device to execute the above is provided.
[0066] According to a sixteenth aspect, there is at least obtaining, by a device, a bitstream, the bitstream including averaged gain data based on averaging of gain data, and an identifier related to at least one audio signal, or at least one audio signal, and at least one room parameter, configuring at least one reverberator based on the averaged gain data and at least one room parameter, and applying at least one reverberator to at least one audio signal as at least part of rendering of the at least one audio signal. A computer-readable medium including program instructions for causing the above to be executed is provided.
[0067] An apparatus including means for performing the operations of the method described above.
[0068] An apparatus configured to perform the operations of the method described above.
[0069] A computer program including program instructions for causing a computer to execute the method described above.
[0070] A computer program product stored in a medium can cause an apparatus to execute the method described herein.
[0071] An electronic device may include an apparatus as described herein.
[0072] A chipset may include an apparatus as described herein.
Brief Description of the Drawings
[0073] To better understand the present application, reference will now be made, by way of example, to the accompanying drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
DETAILED DESCRIPTION OF THE INVENTION
[0074] Hereinafter, a suitable apparatus and possible mechanisms for parameterizing and rendering a reverberant audio scene will be described in more detail.
[0075] As described above, reverberation can be expressed, for example, using a feedback-delay-network (FDN) reverberator with appropriately adjusted delay line lengths. The FDN can control the reverberation time (RT60) and the energy in various frequency bands individually. Therefore, it can be used to render reverberation based on the characteristics of a room or a modeled space. The reverberation time and the energy at various frequencies are affected by the absorption characteristics that depend on the frequency of the room.
[0076] Furthermore, the directivity of the sound source affects the energy at various frequencies. For example, since a human head and body acoustically cast a shadow, a human speaker is also affected. For this reason, when listening from behind the speaker, a phenomenon may occur where the direct sound is attenuated more than when listening from the front of the speaker. This attenuation depends on the frequency because the shadow cast by the head and body depends on the wavelength. Simplifying, a human speaker is considered to be almost omnidirectional at low frequencies (long wavelengths), while being quite directional at high frequencies (short wavelengths).
[0077] Directivity also affects late reverberation (hereinafter, the example of a human voice is used similarly). At low frequencies, since the sound source (of a human voice) is substantially omnidirectional, reverberation can be directly applied using known frequency-dependent energy and reverberation time (generally determined for omnidirectional sound sources).
[0078] However, at high frequencies, considering all sound sources as omnidirectional does not result in optimal reverberation quality. A human speaker radiates sound with “normal” frequency characteristics towards the front (since the audio signal is usually captured by a microphone in the front, the captured audio signal contains this frequency characteristic), while the sound is significantly attenuated at the back at high frequencies. Thus, in reality, due to the directivity of a human speaker, the contribution of late reverberation attenuates at high frequencies. Generally, the stronger the directivity of a sound source at a certain frequency, the less energy contributes to the reverberation in the room at that frequency.
[0079] Therefore, when rendering late reverberation, it is necessary to consider the room's frequency-dependent reverberation time and energy, as well as the directivity dependent on the frequency of the sound source.
[0080] The directivity of a sound source can be obtained in many ways. As an example, the amplitude frequency response of the sound source for various directions around the sound source can be measured (or modeled), and the ratio between these amplitude frequency responses and the amplitude frequency response for the front direction can be calculated. Thus, it can be described how much the sound is attenuated at various frequencies for these directions.
[0081] If this directivity data is available for all directions with infinite resolution, a spatially uniform (or actually pseudo-uniform) distribution can be selected for all 3D directions (for example, 100 data points evenly distributed in 3D). Then, their average value can be calculated, and the reverberation signal can be processed with the resulting amplitude frequency response (or filtered with the corresponding filter).
[0082] However, it is almost never the case that the directivity data is available with infinite resolution. The directivity data is usually available only for a limited number of directions. Furthermore, the distribution of the data points may not be (spatially) uniform. For example, the directivity data may be available for many directions in front, but only for some directions behind the sound source. In such a case, calculating a simple average value may cause a large bias in the amplitude-frequency characteristics.
[0083] In some cases, it is also possible for an audio engineer to manually adjust the appropriate amplitude-frequency characteristics based on the available directivity information. However, this is impossible for an automatic system and also impossible without some manual (and probably artistic) work by an audio engineer or the like.
[0084] Thus, it is necessary to be able to determine and render the effect of the directivity of the sound source on the amplitude-frequency response of the late reverberation in an effective way and without significant or any user input (or interaction).
[0085] The concepts in the embodiments described in more detail herein relate to the reproduction of the late reverberation component of a sound source. In these embodiments, an apparatus and method are proposed that enable the rendering of late reverberation based on the directivity data of the sound source in order to take into account the spectral effects caused by directivity. This is achieved, in some embodiments, by obtaining directivity data for a plurality of different directions, determining whether the directivity data is two-dimensional or three-dimensional from a directivity model, estimating the spatial region of the directivity data based on the directivity model, estimating frequency-dependent directivity-affected reverberation gain data based on the spatial region, and rendering the late reverberation based on the directivity-affected reverberation gain data (and audio signal(s) and room-related parameters such as frequency-dependent reverberation time and energy).
[0086] Furthermore, in some embodiments, the reverberation gain data affected by the determined directivity pools sound sources that are close to each other (the difference is below a threshold), determines the reverberation gain data affected by the average directivity for each pool, and can apply the reverberation gain data affected by the average directivity only once to the sum of the audio signals of the pool, so that the calculation efficiency can be improved.
[0087] In some embodiments, the reverberation gain data affected by the directivity may be determined in the encoder, which is transmitted to the decoder (e.g., as graphic equalizer coefficients), and the decoder may apply them when rendering the late reverberation. Furthermore, in some embodiments in a situation where there is pooled reverberation gain data affected by the directivity, only the reverberation gain data affected by the average directivity is transmitted, and an index to the various reverberation gain data affected by the average directivity may be transmitted for each sound source, minimizing the required bitrate for transmission.
[0088] In some embodiments, the reverberation gain data affected by the directivity of various audio elements can be collated into a common source directivity pattern having a unique identifier. For example, each audio element (audio object or channel) can have a source directivity pattern with a unique identifier, and all audio sources having the same source directivity pattern can be pooled by the renderer. In other words, in some embodiments, the same source directivity pattern identifier can be pooled by the renderer.
[0089] In MPEG-I Audio Phase 2, the bitstream and the rendering process are standardized normatively. Although a reference implementation of the encoder is also performed, the encoder implementation can be changed as long as the output bitstream conforms to the normative specification. Thereby, the quality of the codec can be improved with a new encoder implementation after standardization.
[0090] In some embodiments, an encoder reference implementation is configured to receive an encoder input format description having one or more sound sources with directivity and room-related parameters. Further in some embodiments, the encoder is configured to estimate frequency-dependent reverberation gain data based on the directivity of the sound source. Next, embodiments may be configured to estimate reverberator parameters based on the room-related parameters. Further embodiments may be configured to write a bitstream description including the reverberator parameters and the frequency-dependent reverberation gain data.
[0091] Further in some embodiments, a canonical bitstream is configured to include frequency-dependent reverberation gain data and reverberator parameters described using the syntax described herein. A canonical renderer in some embodiments is configured to decode the bitstream to obtain the frequency-dependent reverberation gain data and the reverberator parameters, use the parameters to initialize processing components for reverberation rendering, and use the initialized processing components to perform reverberation rendering using the presented method.
[0092] Thus, in some embodiments, for VR rendering, the reverberator parameters are derived at the encoder and transmitted in the bitstream. For AR rendering, the reverberator parameters are derived at the renderer based on a listening space description format (LSDF) file or corresponding representation. Source directivity data in some embodiments is available at the encoder. Embodiments described herein do not exclude implementations where a new sound source is provided directly to the renderer, which would also mean that the source directivity data arrives directly at the renderer.
[0093] With respect to FIG. 2, an exemplary implementation of a directivity-sensitive reverberator 299 according to some embodiments is shown.
[0094] In some embodiments, the directivity-affected reverberator 299 is configured to receive directivity data 200, an audio signal 204, and room parameters 206. Further, the directivity-sensitive reverberator 299 is configured to impart directivity-affected reverberation to the audio signal 204 based on the room parameters 206 and the directivity data 200, and output a directivity-affected reverberated audio signal, or a reverberated audio signal incorporating the effects of source directivity (or, generally, a reverberated audio signal) 208. These reverberated audio signals 208 can be in any suitable output format. For example, a multi-channel output format can be a 7.1+4 channel system format, a binaural audio signal, or a monaural audio signal.
[0095] The directivity data 200 is transferred to a directivity-affected reverberation gain determiner 201. In some embodiments, the directivity data 200 is in the form of gain values g dir (i,k) for a number of directions θ(i),φ(i), where i is the index of the data point, k is the frequency, θ is the azimuth angle, and φ is the elevation angle. Optimally, the directions should cover the entire sphere around the sound source evenly or uniformly, but the distribution in some embodiments may not be even or uniform, or may include only a small number of data points.
[0096] In some embodiments, the directivity-affected reverberator 299 includes a directivity-affected reverberation gain determiner 201. The directivity-affected reverberation gain determiner is configured to obtain or receive the directivity data 200 and determine a directivity-affected reverberation gain 202 g dir,rev (k) that describes how the directivity of the sound source affects the amplitude-frequency response of the late reverberation. The operation of the directivity-affected reverberation gain determiner 201 will be presented in more detail later.
[0097] As a result, the directivity-affected reverberation gain 202 is transferred to the reverberator 203.
[0098] In some embodiments, the reverberator 299 affected by directivity includes a reverberator 203. The reverberator receives a reverberation gain 202 affected by directivity and also an audio signal 204 s in (t) (where t is time) and room parameters 206 are configured to be received.
[0099] The room parameters can be in various forms. For example, in some embodiments, the room parameters 206 include energy in a frequency band k (typically as the diffusion-to-total ratio DDR or the reverberation-to-direct ratio RDR) and reverberation time (typically as RT60).
[0100] The reverberator 203 is configured to reverberate the audio signal 204 based on the room parameters 206 and the reverberation gain 202 affected by directivity. For example, the reverberator includes an FDN reverberator implementation configured in a way that will be described in more detail later.
[0101] As a result, a reverberated audio signal 208 s rev (j,t) (where j is the output audio channel index) is output. The output reverberated audio signal may be rendered for a multi-channel loudspeaker setup (such as 7.1+4) in some embodiments. This reverberation can be based on not only the directivity data 200 but also the room parameters 206.
[0102] Regarding FIG. 3, a flowchart showing the operation of the exemplary reverberator 299 affected by directivity shown in FIG. 2 is shown.
[0103] In FIG. 3, as indicated by step 301, the first operation can be to obtain an audio signal, directivity data, and room (reverberation) parameters.
[0104] And in FIG. 3, as shown by step 303, the reverberation gain affected by directivity can be determined.
[0105] In FIG. 3, as shown by step 305, after determining the reverberation gain affected by directivity, the reverberation audio signal affected by directivity is generated from the audio signal based on the reverberation gain affected by directivity and room parameters.
[0106] Next, in FIG. 3, as shown by step 307, the reverberation audio signal affected by directivity is output.
[0107] FIG. 4 shows the reverberation gain determiner 201 affected by directivity in more detail according to some embodiments. In some embodiments, the reverberation gain determiner 201 affected by directivity is configured to receive directivity data 200. The directivity data 200 in some embodiments includes the gain g dir (i,k) for the directions θ(i), φ(i).
[0108] In some embodiments, the reverberation gain determiner 201 affected by directivity includes a directivity model determiner 401. The directivity model determiner 401 is configured to analyze the input directivity data 200 to determine whether the data is three-dimensional or two-dimensional. This is based on the gain p(i)=[x(i), y(i), z(i)] in all directions i provided by the directivity data 200 TIt can be implemented by analyzing the axes of an array consisting of intersection points. If none of the three axes are all zero, it means that the directivity data 200 has three dimensions (3D), and thus the directivity model is three-dimensional. If one of the axes is all zero, the directivity data includes data provided in two dimensions (2D), and the directivity model is two-dimensional. The information of the resulting directivity model 402 in some embodiments is transferred to the (spatial) area weighted gain determiner 403. The area weighted gain can also be referred to as gain data.
[0109] In some embodiments, the reverberation gain determiner 201 affected by directivity includes the area weighted gain determiner 403. The area weighted gain determiner 403 is configured to receive the directivity model 402 information and divide the entire area, such as a sphere (in the case of a 3D model) or a plane (in the case of a 2D model), into small areas. The area weighted gain determiner 403 is further configured to receive the directivity data 200 and assign the directivity data to the sub-areas corresponding to the provided directivity values.
[0110] In some embodiments for a three-dimensional (3D) directivity model, a spherical Voronoi cover is formed from the intersection points p(i) of the gain in all directions. The Voronoi cover divides the sphere into regions close to each of the points p(i).
[0111] For the square of each gain, the magnitude of the area weighting is calculated as follows.
Equation
[0112] In some embodiments, the total area is calculated by summing all the Voronoi cell areas.
Equation
[0113] In some embodiments related to the two-dimensional (2D) directivity model, the Cartesian coordinates of the centroid of the planar region are calculated as follows.
Equation
[0114] In the above example, the non-zero Cartesian elements are represented by x and y, indicating that all z values were zero. However, it is not necessary to be such a case, and the method shown here always obtains two axes of non-zero elements, regardless of whether it is x, y, or z.
[0115]
Equation
[0116] Next, the area of the triangle can be calculated as follows in some embodiments.
Equation
[0117] For the square of each gain, the magnitude of the region weighting is calculated as follows.
Equation
[0118] Since the two-dimensional directivity model is not necessarily circular, in this embodiment, a method that can be used as a general method for obtaining the estimated area from the directivity data 200 of any two-dimensional shape is shown.
[0119] An alternative for obtaining the estimated area of a two-dimensional shape is to use, instead of the area of a triangle, the arc length from each midpoint (on the circle) between two directional samples to the next midpoint, i.e., the angle in radians.
[0120] As a result, the gain weighted by the spatial region (or region-weighted gain) 404 g dir,area (i,k) 2 can be transferred to the average gain determiner 405.
[0121] In some embodiments, the reverberation gain determiner 201 affected by directivity includes the average gain determiner 405. The average gain determiner 405 is configured to receive the gain weighted by the spatial region and determine the reverberation gain 202 affected by directivity. In some embodiments, the reverberation gain affected by directivity can be determined by calculating the average of the gains weighted by the spatial region. For example, in some embodiments, the reverberation gain 202 affected by directivity is determined as follows. [Number]
[0122] The reverberation gain g dir,rev (k) 202 at the output. The reverberation gain g dir,rev (k) is the original gain data g dir (i,k) spatially averaged over the (spatial) direction θ(i)φ(i) in which it is provided, and can also be referred to as the averaged gain. Note that the averaged gain data g dir,rev (k) no longer depends on the direction and depends only on the frequency k.
[0123] Also, the averaged gain data g dir,rev (k) can be directly represented in the bitstream and encoded using the original frequency k. Alternatively, 20*log10(g dir,rev(k)) can be calculated and converted to decibels. Alternatively, in addition to that, the averaged gain data can be expressed in other frequency resolutions such as octave or third-octave frequencies. As yet another alternative, the averaged gain data can be expressed in terms of the coefficients of a graphic equalizer filter that includes the coefficients of a cascade filter bank of second-order section IIR filters. Such a filter bank can be designed so that its amplitude response is similar to the input command gain in decibels, which can be set to be equal to 20*log10(g dir,rev (b)), where g dir,rev (b) represents the averaged gain evaluated at the filter bank center frequency (e.g., octave center frequency).
[0124] Regarding FIG. 5, it is a flowchart showing the operation of the reverberation gain determiner 201 affected by the directivity of the example shown in FIG. 4.
[0125] In FIG. 5, as shown by step 501, the first operation is to acquire the directivity data.
[0126] In FIG. 5, as shown by step 503, after acquiring the directivity data, it is used to determine the directivity model.
[0127] And in FIG. 5, as shown by step 505, based on the determined directivity model and the directivity data, the gain weighted by the spatial region is determined.
[0128] In FIG. 5, as shown by step 507, after determining the gain weighted by the spatial region, the average gain is determined.
[0129] And in FIG. 5, as shown by step 509, the reverberation gain (average gain) affected by the directivity is output.
[0130] With respect to FIG. 6, an exemplary reverberator 203 according to some embodiments is shown. The reverberator 203 may be implemented as any suitable digital reverberator 600 affected by directivity that is enabled or configured to generate reverberation whose characteristics match the room parameters. Exemplary reverberator implementations include a feedback delay network (FDN) reverberator that enables playback of reverberation with a desired frequency-dependent RT60 time and level, and a filter affected by directivity. The room parameters 206 are used to adjust the FDN reverberator parameters to generate the desired RT60 time and level. Examples of level parameters can include the direct-to-diffusion ratio (DDR) (or the diffusion-to-total energy ratio as used in MPEG-I). The directivity-affected reverberation gain 202 is input to the reverberator and applied to the input or output of the reverberator so that the reverberation spectrum (level) is appropriately adjusted according to the source directivity. The input to the directivity-affected FDN reverberator 600 can be an audio signal 204 that can be a monophonic input or a multichannel input or an ambisonics input. The output from the directivity-affected FDN reverberator 600 is a directivity-affected reverberation audio signal 208, which, in the case of binaural headphone playback, is then reproduced as two output signals, and in the case of loudspeaker output, typically means two or more output audio signals. Reproducing some outputs such as the FDN delay line output 15 as binaural outputs can be done, for example, via HRTF filtering.
[0131] FIG. 7 shows in more detail an exemplary directivity-affected FDN reverberator 600 that can be used to generate D uncorrelated output audio signals. In this example, each output signal can be rendered at a specific spatial position around the listener for an enveloping reverberation effect.
[0132] An exemplary implementation of the directivity-based FDN reverberator 600 processes the reverberation parameters to obtain the coefficients GEQ of each attenuation filter 761 d (GEQ1, GEQ2, ··· GEQ D) Coefficients A of the feedback matrix 757, length m of the delay line 759 d (m1, m2, ··· m D ) and the coefficients GEQ of the reverberation filter 753 based on directivity dir It includes an FDN reverberator 601 configured to generate. In this way, the exemplary FDN reverberator 601 shows a D-channel output by providing the output from each FDN delay line as a separate output. The exemplary FDN reverberator 600 affected by the directivity in FIG. 7 further includes a filter GEQ dir 753 affected by a single directivity, but in some embodiments, there are multiple such filters affected by directivity.
[0133] In some embodiments, each attenuation filter GEQ d 761 is implemented as a graphic EQ filter using M biquad IIR band filters. Therefore, in the octave band M = 10, the parameters of each graphic EQ include the feedforward coefficients and feedback coefficients of the biquad IIR filter, the gain of the biquad band filter, and the overall gain. In some embodiments, any suitable method can be implemented to determine the FDN reverberator parameters. For example, the method described in the GB patent application GB2101657.1 can be implemented to derive the FDN reverberator parameters so as to reproduce the desired RT60 time of the virtual / physical scene.
[0134] The reverberator uses a network of delays 759, feedback elements (shown as attenuation filters 761, feedback matrix 757 and combiner 755, and output gain 763) to generate a very dense impulse response for the late part. The input sample 751 is input to the reverberator to generate a reverberant audio signal component, and then can be output.
[0135] The FDN reverberator includes a plurality of recirculation delay lines. The unitary matrix A757 is used to control the recirculation within the network. In some embodiments, the attenuation filter 761, which can be implemented as a graphic EQ filter implemented as a cascade of second - order section IIR filters, can facilitate the control of the energy attenuation rate at various frequencies. The filter 761 is designed to attenuate a desired amount in decibels for each pulse passing through the delay line to obtain a desired RT60 time. Thus, the input to the encoder can provide the desired RT60 time for each specified frequency f, denoted as RT60(f). For a frequency f, the desired attenuation per signal sample is calculated as attenuationPerSample(f)= - 60 / (samplingRate*RT60(f)). The attenuation in decibels for a delay line of length m d is attenuationDb(f)=m d *attenuationPerSample(f).
[0136] The attenuation filter is designed as a cascade graphic equalizer filter as described in V. Valimaki and J. Liski, “Accurate cascade graphic equalizer,” IEEE Signal Process. Lett., vol. 24, no. 2, pp. 176 - 180, Feb. 2017 for each delay line. The design procedure outlined in the paper referenced above takes as input a set of command gains in octave bands. Also, as described in Third - Octave and Bark Graphic - Equalizer Design with Symmetric Band Filters, https: / / www.mdpi.com / 2076 - 3417 / 10 / 4 / 1222 / pdf, there are methods related to a similar graphic EQ structure that supports third - octave bands, increases the number of biquad filters to 31, and provides better conformity to a detailed target response.
[0137] In some further embodiments, the design procedure of V. Valimaki and J. Liski, “Accurate cascade graphic equalizer,” IEEE Signal Process. Lett., vol. 24, no. 2, pp. 176 - 180, Feb. 2017 is also used to design the parameters of the reverberation - directional filter GEQ dir The input to the design procedure is the reverberation gain 202 affected by the directivity in decibels.
[0138] The parameters of the FDN reverberator 601 can be adjusted to generate a reverberation having characteristics that match the input room parameters. In the case of this reverberator 601, the parameters include the coefficients 761 of each attenuation filter GEQ d the coefficients A757 of the feedback matrix, the length m of the D - delay lines 759 d and the spatial positions of the delay lines d.
[0139] Also, the coefficients of the directivity gain filter 753 GEQ dir are obtained based on the reverberation gain 202 affected by the directivity. In the present invention, each attenuation filter GEQ d and the directivity gain filter GEQ dir are graphic EQ filters using M bi - quadratic IIR band filters. Note that there are as many directivity gain filters GEQ dir as there are distinct directivity patterns for the input signal. It should be noted that in embodiments, the number of bi - quadratic filters in various graphic EQ filters can be different, and it is not necessary for the number to be the same in the delay - line attenuation filter and the reverberation gain filter affected by the directivity.
[0140] The number D of delay lines can be adjusted according to the quality requirements and the preferred trade-off between the reverberation quality and the computational complexity. In an embodiment, an efficient implementation with D = 15 delay lines is used. This enables the definition of the feedback matrix coefficient A proposed by Rocchesso: Maximally Diffusive Yet Efficient Feedback Delay Networks for Artificial Reverberation, IEEE Signal Processing Letters, Vol. 4, No. 9, Sep. 1997, from the perspective of a Galois sequence that facilitates an efficient implementation.
[0141] The length m of the delay line d d can be determined based on the size of the virtual room. For example, a room in the shape of a shoebox (or cuboid) can be defined with sizes xDim, yDim, and zDim. If the room is not cube-shaped (or shoebox-shaped), a shoebox or cube can be fitted into the room, and the size of the fitted shoebox can be used as the delay line length. Alternatively, it is also possible to obtain the dimensions as the three longest dimensions in a room that is not in the shape of a shoebox, and other appropriate methods can be used.
[0142] The delay can be set proportional to the standing wave resonance frequency in a virtual or physical room in some embodiments. The delay line length m d can be further configured to be relatively prime to each other in some embodiments.
[0143] Figure 8 schematically shows in more detail a filter 753 affected by directivity according to some embodiments. The purpose of this example is to group sources having the same or similar directivity patterns such that there can be a number B of directivity buses that is less than the number of sources S. A simple grouping would combine into one those sources that share the same directivity pattern since they have the same reverberation gain 202 affected by the same directivity. Further, in some embodiments, B can be made less than the number of distinct directivity patterns of the S sources. In this case, the grouping method combines into one those sources having reverberation gains affected by directivity that are close to each other. In some embodiments, closeness can be defined as the average absolute difference in decibels of the reverberation gains affected by directivity, or other suitable metrics such as the logarithmic spectral distortion of the average directivity pattern.
[0144] The criterion for closeness may depend on the available computing power and the number of sources. Thus, as the computing power decreases, the threshold for combining two sound sources having close directivity patterns can be increased. As the number of sound sources increases, the threshold for combining two sound sources with close directivity can be increased.
[0145] Thus, as shown in Figure 8, a first set of combiners that receive inputs of audio sources is shown. For example, a first set of sources is shown that includes audio source 1 8001, audio source 2 8002, and audio source 3 8003 input to the first combiner 8011 (since sources 1, 2, and 3 have similar directivity patterns or reverberation gains affected by the same directivity). Further, a second set of sources is shown that includes audio source 4 8004 and audio source 5 8005 input to the second combiner 8012 (since sources 4 and 5 have similar or the same reverberation gains affected by the same directivity because they have directivity patterns). Further, the Bth combiner 801 B has an audio source S-1 800 S-1and the audio source S800 S shows the B-th set of sources including (since source S-1 and S have reverberation gains with similar or the same directivity patterns affected by directivity).
[0146] And the output of each combiner 801 forms the input of the filter affected by directivity. Therefore, the first group of filters affected by directivity GEQ dir ,1 8031 has input 1 8021, and the second group of filters affected by directivity GEQ dir ,2 8032 has input 2 8022, and the B-th group of filters affected by directivity GEQ dir ,B 803 B has input B 802 B .
[0147] And the output of the filter 803 affected by directivity in each group is passed to the combiner 805.
[0148] The filter 753 affected by directivity can further include a combiner 805 that receives the outputs of the filters affected by the directivity of the group and then combines them to generate an input to the FDN reverberator 601.
[0149] Regarding Figure 9, a flowchart showing the operation of the configuration of the filter 753 / FDN601 affected by directivity shown in Figures 7 and 8 is shown.
[0150] For example, in Figure 9, as shown by step 901, the first operation is to obtain the directivity data of the sound source.
[0151] Next, in Figure 9, as shown by step 903, calculate or determine the reverberation gain affected by directivity regarding the sound source.
[0152] And in Figure 9, as shown by step 905, the reverberation gains affected by directivity regarding the sound sources are compared with each other.
[0153] And in FIG. 9, as indicated by step 907, when the reverberation gain data affected by the directivity is close to the reverberation gain data affected by the directivity of other sound sources, the sound source and other sound sources are grouped to be configured to use the reverberation gain data affected by the same directivity.
[0154] FIG. 10 schematically shows an apparatus showing an implementation example in which an encoder device is configured to implement a part of the function of a reverberator. For example, as shown in FIG. 10, the encoder is configured to generate a reverberation gain affected by directivity, write this information into a bitstream together with an audio signal and room parameters, and transmit it to a renderer (and / or save this information for later consumption).
[0155] In this embodiment, there are three sound sources as inputs. Therefore, a first sound source having directivity data 2001 and an audio signal 2041, a second sound source having directivity data 2002 and an audio signal 2042, and a third (q-th) sound source having directivity data 200 q and an audio signal 204 q exist. However, the number of input sound sources may be any number. The directivity data of each sound source is passed to a reverberation gain determiner affected by the relevant directivity (for example, a first reverberation gain determiner 2011 affected by the first directivity related to the first audio source (directivity data 2001), a second reverberation gain determiner 2012 affected by the second directivity related to the second audio source (directivity data 2002), and a q-th reverberation gain determiner 201 q ) affected by the q-th directivity related to the q-th audio source (directivity data 200 q ).
[0156] Each reverberation gain determiner 2011, 2012, and 201 q is related to the relevant audio signals 2041, 2042, and 204 qand a set 2021, 2022, and 202 of reverberation gains affected by directivity that can be encoded / quantized and combined into a bitstream with room parameter 206 that can then be passed to reverberator 203 q configured to output.
[0157] In some embodiments, the conversion from room parameters to reverberator parameters is performed by an encoder device, in which case the reverberator parameters are sent in a signal from the encoder to the renderer.
[0158] In some embodiments, a filter GEQ affected by directivity dir,j The "room parameters" mapped to digital reverberator parameters with filter parameters are described in the following bitstream definition.
[0159]
Table 1
[0160] Summarizing the above example of AudioObjectsStruct(), it is as follows. numberOfAudioObjects defines the number of audio objects in the audio scene. id uniquely identifies the audio object within the audio scene. When directivityPresentFlag is equal to 1, it indicates that there is directivity associated with the audio object. When directivityPresentFlag is equal to 1, directivityId is the directivity profile description identifier for each of the audio objects. Each directivity file present in the audio scene description has a unique directivityId. LocationStruct() provides information about the location of audio objects in an audio scene. This can be provided in an appropriate coordinate system (e.g., Cartesian, polar, etc.). This data structure can also provide direction information for audio objects, which may be more relevant for audio objects that are not omnidirectional point sources.
[0161]
Table 2
[0162] Summarizing the above example of AudioChannelSourcesStruct(), it is as follows. numberOfAudioChannelSources defines the number of channel sources in the audio scene. numberOfLoudspeakers defines the number of loudspeakers for a specific channel source. id defines the channel source with a unique identifier. channel_index defines the index of each channel for a given channel source. When directivityPresentFlag is equal to 1, it indicates that directivity is associated with that channel. directivityId is the directivity profile description identifier for the channels of a channel source where directivityPresentFlag is equal to 1. This identifier is unique for each directivity file present in the audio scene description.
[0163]
Table 3
[0164] Summarizing the above example of reverbPayloadStruct(), it is as follows. The numberOfSpatialPositions defines the number of output delay line positions for the late reverberation payload. This value is defined using an index corresponding to a specific number of delay lines. The bit string value "0b00" notifies the renderer of 15 spatial direction values for the delay lines. The other three values "0b01", "0b10", and "0b11" are reserved. The azimuth defines the azimuth angle of the delay line with respect to the listener. The range is from -180 degrees to 180 degrees. The elevation defines the elevation angle of the delay line with respect to the listener. The range is from -90 degrees to 90 degrees. The numberOfAcousticEnvironments defines the number of acoustic environments in the audio scene. The reverbPayloadStruct() carries information about one or more acoustic environments present in the audio scene at that time. An acoustic environment has specific "room parameters" such as the RT60 time used to obtain FDN reverberation parameters. environmentId This value defines the unique identifier of the acoustic environment. The delayLineLength defines the length of the graphic equalizer (GEQ) filter in sample units used for the settings of the delay line attenuation filter. The lengths of different delay lines corresponding to the same acoustic environment are prime to each other. The filterParamsStruct() describes the graphic equalizer cascade filter for configuring the attenuation filter of the delay line. Also, the same configuration is later used for configuring the filters for the diffusion-to-direct reverberation ratio and the reverberant source directivity gain. The details of this configuration are described in the following table.
[0165] The above source directivity processing examples can be summarized as follows. If directivitiesPresent is equal to 1, it indicates that there are audio elements with source directivity in the acoustic environment. If the value is equal to 0, the source directivity processing metadata may not be present in the late reverberation metadata. The numberOfDirectivities indicates the number of source directivities present in a particular acoustic environment. The directivities can be applied to one or more audio elements within the acoustic environment.
[0166] In some embodiments, the directivitiesPresent flag and related checks can be skipped. The directivity processing metadata may be present directly.
[0167] Summarizing the above example of filterParamsStruct(), it is as follows. The SOSLength is the length of each of the second - order section filter coefficients. The b1, b2, a1, a2 filters are composed of the coefficients b1, b2, a1, a2. These are the feed - forward and feedback IIR filter coefficients of a second - order section IIR filter. The globalGain specifies the gain coefficient of the GEQ in decibels. The levelDB specifies the sound level offset of each delay line in decibels.
[0168] The association between the source directivity profile and the reverberation payload directivity processing metadata in some embodiments is performed by the renderer / decoder. This can be done in some embodiments by first checking the relevant audio sources for a particular acoustic environment (e.g., included within the acoustic environment range). The relevant audio sources fed to the reverberation are checked for the presence of source directivity information (e.g., numberOfDirectivities is greater than 0 and directivitiesPresentFlag is equal to 1). Then, the reverberation metadata is checked for the presence of the corresponding reverbDirectivityGainFilterId. Then, the relevant source directivity filtering is applied before the audio is fed for late - reverberation rendering.
[0169] As described above, it can be seen that AudioChannelSourcesStruct and AudioObjectsStruct have a directivityId, and the reverb metadata payload has a reverbDirectivityGainFilterId. In some embodiments, the directivityId and the reverbDirectivityGainFilterId may be the same. In such a scenario, the number of directivityIds in the audio scene corresponding to the audio elements in a particular acoustic environment shall be equal to the numberOfDirectivities in the reverb payload metadata. In other embodiments, when some directivities in the description of the audio scene (represented by unique directivityIds) are determined to exceed a similar or equivalent threshold and can thus be clustered or combined using the method shown in FIG. 9, there may be a smaller number of reverbDirectivityGainFilterIds corresponding to fewer GEQs for performing source directivity-related filtering. Such clustering of multiple directivityIds of an audio scene for fewer reverbDirectivityGainFilterIds can be utilized by the renderer to obtain higher computational efficiency, as illustrated in the embodiment of FIG. 8.
[0170] In the case of mapping multiple directivityIds to one reverbDirectivityGainFilterId, an additional data structure can be provided in the bitstream. In such a situation, the clustering to obtain fewer reverbDirectivityGainFilterIds is performed by the encoder, and the information included in the bitstream can be obtained. In another implementation, such remapping can also be implemented by the renderer after performing a unique analysis to combine multiple directivityIds into a smaller number of reverbDirectivityGainFilterIds.
[0171]
Number
[0172] Summarizing the above example of reverbDirectivityGainFilterMappingStruct(), it becomes as follows. numReverbDirectivityGainFilters is the number of directivity gain filters included in the reverb metadata. reverbDirectivityGainFilterId specifies the GEQ filter identifier for the source directivity gain control of the reverb. numDirectivityIds specifies the number of source direcitivityIds mapped to one reverbDirectivityGainFilterId.
[0173] An example of the metadata that the encoder can provide so that the renderer can select whether to render each directivityId or combine multiple directivities with a single reverb directivity gain filter (GEQ) specified by a single reverbDirectivityGainFilterId based on the audio scene and computational load is shown below. In this way, this function provides flexibility for the renderer to make decisions at runtime.
[0174]
Number
[0175] Summarizing the above example of reverbDirectivitySimilarityStruct(), it becomes as follows. numDirectivityIds is the number of directivities for which similarity data exists. directivityId is the identifier of the source directivity specified in the audio scene description of one or more audio elements. The similarity_index specifies a numerical value that describes the characteristics of the directivity profile specified by the corresponding directivityId. This can be derived based on values derived with an appropriate similarity metric. As an example of such, for various frequency bins, it can be cited that the difference in gain is smaller than a preset threshold value. The smaller the threshold value, the larger the similarity index. That is, if the directivities are the same, the similarity_index becomes 255, and if they are most different, the similarity_index becomes 0. Other methods for measuring the similarity_index can be derived based on the requirements of the application. In an embodiment, the similarity_index may be the result of step 905 in FIG. 9.
[0176] In some embodiments, a determination or check as to whether a sound source has a filter affected by directivity is performed at the time of initialization of the reverb instance. A sound source having a filter affected by directivity receives a valid pointer to the filter instance affected by directivity, and a sound source not affected by directivity can receive a null pointer as the pointer to the filter instance affected by directivity.
[0177] In some embodiments, all filterParamsStruct() are deserialized into the GEQ object in the renderer, and an association between directivity and GEQ is formed. The renderer associates the directivity model of each audio object with the corresponding GEQ used to apply filtering to each audio item.
[0178] In some embodiments, the implementation of the reverberation directivity gain filtering in the renderer can be performed as follows. Initialize the directivity filter for bus B. An input signal having a directivity gain filter has a pointer to the directivity filter. Each directivity gain filter has an input bus. Also, the digital reverb has an input bus.
[0179] In each rendering loop through which the input audio signal passes to the digital reverb, the method first sets the input buffers of all the directional gain filters to zero. The method also resets the status flags of each directional filter indicating whether a signal has been added to the input bus of each directional gain filter.
[0180] When an audio signal is selected to be input to the reverb, the method first checks whether the audio signal is associated with a directional gain filter. This can be done by checking whether the directional gain filter pointer associated with the input audio signal has a valid value. If there is a valid value, the method adds the input audio signal to the input bus of the corresponding directional gain filter. A status flag indicating that an audio signal has been added to the input bus of this directional filter is set. If the pointer is NULL, the input audio signal is added directly to the reverb input bus.
[0181] When all input audio signals have been added either to the reverb input bus (without the gain filter affected by directivity) or to the input bus of the gain filter affected by directivity, the method performs filtering with the filter affected by directivity for which at least one audio signal has been added to its input bus. The method loops through the filters affected by directivity and for each filter affected by directivity, determines from the status flag whether at least one audio signal has been added to the input bus of this filter affected by directivity. If at least one audio signal has been added to the input bus, filtering is performed by this filter affected by directivity, and the output of this filter affected by directivity is added to the reverb input bus. The filter affected by directivity to which no audio signal has been added to the input bus may remain unprocessed.
[0182] Finally, a digital reverberator is used to process the reverberator input bus signal and generate an output signal.
[0183] FIG. 11 is a diagram showing an implementation example of the system of the embodiment as described above. The encoder unit 1101 can be implemented, for example, in a computer of an appropriate content creator and / or a network server computer.
[0184] The encoder 1101 is configured to receive a description 1100 of a virtual scene and an audio signal 1102. The description of the virtual scene may be provided in the MPEG-I encoder input format (EIF) or other appropriate format. Generally, the description of the virtual scene includes a description acoustically related to the content of the virtual scene, for example, scene geometry as a mesh, acoustic materials, an acoustic environment with reverberation parameters, the positions of sound sources, and other audio element-related parameters such as whether reverberation is rendered for the audio elements. The encoder 1101 in some embodiments includes a reverberation parameter acquirer 1103 configured to receive the description 1100 of the virtual scene and configured to acquire reverberation parameters. The reverberation parameters can be acquired from the RT60, DDR, and pre-delay from the acoustic environment in the embodiment.
[0185] The encoder 1101 further includes a reverberation gain determiner 1105 affected by directivity in some embodiments. The reverberation gain determiner 1105 affected by directivity is configured to receive the description 1100 of the virtual scene, more specifically, the directivity data of the sound sources it includes, and generate a reverberation gain affected by directivity that can be passed to a reverberation gain combiner 1107 and a reverberation parameter encoder 1108 affected by directivity.
[0186] In some embodiments, the encoder 1101 further comprises a reverberation gain combiner 1107 affected by directivity. The reverberation gain combiner 1107 affected by directivity obtains a reverberation gain affected by directivity and determines whether any gain grouping should be applied. This information can be passed to the reverberation parameter encoder 1108. The combiner 1107 is optional.
[0187] In some embodiments, the encoder 1101 further comprises a reverberation parameter encoder 1108 affected by directivity. The reverberation parameter encoder 1108 affected by directivity in some embodiments is configured to obtain a reverberation gain affected by directivity and optionally combiner information and write a bitstream description including the reverberator parameters and frequency-dependent reverberation gain data. This can then be output to the bitstream encoder 1109.
[0188] In some embodiments, the encoder 1101 further includes a bitstream encoder 1109 configured to receive the output of the reverberation parameter encoder 1109 and an audio signal and generate a bitstream 1111 that can be passed to the bitstream decoder 1123. In other words, a canonical bitstream can be configured to include frequency-dependent reverberation gain data and reverberator parameters described using the syntax described herein. The bitstream 1111 in some embodiments can be streamed to an end-user device or made available for download or stored.
[0189] The output of the encoder is the bitstream 1111 that is made available for download or streaming. The function of the decoder / renderer 1121 operates on an end-user device that is any suitable system that uses audio, such as a mobile terminal, a personal computer, a soundbar, a tablet computer, a car media system, a home HiFi or theater system, a head-mounted display for AR or VR, a smartwatch, or any other appropriate system.
[0190] In some embodiments, the decoder 1121 includes a bitstream decoder 1123 configured to decode the bitstream to obtain frequency-dependent reverberation gain data and reverberator parameters.
[0191] The decoder 1121 may further include a reverberation parameter decoder 1127 configured to obtain the encoded frequency-dependent reverberation gain data and reverberator parameters from the bitstream decoder 1123 and decode them in an operation opposite to or inverse of that of the reverberation parameter encoder 1108.
[0192] In some embodiments, the decoder 1121 includes a reverberation directivity gain filter creation unit 1125 that receives the output of the reverberation parameter decoder 1127, generates a gain filter affected by reverberation directivity, and passes this to the reverberation directivity gain filter 1131.
[0193] In some embodiments, the decoder 1121 includes a gain filter 1131 affected by reverberation directivity configured to filter the directivity gain affected by reverberation and provide an input to the FDN reverberator 1133. The FDN reverberator 1133 may be initialized with the reverberator parameters provided by the reverberation parameter decoder 1127.
[0194] The decoder 1121 is configured to apply the FDN reverberator 1133 to generate a late reverberation audio signal that is passed to the head-related transfer function (HRTF) processor 1135.
[0195] In some embodiments, decoder 1121 includes an HRTF processor 1135 configured to apply HRTF processing to the late reverberation audio signal to generate a binaural audio signal and output this to a binaural signal combiner 1139.
[0196] Furthermore, decoder / renderer 1121 is configured to receive the audio signal decoded from bitstream decoder 1123 and includes a direct sound processor 1129 configured to perform any direct sound processing such as air absorption and distance-gain attenuation, and can be passed to an HRTF processor 1137 that can generate a direct sound component along with head direction determination, and passed to a binaural signal combiner 1139 along with the reverberation component from HRTF processor 1135. The binaural signal combiner 1139 is configured to combine the direct sound portion and the reverberant sound portion to generate an appropriate output (e.g., for headphone playback).
[0197] In some further embodiments, the decoder includes a head direction determiner 1141 that passes head direction information to HRTF processor 1137.
[0198] The decoder further includes a binaural signal combiner configured to take in the inputs from HRTF processor 1135 and HRTF processor 1137 and generate a binaural audio signal that can be output to an appropriate transducer set such as a headphone / speaker set. Although not shown, various other audio processing methods can be applied, such as early reflection rendering combined with the proposed method.
[0199] The described MPEG-I Audio Phase 2 is configured to standardize bitstream and renderer processing in a normative manner. There are also implementations based on the encoder, but as long as the output bitstream follows the normative specification, it can be changed later. This allows the quality of the codec to be improved by new encoder implementations even after the standard is finalized.
[0200] In the main embodiments of the present invention, the parts directed to different parts of the MPEG-I standard are, as shown in FIG. 11, as follows.
[0201] · The reference implementation of the encoder includes the following. 〇 Receiving an encoder input format description including a virtual scene description having one or more sound sources with directivity and room-related parameters. 〇 Obtaining reverberator parameters from room-related parameters. 〇 Determining a reverberation gain affected by directivity. 〇 Optionally, combining reverberation gains affected by directivity. 〇 Writing a bitstream description including reverberator parameters and frequency-dependent reverberation gain data.
[0202] · The normative bitstream shall include frequency-dependent reverberation gain data and reverberator parameters described using the syntax described herein. The bitstream shall be streamed to terminal devices or be available for download or storage.
[0203] · The normative renderer shall decode the bitstream to obtain frequency-dependent reverberation gain data and reverberator parameters, initialize processing components for reverberation rendering using the parameters, and perform reverberation rendering in the presented manner using the initialized processing components. 〇 In VR rendering, as shown in FIG. 11, the encoder derives the reverberator parameters and transmits them in the bitstream. 〇In AR rendering, the reverberator parameters are derived by the renderer based on a Listening Space Description Format (LSDF) file or a corresponding representation (not shown in FIG. 11). 〇The sound source directivity data is available to the encoder, and there is no use case at present for directly providing a new sound source to the renderer. However, in the future, such a use case may appear, and it is possible that the renderer will determine the reverberation gain affected by the directivity.
[0204] · Also, a complete canonical renderer obtains other parameters related to room acoustics and sound source characteristics from the bitstream and uses them to render, in addition to the diffuse late reverberation, the direct sound, early reflections, diffraction, the spatial range or width of the sound source, and other acoustic effects. The invention introduced here focuses on the rendering of the diffuse late reverberation part, particularly on the method of adjusting the diffuse late reverberation spectrum based on the directivity characteristics of the sound source.
[0205] Regarding FIG. 12, it is an exemplary electronic device that can be used as any of the device parts of the system as described above. The device can be any suitable electronic device or apparatus. For example, in some embodiments, the device 2000 is a mobile terminal, a user device, a tablet computer, a computer, an audio playback device, etc. The device may be configured to implement, for example, an encoder or a renderer or any of the functional blocks as described above.
[0206] In some embodiments, the device 2000 includes at least one processor or central processing unit (CPU) 2007. The processor 2007 can be configured to execute various program codes such as the methods described herein.
[0207] In some embodiments, device 2000 includes a memory (MEM) 2011. In some embodiments, at least one processor 2007 is connected to the memory 2011. The memory 2011 can be any suitable storage means. In some embodiments, the memory 2011 includes a program code portion for storing program code executable by the processor 2007. Further, in some embodiments, the memory 2011 can further include a stored data portion for storing data, for example, data processed or to be processed according to the embodiments described herein. The implemented program code stored within the program code portion and the data stored within the stored data portion can be retrieved by the processor 2007 at any time as needed via the memory-processor coupling.
[0208] In some embodiments, device 2000 includes a user interface (UI) 2005. The user interface 2005 can be connected to the processor 2007 in some embodiments. In some embodiments, the processor 2007 can control the operation of the user interface 2005 and receive inputs from the user interface 2005. In some embodiments, the user interface 2005 can enable a user to input commands to the device 2000, for example, via a keypad. In some embodiments, the user interface 2005 can enable a user to obtain information from the device 2000. For example, the user interface 2005 can include a display configured to display information from the device 2000 to the user. The user interface 2005 can be configured as a touch screen or touch interface that enables both inputting information into the device 2000 and displaying information to the user of the device 2000 in some embodiments. In some embodiments, the user interface 2005 can be a user interface for communication.
[0209] In some embodiments, device 2000 includes an input / output port 2009. The input / output port 2009 in some embodiments includes a transceiver. The transceiver in such embodiments is coupled to processor 2007 and can be configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitting means and / or receiving means can be configured to communicate with other electronic devices or apparatuses via a wiring or a wired connection in some embodiments.
[0210] The transceiver can communicate with further devices by any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth®, or an infrared data communication path (IRDA).
[0211] The input / output port 2009 may be configured to receive signals.
[0212] In some embodiments, device 2000 may be employed as at least a part of a renderer. The input / output port 2009 may be connected to, for example, headphones (which may be head-tracked headphones or non-tracked headphones).
[0213] In general, various embodiments of the present invention can be implemented in hardware or special-purpose circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software executable by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Various aspects of the present invention can be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, but these blocks, devices, systems, techniques, or methods described herein are, by way of non-limiting example, implemented in hardware, software, firmware, special-purpose circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0214] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile terminal, such as within a processor entity, or by hardware, or by a combination of software and hardware. In this regard, it should be noted that any block of the logic flow as illustrated in the figures can represent a program step, or interconnected logic circuits, blocks, and functions, or a combination of program steps and logic circuits, blocks, and functions. The software can be stored in a physical medium, such as a memory chip, or a memory block implemented within a processor, a magnetic medium such as a hard disk or a floppy disk, and an optical medium such as, for example, a DVD and its data variations, a CD.
[0215] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processing device can be of any type suitable for the local technical environment and can include, by way of non-limiting example, one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processing architecture.
[0216] Embodiments of the present invention can be implemented in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0217] In programs provided by Synopsys in Mountain View, California, and Cadence Design in San Jose, California, established design rules and a library of pre-stored design modules are used to automatically route conductors and place components on a semiconductor chip. When the design of a semiconductor circuit is complete, the resulting design in a standardized electronic format (such as Opus, GDSII) may be sent to a semiconductor manufacturing facility or "fab" for manufacturing.
[0218] The foregoing description has been a complete and helpful illustration of exemplary embodiments of this invention by way of illustrative and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the relevant art upon consideration of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined by the appended claims.
Claims
1. An apparatus for spatial rendering for indoor acoustics, the apparatus comprising at least one processor and at least one memory containing computer program code, wherein the at least one memory and the computer program code, together with the at least one processor, cause the apparatus to at least acquire a bitstream, the bitstream including gain data averaged based on averaging of gain data, an identifier associated with at least one audio signal, or at least one audio signal, and at least one room parameter, configure at least one reverberator based on the averaged gain data and the at least one room parameter, process the at least one audio signal with the at least one reverberator as at least part of rendering of the at least one audio signal, be configured to perform, the averaged gain data includes frequency-dependent gain data, the apparatus is adapted to perform processing of the at least one audio signal by the at least one reverberator as at least part of the rendering of the at least one audio signal, and the apparatus further applies the averaged gain data to the at least one audio signal to generate a directionally affected audio signal, applies the configured at least one reverberator to the directionally affected audio signal based on the at least one room parameter to generate a directionally affected reverberated audio signal, the averaged gain data includes at least one set of gains grouped based on a common directivity pattern, apparatus.
2. The apparatus according to claim 1, wherein the at least one room parameter includes at least one digital reverberator parameter.
3. The apparatus according to claim 1, wherein the frequency-dependent gain data are graphic equalizer coefficients.
4. The apparatus according to claim 1, wherein the averaged gain data are spatially averaged gain data.
5. The apparatus according to claim 1, wherein the common directivity pattern includes differences between directivity patterns smaller than a determined threshold.
6. The apparatus according to claim 1, wherein the at least one reverberator comprises a feedback-delay-network reverberator.
Citation Information
Patent Citations
Spatial Audio for Two-Way Audio Environments
JP2021528001A
Encoding reverberator parameters from virtual or physical scene geometry and desired reverberation characteristics and rendering using these
US20210287651A1
Rendering reverberation
WO2021186102A1
Rendering encoded 6DOF audio bitstream and late updates
WO2021186104A1