Determination of early reflex parameters
By using spatial room impulse responses and room shape to determine absorption coefficients, the method addresses the challenge of accurately rendering early reflections in virtual acoustic systems, enabling immersive augmented reality audio experiences.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2024-04-08
- Publication Date
- 2026-06-04
AI Technical Summary
Current acoustic simulation methods fail to accurately reproduce early reflections in virtual acoustic rendering systems due to the lack of efficient methods to determine absorption coefficients for room surfaces, which is necessary for realistic rendering of early reflections, especially in systems requiring automated processing like augmented reality.
A method and apparatus that automatically determine absorption coefficients for room surfaces using spatial room impulse responses (SRIR) and room shape, allowing for the accurate rendering of early reflections by determining reverberation time, image source positions, and reflection coefficients.
Enables realistic rendering of early reflections, providing immersive augmented reality audio experiences by accurately reproducing the acoustic properties of a room without manual measurement of absorption coefficients.
Smart Images

Figure 2026518109000001_ABST
Abstract
Description
[Technical Field]
[0001] This application relates to an apparatus and method for determining early reflection parameters, but is not limited to determining early reflection parameters from spatial intra-intra [Background technology]
[0002] Reverberation refers to the persistence of sound in a space after the actual sound source has ceased. Different spaces are characterized by different reverberation characteristics. Perceptually accurate reproduction of reverberation is important to convey the spatial impression of an environment. Room acoustics are often modeled using statistical models for individually synthesized early reflections and diffuse late reverberation. Figure 1 shows an example of a synthesized room impulse response, where the direct sound 101 is followed by discrete early reflections 103 with a direction of arrival (DOA), and diffuse late reverberation 105 that can be artificially synthesized, corresponding to the temporal, spectral, and directional statistics of actual late reverberation in a real or virtual acoustic space. The delay d1(t) 102 in Figure 1 can be seen as representing the arrival delay of the direct sound from the source to the listener, and the delay d2(t) 104 can be seen as representing the delay from the source to the listener for one of the early reflections (in this case, the first arriving reflection). In such an example, after the direct sound, the listener hears directional early reflections. After a certain point, individual reflections are no longer perceptible, but the listener hears diffused, late-reverberation.
[0003] Late reverberation can be rendered, for example, using a feedback-delay-network (FDN) reverberation device with appropriate adjustment of the delay line length. The FDN allows for individual control of reverberation time (RT60) and energy in different frequency bands. Therefore, the FDN can be used to render reverberation based on the characteristics of the room. Furthermore, the reverberation time and energy in different frequency bands are influenced by the room's frequency-dependent absorption characteristics.
[0004] Early reflections can be rendered as delayed and filtered copies of the original signal, for example, modeling specular reflection that occurs when the original signal reflects off a wall. However, if a recording is made in a real room (for example, by reproducing a test signal through a speaker), and the same signal is then rendered as an object signal using a room simulation, the simulation results do not produce output of comparable quality to current computationally efficient (i.e., real-time interactive rendering) methods. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Lehmann, EA, and Johansson, AM (2008), "Prediction of energy decay in room impulse responses simulated with an image-source model," Journal of the Acoustical Society of America, 124(1), 269-277. [Non-Patent Document 2] J.B. Allen and D.A. Berkley, "Image method for efficient simulating small-room acoustics," J. Acoust. Soc. Am., Vol. 65, pp. 943-950, April 1979. [Non-Patent Document 3] J. Borish, "Extension of the image model to arbitrary polyhedra," Journal of the Acoustical Society of America 75.6(1984):1827-1836. [Overview of the project]
[0006] According to a first aspect, a method is provided for reproducing at least one reflection of reverberation in a virtual acoustic rendering system, the method comprising: a step of obtaining the shape of a room and at least one spatial room impulse response for a room, wherein the at least one spatial room impulse response is obtained based on measurements in the room; a step of determining reverberation time parameters based on the at least one spatial room impulse response; a step of determining a room absorption area based on the determined reverberation time parameters and the shape of the room; a step of determining at least one parameter of early reflection from the at least one spatial room impulse response based on at least one image source location, wherein the at least one image source location is associated with the early reflection of the at least one spatial room impulse response; a step of extracting at least one reflection from the at least one spatial room impulse response; a step of extracting at least one reflection coefficient for at least one surface in the room based on the extracted at least one reflection; a step of assigning an absorption coefficient for at least one surface in the room based on the at least one reflection coefficient and the determined room absorption area; and a step of rendering at least one early reflection based on the determined absorption coefficient.
[0007] The step of determining at least one parameter may include determining at least one of the following: the time of early reflection from at least one intra-space impulse response based on at least one image source position, and the direction of arrival of the early reflection from at least one intra-space impulse response based on at least one image source position.
[0008] The step of extracting at least one reflection from at least one spatial intra-intra
[0009] The step of obtaining at least one spatial intra-intra
[0010] At least one spatial intra-intra
[0011] For interiors, the step of obtaining the shape of the interior is the surface area S i The process may include obtaining an indexed list of planar surfaces containing each of these, and metadata indicating the locations of the receiver and source within the shape of the room as image sources.
[0012] For interiors, the step of acquiring the shape of the interior may further include the step of acquiring metadata indicating a source-directed pattern.
[0013] The step of determining the room absorption area based on the determined reverberation time parameters and the shape of the room may include: the step of determining the total absorption area from a combination of surface areas, and the step of determining the room absorption area based on the total absorption area and the determined reverberation time parameters.
[0014] According to a second aspect, an apparatus is provided for reproducing at least one reflection of reverberation in a virtual acoustic rendering system, the apparatus comprising means configured to: acquire the shape of a room and at least one spatial room impulse response for a room, the at least one spatial room impulse response acquired based on measurements in the room, determine reverberation time parameters based on the at least one spatial room impulse response, determine a room absorption area based on the determined reverberation time parameters and the shape of the room, determine at least one parameter of early reflection from the at least one spatial room impulse response based on at least one image source position, the at least one image source position associated with the early reflection of the at least one spatial room impulse response, extract at least one reflection from the at least one spatial room impulse response, extract at least one reflection coefficient for at least one surface in the room based on the extracted at least one reflection, assign an absorption coefficient for at least one surface in the room based on the at least one reflection coefficient and the determined room absorption area, and render at least one early reflection based on the determined absorption coefficient.
[0015] Means configured to determine at least one parameter may be configured to determine at least one of: the time of early reflection from at least one spatial intra-intra
[0016] Means configured to extract at least one reflection from at least one spatial intra-intra
[0017] Means configured to acquire at least one spatial intra-intra
[0018] At least one spatial intra-intra
[0019] A means configured to acquire the shape of the interior of a room is the surface area S i It may be configured to retrieve an indexed list of planes containing each of these, and metadata indicating the location and source of the receiver within the shape of the room.
[0020] Means configured to acquire the shape of a room may be further configured to acquire metadata indicating a source-directed pattern.
[0021] Means configured to determine the room absorption area based on determined reverberation time parameters and the shape of the room may be configured to: determine the total absorption area from a combination of surface areas, and determine the room absorption area based on the total absorption area and the determined reverberation time parameters.
[0022] According to a third aspect, an apparatus is provided for reproducing at least one reflection of reverberation in a virtual acoustic rendering system, the apparatus comprising at least one processor and at least one memory for storing instructions, the instructions, when executed by at least one processor, to the system: to acquire, for a room, the shape of the room and at least one spatial room impulse response, the at least one spatial room impulse response acquired based on measurements in the room; to determine reverberation time parameters based on the at least one spatial room impulse response; to determine the room absorption area based on the determined reverberation time parameters and the shape of the room; and at least one image saw The method involves determining at least one parameter of early reflection from at least one spatial intra-intra
[0023] An apparatus tasked with determining at least one parameter may be tasked with determining at least one of the following: the time of early reflection from at least one spatial intra-intra
[0024] An apparatus capable of extracting at least one reflection from at least one spatial intra-intra
[0025] An apparatus that is configured to acquire at least one spatial intra-intra
[0026] At least one spatial intra-intra
[0027] A device that is used to acquire the shape of an interior space has a surface area S i The system may obtain an indexed list of planes containing each of these planes, and metadata indicating the location and source of the receiver within the shape of the room.
[0028] A device that acquires the shape of a room may also acquire metadata indicating a source-directed pattern.
[0029] An apparatus that determines the room absorption area based on the determined reverberation time parameters and the shape of the room may also determine the total absorption area from a combination of surface areas, and determine the room absorption area based on the total absorption area and the determined reverberation time parameters.
[0030] According to a fourth aspect, an apparatus is provided for reproducing at least one reflection of reverberation in a virtual acoustic rendering system, the apparatus comprising: an acquisition circuit configured to acquire the shape of a room and at least one spatial room impulse response for a room, the acquisition circuit being acquired based on measurements in the room; a determination circuit configured to determine reverberation time parameters based on the at least one spatial room impulse response; a determination circuit configured to determine the room absorption area based on the determined reverberation time parameters and the shape of the room; and a determination circuit for determining at least one parameter of early reflection from the at least one spatial room impulse response based on at least one image source position. A determination circuit configured to include: a determination circuit in which at least one image source position is associated with an early reflection of at least one spatial-intra
[0031] According to a fifth aspect, a computer program (or computer-readable medium containing instructions) is provided which includes instructions for an apparatus to reproduce at least one reverberation reflection in a virtual acoustic rendering system, wherein the apparatus: acquires the shape of a room and at least one spatial room impulse response, the at least one spatial room impulse response is acquired based on measurements in the room; determines reverberation time parameters based on the at least one spatial room impulse response; determines the room absorption area based on the determined reverberation time parameters and the shape of the room; and, based on at least one image source position, at least one The method involves determining at least one parameter of early reflection from a spatial-intra
[0032] According to the sixth aspect, an apparatus is provided for reproducing at least one reflection of reverberation in a virtual acoustic rendering system, the apparatus comprising means for acquiring the shape of a room and at least one spatial room impulse response, wherein the at least one spatial room impulse response is acquired based on measurements in the room; means for determining reverberation time parameters based on the at least one spatial room impulse response; means for determining the room absorption area based on the determined reverberation time parameters and the shape of the room; and at least one early reflection from the at least one spatial room impulse response based on the position of at least one image source Means for determining 1 parameter, wherein at least one image source position is associated with an early reflection of at least one spatial intra-intra
[0033] According to the seventh aspect, for the reproduction of at least one reflection of reverberation in a virtual acoustic rendering system, the apparatus includes at least the following: acquiring the shape of a room and at least one spatial room impulse response for a room, wherein the at least one spatial room impulse response is acquired based on measurements in the room; determining reverberation time parameters based on the at least one spatial room impulse response; determining the room absorption area based on the determined reverberation time parameters and the shape of the room; and determining at least one parameter of early reflection from the at least one spatial room impulse response based on the position of at least one image source. A non-temporary computer-readable medium is provided which includes program instructions for causing the following to be performed: at least one image source location to be associated with an early reflection of at least one spatial intra-intra
[0034] According to the eighth aspect, for the reproduction of at least one reflection of reverberation in a virtual acoustic rendering system, the apparatus includes at least the following: for a room, the shape of the room and at least one spatial room impulse response, wherein the at least one spatial room impulse response is obtained based on measurements in the room; determining reverberation time parameters based on the at least one spatial room impulse response; determining the room absorption area based on the determined reverberation time parameters and the shape of the room; and at least one parameter of early reflection from the at least one spatial room impulse response based on the position of at least one image source. A computer-readable medium is provided which includes instructions for determining that at least one image source position is associated with an early reflection of at least one spatial intra-intra
[0035] A device equipped with means for performing actions in the manner described above.
[0036] A device configured to perform actions in the manner described above.
[0037] A computer program that contains program instructions to cause a computer to perform the methods described above.
[0038] A computer program product stored on a medium may cause the device to perform actions in the manner described herein.
[0039] The electronic device may include any apparatus as described herein.
[0040] The chipset may include devices such as those described herein.
[0041] The embodiments of this application aim to address problems related to the prior art.
[0042] For a better understanding of this application, references to the attached drawings are made here as an example. [Brief explanation of the drawing]
[0043] [Figure 1] This figure shows a model of room acoustics and room impulse response. [Figure 2] This figure schematically illustrates an exemplary device in which several embodiments may be implemented. [Figure 3] Figure 2 is a flowchart illustrating the operation of an exemplary device. [Figure 4] This figure provides a more detailed schematic representation of an exemplary absorption coefficient determination system, such as the one shown in Figure 2, according to several embodiments. [Figure 5] Figure 4 shows a flowchart illustrating the operation of an exemplary absorption coefficient determinationer in several embodiments. [Figure 6] This figure schematically illustrates an exemplary early reflection parameter processor according to several embodiments. [Figure 7] Figure 6 shows a flowchart illustrating the operation of an early reflection parameter processor in several embodiments. [Figure 8] This figure shows an exemplary environment illustrating the source and image source shapes relative to the listener and reflective surface. [Figure 9] This figure shows an exemplary system in which several embodiments may be implemented. [Figure 10] This figure shows an exemplary device suitable for implementing the device shown in the previous figure. [Modes for carrying out the invention]
[0044] Below, we will describe in more detail appropriate apparatuses and possible mechanisms for determining early reflection parameters from spatial intra-intra
[0045] As mentioned above, there is a difference in quality between capturing reality and acoustic simulations based on current simulation methods. The reason for this discrepancy between efficient simulation and capturing reality is that it is not possible to efficiently capture the substantial amounts of different effects that occur in a room (absorption by materials and air, diffraction, and scattering from walls) that contribute to the density and spectral quality of reflections.
[0046] For example, individual early reflections are typically filtered using composite material filters, such as those implemented as low-order infinite impulse response filters. These filters emulate to some extent the frequency-dependent absorption characteristics of different materials, but more complex acoustic effects are ignored in this approach.
[0047] Because early reflections, when added to direct sound within the listener's ear, produce a distinct comb-filtering effect, the discrepancy between efficient acoustic simulation and the capture of reality has a greater impact on early reflections than on late reverberation. This allows the listener not only to accurately perceive space but also to apply spectral coloration.
[0048] It is known that material or absorption filter parameters can be measured from spatial-in-room impulse response (SRIR), and that the spectral colors induced by the room can be accurately reproduced using SRIR (and room shape). Using these methods, early reflections can be accurately rendered.
[0049] However, these methods fail to define the relationship between absorption or material filter parameters and parts of the room geometry during parameter analysis, requiring a separate step to relate the measured parameters to similar provided material properties of the room (based on either the provided material properties of the room or an image).
[0050] Therefore, the current approach requires information that may not be available in many use cases, such as AR rendering. For example, a user who wants to capture the features of an interior room for use in subsequent AR rendering needs to know the absorption coefficient of the surfaces in their room. Obtaining the absorption coefficient requires a tremendous amount of manual work and measurement, and therefore such a method is not suitable for a system that allows the user to adapt their room or switch rooms adaptively.
[0051] In other words, current approaches to determining absorption parameters for early reflection rendering require knowing the absorption coefficient of a surface, which is not readily available without considerable manual effort. Therefore, they are unsuitable for systems that need to automatically capture reverberation properties for spatial rendering of early reflections (e.g., in AR rendering).
[0052] Without prior determination, the assignment of absorption parameters to surfaces for early reflection rendering is arbitrary unless the user provides additional information (such as images). With arbitrary assignment of absorption to walls, the acoustic effects of early reflections typically do not work as intended; for example, strong reflections from soft (absorbent) walls may be perceived, or vice versa, hindering the immersion of virtual sources in the room. Good audio quality can be achieved when the user provides additional information, but the user must take additional steps to make the acoustic effects of early reflection synthesis work.
[0053] The concepts discussed in the embodiments described in detail below relate to the reproduction of early reflections of reverberation in a (virtual acoustic) rendering system, and provide an apparatus and method that can automatically determine the absorption coefficient for the surface of a room based on the spatial room impulse response (SRIR) measured in the room and the shape of the room. These embodiments can then render early reflections having acoustic properties that match the acoustic properties of early reflections in the room.
[0054] In some embodiments, this can be achieved by: Determining reverberation time (e.g., RT60 hours) from SRIR, Based on RT60 and the interior shape, determine at least one interior absorption area. Based on the interior shape, determine the position of the image source to correspond to premature reflection from the surface. Based on the position of the image source, determine the time and direction of arrival of the early reflection to the SRIR. Applying beamforming to SRIR to extract at least one reflection, Extracting the reflection coefficient from the reflection for at least one surface, Assigning an absorption coefficient to at least one surface based on the reflectance coefficient and absorption area, and Render early reflections using the determined absorption coefficient.
[0055] Embodiments discussed herein attempt to offer advantages over conventional approaches by allowing the absorption coefficient for interior surfaces to be determined as an automated process using SRIR and the shape of the interior. Therefore, the embodiments discussed are suitable for systems requiring automated processing, such as augmented reality (AR) rendering.
[0056] As a result, realistic rendering of early reflections can be achieved by using the embodiments discussed herein. For example, using the proposed method, a user can "scan" the room in which they are located and then have realistic augmented audio objects rendered for them.
[0057] When realistic rendering of early reflections is achieved using the proposed embodiment, users can experience immersive AR audio scenes in which virtual sources can be reproduced in the AR audio scene as if they were real physical sources in space.
[0058] As another example, the embodiments described herein may be used on the encoder side of an MPEG-I audio rendering system, such that a virtual space (or physical space) is measured to provide SRIR, and this SRIR information is subsequently provided to an MPEG-I audio encoder, which derives acoustic parameters therefrom, and these parameters are written to a bitstream and used to render audio according to the input spatial acoustics.
[0059] Figure 2 shows a schematic diagram of an apparatus suitable for carrying out several embodiments. The apparatus or system may include an absorption coefficient determination unit 203 configured to receive room shape information 208 and spatial room impulse response information 206 as inputs. The absorption coefficient determination unit 203 is configured to generate absorption coefficients 210 and pass them to an early reflection renderer 203.
[0060] Space-indoor impulse response 206 h SRIR (n,j) may include an impulse response with J channels measured in a room (where n is the sample time and j is the impulse response channel). For example, the J channels may be channels of a (first- or higher-order) ambisonic room impulse response (which can be obtained by measuring the impulse response using a dedicated microphone such as an Eigenmike).
[0061] Interior shape 208 includes an indexed list of i=1,...,I planar surfaces describing the walls, floors, and ceilings of the interior. Each such surface has a surface area S i It has a total interior area of S tot =Σ i S i That is the case.
[0062] The indoor shape 208 can further include metadata indicating the positions of the receiver and the source within the indoor shape, as well as optionally the orientation and rotation, while the SRIR measurement is being performed.
[0063] The indoor shape 208 can further include metadata indicating the source directivity (if available). The directivity metadata can provide the source directivity pattern D(θ,k), where θ is the angle with respect to the defined (main) axis of the speaker.
[0064] The absorption parameter 204, in some embodiments, includes the characteristics of each surface of the indoor shape and is described by the average absorption coefficient α i (k) at k = 1,...,K frequency bands. In some embodiments, the coefficients are described with respect to octave bands having center frequencies [125, 250, 500, 1000, 2000, 4000]. The absorption coefficient determiner 201 can thus determine the surface absorption coefficient 210 α i (k) for each i-th surface in the room.
[0065] The apparatus can further include an early reflection renderer 203. The early reflection renderer 203 is configured to receive as input the audio signal x(n) 200, the indoor shape 208, the listener position 202, the source position 204, and the absorption coefficient α i (k) 210. The early reflection renderer 203 is configured to render as output a reverberant binaural signal x bin (n,j bin ) 212 that includes early reflections that are perceived as if a sound source actually exists in the room where the spatial indoor impulse response 202 was measured. In the following example, early reflection rendering is described in detail, but the direct sound and late reverberation rendering can be considered separately, and any suitable direct sound and / or late reverberation rendering method can be used.
[0066] Regarding Figure 3, an illustrative flowchart of the operation of the system shown in Figure 2 is presented.
[0067] Therefore, as shown in 301, the spatial intra-intra
[0068] Subsequently, as shown in 303, the absorption coefficient is determined from the spatial intra-intra
[0069] Additionally, there are operations to acquire the audio signal as shown by 305, to acquire the listener's position as shown by 307, and to acquire the source position as shown by 309.
[0070] Subsequently, after determining the absorption coefficient 303, obtaining the room shape, audio signal, listener position, and source position, there is the generation of a reverberant (binaural) audio signal for early reflections, as shown in 311.
[0071] Subsequently, as shown in 313, there is an output of the generated reverberation (binaural) audio signal for the early reflections.
[0072] Figure 4 shows a schematic diagram of the absorption coefficient determination system 201.
[0073] The absorption coefficient determination unit 201 determines the spatial intra-intra-space impulse response 206 h SRIR It is configured to receive (n,j) and the interior shape 208 as input.
[0074] The absorption coefficient determination unit 201 determines the spatial intra-intra-space impulse response 206 h SRIR The RT60 determinator 401 is configured to receive (n,j), and from these, the reverberation time 412 T in at least one frequency band is provided. 60 Determine (k) (where k is the frequency band index). The frequency band can be any suitable bandwidth distribution. For example, the frequency band could be the octave band, the ERB band, or the Bark band.
[0075] In some embodiments, this can be achieved by decomposing the spatial intra-intra-intra-spatial impulse response (SRIR) 206 into several subbands determined by desired parameterization. For example, this decomposition could be octave bands [125, 250, 500, 1000, 2000, 4000]. Thus, the SRIR is decomposed into six separate SRIRs using an appropriate octave band filter bank, h SRIR This yields (n,k,j), where k is the subband index. The remaining parameters are calculated independently across the frequency bands, and therefore the process is explained for one of the subbands.
[0076] Omnidirectional representation of RIR, for example, the first channel of a B-format SRIR, h omni (n,k=h SRIR Assuming there is access to (n,k,1), the reverberation time T 60 (k) can be processed by omni RIR. Estimation is h omni The energy decay curve EDC(n,k) for the analyzed subband, generated by inverse integration of (n,k):
[0077]
number
[0078]
number
[0079] N is the total length of the SRIR in the sample. Since the examples and embodiments deal with measured RIRs, there is likely to be measurement noise in the SRIR that could bias the estimation of these parameters. Noise power can be estimated from the environment or from the microphone performing the measurement recording and is assumed to be uncorrelated with the additive, stationary, and clean RIRs. EDC(n,k) = EDCRIR (n,k)+EDC noise (n,k) It brings about.
[0080] Noise is power P per sample noise Since (k) is assumed to be stationary, the modeled noise EDC is linear, EDC noise (n,k)=(Nn)P noise (k) It can be given by, On the other hand, RIR's EDC is type
[0081]
number
[0082]
number
[0083] Since the RIR attenuates to below the noise floor after a while, the noise EDC can only be calculated for some later segments of the measurement, for example, for t > 1 second in a normal room. It can then be assumed that the EDC for this segment is primarily dominated by the noise term, and a linear equation can be fitted to extrapolate the noise EDC over the entire range of the RIR. Denoised EDC RIR (n,k) is the modeled noise EDC noise It can then be estimated by subtracting (n,k).
[0084] Finally, by using the denoised EDC and expressing the EDC in decibels, the overall T 60(k) is estimated, and the exponential form is converted to a linear form. The line is then fitted to the denoised EDC, starting from a point, for example, -5dB to -10dB from the maximum EDC, avoiding the early part (which usually deviates from the exponential model of the later part). The endpoint can be taken at -20dB or -30dB from the first point, and the corresponding time interval between the two is tripled or doubled, respectively, to T 60 Obtain the scaling to (k).
[0085] In some embodiments, the absorption coefficient determination unit 201 includes an absorption area determination unit 415. The absorption area determination unit 415 determines the reverberation time 412 T 60 The absorption area determiner 415 is configured to receive (k) and the interior shape 208. The absorption area determiner 415 is then configured to determine the total absorption area 414 A(k) for at least one frequency band. The total absorption area 414 is given by the sum A(k) = Σ i a i (k)·S i This corresponds to the following. As input to the method, the shape of the room is the surface area S. i S tot This is required in the form of the total room volume V, which is potentially calculated from the room dimensions. The method for determining these parameters depends on the type of room shape data 208 and the shape of the room. For example, the Sabine formula may be used:
[0086]
number
[0087] Alternatively, an eye-ring type may be used.
[0088]
number
[0089] Alternatively, if the room is rectangular, the Fitzroy formula is
[0090]
number
[0091] In other cases, other methods may be used, such as numerical methods that can return an estimate of the total absorption area from the room shape and reverberation time, as described in Lehmann, EA, and Johansson, AM (2008), "Prediction of energy decay in room impulse responses simulated with an image-source model," Journal of the Acoustical Society of America, 124(1), 269-277.
[0092] In some embodiments, the absorption region determiner 415 can return a highly accurate estimate of the absorption area through a brute-force search by simulating, for example, about 100 all image source omni RIRs, each having 100 different uniform absorption coefficients for every 0.01 value, and can select the one that gives the RT60 closest to the measured value.
[0093] In all of the above cases, the average absorption coefficient per band.
[0094]
number
[0095] The absorption coefficient determinator 201 may include an image source positioning unit 405 (or an isolated echo detector). The image source positioning unit 405 is configured to receive the room shape 208 and refine the initial uniform absorption coefficient by using knowledge of the shape and spatial information in the SRIR to extract individual absorption profiles corresponding to isolated echoes based on these inputs.
[0096] In some embodiments, isolated echoes are defined as early echoes that occur for a short time, e.g., within a window of about 1.5 milliseconds, without the presence of other echoes arriving within the same time window. Typically, these echoes are mostly second-order reflections. Processing isolated echoes in SRIR produces more robust absorption estimates than those obtained by processing time windows in which mixed echoes from different directions exist.
[0097] Therefore, rather than attempting to blindly detect the presence of these isolated echoes solely from the SRIR signal, prior information obtained from the room geometry is used. In some embodiments, the position of the image source corresponding to such reflections is calculated through an image source positioner 405. The image source positioner 405 thus comprises an image source calculator that determines the image source and returns the coordinates of the image source relative to the first-order (and potentially higher-order) reflections from each surface. The image source calculator can be implemented in any suitable way and is not described in further detail. However, several well-known methods are presented, for example, in J. B. Allen and D. A. Berkley, "Image method for efficient simulating small-room acoustic," J. Acoust. Soc. Am., Vol. 65, pp. 943-950, April 1979, and in J. Borish, "Extension of the image model to arbitrary polyhedra," Journal of the Acoustical Society of America 75.6(1984):1827-1836.
[0098] In the image source method, the source position is mirrored against each reflective surface of the room shape 208 in order to acquire the image source.
[0099] Figure 8 shows an example of mirroring. In the example shown in Figure 8, the interior shape is a box-shaped or rectangular space having reflective surfaces 800, 802, 804, and 806. A source 820 and a listener 810 are located within the space. The direction of early reflection between source 820 and listener 810 is shown when the reflection point and / or absorption point 840 is on the reflective surface 806 between source 820 and listener 810. Mirroring of source 820 with respect to reflective surface 806 can be used to establish image source 830. A line connecting image source 830 to listener 810 can then be used to establish the reflection and / or absorption point 840 and the direction of arrival (DOA) of the reflection relative to the listener. The delay applied to synthesize the early reflection is obtained based on the distance of the reflection path (the path from the image source to the listener, equal to the length of the path from source 820 to listener 810). The absorption is the absorption α corresponding to the reflective surface 806 (reflection and / or absorption point 840) to which this early reflection is reflected. i (k) is obtained. Distance attenuation is set proportionally to 1 / distance, where the distance is equal to the length of the reflection path from the source to the listener. Air absorption may also be added to the attenuation. The DOA of early reflection is set based on the angle of arrival (DOA) from the reflection point to the listener.
[0100] Primary reflections occur from a single wall, while higher-order reflections occur from multiple walls. Higher-order reflections can be acquired using a higher-order image source that is mirrored sequentially for each reflective surface.
[0101] Therefore, in some embodiments, the output of the image source positioning unit 405 is [r0,r1,...,r I ,r 1,1 ,...,r 1,I ,r i,i,... This is a list of image source locations such as ], r i,i... =[xi,i ,...,y i,i ,...,z i,i,... ] represents the coordinates of the image source reflected by the i-th subsequent surface in each order of reflection. For example, r0=[x0,y0,z0] is the coordinate of the actual source relative to the receiver, and r1=[x1,y1,z1] is the coordinate of the first reflection reflected from surface i=1.
[0102] For example, in a rectangular room, the primary reflections are [r1,r2,r3,r4,r5,r6]. In some embodiments, the direct path is [r1,...,r I ,r1,1,...,r 1,I ,...,r I,I Only reflections up to the secondary reflection, excluding ], are determined. In some embodiments, the list of distances of the image source [d1,...,d I d 1,1 ,...,d 1,I ,...,d 1,I ] was decided, d i,i,... ha||r i,i,... In some embodiments, each image source in the list is iterated over, and by assuming an internal standard speed of sound such as c = 343 m / s, an acoustic distance window of approximately 0.69 m with length L = 2 milliseconds * 343 m / s centered on the position of the image source is determined.
[0103] In some embodiments, if any other image source distance in the list falls within its window, all such images are removed from the list. The remaining image sources that pass this test are the isolated echoes detected.
[0104] In some embodiments, up to two simultaneous reflections may be permitted within a single window centered on each of the two reflections, as long as they are sufficiently spatially separated. In some embodiments, this is related to the angle between the two reflections.
[0105]
number
[0106] γ i,i’ If the angle is >90°, those image sources are considered to be spatiotemporally isolated.
[0107] List of detected isolated image sources P iso After the initial selection is determined, the final selection stage is applied based on the relationship between their absorption coefficients. Each image source corresponds to the product of the absorption coefficients of the surface from which it is reflected. For example, if the selection of the isolated list is P iso =[r2,r 1,2 ,r4,r 6,3 If ] (ordered based on their distance from the receiver), then it corresponds to the absorption coefficient [a2, a1·a2, a4, a6·a3]. From the secondary image sources, only those image sources that have at least one common absorption coefficient with one of the primary isolated image sources in the list are retained. For example, since image source r2 is also in the list, from the list of the example above, image source r 1,2 Only is retained, and the absorption coefficients a1 and a2 can be calculated. Since neither primary image source r6 nor r3 is in the list, the image source r 6,3 It is not retained. Therefore, the final isolated image echo list after this stage is
[0108]
number
[0109] This can then be passed to the source direction determination unit 407.
[0110] In some embodiments, the absorption coefficient determiner 201 includes a source directivity determiner 407 configured to receive the room shape 208 and the image source location and to incorporate the influence of arbitrary source directivity into the process of selecting isolated echoes to be processed.
[0111] For example, if the speaker used during measurement is not omnidirectional, some echoes may be attenuated too much due to source directivity, making it useful for estimating the absorption profile.
[0112] Therefore, assuming that the source directivity is known, for example, from either the speaker manufacturer's specifications or measurement results, this information can be taken into account to exclude suppressed echoes. In some embodiments, the source directivity is assumed to be nearly axially symmetric. In these embodiments, the directivity may be described as D(φ,k), where φ is the angle with respect to the speaker's principal axis. The principal axis is usually oriented in the speaker's "line of sight" direction, in other words, the direction in which the main driver is directed. Furthermore, the orientation of the speaker may be represented using a vector u representing the direction of the principal axis.
[0113] In some embodiments, the speaker used to capture the SRIR is oriented toward the microphone, and therefore u0 = -r0 / ||r0|| is opposite the speaker's DOA. The direction of radiation (DOR) θ0 = u0 from the source to the receiver coincides with the orientation of the speaker, the angle between the two vectors is zero, and φ0 = ∠(θ0, u0) = 0.
[0114] i-th echo D(φ i To evaluate the directional gain of ,k), the source directivity determiner 407 determines the orientation of the image source speaker u i And each DORθ i It is possible to determine that. DOR is simply θ i =-r i / ||r i||. The new direction of directivity can be given by reflecting vector u0 with respect to the reflective surface of the i-th image source. In the case of a rectangular room where the walls are parallel to the xy, yz, and zy planes, the direction of the image source is,
[0115]
number
[0116] Finally, after calculating the orientation and DOR of the image source, the directional angle is φ i =∠(θ i ,u i ) = acos(θ i ·u i ) is provided as, and the directivity value for the k-th band is D(φ i D(φ) is the case. The decay is compared to the threshold τ. i If τ = 0.5, then the image source is the echo list P iso Discarded if not retained otherwise.
[0117] These image source locations 404 can be passed to the arrival time and DOA estimator 409.
[0118] In some embodiments, the absorption coefficient determiner 201 includes a sound speed estimator 403. The sound speed estimator 403 is configured to receive the room shape 208 and the spatial room impulse response 206 and generate an estimated sound speed 402 therefrom. Knowledge of the sound speed c in the measurement results is important for combining the information from the room shape 208 and the measured SRIR 206. Assuming that the distance d0 between the source and the microphone is known and the system delay has been removed from the measurement results, the time taken for the direct sound to reach the microphone can be determined by finding the sample of the maximum peak of the amplitude |h OMNI (n)|. To increase the accuracy, the response can be oversampled, for example, 8 times. In this case, the travel time of the impulse from the source to the receiver is t0 = n direct / (8f s ), where n direct is the sampled index of the maximum peak and f s is the original sampling rate.
[0119] The sound speed during the recording conditions can then be determined as c est = d0 / t0 .
[0120] This estimated value 402 of the estimated sound speed can be passed to the arrival time and DOA estimator 409.
[0121] In some embodiments, the absorption coefficient determiner 201 includes an arrival time and DOA estimator 409. The arrival time and DOA estimator 409 is configured to receive a list of isolated echo distances from the image source positioner 405 via the source directivity determiner 407 in order to focus on the correct spatio-temporal region within the SRIR and refine the time and spatial position estimates of each echo.
[0122] List
[0123]
Number
[0124] Furthermore, in some embodiments, oversampling may be performed, for example, by 8 times, to improve accuracy, and the echo time is ti=n peak / (8fs) That is the case.
[0125] The direction of arrival (DOA) of each echo can also be refined in the same way.
[0126]
number
[0127]
number
[0128]
number
[0129]
number
[0130] The pseudo-spectrum can be calculated for each of these limited sets of points, for example, by using a beamforming steering response power, or a subspace method such as Multiple Signal Classification (MUSIC). In some embodiments, the index of the grid point having the maximum value is the refined DOA of the echo i is shown.
[0131] In some further embodiments, if up to two simultaneous reflections are allowed within a single window centered on each of two, provided they are spatially well separated based on the previously described condition γ i,i’ > 90°, the DOA refinement process is performed twice within a spatial window centered on each of the two DOAs. Two maximum peaks, one in each spatial window, are then selected, and the indices of the two grid points are the refined DOAs of the two echoes i , DOA i’ is shown.
[0132] The arrival time and DOA estimator 409 is configured to output the estimated values 406 of the refined DOA and arrival time to the reflection beamformer 411.
[0133] The absorption coefficient determiner 201 includes, in some embodiments, the reflection beamformer 411. The reflection beamformer 411 is configured to better isolate the echo components in the measured SRIR after refinement of the spatio-temporal information of each isolated echo. For this purpose, any suitable beamforming structure having no distortion with respect to the DOA of the echo can be used. The resulting beamforming filter w(n,DOA) is then used to extract the beamformed SRIR as
[0134]
Number
[0135] For first-order ambisonic (FOA) or higher-order ambisonic (HOA) SRIR, the beamforming weights are: w j (n,DOA=Y nm (DOA) Simplified to, Y nm (DOA) is the nth-order, m-degree real spherical harmonic value corresponding to the ambisonic channel j.
[0136]
number
[0137] The beamformed SRIR is further passed through an octave band filter bank similar to that used in the reverberation time determination 401 (RT60 estimation stage) in some embodiments. The subband-resolved beamformed SRIR is then subjected to h for K subbands. beam (n,k). After removing the group delay derived by the filter bank, for example, the refined echo arrival time t i By selecting a time window of approximately 2 milliseconds centered on this point, the subband echo signal can be extracted.
[0138]
number
[0139] In some embodiments where two spatially isolated reflections are allowed within the same time window, the beamforming weights are adjusted so that each of the two reflections has a distortion-free constraint in one direction and a null in the other. The beamforming operation is repeated twice, swapping the null constraint and the distortion-free constraint, to extract the signals from the two echoes.
[0140] In some embodiments, the absorption coefficient determination unit 201 includes a reflection coefficient extractor 413, which is configured to receive refined extracted isolated echo signals, but the refined extracted isolated echo signals are decomposed into subbands and subsequently used to estimate the average power reflection coefficients βi(k) within the subbands, β i (k) = 1 - a i (k) is linked to or associated with the absorption coefficient.
[0141] In some embodiments, each echo is assumed to be an alteration of the direct sound signal, modified by the absorption filter of the surface it is reflected from and attenuated by the propagation law inversely proportional to distance. Furthermore, assuming that the absorption coefficient is nearly invariant within each subband, the relationship is:
[0142]
number
[0143]
number
[0144] direct sound h direct The (n,k) subband signal is extracted in the same way as a single echo, but without beamforming operation, for example, the omnidirectional signal h omni Extracted by selecting a time window around the direct sound arrival time t0 for (n,k).
[0145] Extraction is a list
[0146]
number
[0147]
number
[0148] In embodiments where source directivity D(θ,k) is taken into consideration and directivity values for all image sources in the list are calculated in a previous echo selection stage,
[0149]
number
[0150]
number
[0151] The extracted absorption coefficient 410 is then passed to the absorption coefficient refiner 417.
[0152] In some embodiments, the absorption coefficient determination unit 201 includes an absorption coefficient refiner 417. In some embodiments, the absorption coefficient refiner 417 is configured to receive the total absorption area 414, the reflectance coefficient 410, and the room shape 208, and to generate a refined absorption coefficient 210.
[0153] In other words, the reflection coefficient β for surfaces I' ≤ I i (k) After the list of 410 is estimated, each a i (k) = 1 - β i The absorption coefficient of (k), 210, is assigned to each I' surface. The newly assigned coefficient is the initial uniform absorption distribution.
[0154]
number
[0155]
number
[0156]
number
[0157]
number
[0158] The total absorption area estimated from the reverberation equation may deviate from the true total absorption area depending on the room characteristics and the reverberation equation used; therefore, the average absorption coefficient should be set to a reasonable value, for example,
[0159]
number
[0160]
number
[0161] In some embodiments, the average absorption coefficient is the remaining non-isolated echoes.
[0162]
number
[0163]
number
[0164]
number
[0165] The resulting absorption coefficient is 210 α i (k) is then output.
[0166] Regarding Figure 5, a flowchart of the exemplary absorption coefficient determination system shown in Figure 4 is presented.
[0167] Therefore, Figure 5 shows that, according to 501, the spatial intra-intra
[0168] Subsequently, Figure 5 shows the estimated speed of sound, as shown by 503.
[0169] Furthermore, Figure 5 shows the determination of the image source position, as indicated by 505.
[0170] Following the determination of the image source position, Figure 5 shows the correction due to source orientation, as shown by 507.
[0171] Subsequently, Figure 5 shows the estimated arrival time and direction of arrival, as indicated by 509.
[0172] Subsequently, the reflections can be beamformed by 511 to further extract echoes from the SRIR, as shown in Figure 5.
[0173] By 513, the reflection coefficient can then be extracted, as shown in Figure 5.
[0174] Furthermore, the determination of the reverberation time (RT60) is shown in Figure 5 by 502.
[0175] The reverberation time, along with the room geometry, is then used to help determine the absorption area, as shown in Figure 5 by 508.
[0176] Subsequently, the absorption coefficient can be refined by 515, as shown in Figure 5.
[0177] The absorption coefficient can then be output by 517, as shown in Figure 5.
[0178] With respect to Figure 6, schematic diagrams of early reflection renderers 203 suitable for several embodiments are shown in more detail. Early reflection rendering uses the room shape 208, listener position 202, source position 204, and absorption coefficient 210 as inputs to synthesize a reverberation binaural signal 212 (at least one early reflection signal) using an audio signal 200. The input audio signal 200 is first fed to a delay line 603, which buffers past audio signal samples and allows for the selection of past sample regions of the audio signal 200.
[0179] The early reflection renderer further comprises an early reflection parameter determiner 601 configured to determine parameters for one or more early reflections using the room shape 208, listener position 202, source position 204, and absorption coefficient 210. There are several methods for calculating or simulating early reflections.
[0180] There are several methods for calculating or simulating early reflections. For example, the image source method can be used, as discussed in J. Allen and D. Berkley, "Image method for efficient simulating small-room acoustic," J. Acoust. Soc. Am., Vol. 65, pp. 943-950, April 1979, and in J. Borish, "Extension of the image model to arbitrary polyhedra," Journal of the Acoustical Society of America 75.6(1984):1827-1836.
[0181] The exemplary early reflection renderer 203 shown in Figure 6 is configured to generate control parameters such as delay 600, absorption 602, attenuation 604, and direction of arrival (DOA) 606, and to pass these to a processor which will be described later.
[0182] The delay 600 applied to synthesize the early reflection is, in some embodiments, obtained based on the distance of the reflection path (the path from the image source to the listener). The absorption is the absorbance α corresponding to the reflective surface to which this early reflection was reflected. i (k) is obtained from the given value. Attenuation is set proportionally to 1 / distance, where distance is equal to the distance traveled by early reflection. Air absorption can also be simulated by appropriate low-pass filtering, which will attenuate high frequencies depending on the distance traveled by reflection.
[0183] The early reflection signal acquirer 605 can receive the output of the delay line 603 and the delay parameter 600. The early reflection signal acquirer is configured to acquire a delayed signal by acquiring past signal samples based on the delay 600.
[0184] The early reflection absorption processor 607 then filters the selected past signal samples and applies an equalizer filter to absorb the early reflection α. i (k) Data 602 can be modeled to obtain a delayed and absorbed filtered signal.
[0185] The early reflection attenuation processor 609 can obtain a delayed, absorbed, and attenuated signal by subsequently attenuating the delayed and absorbed filtered signal by applying 1 / distance attenuation and optionally air absorption based on the attenuation parameter 604.
[0186] Finally, the early reflection spatializer 611 may be configured to spatialize the delayed, absorbed, filtered, and attenuated signal using HRTF filtering with left and right HRTF filters corresponding to the desired DOA 606 for this early reflection to obtain a reverberant binaural signal 212 that includes the synthesized early reflection portion. In some situations, the early reflection spatializer 611 may implement a binauralizer to generate the binaural signal 212.
[0187] Regarding FIG. 7, a flowchart showing the operation of an exemplary early reflection renderer 203 according to some embodiments is shown.
[0188] Thus, as shown in FIG. 7 by 701, an audio signal, listener position, source position, room shape, and absorption coefficient are obtained.
[0189] Thereafter, as shown in FIG. 7 by 703, the early reflection parameters of delay, absorption, attenuation, and DOA are determined.
[0190] Furthermore, as shown in FIG. 7 by 705, a delay line is applied to the audio signal.
[0191] After this, as shown in FIG. 7 by 707, an early reflection audio signal is obtained.
[0192] Furthermore, as shown in FIG. 7 by 709, early reflection absorption is applied to the early reflection audio signal.
[0193] As shown in FIG. 7 by 711, early reflection attenuation is applied to the absorbed early reflection audio signal.
[0194] Thereafter, as shown in FIG. 7 by 713, early reflection spatialization is applied to the attenuated and absorbed early reflection audio signal.
[0195] Subsequently, unit 715 outputs a reverberation binaural audio signal, as shown in Figure 7.
[0196] In some embodiments, for a given source-listener position pair, there may be no primary reflection for some room shapes. In such situations, the absorption coefficient cannot be obtained from the primary reflection using the embodiments discussed above. However, the absorption coefficient a of other surfaces can be obtained. i The output may be generated using the average of (k). In some embodiments, secondary reflections may be used to calculate the absorption coefficient, as presented above for a directional source.
[0197] In the example presented above, a first-order ambisonics impulse response was used as an example of a spatial-intra-room impulse response. In other embodiments, other types of impulse responses may be used. For example, the impulse response may be from a microphone array attached to a mobile device. In such examples, some details of the embodiment may differ. For example, DOA analysis may be optimized for mobile devices. Thus, the method presented in UK Patent Application GB1619573.7 can be used for the analysis of DOA in a frequency band. Although the method in UK Patent Application GB1619573.7 is presented for continuous audio, in this situation, the impulse response h SRIR (n,j) (or a variant H in its time-frequency domain) SRIR This can be applied to (b,n,j)).
[0198] In the exemplary embodiments presented above, the calculation was implemented in frequency band k. In some embodiments, the calculation may also be performed in bin b, and the results may be averaged over subbands in a subsequent step.
[0199] Figure 9 schematically shows an exemplary system in which an embodiment may be implemented. The system includes an encoder device 1901 as part of a server computer 911, which writes data to a bitstream 1921 and transmits it to a renderer / playback device 941, which decodes the bitstream, performs reverberation processing according to the embodiment, and outputs audio for listening with headphones.
[0200] The encoder side 901 in Figure 9 may be located on a content creator computer and / or a network server computer. The encoder output is a bitstream 921 made available for download or streaming. The decoder / renderer function runs on a playback device 941, which may be a mobile device, personal computer, soundbar, tablet computer, car media system, home HiFi or theater system, AR / VR head-mounted display, smartwatch, or any system suitable for audio reproduction.
[0201] Encoder 901 is configured to receive a virtual scene description 900 and an audio signal 904. The virtual scene description 900 may be provided in MPEG-I encoder input format (EIF) or in another suitable format. Generally, the virtual scene description encompasses an acoustic-related description of the content of the virtual scene, including, for example, the scene shape as a mesh, acoustic materials, an acoustic environment with reverberation parameters, the location of the sound source, and other audio element-related parameters such as whether reflections are rendered for the audio elements. In some embodiments, encoder 901 includes a scene and reverberation payload encoder 913 configured to generate reflection parameters.
[0202] The encoder 901 further comprises an MPEG-H 3D audio encoder 914 configured to acquire audio signals 904, MPEG-H encode them, and pass them to a bitstream encoder 915.
[0203] In some embodiments, encoder 901 further comprises a bitstream encoder 915 configured to receive the output of scene and reverberation payload encoder 913 and encoded audio signals from MPEG-H encoder 914, and to generate a bitstream 921 which can be passed to bitstream decoder 951. In some embodiments, bitstream 921 may be streamed to an end-user device or made available for download or storage.
[0204] In some embodiments, the decoder / renderer includes a bitstream decoder 951 configured to decode a bitstream.
[0205] The decoder / renderer may further include a scene decoder 953 configured to take encoded reverberation parameters and decode them in the opposite or reverse operation to the reverberation payload encoder 913.
[0206] Furthermore, the head pose generator 957 receives information from the head-mounted device 970 and other sources and generates head pose information or parameters that can be passed to the early reflex renderers 990 / 203 and the direct sound binaural renderer 963.
[0207] Decoder 941 includes an MPEG-H 3D audio decoder 954 configured to decode audio signals and pass them to the early reflection renderers 990 / 203 and the direct sound processing 965.
[0208] Furthermore, the decoder / renderer 941 includes a direct sound processor 1965, which is configured to receive a decoded audio signal and implement any direct sound processing such as air absorption and distance gain attenuation, and can pass it to a direct sound binaural renderer 1963, which can generate a direct sound component using head orientation determination (from appropriate sensors), and the direct sound component, along with the reverberation component, is passed to a binaural signal combiner 967. The binaural signal combiner 967 is configured to combine the direct part and the early reflection part to produce an appropriate output (for example, for headphone reproduction).
[0209] Furthermore, in some embodiments, the decoder includes a head orientation determiner that passes head orientation information to a head pose generator 1957.
[0210] In some embodiments, the output is a multi-channel speaker setup (such as a 5.1 or 7.1+4ch multi-channel speaker setup). In this case, the proposed process determines the speaker positions of the actual speakers in the speaker setup θ. ls (i), φ ls (i) can be used and modified by omitting the binaural renderer and reproducing the reverberation audio signal from the corresponding speakers in the speaker setup.
[0211] Referring to Figure 9, in the case of speaker output, instead of the early reflection renderers 990 / 203, there is a speaker renderer (or panner), which in its simplest form simply passes the speaker signal to a speaker signal combiner, which would replace the binaural signal combiner 967. Accordingly, the direct sound portion and the early reflection portion are spatialized using a panner such as VBAP implemented instead of the binaural processor.
[0212] Embodiments can be applied, for example, to AR rendering. In this case, the inputs to the renderer are indoor shape information and SRIR, and the above-described embodiments are executed on a playback device. SRIR can describe acoustic measurement results associated with the room in which the listener is listening to the audio scene. In addition to reverberation rendering, the playback device can perform direct sound rendering and then combine (mix) the direct sound portion and the reverberant sound portion. Playback can be head-tracked such that the rendered audio rotates and moves according to the user's position in the space where the reverberation is rendered.
[0213] In the example shown with respect to FIG. 9, only early reflection rendering is depicted. The overall system can also include a diffuse (late) reverberation renderer that synthesizes the late reverberation portion. The synthesized late reverberation portion can be binauralized and added to the binauralized early reflection portion. In some embodiments, the speaker output is generated via panning instead of binauralization using HRTF.
[0214] In some embodiments, SRIR and indoor shape information, including the source position and receiver position of the SRIR measurement, can be included in the listening space description for the renderer. The purpose of the LSDF interface is to provide appropriate information so that the renderer can render content that matches the real environment for augmented reality and mixed reality applications.
[0215] LSDF can provide speaker setup parameters that can include the directivity, position, and orientation of the speakers or other sound sources used during SRIR measurement.
[0216] LSDF can include the following elements for describing the listening space (see MPEG-I Immersive Audio Augmented Reality Listener Space Description Format, version 2).
[0217] Table 1
[0218] Table 1, <audioscene>This declares a listening space.
[0219] <mesh>It can be used to describe the shape of the listening space (indoor). It includes surfaces and vertices.
[0220] <acousticmaterial>This can be used to explain prior information about wall materials.
[0221] <acousticenvironment>It can retain RT60 and the reverberation-to-direct sound ratio parameter (if available).
[0222] <aranchor>This can be used to match virtual content (encoder input format, EIF content) to the listening space.
[0223] Furthermore, new entries <roomimpulseresponse>The following may be added. An example description may be as follows:
[0224] <acousticenvironment>and <acousticparameters>This is as defined in the MPEG-I encoder input format (EIF), but there are differences as explained below. <acousticenvironment>teeth, <acousticparameters>or <roomimpulseresponse>It must include one of them, not both.
[0225] [Table 2]
[0226] To illustrate the acoustic effects of the listening space, instead of acoustic parameters, recordings of the spatial-indoor impulse response may be provided. The measurement data <roomimpulseresponse>It is held there.
[0227] [Table 3] TIFF2026518109000039.tif83169
[0228] In LSDF, SRIR can be described, using the example above, as a B-format or ambisonics signal containing metadata describing the position and orientation of the microphone (microphone or receiver). Furthermore, the position and orientation relative to the source (the speaker used during the measurement) are provided. Also included is speaker directivity data, for example, as a SOFA format file, binary file, or other file containing directivity coefficients, or as a pointer to a directivity pattern file from which directivity information can be obtained. Such directivity information may, in some embodiments, be used to compensate for the effect of directivity in the calculation of the reflection coefficient.
[0229] Since the proposed analysis method generates an absorption coefficient for a wall, it should be noted that it is also possible to define an LSDF interface to receive walls with frequency-dependent absorption coefficients. In this case, the function of receiving SRIR and then analyzing the absorption coefficient would be the reference part of the standard, while the normative part would begin with the LSDF interface.
[0230] It should be noted that the LSDF interface and / or EIF interface described above may provide further information to improve the spatial rendering of reverberation. One example is providing RT60 values and / or DDR / DSR / RDR values in a space-dependent manner. That is, there may be multiple RT60 values and / or DDR / DSR / RDR values for each spatial direction within the AcousticEnvironment. For example, in some embodiments, there may be four RT60 values and / or DDR / DSR / RDR values for each of four spatial directions within the AcousticEnvironment, such as azimuth angles of 0 degrees, 90 degrees, -90 degrees, and 180 degrees. These values may be used to adjust the directional reverberation characteristics for the AcousticEnvironment reverberator. One method that may be used is to adjust the directional output gain for the reverberator so that the directional RT60 value can be approximated.
[0231] With respect to Figure 10, an exemplary electronic device which may be used as one of the device components of the system as described above. The device may be any suitable electronic device or apparatus. For example, in some embodiments, device 2000 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. The device may be configured to implement, for example, an encoder, renderer, or any functional block as described above.
[0232] In some embodiments, the device 2000 comprises at least one processor or central processing unit 2007. The processor 2007 may be configured to execute various program code, such as those described herein.
[0233] In some embodiments, device 2000 comprises memory 2011. In some embodiments, at least one processor 2007 is connected to memory 2011. Memory 2011 can be any suitable storage means. In some embodiments, memory 2011 includes a program code section for storing program code that can be implemented on processor 2007. Furthermore, in some embodiments, memory 2011 may further include a stored data section for storing data, for example, data processed or to be processed according to embodiments such as those described herein. The implemented program code stored in the program code section and the data stored in the stored data section can be retrieved by processor 2007 as needed via memory-processor coupling.
[0234] In some embodiments, device 2000 includes a user interface 2005. In some embodiments, the user interface 2005 may be coupled to a processor 2007. In some embodiments, the processor 2007 can control the operation of the user interface 2005 and can receive input from the user interface 2005. In some embodiments, the user interface 2005 may allow a user to input commands to device 2000, for example, via a keypad. In some embodiments, the user interface 2005 may allow a user to retrieve information from device 2000. For example, the user interface 2005 may include a display configured to show information from device 2000 to the user. In some embodiments, the user interface 2005 may include a touchscreen or touch interface capable of both allowing information to be input to device 2000 and, further, displaying the information to the user of device 2000. In some embodiments, the user interface 2005 may be a user interface for communication.
[0235] In some embodiments, device 2000 includes an input / output port 2009. In some embodiments, the input / output port 2009 includes a transceiver. In such embodiments, the transceiver may be coupled to a processor 2007 and configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. The transceiver, or any suitable transceiver, or transmitter and / or receiver, may, in some embodiments, be configured to communicate with other electronic devices or devices via a wired or wired connection.
[0236] The transceiver can communicate with further devices by any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an Infrared Data Pathway (IRDA).
[0237] Input / output port 2009 may be configured to receive signals.
[0238] In some embodiments, device 2000 may be used as at least part of a renderer. The input / output port 2009 may be coupled to headphones (with or without head tracking), etc.
[0239] In general, various embodiments of the present invention may be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some embodiments may be implemented in hardware, while others may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Various embodiments of the present invention may be illustrated and described as block diagrams, flowcharts, or using any other illustrative representation, but it will be understood that these blocks, devices, systems, techniques, or methods described herein may be implemented, in non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or any combination thereof.
[0240] Embodiments of the present invention may be implemented by computer software executable by a data processor in a mobile device, such as a processor entity, by hardware, or by a combination of software and hardware. Furthermore, it should be noted that any block of the logic flow shown in the figure may represent a program step, or an interconnected set of logic circuits, blocks, and functions, or a combination of program steps and logic circuits, blocks, and functions. The software may be stored on a physical medium such as a memory chip or a memory block implemented in the processor, a magnetic medium such as a hard disk or floppy disk, or an optical medium such as a DVD and its data variants or a CD.
[0241] Memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technical environment and may include, in non-limiting examples, one or more of the following: general-purpose computers, dedicated computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multi-core processor architectures.
[0242] Embodiments of the present invention may be put into practice in various components, such as integrated circuit modules. Designing integrated circuits is generally a highly automated process. Complex and powerful software tools are available to translate logic-level designs into semiconductor circuit designs ready for etching and forming on semiconductor substrates.
[0243] Programs such as those offered by Synopsys, Inc. in Mountain View, California, and Cadence Design, Inc. in San Jose, California, automatically route conductors and place components on a semiconductor chip using not only established design rules but also a library of pre-stored design modules. Once the design for the semiconductor circuit is complete, the resulting design may be sent to a semiconductor manufacturing facility or "fab" for production in a standardized electronic format (e.g., Opus or GDSII).
[0244] The foregoing description has provided a sufficient and informative explanation of exemplary embodiments of the invention in an illustrative and non-limiting manner. However, various modifications and applications will become apparent to those skilled in the art in light of the foregoing description when read in conjunction with the accompanying drawings and the attached claims. Nevertheless, all such modifications and similar modifications of the teachings of the invention will still fall within the scope of the invention as defined in the attached claims.< / roomimpulseresponse> < / roomimpulseresponse> < / acousticparameters> < / acousticenvironment> < / acousticparameters> < / acousticenvironment> < / roomimpulseresponse> < / aranchor> < / acousticenvironment> < / acousticmaterial> < / mesh> < / audioscene>
Claims
1. A method for reproducing at least one reflection of reverberation in a virtual acoustic rendering system, The method involves acquiring the shape of the room and at least one spatial intra-intra Determining reverberation time parameters based on at least one spatial-in-room impulse response, The room absorption area is determined based on the determined reverberation time parameters and the shape of the room, Determining at least one parameter of early reflection from at least one intra-spatial impulse response based on at least one image source position, wherein at least one image source position is associated with the early reflection of at least one intra-spatial impulse response, Extracting at least one reflection from at least one spatial intra-intra-space impulse response, Extracting at least one reflection coefficient based on at least one extracted reflection for at least one surface in the room, Assigning an absorption coefficient to at least one surface in the room based on at least one reflectance coefficient and the determined indoor absorption area, Render at least one early reflection based on the determined absorption coefficient. Methods that include...
2. The step of determining at least one parameter is, The time of early reflection from at least one spatial intra-intra-space impulse response based on at least one image source position, The direction of arrival of early reflections from at least one spatial intra-intra-space impulse response based on at least one image source position and The method according to claim 1, comprising determining at least one of the following.
3. The method according to claim 1 or 2, wherein extracting at least one reflection from at least one spatial intra-intra
4. The method according to any one of claims 1 to 3, wherein obtaining at least one spatial intra-intra
5. The method according to claim 4, wherein at least one spatial intra-intra
6. Obtaining the shape of the interior of a room is equivalent to obtaining the surface area S. i The method according to any one of claims 1 to 5, comprising obtaining an indexed list of planar surfaces each containing and metadata indicating the location and source of a receiver within the shape of a room.
7. The method according to claim 6, further comprising obtaining metadata indicating a source-directivity pattern for a room.
8. Determining the indoor absorption area based on the determined reverberation time parameters and the shape of the room is possible. Determining the total absorption area from the combination of surface areas, The room absorption area is determined based on the total absorption area and the determined reverberation time parameters. The method according to claim 6 or 7, including the method described in claim 6 or 7.
9. A device for reproducing at least one reflection of reverberation in a virtual acoustic rendering system, The method involves acquiring the shape of the room and at least one spatial intra-intra Determining reverberation time parameters based on at least one spatial-in-room impulse response, The room absorption area is determined based on the determined reverberation time parameters and the shape of the room, Determining at least one parameter of early reflection from at least one intra-spatial impulse response based on at least one image source position, wherein at least one image source position is associated with the early reflection of at least one intra-spatial impulse response, Extracting at least one reflection from at least one spatial intra-intra-space impulse response, Extracting at least one reflection coefficient based on at least one extracted reflection for at least one surface in the room, Assigning an absorption coefficient to at least one surface in the room based on at least one reflectance coefficient and the determined indoor absorption area, Render at least one early reflection based on the determined absorption coefficient. An apparatus comprising means configured to perform a certain action.
10. A means configured to determine at least one parameter, The time of early reflection from at least one spatial intra-intra-space impulse response based on at least one image source position, The direction of arrival of early reflections from at least one spatial intra-intra-space impulse response based on at least one image source position and The apparatus according to claim 9, configured to determine at least one of the following.
11. The apparatus according to claim 9 or 10, wherein means configured to extract at least one reflection from at least one spatial intra-intra
12. The apparatus according to any one of claims 9 to 11, wherein means configured to acquire at least one spatial intra-intra
13. The apparatus according to claim 12, wherein at least one spatial intra-intra
14. A means configured to acquire the shape of an interior space is, i The apparatus according to any one of claims 9 to 13, configured to obtain an indexed list of planar surfaces, each containing such a surface, and metadata indicating the location and source of a receiver within the shape of a room.
15. The apparatus according to claim 14, wherein means configured to acquire the shape of a room are further configured to acquire metadata indicating a source-directed pattern.
16. A means configured to determine the room absorption area based on the determined reverberation time parameters and the shape of the room, Determining the total absorption area from the combination of surface areas, The room absorption area is determined based on the total absorption area and the determined reverberation time parameters. The apparatus according to claim 14 or 15, configured to perform the following:
17. An apparatus for reproducing at least one reflection of reverberation in a virtual acoustic rendering system, the apparatus comprising at least one processor and at least one memory for storing instructions, wherein when an instruction is executed by the at least one processor, the apparatus provides at least one For a room, the shape of the room and at least one spatial intra-room impulse response are acquired, and at least one spatial intra-room impulse response is acquired based on measurements taken in the room. The reverberation time parameters are determined based on at least one spatial-in-room impulse response, Based on the determined reverberation time parameters and the shape of the room, the room absorption area is determined. Based on at least one image source position, at least one parameter of early reflection is determined from at least one intra-spatial impulse response, and at least one image source position is associated with the early reflection of at least one intra-spatial impulse response. Extract at least one reflection from at least one spatial intra-intra-intra-space impulse response, For at least one surface in the room, at least one reflection coefficient is extracted based on at least one extracted reflection, Based on at least one reflectance coefficient and the determined indoor absorption area, an absorption coefficient is assigned to at least one surface in the room. Render at least one early reflection based on the determined absorption coefficient. A device that causes something to happen.
18. A device for reproducing at least one reflection of reverberation in a virtual acoustic rendering system, An acquisition circuit configured to acquire the shape of a room and at least one spatial intra-intra A decision circuit configured to determine reverberation time parameters based on at least one spatial-in-room impulse response, A determination circuit configured to determine the room absorption area based on the determined reverberation time parameters and the shape of the room, A decision circuit configured to determine at least one parameter of early reflection from at least one intra-space impulse response based on at least one image source position, wherein the at least one image source position is associated with the early reflection of at least one intra-space impulse response, An extraction circuit configured to extract at least one reflection from at least one spatial intra-intra-space impulse response, An extraction circuit configured to extract at least one reflection coefficient based on at least one extracted reflection for at least one surface in the room, An assignment circuit configured to assign an absorption coefficient to at least one surface in a room based on at least one reflection coefficient and a determined indoor absorption area, A rendering circuit configured to render at least one early reflection based on the determined absorption coefficient, A device equipped with the following features.
19. A computer program (or computer-readable medium containing instructions) that includes instructions to cause a device to reproduce at least one reverberation reflection in a virtual acoustic rendering system, wherein the device For a room, the shape of the room and at least one spatial intra-room impulse response are acquired, and at least one spatial intra-room impulse response is acquired based on measurements taken in the room. The reverberation time parameters are determined based on at least one spatial-in-room impulse response, Based on the determined reverberation time parameters and the shape of the room, the room absorption area is determined. Based on at least one image source position, at least one parameter of early reflection is determined from at least one intra-spatial impulse response, and at least one image source position is associated with the early reflection of at least one intra-spatial impulse response. Extract at least one reflection from at least one spatial intra-intra-intra-space impulse response, For at least one surface in the room, at least one reflection coefficient is extracted based on at least one extracted reflection, Based on at least one reflectance coefficient and the determined indoor absorption area, an absorption coefficient is assigned to at least one surface in the room. Render at least one early reflection based on the determined absorption coefficient. A computer program (or computer-readable medium) that causes something to happen.
20. A device for reproducing at least one reflection of reverberation in a virtual acoustic rendering system, For a room, the shape of the room and at least one spatial intra-room impulse response are acquired, and at least one spatial intra-room impulse response is acquired based on measurements taken in the room. The reverberation time parameters are determined based on at least one spatial-in-room impulse response, Based on the determined reverberation time parameters and the shape of the room, the room absorption area is determined. Based on at least one image source position, at least one parameter of early reflection is determined from at least one intra-spatial impulse response, and at least one image source position is associated with the early reflection of at least one intra-spatial impulse response. Extract at least one reflection from at least one spatial intra-intra-intra-space impulse response, For at least one surface in the room, at least one reflection coefficient is extracted based on at least one extracted reflection, Based on at least one reflectance coefficient and the determined indoor absorption area, an absorption coefficient is assigned to at least one surface in the room. Render at least one early reflection based on the determined absorption coefficient. A device equipped with means for doing so.
21. For the reproduction of at least one reflection of reverberation in a virtual acoustic rendering system, the device shall have at least the following: The method involves acquiring the shape of the room and at least one spatial intra-intra Determining reverberation time parameters based on at least one spatial-in-room impulse response, The room absorption area is determined based on the determined reverberation time parameters and the shape of the room, Determining at least one parameter of early reflection from at least one intra-spatial impulse response based on at least one image source position, wherein at least one image source position is associated with the early reflection of at least one intra-spatial impulse response, Extracting at least one reflection from at least one spatial intra-intra-space impulse response, Extracting at least one reflection coefficient based on at least one extracted reflection for at least one surface in the room, Assigning an absorption coefficient to at least one surface in the room based on at least one reflectance coefficient and the determined indoor absorption area, Render at least one early reflection based on the determined absorption coefficient. A non-temporary computer-readable medium containing program instructions for performing an action.