Determining Virtual Audio Source Position
The method and apparatus efficiently determine virtual audio source positions by generating mirror rooms and using mapping matrices to reduce complexity, improving the accuracy and realism of audio reflections in virtual reality applications.
Patent Information
- Application Number
- JP2024506191
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-05
- Filing Date
- 2022-07-26
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2042-07-26
AI Technical Summary
Existing models for determining the location of virtual audio sources representing reflections in a room are complex, computationally intensive, and limit the accuracy and quality of the audio experience in virtual reality applications.
A method and apparatus for determining virtual audio source positions by generating mirror rooms through mirroring operations, using mapping matrices to efficiently calculate relative position offsets, reducing computational complexity and resource requirements.
Facilitates accurate and efficient generation of image source models for early reflections, allowing dynamic adaptation to audio source movements with reduced computational load, enhancing the realism of the audio experience.
Smart Images

Figure 0007787283000020 
Figure 0007787283000021 
Figure 0007787283000022
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus and method for determining the position of virtual audio sources representing reflections of audio sources in a room, particularly but not exclusively for rendering audio in augmented / virtual reality applications. [Background technology]
[0002] In recent years, the wide variety of experiences based on audiovisual content has increased significantly as new services and methods for using and consuming such content are continually developed and introduced. In particular, many spatial and interactive services, applications, and experiences are being developed to provide users with more engaging and immersive experiences.
[0003] Examples of such applications are virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications, which are rapidly becoming mainstream with many solutions targeting the consumer market. Many standards are also under development by a number of standards bodies. These standardization efforts are actively developing standards for various aspects of VR / AR / MR systems, including, for example, streaming, broadcasting, rendering, etc.
[0004] VR applications tend to provide a user experience corresponding to the user being in a different world / environment / scene, whereas AR (including mixed reality MR) applications tend to provide a user experience corresponding to the user being in a current environment but with additional information or virtual objects or information added. Thus, VR applications tend to provide fully immersive synthetically generated worlds / scenes, whereas AR applications tend to provide partially synthetic worlds / scenes that are overlaid on a real scene in which the user is physically present. However, these terms are often used interchangeably and have a high degree of overlap. In the following, the term virtual reality / VR will be used to refer to both virtual reality and augmented / mixed reality.
[0005] As an example, an increasingly popular service is the presentation of images and audio in a manner that allows the user to actively and dynamically interact with the system, changing the parameters of the rendering to adapt to movement and changes in the user's position and orientation. A very attractive feature in many applications is the ability to change the viewer's effective viewing position and direction, for example, allowing the viewer to move and "look around" within the scene being presented.
[0006] Such features in particular enable a virtual reality experience to be provided to the user, which may allow the user to move around (relatively) freely in the virtual environment and dynamically change their position and their point of view. Typically, such virtual reality applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well known from gaming applications, such as in the class of first-person shooter games for computers and consoles.
[0007] It is also desirable, particularly for virtual reality applications, that the presented images be three-dimensional images. Indeed, to optimize the viewer's immersion, it is typically preferred that the user experience the presented scene as a three-dimensional scene. Indeed, it is preferred that the virtual reality experience allow the user to select their position, camera viewpoint, and moment in time relative to the virtual world.
[0008] In addition to the visual rendering, most VR / AR applications also provide a corresponding audio experience. In many applications, the audio preferably provides a spatial audio experience, such that audio sources are perceived as coming from positions corresponding to the positions of corresponding objects in the visual scene. In this way, the audio scene and the video scene are preferably perceived as coherent, and together provide a complete spatial experience.
[0009] For example, many immersive experiences are provided by generating virtual audio scenes through headphone playback using binaural audio rendering techniques. In many scenarios, such headphone playback can be based on head tracking, allowing the rendering to occur in response to the user's head movements, which greatly increases the sense of immersion.
[0010] However, in order to provide the user with a highly immersive, personalized and natural experience, it is important that the rendering of the audio scene is as realistic as possible, and for combined audio-visual experiences such as many VR experiences, it is important that the audio experience closely matches that of the visual experience, i.e., that the rendered audio and video scenes closely match.
[0011] To provide a high-quality experience, and especially for the audio to be perceived as realistic, it is important that the acoustic environment be characterized by an accurate and realistic model. This is required whether the audio scene being presented is a purely virtual scene, or whether it is desired that the scene correspond to a particular real-world scene.
[0012] When simulating room acoustics, or more generally environmental acoustics, reflections of sound waves off the walls, floor, and ceiling (if present) of the environment cause delayed and attenuated (usually frequency-dependent) versions of the audio source signal to arrive at the listener from different directions, resulting in impulse responses called room impulse responses (RIRs).
[0013] As shown in Figure 1, a room impulse response consists of a direct / anechoic part that depends on the distance from the audio source to the listener, followed by a reverberant part that characterizes the acoustics of the room. The size and shape of the room, the position of the audio source and listener within the room, and the reflective properties of the room's surfaces all affect the characteristics of this reverberant part.
[0014] The reverberant part can be decomposed into two usually overlapping temporal regions. The first region contains the so-called early reflections, which are isolated reflections of the audio source on the walls or obstacles in the room before reaching the listener. As the time delay increases, the number of reflections present within a given period increases, including second-order and higher-order reflections.
[0015] The second region in the reverberant section is where the density of these reflections increases to the point where they cannot be separated and differentiated by the human brain. This region is called diffuse reverberation, late reverberation, or reverberation tail.
[0016] The reverberant part contains clues that give the auditory system information about the distance of the audio source, the size and acoustic properties of the room. The energy in the reverberant part relative to the energy in the anechoic part largely determines the perceived distance of the audio source. The level and delay of the earliest reflections can give clues as to how close the audio source is to a wall, and anthropometric filtering of the reflections can enhance the estimation of what wall, floor, or ceiling it is.
[0017] The density of (early) reflections affects the perceived size of a room. It takes the reflections to reduce the energy level by 60 dB, T 60 The time given by (the reverberation time) is a measure of how quickly reflections dissipate in a room. Reverberation time provides information about the acoustic properties of a room; whether the walls of the room are highly reflective (e.g., a bathroom) or highly sound absorbing (e.g., a bedroom with furniture, carpet, and curtains).
[0018] For reverberation to provide an immersive experience, multiple RIRs are required to represent the directions from which reflections reach the listener. These can be associated with a loudspeaker system, where each RIR is associated with one of several known positions on the loudspeakers. A panning algorithm, such as VBAP, can be employed to generate RIRs from known reflection directions of multiple reflections.
[0019] Furthermore, the immersive RIR may depend on the anthropometric characteristics of the user, i.e., the Head-Based Impulse Response (HRIR), if the RIR is part of the Binaural Room Impulse Response (BRIR), as the RIR is filtered by the head, ears, and shoulders.
[0020] Since the reflections in the late reverberation are no longer separable, they can be simulated parametrically using parametric reverberators, e.g., feedback delay networks like the Giotto reverberator. For early reflections, the delay, which depends on the incidence direction and distance, is an important clue for humans to extract information about the relative position of the room and the audio source. Therefore, for a realistic immersive experience, the simulation of early reflections must be even clearer than the late reverberation.
[0021] One approach to modeling early reflections is to mirror audio sources at each boundary of the room to generate virtual audio sources representing the reflections. Such models, known as image source models, are described in Allen JB, Berkley DA, “Image method for efficiently simulating small-room acoustics,” The Journal of the Acoustical Society of America, 1979; 65(4):943-50. European Patent Application Publication No. 3828882 discloses an example implementation of such a method. However, while such models can provide efficient, high-quality modeling of early reflections compared to approaches that are less restrictive on room geometry, such as ray tracing or finite element modeling, they tend to have several drawbacks. Specifically, such modeling approaches tend to be relatively complex and have high computational resource requirements, especially for locating the virtual source of reflections. The processing required to determine mirror positions, and the geometric calculations required for mirroring, especially around room boundaries, tend to be complex and resource-intensive. These drawbacks tend to increase the larger the number of reflections considered, and in many practical applications the number of reflections is limited accordingly, reducing the accuracy and quality of the model and, consequently, the user experience. Summary of the Invention [Problem to be solved by the invention]
[0022] Therefore, improved models and approaches for determining the location of virtual audio sources representing reflections would be advantageous. In particular, approaches / models that enable improved operation, increased flexibility, reduced complexity, facilitated configuration, improved audio experience, reduced complexity, reduced computational load, improved audio quality, improved model accuracy and quality, and / or improved performance and / or operation would be advantageous.
[0023] SUMMARY OF THE INVENTION Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]
[0024] According to one aspect of the present invention, there is provided a method for determining a virtual audio source position of an audio source representing a reflection of a first audio source in a first room, the method comprising: receiving data describing a boundary of the first room; generating a group of mirror rooms for the first room, each mirror room resulting from a plurality of mirrorings, each mirroring being a mirroring of a previous mirror room to (around) a boundary of the previous mirror room, the first previous mirror room of the plurality of mirrorings being the first room; mapping, for at least a first mirror room of the group of mirror rooms, directions in the first room to directions in the first mirror room; determining a source reference position in the first room; determining a first mirror reference position in the first mirror room, the first mirror reference position being , a position within the first mirror room resulting from applying mirroring to a first source reference position to produce a first mirror room; determining a source position offset in the source room relative to the first audio source, the source position offset representing a position offset between the source reference position and a position within the first room of the first audio source; determining a first mirror position offset in the first mirror room relative to the first audio source; determining a mirror position of the first audio source within the first mirror room from the first mirror reference position and the first mirror position offset; wherein the step of determining the first mirror position offset includes determining the first mirror position offset by applying a mapping to the source position offset.
[0025] The present invention may result in improved and / or facilitated determination of mirror / virtual positions for virtual audio sources representing reflections in a room. The approach may allow for facilitated and / or more efficient generation of models for early reflections in a room. The approach, in many embodiments, can significantly facilitate determination of mirror positions and significantly reduce the number and / or complexity of calculations required to determine mirror positions corresponding to reflections. The approach, in many embodiments, allows accurate models representing reflections in a room to be generated with reduced computational requirements and / or complexity.
[0026] This approach can be used to generate / update image source models.
[0027] The approach may be particularly suited to dynamic applications, for example, where one or more audio sources within a room may move, resulting in changes in the reflections of the audio sources. Corresponding changes in mirror position can typically be determined with low complexity, thereby facilitating dynamic adaptation to audio source movement, including accurate representation of resulting changes in the characteristics of early reflections. In particular, the approach may, in many embodiments, enable the relatively complex computations required to determine and evaluate mapping / mirroring to different mirror rooms to represent two or more reflections to be performed only once at initialization, allowing later effects resulting from movement to be determined with significantly lower complexity.
[0028] The approach results in a low complexity and low resource demanding process for representing the reflections in the first room, and may in particular allow image source models of a given quality / complexity to be generated / updated with significantly lower computational load.
[0029] The method may be a computer-implemented method, which may be executed by a computer / processor, and in many embodiments may include generating an audio signal that includes a component from an audio source, such as one located at the mirror position.
[0030] A data processing apparatus / device / system may be provided comprising means for carrying out the above method.
[0031] A mirror chamber may be a chamber resulting from one or more mirror operations on (around) the edge / edge / boundary of an original chamber and / or one or more previously generated mirror chambers. A mirror position may be a position within a mirror chamber. A mirror chamber / mirror position may be a virtual position.
[0032] The first chamber (and mirror chamber) can be represented as a two-dimensional rectangle and / or a three-dimensional rectangle. The first chamber can be a two-dimensional or three-dimensional orthotope, also known as a right-angled rectangular prism, a rectangular cube, or a rectangular parallelepiped, and sometimes referred to in the art as a shoebox chamber.
[0033] A room boundary may be a planar element that divides / defines / delimits the room, such as a wall, floor, ceiling, etc. A boundary may be an acoustically reflective element, possibly a virtual or theoretical (arbitrary) depiction of a boundary, e.g., where no (significant) acoustically reflective element is present.
[0034] A room can be any acoustic environment that is substantially planar and typically segmented by acoustically reflective elements. The planar elements can be parallel in pairs; a two-dimensional room can have two such parallel pairs, and a three-dimensional room can have three such parallel pairs (corresponding to the four walls, floor, and ceiling).
[0035] Mirroring an audio source across a boundary may correspond to determining a mirrored audio source position by mirroring the audio source position of the audio source relative to the boundary. The previous mirrored room may be a room belonging to a set that includes the first room and a mirrored room generated by a previous mirroring in the plurality of mirrorings.
[0036] The mirrored audio source position of the reflected audio source in the nearby mirrored room may correspond to the position that results from mirroring the audio source position of the audio source in the source / first room relative to the boundary.
[0037] In accordance with an optional feature of the invention, the mapping is represented by a mapping matrix, and applying the mapping to the source position offsets comprises multiplying the mapping matrix by the source position offsets.
[0038] This configuration can enable improved performance and / or operation in many embodiments. This configuration can enable more efficient operation in many embodiments and can significantly reduce computational resource burden / requirements. The source position offset, including being represented as a coordinate offset, can be represented by multiple coordinates, and a matrix multiplication can map these coordinate offsets to a coordinate offset corresponding to the first mirror position offset.
[0039] In accordance with an optional feature of the invention, the source position offset is a two-dimensional offset and the mapping matrix is a 2x2 matrix.
[0040] This configuration may provide improved and / or facilitated operation in many embodiments, and the approach results in less complex and less resource-demanding processing for representing reflections in the first room.
[0041] In accordance with an optional feature of the invention, the source position offset is a three-dimensional offset and the mapping matrix is a 3x3 matrix.
[0042] This configuration may provide improved and / or facilitated operation in many embodiments, as the approach results in less complex and less resource-demanding processing for representing reflections in the first chamber.
[0043] In accordance with an optional feature of the invention, the number of mirrorings for the first room is at least two, and the mapping matrix is a combination of multiple boundary mirror mapping matrices, each boundary mirror mapping matrix representing a change in orientation resulting from mirroring for a single room boundary.
[0044] This arrangement may provide improved and / or facilitated operation. In many embodiments, the arrangement may provide efficient determination and representation of reflections originating from room boundaries. In particular, the arrangement may facilitate and / or improve the representation of multiple reflections of the same audio source within a first room.
[0045] According to an optional feature of the invention, each room boundary of the first room is linked to one boundary mirror mapping matrix, and the mapping is a combination of boundary mirror mapping matrices linked to room boundaries of mirrorings among the plurality of mirrorings for the first mirror.
[0046] This configuration may facilitate and / or improve operation for representing multiple reflections in the first chamber.
[0047] In accordance with an optional feature of the invention, parallel chamber boundaries of the first chamber are linked to the same boundary mirror mapping matrix.
[0048] This arrangement may provide improved and / or facilitated operation. In many embodiments, the arrangement may provide efficient determination and representation of reflections arising from chamber boundaries. In particular, the arrangement may reduce computational resource requirements for determining the mapping and / or first mirror position offset.
[0049] In accordance with an optional feature of the invention, the mapping is a distance-preserving mapping.
[0050] This configuration can provide improved performance and / or operation, and in many embodiments can achieve reduced computational resource requirements for a given accuracy of the acoustically reflective environment of the first chamber.
[0051] In accordance with an optional feature of the invention, a distance of the first mirror position offset is equal to a distance of the source position offset.
[0052] This configuration can provide improved performance and / or operation in many embodiments. The configuration can achieve reduced computational resource requirements for a given accuracy of the acoustically reflective environment of the first room. The distance of the offset can be the size / magnitude of the offset (specifically, of a vector representing the offset).
[0053] According to an optional feature of the invention, the method further comprises: determining a chamber boundary position relative to a boundary of the first chamber; determining a boundary position offset in the first chamber relative to the chamber boundary position, the boundary position offset representing a position offset between a source reference position and the chamber boundary position; determining a boundary position offset in the first mirror chamber by applying a mapping to the boundary position offset; and determining a mirror boundary position relative to the first mirror chamber from the first mirror reference position and the boundary position offset.
[0054] This configuration can provide improved performance and / or operation in many embodiments. It can also achieve reduced computational resource requirements for a given accuracy of the acoustically reflective environment of the first chamber, and can facilitate, among other things, the determination of geometric characteristics of the mirror chamber. The chamber boundary locations can be locations that are included within one or more boundaries, particularly corner locations of the chamber.
[0055] In accordance with an optional feature of the invention, the first chamber is an orthotope, and the source reference position and source position offset are expressed in terms of coordinates of coordinate axes not aligned with sides of the orthotope.
[0056] This configuration can provide improved performance and / or operation in many embodiments.
[0057] In accordance with an optional feature of the invention, the method further comprises determining a room response function for the first room, the room response function including a reflected component representative of audio from an audio source as located at the mirror position.
[0058] The method may, in many embodiments, allow for an improved representation of the reflectance properties of a room and / or reduced complexity.
[0059] In accordance with an optional feature of the invention, the method includes rendering the audio output signal including a component from the audio source as being located at the mirror position.
[0060] The method can provide improved and / or facilitated rendering of audio, which typically provides an improved user experience with a more realistic perception of the environment.
[0061] According to one aspect of the present invention, there is provided an apparatus for determining a virtual audio source position of an audio source representing a reflection of a first audio source in a first room, the apparatus comprising: receiving data describing a boundary of the first room; generating a group of mirror rooms for the first room, where each mirror room results from a plurality of mirrorings, each mirroring being a mirroring of a previous mirror room relative to a boundary of the previous mirror room, the first previous mirror room of the plurality of mirrorings being the first room; mapping, for at least a first mirror room of the group of mirror rooms, a direction in the first room to a direction in the first mirror room; determining a source reference position in the first room; and determining a first mirror reference position in the first mirror room, where the first mirror reference position is used to generate the first mirror room. a position in the first mirror room resulting from applying the mapping to the first source reference position; determining a source position offset in the source room relative to the first audio source, wherein the source position offset represents a position offset between the source reference position and a position in the first room of the first audio source; determining a first mirror position offset in the first mirror room relative to the first audio source; and determining a mirror position of the first audio source in the first mirror room from the first mirror reference position and the first mirror position offset, wherein the operation of determining the first mirror position offset includes an operation of determining the first mirror position offset by applying the mapping to the source position offset.
[0062] This apparatus may, in many embodiments, result in improved performance and / or reduced complexity / resource usage, and typically allows improved models to be generated with less complexity and use of computational resources.
[0063] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.
[0064] Embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which: [Brief explanation of the drawings]
[0065] [Figure 1] FIG. 1 shows an example of the components of an acoustic room response. [Figure 2] FIG. 2 shows an example of elements of an apparatus according to some embodiments of the present invention. [Figure 3] FIG. 3 shows an example of mirroring to model acoustic reflections at room boundaries. [Figure 4] Figure 4 shows an example of mirroring to model the acoustic reflection of two boundaries of a room. [Figure 5] FIG. 5 shows an example of room mirroring for an image source model of acoustic reflections in a room. [Figure 6] FIG. 6 illustrates an example of elements of a method according to some embodiments of the present invention. [Figure 7] FIG. 7 illustrates an example of elements of a method according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0066] Audio rendering, which aims to provide a natural and realistic effect to the listener, usually involves rendering the acoustic environment. The rendering is based on a model of the acoustic environment, which typically includes modeling the direct path, (early) reflections, and reverberation. The following description focuses on an efficient approach to generate an appropriate model of the (early) reflections in a real or virtual room.
[0067] The approach will be described with reference to an audio rendering apparatus having the elements shown in FIG. 2. The audio rendering apparatus 201 of FIG. 2 includes a receiver 201 configured to receive room data characterizing a first room representing the acoustic environment to be emulated by the rendering. The room data, in particular, describes the boundaries of the first room. The receiver 201 may also receive data regarding the audio source position of at least one audio source within the room. Typically, data indicating the audio source positions of multiple audio sources may be received. Furthermore, in many embodiments, these positions may change dynamically, and the system may be configured to adapt to such position changes. The receiver may also receive audio data of an audio source, the audio data representing audio generated by the audio source, and may render the audio from the audio source. The audio rendering apparatus is configured to render the audio with characteristics (in particular, early reflections, and typically direct path and reverberant components) such that the audio is perceived as realistic audio of the room.
[0068] In the following, to distinguish between the virtual (mirrored) room that is generated and the virtual (mirrored) audio source that is generated due to the described reflection model, the room will also be referred to as the original room (or first room or source room), and the audio source in the original room will also be referred to as the original audio source.
[0069] The receiver 201 may be implemented in any suitable manner, including, for example, using discrete or dedicated electronic circuitry. The processing circuitry 203 may be implemented as an integrated circuit, such as, for example, an application specific integrated circuit (ASIC). In some embodiments, the circuitry may be implemented as a programmed processing unit, such as firmware or software running on a suitable processor, such as, for example, a central processing unit, a digital signal processing unit, or a microcontroller. It will be understood that in such embodiments, the processing unit may include on-board or external memory, clock driving circuits, interface circuits, user interface circuits, etc. Such circuitry may further be implemented as part of the processing unit, as an integrated circuit, and / or as discrete electronic circuitry.
[0070] The receiver 201 can receive data, particularly room data and / or audio data, from any suitable source in any suitable format (e.g., as part of an audio signal). Room data can be received from an internal or external source. The receiver 201 can be configured, for example, to receive room / audio data via a network connection, a wireless connection, or other suitable connection to an internal source. In many embodiments, the receiver can receive the data from a local source, such as local memory. In many embodiments, the receiver 201 can be configured, for example, to retrieve the room data from local memory, such as local RAM or ROM memory.
[0071] The boundaries define the contours of a room and typically represent walls, ceiling, and floor (or, in the case of 2D applications, usually just walls). The room is a 2D or 3D rectangular volume, such as a 2D rectangle or a 3D rectangle (shoebox shape). The boundaries are pairwise parallel and substantially planar. Furthermore, the boundary in one pair of parallel boundaries is perpendicular to the parallel boundary in another pair. These boundaries define a particular rectangular volume (2D or 3D). These boundaries may reflect any physical properties, such as any material. They may also represent any acoustic properties.
[0072] The room described by the room data corresponds to the intended acoustic environment for rendering and may therefore represent a real room / environment or a virtual room / environment. A room may be any area / region / environment that can be bounded / delimited by four (in the 2D case) or six (in the 3D case) substantially planar boundaries that are pairwise parallel and substantially perpendicular between pairs. The room data may, in some embodiments, represent a suitable approximation of the intended room that is not pairwise parallel and / or does not exhibit right angles between connected boundaries.
[0073] In most embodiments, the room data may further include acoustic data for one, more, or typically all, of the boundaries. The acoustic property data may specifically include a reflection attenuation measure for each wall, which indicates the attenuation caused by the boundary when sound is reflected from the boundary. As another example, a reflection coefficient may indicate the fraction of signal energy that is specularly reflected from the boundary. In many embodiments, the attenuation measure may be frequency-dependent to model that reflections may differ for different frequencies. Furthermore, the acoustic property may depend on the position on the boundary.
[0074] The receiver 201 is coupled to a processing circuit 203, which is configured to generate a reflection model for the room / acoustic environment that represents (early) reflections in the room and allows these to be emulated during rendering. In particular, the processing circuit 203 is configured to determine virtual audio sources that represent reflections of the original audio source in the original room.
[0075] The processing circuitry 203 may be implemented in any suitable form, including, for example, using discrete or dedicated electronic circuitry. The processing circuitry 203 may be implemented as an integrated circuit, such as, for example, an application specific integrated circuit (ASIC). In some embodiments, the circuitry may be implemented as a programmed processing unit, such as firmware or software running on a suitable processor, such as, for example, a central processing unit, a digital signal processing unit, or a microcontroller. In such embodiments, the processing unit may include on-board or external memory, clock driver circuits, interface circuits, user interface circuits, etc. Such circuitry may further be implemented as part of the processing unit, as an integrated circuit, and / or as discrete electronic circuitry.
[0076] The processing circuitry 203 is coupled to a rendering circuitry 205, which is configured to render audio signals representing the audio source and typically also multiple other audio sources to provide a rendering of an audio scene. The rendering circuitry 205 can specifically receive audio data characterizing audio from an original audio source and render it according to any suitable rendering methodology and technique. Rendering the original audio source may include generating reflected audio based on a reflection model generated by the processing circuitry 203. Furthermore, signal components corresponding to the direct path and reverberation of the original audio source will typically also be rendered. Those skilled in the art will be aware of many different approaches for rendering audio, including for spatial speaker configurations and headphones using, for example, binaural processing, which will not be described in further detail for the sake of brevity.
[0077] In this way, the rendering circuitry 205 can generate an audio output signal that includes audio from (at least one of) the room's audio sources. The audio from a given audio source is processed according to the room's characteristics. In particular, the audio output signal is generated to include at least one component representing a reflection of the given audio source, which component is determined as the component that would result from the audio source if it were located at a mirror position (of the reflection model). For example, the component can be generated to be perceived as arriving from a direct path from the mirror position, possibly with a (possibly frequency-selective) attenuation corresponding to the wall / boundary reflection component of the reflection being modeled.
[0078] In many embodiments, a room response function is generated for the original room, which includes a reflected component representing audio from an audio source located at the mirror. For example, the room response may include a contribution representing the acoustic transfer function (e.g., time-of-flight delay, distance attenuation, and head-related impulse response) of the direct path from the mirror to the listening position, modified to include the reflective properties of the walls involved in the reflection being modeled. Typically, the room response is generated from many such contributions, each representing one early reflection.
[0079] The rendering circuit 205 may be configured to filter each audio source with a room response determined for the room / position and including such contributions representing early reflections.
[0080] The rendering circuitry 205 may be implemented in any suitable form, including, for example, using separate or dedicated electronic circuitry. The rendering circuitry 205 may be implemented as an integrated circuit, such as, for example, an application specific integrated circuit (ASIC). In some embodiments, the circuitry may be implemented as a programmed processing unit, such as firmware or software running on a suitable processor, such as, for example, a central processing unit, a digital signal processing unit, or a microcontroller. It will be appreciated that in such embodiments, the processing unit may include on-board or external memory, clock driving circuits, interface circuits, user interface circuits, etc. Such circuitry may further be implemented as part of the processing unit, as an integrated circuit, and / or as separate electronic circuitry.
[0081] The processing circuitry 203 is particularly configured to generate a mirror source model for reflections. In the mirror source model, reflections are modeled by individual virtual audio sources, where each virtual audio source is a replica of the original audio source, located at a (virtual) location outside the original room, but in such a way that a direct path from the virtual location to the listening position exhibits the same characteristics as a reflected path from the original audio source to the listening position. Specifically, the path length of the virtual audio source representing the reflection will be equal to the path length of the reflected path from the original audio source to the listening position. Furthermore, the direction of arrival at the listening position for the virtual sound source path will be the same as the direction of arrival for the reflected path. Furthermore, for each reflection by a boundary (e.g., a wall) for the reflected path, the direct path will pass through a boundary corresponding to the reflecting boundary. Therefore, transmission through the model boundary can be used to directly model the reflection effect; for example, an attenuation corresponding to the boundary's reflection attenuation can be assigned to the transmission through the corresponding model boundary.
[0082] A particularly important property of the mirror source model is its independence from the listening position. The determined position and room configuration are such that it provides correct results for all positions in the original room. Specifically, a virtual mirror audio source and a virtual mirror room are generated that can be used to model the reflection behavior for any position in the original room. That is, they can be used to determine the path length, reflections, and direction of arrival for any position in the original room. In this way, the generation of the mirror source model can be performed during the initialization process, and the generated model can be used and evaluated continuously and dynamically, for example, as a user is expected to move around (translate and / or rotate) in the original room. In this way, the generation of the mirror source model is performed without any consideration of the actual listening position; rather, a more general model is generated.
[0083] However, if the audio source moves, the model must be updated to reflect how the movement affects the reflection. That is, as the position of the audio source changes in the original room, the position of the corresponding mirror of the audio source, which represents the reflection, also changes. This requires that the model be updated to reflect the new mirror position. In many embodiments, the model may need to be updated relatively frequently to reflect new positions and mirror positions. However, the process of determining mirror positions tends to be complex and resource-intensive, and therefore, such updates also tend to be complex and resource-intensive. In many cases, the perceived accuracy of the rendered audio and the resulting user experience may be limited by available computational resources. An improved trade-off between the perceived quality of the rendered audio (and, in particular, how realistic it appears) and computational resource requirements is desired. Therefore, an efficient, high-performance approach for determining / updating the mirror positions of audio sources in the original room is highly desirable.
[0084] The processing circuitry 203 can generate a mirror audio source model that can emulate reflections in the original room with a direct path from the virtual mirror audio source.
[0085] As shown in Figure 3, the reflected sound components can be rendered as a direct path of the mirrored audio source, which represents the correct distance and incident direction to the listener. This applies to all positions in the original room, and there is no need to determine new mirror audio source positions for different listening positions. Rather, the virtual mirror source is valid for all user positions in the original room. However, if the audio source position changes, new mirror positions will need to be determined to accurately represent the reflections.
[0086] When generating this virtual mirror source, reflection effects can be taken into account, as previously mentioned, which is typically achieved by assigning to each transition between rooms an attenuation or frequency-dependent filtering that represents the portion of the audio source's energy that is specularly reflected by the surface of the boundary that is traversed.
[0087] Because sound may reach the user via multiple boundary reflections, the approach can be iterative, as shown in FIG. 4. Processing circuitry 203 can, for example, determine multiple "layers" of mirror chambers and mirror sources to be generated, thereby allowing multiple reflections to be modeled. Each "layer" increases the number of reflections along the path; that is, the first iteration represents sound components that reach the listening position via one reflection, the second iteration represents sound components that reach the listening position via two reflections, and so on. Each "layer" can correspond to a reflection order; thus, for example, the first layer represents a single boundary reflection and is represented by a mirror chamber resulting from mirroring the original chamber, while the second layer represents reflections across two boundaries and is represented by a mirror chamber resulting from mirroring the first "layer" against its own boundary. Third, fourth, etc. layers of mirror chambers can be used to represent higher-order reflections, as appropriate.
[0088] This approach typically results in a diamond-shaped representation of the original and mirrored chambers when successively mirroring the chambers until a certain order / layer (a certain number of mirrorings) is reached that represents a given order of reflection. This is shown in 2D in Figure 5 for up to second order, i.e., with two mirror layers. In 3D, a similar structure would exist if viewed in cross section through the original chamber (i.e., the same pattern would be seen in a vertical plane through a row of five chambers).
[0089] However, while the principles of the above-described approach may appear relatively straightforward, the practical implementation is not, and in fact practical considerations are important to the performance of the approach.
[0090] For example, in many applications, the coordinate system used to represent the room and audio sources may not be aligned with the boundary direction. This makes mirroring less straightforward to calculate, since it affects more than one dimension at a time. In such cases, either the room boundaries and audio sources need to be rotated to align with the coordinate system, and all subsequently determined virtual mirror sources must be rotated in the opposite direction, or the mirroring itself needs to be performed in more than one dimension (e.g., using the boundary normal vector). In many situations, the latter approach will be more efficient.
[0091] Furthermore, the complexity and computational resource requirements for determining mirror positions tend to be high, and thus the complexity problem is exacerbated in applications where the location of the audio source may change. Such changes are likely in many current practical applications, such as AR, VR, games, or other immersive experiences. These changes may result, for example, from animated elements, such as characters or other users walking around and talking, or other active elements (animals, robots, vehicles, etc.). Also, new sources may be introduced into the room or become active at a later point in time.
[0092] For applications where the location of an audio source may change or new sources may be introduced, the mirror locations must be updated, however the process of performing all the iterative mirroring tends to require very large computational resources.
[0093] For example, generating mirrored audio sources for all audio sources in a room with sufficient time resolution to realistically track source movement and with a sufficient order of reflection (typically, 4th to 5th order is required for good quality) requires significant computational power. In fact, 5th order reflections require computing 230 unique mirrored sources per original source in the room. For smooth motion, a 90 Hz animation update rate is typically recommended. Therefore, for each moving source in a simulated room, 230 mirrored sources must be recomputed 90 times per second. In addition to mirroring the sound sources, the rooms also need to be mirrored; specifically, the room definition is mirrored to generate the mirrored rooms. This adds four additional positions to be processed with each room update.
[0094] The number of mirroring operations per second is:
number
[0095] A typical mirroring operation, such as that used in the method of EP 3828882, requires 11 operations per mirrored room plus an additional 25 operations (including square roots) to initialize the mirroring for a particular surface. Therefore, the minimum number of operations per second is:
number
[0096] This ranges from 1.7 to 2.6 MOPS when updating 1 to 5 sources. Additionally, the management and logic for finding unique rooms must be performed. Rerunning the image source method at this rate is very computationally intensive.
[0097] The processing circuitry 203 of Figure 2 is configured to use a highly efficient approach that can significantly reduce the use of computational resources, and in particular, can allow for faster and / or easier determination and updating of mirror positions representing the reflection of the room.
[0098] The approach is not based on directly mirroring individual positions until the desired mirror room is reached, but instead on determining relative position offsets and directly mapping these to the associated mirror room(s). The approach is based on determining corresponding reference positions in the original room and the individual mirror rooms, and then performing a typical (position-independent, distance-preserving) mapping between the relative position offsets in the original room and the relative offsets in the mirror room. The position in the mirror room is then determined from the resulting relative offsets and the mirror reference positions. In this way, the approach is based on mapping relative offsets that change with audio source position, and reference positions that are not (necessarily) dependent on audio source position.
[0099] The mapping may be, in particular, position and / or offset independent. That is, the mapping applied to the source position offset may be independent of the source position offset itself, the reference position, and the audio source position in the original room. Furthermore, the mapping may be distance independent / conserving; in particular, the mapping may be such that the length / magnitude of the mirror position offset is the same as the source position offset. In many embodiments, the mapping may depend only on the geometry of the original room. In many embodiments, the mapping may be independent of the characteristics of the audio source.
[0100] An exemplary approach is described in more detail with reference to the flowchart of FIG.
[0101] In step 601, processing circuitry 203 receives data describing the boundaries of the original room from receiver 201, and possibly also audio data and location data regarding the audio source. In some embodiments, different types of data may be received from different sources.
[0102] Step 601 is followed by step 603 in which processing circuitry 203 generates a group of mirror rooms, each resulting from mirroring an original or previous mirror room. The mirroring to generate the mirror rooms is by mirroring relative to the boundaries of the previous mirror room (which may in particular be the original room). The process will start with the original room being the previous mirror room.
[0103] For example, a first mirror chamber may be generated by mirroring a corner / boundary / wall of the original chamber relative to a first boundary / wall of the original chamber, then a second mirror chamber may be generated by mirroring a corner / boundary / wall of the original chamber relative to another boundary / wall of the original chamber, then a third mirror chamber may be generated by mirroring a corner / boundary / wall of the original chamber relative to a third boundary / wall of the original chamber, etc. Processing circuitry 203 may then proceed to generate a second layer of mirror chambers by mirroring the previously generated mirror chambers.
[0104] For example, a mirror room can be generated by mirroring a corner / boundary / wall of a first previously generated mirror room against the first corner / boundary / wall of the first previously generated mirror room, another mirror room can be generated by mirroring the corner / boundary / wall of the mirror room thus generated, etc. This process can be repeated for all other mirror rooms generated in the previous iteration. This process can then be repeated based on the newly generated mirror room, etc. A given mirror room can typically result from different sequences of mirroring, and the processing circuit 203 can be configured to generate only one copy of each, for example by discarding sequences / mirroring that result in a mirror room identical to one already generated. EP 3828882 discloses a particularly efficient method for generating mirror rooms and selecting between different sequences.
[0105] Thus, step 603 describes a set of mirror rooms in the original room, where each mirror room corresponds to a reflection across a set of boundaries of the original room (the boundaries corresponding to the boundaries along which the mirroring was performed). For a given audio source in the original room, each mirror room has a single reflected audio source such that the direct path from the mirror room location to the listener location matches the reflected path in the original room.
[0106] In the above example, the coordinates / characteristics (such as corner coordinates) of the mirror room may be determined and stored for processing. In some embodiments and applications, each mirror room may be represented only by a reference position and mapping, e.g., as further described below. In many embodiments, transfer function characteristics may be stored, in particular, transfer functions related to one or more reflection coefficients / reflection characteristics of boundaries / walls included in the mirror sequence (and thus corresponding to the reflections represented by the mirror room) may be stored and used for audio.
[0107] In some embodiments, mirror chambers are not generated during an initialization procedure, but rather, a procedure for generating new mirror chambers may be performed when and if needed / desired. Specifically, properties for a given mirror chamber representing a given reflection may only be generated if that reflection has been modeled.
[0108] Thus, each mirror chamber may be generated by / correspond to one or more mirrorings of an original chamber, each mirroring being relative to the boundaries of the original chamber or mirror chamber.
[0109] Step 603 is followed by step 605, in which a source reference position is determined in the first room. The source reference position can be, for example, the center of the room, a corner of the room, or, for example, any position within the room. In many embodiments, the center of the room can be advantageously used as the source reference position, as this can allow the average offset to the source reference position to be minimized (for most practical audio source distributions). The location of the source reference position is typically not important, and in many embodiments, any position within the original room can be determined as the source reference position. Indeed, in some embodiments, the source reference position can be updated or changed (although this will typically be at a low update rate, as this would require a new determination of the corresponding reference position in the mirror room).
[0110] Step 605 is followed by step 607, in which a reference position is determined in the mirror chamber. The mirror reference position for a given mirror chamber is the position within the mirror chamber obtained by applying the mirroring occurring in the mirror chamber to the source reference position. Thus, applying a series of one or more mirrorings relative to the boundary of the previous mirror chamber (the original chamber is the first mirror chamber) to the source reference position results in the mirror reference position.
[0111] The source reference position determination (605) and the mirror chamber reference position determination (607) can be combined with the mirror chamber determination (603). In particular, in some embodiments, one of the corners of the original chamber can be the source reference position, and the corresponding mirrored corner can be the reference position in the mirror chamber.
[0112] Step 607 is followed by step 609, in which a mirror room is selected. The method then determines a mirror position corresponding to mirroring the audio source position of the original room into the mirror room. In this way, the mirror position corresponds to the position of the direct path to the listening position that matches the reflected path in the original room. In this approach, the determination of the mirror position is not based on (iterative) mirror operations being applied to audio source positions in different rooms. Rather, a relative position offset is determined, a (direct) mapping to the mirror room is applied, and the mirror position is determined based on the mapped offset.
[0113] In particular, step 609 is followed by step 611, in which an offset mapping for the current mirror chamber is determined. In many embodiments, the mapping is retrieved from storage / memory. For example, a mapping may be determined in step 603 and stored in memory for each generated mirror chamber. In this manner, step 611 can simply extract the appropriate mapping for the current mirror chamber stored in memory. The mapping data typically includes a mapping matrix, a source reference position in the original chamber, and a reference position in the current mirror chamber.
[0114] Step 611 is followed by step 613, in which the current audio source for which the mirror position is to be determined is selected.
[0115] Step 613 is followed by step 615 in which the mirror position of the selected audio source within the selected mirror room is determined using the current mirror room mapping data.
[0116] The method then returns to step 613, where the next audio source is selected, and then the method proceeds to step 615, where the mirror position for the new audio source is determined. Step 613 may advantageously select only sources that have moved a significant amount since the previous mirror source calculation for the current room, or that are newly introduced. If no further audio sources need to be processed, the method returns to step 609, where the next mirror room is selected, and then the method determines the mirror position for the next mirror room. If no further mirror rooms need to be processed, the method proceeds to step 617.
[0117] It will be understood that in different embodiments, different criteria for determining how (or what) mirror rooms and audio sources should be processed may be used, depending on the preferences and requirements of a particular implementation. For example, in many embodiments, all mirror rooms up to a given number N of mirror rooms that can be generated by mirroring may be processed (and generated), and the mirror positions in all these mirror rooms for all point audio sources may be determined. In other embodiments, other selections may be used, such as, for example, including only audio sources whose volume / level exceeds a given threshold, or reflections whose distance is below a given threshold. It will also be understood that the order in which the mirror rooms and audio sources are passed through may be different in different embodiments (the order / nesting of loops may be changed), and / or parallel processing may be applied, for example.
[0118] Step 615 is followed by step 617, in which an audio model / room response is generated for the room and the current audio source and position. Specifically, the model may determine, for each (or perhaps only a subset) of the audio sources, a transfer function that includes a representation of the direct path, early reflections, and reverberation. Early reflections are represented by contributions determined based on the mirror position (e.g., as a single tap for each reflection). That is, reflections may be represented by contributions corresponding to the direct path from the mirror position and corresponding to the audio source modified by a transfer function that takes into account the reflective properties of the boundaries.
[0119] The model may be two impulse responses representing binaural impulse responses that depend on the listener's position and orientation, or in many embodiments, the model may be multiple impulse responses per source, each impulse associated with a different position in the original room or loudspeaker in the user's playing area.
[0120] The model will typically also include reflection coefficients / filters or similar material properties to model the attenuation due to reflections at the boundaries of the room (possibly frequency dependent).
[0121] Step 617 is followed by step 619, in which an audio signal is generated based on the audio sources and the determined mirror positions. Specifically, the audio signal may be generated by determining the audio signal contribution for each audio source using the transfer function determined in step 617 and generating the audio signal by combining these audio components.
[0122] An audio signal may be generated based on the impulse responses generated in step 617. This may be done by convolving the signal with the impulse responses, possibly in two or more processing steps. For example, the source signal may first be processed with the impulse responses, and then HRTF processing of the resulting signals may be applied with HRTFs corresponding to positions associated with each impulse response in the original room relative to the current user position. The latter approach is particularly advantageous when the user is wearing headphones and is head-tracked. This may, for example, avoid the need to rapidly recalculate the impulse responses.
[0123] In many embodiments, a binaural stereo signal containing directional cues can be generated, which can be generated using binaural filtering and processing, such as HRTF or BRIR processing, as is well known in the art.
[0124] It will be appreciated that many different approaches, algorithms and processes are known for generating audio signals based on audio source and location, and any suitable approach may be used without departing from the invention.
[0125] The approach uses an approach for determining mirror positions that is particularly efficient and allows for significantly reduced computational requirements, which may allow for faster updates of positions and / or more realistic processing for a given computational resource, and may allow for significantly faster audio processing that may result in a significantly improved user experience.
[0126] The process of step 615 may specifically implement the steps of the approach of Figure 7. The process begins in step 701, where a source position offset is determined for the current audio source. The source position offset represents the spatial offset / difference between the source reference position and the position of the audio source. Specifically, the source position offset may be determined as a vector from the source reference position to the audio source position (or, in some embodiments, from the source position offset to the source reference position). For example, the processing circuit 203 may subtract the coordinates of the source reference position from the coordinates of the audio source position to generate the source position offset. The offset is typically represented by a vector.
[0127] In some embodiments, a (typically monotonic) function can be applied, for example, to the above coordinates or differences, to determine a source position offset that reflects the difference between the source position offset and the source reference position, but is not directly a vector between these positions (e.g., a (possibly non-linear) scaling can be applied).
[0128] Step 701 is followed by step 703, in which the retrieved mapping for the current mirror room is applied to the source position offset to generate a mirror position offset. The mapping is such that the mirror position offset corresponds to the source position offset but takes into account the change in direction resulting from mirroring the source position offset according to the series of mirrorings that result in the mirror room. Specifically, in many embodiments, the mapping may be a distance-preserving mapping such that the size of the mirror position offset (vector) is the same as that of the source position offset (vector). The mapping may perform a directional mapping from the direction of the source position offset to the direction of the mirror position offset resulting from mirroring the source position offset according to the series of mirrorings.
[0129] In some embodiments, such as when generating the source position offset involves scaling or other functions, the mapping may take such scaling into account and may include an inverse function, or the inverse function may be applied later, and indeed such functions and inverse functions may be considered part of the mapping and included in the mapping operation.
[0130] Step 703 is followed by step 705, in which the mirror position of the audio source (i.e., the position within the mirror room that represents the reflection modeled by the mirror room) is determined from the mirror reference position and the mirror position offset. Specifically, the mirror reference position may be offset according to the mirror position offset to yield the mirror position. For example, in many embodiments, the mirror position offset may be a vector that is added to the coordinates of the mirror reference position to yield the coordinates of the mirror position.
[0131] In this way, rather than performing individual iterative mirroring of the audio source position to determine the mirror position, the approach instead determines the relative offset with respect to a reference position in the original room, which is then directly mapped to an offset in the mirror room, which is combined with the mirror reference position to determine the mirror position. The inventors have realized that by using the position offset relative to the reference position in the original room, a direct mapping can be applied to the offset. In particular, the mapping applied to the offset is independent of the position (or offset) of the audio source and therefore can be determined once and applied to all positions. Furthermore, the mapping can reflect mirroring without the need for a mirroring operation to be performed by reflecting the change in orientation resulting from mirroring. In particular, in many embodiments, the distance / size of the mirror position offset and the source position offset can be the same, and the mapping only needs to map the orientation that reflects the mirroring between the original room and the mirror room.
[0132] In some embodiments, the mapping may be implemented, for example, as a predetermined function applied to the source position offset. Specifically, in many embodiments, the mapping may be implemented as a look-up table (LUT), where the source position offset, represented by a set of coordinates, is the input to the table lookup, and the output of the LUT corresponds to the coordinate of the mirror position offset.
[0133] In many embodiments, the mapping can be represented by a mapping matrix, and the mapping can be applied to a source position offset by multiplying the source position offset, represented as a vector, by the mapping matrix. Thus, multiplying the mapping matrix by the source position offset vector results in a mirror position offset vector. In many embodiments, the process is performed in three-dimensional space using a vector containing three coordinates and a mapping matrix that is a 3x3 matrix, thereby mapping a three-component (specifically three-dimensional) vector to another three-component (specifically three-dimensional) vector.
[0134] In some embodiments, the processing need not be three-dimensional, but may be two-dimensional, for example. For example, it may be assumed that all audio sources and listening positions are at the same height, and therefore only the horizontal positions may vary while the vertical positions remain constant. In such cases, the offsets may be two-dimensional, specifically representing offsets only in the two-dimensional horizontal plane. In such cases, the mapping may be two-dimensional, specifically, the mapping matrix may be a 2×2 matrix that maps two-dimensional source position offsets to two-dimensional mirror position offsets. While such an approach may not offer the same flexibility as three-dimensional processing, it may offer more efficient processing with reduced complexity and reduced computational resource usage.
[0135] Specifically, each virtual mirror source represents a reflection of the audio source in the original room.
number
number
number
number
number
number
[0136] This approach reflects the inventors' realization that a movement of a certain distance in the original chamber can result in a movement of the same distance in each mirror chamber. The position delta / offset of the corresponding virtual mirror source in the mirrored chamber will not be the same as in the original chamber due to the mirroring operation, especially if the chamber is misaligned with the coordinate axes. As a result, for a given direction in the original chamber, it is not trivial which direction the mirrored source moves. However, in this approach, rather than determining this characteristic by performing a mirroring operation, the method is configured to directly map the offset in the original chamber to the corresponding offset in the mirror chamber. Such a mapping can be advantageously used even when the boundaries are not aligned with the coordinate axes, providing particularly efficient operation and resource conservation in this case.
[0137] The approach may be based on determining, for example, only once (e.g., as part of an initialization phase), for at least one reference position in the original room, a corresponding reference position in the mirror room with appropriate mirroring. Similarly, the mapping for the orientation, in particular the mapping matrix, may be determined only once.
[0138] The use of mapping, and in particular a mapping matrix, allows for reduced update complexity for moving audio sources. The mapping can capture how the coordinates of a point in the mirrored room change as a function of the change along an axis of the corresponding point in the original room. Thus, any known position in the original room and the corresponding virtual mirror source position can be used to calculate the virtual mirror source position for any new position in the original room. The approach can be used to update existing mirror source positions or generate new mirror source positions based on an offset to the reference position and applying a relative correction matrix to the corresponding reference mirror source position.
[0139] This approach allows the mirror source position relative to a moving source to be updated periodically without computing (successive) mirroring operations, and may therefore result in significantly improved performance.
[0140] It will be appreciated that this approach can also be used to determine mirror chamber characteristics. For example, during initialization, the mapping, source reference position, and mirror reference position can be determined. Based on these determinations, processing circuitry 203 can determine mirror chamber characteristics. For example, for each corner, the relative offset between the corner and the source position offset can be determined, the mapping can be applied to the offset, and the resulting offset can be combined with the mirror reference position to determine the corresponding mirror chamber corner. Such an approach can provide an efficient approach to determining mirror chamber characteristics.
[0141] Determining this mapping can be performed in different ways in different embodiments. For example, if the mirror chamber is a single mirroring relative to the boundary of the original chamber, the mapping can be considered by performing a mirror operation on positions within the mirror chamber with offsets corresponding to unit vectors aligned with the axes of the coordinate system. For example, the processing circuit 203 can determine a first position by adding a [1,0,0] vector to the source reference position. The circuit can then determine the offset within the mirror chamber by subtracting the mirror reference position from the mirrored position. This offset represents the mapping of the [1,0,0] offset, i.e., the dependence of the mirror position offset on the first coordinate of the source position offset. In this way, this coordinate can be used as the first row of the mapping matrix for the mirror chamber. The same approach can be applied to offsets of [0,1,0] and [0,0,1], respectively, to determine the second and third rows of the mapping matrix.
[0142] For other mirror rooms resulting from multiple mirroring sequences, the same unit vector can be used for the source room, and the mirroring sequence can be performed with the resulting position within the mirror room subtracted from the mirror reference position to provide the offset for that mirror room, i.e., in the appropriate row of the mapping matrix for that mirror room.
[0143] In many embodiments, reduced complexity and computational load can be achieved by using an iterative / correlated approach. In particular, in the case of shoebox-shaped chambers, mirroring will be relative to the corresponding boundaries. Thus, mirroring relative to one boundary of the mirror chamber will be identical to mirroring relative to the corresponding boundary of the original chamber. Thus, such mirroring will match the corresponding mirroring from the original chamber. In this way, the mapping (matrix) determined for the mirror chamber as a result of mirroring the original chamber also applies to the current mirroring. Therefore, the mapping matrix can also be used to represent this mirroring. That is, the mirroring can be represented by the already determined mapping matrix. In this way, in the case of mirror chambers resulting from a series of mirrorings, each mirroring corresponds to a mapping matrix determined relative to the boundary of the original chamber. Therefore, the overall mapping matrix for the mirror chamber can be determined as a result of multiplying the individual mapping (sub) matrices corresponding to each individual mirroring. Specifically, the mapping matrix for the current mirror room can be determined by multiplying the mapping matrix for the last / new mirroring that will produce the current mirror room (this mapping matrix was determined for the corresponding boundary of the original room) by the mapping matrix determined for the mirror room in which this mirroring will be performed. In this way, an iterative approach can be used.
[0144] In many embodiments, a set of boundary mirror mapping matrices can be determined that represent the change in orientation resulting from mirroring relative to a single chamber boundary. The boundary mirror mapping may represent the mapping to be applied to the resulting offset of the single mirroring relative to the boundary of the original chamber. Such a single boundary imaging mapping matrix may be referred to as a boundary mirror mapping matrix. The set of boundary mirror mapping matrices may specifically include mapping matrices for mirror chambers resulting from a single mirroring of the source chamber.
[0145] In many embodiments, several of the mirrorings for the boundaries may result in the same boundary mirror mapping matrix. In particular, for shoebox-shaped rooms, the boundary mirror mapping matrices for parallel boundaries (e.g., opposing walls) are the same. Thus, in some embodiments, parallel room boundaries of a first room are linked with the same boundary mirror mapping matrix. In some embodiments, at least two parallel room boundaries of a first room are linked with a single boundary mirror mapping matrix.
[0146] Thus, in some embodiments, the set of boundary mirror mapping matrices may include fewer matrices than the number of boundaries, e.g., for a three-dimensional process, the set of boundary mirror mapping matrices may include three boundary mirror mapping matrices, while for a two-dimensional process, the set of boundary mirror mapping matrices may include two boundary mirror mapping matrices.
[0147] This approach may reduce the number of mirroring operations that need to be performed and instead make extensive use of mapping operations that can be performed with significantly less computational resources.
[0148] Specifically, during initialization, a single source position offset within the simulated room can be mirrored to one or more orders of magnitude relative to its boundaries (walls, floor, ceiling).
[0149] Typically, this is achieved by the normalized normal vector n of the mirrored boundary, especially if the room is not axis-aligned. - (||n - (column vector with || = 1). For example, the values d, n(1)·x+n(2)·y+n(3)·z=d By first finding
[0150] This can be easily found by entering the x, y, z coordinates of a point within the boundary (for example, one of its corners). - To mirror the boundary surface, the nearest point is p, which conforms to the boundary surface equation. - +α·n - This is done by finding α in α=dp -T n - This is achieved by (·) T denotes transposition. In this case, the mirrored point is p - m =p - +2·α·n - is given as:
[0151] This principle can be used to mirror any location within a room. However, the computations are complex, and in many embodiments, the described approach can be used only to determine the mirror reference position. In particular, other locations (e.g., corner locations) can be determined using the described mapping approach, including, for example, the location of the room itself.
[0152] The mapping matrix can be derived from the normal vector of the mirror operation. Furthermore, the contributions for subsequent mirror operations can be combined by multiplying the matrices accordingly.
[0153] The matrix for a single mirroring operation is specifically:
number
number
[0154] The mapping matrix for a single mirroring operation is symmetric, with each row (and column) representing the resulting change in the mirror chamber when the position in the source chamber changes by 1. That is, row 1 corresponds to the (x,y,z) change in the mirror chamber from a change of +1 in the x direction of the chamber on the other side of the boundary under consideration, row 2 corresponds to a change of +1 in the y direction, and row 3 corresponds to a change of +1 in the z direction.
[0155] Each mirroring operation produces a new mirrored room and therefore a new mirrored source with a corresponding mapping matrix, e.g., M'. The mirrored room can then be mirrored again to obtain a mirrored source representing a higher order of reflection. The matrix calculated from the subsequent mirroring operation M" can then be combined with the matrix from the room that was mirrored to provide a mapping matrix representing the mirroring sequence M = M' · M".
[0156] By combining the mapping matrices from all mirroring operations in the sequence that produced a given mirror room from the original room, the combined mapping matrix will relate changes in the original room to changes in the corresponding mirrored room.
[0157] The mirror mapping matrix for a particular mirror chamber can be efficiently determined by multiplying only at most one boundary mirror mapping matrix from each of the different parallel pairs and the boundary mirror mapping matrices of mirrors that occur an odd number of times in the corresponding mirroring sequence. A boundary mirror mapping matrix becomes the identity matrix when multiplied by itself an even number of times.
[0158] After initialization, the reference position in the original room
number
number
[0159] As mentioned above, each virtual mirror source
number
number
number
number
number
number
[0160] In a typical application, i ranges over a large set of I chambers (e.g., 62 for 3rd order, 128 for 4th order, 230 for 5th order, etc.).
number
[0161] Specifically, for any additional or updated source positions in the original chamber, only one offset vector calculation is required, followed by multiplication of this offset vector with a matrix and summation with the corresponding mirror reference position vector. In addition to requiring few operations, these operations are particularly suitable for high-speed processing using modern processor architectures, and may be suitable for, for example, parallel processing architectures. They can be performed faster than approaches that repeat the mirroring technique for new or updated source positions.
[0162] For example, in the example given above where the 5th order reflections are modeled with an update rate of 90Hz: MOPS invention =90·(230·16·N sources +3) This results in a computational load of 0.3-1.7 MOPS to update 1-5 sources, with no room-finding bookkeeping or logic required. For a typical number of animation sources in a room, a reduction of at least 35-80% can be achieved.
[0163] Furthermore, while the algorithm for updating the mirrored source can be implemented on modern processors to run very efficiently using parallel hardware structures, re-execution of the more complex image source method algorithm using mirroring must be performed almost serially, moving from one room (i.e., one of 230 mirrored rooms) to the next in a small set (e.g., three). This further greatly increases the benefit of the proposed approach, which would require even less throughput time than the 35-80% reduction based on MOPS alone. For real-time processing of audio in AR / VR, this reduction is significant and may enable new applications and / or greatly improved user experiences for a given processing unit.
[0164] Where the terms audio and audio source are used above, it will be understood that this is equivalent to the terms sound and sound source, and any reference to the term "audio" can be replaced with a reference to the term "sound".
[0165] It will be understood that the above description, for clarity, has described embodiments of the invention in terms of different functional circuits, units, and processors. However, it will be apparent that any suitable distribution of functionality between different functional circuits, units, or processors may be used without departing from the invention. For example, functionality described as being performed by separate processors or controllers may be performed by the same processor or controller. Accordingly, references to specific functional units or circuits should not be considered to indicate a strict logical or physical structure or organization, but merely as references to suitable means for providing the described functionality.
[0166] The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the invention may be physically, functionally and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in multiple units or as part of other functional units. Thus, the invention may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits and processors.
[0167] Although the present invention has been described in connection with several embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Moreover, while features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0168] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Moreover, although individual features may be included in different claims, they may be advantageously combined, and their inclusion in different claims does not imply that combining the features is not possible and / or advantageous. Furthermore, the inclusion of a feature in one claim does not imply limitation to that claim, but rather indicates that the feature may be equally applicable to other claim classes, as appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features must be performed, and in particular the order of individual steps in a method claim does not imply that the steps must be performed in that order. Rather, the steps may be performed in any suitable order. Furthermore, singular references do not exclude pluralities. Thus, references to "a," "an," "first," "second," etc. do not exclude pluralities. Reference signs in the claims are provided merely as a clarifying example and are not to be construed as limiting the scope of the claims in any way.
Claims
1. 1. A computer-implemented method for determining a virtual audio source position of an audio source representing a reflection of a first audio source in a first room, the method comprising: receiving data describing a boundary of a first chamber; generating a group of mirror rooms for the first room, each mirror room resulting from a plurality of mirrorings, each mirroring being a mirroring of a previous mirror room relative to a boundary of the previous mirror room, the first previous mirror room for the plurality of mirrorings being the first room, the method comprising: for at least a first mirror chamber of the group of mirror chambers, mapping directions in the first chamber to directions in the first mirror chamber; determining a source reference position within the first chamber; determining a first mirror reference position within the first mirror chamber, the first mirror reference position being a position within the first mirror chamber that results from applying mirroring to the source reference position that results in the first mirror chamber; determining a source position offset in the first room for the first audio source, the source position offset representing a position offset between the source reference position and a position of the first audio source within the first room; determining a first mirror position offset in the first mirror chamber relative to the first audio source; determining a mirror position of the first audio source within the first mirror room from the first mirror reference position and the first mirror position offset; rendering an audio output signal including a component from the first audio source as if positioned at the mirror position; and 10. The method of claim 9, wherein determining the first mirror position offset comprises determining the first mirror position offset by applying the mapping to the source position offset.
2. The method of claim 1 , wherein the mapping is represented by a mapping matrix, and applying the mapping to the source position offset comprises multiplying the mapping matrix by the source position offset.
3. The method of claim 2 , wherein the source position offset is a two-dimensional offset and the mapping matrix is a 2×2 matrix.
4. The method of claim 2 , wherein the source position offset is a three-dimensional offset and the mapping matrix is a 3×3 matrix.
5. 5. The method of claim 2, wherein the number of mirrorings for the first room is at least two, and the mapping matrix is a combination of multiple boundary mirror mapping matrices, each boundary mirror mapping matrix representing the change in direction resulting from mirroring for a single room boundary.
6. 6. The method of claim 5, wherein each room boundary of the first room is linked to one boundary mirror mapping matrix, and the mapping is a combination of the boundary mirror mapping matrices linked to room boundaries of mirrorings among the plurality of mirrorings for the first mirror room.
7. The method of claim 6 , wherein parallel chamber boundaries of the first chamber are linked with the same boundary mirror mapping matrix.
8. The method of claim 1 , wherein the mapping is a distance-preserving mapping.
9. The method of claim 1 , wherein the distance of the first mirror position offset is equal to the distance of the source position offset.
10. determining a chamber boundary location relative to a boundary of the first chamber; determining a boundary position offset in the first chamber relative to the chamber boundary position, the boundary position offset representing a position offset between the source reference position and the chamber boundary position; determining a boundary position offset in the first mirror chamber by applying the mapping to the boundary position offset; determining a mirror boundary position for the first mirror chamber from the first mirror reference position and the boundary position offset; 10. The method of claim 1, further comprising:
11. 11. The method of claim 1, wherein the first chamber is a rectangular parallelepiped, and the source reference position and the source position offset are represented by coordinates of coordinate axes that are not aligned with sides of the rectangular parallelepiped.
12. 12. The method of claim 1, further comprising determining a room response function for the first room, the room response function including a reflected component representative of audio from the audio source as located at the mirror position.
13. A computer program comprising computer program code means for performing all the steps of the method according to any one of claims 1 to 12 when said computer program is run on a computer.
14. 1. An apparatus for determining a virtual audio source position of an audio source representing a reflection of a first audio source in a first room, the apparatus comprising: receiving data describing a boundary of a first chamber; generating a group of mirror chambers for the first chamber, wherein each mirror chamber results from a plurality of mirrorings, each mirroring being a mirroring of a previous mirror chamber relative to a boundary of the previous mirror chamber, and a first previous mirror chamber for the plurality of mirrorings being the first chamber; In an apparatus having a processing circuit, the processing circuit further for at least a first mirror chamber of the group of mirror chambers, performing a mapping of directions in the first chamber to directions in the first mirror chamber; determining a source reference position within the first chamber; determining a first mirror reference position within the first mirror chamber, wherein the first mirror reference position is a position within the first mirror chamber that results from applying mirroring that results in the first mirror chamber to the source reference position; determining a source position offset in the first room for the first audio source, wherein the source position offset represents a position offset between the source reference position and a position of the first audio source within the first room; determining a first mirror position offset in the first mirror chamber relative to the first audio source; determining a mirror position of the first audio source within the first mirror room from the first mirror reference position and the first mirror position offset; rendering an audio output signal including a component from the first audio source as if positioned at the mirror position; 10. The apparatus of claim 9, wherein determining the first mirror position offset comprises determining the first mirror position offset by applying the mapping to the source position offset.
Citation Information
Patent Citations
Apparatus and method for determining virtual sound sources
EP3828882A1
Real-time sound reproduction system
JP2005080124A