Methods, apparatus, and systems for modeling audio objects with spatial properties.
By modeling spread audio objects using simplified parameters based on user position and geometric form, the method addresses the complexity of 6DoF rendering, enabling efficient and immersive audio experiences in virtual and augmented reality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- DOLBY INTERNATIONAL AB
- Filing Date
- 2022-04-28
- Publication Date
- 2026-05-15
AI Technical Summary
Existing audio rendering systems struggle to provide a realistic and immersive 6DoF experience due to high computational complexity from complex interactions between changing listening positions and audio object spreads, particularly in virtual reality environments, as they do not adequately account for significant translational movements of the listener.
A method and device for modeling spread audio objects by determining a spread parameter based on user position, using geometric form and relative points, allowing for adaptive modeling of audio objects with simplified parameters, reducing the need for detailed data processing.
This approach simplifies the modeling of audio objects, enabling efficient 6DoF rendering by converting complex 6DoF data into simpler parameters, thereby reducing computational complexity and enhancing the realism of audio experiences in virtual and augmented reality environments.
Smart Images

Figure 0007859616000001 
Figure 0007859616000002 
Figure 0007859616000003
Abstract
Description
[Technical Field]
[0001] Cross-references to related applications This application claims priority to the following priority applications: U.S. Provisional Application No. 63 / 181,865 (reference number: D21045USP1) filed on 29 April 2021, U.S. Provisional Application No. 63 / 247,156 (reference number: D21045USP2) filed on 22 September 2021, and European Patent Application No. 21200055.8 (reference number: D21045EP) filed on 30 September 2021.
[0002] Technical field This paper concerns object-based audio rendering, more specifically, rendering audio objects with spatial capabilities in virtual reality (VR) environments. [Background technology]
[0003] The new MPEG-I standard enables audio experiences from different viewpoints and / or perspectives or listening positions by supporting full 6 degrees of freedom (6DoF) in virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or extended reality (XR) applications. 6DoF interaction extends the 3DoF spherical video / audio experience, which is limited to head rotation (pitch, yaw, and roll), to include translational movement (forward, backward, up, down, and left, right), allowing for navigation within the virtual environment (e.g., physically walking in a room) in addition to head rotation.
[0004] For audio rendering in VR applications, an object-based approach is widely adopted, which involves representing complex auditory scenes as multiple distinct audio objects, each associated with parameters or metadata defining its location and trajectory within the scene. Instead of being point sources, audio objects may be given a spatial extent that reflects the auditory perception derived from them. Such audio objects may emit one or more sound sources to be rendered in the VR implementation.
[0005] To create a natural and realistic 6DoF experience for the listener, the listener's experience of the directionality and spatial extent of sound or audio sources (objects) is crucial for 6DoF rendering, particularly for realizing the experience of navigation around virtual audio sources through a scene. Since 6DoF rendering involves even greater translational changes in the listener's listening position, the complex interaction between constantly changing listening positions and the sophisticated structure of audio object spreads can present challenges in 6DoF rendering implementations. In particular, modeling such position-object interactions requires many more parameters, which leads to very high computational complexity in the corresponding audio processing.
[0006] It should be noted that available audio rendering systems (such as MPEG-H 3D audio renderers) are typically limited to 3DoF rendering (i.e., rotational motion of the audio scene caused by the listener's head movement), which does not take into account translational changes in the listener's listening position. Even 3DoF+ only adds small translational changes in the listener's listening position and does not take into account larger translational movements of the listener. Therefore, existing techniques that do not take into account larger translational movements of the listener may encounter difficulties in truly immersive rendering of 6DoF sound. [Overview of the project] [Problems that the invention aims to solve]
[0007] Therefore, there is a need to provide a simple way to implement 6DoF rendering of audio objects. In particular, it is sometimes desirable to simplify the modeling of the (spatial) extent of audio objects, taking into account significant user movement for 6DoF rendering. [Means for solving the problem]
[0008] In one aspect, a method for modeling (e.g., computer-implemented) an extended audio object for audio rendering in a virtual reality or augmented reality environment (or more generally, a computer-mediated reality environment) is described. The method may include the step of obtaining an extended representation showing the geometric form of the extended audio object and information about one or more first audio sources associated with the extended audio object. One or more first audio sources may be captured as recorded audio sources associated with the extended audio object using audio sensors. Specifically, the method may include the step of obtaining a relative point using the geometric form of the extended audio object (e.g., the extended representation showing the geometric form) based on the user's position in the virtual or augmented reality environment (i.e., the listener's listening position). Furthermore, the method may include the step of determining an extended parameter of the extended representation based on the user's position and the relative point.
[0009] In particular, the spread parameter can describe the spatial extent of a spread audio object as perceived at the user's position. Therefore, it can be understood that such spatial extent of a spread audio object can vary according to the user's position, and that spread audio objects can be adaptively modeled for various user positions. To effectively model spread audio objects, the method may also include determining the positions of one or more second audio sources relative to the user's position. Such one or more second audio sources can be considered virtual played audio sources for modeling the spread audio object at the corresponding user position. Furthermore, the method may include outputting a modified representation of the spread audio object to model it. Note that the modified representation includes the determined spread parameter and the positions of one or more second audio sources.
[0010] The proposed method, configured as described above, allows for the modeling of spread audio objects using simple parameters. In particular, using knowledge of the spatial extent of a spread audio object and the corresponding positions of a second (virtual) audio source(s) calculated for a given user position, the spread audio object can be effectively modeled to have an appropriate (perceived) size corresponding to a given user position, which may be applicable to subsequent renderings of the spread audio object (e.g., 6DoF). This reduces the complexity of the audio rendering calculations, as detailed information about the form / position / orientation of the audio object and the movement of the user position may not be required.
[0011] In other words, the proposed method effectively converts 6DoF data (e.g., input audio object source, user position, object extent geometry, object extent position / orientation, etc.) into simple information for use as input for an extent modeling interface / tool. This allows for efficient 6DoF rendering of audio objects without requiring the processing of massive amounts of data. In one embodiment, the extent representation, which shows the geometric form of the extent audio object, corresponds to (matches) the geometric form of the extent audio object. For example, for relatively simple geometric forms, the geometric form of the extent audio object can be used as the extent representation.
[0012] In one embodiment, the spread parameter may be determined further based on the position and / or orientation of the spread audio object. The method may also further include determining one or more second audio sources for modeling the spread audio object based on one or more first audio sources. According to this embodiment, the method may further include determining the extent angle based on the user position, a relative point, and the position and / or orientation of the spread audio object. For example, the extent angle may be an arc measure indicating the spatial extent of the spread audio object as perceived at the user position. Thus, the extent angle may refer to a relative arc measure (i.e., relative extent angle) that depends on the relative point and the position and / or orientation of the spread audio object. In this case, the spread parameter may be determined based on the (relative) extent angle.
[0013] Structured as described above, the proposed method provides a simplified approach to obtaining accurate estimations of the spatial extent / perceived size of audio objects at different user positions, thereby improving performance when modeling audio objects using simple parameters.
[0014] In some embodiments, determining the position of one or more second audio sources may include determining an arc based on the user position, a relative point, and the geometric form of the spread audio object. In addition, determining the position of one or more second audio sources may further include positioning the determined one or more second audio sources on the arc. Furthermore, the arc may include an arc related to the (relative) spread angle as a corresponding arc measure at the user position, and may be determined based on the spread angle and the user position. In some embodiments, positioning may include distributing all second audio sources equally spaced on the arc. Positioning may also depend on the correlation level between the second audio sources and / or the content creator's intent. In other words, the second audio sources may be placed on the arc at appropriate distance intervals determined based on the correlation level between the second audio sources and / or the content creator's intent.
[0015] In one embodiment, the spread parameter may be determined further based on the number (i.e., count) of one or more determined second audio sources. In particular, the number of one or more determined second audio sources may be a predetermined constant independent of the user position and / or relative point. Alternatively, determining one or more second audio sources for modeling a spread audio object may include determining the number of one or more second audio sources based on a (relative) spread angle. In this case, the number of one or more second audio sources may increase as the spread angle increases (i.e., the number may be positively correlated with the spread angle). More specifically, determining one or more second audio sources for modeling a spread audio object may further include duplicating one or more first audio sources, or adding a weighted mixture of one or more first audio sources, and applying a decorrelation process to the duplicated or added first audio sources. That is, to obtain a determined number of second audio sources, one or more first audio sources may be duplicated, or a weighted mixture thereof may be added.
[0016] Configured as described above, and by appropriately defining a second (virtual) audio source, this method allows for the accurate and adaptable modeling of audio objects. In particular, modeling can be effectively performed for various user positions, for audio objects with different input sources and forms / positions, and for the intent of the content creator.
[0017] In one embodiment, the spread representation may indicate a two - dimensional or three - dimensional geometric form for representing the spatial spread of a spreading audio object. Further, the spreading audio object may be oriented in two dimensions or three dimensions. Also, the spatial spread of the spreading audio object perceived at the user position may be described as the perceived width, size, and / or massiveness of the spreading audio object.
[0018] The relative point may be obtained at the location closest to the user position within the virtual or augmented reality environment, using the geometric form of the spreading audio object (e.g., the spread representation indicating it).
[0019] In particular, the relative point may be a point on the geometric form of the spreading audio object (e.g., the spread representation indicating it) that is closest to the user position.
[0020] The inventors have surprisingly found that using the relative point closest to the user position to model a spreading audio object leads to better control of the audio - level attenuation for different user positions with respect to the spreading audio object, and thus leads to better modeling of the spreading audio object.
[0021] In one embodiment, obtaining the relative point using the spread representation indicating the geometric form of the spreading audio object includes obtaining the relative point on the geometric form or obtaining the relative point at a distance from the geometric form or the spread representation. For example, the relative point may be located on the geometric shape. Alternatively, the relative point may be located at a distance from the spread representation or the geometric form. For example, the relative point may be located at a distance from the boundary or origin of the spread representation or the geometric form.
[0022] In one embodiment, the method may further include obtaining a vertical projection of the extended audio object on a projection plane that is orthogonal to a first line connecting the user position and the relative point. The method may also include determining, on the vertical projection, a plurality of boundary points that identify the projected size of the extended audio object. In this case, the (relative) spread angle may be determined using the user position and the plurality of boundary points. For example, the spread angle may be determined by connecting two boundary points to the user position and determining the angle formed by two straight lines connecting each boundary point to the user position as the spread angle.
[0023] Specifically, determining the plurality of boundary points may include obtaining a second line related to the horizontal projected size of the extended audio object. Thus, the plurality of boundary points may include the leftmost boundary point and the rightmost boundary point of the vertical projection on the second line. Depending on the orientation of the extended audio object, the horizontal projected size may be the maximum size of the extended audio object. In some embodiments, the extended audio object may have a complex geometric shape, and the vertical projection may include a simplified projection of the extended audio object having a complex geometric shape. In this case, the method may further include obtaining a simplified geometric shape of the extended audio object for use in determining the relative point before obtaining the relative point on the geometric shape of the extended audio object.
[0024] Configured as described above, the method simplifies the estimation of the spatial spread / perceived size of audio objects having various geometric shapes while providing sufficient accuracy when modeling the audio object using simple parameters.
[0025] In one embodiment, the method may further include rendering a spread audio object based on a modified representation of the spread audio object. The spread audio object may be rendered using the determined positions and spread parameters of one or more second audio sources. In particular, the rendering may include 6DoF audio rendering. The method may further include obtaining the user position, the position and / or orientation and geometry of the spread audio object for rendering.
[0026] In one embodiment, the method may further include controlling the perceived size of a spread audio object using a spread parameter. Thus, a spread audio object can be modeled as a point source or a wide source by controlling the perceived size of the spread audio object. Note that the positions of one or more second audio sources may be determined such that all second audio sources have the same reference distance from the user position.
[0027] The above configuration allows for the effective modeling of expansive audio objects using simple parameters. In particular, the spatial extent of the expansive audio object and the corresponding position of the second (virtual) audio source are calculated for a given user position to allow for accurate estimation of the appropriate (perceived) size of the audio object corresponding to a given user position (i.e., the spatial extent size that can be perceived at that user position). Detailed information regarding the morphology / position / orientation of the audio object and the movement of the user position may not be necessary for modeling, so the processing for subsequent rendering of the audio object (e.g., 6DoF rendering) may be simplified accordingly.
[0028] In other words, the proposed method provides an automatic conversion of 6DoF data for audio spread modeling that may require simple parameters as input interface data (e.g., input audio object source, user position, object spread geometry, object spread position / orientation, etc.), further allowing for efficient 6DoF rendering of audio objects without complex data processing.
[0029] In another aspect, a device for modeling spread audio objects for audio rendering in a virtual or augmented reality environment (or more generally, a computer-mediated reality environment) is described. The device may include a processor and memory coupled to the processor that stores instructions for the processor. The processor may be configured to acquire a spread representation showing the geometric form of the spread audio object and information about one or more first audio sources associated with the spread audio object. One or more first audio sources may be captured using audio sensors as recorded audio sources associated with the spread audio object. Specifically, the processor may be configured to acquire relative points on the geometric form of the spread audio object based on the user's position in the virtual or augmented reality environment. In addition, the processor may be configured to determine spread parameters for the spread representation based on the user's position and the relative points.
[0030] In particular, the spread parameter can describe the spatial extent of a spread audio object as perceived at the user's position. Thus, it can be understood that such spatial extent of a spread audio object may vary according to the user's position, and that spread audio objects can be adaptively modeled for various user positions. Furthermore, the processor may be configured to determine the positions of one or more second audio sources relative to the user's position in order to model a spread audio object. Such one or more second audio sources can be considered virtual audio sources to be played for modeling a spread audio object at the corresponding user position. The processor may also be configured to output a modified representation of a spread audio object in order to model the spread audio object. In particular, the modified representation may include the spread parameter and the positions of one or more second audio sources.
[0031] Configured as described above, the proposed device effectively converts 6DoF data (e.g., input audio object source, user position, object spread geometry, object spread position / orientation, etc.) into simple information / parameters as input for a spread modeling interface / tool. This allows for efficient 6DoF rendering of audio objects without requiring the processing of massive amounts of data.
[0032] In particular, by using knowledge of the spatial extent of a spread audio object and the corresponding position of a second (virtual) audio source calculated for a given user position, the spread audio object can be effectively modeled to have an appropriate (perceived) size corresponding to a given user position, which may be applicable to subsequent renderings (e.g., 6DoF) of the spread audio object. This reduces the complexity of the audio rendering calculation, as detailed information about the shape / position / orientation of the audio object and the movement of the user position may not be required.
[0033] In another aspect, a system for implementing audio rendering in a virtual or augmented reality environment (or more generally, a computer-mediated reality environment) is described. The system may include the proposed apparatus (for example, as described above) and a spread modeling unit. The spread modeling unit may be configured to receive information from the apparatus regarding a modified representation of a spread audio object, as described above. Furthermore, the spread modeling unit may be configured to further control the spread size of the spread audio object based on the information regarding the modified representation (for example, spread parameters included in the modified representation). In some embodiments, the system may be a user virtual reality console (for example, a headset, computer, mobile phone, or any other audio rendering device for rendering audio in a virtual and / or augmented reality environment), or part thereof. In some embodiments, the system may be configured to transmit the information regarding the modified representation of the spread audio object and / or the controlled spread size to an audio output.
[0034] In a further aspect, a computer program is described. The computer program may include executable instructions for performing method steps outlined throughout this disclosure when executed by a computing device.
[0035] In another aspect, a computer-readable storage medium is described. The storage medium can store a computer program adapted for execution on a processor and, when executed on the processor, to perform the method steps outlined throughout this disclosure.
[0036] It should be noted that methods and systems including preferred embodiments thereof as outlined in this patent application may be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this patent application may be combined in any way. In particular, the features of the claims can be combined with each other in any way.
[0037] It will be understood that the features of the apparatus and the steps of the method can be replaced in many ways. In particular, the details of the disclosed method can be implemented by the corresponding apparatus, as those skilled in the art will understand, and vice versa. Furthermore, it will be understood that any of the above statements made with respect to the method (and, for example, its steps) are equally applicable to the corresponding apparatus (and, for example, its blocks, stages, units), and vice versa. [Brief explanation of the drawing]
[0038] The present invention is described below illustratively with reference to the accompanying drawings. [Figure 1] This is a conceptual diagram of an exemplary spread modeling tool according to an embodiment of the present disclosure. [Figure 2a]This embodiment of the disclosure shows an exemplary audio scene including different user positions for implementing audio rendering of a spread audio object. [Figure 2b] This figure shows the spread level for the corresponding user position in the exemplary audio scene shown in Figure 2(a) according to embodiments of the present disclosure. [Figure 3] An exemplary flowchart for implementing audio rendering of a spread audio object according to embodiments of this disclosure is shown. [Figure 4] An exemplary block diagram for implementing audio rendering of a spread audio object according to embodiments of this disclosure is shown. [Figure 5a] A schematic diagram is shown for determining a modified representation of an expansive audio object, as performed in Method 300 according to embodiments of the present disclosure. [Figure 5b] A schematic diagram is shown for determining a modified representation of an expansive audio object, as performed in Method 300 according to embodiments of the present disclosure. [Figure 5c] A schematic diagram is shown for determining a modified representation of an expansive audio object, as performed in Method 300 according to embodiments of the present disclosure. [Figure 6] This is another schematic diagram for determining a modified representation of an expansive audio object, as performed in Method 300 according to an embodiment of the present disclosure. [Figure 7] (a) and (b) show the definition of a reference distance for an object source with extent. [Figure 8a] The embodiments of this disclosure show the modified representations obtained for each of the spread audio objects for different user positions, as shown in Figure 2. [Figure 8b] The embodiments of this disclosure show the modified representations obtained for each of the spread audio objects for different user positions, as shown in Figure 2. [Figure 8c] The embodiments of this disclosure show the modified representations obtained for each of the spread audio objects for different user positions, as shown in Figure 2. [Modes for carrying out the invention]
[0039] As outlined above, this disclosure relates to the effective modeling of spread audio objects for audio rendering in virtual and / or augmented reality environments (or generally computer-mediated reality environments). Figure 1 shows a conceptual diagram of an exemplary spread modeling tool according to an embodiment of this disclosure. Here, the spread audio object 101 to be modeled is associated with one or more audio sources 102. The audio sources 102 may be captured as recorded audio sources using audio sensors (e.g., microphones). Generally, the spread audio object 101 can be considered a spread audio object having a geometric form. It may be provided with a spread representation showing the geometric form and information about one or more audio sources 102. In addition, one or more audio sources 102 may include, for example, one or more point source signals associated with the spread audio object 101. The spread modeling tool 103 can model the spread audio object 101 based on the form of the spread audio object 101 and information about one or more audio sources 102. For example, the positions of one or more audio sources 102 may be included in the spread expression. The one or more audio sources 102 themselves may or may not be included in the spread expression.
[0040] The spread modeling tool 103 may also model spread audio objects 101 based on the user's position in a virtual and / or augmented reality environment (e.g., the listener's listening position). That is, depending on the user's position, spread audio objects 101 can be modeled as audio sources (e.g., wide sources or point sources) with different spread sizes. This can be achieved by providing modified representations of spread audio objects 101 for specific user positions based on the (original) spread representation. Thus, spread audio objects 101 can be effectively modeled as having different spread sizes experienced / perceived at different user positions through their respective modified representations.
[0041] Figure 2(a) shows an exemplary audio scene including different user positions for implementing audio rendering of a spread audio object according to embodiments of the present disclosure. For example, the spread audio object may include a “beachfront” with large waves. Other examples may be known to those skilled in the art. The exemplary audio scene may be applied to implement 6DoF audio rendering in a virtual or augmented reality environment. In this embodiment, the object spread 201 is shown as a spread audio object having a two- or three-dimensional geometric form that can be oriented in two or three dimensions. An (original) spread representation showing the geometric form of the object spread 201 is obtained and includes information about the geometric shape, position, and orientation of the spread, as well as information about the audio source 202 associated with the object spread. The audio source 202 is shown here as two point sources, but any number and other types of audio sources may be realized in the context of the present disclosure. As described above, the audio source 202 may be a recorded audio source captured using an audio sensor (e.g., a microphone). Furthermore, user positions 203a, 203b, and 203c (which may be, for example, relative to object spread 201) are also obtained. In the illustrated exemplary scene, users 203a and 203b are located in front of object spread 201 but at different distances from it, and user 203c is located on one side of object spread 201. However, any other positions may also be included in the scene as user positions.
[0042] Figure 2(b) shows the spread levels for corresponding user positions in the exemplary audio scene shown in Figure 2(a) according to embodiments of the present disclosure. Here, spread levels 204a, 204b, and 204c represent the perception of object spread 201 at user positions 203a, 203b, and 203c, respectively. In these embodiments, spread levels may be defined using spread parameters that describe the spatial spread of object spread (spreading audio objects) as perceived at a particular user position. In particular, the perception of object spread 201 (and therefore the spread parameters) may depend on the relative geometric position and orientation of the user versus spread (e.g., the orientation of the user position relative to the object spread and / or the spreading object). For example, as shown in Figure 2(b), a user may experience a greater spread level 204a at user position 203a than at user positions 203b and 203c, which have corresponding spread levels 204b and 204c, respectively. Therefore, it may be beneficial to model spread audio objects in order to relate such spread levels to the user's position and simply use the spread parameter to render audio scenes where significant changes in the user's position (i.e., large translational motion) can occur.
[0043] Figure 3 shows an exemplary flowchart for implementing audio rendering of a spread audio object according to embodiments of the present disclosure. As illustrated, Method 300 may be performed to model a spread audio object, such as a spread audio object 101 or object spread 201, for audio rendering in a virtual and / or augmented reality environment. In step 301, a spread representation showing the geometric form of the spread audio object and information about one or more first audio sources associated with the spread audio object are obtained. In step 302, relative points on the geometric form of the spread audio object (e.g., the spread representation showing the geometric form) are obtained (e.g., determined) based on the user's position in a virtual or augmented reality environment. In step 303, a spread parameter (e.g., indicating a spread level representing the perceived / spatial spread of the spread audio object) is determined for the spread representation based on the user's position and the relative points. As described above, the spread parameter can describe the spatial spread of the spread audio object as perceived at the user's position.
[0044] In step 304, the positions of one or more second audio sources relative to the user's position are determined in order to model a spread audio object. Unlike the first audio source, which may have been captured through direct recording, the one or more second audio sources may be virtual, reproduced audio sources determined based on the first audio source(s), for example, through duplication and / or audio processing (including filtering), as will be described in detail below. Then, in step 305, a modified representation of the spread audio object is output in order to model the spread audio object. Note that the modified representation may include a spread parameter and the determined positions of one or more second audio sources for a given user position. Thus, the spread audio object can be effectively modeled for a particular user position using simple parameters that include knowledge of the spatial spread of the spread audio object and / or the corresponding positions of the second audio sources, calculated for this particular position.
[0045] In other words, the proposed method 300 effectively converts 6DoF data (e.g., input audio object source, user position, object spread geometry, object spread position / orientation, etc.) into simpler information (e.g., spread parameters and the position of a second audio source included in the modified representation) as input for a spread modeling interface / tool (e.g., spread modeling tool 103), which in some implementation forms may be a legacy interface / tool.
[0046] Furthermore, subsequent rendering of the spread audio object may be performed based on a modified representation of the spread audio object. In this case, the spread audio object may be rendered using the determined positions and spread parameters of one or more second audio sources. In some embodiments, the rendering may be a 6DoF audio rendering of the spread audio object. In this case, in addition to the user position, the position and / or orientation and geometry of the spread audio object may be obtained for rendering. Thus, the spread parameters may be determined further based on the position and / or orientation of the spread audio object.
[0047] Figure 4 shows an exemplary block diagram for implementing audio rendering of a spread audio object according to embodiments of the present disclosure. In particular, System 400 comprises a device for modeling a spread audio object for audio rendering in a virtual and / or augmented reality environment. In some embodiments, System 400 may be or part thereof a user virtual reality console such as a headset, a computer, a mobile phone, or any other audio rendering device for rendering audio in a virtual and / or augmented reality environment.
[0048] In some embodiments, the apparatus may take the form of a parameter conversion unit 401, which comprises, for example, a processor configured to perform all the steps of method 300, and a memory coupled to the processor and storing instructions for the processor. In particular, the parameter conversion unit 401 may be configured to receive audio scene data such as 6DoF data, which includes information about the spread geometry, position, and / or orientation of an input audio object source, user position 403, and a spread audio object 402 (e.g., a spread audio object 101 or object spread 201). The parameter conversion unit 401 may be further configured to perform steps 301-305 of method 300 as described above. Thus, the parameter conversion unit 401 converts the received audio scene data into simplified information (e.g., as a modified representation of the spread audio object). The simplified information includes information about a second (virtual) audio object 404 (e.g., object location and signal data of a second audio source(s)) and a spread parameter 405 indicating a spread level representing the perceived / spatial spread of the spread audio object experienced at a particular location. The parameter conversion unit 401 may transmit this (simplified) information, directly or via other processing components, to an audio rendering unit (e.g., internal or external to system 400) for outputting audio to the user (or alternatively, as part of an audio rendering unit that outputs audio to the user, for example, via a suitable device speaker). Thus, when system 400 is part of an audio rendering device, it may transmit the converted parameters (e.g., the simplified information described above relating to the modified representation) to the audio output of the audio rendering device.
[0049] In some embodiments, the simplified information output by the parameter conversion unit 401 may then be provided to the spread modeling unit 406 as input interface data. The spread modeling unit 406 (also known as the spread modeling tool, e.g., spread modeling tool 103) can control the spread size of the spread audio object (e.g., to be rendered) based on the spread parameters contained in the simplified information. For example, the perceived size of the spread audio object may be controlled using the spread parameter, which may model the spread audio object as a point source or a wide source. Thus, an appropriate (perceived) size corresponding to a particular user position can be provided for subsequent rendering (e.g., 6DoF rendering) of the spread audio object simply by adjusting the spread parameter (e.g., spread level). This provides a simplified system for implementing 6DoF rendering of spread audio objects. As a result, detailed information about the form / position / orientation of audio objects and the (translational) movement of the user's position may not be required for rendering / modeling, which further allows 6DoF rendering to be performed by existing audio rendering techniques (e.g., those suitable for 3DoF rendering), thereby reducing the computational complexity of 6DoF audio rendering.
[0050] In other words, the automatic conversion of 6DoF scene data is provided by the proposed method 300 and / or system 400 for audio spread modeling, which may require simple parameters as input interface data, allowing for efficient 6DoF rendering of audio objects using existing systems available without requiring complex data processing for rendering.
[0051] Figures 5(a) to 5(c) show schematic diagrams for determining a modified representation of a spread audio object, as performed in Method 300 according to embodiments of the present disclosure. The spread representation is assumed to represent a three-dimensional geometric form for representing the spatial spread of a spread audio object oriented in three dimensions. In the exemplary embodiment shown in Figure 5, the spread representation represents a cuboid representing the spatial spread of a spread audio object as perceived at a given user position, which may be described as the perceived width, size, and / or weight of the spread audio object. However, the spread representation may also represent other three-dimensional shapes or more complex geometric shapes for representing the object's spread. An example of the first stage of translating a three-dimensional (3D) geometric shape into a two-dimensional (2D) user observation domain is shown in Figure 5(a). Subsequently, the second stage in the 2D user observation domain determines a spread parameter (e.g., spread level) for a given user position, as shown in Figure 5(b), and the third stage determines the positions of one or more second audio sources according to a one-dimensional (1D) view, as shown in the example in Figure 5(c).
[0052] As shown in the example in Figure 5(a), user 501 is positioned in front of a spread audio object 503 that is oriented in three dimensions. The spread audio object 503 is represented by a 3D geometric form that indicates the spatial extent of the spread audio object 503. According to an exemplary embodiment, a point 502 on the geometric form of the spread audio object 503 closest to user 501 may be obtained (e.g., determined) (e.g., as a relative point). Optionally, a projection plane 504 orthogonal to a first line 505 connecting the user position 501 and point 502 may be obtained (e.g., determined). On the projection plane 504, a vertical projection 506 of the spread audio object 503 may be obtained (e.g., determined). Subsequently, a second line 507 characterizing the horizontal size of the spread audio object 503 (e.g., on the projection) may be obtained (e.g., determined). The second line 507, together with the first line 505, can form a plane 508 (for example, as an observation plane). Thus, the first stage shown in Figure 5(a) transforms the 3D geometry of the extended audio object 503 into a 2D observation plane 508.
[0053] As seen in the example in Figure 5(b), multiple boundary points 509, 510 can be determined on the vertical projection 506. In particular, boundary points 509, 510 can include the leftmost and rightmost boundaries of the perspective projection 506 on the second line 507. Thus, multiple boundary points 509, 510 can identify the projection size of the spread audio object 503. Using the determined boundary points 509, 510 and user position 501, the spread angle x representing the spread level (for example, as a spread parameter) can be determined. 0 However, it can be calculated accordingly, for example, using trigonometry. That is, the spread angle x 0 This can be determined as the angle between two lines connecting the user position 501 and boundary points 509 and 510, respectively. Therefore, as described above, the spread angle x 0can refer to a relative point and a relative angular measure (i.e., relative spread angle) that depends on the position and / or orientation of the extended audio object. Further, the arc 513 may be determined based on the geometry of the user position 501, the relative point 502, and the extended audio object 503. In particular, the arc 513 is the spread angle x as the corresponding angular measure at the user position 501 0 and may be an arc related to the spread angle x 0 and may be determined based on the spread angle x
[0054] As described above, the position and / or orientation of the extended audio object may be used to determine the spread parameter. More specifically, the user position 501, the relative point 502, and the position and / or orientation of the extended audio object 503 may be used to determine the spread angle x on which the spread parameter may be based 0 As described above, the spread angle x 0 may be determined (e.g., calculated) by using trigonometric operations. It can be further understood that the determined spread level (e.g., as a scalar quantity) and thus the corresponding arc obtained at this stage can be used for audio source positioning performed in the third stage described below
[0055] After determining the spread angle x 0 and the corresponding arc 513, one or more second audio sources 511 may be positioned on the arc 513 as shown in the example of FIG. 5(c). Note that the audio sources positioned on the arc 513 may be of equal volume (e.g., perceived as of equal magnitude) to the user 501 and may have the same reference distance. Optionally, the number (count) of the second audio sources 511 is based on the spread angle x 0It may be determined based on the spread angle x. For example, for small angles, only one second audio source (N=1) may be applied, and for large angles, two or more second audio sources (N>1) may be applied. That is, the number of second audio sources 511 is determined by the spread angle x 0 It can increase as increases. Alternatively, the number N of the second audio sources 511 does not depend on the user position 501 and / or relative point 502 (for example, the spread angle x). 0 It can be a predetermined constant (independent of the spread level 512 and the length of the arc 513).
[0056] Subsequently, spread level 512 is used to model the spread audio object 503, with a (relative) spread angle x 0and may be set / determined depending on the number N of second audio sources 511. In particular, the second audio sources 511 may be placed / positioned on the arc 513. In the case of two or more second audio sources (i.e., N>1), these available N audio sources 511 may be positioned on the arc 513 such that all second audio sources 511 are of equal volume to the user (i.e., perceived as equal in magnitude) and / or have the same reference distance calculated from a point on the arc 513 (e.g., the distance from the user's position to the relative point 502) for appropriate distance attenuation. For example, the second audio sources 511 may be distributed at equal intervals on the arc 513, i.e., they may be placed / positioned on the arc 513 such that adjacent second audio sources 511 are separated from each other by the same distance. In some embodiments where two or more second audio sources are considered, positioning may depend on the correlation level between the second audio sources 511 and / or the content creator's intent, as shown in the case of N=2 in the example of Figure 5(c). For example, each pair of adjacent second audio sources 511 may have different levels of decorrelation (e.g., their original recordings or copies, decorrelation filtering, etc.). The distance D2 between a pair of second audio signals 511 with a high (or higher) level of decorrelation may be greater than the distance D1 between a pair of second audio signals 511 with a low (or lower) level of decorrelation.
[0057] It should be further noted that one or more second audio sources 511 can be determined from the (original) first audio source, for example, by increasing the number of first audio sources. This can be achieved by duplicating one or more first audio sources and / or adding a weighted mixture of one or more first audio sources, and then applying a decorrelation process to the duplicated and / or added first audio sources. For example, when only one or a few first audio sources are recorded / captured for an extended audio object, the number of audio sources can be increased by duplicating one or a few first audio sources to determine the second audio source. Alternatively, in the case of multiple first audio sources, the second audio source can be determined by adding a weighted mixture of them. Subsequent application of a signal decorrelation process can be performed to obtain the final second audio source.
[0058] Figure 6 shows another schematic diagram of an example of determining a modified representation of a spread audio object (for example, as performed in Method 300) according to embodiments of the present disclosure. While Figure 5 shows a cuboid for representing the spatial spread of a spread audio object 503, the example in Figure 6 shows a case where the spread audio object 603 may have a complex geometric shape (for example, a vehicle). Similar to the embodiment in Figure 5, the user 601 is positioned in front of the spread audio object 603 oriented in three dimensions. However, in this exemplary embodiment, in order to simplify steps 301 and 303, a simplified spread representation 605, for example, showing an ellipsoid, may be obtained before applying step 301 of the proposed Method 300. In other words, for embodiments in which the spread object has a complex geometric shape, Method 300 may further include obtaining a simplified geometric shape (simplified spread geometry) 605 of the spread audio object to be used in determining the relative points, before or in order to obtain relative points on the geometric shape of the spread audio object. Therefore, a simplified vertical projection 606 of the spread audio object 603 may be obtained for the subsequent determination of the spread angle / level 604 described above.
[0059] Returning to Figure 5, the second audio source 511 may be positioned on the arc 513 such that it has the same reference distance from the user position 501. It will be understood that the reference distance specifies the distance at which the calculated attenuation of the audio source element is minimized, for example, to 0 dB (e.g., the user-source distance), regardless of the distance attenuation law used. For object sources with spread, such user-source distance may be measured relative to the origin of the spread (e.g., its "position" attribute) or relative to the spread itself, as shown in the examples of Figures 7(a) and 7(b), respectively. In the example of Figure 7(a), the reference distance D ref This is defined with respect to the origin 702a of the object spread 701a, while in the example of Figure 7(b), the reference distance D refThis is defined relative to the nearest point of the object spread 701b, as also shown by relative point 502 in the example in Figure 5. Therefore, referring to Figures 7(a) and 7(b), the relative point closest to the user's position lies on the dashed line. In the example in Figure 7(a), the relative point is a reference distance D from the origin 702a of the object spread 701a. ref It is located at the following point. In the example in Figure 7(b), the relative point is at the reference distance D from the object spread 701b. ref It is located at this point. Referring to Figure 5(b), since relative point 502 (for example, as the point closest to user 501) is also located on arc 513, a second audio source placed on arc 513 has the same reference distance as relative point 502, where attenuation can be minimized.
[0060] Figures 8(a)–(c) show examples of modified representations resulting from the spread audio object for different user positions as shown in Figure 2, according to embodiments of the present disclosure. Similar to the audio scene shown in Figure 2, users 803a and 803b are positioned in front of the object spread 801, but at different distances from it, and user 803c is positioned on one side of the object spread 801. It can be understood that any other position may also be included in the scene as a user position. Thus, the resulting spread levels 804a, 804b, and 804c represent the respective perceptions (e.g., spatial spread) of the object spread 801 at user positions 803a, 803b, and 803c. As shown in the example in Figure 8, the resulting spread level 804a at user position 803a is greater than the resulting spread level 804b at user position 803b. This also allows for a larger number of second audio sources 802a to be determined (and placed) in order to model the object spread 801. Similarly, the resulting spread level 804b at user position 803b is greater than the resulting spread level 804c at user position 803c, which also allows for a larger number of second audio sources 802b to be determined (and placed) in order to model the object spread 801. In this example, the modified representation for user position 803a includes five second audio sources 802a, the modified representation for user position 803b includes two second audio sources 802b, and the modified representation for user position 803c includes only one second audio source 802c.
[0061] interpretation Aspects of the systems described herein may be implemented in a suitable computer-based sound processing network environment (e.g., a server or cloud environment) for processing digital or digitized audio files. The components of an adaptive audio system may include one or more networks containing any desired number of individual machines, including one or more routers (not shown) that buffer and route data transmitted between computers. Such networks may be built on various different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0062] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls the execution of the system's processor-based computing device. It should also be noted that the various functions disclosed herein may be described, with respect to their behavior, register transfers, logical components, and / or other characteristics, as data and / or instructions embodied using any number of combinations of hardware, firmware, and / or in various machine-readable or computer-readable media. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, various forms of physical (non-temporary) non-volatile storage media, such as optical storage media, magnetic storage media, or semiconductor storage media.
[0063] Specifically, embodiments may include hardware, software, and electronic components or modules, and it should be understood that, for the sake of discussion, the majority of the components may be illustrated and described as if they were implemented solely in hardware. However, those skilled in the art will recognize, based on reading this detailed description, that in at least one embodiment, the electronic-based aspects may be implemented in software (for example, stored on a non-temporary computer-readable medium) that can be executed by one or more electronic processors, such as microprocessors and / or application-specific integrated circuits ("ASICs"). Thus, it should be noted that multiple hardware and software-based devices, as well as multiple different structural components, may be utilized to implement the embodiments. For example, a "content activity detector" described herein may include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (for example, a system bus) connecting the various components.
[0064] While one or more implementations have been described, as examples, with respect to specific embodiments, it should be understood that one or more implementations are not limited to the disclosed embodiments. Rather, they are intended to cover a variety of modifications and similar configurations that would be obvious to those skilled in the art. Accordingly, the appended claims should be given the broadest possible interpretation to encompass all such modifications and similar configurations.
[0065] Furthermore, it should be understood that the expressions and terms used herein are for illustrative purposes only and should not be considered limiting. The use of “includes,” “equips,” or “has,” and their variations, is intended to encompass the enumerated items and their equivalents, as well as any additional items. Unless otherwise specified or limited, the terms “attached,” “connected,” “supported,” and “joined,” and their variations, are used broadly to encompass both direct and indirect attachment, connection, support, and joining. Several aspects are described below. [Aspect 1] A computer-implemented method for modeling expansive audio objects for audio rendering in a virtual or augmented reality environment, the method being: Steps include obtaining a spread representation showing the geometric form of a spread audio object, and information relating to one or more first audio sources associated with the spread audio object; The steps include: obtaining the relative point closest to the user's position in the virtual or augmented reality environment using the spread representation that shows the geometric form of the spread audio object; A step of determining the spread parameter of the spread expression based on the user position and the relative point, wherein the spread parameter describes the spatial spread of the spread audio object as perceived at the user position; The steps include determining the position of one or more second audio sources relative to the user's position for modeling the aforementioned expansive audio object; A step of outputting a modified representation of the spread audio object for modeling the spread audio object, wherein the modified representation includes the spread parameter and the position of one or more second audio sources. The method of computer implementation. [Aspect 2] A computer-implemented method according to Embodiment 1, further comprising the step of rendering the spread audio object based on the modified representation of the spread audio object, wherein the spread audio object is rendered using the determined positions of the one or more second audio sources and the spread parameters. [Aspect 3] The computer-implemented method according to Embodiment 2, wherein the rendering includes 6DoF audio rendering, and further includes the step of obtaining the user position, the position and / or orientation and geometry of the spread audio object for the rendering. [Aspect 4] A computer-implemented method according to embodiment 1 or 2, further comprising the step of determining one or more second audio sources for modeling the expansive audio object based on one or more first audio sources. [Aspect 5] The computer-implemented method according to any one of embodiments 1 to 4, wherein the spread parameter is further determined based on the position and / or orientation of the spread audio object. [Aspect 6] A computer-implemented method according to embodiment 5, further comprising the step of determining a relative spread angle based on the user position, the relative point, and the position and / or orientation of the spread audio object, wherein the spread parameter is determined based on the relative spread angle. [Aspect 7] Determining the location of one or more second audio sources is: An arc is determined based on the user position, the relative point, and the geometric shape of the spreading audio object; This includes positioning one or more of the determined second audio sources on the arc, A computer-implemented method as described in any one of the embodiments 1 to 6. [Aspect 8] The computer-implemented method according to embodiment 7, wherein the positioning includes distributing all the second audio sources at equal intervals along the arc. [Aspect 9] The computer-implemented method according to embodiment 7 or 8, wherein the positioning depends on the correlation level between the second audio sources and / or the intent of the content creator. [Aspect 10] The computer-implemented method according to any one of embodiments 1 to 9, as referenced to embodiment 4, wherein the spread parameter is further determined based on the number of one or more second audio sources determined. [Aspect 11] The computer-implemented method according to embodiment 10, wherein the number of one or more second audio sources determined is a predetermined constant independent of the user position and / or the relative point. [Aspect 12] A computer-implemented method according to Embodiment 10, as referenced to Embodiment 4, which includes determining the one or more second audio sources for modeling the spread audio object, by determining the number of the one or more second audio sources based on the relative spread angle. [Aspect 13] The computer-implemented method according to embodiment 12, wherein the number of the one or more second audio sources increases as the relative spread angle increases. [Aspect 14] Determining the one or more second audio sources for modeling the aforementioned expansive audio object is: To duplicate one or more of the first audio sources, or to add a weighted mix of the one or more first audio sources; A computer-implemented method according to any one of embodiments 1 to 13, as referenced to embodiment 4, further comprising applying a decorrelation process to a duplicated or added first audio source. [Aspect 15] The computer-implemented method according to any one of embodiments 1 to 14, wherein the spread representation indicates a two-dimensional or three-dimensional geometric form for representing the spatial spread of the spread audio object. [Aspect 16] The computer-implemented method according to any one of embodiments 1 to 15, wherein the expansive audio object (3) is oriented in two or three dimensions. [Aspect 17] The computer-implemented method according to any one of embodiments 1 to 16, wherein the spatial extent of the expansive audio object as perceived at the user's position is described as the perceived width, size, and / or weight of the expansive audio object. [Aspect 18] The steps include obtaining a vertical projection of the extended audio object on a projection plane perpendicular to a first line connecting the user position and the relative point; The further step includes determining a plurality of boundary points on the vertical projection that identify the projection size of the extended audio object, The relative spread angle is determined using the user position and the plurality of boundary points. A computer-implemented method as described in any one of the embodiments 1 to 17 when referring to embodiment 6. [Aspect 19] The computer-implemented method according to embodiment 18, wherein determining the plurality of boundary points includes obtaining a second line relating to the horizontal projection size of the spread audio object, the plurality of boundary points including the leftmost and rightmost boundary points of the vertical projection on the second line. [Aspect 20] The computer-implemented method according to embodiment 19, wherein the horizontal projection size is the maximum size of the spreading audio object. [Aspect 21] The computer-implemented method according to any one of embodiments 18 to 20, wherein the vertical projection includes a simplified projection of the expansive audio object having a complex geometric shape. [Aspect 22] A computer-implemented method according to any one of embodiments 1 to 21, further comprising using the spread parameter to control the perceived size of the spread audio object. [Aspect 23] The computer-implemented method according to aspect 23, wherein the spreading audio object is modeled as a point source or a wide source by controlling the perceived size of the spreading audio object. [Aspect 24] The computer-implemented method according to any one of embodiments 1 to 23, wherein the positions of the one or more second audio sources are determined such that all of the second audio sources are at the same reference distance from the user position. [Aspect 25] A computer-implemented method according to any one of embodiments 1 to 24, wherein obtaining the relative point using the spread representation that shows the geometric form of the spread audio object includes obtaining the relative point on the geometric form or obtaining the relative point at a certain distance from the spread representation. [Aspect 26] An apparatus for modeling expansive audio objects for audio rendering in a virtual or augmented reality environment, the apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is configured to perform all steps of the computer-implemented method described in any one of embodiments 1 to 25. [Aspect 27] A system for implementing audio rendering in a virtual or augmented reality environment, wherein the system is: An apparatus for modeling expansive audio objects for audio rendering in a virtual or augmented reality environment, the apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is configured to perform all steps of the computer-implemented method described in any one of embodiments 1 to 25; The device comprises a spread modeling unit, the spread modeling unit is configured to receive information from the device regarding the modified representation of the spread audio object, and to control the spread size of the spread audio object based on the information regarding the modified representation. The system is configured to transmit the information relating to the modified representation of the spread audio object and / or the controlled spread size to the audio output. system. [Aspect 28] The system according to aspect 27, wherein the system is a user virtual reality console or a part thereof. [Aspect 29] The system according to embodiment 27 or 28, wherein the system is configured to transmit the information relating to the modified representation of the spread audio object and / or the controlled spread size to the audio output. [Aspect 30] A computer program having instructions that, when executed by a computing device, cause the computing device to perform all steps of the method according to any one of embodiments 1 to 25. [Aspect 31] A computer-readable storage medium storing the computer program described in aspect 30.
[0066] Bulleted Exemplary Embodiments Various aspects and implementations of this disclosure can also be understood from the following enumerated example embodiments (EEEs) that are not claims. [EEE1] A method for modeling a spread audio object for audio rendering in a virtual reality or augmented reality environment, comprising: obtaining a spread representation showing the geometric form of the spread audio object and information relating to one or more first audio sources associated with the spread audio object; obtaining relative points on the geometric form of the spread audio object based on a user position in the virtual or augmented reality environment; determining a spread parameter of the spread representation based on the user position and the relative point, wherein the spread parameter describes the spatial spread of the spread audio object as perceived at the user position; determining the positions of one or more second audio sources for modeling the spread audio object relative to the user position; and outputting a modified representation of the spread audio object for modeling the spread audio object, wherein the modified representation includes the spread parameter and the positions of the one or more second audio sources. [EEE2] The method of EEE1, further comprising the step of determining one or more second audio sources for modeling the expansive audio object based on one or more first audio sources. [EEE3] The method according to EEE1 or 2, wherein the spread parameter is further determined based on the position and / or orientation of the spread audio object. [EEE4] The method of EEE3, further comprising the step of determining a relative spread angle based on the user position, the relative point, and the position and / or orientation of the spread audio object, wherein the spread parameter is determined based on the relative spread angle. [EEE5] The method according to any one of EEE1 to 4, wherein determining the location of the one or more second audio sources includes determining an arc based on the user position, the relative point, and the geometric shape of the spreading audio object; and positioning the determined one or more second audio sources on the arc. [EEE6] The method according to EEE5, wherein the positioning includes distributing all of the second audio sources at equal intervals along the arc. [EEE7] The positioning according to the method of EEE5 or 6, wherein the positioning depends on the correlation level between the second audio sources and / or the intent of the content creator. [EEE8] The method according to any one of EEE2 to 7, wherein the spread parameter is further determined based on the number of one or more second audio sources determined. [EEE9] The method according to EEE8, wherein the number of one or more second audio sources determined is a predetermined constant independent of the user position and / or the relative point. [EEE10] The method according to EEE8, as referenced to EEE4, wherein determining the one or more second audio sources for modeling the spread audio object includes determining the number of the one or more second audio sources based on the relative spread angle. [EEE11] The method according to EEE10, wherein the number of the one or more second audio sources increases as the relative spread angle increases. [EEE12] The method according to any one of EEE2 to 11, wherein determining the one or more second audio sources for modeling the spread audio object further comprises duplicating the one or more first audio sources or adding a weighted mixture of the one or more first audio sources; and applying a decorrelation process to the duplicated or added first audio sources. [EEE13] The method according to any one of EEE1 to 12, wherein the spread representation indicates a two-dimensional or three-dimensional geometric form for representing the spatial spread of the audio object having a spread. [EEE14] The method according to any one of EEE1 to 13, wherein the aforementioned expansive audio object (3) is oriented in two or three dimensions. [EEE15] The spatial extent of the expansive audio object as perceived at the user's position is described as the perceived width, size, and / or weight of the expansive audio object, according to any one of EEE1 to 14. [EEE16] The method according to any one of EEE1 to 15, wherein the relative point is the point on the geometric form of the spreading audio object that is closest to the user's position. [EEE17] The method according to any one of EEE4 to 16, further comprising: obtaining a vertical projection of the spread audio object on a projection plane perpendicular to a first line connecting the user position and the relative point; and determining a plurality of boundary points on the vertical projection that identify the projection size of the spread audio object, wherein the relative spread angle is determined using the user position and the plurality of boundary points. [EEE18] The method according to EEE17, wherein determining the plurality of boundary points includes obtaining a second line relating to the horizontal projection size of the spread audio object, the plurality of boundary points including the leftmost and rightmost boundary points of the vertical projection on the second line. [EEE19] The method according to EEE18, wherein the horizontal projection size is the maximum size of the extended audio object. [EEE20] The method according to any one of EEE17 to 19, wherein the vertical projection includes a simplified projection of the expansive audio object having a complex geometric shape. [EEE21] The method according to EEE20, further comprising the step of obtaining a simplified geometric form of the spreading audio object for use in determining the relative point, before obtaining the relative point on the geometric form of the spreading audio object. [EEE22] The method according to any one of EEE1 to 21, further comprising the step of rendering the spread audio object based on the modified representation of the spread audio object, wherein the spread audio object is rendered using the determined positions of the one or more second audio sources and the spread parameters. [EEE23] The method according to EEE22, wherein the rendering includes 6DoF audio rendering, and further includes the step of obtaining the user position, the position and / or orientation and geometry of the spread audio object for the rendering. [EEE24] The method according to any one of EEE1 to 23, further comprising using the spread parameter to control the perceived size of the spread audio object. [EEE25] The method according to EEE24, wherein the spreading audio object is modeled as a point source or a wide source by controlling the perceived size of the spreading audio object. [EEE26] The method according to any one of EEE1 to 25, wherein the positions of the one or more second audio sources are determined such that all of the second audio sources are at the same reference distance from the user position. [EEE27] An apparatus for modeling expansive audio objects for audio rendering in a virtual or augmented reality environment, the apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is configured to perform all steps of the method described in any one of EEE1 to 26. [EEE28] A system for implementing audio rendering in a virtual or augmented reality environment, the system comprising: an apparatus described in EEE27; and a spread modeling unit, the spread modeling unit being configured to receive information from the apparatus regarding the modified representation of the spread audio object and to control the spread size of the spread audio object based on the information regarding the modified representation. [EEE29] The system in question is a user virtual reality console, or a part thereof, as described in EEE28. [EEE30] The system according to EEE28 or 29, wherein the system is configured to transmit the information relating to the modified representation of the spread audio object and / or the controlled spread size to the audio output. [EEE31] A computer program having instructions that, when executed by a computing device, cause the computing device to perform all the steps of the method described in any one of EEE1 to 26. [EEE32] A computer-readable storage medium that stores computer programs as described in EEE31.
Claims
1. A computer-implemented method for modeling expansive audio objects for audio rendering in a virtual or augmented reality environment, the method being: The steps include obtaining a spread representation showing the geometric form of a spread audio object, and information relating to one or more first audio sources associated with the spread audio object; A step of obtaining the relative point closest to the user's position in the virtual or augmented reality environment using the spread representation showing the geometric form of the spread audio object, wherein obtaining the relative point using the spread representation showing the geometric form of the spread audio object includes obtaining the relative point on the geometric form; A step of determining a spread parameter of the spread representation based on the user position and the relative point, wherein the spread parameter describes the spatial spread of the spread audio object as perceived at the user position; The steps include determining the position of one or more second audio sources relative to the user's position for modeling the aforementioned expansive audio object; A step of outputting a modified representation of the spread audio object for modeling the spread audio object, wherein the modified representation includes the spread parameter and the positions of one or more second audio sources. The method of computer implementation.
2. The computer-implemented method according to claim 1, further comprising the step of rendering the spread audio object based on the modified representation of the spread audio object, wherein the spread audio object is rendered using the determined positions of the one or more second audio sources and the spread parameters.
3. The computer-implemented method according to claim 2, wherein the rendering includes 6DoF audio rendering, and further includes the step of obtaining the user position, the position and / or orientation and geometry of the spread audio object for the rendering.
4. The computer-implemented method according to claim 1, further comprising the step of determining one or more second audio sources for modeling the expansive audio object based on one or more first audio sources.
5. The computer-implemented method according to claim 1, wherein the spread parameter is further determined based on the position and / or orientation of the spread audio object.
6. The computer-implemented method according to claim 5, further comprising the step of determining a relative spread angle based on the user position, the relative point, and the position and / or orientation of the spread audio object, wherein the spread parameter is determined based on the relative spread angle.
7. Determining the location of the one or more second audio sources is: An arc is determined based on the user position, the relative point, and the geometric shape of the spreading audio object; This includes positioning one or more of the determined second audio sources on the arc, The computer-implemented method according to claim 1.
8. The computer-implemented method according to claim 7, wherein the positioning includes distributing all of the second audio sources at equal intervals along the arc.
9. The computer-implemented method according to claim 7, wherein the positioning depends on the correlation level between the second audio sources and / or the intent of the content creator.
10. The computer-implemented method according to claim 4, wherein the spread parameter is further determined based on the number of one or more second audio sources to be determined.
11. The computer-implemented method according to claim 10, wherein the number of one or more second audio sources determined is a predetermined constant independent of the user position and / or the relative point.
12. The method is: The process further includes determining a relative spread angle based on the user position, the relative point, and the position and / or orientation of the spread audio object, The computer-implemented method according to claim 4, wherein determining the one or more second audio sources for modeling the spread audio object includes determining the number of the one or more second audio sources based on the relative spread angle.
13. The computer-implemented method according to claim 12, wherein the number of the one or more second audio sources increases as the relative spread angle increases.
14. Determining the one or more second audio sources for modeling the aforementioned spread audio object is: To duplicate one or more of the first audio sources, or to add a weighted mix of the one or more first audio sources; The computer-implemented method according to claim 4, further comprising applying a decorrelation process to a duplicated or added first audio source.
15. The computer-implemented method according to claim 1, wherein the spread representation indicates a two-dimensional or three-dimensional geometric form for representing the spatial spread of the spread audio object.
16. The computer-implemented method according to claim 1, wherein the expansive audio object is oriented in two or three dimensions.
17. The computer-implemented method according to claim 1, wherein the spatial extent of the expansive audio object as perceived at the user's position is described as the perceived width, size, and / or weight of the expansive audio object.
18. The steps include obtaining a projection of the extended audio object on a projection plane perpendicular to a first line connecting the user position and the relative point; The process further includes the step of determining a plurality of boundary points on the projection plane that identify the projection size of the extended audio object, The relative spread angle is determined using the user position and the plurality of boundary points. The computer-implemented method according to claim 6.
19. The computer-implemented method according to claim 18, wherein determining the plurality of boundary points includes obtaining a second line relating to the horizontal size of the spread audio object on the projection plane, the plurality of boundary points including the leftmost and rightmost boundary points of the projection on the second line.
20. The computer-implemented method according to claim 19, wherein the horizontal size is the maximum size of the spreading audio object.
21. The computer-implemented method according to claim 18, wherein the projection includes a simplified projection of the expansive audio object having a complex geometric shape.
22. The computer-implemented method according to claim 1, further comprising using the spread parameter to control the perceived size of the spread audio object.
23. The computer-implemented method according to claim 22, wherein the spreading audio object is modeled as a point source or a wide source by controlling the perceived size of the spreading audio object.
24. The computer-implemented method according to claim 1, wherein the positions of the one or more second audio sources are determined such that all of the second audio sources are at the same reference distance from the user position.
25. An apparatus for modeling expansive audio objects for audio rendering in a virtual or augmented reality environment, the apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is configured to perform all steps of the computer-implemented method described in claim 1.
26. A system for implementing audio rendering in a virtual or augmented reality environment, wherein the system: An apparatus for modeling expansive audio objects for audio rendering in a virtual or augmented reality environment, the apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is configured to perform all steps of the computer-implemented method described in claim 1; The device comprises a spread modeling unit, the spread modeling unit is configured to receive information from the device regarding the modified representation of the spread audio object, and to control the spread size of the spread audio object based on the information regarding the modified representation. The system is configured to transmit the information relating to the modified representation of the spread audio object and / or the controlled spread size to the audio output. system.