Rendering audio
By configuring speaker placement and the apparent position mapping of audio objects in multi-user scenarios, the problem of poor matching between audio and visual objects in speaker-rendered audio is solved, thus improving the user's immersive experience.
Patent Information
- Application Number
- CN202510491095.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-23
- Filing Date
- 2025-04-18
- Publication Date
- 2025-10-24
Smart Images

Figure CN120835264A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to rendering audio, in particular to rendering audio at loudspeaker(s). BACKGROUND
[0002] It is known to render audio at a loudspeaker. There remains a need for improvements in rendering audio at a loudspeaker in a multi-user setting. SUMMARY
[0003] In a first aspect, the present disclosure provides an apparatus comprising: means for rendering audio of a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the means for rendering comprises: means for determining a number of users that are using the three-dimensional scene; and means for, based on determining that more than one user is using the three-dimensional scene, configuring an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
[0004] In some examples, configuring the apparent position comprises: rendering audio corresponding to the first audio object from the first loudspeaker.
[0005] Some examples comprise: means for determining whether the three-dimensional scene comprises a second audio object in addition to the first audio object; wherein the means for rendering further comprises: means for, based on determining that the three-dimensional scene comprises at least the first audio object and the second audio object, configuring an apparent position of the second audio object to be mapped to a position of a second loudspeaker of the plurality of loudspeakers.
[0006] In some examples, the means for configuring the apparent positions of the first and second audio objects to be mapped to the positions of the first and second loudspeakers comprises: means for performing one or more of translating, rotating and / or scaling the three-dimensional scene. In some examples, the translating comprises: translating one or more audio objects within the three-dimensional scene to be mapped to the first and second loudspeakers.
[0007] Some examples include means for determining whether the three-dimensional scene includes more than a threshold number of audio objects; wherein, if the three-dimensional scene includes more than the threshold number of audio objects, the means for rendering includes means for configuring an apparent position of at least one audio object to be mapped outside of the first loudspeaker arrangement. In some examples, the threshold number is two. In some examples, the apparent position of each of the at least two audio objects is configured to be mapped to a position of each of the at least two loudspeakers, respectively. In some examples, rendering audio corresponding to the audio object to be mapped outside of the first loudspeaker arrangement includes rendering the audio using a spatial range.
[0008] In some examples, the means for rendering includes means for rendering the audio independent of changes in the position of the more than one user.
[0009] Some examples include means for configuring one or more of a position, a direction, and / or a size of a visual representation of any audio object based at least in part on an apparent position of the respective audio object.
[0010] In some examples, the means for rendering further includes means for rendering the audio from one or more of the plurality of loudspeakers based at least in part on an initial position of the first audio object and a position of the first user based on a determination that a single first user is using the three-dimensional scene. In some examples, the means for rendering further includes means for changing a direction and / or a volume of the audio based at least in part on changes in the position of the first user.
[0011] In some examples, the three-dimensional scene is one or more of a virtual reality scene, an augmented reality scene, or a mixed reality scene.
[0012] In some examples, the first loudspeaker is selected from the plurality of loudspeakers based on which loudspeaker of the plurality of loudspeakers has a minimum distance to the first audio object.
[0013] In some examples, the distance involves one or more of a horizontal distance, a vertical distance, or an angular distance.
[0014] The means can include at least one processor; and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform.
[0015] In a second aspect, the disclosure describes a method comprising: rendering audio of a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene includes a first audio object, wherein the rendering includes: determining a number of users that are using the three-dimensional scene; and based on determining that more than one user is using the three-dimensional scene, configuring an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
[0016] In some examples, configuring the apparent position includes: rendering audio corresponding to the first audio object from the first loudspeaker.
[0017] Some examples include: determining whether the three-dimensional scene includes a second audio object in addition to the first audio object; wherein the means for rendering further includes: means for, based on determining that the three-dimensional scene includes at least the first audio object and the second audio object, configuring an apparent position of the second audio object to be mapped to a position of a second loudspeaker of the plurality of loudspeakers.
[0018] In some examples, configuring the apparent positions of the first and second audio objects to be mapped to the positions of the first and second loudspeakers includes: performing one or more of shifting, rotating, and / or scaling the three-dimensional scene. In some examples, the shifting includes: shifting one or more audio objects within the three-dimensional scene to be mapped to the first and second loudspeakers.
[0019] Some examples include: determining whether the three-dimensional scene includes more than a threshold number of audio objects; wherein, if the three-dimensional scene includes more than the threshold number of audio objects, the rendering includes: means for configuring an apparent position of at least one audio object to be mapped to outside of the first loudspeaker arrangement. In some examples, the threshold number is two. In some examples, the apparent position of each of the at least two audio objects is configured to be mapped to a position of each of the at least two loudspeakers, respectively. In some examples, rendering audio corresponding to the audio object to be mapped to outside of the first loudspeaker arrangement includes: rendering the audio using a spatial range.
[0020] In some examples, the means for rendering includes: means for rendering the audio independent of changes in positions of the more than one user.
[0021] Some examples include: means for configuring one or more of a position, a direction, and / or a size of a visual representation of any audio object to be based at least in part on an apparent position of the respective audio object.
[0022] In some examples, the means for rendering further includes means for rendering the audio from one or more of the plurality of speakers based at least in part on a determination that a single first user is using the three-dimensional scene, based on an initial position of the first audio object and a position of the first user. In some examples, the means for rendering further includes means for changing a direction and / or a volume of the audio based at least in part on a change in the position of the first user.
[0023] In some examples, the three-dimensional scene is one or more of a virtual reality scene, an augmented reality scene, or a mixed reality scene.
[0024] In some examples, the first speaker is selected from the plurality of speakers based on which speaker of the plurality of speakers has a minimum distance to the first audio object.
[0025] In some examples, the distance involves one or more of a horizontal distance, a vertical distance, or an angular distance.
[0026] In a third aspect, the disclosure describes an apparatus configured to perform any of the methods as described with reference to the second aspect.
[0027] In a fourth aspect, the disclosure describes computer-readable instructions that, when executed by a computing apparatus, cause the computing apparatus to perform any of the methods as described with reference to the second aspect.
[0028] In a fifth aspect, the disclosure describes a computer program comprising instructions for causing an apparatus to perform at least the following: rendering audio of a three-dimensional scene at one or more of a plurality of speakers, wherein the plurality of speakers are arranged in a first speaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users using the three-dimensional scene; and based on a determination that more than one user is using the three-dimensional scene, configuring an apparent position of the first audio object to be mapped to a position of a first speaker of the plurality of speakers.
[0029] In a sixth aspect, the disclosure describes a computer-readable medium (e.g., a non-transitory computer-readable medium) comprising program instructions stored thereon for at least causing an apparatus to perform the following: rendering audio of a three-dimensional scene at one or more of a plurality of speakers, wherein the plurality of speakers are arranged in a first speaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users using the three-dimensional scene; and based on a determination that more than one user is using the three-dimensional scene, configuring an apparent position of the first audio object to be mapped to a position of a first speaker of the plurality of speakers.
[0030] In a seventh aspect, the disclosure describes an apparatus comprising: at least one processor; and at least one memory including computer program code, the computer program code, when executed by the at least one processor, causes the apparatus to: render audio of a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users that are using the three-dimensional scene; and based on determining that more than one user is using the three-dimensional scene, configuring an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
[0031] In an eighth aspect, the disclosure describes an apparatus comprising: a first module configured to: render audio of a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining, by a second module, a number of users that are using the three-dimensional scene; and based on determining, by a third module, that more than one user is using the three-dimensional scene, configuring an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers. BRIEF DESCRIPTION OF DRAWINGS
[0032] Example embodiments will now be described, by way of example only, with reference to the following schematic drawings, in which:
[0033] Figures 1 to 5 is a block diagram of an example system;
[0034] Figure 6 is a flow diagram of an algorithm according to an example embodiment;
[0035] Figure 7 is a block diagram of a system according to an example embodiment;
[0036] Figure 8 is a flow diagram of an algorithm according to an example embodiment;
[0037] Figure 9 is a block diagram of a system according to an example embodiment;
[0038] Figure 10 is a flow diagram of an algorithm according to an example embodiment;
[0039] Figure 11 is a block diagram of a system according to an example embodiment;
[0040] Figure 12 is a block diagram of components of a system according to an example embodiment; and
[0041] Figure 13 An example of a tangible medium for storing computer readable code, which when run by a computer, can perform the method according to the example embodiments described above is shown. DETAILED DESCRIPTION
[0042] The scope of protection sought for various embodiments of the present application is defined by the independent claims. The embodiments and features that are not contained in the independent claims, if any, are to be interpreted as examples useful in understanding various embodiments of the present application.
[0043] In the description and drawings, like numbers refer to like elements throughout.
[0044] Figure 1 is a block diagram of an example system (generally indicated by reference numeral 10). The system 10 shows a three-dimensional scene, such as a virtual reality scene, an augmented reality scene, and / or a mixed reality scene, in which the system 10 includes a plurality of loudspeakers 11a-11e. For example, the loudspeakers 11 can be arranged to provide six degrees of freedom (6DoF) audio to a user(s) experiencing a scene, such as a virtual reality, augmented reality, and / or mixed reality scene (herein referred to as a VR scene). The audio can be processed (e.g., as with MPEG-I processing) using the plurality of loudspeakers 11a-11e for being rendered as 6DoF audio (e.g., instead of being rendered with headphones / other in-ear devices).
[0045] The system 10 also includes a user 12 (the user 12 is shown at a location 12a) and an audio object 13a. The VR scene can be defined such that the audio object 13a corresponds to a visual object shown to the user 12 at the location of the audio object 13a, where the audio object is located within the periphery of the speaker arrangement of the speakers 11a-11e. An apparent location 13b of the audio object 13a can be configured based at least in part on the location 12a of the user 12 and the positioning of one or more of the plurality of speakers 11a-11e. For example, the line 15 represents the direction of the audio object 13a relative to the user location 12a, and thus the apparent location 13b is configured to be placed between the speakers 11a and 11b (represented by the line 14) along the line 15. In other words, the speakers 11a and 11b can render audio such that the apparent location 13b of the audio object 13a is configured to move along the line 15. Thus, the user 12 can perceive the audio to be generated from the direction of the audio object 13a. In one example, it is determined which speakers (e.g., speaker pairs) can be selected for allowing the apparent location 13b to be behind the location of the audio object 13a, and thus, on this basis, the speakers 11a and 11b are selected for rendering the audio. The audio can be panned based on movement of the user 12 such that the apparent location 13b can move along the line 14 in order to perceive the audio rendered from the direction of the audio object 13a.
[0046] Figure 2 is a block diagram of an example system (generally indicated by reference numeral 20). The system 20 shows Figure 1 a VR scene of the user 12 moving from the location 12a to the location 12b, and the line 16 represents the direction of the audio object 13a relative to the user location 12b. Thus, based on the updated location 12b of the user 12, the audio object 13a has an updated apparent location 13c along the line 16. The apparent location 13c is configured to be placed between the speakers 11a and 11e (represented by the line 17) along the line 16. In other words, the speakers 11a and 11e can render audio such that the apparent location 13c of the audio object 13a is configured to move along the line 16. Thus, the user 12 can perceive the audio to be generated from the direction of the audio object 13a.
[0047] Figure 3 is a block diagram of an example system (generally indicated by reference numeral 30). The system 30 shows Figure 1, wherein the position of user 12 moves from position 12a to position 12c, and line 15 represents the direction of audio object 13a relative to the user's position 12c. Thus, user 12 is shown moving closer to audio object 13a along line 15. As user 12 moves closer to audio object 13a, the gains of speakers 11a and 11b may be adjusted to simulate a distance gain effect. For example, audio may be emitted by speakers 11a and 11b at position 13d at a higher frequency than at position 13d. Figure 1 Rendering is done with the larger volume shown in (indicated by the larger circle).
[0048] Figure 4 is a block diagram of an example system generally indicated by reference numeral 40. System 40 illustrates a system similar to Figure 1 VR scene of , in which speakers 11a-11e are arranged in a similar manner. System 40 also shows audio object 41a so that the direction of audio object 41a relative to user 12 (shown by line 42) matches the direction of speaker 11a relative to user 12. Therefore, apparent position 41b is placed behind audio object 41a along line 42 so that apparent position 41b can fall at the position of speaker 11a. Therefore, the audio of audio object 41a can be rendered from a single speaker 11a. In some examples, if the audio object is placed at the same position as the speaker, the apparent position of the audio object does not need to change based on the user's movement, because the audio from the audio object can appear to be generated from the direction of the corresponding speaker regardless of the change in the user's position. However, the gain of the speaker can be adjusted to simulate a distance gain effect.
[0049] Figure 5 is a block diagram of an example system (generally indicated by reference numeral 50). System 50 illustrates a system similar to Figure 1 , but with another user 51 and audio object 52a. In this scenario, it may not be feasible to render the audio of audio object 52a based on the positions of both user 12 and user 51 because the two users are at different positions and they will perceive the audio from different relative directions. For example, if the position of user 12 is considered, line 54 represents the direction of audio object 52a relative to user 12. As shown in FIG. Figure 1As described, the apparent position 52b of the audio object 52a is configured to appear behind the audio object 52a along the line 54, thus, the audio is rendered by the speakers 11a and 11b and the audio can translate along the line 55 based on the movement of the user 12. However, if the apparent position of the audio object is at the apparent position 52b, the user 51 can perceive the audio object 52a to be placed in an incorrect position. The line 56 represents the direction of the apparent position 52b relative to the user 51. For example, when the audio of the audio object 52a appears to be generated from the apparent position 52b, the user 51 can perceive the audio object to be located at, for example, the position 53 along the line 56. This can result in the user 51 perceiving a mismatch between the audio direction and the position corresponding to the visual representation of the audio object 52a (e.g., a visual object placed at the same position as the audio object 52a). This mismatch can be more pronounced if the user 51 is stationary while the user 12 moves, as the movement of the user 12 can also cause the apparent position 52b to move, thus, causing the audio object position 53 to move back and forth for the user 51 even though the user 51 is not moving. This mismatch can not be desirable and can interfere with the immersive experience of the user 51.
[0050] The example embodiments described below can aim to solve the problems arising from Figure 5 the situations described in
[0051] Figure 6 is a flowchart of an algorithm (generally indicated by reference numeral 60) in accordance with an example embodiment. For a better understanding, Figure 6 can be seen in Figure 7
[0052] Figure 7 is a block diagram of a system (generally indicated by reference numeral 70) in accordance with an example embodiment. The system 70 shows a scene 71 with a single user 12 and a scene 72 with multiple users (users 12 and 78). The scene 71 can be similar to the scene as described with reference to Figure 1
[0053] The algorithm 60 describes a method for rendering audio of a three-dimensional scene (e.g., a VR scene) at one or more of a plurality of speakers 11a-11b. The plurality of speakers are arranged in a first speaker arrangement (e.g., shown by the speakers 11a-11e) suitable for providing spatial audio corresponding to the three-dimensional scene. The three-dimensional scene can include a first audio object 73a.
[0054] The algorithm 60 can start at operation 601, determining a number of users that are using the three-dimensional scene.
[0055] If it is determined at operation 602 that more than one user (e.g., user 12 and user 78) is using the scene (e.g., scene 72), the algorithm moves to operation 604 and the apparent position 76 of the first audio object 73a is configured to be mapped (e.g., shown by arrow 77) to the location of the first loudspeaker (e.g., loudspeaker 11a of the plurality of loudspeakers). Configuring the apparent position can include rendering audio corresponding to the first audio object 73a from loudspeaker 11a independent of changes in the location of user 12 and / or 78. The rendering can include rendering the audio independent of changes in the location of user 12 and / or 78.
[0056] If it is determined at operation 602 that a single user (e.g., user 12) is using the scene, the algorithm can move to operation 603 and the apparent position (e.g., 73b) of the first audio object 73a is configured to be determined based on the initial position of the audio object (e.g., the position of audio object 73a) and / or the relative position of the user (e.g., the position of user 12). For example, the audio can be rendered from one or more of the plurality of loudspeakers based at least in part on the initial position of the first audio object 73a and the position of the first user 12, and optionally based on the positioning of one or more of the plurality of loudspeakers 11a-11e. For example, line 74 represents the direction of audio object 73a relative to the position of user 12, and thus, the apparent position 73b is configured to be placed between loudspeakers 11a and 11b along line 74 (represented by line 75). In other words, loudspeakers 11a and 11b can render audio such that the apparent position 73b of audio object 73a is configured to move along line 75. Thus, user 12 can perceive that the audio is generated from the direction of audio object 73a. This can be similar to the audio rendering described with reference to Figure 1 In one example, rendering the audio in the scene can further include changing the direction and / or volume of the audio based at least in part on changes in the position of the first user 12 (e.g., as described with reference to Figure 2 and Figure 3 In one example, rendering the audio in the scene can further include changing the direction and / or volume of the audio based at least in part on changes in the position of the first user 12 (e.g., as described with reference to
[0057] Moving to scene 72, since apparent position 76 is configured to be mapped to the location of loudspeaker 11a, both users 12 and 78 can perceive in a similar manner that the audio is rendered from the location of loudspeaker 11a, thus preventing the case of audio and visual mismatch as described with reference to Figure 5 Further, apparent position 76 can remain unchanged even if one of the users moves, or the users move in different directions, since apparent position 76 is not determined according to the position of the users.
[0058] In one example, one or more of the position, direction, and / or size of the visual representation of the audio object 73a can be configured to be based at least in part on the apparent position 76 of the respective audio object 73a. Thus, the visual representation of the audio object 73a is further configured to appear at the apparent position 76 in order to improve consistency between the positions of the audio and visual objects.
[0059] In one example, when a single user 12 is using the scene 71 and a second user 78 joins, the apparent position 73b can gradually change to the apparent position 76 so that the user 12 can perceive the change as gradual rather than a sudden change.
[0060] In example embodiments, a first loudspeaker (e.g., loudspeaker 11a) can be selected based at least in part on which of the plurality of loudspeakers has a minimum distance to the first audio object. For example, loudspeaker 11a is selected from the plurality of loudspeakers because loudspeaker 11a can be closest to the audio object 73a. The distance can relate to one or more of: a horizontal distance between the loudspeaker(s) and the audio object 73a, a vertical distance between the loudspeaker(s) and the audio object 73a, or an angular distance between the loudspeaker(s) and the audio object 73a.
[0061] For example, when considering the vertical distance, the loudspeaker (e.g., loudspeaker 11a) whose distance to the floor matches the height of the audio object (e.g., the y coordinate in MPEG-I audio) can be selected. This way, when the VR scene is moved, the floor of the VR scene can match the real-life room floor as closely as possible. In another example, if the matching of the height of the loudspeaker to the height of the audio object is not feasible, the VR scene can be modified by adjusting the height of the audio object (and any associated visual object) so that the VR scene floor matches the real-life room floor location when placed at the loudspeaker location.
[0062] In example embodiments, when more than one user is experiencing the VR scene, any changes to the audio rendering based on the relative positions of the users (e.g., with reference to Figures 1 to 3 The described changes in position / volume, such as loudspeaker directivity compensation, can be paused and / or cancelled.
[0063] In one example, one or more of the position, direction, and / or size of the visual representation of any audio object can be configured to be based at least in part on the apparent position of the respective audio object.
[0064] Figure 8 is a flowchart of an algorithm (generally indicated by reference numeral 80) in accordance with example embodiments. For better understanding, it can be helpful to consider Figure 9 in conjunction withFigure 8 .
[0065] Figure 9 is a block diagram of a system (generally indicated by reference numeral 90) in accordance with example embodiments. The system 90 illustrates a three-dimensional scene (e.g., scene 91) with multiple loudspeakers 11a-11e and two audio objects 92a (similar to the first audio object 73a) and 93a. The audio can be rendered so that the scene is being used by more than one user (such as users 12 and 78) (for simplicity, the users are not depicted here). Scenes 97, 98, and 99 illustrate how the scene 91 can be modified in accordance with example embodiments.
[0066] The algorithm 80 can be executed after operation 604 of the algorithm 60, where the apparent position of the first audio object (73a, 92a) is mapped to the position of the first loudspeaker of the plurality of loudspeakers. The algorithm 80 can begin at operation 801, where it is determined that the three-dimensional scene includes more than one audio object, such as the second audio object 93a, in addition to the first audio object 92a. If it is determined that the scene includes more than one audio object, the algorithm can move to operation 802.
[0067] At operation 802, the apparent position of the second audio object 93a can be configured to be mapped to the position of a second loudspeaker of the plurality of loudspeakers.
[0068] In one example, the mapping can include performing one or more of a translation, a rotation, and / or a scaling of the three-dimensional scene (e.g., one or more audio objects of the three-dimensional scene). For example, the scene can be translated as shown by arrow 94 so that the audio object 92a is mapped to the position of the loudspeaker 11a (e.g., similar to how the audio object 73a is mapped to the loudspeaker 11a). The new positions of the audio objects are shown by the apparent positions 92b and 93b in the scene 97. Next, as shown in the scene 97, the scene can be rotated as shown by arrow 95 so that the positions of the audio objects can be updated to the apparent positions 92c and 93c as shown in the scene 98. Next, the scene can be translated as shown by arrow 96 so that the positions of the audio objects can be updated to the apparent positions 92d and 93d as shown in the scene 99. Thus, the audio objects 92a and 93a are mapped to the positions of the loudspeakers 11a and 11b, respectively. Accordingly, the multiple users can perceive the audio from the audio object 92a in a manner similar to how it is rendered from the position of the loudspeaker 11a and perceive the audio from the audio object 93a in a manner similar to how it is rendered from the position of the loudspeaker 11b, thereby preventing the artifacts as described with reference to FIG. 6. Figure 5The described audio and visual mismatch. Further, the apparent locations 92d and 93d can remain unchanged even if one of the users moves, or the users move in different directions or different distances, as the apparent locations 92d and / or 93d are not determined dependent on the locations of the users. In one example, one or more of the position, direction, and / or size of the visual representation of the audio object 92a and 93a can be configured to be based at least in part on the apparent locations 92d and 93d, respectively. Thus, the visual representation of the audio object 92a is configured to appear at the apparent location 92d, and the visual representation of the audio object 93a is configured to appear at the apparent location 93d, in order to improve consistency between the positions of the audio objects and the visual objects.
[0069] In an example embodiment, the loudspeaker 1 lb is selected for rendering audio from the audio object 93a based on the loudspeaker 1 lb being the closest one of the plurality of loudspeakers (e.g., smallest horizontal, vertical, and / or angular distance) relative to the second audio object 93a.
[0070] Figure 10 is a flowchart of an algorithm (generally indicated by reference numeral 100) according to an example embodiment. For a better understanding, it can be helpful to consider Figure 11 in conjunction with Figure 10 .
[0071] Figure 11 is a block diagram of a system (generally indicated by reference numeral 110) according to an example embodiment. The system 110 shows a three-dimensional scene (e.g., scene 111a) with a plurality of loudspeakers 11a-11e and a plurality of audio objects 112a, 113a, 114a, 115a, 116a, and 117a. Audio can be rendered so that the scene is being used by more than one user (e.g., users 12 and 78, which are not shown here for simplicity). The scenes 111b-111d show how the scene 111a is modified according to an example embodiment.
[0072] The algorithm 100 can start at operation 1001, where it is determined that a three-dimensional scene includes more than a threshold number of audio objects, such as the audio objects 112a-117a. In one example, the threshold number can be 2 or 3. In another example, the threshold can depend on the number of loudspeakers. For example, the threshold number can be half of the number of loudspeakers (e.g., rounded up or down to the nearest whole number), the threshold number can be equal to the number of loudspeakers, the threshold number can be twice the number of loudspeakers, the threshold number can be the number of loudspeakers plus or minus a predefined number, etc.
[0073] If it is determined at operation 1001 that the scene includes more than a threshold number of audio objects, the algorithm can move to operation 1002. In one example, the threshold number can be 2, such that when there are more than two audio objects, operation 1002 can be performed. At operation 1002, the apparent position of at least one audio object is configured to be mapped outside of the first loudspeaker arrangement. For example, the threshold number (e.g., two) of audio objects can be mapped onto respective loudspeakers, and the remaining audio objects, other than the threshold number, can be mapped outside of the first loudspeaker arrangement.
[0074] For example, scene 111b illustrates how audio objects 112a-117a can initially be positioned relative to one another (e.g., using the dashed lines between positions 112b-117b to illustrate the constellation of audio objects). Since the number of audio objects 112a-117a can be more than the threshold number (e.g., 2), operation 1002 can be performed accordingly. Operation 1002 can be performed, for example, by shifting, translating, and / or scaling one or more audio objects of the scene. Scene 111c illustrates that audio objects 112a (at position 112c) and 116a (at position 116c) are mapped onto loudspeakers 1 lb and 11c, respectively, and further illustrates that audio objects 113a, 114a, 115a, and 117a (at positions 113c, 114c, 115c, and 117c, respectively) are mapped outside of the loudspeaker arrangement of loudspeakers 11a-11e. In example embodiments, rendering audio corresponding to audio objects mapped outside of the first loudspeaker arrangement can include rendering the audio with a spatial extent.
[0075] In one example, loudspeakers 1 lb and 11c can be selected based on the distance (e.g., horizontal, vertical, and / or angular distance) of the respective loudspeaker from the audio object to be used for audio objects 112a (at position 112d) and 116a (at position 116d), respectively, as described with reference to Figure 9
[0076] In one example, audio objects 112a and 116a can be selected for mapping onto the loudspeakers (11b, 11c, respectively) because the two audio objects are located next to each other at the edge of the audio object constellation (e.g., based on a triangulation of the audio object locations). The scene can then be shifted, rotated, and scaled so that these audio objects 112d and 116d are aligned with the two loudspeakers 11b and 11c. For audio objects that do not match up with the locations of the loudspeakers, such as audio objects 113c, 114c, 115c, and 117c, the audio processing can be modified. For example, a spatial spread can be applied to audio objects 113c, 114c, 115c, and 117c so that the user can perceive the rendered audio as coming from a general direction of the loudspeakers 11b and 11c (e.g., so that any mismatch between the audio objects and the visual objects can be less apparent). In one example, the spatially spread source can be rendered via the creation of multiple uncorrelated N "auxiliary" audio objects from the original audio object and placing them around the original audio object.
[0077] The example embodiments described above provide a system for loudspeaker rendering of 6DoF VR content to multiple users by modifying the VR scene based on the loudspeaker locations and independent of the location of each individual user so that any mismatch between the user's perception of the audio objects and the corresponding visual objects can be minimized. This can be achieved by configuring the apparent locations of one or more audio objects to be mapped onto the locations of the loudspeakers, respectively.
[0078] For completeness, Figure 12 is a schematic diagram of components of one or more of the example embodiments described previously (which are generally referred to below as processing system 300). Processing system 300 can have a processor 302, a memory 304 (which is tightly coupled to the processor and includes RAM 314 and ROM 312), and optionally a user input 310 and a display 318. Processing system 300 can include one or more network / device interfaces 308 for connecting to a network / device, e.g., a modem, which can be wired or wireless. Interface 308 can also be used as a connection to other devices, such as equipment / devices that are not network-side devices. Thus, direct connections between equipment / devices that do not require network participation are possible.
[0079] Processor 302 is connected to each of the other components in order to control their operation.
[0080] The memory 304 can comprise non-volatile memory such as a hard disk drive (HDD) or a solid state drive (SSD). The ROM 312 in the memory 304 stores an operating system 315 and can store software applications 316. The RAM 314 in the memory 304 is used by the processor 302 for the temporary storage of data. The operating system 315 can contain computer program code which, when executed by the processor, implements aspects of the algorithms 60, 80, 90 described above. Note that in the case of small devices / apparatuses, the memory can be most suitable for small size use, i.e. not always a hard disk drive (HDD) or solid state drive (SSD) is used.
[0081] The processor 302 can take any suitable form. For example, it can be one microcontroller, a plurality of microcontrollers, one processor, or a plurality of processors.
[0082] The processing system 300 can be a standalone computer, a server, a console, or a network thereof. The processing system 300 and the required structural components can all be within a device / apparatus, such as an loT device / apparatus, i.e. embedded to very small size.
[0083] In some example embodiments, the processing system 300 can also be associated with external software applications. These applications can be applications stored on a remote server device / apparatus and can be run partly or exclusively on the remote server device / apparatus. These applications can be referred to as cloud-hosted applications. The processing system 300 can communicate with the remote server device / apparatus in order to utilize the software applications stored there.
[0084] Figure 13 A tangible medium storing computer readable code which, when run by a computer, can perform a method according to the example embodiments described above is shown, in particular a removable storage unit 365. The removable storage unit 365 can be a memory stick, such as a USB memory stick, having an internal memory 366 for storing the computer readable code. The internal memory 366 is accessible by a computer system via a connector 367. Other forms of tangible storage medium can be used. The tangible medium can be any device / apparatus capable of storing data / information which can be exchanged between devices / apparatuses / networks.
[0085] Embodiments of the application can be implemented in software, hardware, application logic or a combination of software, hardware and application logic. The software, application logic and / or hardware can reside on memory, a computer, or any computer medium. In an example embodiment, the application logic, software or an instruction set is maintained on any one of various conventional computer-readable media. In the context of this document, a "computer-readable medium" can be any non-transitory medium that can contain, store, communicate, propagate or transport the program for use by or in connection with the instruction execution system, apparatus or device (e.g., a computer).
[0086] In related contexts, references to "computer-readable storage medium", "computer program product", "tangibly embodied computer program" etc., or a "processor" or "processing circuitry" etc. should be understood to encompass not only computers having differing architectures such as single / multi-processor architectures and sequencers / parallel architectures, but also specialized circuits such as FPGAs, ASICs, signal processing devices / apparatus and other devices / apparatus. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of FPGAs, ASICs, signal processing devices / apparatus and other devices / apparatus, whether referred to as software or firmware.
[0087] As used in this application, the term "circuitry" refers to all of the following: (a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of circuits and software (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions of processor(s) / software (including digital signal processors), software, and memory that work together to cause an apparatus, such as a server, to perform various functions) and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even if the software or firmware is not physically present.
[0088] If desired, the different functions discussed herein can be performed in a different order and / or concurrently with each other. Furthermore, if desired, one or more of the functions will be optional. Similarly, it will be further appreciated that, Figure 6 , Figure 8 and Figure 10 The flow diagrams depicted herein are merely examples, and that various operations depicted can be omitted, reordered and / or combined.
[0089] It will be understood that the example embodiments described above are only illustrative and that modifications and alterations can become apparent to those skilled in the art from a reading and understanding of this specification. Other changes and modifications of the example embodiments described herein will also become apparent to those skilled in the art from a reading and understanding of this specification.
[0090] Furthermore, the disclosure herein is to be understood as being without limitation in regard to any novel feature or any novel combination of features, or any generalization thereof, that is either expressly disclosed herein or inherent to those with ordinary skill in the art, and that could be claimed in any application derived therefrom, as a new claim can be formulated during prosecution of an application to encompass any such features and / or combinations of features.
Claims
1. An apparatus comprising: means for rendering audio of a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the means for rendering comprises: means for determining a number of users that are using the three-dimensional scene; and means for, based on determining that more than one user is using the three-dimensional scene, configuring an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
2. The apparatus of claim 1, wherein, configuring the apparent position comprises rendering audio corresponding to the first audio object from the first loudspeaker.
3. The apparatus of any one of the preceding claims, further comprising: means for determining whether the three-dimensional scene comprises a second audio object in addition to the first audio object; wherein the means for rendering further comprises means for, based on determining that the three-dimensional scene comprises at least the first audio object and the second audio object, configuring an apparent position of the second audio object to be mapped to a position of a second loudspeaker of the plurality of loudspeakers.
4. The apparatus of claim 3, wherein, the means for configuring the apparent positions of the first audio object and the second audio object to be mapped to the positions of the first loudspeaker and the second loudspeaker comprises means for performing one or more of shifting, rotating, and / or scaling the three-dimensional scene.
5. The apparatus of claim 4, wherein, the shifting comprises shifting one or more of the audio objects within the three-dimensional scene to be mapped to the first loudspeaker and the second loudspeaker.
6. The apparatus of any one of the preceding claims, further comprising: means for determining whether the three-dimensional scene comprises more than a threshold number of audio objects; wherein, if the three-dimensional scene comprises more than the threshold number of audio objects, the means for rendering comprises means for configuring an apparent position of at least one of the audio objects to be mapped to other than the first loudspeaker arrangement.
7. The apparatus of claim 6, wherein, the apparent position of each of at least two of the audio objects is configured to be mapped to a position of each of at least two loudspeakers, respectively.
8. The apparatus of claim 6, wherein, rendering audio corresponding to an audio object to be mapped to other than the first loudspeaker arrangement comprises rendering the audio with a spatial extent.
9. The device according to any one of the preceding claims, wherein the means for rendering comprises means for rendering the audio independently of changes in position of the more than one user.
10. The apparatus of any of the preceding claims, further comprising: the means for rendering further comprises:
11. The device of any of the preceding claims, wherein, means for, based on determining that a single first user is using the three-dimensional scene, rendering audio from one or more of the plurality of loudspeakers based at least in part on an initial position of the first audio object and a position of the first user. 12. The apparatus of claim 11, wherein, The means for rendering further comprises changing a direction and / or a volume of the audio based at least in part on a change in a location of the first user.
13. The device of any of the preceding claims, wherein, The first loudspeaker is selected from the plurality of loudspeakers based on which loudspeaker of the plurality of loudspeakers has a minimum distance to the first audio object.
14. A method comprising: rendering audio of a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users that are using the three-dimensional scene; and based on determining that more than one user is using the three-dimensional scene, configuring an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.
15. A computer program comprising instructions which, when executed by an apparatus, cause the apparatus to: rendering audio of a three-dimensional scene at one or more of a plurality of loudspeakers, wherein the plurality of loudspeakers are arranged in a first loudspeaker arrangement suitable for providing spatial audio corresponding to the three-dimensional scene, wherein the three-dimensional scene comprises a first audio object, wherein the rendering comprises: determining a number of users that are using the three-dimensional scene; and based on determining that more than one user is using the three-dimensional scene, configuring an apparent position of the first audio object to be mapped to a position of a first loudspeaker of the plurality of loudspeakers.