Spatialized audio chat in the virtual metaverse

Spatialized audio in virtual metaverses transforms audio streams based on avatar positions and scene information to enhance immersion, addressing the lack of spatial depth in existing audio technologies.

JP7841074B2Active Publication Date: 2026-04-06ROBLOX CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-15
Publication Date
2026-04-06

AI Technical Summary

Technical Problem

Existing audio technologies in virtual environments provide monaural or stereo audio that lacks spatial depth, detracting from the immersive experience as the virtual experience becomes more visually immersive.

Method used

Implementing spatialized audio in virtual metaverses by transforming audio streams based on avatar positions, velocities, and scene information, combining them to create a spatialized audio stream that accounts for proximity, orientation, and other factors.

Benefits of technology

Enhances the immersive experience by providing spatialized audio that accurately reflects the virtual environment, prioritizing audio streams based on proximity and movement, and reducing computational complexity and bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007841074000001
    Figure 0007841074000001
  • Figure 0007841074000002
    Figure 0007841074000002
  • Figure 0007841074000003
    Figure 0007841074000003
Patent Text Reader

Abstract

Implementations described herein relate to methods, systems, and computer-readable media for providing spatialized audio in a virtual experience. Spatialized audio may be used, for example, in audio communications such as voice and / or video chat. Chat may include spatialized audio that is combined and directed to a particular user at a client device or online experience platform. Individual audio streams may be collected from multiple avatars and other objects and combined based on a target user. Audio may also include background and / or environmental sounds to provide a rich and immersive audio stream in the virtual experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Provisional Application No. 63 / 222,304, filed Jul. 15, 2021, entitled "SPATIALIZED AUDIO CHAT IN A VIRTUAL METAVERSE", the entire content of which is incorporated herein by reference.

[0002] Embodiments generally relate to audio output via a computer device, and more particularly to methods, systems, and computer - readable media for providing spatialized audio in a virtual immersive environment such as a metaverse place in a virtual metaverse.

Background Art

[0003] Computer audio (e.g., a chat between users of a computer device) often consists of providing monaural or stereo audio when the chat is received from a listening device or microphone. The audio provided is generally either unfiltered or minimally filtered and may sound flat or direct, regardless of the actual virtual positions of the two avatars representing the users participating in the chat. Thus, as the virtual experience becomes more visually immersive, the simplistic nature of the provided audio can impede and / or detract from the immersive experience, e.g., causing the user to be pulled out of the experience.

[0004] The description of the background art provided herein is for the purpose of presenting the context of the present disclosure. The research of the inventors named herein within the scope described in this background art section, and aspects of the description that may not be considered prior art at the time of filing, are not admitted as prior art to the present disclosure, either explicitly or implicitly.

Summary of the Invention

[0005] The implementation of this application relates to providing spatialized audio in a virtual metaverse.

[0006] According to one embodiment, a computer-based method for spatialized audio in a virtual metaverse, comprising the steps of: receiving a request from a first user among a plurality of users to receive audio related to a metaverse place in the virtual metaverse, wherein the first user is associated with a user device, and the plurality of users are associated with each of the plurality of avatars in the metaverse place; retrieving a data model related to the metaverse place, wherein the data model includes one or more spatial parameters representing one or more physical laws applied to the metaverse place; and extracting avatar information and scene information from the data model, wherein the avatar information includes the position and velocity of the plurality of avatars in the metaverse place, including the first avatar associated with the first user. A computer-implemented method is disclosed, which includes the steps of: the scene information includes one or more of the following: an occlusion, reverberation, or a virtual wall that virtually approaches a first avatar in the metaverse place; the steps of: transforming each audio stream received from each of a plurality of users based on the avatar information and scene information, and transforming one or more audio characteristics of at least one of the audio streams based on one or more spatial parameters, in order to create a spatialized audio stream; combining the spatialized audio streams to create a combined spatialized audio stream; and providing the combined spatialized audio stream to a user device.

[0007] Various implementations of the method performed by computer are described herein.

[0008] In some implementations, spatial parameters include distance attenuation parameters to attenuate audio based on the distance between avatars.

[0009] In some implementations, each audio stream received from multiple users includes the mono audio received by the microphone device, while the combined spatialized audio stream includes stereo audio.

[0010] In some implementations, the combined spatialized audio stream includes stereo audio generated by placing each user's mono audio at the location of their respective avatar.

[0011] In some implementations, the combined spatialized audio stream includes spatial audio based on audio streams received from users other than the first user among multiple users, and background audio, the background audio being generated based on one or more of the following: audio received from users other than the first user, and audio generated based on the movement of avatars in the metaverse place.

[0012] In some implementations, the method performed by the computer further includes the step of determining a set of prioritized audio streams received from each of multiple users, and the step of transforming each audio stream further includes transforming the set of prioritized audio streams in order to create a spatialized audio stream.

[0013] In some implementations, the step of determining a set of prioritized audio streams includes prioritizing audio streams received from each of several users based on one or more of the following: proximity of avatars in the metaverse place, velocity of avatars in the metaverse place, orientation of avatars in the metaverse place, virtual objects adjacent to avatars in the metaverse place, capabilities of the user device, or the user preferences of the first user.

[0014] In some implementations, audio streams associated with avatars closer to the receiving avatar take precedence over audio streams associated with avatars further away from the receiving avatar; audio streams associated with avatars directed towards the receiving avatar take precedence over audio streams associated with avatars directed away from the receiving avatar; and audio streams associated with avatars moving toward the receiving avatar take precedence over audio streams associated with avatars moving away from the receiving avatar.

[0015] In another embodiment, a computer-implemented method for providing spatialized audio in a virtual metaverse is disclosed, comprising the steps of: receiving a request from a first user among a plurality of users to receive audio related to a metaverse place in the virtual metaverse, wherein the first user is associated with a user device, and the plurality of users are associated with each avatar of a plurality of avatars in the metaverse place; determining a set of prioritized audio streams received from each of the plurality of users; transforming the set of prioritized audio streams to create a spatialized audio stream; combining the spatialized audio streams to create a combined spatialized audio stream; and providing the combined spatialized audio stream to the user device.

[0016] Various implementations of the method performed by computer are described herein.

[0017] In some implementations, the step of determining a set of prioritized audio streams includes prioritizing audio streams received from each of several users based on one or more of the following: proximity of avatars in the metaverse place, velocity of avatars in the metaverse place, orientation of avatars in the metaverse place, virtual objects adjacent to avatars in the metaverse place, capabilities of the user device, or the user preferences of the first user.

[0018] In some implementations, audio streams associated with avatars closer to the receiving avatar take precedence over audio streams associated with avatars further away from the receiving avatar; audio streams associated with avatars directed towards the receiving avatar take precedence over audio streams associated with avatars directed away from the receiving avatar; and audio streams associated with avatars moving toward the receiving avatar take precedence over audio streams associated with avatars moving away from the receiving avatar.

[0019] In another embodiment, the processing device includes a memory storing instructions and a processing device coupled to the memory, configured to access the memory, wherein when the instructions are executed by the processing device, the processing device receives a request from a first user among a plurality of users to receive audio related to a metaverse place of a virtual metaverse, the first user being associated with a user device, and the plurality of users being associated with each of the avatars of a plurality of avatars in the metaverse place; retrieves a data model related to the metaverse place, the data model including one or more spatial parameters representing one or more physical laws applied to the metaverse place; and extracts avatar information and scene information from the data model, the avatar information being provided to the first user. A system is disclosed that performs operations including extracting scene information which includes one or more of the positions, velocities, or directions of multiple avatars in a metaverse place including a related first avatar, and which includes one or more of the occlusions, reverberations, or virtual walls that are virtually close to the first avatar in the metaverse place; transforming each audio stream received from each of multiple users based on the avatar information and scene information to create a spatialized audio stream, transforming one or more audio characteristics of at least one of the audio streams based on one or more spatial parameters; combining the spatialized audio streams to create a combined spatialized audio stream; and providing the combined spatialized audio stream to a user device.

[0020] Various implementations of the system are described herein.

[0021] In some implementations, spatial parameters include distance attenuation parameters to attenuate audio based on the distance between avatars.

[0022] In some implementations, each audio stream received from multiple users includes the mono audio received by the microphone device, while the combined spatialized audio stream includes stereo audio.

[0023] In some implementations, the combined spatialized audio stream includes stereo audio generated by placing each user's mono audio at the location of their respective avatar.

[0024] In some implementations, the combined spatialized audio stream includes spatial audio based on audio streams received from users other than the first user among multiple users, and background audio, the background audio being generated based on one or more of the following: audio received from users other than the first user, and audio generated based on the movement of avatars in the metaverse place.

[0025] In some implementations, the operation further involves determining a set of prioritized audio streams received from each of multiple users, and transforming each audio stream further involves transforming the set of prioritized audio streams to create a spatialized audio stream.

[0026] In some implementations, determining a set of prioritized audio streams involves prioritizing audio streams received from each of multiple users based on one or more of the following: proximity of avatars in the metaverse place, velocity of avatars in the metaverse place, orientation of avatars in the metaverse place, virtual objects adjacent to avatars in the metaverse place, capabilities of the user device, or the user preferences of the first user.

[0027] In some implementations, the audio stream associated with an avatar closer to the receiving avatar is prioritized over the audio stream associated with an avatar farther from the receiving avatar, and the audio stream associated with an avatar facing the receiving avatar is prioritized over the associated audio stream.

[0028] In some implementations, the system further includes a spatial audio manager configured to transform each audio stream received from each of a plurality of users, and an audio device override module configured to disable non-spatialized audio on the user device before providing the combined spatial audio stream to the user device.

[0029] According to another aspect, a non - transient computer - readable medium is provided. In response to execution by a processing device, the processing device retrieves a data model related to a metaverse place in a virtual metaverse, the data model including one or more spatial parameters representing a group of physical laws applied to the metaverse place; receives a request to participate in the metaverse place from a first user among a plurality of users, the first user being associated with a first avatar and a user device, and the plurality of users being associated with a plurality of avatars within the metaverse place; extracts avatar information and scene information from the data model in response to the request, the avatar information including one or more of the position, velocity, or direction of the first avatar and the plurality of avatars within the metaverse place, and the scene information including one or more of occlusion, reverberation, or virtual walls that are virtually proximate to the first avatar; uses the spatial parameters to transform each respective audio stream received from each user among the plurality of users based on the avatar information and the scene information, including modifying one or more audio characteristics to create a spatialized audio stream; combines the spatialized audio streams to create a combined spatialized audio stream; and stores instructions for causing the processing device to perform operations including providing the combined spatialized audio stream to the user device.

[0030] Various implementations of the non - transient computer - readable medium are described herein.

[0031] According to yet another aspect, details of a system, method, and non - transient computer - readable medium, parts, features, and implementations may be combined to form additional aspects that include omitting and / or modifying some or a portion of individual components or features, and including additional components or features, and / or other modifications, and all such modifications are within the scope of this disclosure. [Brief explanation of the drawing]

[0032] [Figure 1] This is a diagram illustrating an exemplary network environment for providing spatialized audio chat in a virtual metaverse, based on some implementations. [Figure 2A] This is a diagram illustrating an exemplary network environment for providing spatialized audio chat in a virtual metaverse, based on some implementations. [Figure 2B] This is a diagram illustrating an exemplary network environment for providing spatialized audio chat in a virtual metaverse, based on some implementations. [Figure 3] This is a diagram illustrating an exemplary network environment for prioritizing spatialized audio streams in a virtual metaverse, as implemented in some cases. [Figure 4] This figure shows an example of a 3D metaverse place within a virtual experience, as demonstrated by some implementations. [Figure 5] This figure shows an example of a 3D metaverse place within a virtual experience, as demonstrated by some implementations. [Figure 6] This is a flowchart illustrating an exemplary method for providing spatialized audio chat in a virtual metaverse, based on some implementations. [Figure 7] This is a flowchart illustrating an exemplary method for prioritizing spatialized audio streams in a virtual metaverse, as used in some implementations. [Figure 8] This block diagram shows an exemplary computing device that may be used to implement one or more features described herein, depending on the implementation. [Modes for carrying out the invention]

[0033] One or more implementations described herein relate to spatialized audio related to online game platforms. Features may include automatically prioritizing spatialized audio streams based on position, velocity, and / or other factors related to virtual objects, avatars, and other items within a metaverse place of a virtual metaverse, and providing spatialized audio.

[0034] The features described herein provide spatialized audio for output on client devices connected to an online platform, such as an online experience platform or an online game platform. The online platform may provide a virtual metaverse having multiple associated metaverse places. A virtual avatar associated with a user can move between metaverse places and interact with them, as well as with items, characters, other avatars, and objects within them. The avatar can move from one metaverse place to another while experiencing spatialized audio that provides a more immersive and enjoyable experience. Spatialized audio streams from multiple users (for example, avatars associated with multiple users) can be prioritized based on many factors so that rich audio can be provided, taking into account the position, velocity, movement, and actions of avatars and characters, as well as the bandwidth, processing, and other capabilities of the client device.

[0035] By prioritizing and combining different audio streams, a combined spatialized audio stream can be provided for output on client devices, offering a richer user experience, reduced computational complexity, and lower bandwidth requirements for spatialized audio without compromising the virtual immersion experience. Furthermore, a spatial audio application programming interface (API) is defined, enabling users and developers to implement spatialized audio for virtually any online experience, thereby allowing the creation of high-quality online virtual experiences, games, metaverse places, and other interactions with immersive audio while requiring lower technical proficiency from users and developers.

[0036] Online experience platforms and online game platforms (also known as “user-generated content platforms” or “user-generated content systems”) provide various ways for users to interact with each other. For example, users of an online experience platform may create games or other content or resources within the online platform, such as characters, graphics, and items for use in gameplay and / or within a virtual metaverse.

[0037] Users of online experience platforms may collaborate towards common goals in metaverse places, games, or game creation, share various virtual items (e.g., inventory items, game items), participate in audio chat (e.g., spatialized audio chat), and send electronic messages to each other. Users of online experience platforms may interact with other users, for example, by playing games that include characters (avatars) or other game objects and mechanisms. Online experience platforms may also enable users of the platform to communicate with each other. For example, users of online experience platforms may communicate with each other using voice messages (e.g., voice chat with spatialized audio), text messaging, video messaging (e.g., including spatialized audio), or a combination of the above. Some online experience platforms may provide virtual 3D environments, or multiple linked environments within the metaverse, within which users can interact with each other or play online games.

[0038] To enhance the entertainment value of an online experience platform, the platform can provide rich audio for playback on the user's device. This audio may include, for example, different audio streams from different users and background audio. According to the various implementations described herein, different audio streams can be converted into spatialized audio streams. Spatialized audio streams may be combined, for example, to provide a combined spatialized audio stream for playback on the client device. Furthermore, prioritized audio streams may be provided to reduce bandwidth while still providing immersive spatialized audio. Additionally, background audio streams may be combined with spatialized audio so that realistic background noise / effects are also played for the user. Furthermore, characteristics of the metaverse place, such as the surrounding medium (air, water, etc.), reverberation, reflections, hole sizes, wall density, ceiling height, doorways, corridors, object placement, non-player objects / characters, and other properties, are utilized to create spatialized and / or background audio to enhance realism and immersion within the online virtual experience.

[0039] Figures 1-3: System Architecture Figure 1 shows an exemplary network environment 100 based on a partial implementation of the present disclosure. The network environment 100 (also referred to herein as the “System”) includes an online experience platform 102, a first client device 110, and a second client device 116 (collectively referred to herein as “Client Devices 110 / 116”), all connected via a network 122. The online experience platform 102 may include, among other things, a game engine 104, one or more games 105, a spatialized audio API 106, and a data store 108. Client device 110 may include a game application 112, and client device 116 may include a game application 118. Users 114 and 120 can use client devices 110 and 116, respectively, to interact with the online experience platform 102 and other users utilizing the online experience platform 102.

[0040] The network environment 100 is provided for illustrative purposes. In some implementations, the network environment 100 may contain the same, fewer, more, or different elements configured in the same or different ways as shown in Figure 1.

[0041] In some implementations, network 122 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or wide area network (WAN)), a wired network (e.g., an Ethernet network), a wireless network (e.g., an 802.11 network, a Wi-Fi® network, or a wireless LAN (WLAN)), a cellular network (e.g., a Long-Term Evolution (LTE) network), a router, a hub, a switch, a server computer, or a combination thereof.

[0042] In some implementations, the datastore 108 may be non-temporary computer-readable memory (e.g., random-access memory), a cache, a drive (e.g., a hard drive), a flash drive, a database system, or another type of component or device capable of storing data. The datastore 108 may also include multiple storage components (e.g., multiple drives or multiple databases) that may span multiple computing devices (e.g., multiple server computers).

[0043] In some implementations, the online experience platform 102 may include a server having one or more computing devices (e.g., a cloud computing system, a rack-mount server, a server computer, a cluster of physical servers, a virtual server, etc.). In some implementations, the server may be included in the online experience platform 102, a separate system, or part of another system or platform.

[0044] In some implementations, the online experience platform 102 may include one or more computing devices (such as rack-mount servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, and desktop computers), data stores (e.g., hard disks, memory, databases), networks, software components, and / or hardware components that run on the online experience platform 102 and may be used to provide users with access to the online experience platform 102. The online experience platform 102 may also include a website (e.g., one or more web pages) or application backend software that may be used to provide users with access to content provided by the online experience platform 102. For example, users 114 / 120 may access the online experience platform 102 using a game application 112 / 118 on a client device 110 / 116.

[0045] In some implementations, the online experience platform 102 may include some kind of social network that provides connections between users, or some kind of user-generated content system that enables users (e.g., end users or consumers) to communicate with other users through the online experience platform 102, and the communication may include voice chat (e.g., synchronous and / or asynchronous voice communication with or without spatialized audio), video chat (e.g., synchronous and / or asynchronous video communication with or without spatialized audio), or text chat (e.g., synchronous and / or asynchronous text-based communication).

[0046] In some implementations of this disclosure, “User” may represent a single individual. However, other implementations of this disclosure may include the fact that “User” (e.g., a creator user) is an entity controlled by a set of users or an automated source. For example, a set of individual users united as a community or group within a user-generated content system may be considered “User.”

[0047] In some implementations, the online experience platform 102 may be a virtual game platform. For example, the game platform may provide single-player or multiplayer games to a community of users who may access or interact with games (e.g., user-generated games or other games) using client devices 110 / 116 via the network 122. In some implementations, games (also referred herein as “video games,” “online games,” “metaverse places,” or “virtual experiences”) may be, for example, two-dimensional (2D) games, three-dimensional (3D) games (e.g., 3D user-generated games), virtual reality (VR) games, or augmented reality (AR) games. In some implementations, users may search for games and game items and participate in gameplay with other users in one or more games. In some implementations, games may be played in real time with other users of the game. Similarly, some users may participate in real-time voice or video chats with other users of the game. As described herein, real-time voice or video chats may include spatialized audio.

[0048] In some implementations, other collaboration platforms may be used in place of, or in addition to, the online experience platform 102 and / or spatialized audio API 106, along with the features described herein. For example, social networking platforms, purchase platforms, messaging platforms, and creation platforms may be used with spatial audio features so that immersive spatialized audio is delivered to users outside the game.

[0049] In some implementations, gameplay may refer to interactions between one or more players using client devices (e.g., 110 and / or 116) within a game (e.g., 105), or to representations of interactions on the display or other output devices of client devices 110 or 116. In some implementations, gameplay may instead refer to interactions within a virtual experience or metaverse place, which may be similar to, different from, or have the same purpose as some games. Furthermore, although referred to as “players,” the terms “avatar,” “user,” and / or other terms may be used to refer to users who engage in and / or interact with an online virtual experience.

[0050] One or more games 105 are provided by an online experience platform. In some implementations, a game 105 may include electronic files that can be executed or loaded using software, firmware, or hardware configured to present game content (e.g., digital media items) to an entity. In some implementations, a game application 112 / 118 may run in conjunction with a game engine 104, and a game 105 may be rendered in conjunction with the game engine 104. In some implementations, a game 105 may have a common set of rules or common goals, and the virtual environment of a game 105 may share a common set of rules or common goals. In some implementations, different games may have different rules or goals from each other. It should be noted that, although specifically referred to as “games” or game-related, the game application 112 / 118, game 105, and game engine 104 may also be referred to as a virtual experience application 112 / 118, a virtual experience 105, and / or a virtual experience engine 104.

[0051] In some implementations, a game and / or virtual experience may have one or more environments (also referred to herein as “game environments,” “metaverse places,” or “virtual environments”), and multiple environments may be linked. An example of an environment may be a three-dimensional (3D) environment. One or more environments of game 105 or a virtual experience may collectively be referred herein as a “world,” “game world,” “virtual world,” “universe,” or “metaverse.” An example of a world may be a 3D metaverse place in game 105. For example, a user may construct a metaverse place that is linked to another metaverse place created by a different user than the first user. A character in a virtual experience may cross a virtual boundary to enter an adjacent metaverse place. Furthermore, sounds, theme music, and / or background music may also cross a virtual boundary so that an avatar standing near a virtual boundary may hear spatialized audio that includes at least some of the sounds emanating from an adjacent metaverse place. In this way, spatialized audio may enable a fully immersive experience, including virtual audio that represents the similarity of sound propagation to that of a real-world environment.

[0052] It should be noted that 3D environments or 3D worlds use graphics that utilize a three-dimensional representation of geometric data representing the content (or at least present the content in a way that makes it appear as 3D content, regardless of whether a three-dimensional representation of geometric data is used). 2D environments or 2D worlds use graphics that utilize a two-dimensional representation of geometric data representing the game content.

[0053] In some implementations, the online experience platform 102 may host one or more games 105 and allow users to interact with the games 105 using game applications 112 / 118 on client devices 110 / 116 (for example, to search for games, game-related content, or other content). Users of the online experience platform 102 (for example, 114 and / or 120) may play, create, interact with, or build games 105, search for games 105, communicate with other users, create and build objects for games 105 (for example, also referred to herein as “items,” “game objects,” or “virtual game items”), and / or search for objects. For example, when creating user-generated virtual items, users may, among other things, create characters, decorations for characters, one or more virtual environments for an interactive game, or build structures used within game 105.

[0054] In some implementations, users may buy, sell, or trade virtual game objects of a game, such as in-platform currency (e.g., virtual currency), with other users of the online experience platform 102. In some implementations, the online experience platform 102 may transmit game content to a game application (e.g., 112). In some implementations, game content (also referred to herein as “Content”) may refer to any data or software instructions related to the online experience platform 102 or the game application (e.g., game objects, games, user information, videos, images, commands, media items, etc.).

[0055] In some implementations, a game object (for example, also referred to herein as “item,” “object,” or “virtual game item”) may refer to an object used, created, shared, or otherwise depicted in a game application 105 on an online experience platform 102 or in a game application 112 or 118 on a client device 110 / 116. For example, a game object may include parts, models, characters, tools, weapons, clothing, buildings, vehicles, currency, flora, fauna, and components of the aforementioned (for example, windows of a building).

[0056] It should be noted that the online experience platform 102 hosting game 105 is provided for illustrative purposes only, not limitation. In some implementations, the online experience platform 102 may host one or more media items that may contain communication messages from one user to one or more other users. Media items may include, but are not limited to, digital video, digital movies, digital photographs, digital music, audio content, melodies, website content, social media updates, ebooks, e-magazines, digital newspapers, digital audiobooks, e-journals, web blogs, real simple syndication (RSS) feeds, e-comics, and software applications. In some implementations, media items may be electronic files that can be executed or loaded using software, firmware, or hardware configured to present digital media items to entities.

[0057] In some implementations, a game 105 may be associated with a specific user or a specific group of users (for example, a private game), or it may be made available to users of the online experience platform 102 (for example, a public game). In some implementations where the online experience platform 102 associates one or more games 105 with a specific user or group of users, the online experience platform 102 may associate a specific user with a game 105 using user account information (for example, a user account identifier such as a username and password). Similarly, in some implementations, the online experience platform 102 may associate a specific developer or group of developers with a game 105 using developer account information (for example, a developer account identifier such as a username and password).

[0058] In some implementations, the online experience platform 102 or client devices 110 / 116 may include a game engine 104 or game applications 112 / 118. The game engine 104 may include game applications similar to game applications 112 / 118. In some implementations, the game engine 104 may be used for deploying or running a game 105. For example, the game engine 104 may include, among other features, a rendering engine ("renderer") for 2D, 3D, VR, or AR graphics, a physics engine, a collision detection engine (and collision response), a sound engine, a spatialized audio manager / engine, an audio mixer, an audio subscription exchange, an audio subscription logic, an audio subscription prioritizer, a real-time communication engine, scripting capabilities, an animation engine, an artificial intelligence engine, networking capabilities, streaming capabilities, memory management capabilities, threading capabilities, scene graph capabilities, or video support for cinematic techniques. The components of the game engine 104 may generate commands (e.g., rendering commands, collision commands, physics commands, etc.) that help compute and render the game, and may convert audio (e.g., converting mono or stereo sound to a spatialized audio stream, etc.). In some implementations, the game applications 112 / 118 on client devices 110 / 116 may work independently, in cooperation with the game engine 104 on the online experience platform 102, or in a combination of both.

[0059] In some implementations, both the online experience platform 102 and client devices 110 / 116 run game engines (104, 112, and 118, respectively). The online experience platform 102, using game engine 104, may run some or all of the game engine's functions (e.g., generate physics commands, rendering commands, spatialized audio commands, etc.) or offload some or all of the game engine's functions to the game engine 104 on client device 110. In some implementations, each game 105 may have different ratios between the game engine's functions run on the online experience platform 102 and the game engine's functions run on client devices 110 and 116.

[0060] For example, the game engine 104 of the online experience platform 102 may be used to generate physics commands when a collision exists between at least two game objects, while further game engine functions (e.g., generating rendering commands or combining spatialized audio streams) may be offloaded to the client device 110. In some implementations, the ratio of game engine functions executed on the online experience platform 102 to those executed on the client device 110 may be changed (e.g., dynamically) based on gameplay conditions. For example, if the number of users participating in the gameplay of game 105 exceeds a threshold number, the online experience platform 102 may execute one or more game engine functions previously executed by the client device 110 or 116.

[0061] For example, a user may be playing game 105 on client devices 110 and 116 and may send control commands to the online game platform 102 (e.g., user input such as right, left, up, down, user selection, or character position and velocity information). After receiving control commands from client devices 110 and 116, the online experience platform 102 may send gameplay commands to client devices 110 and 116 based on the control commands (e.g., position and velocity information for characters participating in group gameplay, or commands such as rendering commands, collision commands, spatialized audio commands). For example, the online experience platform 102 may perform one or more logical actions according to the control commands (e.g., using the game engine 104) to generate gameplay commands for client devices 110 and 116. Otherwise, the online experience platform 102 may pass one or more control commands from one client device 110 to other client devices (e.g., 116) participating in game 105. Client devices 110 and 116 may use gameplay commands to render gameplay for presentation on the displays of client devices 110 and 116. Client devices 110 and 116 may also use gameplay commands to create, modify, and / or combine spatialized audio streams for output to the audio output devices of client devices 110 and 116.

[0062] In some implementations, control commands may refer to commands that indicate in-game actions for the user's character. For example, control commands may include user input, user selection, gyroscope position and orientation data, force sensor data, etc., to control in-game actions such as right, left, up, down. Control commands may also include character position and velocity information. In some implementations, control commands are sent directly to the online experience platform 102. In other implementations, control commands may be sent from client device 110 to another client device (e.g., 116), which then generates gameplay commands using the local game engine 104. Control commands may include commands to play voice communication messages or other sounds from another user through an audio device (e.g., speaker, headphones, etc.).

[0063] In some implementations, gameplay instructions may refer to instructions that enable a client device 110 (or 116) to render gameplay for a game, such as a multiplayer game. Gameplay instructions may include one or more of the following: user input (e.g., control instructions), character position and velocity information, or commands (e.g., physics commands, rendering commands, collision commands, etc.). As described in more detail herein, character position and velocity information may be used to determine an appropriate head-related transfer function (HRTF) for that other character so that a spatialized audio stream representing the propagation of sound in the real world can be created for that other character. The relevant HRTF, position information, velocity information, Baum-Welch (BW) algorithm data, virtual auditory display (VAD) data, and / or other data may be stored in a data store 108 by the online experience platform 102.

[0064] In some implementations, a character (or more broadly, a game object) is constructed from components that automatically combine to assist the user in editing, and one or more of these components may be selected by the user. One or more characters (also referred herein as “avatars” or “models”) may be associated with a user, and the user may control the characters to facilitate the user’s interaction with the game 105. In some implementations, a character may include components such as body parts (e.g., head, hair, arms, legs, etc.) and accessories (e.g., T-shirts, glasses, decorative images, tools, etc.). In some implementations, customizable character body parts may include, among other things, head type, body part type (arms, legs, torso, and hands), face type, hair type, and skin type. In some implementations, customizable accessories may include clothing (e.g., shirts, trousers, hats, shoes, glasses, etc.), weapons, or other tools.

[0065] In some implementations, the user may also control the character's scale (e.g., height, width, or depth) or the scale of the character's components. In some implementations, the user may also control the character's proportions (e.g., blocky, anatomical). In some implementations, a character may not include a character game object (e.g., body parts), but it should be noted that the user may control the character (without a character game object) to facilitate user interaction with the game (for example, a puzzle game where there is no rendered character game object, but the user still controls the character to control in-game actions).

[0066] In some implementations, components such as body parts may be basic geometric shapes such as blocks, cylinders, spheres, or any other basic shapes such as wedges, rings, tubes, or channels. In some implementations, the creator module may expose the user's character for viewing or use by other users of the online experience platform 102. In some implementations, creating, modifying, or customizing characters, other game objects, games 105, or game environments may be done by the user using a user interface (e.g., a developer interface), with or without scripting (or with or without an application programming interface (API)). It should be noted that, for illustrative purposes only and not limitation, characters are described as having the form of a humanoid robot. It should be further noted that characters may have any form, such as a vehicle, an animal, an inanimate object, or any other creative form.

[0067] In some implementations, the online experience platform 102 may store user-created characters in a data store 108. In some implementations, the online experience platform 102 maintains a character catalog and a game catalog that may be presented to the user via the game engine 104, the game 105, and / or client devices 110 / 116. In some implementations, the game catalog includes images of games stored in the online experience platform 102. In addition, the user may select a character (for example, a character created by the user or another user) from the character catalog to participate in a selected game. The character catalog includes images of characters stored in the online experience platform 102. In some implementations, one or more characters in the character catalog may be created or customized by the user. In some implementations, the selected character may have character settings that define one or more of the character's components.

[0068] In some implementations, a user's character may include the configuration of components, and the configuration and appearance of these components, as well as the character's appearance more broadly, may be defined by character settings. In some implementations, the user's character settings may be selected by the user, at least partially. In other implementations, the user may select a character that has default character settings or other user-selected character settings. For example, the user may select a default character with predefined character settings from a character catalog, and furthermore, the user may customize the default character by changing some of the character settings (for example, by adding a shirt with a customized logo). Character settings may be associated with a particular character by the online experience platform 102.

[0069] In some implementations, client devices 110 or 116 may include, respectively, computing devices such as personal computers (PCs), mobile devices (e.g., laptops, mobile phones, smartphones, tablet computers, or netbooks), network-connected televisions, and game consoles. In some implementations, client devices 110 or 116 may also be referred to as “user devices.” In some implementations, one or more client devices 110 or 116 may connect to the online experience platform 102 at any time. It should be noted that the number of client devices 110 or 116 is given as an example, not an limitation. In some implementations, any number of client devices 110 or 116 may be used.

[0070] In some implementations, each client device 110 or 116 may contain an instance of the game application 112 or 118. In one implementation, the game application 112 or 118 may enable the user to use and interact with the online experience platform 102, such as searching for games, experiences, or other content, controlling virtual characters in virtual experiences hosted by the online experience platform 102, or viewing or uploading content such as games 105, images, video items, web pages, or documents. In one example, the game application may be a web application (e.g., an application that works in conjunction with a web browser) that can access, retrieve, present, or navigate content provided by a web server (e.g., virtual characters in a virtual environment). In another example, the game application may be a native application (e.g., a mobile application, app, or game program) that is installed and runs locally on the client device 110 or 116 and enables the user to interact with the online experience platform 102. Game applications may render, display, or present content to the user (e.g., web pages, user interfaces, media viewers, audio streams). In implementation, game applications may also include embedded media players embedded within web pages.

[0071] In aspects of this disclosure, the game application 112 / 118 may be an online experience platform application for users to build, create, and edit content, upload it to the online experience platform 102, and interact with the online experience platform 102 (for example, to play a game 105 hosted by the online experience platform 102). Thus, the game application 112 / 118 may be provided to a client device 110 or 116 by the online experience platform 102. In another example, the game application 112 / 118 may be an application downloaded from a server.

[0072] In some implementations, users may log in to the online experience platform 102 via a game application. Users may access their user account by providing user account information (e.g., username and password), and the user account is associated with one or more characters available to participate in one or more games 105 on the online experience platform 102.

[0073] In general, the functions described as being performed by the online experience platform 102 may also be performed by client devices 110 or 116 or a server in other implementations, as appropriate. In addition, functions attributed to specific components may be performed by different or multiple components working together. Furthermore, the online experience platform 102 may be accessed as a service provided to other systems or devices through appropriate application programming interfaces (APIs), and is therefore not limited to use on a website.

[0074] In some implementations, the online experience platform 102 may include a spatialized audio API 106. In some implementations, the spatialized audio API 106 may be a set of computer-executable code that provides functionality to users and / or developers in the form of function calls that enable software components to communicate and / or provide / receive data. The spatialized audio API includes several defined software functions related to spatialized audio, which can be used by developers to enable spatialized audio functionality in user-generated content and may include any functions related to audio playback on the user device.

[0075] In at least one implementation, the Spatial Audio API 106 includes many functions, events, and properties that enable spatial audio. For example, the Spatial Audio API 106 may include functions that create and destroy audio channels. These functions may allow the creation of new audio channels associated with a particular server, and / or the creation of global audio channels shared among servers in the same metaverse place. The functions may also allow the deletion / destruction of previously created audio channels.

[0076] Spatialized audio API 106 may also include functionality that includes adding and removing players, as well as retrieving players associated with an audio channel. These functions may allow adding one or more players to a specific audio channel, removing players from a specific audio channel, and / or retrieving a list of players associated with an audio channel within a metaverse place. In some implementations, these functions may trigger events, such as when a player joins and / or leaves an audio channel.

[0077] The Spatialized Audio API 106 may also include functionality that includes creating audio channels not associated with a specific player. In this way, these functions may allow non-player characters, objects, and other virtual items to emit sounds used in the spatialized audio stream. For example, a speaker object that emits sound, such as representing a functioning jukebox with a specific location in the metaverse place, may be created. An avatar in the vicinity of the speaker object may then receive a spatialized audio stream containing a transformed audio stream that includes the sound from the speaker object. Sounds produced by non-player characters, objects, and other virtual items may also be incorporated into the background audio stream. This background audio stream may also include sounds produced by several (e.g., one or more) other avatars or player characters.

[0078] The Spatialized Audio API 106 may also include properties that include parameters or properties related to sound propagation. In this way, properties may include properties such as propagation medium (e.g., water, air, etc.), sound source (e.g., player or non-player sound source), volume (e.g., representing the volume of the sound source), decay distance (e.g., the distance at which sound begins to decay), maximum audible distance (e.g., if the avatar is further than this distance, this audio stream is not included in the spatialized combination), linear or logarithmic roll-off (e.g., with respect to a particular roll-off mode), playback loudness, connection state (e.g., of an audio channel), mute state (e.g., if the player or sound source is muted), and other properties.

[0079] The Spatial Audio API 106 may further include additional features, variables, properties, and / or parameters that enable the use of rich, immersive spatial audio in user-created content and / or games. The operation of the online experience platform 102 in providing spatial audio (or a combined spatial audio stream) using the Spatial Audio API 106 will be described more fully hereafter with reference to Figures 2A and 2B.

[0080] Figure 2A is a diagram of an exemplary network environment 200 (e.g., a subset of network environment 100) for providing spatialized audio chat in a virtual metaverse, as in some implementations. Network environment 200 is provided for illustrative purposes only. In some implementations, network environment 200 may contain the same, fewer, more, or different elements configured in the same or different ways as shown in Figure 2A.

[0081] As shown in Figure 2A, the online experience platform 102 may communicate with the client device 110 (for example, via a network 122, not shown) so that a user audio stream 232 is received from the client device 110 (for example, a signal 230 from the system audio input 216) and a combined spatialized audio stream 250 is provided for output at the client device 110 (for example, via the system audio output 214).

[0082] The online experience platform 102 may include a media server 202 and a data model 206 in addition to the components shown in Figure 1. The client device 110 may include an audio mixer 204, a spatialized audio manager 205, and a sound engine 260 in addition to the components shown in Figure 1.

[0083] Generally, the media server 202 is a logical server built for a specific purpose, configured to connect components of a network environment 100 and transmit audio streams (or other data) between those components. The media server 202 may, for example, facilitate real-time communication between a client device and an online game server 102 and vice versa.

[0084] The audio mixer 204 may be a software module configured to extract audio streams from multiple player or non-player objects for conversion into spatialized audio streams. The audio mixer 204 may include an audio mixer override component 208, an echo canceling component 210, and / or an audio device module override component 212.

[0085] The audio mixer override component 208 may be configured to override the underlying audio provided by the online experience platform 102 so that spatialized audio is enabled. For example, the audio mixer override component 208 may receive the user audio stream 234 as well as a copy 251 of the spatialized audio output 250, provide the individual audio streams 238 to the spatialized audio manager 205, and provide a typical audio stream 236 as an output. In this way, if the audio mixer override component 208 is not initialized, the client device 110 may function to provide normal non-spatialized audio 239. Similarly, if the audio mixer override component is initialized (for example, if spatialized audio is enabled in the online experience platform 102 or the client device 110 for a particular game 105), the individual audio streams 238 for spatialization conversion are provided to the spatialized audio manager 205.

[0086] The echo-canceling component 210 may be configured to cancel out echoes and / or other undesirable sound artifacts from the audio stream. A filtered output 232 (for example, echo-canceled based on the audio stream 236) may be provided from the echo-canceling component 210 to the media server 202. In this way, the echo-canceling component 210 may establish filters or other functions to help provide a high-quality audio stream.

[0087] The audio device module override component 212 may be configured to override and / or disable the system audio output of the client device 110 so that spatialized audio is output instead of standard system audio (for example, if spatialized audio is not enabled for the online experience platform 102 or the client device 110). The audio device module override component 212 may otherwise output normal audio 239 when spatialized audio is not enabled.

[0088] The spatialized audio manager 205 may be a software component configured to take one or more user audio streams 238 as input and convert the streams into spatialized audio streams 242 using the spatialized audio API 106 and related software functions. For example, the spatialized audio manager 205 may convert each individual user audio stream 238 into a new individual spatialized audio stream 242. Alternatively, the spatialized audio manager 205 may provide the individual user audio streams 238 to another component for spatialization conversion in the client device 110.

[0089] Furthermore, each individual user audio stream 238 may be used to augment physical and / or action commands associated with each avatar. In this way, each individual user audio stream 238 may be used to implement audio-synchronized facial animations to create a more realistic and / or immersive experience for the user. For example, when facial animations are synchronized with spatialized audio, distance attenuation is achieved while also having visible facial movements that allow the user to identify the avatar making the sound, thereby further enhancing the user experience. Similarly, individual audio streams 238 and / or spatialized audio streams 242 may be interpreted to extract emotions and / or intentions. In this way, facial animations may be extracted to further enhance the user experience.

[0090] Furthermore, each individual user audio stream 238 may be used for user moderation on the online experience platform 102. For example, since each individual audio stream is already isolated, abusive or offensive language can be more easily associated with the relevant user. Subsequently, calls to a “mute” function or a remove function from the audio channel (via API 106) may be used to effectively moderate users associated with abusive behavior. Moderation may be extensible and / or adaptable to machine learning techniques to enable automated moderation tools to analyze vocal behavior, intonation, and shouting, and / or utilize natural language processing techniques to identify abusive behavior and automatically moderate the relevant users.

[0091] Data model 206 may include multiple spatial parameters related to audio conversion. For example, data model 206 may include one or more spatial parameters representing a group of physical laws that apply to the metaverse place. These physical laws may represent or mimic real-world sound propagation environments, exaggerated real-world sound propagation environments, and / or newly defined sound propagation environments. Sound propagation parameters may be defined through the exposed spatialized audio API 106 by the developer assigning specific values ​​to the parameters. For example, different propagation media, roll-off audio parameters, distance attenuation features / parameters, and / or reflection parameters may be defined. Similarly, volume parameters, minimum / maximum audible parameters, and / or propagation parameters may be defined.

[0092] The data model 206 may further include avatar and / or scene information provided by the developer, and the actual positioning of the avatar within the metaverse place. Avatar information may include one or more of the avatar's position, velocity, or orientation within the metaverse place. Scene information may include one or more of the following that are virtually adjacent to the avatar within the metaverse place: occlusions, reverberations, virtual objects, non-player objects, openings, orifices, reflective surfaces, virtual ceilings, virtual floors, virtual corridors, virtual doorways, and / or virtual walls. Scene information may also include information related to the medium of the surrounding environment (e.g., water, air, etc.). This information / data may be separated based on each avatar within the metaverse place (e.g., as individualized space, avatar, and scene information signals 244) and provided for conversion by the spatialized audio manager 205 and / or their respective client devices 110 / 116.

[0093] Subsequently, multiple spatialized audio streams 246 may be combined into a combined spatialized audio stream 250 by the sound engine 260 (or alternatively, the spatialized audio manager 205 and / or the game engine 104) for output to the client device 110. The sound engine 260 may include any suitable sound engine, including a sound effects engine and / or a portion of the game engine 104 dedicated to audio effects. In at least one implementation, the sound engine 260 is a proprietary audio effects engine. In other implementations, the sound engine 260 may be a digital effects engine or a game engine.

[0094] Hereafter, alternative network environments 275 will be described with reference to Figure 2B. Figure 2B is a diagram of an exemplary network environment 275 (e.g., a subset of network environment 100) for providing spatialized audio chat in a virtual metaverse, as used by some implementations. Network environment 275 is provided for illustrative purposes only. In some implementations, network environment 275 may contain the same, fewer, more, or different elements configured in the same or different ways as shown in Figure 2B.

[0095] As shown in Figure 2B, the online experience platform 102 may communicate with the client device 110 (for example, via a network 122, not shown) so that a user audio stream 232 is received from the client device 110 (for example, a signal 230 from the system audio input 216) and a combined spatialized audio stream 250 is provided for output at the client device 110 (for example, via the system audio output 214). It should be noted that, for the sake of brevity, components and / or parts of the network environment 275 that are numbered the same as components and / or parts of the network environment 200 are not described again herein.

[0096] An online experience platform 102 similar to the one shown in Figure 2A may include a media server 202 and a data model 206. A client device 110 similar to the one shown in Figure 2A may include an audio mixer 204 and a spatialized audio manager 205. However, in contrast to the arrangement in Figure 2A, the spatialized audio manager 205 may provide an output 250 according to data received from the game engine 104, individual audio streams 238, and individualized space, avatar, and scene information signals 244 (e.g., combined signals 276). In this way, the spatialized audio manager 205 may include the functionality of a sound engine built into or implemented therein. Alternatively, one or more standalone sound engine components may be used to provide the spatialized audio output 250.

[0097] In this way, the spatialized audio manager 205 provides a combined spatialized audio stream 250 based on the data received from the game engine 104 and the data model 206. For example, the spatialized audio manager 205 may receive each individual user audio stream 238 as a new individual spatialized audio stream 276 based on the spatial, avatar, and scene information signals 244 (combined, for example, in an audio sync).

[0098] As described above with reference to Figures 2A and 2B, multiple spatialized audio streams 246 / 276 (or individual non-spatialized audio streams 238) may be combined to produce a combined spatialized audio stream 250 for output at a particular client device 110. Considering the potentially very large number of audio streams available in any particular virtual metaverse place or virtual environment, some implementations provide audio stream prioritization. Audio stream prioritization (sometimes called "subscription" to audio streams) allows a reduced set of prioritized streams (compared to, for example, all available streams) to be converted into spatialized audio streams. This reduced set of streams provides technical benefits and effects, including reduced computation cycles for generating the combined spatialized audio stream, reduced system resource usage, energy savings, and reduced bandwidth usage.

[0099] The prioritization of spatialized audio streams using the spatialized audio API 106 and available avatar and scene data will be explained in more detail below with reference to Figure 3.

[0100] Figure 3 is a diagram of an exemplary network environment 300 (e.g., a subset of network environment 100) for wire-ranking spatialized audio streams in a virtual metaverse, as in some implementations. Network environment 300 is provided for illustrative purposes only. In some implementations, network environment 300 may contain the same, fewer, more, or different elements configured in the same or different ways as shown in Figure 3.

[0101] As shown in Figure 3, the media server 202 and the client device 110 may communicate (for example, via a network 122 not shown) so that the exposed audio stream 330 is received by the client device 110 and a set of prioritized audio streams 350 are provided to the client device 110 for conversion, combination, and output. It should be noted that in some implementations, the prioritized audio streams 350 may be converted and combined on the online experience platform 102 rather than on the client device 110. All such modifications are within the scope of this disclosure.

[0102] In addition to the components shown in Figures 1 and 2A to 2B, the client device 110 may include a real-time communication module 302 and a subscription logic 312. In addition to the components shown in Figures 1 and 2A to 2B, the media server 202 may include a real-time communication module 306, a subscription exchange component 304, and a subscription prioritizer 310.

[0103] The real-time communication modules 302 / 306 may be software components or instances of a real-time communication server, instantiated in the client device 110 and the media server 202, respectively. In at least one implementation, the real-time communication modules 302 / 306 are instances of a WebRTC server implemented by an exposed WebRTC API (not shown). The real-time communication modules 302 / 306 are configured to pass exposed streams 330 from each connected client device and a set of prioritized audio streams 350 to each connected client device.

[0104] The subscription exchange component 304 is a software component configured to receive head-related transfer function (HRTF) and / or estimates of at least some Baum-Welch (BW) hidden Markov models 332 / 333 from the real-time communication modules 306 / 302, and to prioritize a set of audio streams into a set 331 / 351 for output to client devices.

[0105] The subscription prioritizer 310 retrieves appropriate and / or relevant spatial parameters 336 from the data model 206 and several subscription requests 334 from the spatialized audio API 106. Using the spatial parameters and subscription requests, the subscription prioritizer prioritizes the available audio streams into a prioritized set 350. The prioritization may be based on a number of factors, including, for example, the proximity of avatars in the metaverse place, the velocity of avatars in the metaverse place, the orientation of avatars in the metaverse place, virtual objects close to avatars in the metaverse place, the capabilities of the user device, bandwidth availability, the number of connections to the media server 202, the total number of users using spatialized audio, and / or user preferences related to the client device 110. Further factors may include the available bandwidth (e.g., between the media server 202 and the client device 110), the available processing power (e.g., the media server 202, the client device 110, and / or a combination thereof), the available memory (e.g., the media server 202, the client device 110, and / or a combination thereof), and other factors.

[0106] The subscription requests are based on subscription logic 312, which is configured to issue individual subscription requests 338 based on spatial parameters 340 relating to a specific avatar of a particular user. In this way, each individual client device issues a different subscription request 338 based on its associated avatar and spatial parameters 340, and other factors 342 which may include, for example, HRTF parameters, BW Hidden Markov Model parameters, and / or VAD parameters.

[0107] As described above, the prioritized set of streams may be based on multiple spatial parameters related to avatars, items, objects, and other features of the metaverse place, as well as computational resources, storage resources, bandwidth resources, and other resources. Examples of embodiments related to spatialized audio are provided below with reference to Figures 4 and 5.

[0108] Figures 4 and 5: Examples of spatial audio in a metaverse place Figures 4 and 5 illustrate exemplary virtual environments, such as metaverse place 400 and metaverse place 500, within an online virtual experience, as implemented in some cases. Metaverse place 400 includes a first avatar 402, a second avatar 404, and a third avatar 406. The avatars (402-406) may include avatars controlled by their respective users and / or avatars under the automatic control of the online experience platform 102 (e.g., computer-generated characters).

[0109] In addition to avatars 402-406, the metaverse place 400 includes a first virtual object 408, a second virtual object 410, a third virtual object 412, and a fourth virtual object 422. The virtual objects may represent, among other things, buildings, building components (e.g., walls, windows, doors), bodies of water (e.g., ponds, rivers, lakes, seas), furniture, machinery, vehicles, plants, animals, etc. The virtual objects (408-412) may include relevant data or metadata corresponding to the characteristics of one or more objects, such as material type (e.g., metal, wood, cloth, stone, etc.), the object's location within the metaverse place 400, the object's size, the object's shape, or the object's sound characteristics (e.g., ambient sound emitted by the object, how often the sound is emitted, the volume of the sound, etc.).

[0110] The sound characteristics of an object can be based on the object type, size, shape, or location of the object. For example, an object representing a large adult dog may have sound characteristics typical of an adult dog (object type) that is large (object size) and located in a given position relative to a character (object location). The sound characteristics may include ambient sounds (e.g., barking) emitted by the object, which may be provided by a sound file (e.g., computer-generated or recorded sound). Furthermore, the sound characteristics may include the frequency with which the object makes sounds (e.g., how often the dog barks) and the volume of the ambient sounds (e.g., how loud the dog's bark is at a given distance). The volume of the object's ambient sounds may then be modified as part of the audio spatialization process (e.g., the dog's bark may be louder for dogs closer to the character and quieter for dogs further away).

[0111] In the example shown in Figure 4, objects may include furniture or parts of a building. For example, virtual object 408 could be a doorway or virtual boundary between metaverse places. Virtual objects 410 and 412 could be walls. Virtual object 422 may be a small table or other virtual furniture.

[0112] Avatar 402 is capable of speaking and emitting simulated sounds, as indicated by sound paths 414 and 416. The voice of Avatar 402 along sound path 414 comes from the left of Avatar 406, and sound path 416 comes from the right of Avatar 404 through a virtual doorway 408. The sound path between Avatar 402 and Avatar 406 is mostly direct, while the sound path from Avatar 402 to Avatar 404 is partially reflected by walls 410 and 412, as indicated by sound path 416.

[0113] Furthermore, ambient sounds may be emitted by table 422. For example, a table may contain speakers or other objects on the table that emit sound. Ambient sounds are indicated by sound paths 418 and 419. In some implementations, ambient sounds may include sounds such as wind, rain, music, machinery, animals, item movement, and footsteps. Ambient sounds may also include sounds emitted from objects that may be stationary or moving within the environment (e.g., cars). Ambient sounds may also include sounds produced by other avatars moving objects in environment 400 and / or within environment 400, or interacting with objects in environment 400 and / or within environment 400.

[0114] During operation, an implementation of the audio spatialization techniques described herein may perform one or more of the following operations based on the exemplary metaverse place shown in Figure 4: 1) spatialization of voice communication of avatars (e.g., avatar 402) based on the position of each receiving avatar (e.g., 404 or 406) relative to the speaking avatar; 2) audio spatialization of voice communication based on any object in the virtual environment (e.g., 408, 410, 412, or 422); and 3) audio spatialization of ambient sounds in the virtual environment (e.g., sounds emitted by object 410).

[0115] Now, looking at Figure 5, we see a top-down view of the metaverse place 500. As shown, the simplified schematic diagrams of avatars 502, 504, 506, and 508 include a distance radius 525 measured with respect to avatar 502. In this example, avatar 502 may represent a user requesting spatialized audio, and the distance radius 525 may be a parameter or setting used to prioritize audio streams.

[0116] As further shown, avatar position and velocity data 541, 561, and 581 are represented by arrows emanating from each avatar. In this example, a user device associated with avatar 502 may receive prioritized audio streams associated with avatars 504 and 506, and ambient sounds associated with any object or non-player item within radius 525. However, if avatar 508 continues to approach radius 525 as indicated by arrow 581, data associated with avatar 508 may also be prepared for prioritization. Alternatively, if computing resources or other parameters allow, data associated with avatar 508 may also be included in the prioritized audio stream until the resources or other prioritization parameters change (for example, an additional avatar approaches avatar 502, computing resource usage increases beyond a threshold, or an additional audio stream such as background audio is given higher priority).

[0117] In this way, a subset of available audio streams is prioritized so that computing resources are reduced while rich, immersive audio continues to be effectively delivered to client devices.

[0118] A more detailed examination of the creation of spatialized audio streams and the prioritization of audio streams is provided below with reference to Figures 6 and 7.

[0119] Figure 6: Exemplary method for creating a spatialized audio stream Figure 6 is a flowchart of an exemplary method 600 for creating spatialized audio in a metaverse place, as in some implementations. In some implementations, method 600 may be implemented, for example, on a server system, such as the online experience platform 102 shown in Figure 1. In some implementations, some or all of method 600 may be implemented on a system such as one or more client devices 110 and 116 shown in Figure 1, and / or on both the server system and one or more client systems. In the examples described, the implementing system includes one or more processors or processing circuits and one or more storage devices such as a database or other accessible storage. In some implementations, one or more different components of the server and / or client may execute different blocks or other parts of method 600. Method 600 may begin in block 602.

[0120] In block 602, a request to receive audio related to a metaverse place in the virtual metaverse may be received (for example, from a first user among multiple users). For example, a client device 110 (also called a user device) may be associated with a first user. The first user is associated with a first avatar. Furthermore, multiple users may be associated with multiple avatars in the metaverse place (for example, other avatars involved in the metaverse place). Block 602 is followed by block 604.

[0121] In block 604, a data model related to the metaverse place is retrieved. For example, data model 206 may be stored in data store 108. The data model may include one or more spatial parameters representing a group (e.g., one or more) of physical laws that apply to the metaverse place. These physical laws may be exaggerated (compared to physical laws that may apply on Earth) to enhance spatial effects, or attenuated to implement spatial effects only slightly. The parameters of sound propagation and the underlying physical laws may be adjusted and / or modified through the spatialized audio API 106 described above. Block 604 is followed by block 606.

[0122] In block 606, avatar information and scene information are extracted from the data model in response to a request. For example, the request may be associated with a specific client device and, therefore, a specific avatar. Thus, the extracted avatar information includes the specific avatar and one or more of the positions, velocities, or directions of multiple avatars adjacent to that specific avatar in the metaverse place. Similarly, the scene information includes one or more of the occlusions, reverberations, virtual objects, non-player objects, openings, orifices, reflective surfaces, virtual ceilings, virtual floors, and / or virtual walls that are virtually adjacent to the specific avatar. Block 606 is followed by block 608.

[0123] In block 608, each audio stream received from each of the multiple users is transformed using extracted spatial parameters. The transformation is based on avatar and scene information. The transformation may include modifying one or more audio characteristics to create a spatialized audio stream. For example, attenuation based on a distance attenuation parameter to attenuate audio based on the distance between avatars as defined in the data model may be used to modify the audio characteristics. Similarly, rolling in or rolling out of audio (e.g., using “fading” or “rolling” effects) may be performed to modify the audio characteristics. In addition, volume increase, volume decrease, Doppler shift, reverberation, or reflection may be provided, and / or other characteristics may be modified. In this way, the transformation outputs a spatialized audio stream for each individual avatar and / or virtual object / item in the metaverse place. Block 608 is followed by block 610.

[0124] In block 610, spatialized audio streams are combined to create a combined spatialized audio stream. Combination may be performed on the online experience platform 102, on client devices 110 / 116, and / or by a combination of the online experience platform and client devices. In some implementations, combination may be performed with respect to only a prioritized set of audio streams. In other implementations, a subset of available audio streams is transformed based on proximity to avatars in the metaverse place or a threshold distance. Other modifications and limitations on the number of audio streams transformed are also possible and such modifications are within the scope of this disclosure.

[0125] According to at least one implementation, background audio streams are also combined with spatialized audio streams to create a combined spatialized audio stream. For example, a background or “special” stream that mixes multiple participants into background noise / talking may provide a more realistic experience. For instance, in a room with 50 people talking, an avatar might be conversing with several avatars that are very close by. The audio from the closest participants may be the clearest, but there is also background talk around the avatar (e.g., pure silence would be unrealistic). Therefore, the background stream may be pre-mixed in a relatively simple way to include the overall background talk from the remaining 50 avatars (e.g., a uniform background stream from all participants used by the combined stream to all participants). In this way, background audio may be generated based on one or more of the following: audio received from other users different from the first user, audio generated based on the avatar’s movement within the metaverse place, and / or overall or “special” background streams. Block 610 is followed by Block 612.

[0126] In block 612, the combined spatialized audio stream is provided to the user device for output through an audio output device connected to the user device, such as a set of speakers or headphones. In some implementations, the spatialized audio stream may be provided via an audio output device such as a virtual reality headset, augmented reality headset, or head-mounted device.

[0127] Blocks 602-612 can be executed (or repeated) in a different order than described above, and / or one or more blocks can be omitted. For example, data extraction (blocks 604-606) may be performed independently of audio conversion and combination (blocks 608-612). Furthermore, request reception, related data extraction, and audio conversion may be performed in parallel or by different components in some implementations.

[0128] A more detailed examination of stream prioritization is provided below with reference to Figure 7.

[0129] Figure 7: Exemplary method for prioritizing spatialized audio streams Figure 7 is a flowchart of an exemplary method 700 for providing content to users based on classification, as in some implementations. In some implementations, method 700 may be implemented, for example, on a server system, such as the online experience platform 102 shown in Figure 1. In some implementations, some or all of method 700 may be implemented on a system such as one or more client devices 110 and 116 shown in Figure 1, and / or on both the server system and one or more client systems. In the examples described, the implementing system includes one or more processors or processing circuits and one or more storage devices such as a database or other accessible storage. In some implementations, one or more different components of the server and / or client may execute different blocks or other parts of method 700. Method 700 may begin in block 702.

[0130] In block 702, a request to receive audio related to a metaverse place in the virtual metaverse may be received (for example, from a first user among multiple users). For example, client device 110 may be associated with a first user. The first user is associated with a first avatar. Furthermore, multiple users may be associated with multiple avatars in the metaverse place (for example, avatars involved in the metaverse place). Block 702 is followed by block 704.

[0131] In block 704, a set of prioritized audio streams is determined for the first user. The set of prioritized audio streams is a ranked subset of all audio streams received from each of the multiple users in the metaverse place, as well as all other audio streams (e.g., ambient sounds, non-player character / item sounds, etc.). Prioritization may be based, for example, on threshold distances and / or threshold radii (e.g., shown in Figure 5).

[0132] Prioritization may also be based on the proximity of avatars within the metaverse, their velocity, their orientation, virtual objects adjacent to the avatars, the capabilities of the user device, or the user preferences of the first user. For example, prioritization may take into account whether an avatar is moving toward a target avatar, whether an avatar is facing a target avatar, and other similar prioritization parameters.

[0133] Prioritization may also be based on the processing resources and / or capabilities of the client device. For example, prioritization may take into account memory usage, disk usage, bandwidth availability, processor usage, and other resource usage to determine whether there are sufficient resources to handle a certain number of spatialized audio streams. Prioritization may then prioritize streams so that a certain number of streams are prioritized to avoid contention or further resource utilization exceeding a threshold.

[0134] Prioritization may also be based on the processing resources and / or capabilities of the media server. For example, prioritization may take into account memory usage, storage usage, bandwidth availability, processor usage, number of active connections, number of inactive connections, total number of users, and other resource usage to determine whether there are sufficient resources to handle a particular number of spatialized audio streams. Prioritization may then prioritize streams so that a certain number of streams are prioritized to avoid contention or further resource utilization exceeding a threshold.

[0135] Prioritization may also be based on the processing resources and / or capabilities of the online experience platform. For example, prioritization may take into account memory usage, storage usage, bandwidth availability, processor usage, number of active connections, number of inactive connections, total number of users, number of active online experiences, and other resource usage to determine whether there are sufficient resources to handle a particular number of spatialized audio streams. Prioritization may then prioritize streams so that a certain number of streams are prioritized to avoid contention or further resource utilization exceeding a threshold.

[0136] Other changes to the basis for prioritization are also possible and within the scope of this disclosure. Block 704 is followed by Block 706.

[0137] In block 706, each audio stream of the prioritized audio streams is transformed using extracted spatial parameters. The transformation is based on avatar and scene information. The transformation may include modifying one or more audio characteristics to create a spatialized audio stream. For example, attenuation based on a distance attenuation parameter to attenuate audio based on the distance between avatars as defined in the data model may be used to modify the audio characteristics. Similarly, audio roll-in or roll-out may be performed to modify the audio characteristics. In addition, volume rise, volume fall, Doppler shift, reverberation, reflections, and other characteristics may be modified. In this way, the transformation outputs a spatialized audio stream for each individual avatar and / or virtual object / item associated with the prioritized audio streams. Block 706 is followed by block 708.

[0138] In block 708, spatialized audio streams are combined to create a combined spatialized audio stream. Combination may be performed on the online experience platform 102, on client devices 110 / 116, and / or by a combination of the online experience platform and client devices. As described above, special or background audio streams may also be combined to provide environment and / or background audio to the spatialized audio output. Combination may be performed with respect to only a prioritized set of audio streams. Block 708 is followed by block 710.

[0139] In block 710, the combined spatialized audio stream is provided to the user device for output through an audio output device connected to the user device, such as a set of speakers or headphones.

[0140] Blocks 702–710 can be executed (or repeated) in a different order than described above, and / or one or more blocks can be omitted. Methods 600 and / or 700 can be executed on a server (e.g., 102) and / or client devices (e.g., 110 or 116). Furthermore, parts of methods 600 and 700 may be combined and executed sequentially or in parallel, according to any desired implementation.

[0141] As described above, systems, methods, and computer-readable media may provide spatialized audio in virtual experiences. By providing a robust spatialized audio API, developers may be able to use spatialized audio in virtually any virtual experience they create. The immersive quality of a typical virtual experience enhanced by a spatialized audio API provides a richer user experience, increases user engagement, provides intuitive feedback (e.g., through audio positioning), and significantly reduces the complexity of implementing spatialized audio.

[0142] A more detailed description of various computing devices that may be used to implement the different devices shown in Figures 1 to 3 is provided below with reference to Figure 8.

[0143] Figure 8 is a block diagram of an exemplary computing device 800 that may be used to implement one or more of the features described herein, in some implementations. In one example, device 800 may be used to implement a computer device (e.g., 102, 110, and / or 116 in Figure 1) and to perform an implementation of a suitable method described herein. Computing device 800 can be any suitable computer system, server, or other electronic or hardware device. For example, computing device 800 can be a mainframe computer, desktop computer, workstation, portable computer, or electronic device (portable device, mobile device, cell phone, smartphone, tablet computer, television, TV set-top box, personal digital assistant (PDA), media player, game device, wearable device, etc.). In some implementations, device 800 includes a processor 802, memory 804, an input / output (I / O) interface 806, and audio / video input / output devices 814 (e.g., a display screen, touchscreen, display goggles or glasses, audio speakers, headphones, microphone, etc.).

[0144] Processor 802 can be one or more processors and / or processing circuits for executing program code and controlling the basic operation of device 800. “Processor” includes any suitable hardware and / or software system, mechanism, or component for processing data, signals, or other information. A processor may include a general-purpose central processing unit (CPU), multiple processing units, a system with dedicated circuits for implementing functions, or other systems. Processing is not necessarily limited to a specific geographical location or subject to temporal constraints. For example, a processor may perform its functions in “real-time,” “offline,” “batch mode,” etc. Parts of the processing may be performed at different times and in different locations by different (or the same) processing systems. A computer may be any processor communicating with memory.

[0145] Memory 804 is generally located within device 800 for access by processor 802 and is suitable for storing instructions for execution by the processor. It may be any suitable processor-readable storage medium located separately from and / or integrated with processor 802, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc. Memory 804 can store software that runs on server device 800 by processor 802, including operating system 808, application 810, and associated data 812. In some implementations, application 810 may include instructions that enable processor 802 to perform some or all of the functions described herein, such as the methods shown in Figures 6 and 7.

[0146] For example, memory 804 may contain software instructions for prioritizing and / or providing spatialized audio within an online experience platform (e.g., 102) or metaverse place. Any software in memory 804 may, alternatively, be stored in any other suitable storage location or computer-readable medium. Furthermore, memory 804 (and / or other connected storage devices) may store instructions and data used in the features described herein. Memory 804 and any other type of storage (magnetic disks, optical disks, magnetic tapes, or other tangible media) may be considered “storage” or “storage devices”.

[0147] The I / O interface 806 can provide functionality to enable the server device 800 to interface with other systems and devices. For example, network communication devices, storage devices (e.g., memory and / or datastore 108), and input / output devices can communicate via interface 806. In some implementations, the I / O interface may connect to interface devices including input devices (such as keyboards, pointing devices, touchscreens, microphones, cameras, scanners, etc.) and / or output devices (such as display devices, speaker devices, printers, monitors, etc.).

[0148] For ease of illustration, Figure 8 shows one block for each of the following: processor 802, memory 804, I / O interface 806, software blocks 808 and 810, and database 812. These blocks may represent one or more processors or processing circuits, operating systems, memory, I / O interfaces, applications, and / or software modules. In other implementations, device 800 may not have all of the components shown and / or may have other components, including other types of elements, in place of or in addition to the elements shown herein. Although the online experience platform 102 is described as performing the operation as described in some implementations herein, any preferred component or combination of components of the online experience platform 102 or a similar system, or any preferred one or more processors associated with such a system, may perform the operation described.

[0149] A user device may implement and / or be used in conjunction with the features described herein. An exemplary user device may be a computer device including several components similar to device 800, for example, a processor 802, memory 804, and an I / O interface 806. A suitable operating system, software, and applications for the client device may be provided in memory and used by the processor. The I / O interface for the client device may be connected to network communication devices as well as input and output devices, for example, a microphone for capturing sound, a camera for capturing images or video, an audio speaker device for outputting sound, a display device for outputting images or video, or other output devices. For example, a display device in the audio / video input / output device 814 may be connected to (or included in) device 800 to display pre- and post-processing of images as described herein, and such a display device may include any suitable display device, for example, an LCD, LED, or plasma display screen, a CRT, a television, a monitor, a touchscreen, a 3-D display screen, a projector, or other visual display device. Some implementations may provide audio output devices, such as text-to-speech or speech synthesis.

[0150] The methods, blocks, and / or operations described herein may be executed in a different order than shown or described, and / or concurrently (partially or completely) with other blocks or operations. Some blocks or operations may be executed for one part of the data and then executed again later, for example, for another part of the data. Not all described blocks and operations may be executed in all implementations. In some implementations, blocks and operations may be executed multiple times, in different orders, and / or at different times within a method.

[0151] In some implementations, some or all of the methods may be performed on a system such as one or more client devices. In some implementations, one or more of the methods described herein may be performed, for example, on a server system and / or on both a server system and a client system. In some implementations, one or more different components of a server and / or client may perform different blocks, operations, or other parts of the method.

[0152] One or more methods described herein (for example, methods 600 and / or 700) can be implemented by computer program instructions or code that can be executed on a computer. For example, the code can be implemented by one or more digital processors (for example, microprocessors or other processing circuits) and can be stored in computer program products including non-temporary computer-readable media (for example, storage media), such as magnetic, optical, electromagnetic, or semiconductor storage media, including semiconductor or solid-state memory, magnetic tape, removable computer diskettes, random-access memory (RAM), read-only memory (ROM), flash memory, hard magnetic disks, optical disks, solid-state memory drives, etc. Program instructions can also be contained in electronic signals and provided as electronic signals, for example, in the form of software as a service (SaaS) delivered from a server (for example, a distributed system and / or a cloud computing system). Alternatively, one or more methods can be implemented in hardware (such as logic gates) or in a combination of hardware and software. Exemplary hardware may include programmable processors (e.g., field-programmable gate arrays (FPGAs), complex programmable logic devices), general-purpose processors, graphics processors, and application-specific integrated circuits (ASICs). One or more methods may run as part of or as a component of an application running on the system, or as an application or software running in conjunction with other applications and operating systems.

[0153] One or more methods described herein may be performed as standalone programs that can run on any type of computing device, programs that run on a web browser, or within mobile applications ("apps") that run on mobile computing devices (e.g., cell phones, smartphones, tablet computers, wearable devices (watches, armbands, jewelry, hats, goggles, glasses, etc.), laptop computers, etc.). In one example, a client / server architecture may be used, for example, where a mobile computing device (as a client device) sends user input data to a server device and receives final output data (e.g., for display) from the server. In another example, all computations may be performed within a mobile app (and / or other app) on a mobile computing device. In yet another example, computations may be divided between a mobile computing device and one or more server devices.

[0154] While the explanations are given in relation to specific implementations, these specific implementations are illustrative and not limiting. Concepts shown in the examples may apply to other examples and implementations.

[0155] Where a particular implementation considered herein may acquire or use user data (e.g., user demographics, user behavior data on the platform, user search history, purchased and / or viewed items, user social relationships on the platform, etc.), the user will be provided with choices to control whether, and how, such information is collected, stored, or used. That is, the implementation considered herein will collect, store, and / or use user information only with the explicit authorization of the user and in compliance with applicable regulations.

[0156] Users are enabled to control whether a program or feature collects user information about that particular user or other users associated with that program or feature. Each user from whom information should be collected is presented with options (e.g., through the user interface) that allow the user to exercise control over information collection related to that user, granting permission or authorization regarding whether information is collected and which parts of the information should be collected. Furthermore, certain data may be modified in one or more ways before being stored or used so that personally identifiable information is removed. For example, a user's identity may be modified (e.g., by substitution using pseudonyms, numbers, etc.) so that personally identifiable information cannot be determined. Another example is a user's geographical location being generalized to a larger area (e.g., city, zip code, state, country, etc.).

[0157] It should be noted that the functional blocks, operations, features, methods, devices, and systems described herein may be integrated into or separated into different combinations of systems, devices, and functional blocks, as known to those skilled in the art. Any preferred programming language and programming technique may be used to implement routines in a particular implementation. Different programming techniques, such as procedural or object-oriented programming techniques, may be used. Routines may be executed on a single processing device or on multiple processors. Steps, operations, or calculations may be presented in a specific order, but the order may be modified in different particular implementations. In some implementations, multiple steps or operations shown herein as sequential may be executed simultaneously. [Explanation of symbols]

[0158] 100 Network Environment 102 Online Experience Platforms 104 Game engines, virtual experience engines 105 Games, Virtual Experiences, Game Applications 106 Spatial Audio API, Spatialized Audio API 108 Datastores 110 / 116 Client Devices 110 First client device 112 Game applications, virtual experience applications 114 users 116 Second client device 118 Game applications, virtual experience applications 120 users 122 Network 200 Network Environment 202 Media Server 204 Audio Mixer 205 Spatial Audio Manager 206 Data Models 208 Audio Mixer Override Components 210 Echo Canceling Components 212 Audio Device Module Override Components 214 System Audio Output 216 System Audio Inputs 230 signal 232 user audio streams, filtered output 234 User Audio Streams 236 Typical audio streams 238 individual audio streams 239 Normal non-spatial audio, non-spatial audio stream 242 Spatialized Audio Streams 244 Space, Avatar, and Scene Information Signals 246 Spatialized Audio Streams 250 combined spatialized audio streams, spatialized audio outputs 251 Copy of spatialized audio output 260 Sound Engine 275 Alternative network environments 276 Combined signals 300 Network Environment 302 Real-time communication module 304 Subscription Exchange Components 306 Real-time communication module 310 Subscription Prioritizer 312 Subscription Logic 330 Published audio streams 331 sets 332 Head-Related Transfer Function (HRTF) and / or Estimates of At Least Parts of the Baum-Welch (BW) Hidden Markov Model 333 Head-Related Transfer Function (HRTF) and / or Estimates of At Least Parts of the Baum-Welch (BW) Hidden Markov Model 334 Subscription Request 336 Spatial Parameters 338 Subscription Request 340 spatial parameters 342 Other factors 350 prioritized audio streams, prioritized sets 351 sets 400 Metaverse Places 402 First Avatar 404 Second Avatar 406 The Third Avatar 408 First virtual object, virtual entrance 410 Second virtual object, wall 412 Third virtual object, wall 414 The Path of Sound 416 Sound Paths 418 Sound Paths 419 Sound Paths 422 The fourth virtual object, table 500 Metaverse Places 502 Avatar 504 Avatar 506 Avatars 508 Avatars 525 distance radius 541 Avatar position and velocity data 561 Avatar position and speed data 581 Avatar position and speed data 600 ways 700 methods 800 Computing Devices 802 Processor 804 memory 806 Input / Output (I / O) Interface 808 Operating System 810 Applications 812 Related data, databases 814 Audio / Video Input / Output Devices

Claims

1. A computer-based method for spatialized audio in a virtual metaverse, A step of receiving a request from a first user among multiple users to receive audio related to a metaverse place of the virtual metaverse, wherein the first user is associated with a user device, and each of the multiple users is associated with each of the multiple avatars in the metaverse place; A step of retrieving a data model related to the metaverse place, wherein the data model includes one or more spatial parameters representing one or more physical laws applied to the metaverse place; A step of extracting avatar information and scene information from the data model, wherein the avatar information includes one or more of the positions, velocities, or directions of the plurality of avatars in the metaverse place, including a first avatar associated with the first user, and the scene information includes one or more of the occlusions, reverberations, or virtual walls that are virtually close to the first avatar in the metaverse place. A step of determining a set of prioritized audio streams received from each of the aforementioned users, wherein audio streams associated with avatars moving toward the receiving avatar are given priority over audio streams associated with avatars moving away from the receiving avatar. To create a spatialized audio stream, the process involves transforming a prioritized set of audio streams from the audio streams received from each of the multiple users based on the avatar information and the scene information, and transforming at least one audio characteristic of each of the audio streams based on one or more spatial parameters. The steps include combining the spatialized audio streams to create a combined spatialized audio stream, The steps include providing the combined spatialized audio stream to the user device. Methods performed by computers, including those mentioned above.

2. The computer-based method according to claim 1, wherein the one or more spatial parameters include distance attenuation parameters for attenuating audio based on the distance between avatars.

3. The method performed by a computer according to claim 1, wherein each of the audio streams received from each of the plurality of users includes mono audio received by a microphone device, and the combined spatialized audio stream includes stereo audio.

4. The computer-based method according to claim 3, wherein the combined spatialized audio stream includes stereo audio generated by placing each user's mono audio at the location of each user's avatar.

5. The combined spatialized audio stream includes spatial audio and background audio based on the audio stream received from a user among the plurality of users other than the first user, and the background audio is Audio received from other users different from the first user, and Audio generated based on the movement of avatars within the aforementioned metaverse place A computer-based method according to claim 1, generated based on one or more of the following.

6. The step of converting the aforementioned prioritized set of audio streams is: A computer-based method according to claim 1, comprising the step of prioritizing audio streams received from each of the multiple users based on the speed of each avatar in the metaverse place.

7. The step of determining the set of prioritized audio streams is: A method implemented by a computer according to claim 6, comprising prioritizing audio streams received from each of the multiple users based on one or more of the proximity of avatars in the metaverse place, the orientation of avatars in the metaverse place, virtual objects adjacent to avatars in the metaverse place, the capabilities of the user device, or the user preferences of the first user.

8. The computer-based method according to claim 1, wherein audio streams associated with avatars closer to the receiving avatar are given priority over audio streams associated with avatars further away from the receiving avatar, and audio streams associated with avatars directed toward the receiving avatar are given priority over audio streams associated with avatars directed away from the receiving avatar.

9. A method implemented by a computer that provides spatialized audio in a virtual metaverse, A step of receiving a request from a first user among multiple users to receive audio related to a metaverse place of the virtual metaverse, wherein the first user is associated with a user device, and each of the multiple users is associated with each of the multiple avatars in the metaverse place; A step of determining a set of prioritized audio streams received from each of the aforementioned users, wherein audio streams associated with an avatar moving toward the receiving avatar are given priority over audio streams associated with an avatar moving away from the receiving avatar. To create a spatialized audio stream, the steps include transforming the aforementioned set of prioritized audio streams, The steps include combining the spatialized audio streams to create a combined spatialized audio stream, The steps include providing the combined spatialized audio stream to the user device. Methods performed by computers, including those mentioned above.

10. The step of determining the set of prioritized audio streams is: A method implemented by a computer according to claim 9, comprising prioritizing audio streams received from each of the plurality of users based on one or more of the proximity of avatars in the metaverse place, the velocity of avatars in the metaverse place, the orientation of avatars in the metaverse place, virtual objects adjacent to avatars in the metaverse place, the capabilities of the user device, or the user preferences of the first user.

11. The computer-based method according to claim 9, wherein audio streams associated with avatars closer to the receiving avatar are given priority over audio streams associated with avatars further away from the receiving avatar, and audio streams associated with avatars directed toward the receiving avatar are given priority over audio streams associated with avatars directed away from the receiving avatar.

12. It is a system, The memory where the instructions are stored, A processing device coupled to the memory, configured to access the memory, and when the instruction is executed by the processing device, the processing device is configured to access the memory. Receiving a request from a first user among multiple users to receive audio related to a metaverse place in a virtual metaverse, wherein the first user is associated with a user device, and each of the multiple users is associated with each of the multiple avatars in the metaverse place, The process involves extracting a data model related to the metaverse place, wherein the data model includes one or more spatial parameters representing one or more physical laws applied to the metaverse place. Extracting avatar information and scene information from the data model, wherein the avatar information includes one or more of the positions, velocities, or directions of the plurality of avatars in the metaverse place, including a first avatar associated with the first user, and the scene information includes one or more of the occlusions, reverberations, or virtual walls that are virtually close to the first avatar in the metaverse place. Determining a set of prioritized audio streams received from each of the aforementioned users, wherein audio streams associated with avatars moving toward the receiving avatar are given priority over audio streams associated with avatars moving away from the receiving avatar. To create a spatialized audio stream, the system transforms a set of prioritized audio streams from the audio streams received from each of the multiple users based on the avatar information and the scene information, and transforms at least one audio characteristic of each of the audio streams based on one or more spatial parameters. Combining the spatialized audio streams to create a combined spatialized audio stream, To provide the aforementioned combined spatialized audio stream to the user device. A system that performs actions including those mentioned above.

13. The system according to claim 12, wherein the one or more spatial parameters include a distance attenuation parameter for attenuating audio based on the distance between avatars.

14. The system according to claim 12, wherein each of the audio streams received from each of the plurality of users includes mono audio received by a microphone device, and the combined spatialized audio stream includes stereo audio.

15. The system according to claim 14, wherein the combined spatialized audio stream includes stereo audio generated by placing each user's mono audio at the position of each user's avatar.

16. The combined spatialized audio stream includes spatial audio based on the audio stream received from a user among the plurality of users other than the first user, and background audio, wherein the background audio is Audio received from other users different from the first user, and Audio generated based on the movement of avatars within the aforementioned metaverse place The system according to claim 12, which is generated based on one or more of the following.

17. The aforementioned operation, The system according to claim 12, further comprising prioritizing audio streams received from each of the multiple users based on the speed of each avatar in the metaverse place.

18. Determining the set of prioritized audio streams is The system according to claim 17, comprising prioritizing audio streams received from each of the plurality of users based on one or more of the proximity of avatars in the metaverse place, the orientation of avatars in the metaverse place, virtual objects adjacent to avatars in the metaverse place, the capabilities of the user device, or the user preferences of the first user.

19. The system according to claim 18, wherein audio streams associated with avatars closer to the receiving avatar are given priority over audio streams associated with avatars further away from the receiving avatar, and audio streams associated with avatars directed toward the receiving avatar are given priority over audio streams associated with avatars directed away from the receiving avatar.

20. A spatialized audio manager configured to convert the respective audio streams received from each of the multiple users, Before providing the combined spatialized audio stream to the user device, an audio device override module configured to disable non-spatialized audio on the user device and The system according to claim 12, further comprising:

Citation Information

Patent Citations

  • Network system

    JP2007336481A

  • On-line conversation system, on-line conversation server, on-line conversation control method, and program

    JP2010122826A

  • Web-based videoconference virtual environment with navigable avatars, and applications thereof

    US10979672B1

  • Virtual meeting rooms with spatial audio

    US7346654B1