Distributed processing of sound in virtual environment
By distributing and simulating sounds in a virtual environment and allocating processing tasks with priority values, the problem of insufficient computing power of user equipment is solved, and more accurate sound reproduction and rich user experience are achieved.
Patent Information
- Application Number
- CN202380074347.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-02
- Filing Date
- 2023-10-26
- Publication Date
- 2025-05-30
AI Technical Summary
In a virtual environment, insufficient computing power of the user equipment leads to the inability to simulate a large amount of sound at the same time, resulting in a reduced user experience.
By distributing sound processing and simulations between user equipment and cloud-based computing devices, the processing tasks for each sound are assigned using priority values to ensure that high-priority sounds are generated locally or on remote servers.
It realizes providing more accurate sound reproduction in a virtual environment, supports the virtual experience of a large number of users, and improves the richness and synchronization of the user experience.
Smart Images

Figure CN120077683A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims priority to U.S. Patent Application No. 17 / 979,416, filed on November 2, 2022, with the title "DISTRIBUTED PROCESSING OF SOUNDS IN VIRTUAL ENVIRONMENTS", the entire content of which is incorporated herein by reference. Technical Field
[0002] Embodiments generally relate to computer - based virtual experiences, and more particularly, to methods, systems, and computer - readable media for providing audio for a virtual environment. Background Art
[0003] Some online virtual experience platforms allow users to connect with each other, interact with each other (e.g., within a virtual experience), create virtual experiences, and share information with each other over the Internet. Users of online virtual experience platforms can participate in (e.g., in a multi - player game environment in a virtual three - dimensional environment), design customized environments, design characters and avatars, design, simulate, or create sounds used within the environment, decorate avatars, exchange virtual items / objects with other users, communicate with other users using audio or text messages, etc.
[0004] In view of the above, some embodiments are contemplated. Summary of the Invention
[0005] A system of one or more computers can be used to perform certain operations or actions because the system is installed with software, firmware, hardware, or a combination thereof, which, when running, causes the system to perform the above - mentioned actions. One or more computer programs can be used to perform certain operations or actions because these programs include instructions that, when executed by a data - processing device, cause the device to perform the above - mentioned actions. One general aspect includes computer - implemented. The computer - implemented method further includes: receiving, at a server, a first request to generate a plurality of sounds for a user device, where the user device is associated with a virtual experience hosted by the server; the server obtaining sound source data of a plurality of sound sources in the virtual experience, each sound source being associated with a specific sound among the plurality of sounds; the server obtaining virtual experience state information, the virtual experience state information including the position of a virtual microphone in the virtual experience and at least one of the following: the speed of the virtual microphone in the virtual experience or the orientation of the virtual microphone in the virtual experience; the server generating an audio mix of the plurality of sounds based on the sound source data and the virtual experience state information; and transmitting the audio mix to the user device. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the above - mentioned method.
[0006] Embodiments may include one or more of the following features. In a computer-implemented method, transmitting an audio mix to a user device may include providing an audio mix encoded in a streaming audio format. The first request also includes a priority value for at least one sound source among a plurality of sound sources. The priority value is based on one or more of the following: the loudness of at least one sound source and the distance of at least one sound source from a virtual microphone in a virtual experience. Obtaining virtual experience state information may also include obtaining the head orientation of a user associated with the user device. Generating an audio mix of multiple sounds may include: for each of the plurality of sound sources: generating an audio segment of the sound source based on the corresponding sound source data; and applying to the audio segment at least one of the following: loudness adjustment based on the distance of the sound source from a virtual microphone in the virtual experience, or Doppler adjustment based on the speed of the virtual microphone in the virtual experience; and after the application, combining the audio segments of the multiple sounds to generate an audio mix. Generating an audio mix of multiple sounds may include: for each sound source: applying to the generated audio segment at least one of the following: a second loudness adjustment based on the distance of the sound source from a second virtual microphone in the virtual experience; and a second Doppler adjustment based on the speed of the second virtual microphone in the virtual experience; and after applying at least one of the second loudness adjustment and the second Doppler adjustment, combining the audio segments of the multiple sounds to generate a second audio mix. The plurality of sound sources includes at least one narrative sound source and at least one non-narrative sound source. Generating an audio mix of multiple sounds may include: generating a first set of multiple sounds at a server; and transmitting a request to a second server to generate a second set of multiple sounds, where the first set and the second set are mutually exclusive. Generating the first set of sounds may include generating one or more sounds of a sound source associated with a priority value that meets a predetermined priority value threshold. The computer-implemented method may include: receiving at the second server a request to generate a second set of multiple sounds; generating a portion of the second set of multiple sounds at the second server; and transmitting a request to a third server to generate a third set of multiple sounds. Obtaining the position of the virtual microphone may include: obtaining the position of a virtual camera placed within the virtual experience; and determining the position of the virtual microphone based on the position of the virtual camera. The position of the virtual microphone may include: obtaining the position of an avatar within the virtual experience; and determining the position of the virtual microphone based on the position of the avatar. Obtaining the position of the virtual microphone may include: obtaining the position of a virtual camera placed within the virtual experience; obtaining the position of an avatar within the virtual experience; and determining the position of the virtual microphone based on the position of the virtual camera. Determining the position of the virtual microphone such that the virtual microphone is equidistant from the position of the virtual camera and the position of the avatar. Embodiments of the technology may include hardware, methods or processes, or computer software on a computer-accessible medium.
[0007] The non - transitory computer - readable medium further includes: the server receives a first request to generate a plurality of sounds for a user device, where the user device is associated with an avatar participating in a virtual experience hosted by the server; the server obtains sound source data of a plurality of sound sources associated with the plurality of sounds; the server obtains virtual experience status information, which may include the position of a virtual microphone in the virtual experience and at least one of the following: the speed of the virtual microphone in the virtual experience or the orientation of the virtual microphone in the virtual experience; the server generates an audio mix of the plurality of sounds based on the sound source data and the virtual experience status information; and transmits the audio mix to the user device. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the above - described method.
[0008] The implementation may include one or more of the following features. In the non - transitory computer - readable medium, transmitting the audio mix to the user device may include providing the audio mix encoded in a streaming audio format. The first request further includes a priority value for at least one of the plurality of sound sources. The priority value is based on one or more of the following: the loudness of at least one sound source and the distance between at least one sound source and the virtual microphone in the virtual experience. Embodiments of the technology may include hardware, a method or process, or computer software on a computer - accessible medium.
[0009] The system further includes a memory storing instructions thereon; and other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the above - described method.
[0010] The implementation may include one or more of the following features. In the system, obtaining the virtual experience status information further includes obtaining the head orientation of a user associated with the user device. Generating the audio mix of the plurality of sounds may include: for each of the plurality of sound sources: generating an audio segment of the sound source based on the corresponding sound source data; and applying at least one of the following to the audio segment: loudness adjustment based on the distance between the sound source and the virtual microphone in the virtual experience; and Doppler adjustment based on the speed of the virtual microphone in the virtual experience; and after the application, combining the audio segments of the plurality of sounds to generate the audio mix. The plurality of sound sources includes at least one plot sound source and at least one non - plot sound source. Generating the audio mix of the plurality of sounds may include: generating a first set of the plurality of sounds at the server; and transmitting a request to a second server to generate a second set of the plurality of sounds, where the first set and the second set are mutually exclusive. Embodiments of the technology may include hardware, a method or process, or computer software on a computer - accessible medium. Description of the Drawings
[0011] Figure 1 FIG. is a diagram of an example system architecture for distributed processing of sound in a virtual environment according to some embodiments.
[0012] Figure 2A FIG. illustrates an example embodiment of a system architecture for generation of simulated sound in a virtual environment according to some embodiments.
[0013] Figure 2B FIG. illustrates an example embodiment of a system architecture for generation of simulated sound in a virtual environment including a hierarchical arrangement of sound servers according to some embodiments.
[0014] Figure 3 FIG. is a diagram showing an example scenario within a virtual environment using simulated sound according to some embodiments.
[0015] Figure 4 FIG. is a flowchart showing an example method of providing an encoded sound mix to a user device according to some embodiments.
[0016] Figure 5 FIG. is a flowchart showing an example method of generating an audio mix of multiple sounds according to some embodiments.
[0017] Figure 6 FIG. is a block diagram showing an example computing device according to some embodiments. DETAILED DESCRIPTION
[0018] In the following detailed description, reference is made to the accompanying drawings which form a part hereof. In the drawings, like reference numerals generally identify like components unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein. As generally described herein and shown in the drawings, aspects of the present disclosure may be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are contemplated herein.
[0019] References in the specification to “some embodiments,” “embodiments,” “example embodiments,” etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, these phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, whether or not explicitly described, the above feature, structure, or characteristic may be implemented in connection with other embodiments.
[0020] An online virtual experience platform (also referred to as a "user-generated content platform" or "user-generated content system") provides users with various ways to interact with each other. For example, users of the online virtual experience platform can cooperate towards a common goal, share various virtual experience projects, send electronic messages to each other, etc. Users of the online virtual experience platform can join a virtual experience, such as a game or other experience, as a virtual character and play a specific role. For example, the virtual character can be part of a team or multiplayer game environment, where each character is assigned a certain role and has associated parameters corresponding to that role, such as clothing, armor, weapons, skills, etc. In another example, for instance, when a single player is part of a game, the virtual character can be joined by a computer-generated character.
[0021] The online virtual experience platform can also enable users to experience sounds from the virtual environment. For example, sounds can be generated to simulate the footsteps of an avatar walking around in the virtual environment, sounds can be generated to simulate the sound of a waterfall that is part of the virtual environment, and sounds can also be generated to mimic the sounds of people in a stadium. The generated sounds can include the sounds emitted by various objects in the virtual environment, and can also include other sounds that are not particularly associated with a specific object, such as thunder, etc.
[0022] Many online virtual experiences take place in a simulated three-dimensional reality, a virtual world. The virtual experience includes simulated sounds generated by the environment and each player. For example: the footsteps of an avatar (character), the rustling sounds of the clothes and accessories worn by the avatar, the noises made by the objects used by the avatar, explosion sounds, the roars of monsters, the rumbling sounds of a building collapsing, etc.
[0023] In some scenarios, if the virtual environment includes a large number of participants (players) and / or objects, the number of sounds to be generated may be so large that the player's (user's) computing device cannot generate them while remaining synchronized with the activities within the virtual environment. This situation may occur because the computing power of the user device is relatively low, and the user device may be performing multiple tasks, such as graphics processing, processing user input received from one or more user interfaces, performing physical updates on the objects and / or characters in the virtual environment, etc.
[0024] In some cases, the number of simulated sounds may exceed the computing power of the local user device. In such a case, the online virtual experience platform may only be able to simulate a subset of the sounds to be played, thus reducing the user experience.
[0025] In such a scenario, the virtual experience platform and / or the user device may perform "voice stealing" and play only the highest-priority sounds, e.g., only the N loudest sounds. This can lead to an insufficiently rich voice experience for the user, thereby having a negative impact on the user experience. One technical problem faced by virtual application platform operators is to simulate multiple sounds simultaneously in a virtual environment.
[0026] The present disclosure addresses the above drawbacks by using a cloud-based sound server in combination with a user device. According to the technology of the present disclosure, the processing and / or simulation of sounds is distributed between one or more user devices and cloud-based computing devices (e.g., servers, such as physical servers, or virtual machines configured to run on physical servers). A priority value for each sound source to be simulated can be determined, and the processing of each sound can be assigned to a suitable computing device or process based on the priority value.
[0027] In some embodiments, the priority value can be determined by an appropriate combination of the importance of a game, a storyline, the loudness of a sound, etc. In some embodiments, the priority value can be provided by a user (e.g., a developer of a virtual experience).
[0028] The technology of the present disclosure enables a user device to receive an audio mix including a large number of sounds and play back such a sound mix, thereby providing a more accurate sound reproduction for a virtual environment. A virtual experience platform can utilize the disclosed technology to support virtual experiences with hundreds of thousands of users (e.g., game players, participants in a lecture and / or a concert, etc.), all of whom may provide sounds that can be heard by a large number of other users (e.g., game players).
[0029] Figure 1 is a diagram of an example system architecture for distributed processing of sounds in a virtual environment according to some embodiments. Figure 1 The same reference numerals are used in this and other drawings to identify the same elements. Characters following a reference numeral, e.g., "110", indicate the element in the text specifically referred to having that particular reference numeral. A reference numeral without a character in the text, e.g., "110", refers to any or all of the elements in the drawing having that reference numeral (e.g., "110" in the text refers to reference numerals "110a", "110b", and / or "110n" in the drawing).
[0030] The system architecture 100 (also referred to as "the system" in this document) includes an online virtual experience server 102, a data storage area 120, user devices 110a, 110b, and 110n (commonly referred to as "user devices 110" in this document), and developer devices 130a and 130n (commonly referred to as "developer devices 130" in this document). The virtual experience server 102, the sound server 140, the data storage area 120, the user devices 110, and the developer devices 130 are coupled via a network 122. In some embodiments, the (one or more) user devices 110 and the (one or more) developer devices 130 may refer to the same device or devices of the same type.
[0031] The online virtual experience server 102 may include a virtual experience engine 104, one or more virtual experiences 106, and a graphics engine 108. The user device 110 may include a virtual experience application 112 and an input / output (I / O) interface 114 (e.g., an input / output device). The input / output device may include one or more of a microphone, a speaker, headphones, a display device, a mouse, a keyboard, a game controller, a touch screen, a virtual reality console, etc. The input / output device may also include accessory devices connected to the user device via a cable (wired) or wirelessly.
[0032] The sound server 140 may include an audio engine 144 and a sound controller 146. In some embodiments, the sound server may include multiple servers. In some embodiments, the sound server 150 may be connected to the network 122 and the virtual experience server 102. In some embodiments, the multiple servers may be hierarchically arranged, for example, based on corresponding priority values assigned to sound sources. For example, in some embodiments, the generation of one or more sound sources may be assigned to the hierarchical servers based on the priority values associated with the sound sources.
[0033] In some embodiments, the sound server 140 may be connected to the data storage area 120 and may utilize the data storage area 120 to store data elements associated with the generation of sounds. In some other embodiments, the sound server 140 may include a separate data storage area 148.
[0034] The audio engine 144 may be used to generate one or more sounds associated with the virtual environment. The sound controller 146 may be used to orchestrate the computing resources associated with the generation of sounds, e.g., invoking computing instances for sound generation, load balancing of different processes / instances within a distributed computing environment, etc.
[0035] The developer device 130 may include a virtual experience application 132 and an input / output (I / O) interface 134 (e.g., an input / output device). The input / output device may include one or more of a microphone, a speaker, headphones, a display device, a mouse, a keyboard, a game controller, a touch screen, a virtual reality console, etc.
[0036] The system architecture 100 is provided for illustration. In different embodiments, the system architecture 100 may include the same, fewer, more, or different elements configured in the same or different ways as shown. Figure 1 shown.
[0037] In some embodiments, the network 122 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wired network (e.g., Ethernet), a wireless network (e.g., an 802.11 network, a network, or a wireless LAN (WLAN)), a cellular network (e.g., a 5G network, a long term evolution (LTE) network, etc.), a router, a hub, a switch, a server computer, or a combination thereof.
[0038] In some embodiments, the data store 120 may be a non-transitory computer-readable memory (e.g., a random access memory), a cache, a drive (e.g., a hard disk drive), a flash drive, a database system, a cloud storage system, or another type of component or device capable of storing data. The data store 120 may also include multiple storage components (e.g., multiple drives or multiple databases) that may also span multiple computing devices (e.g., multiple server computers).
[0039] In some embodiments, the online virtual experience server 102 may include a server having one or more computing devices (e.g., a cloud computing system, a rack server, a server computer, a physical server cluster, etc.). In some embodiments, the online virtual experience server 102 may be a stand-alone system, may include multiple servers, or may be part of another system or server.
[0040] In some embodiments, the online virtual experience server 102 may include one or more computing devices (e.g., rack servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, distributed computing systems, etc.), data storage areas (e.g., hard disks, memories, databases), networks, software components, and / or hardware components, which can be used to perform operations on the online virtual experience server 102 and provide users with access to the online virtual experience server 102. The online virtual experience server 102 may also include a website (e.g., a web page) or application backend software that can be used to provide users with access to the content provided by the online virtual experience server 102. For example, a user may access the online virtual experience server 102 using the virtual experience application 112 on the user device 110.
[0041] In some embodiments, the online virtual experience server 102 may be a social network that provides connections between users or a user-generated content system that allows users (e.g., end users or consumers) to communicate with other users on the online virtual experience server 102. In this user-generated content system, the communication may include voice chat (e.g., synchronous and / or asynchronous voice communication), video chat (e.g., synchronous and / or asynchronous video communication), or text chat (e.g., synchronous and / or asynchronous text-based communication). In some embodiments of the present disclosure, a "user" may be represented as a single individual. However, other embodiments of the present disclosure cover "users" as entities controlled by a group of users or an automated source (e.g., creative users). For example, a group of individual users united as a community or group in a user-generated content system may be regarded as a "user".
[0042] In some embodiments, the online virtual experience server 102 may be an online game server. For example, the virtual experience server may provide single-player or multi-player games to a user community, where the user community may access or interact with the games using the user device 110 through the network 122. In some embodiments, the games (also referred to herein as "video games", "online games", or "virtual games") may be, for example, two-dimensional (2D) games, three-dimensional (3D) games (e.g., 3D user-generated games), virtual reality (VR) games, or augmented reality (AR) games. In some embodiments, a user may participate in a game with other users. In some embodiments, the game may be played in real time with other users of the game.
[0043] In some embodiments, gameplay can refer to the interaction of one or more players using a user device (e.g., 110) within a game (e.g., 106), or the presentation of the interaction on a display or other output device (e.g., 114) of the user device 110.
[0044] In some embodiments, the game 106 can include an electronic file that can be executed or loaded using software, firmware, or hardware for presenting game content (e.g., digital media items) to an entity. In some embodiments, the virtual experience application 112 can be executed and the game 106 can be executed in conjunction with the virtual experience engine 104. In some embodiments, the game 106 can have a set of common rules or common goals, and the environments of the game 106 share the set of common rules or common goals. In some embodiments, different games can have rules or goals that are different from each other.
[0045] In some embodiments, a virtual experience can have one or more environments (also referred to herein as "game environments" or "virtual environments") in which multiple environments can be linked. Examples of environments can be three-dimensional (3D) environments. One or more environments of the virtual experience application 106 can be collectively referred to herein as a "world" or "game world" or "virtual world" or "universe". An example of a world can be the 3D world of the game 106. For example, a user can construct a virtual environment that links to another virtual environment created by another user. Characters in a virtual game can cross virtual boundaries and enter adjacent virtual environments.
[0046] It should be noted that 3D environments or 3D worlds use graphics that use a three-dimensional representation of geometric data representing game content (or at least present game content as 3D content whether or not a 3D representation of geometric data is used). 2D environments or 2D worlds use graphics that use a two-dimensional representation of geometric data representing game content.
[0047] In some embodiments, the online virtual experience server 102 may host one or more virtual experiences 106 and may allow users to interact with the virtual experiences 106 using a virtual experience application 112 of the user device 110. Users of the online virtual experience server 102 may play virtual experiences 106, create virtual experiences 106, interact with virtual experiences 106, or build virtual experiences 106, communicate with other users, and / or create and build objects of virtual experiences 106 (e.g., also referred to herein as "items" or "game objects" or "virtual game items"). For example, when generating user-generated virtual items, a user may create characters, decorations for the characters, one or more virtual environments for an interactive game, or build structures used in the game. In some embodiments, a user may purchase, sell, or trade virtual game objects, such as in-platform currency (e.g., virtual currency), with other users of the online virtual experience server 102. In some embodiments, the online virtual experience server 102 may transmit game content to the virtual experience application (e.g., 112). In some embodiments, game content (also referred to herein as "content") may refer to any data or software instructions associated with the online virtual experience server 102 or the virtual experience application (e.g., game objects, games, user information, videos, images, commands, media items, etc.). In some embodiments, game objects (e.g., also referred to herein as "items" or "objects" or "virtual objects" or "virtual game items") may refer to objects used, created, shared, or otherwise depicted in the virtual experience application 106 of the online virtual experience server 102 or the virtual experience application 112 of the user device 110. For example, game objects may include parts, models, characters, decorations, tools, weapons, clothing, buildings, vehicles, currency, flora, fauna, components of the foregoing (e.g., windows of a building), etc.
[0048] It should be noted that the provision of the online virtual experience server 102 hosting the virtual experiences 106 is for illustration and not limitation. In some embodiments, the online virtual experience server 102 may host one or more media items, and the media items may include communication messages from one user to one or more other users. Media items may include, but are not limited to, digital videos, digital movies, digital photos, digital music, audio content, melodies, digital concerts, digital lecture series, website content, social media updates, e-books, e-magazines, digital newspapers, digital audiobooks, e-journals, web blogs, real simple syndication (RSS) feeds, digital comic books, software applications, etc. In some embodiments, the media item may be an electronic file that may be executed or loaded using software, firmware, or hardware for presenting the digital media item to an entity.
[0049] In some embodiments, the virtual application 106 can be associated with a specific user or a specific group of users (e.g., a private game), or can be widely available to users having access to the online virtual experience server 102 (e.g., a public game). In some embodiments, in the case where the online virtual experience server 102 associates one or more virtual experiences 106 with a specific user or group of users, the online virtual experience server 102 can use user account information (e.g., user account identifiers such as a username and password) to associate the specific user with the virtual experience 106.
[0050] In some embodiments, the online virtual experience server 102 or the user device 110 can include a virtual experience engine 104 or a virtual experience application 112. In some embodiments, the virtual experience engine 104 can be used for the development or execution of the virtual experience 106. For example, the virtual experience engine 104 can include a rendering engine (“renderer”) for 2D, 3D, VR, or AR graphics, a physics engine, a collision detection engine (and collision response), a sound engine, a scripting function, an animation engine, an artificial intelligence engine, a network function, a streaming function, a storage management function, a threading function, a scene graph function, or an animated video support, and other functions. The components of the virtual experience engine 104 can generate commands (e.g., rendering commands, collision commands, physics commands, etc.) that assist in calculating and rendering the game. In some embodiments, the virtual experience applications 112 of the user devices 110 / 116 can work independently and / or cooperate with the virtual experience engine 104 of the online virtual experience server 102.
[0051] In some embodiments, both the online virtual experience server 102 and the user device 110 can execute virtual experience engines (104 and 112 respectively). The online virtual experience server 102 using the virtual experience engine 104 can execute some or all of the virtual experience engine functions (e.g., generating physical commands, rendering commands, etc.), or offload some or all of the virtual experience engine functions to the virtual experience engine 104 of the user device 110. In some embodiments, the ratio of the virtual experience engine functions executed on the online virtual experience server 102 to those executed on the user device 110 for each virtual application 106 can be different. For example, the virtual experience engine 104 of the online virtual experience server 102 can be used to generate physical commands in the case of a collision occurring between at least two virtual application objects, while additional virtual experience engine functions (e.g., generating rendering commands) can be offloaded to the user device 110. In some embodiments, the ratio of the virtual experience engine functions executed on the online virtual experience server 102 and the user device 110 can change based on game conditions (e.g., dynamically). For example, if the number of users participating in a particular virtual application 106 exceeds a threshold number, the online virtual experience server 102 can execute one or more virtual experience engine functions previously executed by the user device 110.
[0052] For example, a user can play the virtual application 106 on the user device 110 and can send control instructions (e.g., user inputs such as right, left, up, down, user selection, or character position (location), and speed information) to the online virtual experience server 102. After receiving the control instructions from the user device 110, the online virtual experience server 102 can send gameplay instructions (e.g., the position (location) and speed information of a character participating in a team game, or commands such as rendering commands, collision commands) to the user device 110 based on the control instructions. For example, the online virtual experience server 102 can perform one or more logical operations on the control instructions (e.g., using the virtual experience engine 104) to generate the (one or more) gameplay instructions for the user device 110. In other cases, the online virtual experience server 102 can transfer one or more control instructions from one user device 110 to other user devices participating in the virtual application 106 (e.g., from user device 110a to user device 110b). The user device 110 can use the gameplay instructions and render the game to present on the display of the user device 110.
[0053] In some embodiments, control instructions may refer to instructions that direct in-game actions of a user's character. For example, control instructions may include user input for controlling actions in the game, such as right, left, up, down, user selection, gyroscopic position (orientation) and heading data, force sensor data, etc. Control instructions may include character position (orientation) and speed information. In some embodiments, the control instructions are sent directly to the online virtual experience server 102. In other embodiments, the control instructions may be sent from the user device 110 to another user device (e.g., from user device 110b to user device 110n), where the other user device uses the local virtual experience engine 104 to generate gameplay instructions. Control instructions may include instructions to play a voice communication message or other sound from another user on an audio device (e.g., speakers, headphones, etc.), such as voice communication or other sounds generated using audio spatialization techniques as described herein.
[0054] In some embodiments, gameplay instructions may refer to instructions that allow the user device 110 to render the gameplay of a game (e.g., a multiplayer game). Gameplay instructions may include one or more of user input (e.g., control instructions), character position (orientation) and speed information, or commands (e.g., physics commands, rendering commands, collision commands, etc.).
[0055] In some embodiments, the online virtual experience server 102 may store user-created characters in the data store 120. In some embodiments, the online virtual experience server 102 maintains a character directory and a game directory that can be presented to the user. In some embodiments, the game directory includes images of virtual experiences stored on the online virtual experience server 102. Additionally, the user may select a character from the character directory (e.g., a character created by the user or another user) to participate in the selected game. The character directory includes images of the characters stored on the online virtual experience server 102. In some embodiments, one or more of the characters in the character directory may have been created or customized by the user. In some embodiments, the selected character may have character settings that define one or more components of the character.
[0056] In some embodiments, the role of a user may include the configuration of components, where the configuration and appearance of the components (more generally, the appearance of the role) may be defined by role settings. In some embodiments, the role settings of a user role may be at least partially selected by the user. In other embodiments, the user may select a role with default role settings or role settings selected by other users. For example, the user may select a default role from a role directory with predefined role settings, and the user may further customize the default role by changing some role settings (e.g., adding a shirt with a customized logo). The role settings may be associated with a specific role by the online virtual experience server 102.
[0057] In some embodiments, each of the user devices 110 may include a computing device (e.g., a personal computer (PC)), a mobile device (e.g., a laptop computer, a mobile phone, a smartphone, a tablet computer, or a netbook computer), a network-connected television, a game console, etc. In some embodiments, the user device 110 may also be referred to as a "client device". In some embodiments, one or more of the user devices 110 may be connected to the online virtual experience server 102 at any given moment. It should be noted that the number of user devices 110 provided is for illustration purposes. In some embodiments, any number of user devices 110 may be used.
[0058] In some embodiments, each user device 110 may include an instance of the virtual experience application 112, respectively. In one embodiment, the virtual experience application 112 may allow the user to use and interact with the online virtual experience server 102, such as controlling a virtual character in a virtual game hosted by the online virtual experience server 102, or viewing or uploading content, such as the virtual experience 106, images, video items, web pages, documents, etc. In one example, the virtual experience application may be a web application (e.g., an application operating in conjunction with a web browser) that can access, retrieve, present, or navigate content served by a web server (e.g., virtual characters in a virtual environment, etc.). In another example, the virtual experience application may be a native application (e.g., a mobile application, an app, or a game program) that is installed and executed locally on the user device 110 and allows the user to interact with the online virtual experience server 102. The virtual experience application may render, display, or present content to the user (e.g., a web page, a media viewer). In an embodiment, the virtual experience application may also include an embedded media player (e.g., player) embedded in a web page.
[0059] In some embodiments, the virtual experience application may include an audio engine 116 that is installed on the user device and enables playback of sound on the user device. In some embodiments, the audio engine 116 may cooperate with an audio engine 144 installed on a sound server.
[0060] In accordance with aspects of the present disclosure, the virtual experience application may be an online virtual experience server application for allowing a user to build, create, edit, upload content to, and interact with an online virtual experience server 102 (e.g., participate in a virtual experience 106 hosted by the online virtual experience server 102). In this way, the virtual experience application may be provided by the online virtual experience server 102 to one or more user devices 110. In another example, the virtual experience application may be an application downloaded from a server.
[0061] In some embodiments, each developer device 130 may separately include an instance of the virtual experience application 132. In one embodiment, the virtual experience application 122 may allow a developer user to use and interact with the online virtual experience server 102, such as controlling a virtual character in a virtual game hosted by the online virtual experience server 102, or viewing or uploading content, such as a game 106, an image, a video item, a web page, a document, etc. In one example, the virtual experience application may be a web application (e.g., an application that operates in conjunction with a web browser) that can access, retrieve, present, or navigate content served by a web server (e.g., a virtual character in a virtual environment). In another example, the virtual experience application may be a native application (e.g., a mobile application, an app, or a virtual experience program) that is installed and executed locally on the user device 130 and allows a user to interact with the online virtual experience server 102. The virtual experience application may render, display, or present content to the user (e.g., a web page, a media viewer). In an embodiment, the virtual experience application may also include an embedded media player (e.g., a player).
[0062] In accordance with aspects of the present disclosure, the virtual experience application 132 can be an online virtual experience server application for enabling a user to construct, create, edit content, upload the content to the online virtual experience server 102, and interact with the online virtual experience server 102 (e.g., provide and / or play a game 106 hosted by the online virtual experience server 102). In this way, the virtual experience application can be provided by the online virtual experience server 102 to the user device 130. In another example, the virtual experience application 132 can be an application downloaded from a server. The virtual experience application 132 can be used to interact with the online virtual experience server 102 and obtain access to user credentials, user currency, etc. for one or more virtual applications 106 developed, hosted, or provided by a virtual experience application developer.
[0063] In some embodiments, a user can log in to the online virtual experience server 102 through the virtual experience application. The user can access the user account by providing user account information (e.g., username and password), where the user account is associated with one or more characters that can be used to participate in one or more games 106 of the online virtual experience server 102. In some embodiments, with appropriate credentials, a virtual experience application developer can obtain access to virtual experience application objects such as in-platform currency (e.g., virtual currency), avatars, special abilities, decorations owned or associated with other users.
[0064] Generally, if appropriate, functions described in one embodiment as being performed by the online virtual experience server 102 can also be performed by the (one or more) user devices 110 or a server in other embodiments. Additionally, functions attributed to a particular component can be performed by different components or multiple components operating together. The online virtual experience server 102 can also be accessed as a service provided to other systems or devices through an appropriate application programming interface (API) and is thus not limited to use in a website.
[0065] In some embodiments, the online virtual experience server 102 can include a graphics engine 106. In some embodiments, the graphics engine 106 can be a system, application, or module that allows the online virtual experience server 102 to provide graphics and animation capabilities. In some embodiments, the graphics engine 106 can perform one or more operations described in connection with the flowcharts shown in Figure 4 and Figure 5 below.
[0066] Figure 2A An example system architecture for the simulation of sound in a virtual environment is shown in accordance with some embodiments.
[0067] AsFigure 2A As depicted in Figure 2A , one or more user devices 110 are coupled to a virtual experience server 102 via a network 122. Each user device may include (e.g., have installed) a virtual experience application (e.g., software) that executes on the user device to enable the user to connect to the virtual experience server and participate in games and / or other activities within the virtual environment. The user device also includes a sound library 250 for simulating sounds within the virtual environment.
[0068] In some embodiments, the sound library may be a dedicated library for generating (producing) sounds. The sounds may be generated by the sound library based on triggers and / or signals received from the virtual experience application, and the virtual experience application may make calls to the sound library via an application programming interface (API). The generated sounds may then be played back on the user device, e.g., using the user device's speaker or an auxiliary device (e.g., headphones, earbuds, etc.) connected to the user device. The time points of the sound playback may be configured by the virtual experience application to be synchronized with the virtual environment and with the activities of one or more user avatars and / or virtual objects within the virtual environment.
[0069] In some embodiments, the sound library may be external to the virtual experience application, while in some other embodiments, the sound library may be included within the virtual experience application executable file.
[0070] In some embodiments, the sound library may be implemented as a remote procedure call (RPC) shim that is configured to marshal arguments in addition to generating (creating) sounds and forward these arguments as calls over the network. The remote procedure call system may call a remote processor coupled to the user device via a communication channel; the remote processor may be a server accessible over the network, another CPU core in a multi-core system on the user device, computing resources provided by a distributed computing system, etc. The implementation of the RPC shim is such that, from the perspective of the virtual experience application (local client), the RPC is similar to calling a function in a library built into the virtual experience application (e.g., a game application) that executes on the user device.
[0071] Figure 2A An example sound server 210 for cloud-based sound processing in a virtual environment is depicted. The sound server may also include modules for performing specific processing. For example, the sound server may include a sound source 215 that may include a collection of stored sounds that were previously generated and / or recorded for use within the virtual environment. The sound server also includes a module for processing sound effects 220 (e.g., sound effects such as loudness adjustment, Doppler adjustment, reverb effects, echo effects, etc.).
[0072] In some embodiments, as Figure 2A depicted in Figure 2A , the sound server 210 may be included within the virtual experience server. In some other embodiments, the sound server may be separate from the virtual experience server.
[0073] The sound server may additionally include an audio mixer 225 and an audio encoder 235. The audio mixer 225 may combine audio segments from multiple sources, and the audio encoder 235 may encode the final audio mix for further transmission, such as to a user device.
[0074] In some embodiments, virtual application state information 230 is provided to the sound server 210, and the virtual application state information 230 may be provided by the server-based virtual experience engine 104. The virtual application state information may include information associated with objects, avatars, and activities within the virtual environment. For example, the virtual application state information may include the location (positioning) of one or more avatars within the virtual environment, the rate or speed of one or more avatars within the virtual environment, the orientation of the avatars within the virtual environment, and so on.
[0075] In some embodiments, the virtual application is a game application, and the virtual application state information includes game state information, such as the location and speed of one or more avatars and / or objects within the game, the orientation of the avatars within the game, and so on.
[0076] Although this illustrative example depicts a sound server to be included within the virtual experience server, in some embodiments, the sound server may also be implemented separately. For example, the sound server may be implemented using a distributed computing system, and the RPC for sound library production may be transmitted to a sound server implemented using one or more computing resources (processes) of the distributed computing system.
[0077] In some embodiments, a single sound server process is provided in the cloud for each user device (player). Multiple sound server processes may be executed on a single physical or virtual server.
[0078] In some embodiments, multiple sound server processes may be provided in the cloud for each user device, while in some other embodiments, a single sound process may be provided to serve requests from multiple user devices.
[0079] A sound server can offload at least some of the processing associated with playing audio to a device independent of the local user device and can handle sounds that may exceed the computational or memory limits of the player device, as the sound server can be configured to include additional computing resources. Additionally, the hardware of the server can be customized for audio processing and run optimized hardware-specific code on it. More simultaneously playable sounds can be achieved using a sound server compared to a user device (e.g., a player's computing device).
[0080] However, generating sounds on a remote sound server incurs latency due to the time required for RPC calls, the round-trip communication time between the user device and the sound server, and other delays and wait times that may be introduced into the process. This may be unacceptable for some high-priority sound sources that may have to be played with a relatively low time latency. For example, the sound corresponding to an avatar jumping into water and creating a splash may be best played (rendered) on the local user device when the avatar enters the water surface to achieve an excellent user experience. In such a scenario, the time delay caused by transmitting the request to the sound server, the sound server generating the sound, and transmitting the finally generated sound back to the user device, as well as the sound playback, may be too large to provide an appropriate user experience.
[0081] In some embodiments, a priority value can be determined for various sound sources based on the time sensitivity of sound playback, and this priority value can be used to determine whether a sound should be generated locally or on the sound server.
[0082] For example, in some embodiments, a priority value can be determined for each generated sound source. Based on the priority value, one or more of the sounds to be generated can be generated locally on the user device. The priority value can be based on loudness, the distance of the sound object from the avatar, or other flags set by the user or developer.
[0083] In some embodiments, the priority of a sound can be determined based on the criticality (importance) of the particular sound to the artistic effect of the virtual experience or based on the importance of the particular sound to enhancing the virtual experience (e.g., gameplay). For example, in a first-person shooter game, the footsteps, clothing sounds, and breathing sounds of an enemy may be given priority (e.g., assigned a relatively high priority value) because they are very important for the game objective of determining the possible location of an invisible enemy within the virtual environment. The dialogue spoken by a non-player character (NPC) may be given priority because the dialogue related to the NPC may be very important in the game narrative, and / or because it may include important information required to complete the game, or because it may set the emotional tone of the game.
[0084] In some embodiments, the priority value can be based on the computing capabilities of the local user device to generate one or more sounds. In some cases, the priority value can be specified as a time delay threshold for each sound to be received. Based on network conditions, such as the round-trip transmission time to satisfy a request for generating a sound at the server, etc., the total time to receive the sound can be determined, and this total time is compared with the time delay threshold to determine whether a particular sound should be generated locally or at a remote sound server.
[0085] In some embodiments, the sound server can be configured as a microservice, whereby requests to the sound server are processed by sound server instances provided using a distributed computing environment. When a sound server instance is started (instantiated), registration is performed by an administrative server (not shown, but can be implemented as part of the virtual application server 102 or as a separate server). The administrative server assigns the sound server instance to a user device (virtual application client). Assigning the sound server instance to a user device establishes a source device and a target device from which the sound server instance can receive RPC calls and to which the final audio mixed stream will be transmitted.
[0086] In some embodiments, the administrative server performs the assignment from the sound server to the user device (virtual application client) and provides an identifier to the virtual application client, such as the addresses (e.g., IP addresses, MAC addresses, etc.) of all the sound servers assigned to a particular user device (virtual application client).
[0087] In some other embodiments, the administrative server can act as a proxy server and accept sound server calls (RPCs) from the user device and then forward the above calls to a group of separate sound servers.
[0088] The administrative server can be responsible for starting a remote sound server on behalf of the virtual application client. When the virtual application is executing, the administrative server can start or stop the remote sound server as needed. If there are a large number of sound sources generating sounds simultaneously within the virtual application in the virtual environment, more servers can be started to meet the requirements. Similarly, if the requirements associated with the virtual application result in relatively few sounds to be played back simultaneously, additional remote sound servers can be shut down to save costs. If the administrative server is configured to act as a proxy server for multiple remote sound servers, the administrative server can determine (measure) the demand for the sound server either through signals transmitted from the user device (virtual application client) to the administrative server or by directly measuring the number of sounds triggered within the virtual experience.
[0089] For interoperability, microservices can use standard RPC and marshaling protocols, as well as standard protocols such as WebRTC, real-time transport protocol (RTP), or alternative network standards designed for transmitting audio / video data and optimized for consistent transmission of real-time data for audio streams.
[0090] In some embodiments, the sound server can utilize one or more graphics processing units (GPUs) to process sound. In some other embodiments, the sound server can utilize one or more field programmable gate arrays (FPGAs) or digital signal processing (DSP) cores. In some embodiments, the DSP can be installed on an expansion card to accelerate audio processing in the server.
[0091] In some embodiments, the sound server can trigger external audio hardware, such as a music sampler or synthesizer, and then re-digitalize the generated sound for transmission over the network to a user device.
[0092] The sound server receives an RPC transmitted from a user device and converts the RPC back into a call to a server-based sound library. The server-based sound library creates sounds based on the parameters contained in the RPC and generates an audio mix. In some embodiments, the audio mix is encoded and sent back over the network as streaming audio to a shim library of the user device associated with the transmitted RPC. The streaming audio can be in any suitable format and can include immersive audio formats such as 7.1 surround sound format, 7.4.2 surround sound format, Ambisonics, vector base amplitude panning (VBAP), Dolby Atmos, etc.
[0093] The shim library of the user device receives the audio mix and plays back the audio mix on the user device. If the audio mix is in a rotatable immersive format, the user device can perform a rotation on the rendering of the mix to align with the current orientation of the player's avatar, provided that the orientation of the mix at the time of rendering the audio is different from the orientation of the player avatar. Using a rotatable immersive format for the sound format enables changing the 3D orientation of the sound field when presenting the sound to the user to match the orientation (e.g., a second orientation) at the time of sound playback, even if the orientation is different from the orientation (e.g., a first orientation) at the time of generating the sound in response to a request.
[0094] Each user device (virtual application client) can receive an audio mix from one or more sound servers. Additionally, the user device can generate (create) one or more high-priority sounds at the user device. The high-priority sounds can be determined based on a priority value corresponding to the sound source. The sounds generated at the user device can be combined with the sounds generated at one or more sound servers to create a final audio mix.
[0095] Figure 2B An example embodiment of a system architecture for simulating sounds in a virtual experience including a hierarchical arrangement of sound servers is shown.
[0096] In some embodiments, a hierarchical arrangement of sound servers can be provided such that a greater number of sounds can be generated than can be generated by a single sound server.
[0097] For example, a hierarchical arrangement of sound servers can be provided by arranging remote sound servers in a tree structure, where Tier 0 is the user device (virtual application client). The sound server system is configured such that each sound server can be connected to additional sound servers based on a branching factor L. If each sound server can generate N sounds, then any Nth tier can be connected to up to L^N sound servers, which can render up to K*L^N sounds. This configuration enables a relatively large number of sounds to be achieved even with a relatively small number of tiers and a small branching factor L. By utilizing additional sound servers that can be called (initialized) based on the current demand of the sound servers, a large total number of sounds can be provided that can be generated and played back at the user device.
[0098] Figure 2B The illustrative example depicted includes Tier 0 to Tier N, where Tier 0 is the (local) user device 110. Tier 1 includes sound server 210a; Tier 2 includes servers 210b and 210c; Tier 3 includes sound servers 210a, 210e, 210f, 210g; and Tier N includes 210n, 210m, 210o, and 210p. Different tiers can have different numbers of servers, and the number of tiers can be selected based on various parameters (e.g., the number of sound generation objects in the virtual experience, the number of avatars, etc.).
[0099] Each tier can have an appropriate number of servers based on the number of sound sources that meet the priority threshold of that tier. In some embodiments, the number of servers included in a tier can be based on the sound generation capabilities of the servers.
[0100] In some embodiments, the number of layers and the branching factor can be based on one or more parameters, such as the hardware capacity of the user device, the number of sound sources, network and traffic related parameters, etc.
[0101] In some embodiments, the number of servers in each layer and the total number of layers can be adjusted dynamically. In this illustrative example, the branching factor is 2. For example, each voice server can be connected to at most 2 voice servers in the higher layer.
[0102] The priority values of the sound sources in the multiple sounds to be generated can be utilized to determine how to distribute the sound generation among the hierarchically arranged voice servers.
[0103] In some embodiments, different types of voice servers can be used for different types of sounds. In some embodiments, the sound catalog can be distributed across multiple voice servers based on the capacity of the block storage available in the voice servers. For example, if the sound catalog includes approximately 8TB of sounds and each voice server has a block storage capacity of 2TB, the 8TB of sounds can be distributed across 4 servers. Based on the incoming requests to generate sounds, appropriate voice servers can be utilized and the sound playback can be routed to the voice server that includes the data of the required sound.
[0104] In some embodiments, the branching factor (L) can be determined based on the input / output (I / O) capabilities of the network. For example, if a server can only have 100 network connections, one network connection is allocated for transmitting the generated audio mix outward, and at most 99 additional voice servers can be configured as branch servers that can receive the audio from the above branch servers.
[0105] Figure 3 FIG. is a diagram showing an example scenario within a virtual experience (environment) that utilizes sounds according to some embodiments. The example scenario can be a scene displayed on the display of a user device used by a player playing a game on an online virtual experience platform.
[0106] In this illustrative example, two characters (avatars) are depicted; a first avatar 320 and a second avatar 350. The virtual experience includes a cave 310, a waterfall 330, and a forest experience 340.
[0107] One or more objects and / or avatars within the virtual experience can be associated with sound objects. The sound objects can be associated with the avatars and / or objects in the experience by a developer user, such as a developer using Figure 1 the developer device 130 described above.
[0108] In this illustrative example, the waterfall can be associated with the sound of the waterfall, and the forest can be associated with the sounds of the forest (e.g., including the sounds of crickets, other insects, etc.).
[0109] Developer users can also specify sound effects and settings, such as absolute loudness, roll-off distance specifying the distance from the sound object at which the sound starts to attenuate, a function specifying how loudness varies with the distance of the object (e.g., linear attenuation, logarithmic attenuation, etc.), a reverb effect that provides closely stacked echoes to form a mixture of delayed sounds that are difficult to distinguish, an echo effect based on the distance of an object within an enclosed space, etc.
[0110] In some embodiments, a sound server can be used for diegetic sounds (sound sources in a virtual world) as well as non-diegetic sounds. Non-diegetic sounds are sounds within a virtual experience (e.g., a game), but are not part of the virtual experience (world) because they do not correspond to any virtual physical source in the world. Examples in games and movies are narration and music tracks. Non-diegetic sounds generally do not pan with the perspective of a character (a person in the world) and / or an avatar, and thus can be processed differently from other sounds. Sound sources in the real world can also be included and associated with an avatar within a virtual experience. For example, the speech of a user associated with avatars 310 and 350 is a non-diegetic sound and can also be included as a sound object within the virtual experience.
[0111] In some embodiments, sounds can include voice chat, NPC dialogue, narration, background music, ambient sound, spot sound effects, foley, etc.
[0112] Figure 4 is a flowchart showing an example method of providing an encoded sound mix to a user device according to some embodiments.
[0113] In some embodiments, method 400 can be implemented, for example, on the virtual experience server 102 described in reference Figure 1 In some other embodiments, method 400 can be implemented, for example, on one or more sound servers described in reference to FIG. 2. In the example described, the implementation system includes one or more digital processors or processing circuits (“processors”), and one or more storage devices (e.g., data storage area 120 or other storage devices). In some embodiments, different components of one or more servers and / or clients can execute different blocks or other parts of method 400. In some examples, a first device is described as executing a block of method 400. Some embodiments can have one or more blocks of method 400 executed by one or more other devices (e.g., other client devices or server devices) that can send results or data to the first device.
[0114] Method 400 can start at block 405.
[0115] At block 405, a request to generate multiple sounds for a user device is received at a sound server. The user device may be associated with an avatar participating in a virtual experience.
[0116] In some embodiments, the frequency of the received request to play sounds may be based on a particular virtual experience. For example, if the request is associated with a game, the frequency may be based on a particular game and / or sound design. For example, as the game loop iterates, sound requests may be batched and processed once every one-sixtieth of a second, and each request and / or iteration may include any number of sounds.
[0117] After block 405 may be block 410.
[0118] At block 410, sound source data for a plurality of sound sources associated with each of the multiple sounds is obtained. In some embodiments, the sound source data is obtained from a set of sound objects stored and associated with the virtual experience. The sound source data may be an audio (sound) file uploaded by a user, or a sound object provided by an online virtual experience platform. For example, the sound objects may include sounds such as footsteps, waterfall sounds, weapon sounds, object collision sounds, etc. After block 410 may be block 415.
[0119] At block 415, the remote sound server obtains virtual experience status information.
[0120] In some embodiments, the virtual experience status information may include the position of a virtual microphone in the virtual experience and at least one of the following: the speed of the virtual microphone in the virtual experience or the orientation of the virtual microphone in the virtual experience.
[0121] In some embodiments, for example, the speed of the virtual microphone may be the absolute speed of the virtual microphone within the virtual experience. In some embodiments, the speed of the virtual microphone may be the relative speed of the virtual microphone with respect to one or more sound sources.
[0122] In some embodiments, the position of the virtual microphone may be based on the position (location) of a virtual camera placed within the virtual experience. In some embodiments, the position of the virtual microphone may match the position (location) of a virtual camera placed within the virtual experience, for example, being at the same position as the virtual camera placed within the virtual experience. In some other embodiments, the position of the virtual camera and / or the virtual microphone may be based on the position (location) of a participant's avatar within the virtual experience. For example, in some embodiments, the virtual camera and / or the virtual microphone may be located at a position opposite to the shoulder position (location) of the avatar associated with the user device.
[0123] In some embodiments, the position of the virtual microphone can be equidistant from the position (orientation) of the virtual camera and the position (orientation) of the avatar associated with the user device. In some embodiments, the position of the virtual microphone is based on a developer-specified distance associated with the virtual experience, where the developer specifies the position (orientation) of the virtual microphone relative to the virtual camera and the avatar.
[0124] In some embodiments, the virtual experience can be an online game, and the virtual experience status information can include game status information. For example, the game status information can include one or more parameters such as the position of the avatar in the virtual experience, the speed or rate of the avatar in the virtual experience, and the orientation of the avatar in the virtual experience. In some embodiments, the game status information can include two or more of the following: the position of the avatar in the virtual experience, the speed (or rate) of the avatar in the virtual experience, and the orientation of the avatar in the virtual experience.
[0125] After block 415 can be block 420.
[0126] At block 420, the remote sound server generates an audio mix of the plurality of sounds based on the source data of the plurality of sounds and the virtual experience status information for the first request. Figure 5 An example method of generating an audio mix of the plurality of sounds is shown. After block 420 can be block 425.
[0127] At block 425, the audio mix is transmitted to the user device. For example, the audio mix can first be encoded in a streaming audio format and then transmitted to the user device. Providing the audio mix in a streaming format enables the user device to quickly or instantaneously play back the audio because playback can begin based on the received data. Synchronization of the encoded sound packets can be based on timestamped audio, video, and other events.
[0128] To transmit the sound over the network to the user device, the audio mix can be encoded such that the codec packets include approximately 20 milliseconds of sound, but the encoding and codec size can be configured based on a particular virtual experience.
[0129] In some embodiments, the packet size can be chosen to be as small as possible to reduce latency until the limit of efficiency is reached. For example, if too little audio is sent in each packet, most of the packet data will be network packet headers, resulting in wasted bandwidth.
[0130] If no sound is received, error correction or concealment can be utilized at the user device to handle the lost sound. For example, model synthesis of the sound playback can be used to replace the audio for the lost packets from the sound server.
[0131] Figure 5 FIG. is a flowchart showing an example method of generating an audio mix of multiple sounds according to some embodiments.
[0132] Method 500 may start at block 510 and may be performed for each sound source included in the multiple sound sources for which corresponding sounds are to be generated.
[0133] At block 510, information associated with the sound source may be obtained for each of the multiple sound sources. The sound source may be, for example, an audio file in a suitable format such as WAV, PCM, AIFF, MP3, AAC, OGG, FLAC, ALAC, etc. Based on the sound source, an audio (sound) segment of a time length corresponding to the request may be generated.
[0134] In some embodiments, the sound source file is obtained from a data storage area accessible by a remote sound server. In some embodiments, the sound sources associated with a particular virtual experience or game may be pre-fetched and stored on a storage device associated with the sound server, such as high-bandwidth memory associated with the processor of the source server.
[0135] After block 510 may be block 515.
[0136] At block 515, a particular sound source is selected. After block 515 may be block 520.
[0137] At block 520, a priority value may be determined for the particular sound source.
[0138] After block 520 may be block 525.
[0139] At block 525, a server may be selected or assigned based on the priority value of the sound source.
[0140] The disadvantage of rendering sounds on a remote server is that there may be a delay of several hundred milliseconds between triggering a sound on the player device, rendering the sound on the server, and finally receiving and playing back the sound on the player device.
[0141] To alleviate this problem, the shim library may be extended to a sound library. When a request for a sound is triggered, the priority value of the sound source is used to determine whether to play the sound using the local sound library or to render it on the sound server. The priority value may be sent from a virtual experience (e.g., a game), or calculated from the parameters of the sound (e.g., loudness, distance). In this case, high-priority sounds are played almost immediately, while low-priority sounds are played with some delay.
[0142] In some embodiments, the plurality of sounds included in the first request may not include the sounds to be generated on the user device. Generating sounds on the user device can reduce the latency that may occur when transmitting the request from the user device to the sound server, generating the sounds at the sound server, and transmitting the generated sounds back to the user device for subsequent playback. A priority value can be determined for each sound source, and this priority value can be used to determine whether one or more sounds should be generated on the user device or on a remote sound server.
[0143] In some embodiments, a priority value for each sound source is determined based on one or more parameters associated with an avatar associated with the sound object and / or the user device. For example, the priority value of a sound source can be based on the loudness of the sound, the distance between the avatar and the sound source, the distance between the virtual microphone and the sound source, etc.
[0144] In some embodiments, the priority value of a sound source is based on one or more of the following: the loudness of each sound source associated with the sound and the distance of each sound source from the virtual microphone in the virtual experience.
[0145] In some embodiments, a sorted list of all sounds to be generated can be created, and a threshold number of high-priority sounds can be generated locally on the user device, and the remainder can be included in the first request transmitted to the remote sound server. For example, each sound included in the first request can be associated with a priority value that meets the determined threshold.
[0146] In some embodiments, the threshold can be a predetermined fixed threshold. In some other embodiments, the threshold is based on network rate (bandwidth of the network connection, measured network rate, etc.), round-trip latency of data packets between the user device and one or more remote sound servers, local user device capacity, etc.
[0147] In some embodiments, the priority value is determined on the local device and transmitted to the remote sound server. For example, the first request can include the priority value of each sound source among the plurality of sound sources included in the request.
[0148] In some other embodiments, the priority value is determined on the remote sound server based on parameters associated with each of the sound sources. Parameters associated with each sound can be received from the user device, from the virtual experience server, or from another data storage area associated with the online virtual experience platform.
[0149] In some embodiments, each sound can be sorted, and a certain portion of the sound (e.g., a certain percentage or number of sounds) can be executed on the local user device, and the remainder of the sound can be included in the request transmitted to the remote sound server for generation at the remote sound server.
[0150] After block 525, there may be block 530.
[0151] At block 530, an audio clip may be generated. After block 530, there may be block 535.
[0152] At block 535, a loudness adjustment may be applied to the audio clip based on the distance between the sound source and the virtual microphone and / or the avatar in the virtual experience to generate a loudness-adjusted audio clip.
[0153] After block 535, there may be block 540.
[0154] At block 540, a Doppler adjustment may be applied to the loudness-adjusted audio clip based on the speed of the virtual microphone and / or the avatar in the virtual experience to generate a loudness- and Doppler-adjusted audio clip. In some embodiments, the speed of the virtual microphone and / or the avatar may be a relative speed compared to one or more sound sources within the virtual experience.
[0155] After block 540, there may be block 545.
[0156] At block 545, it is determined whether there is an additional sound source to be processed. If it is determined that there is an additional sound source to be processed, then after block 545, there may be block 515; otherwise, after block 545, there may be block 550.
[0157] At block 550, the loudness- and Doppler-adjusted audio clips of each sound among multiple sounds are mixed to generate an audio mix of the multiple sounds.
[0158] In some embodiments, the transmission of the audio mix is to transmit the audio mix encoded in a streaming audio format to a user device. For example, an audio codec such as VBAP, Dolby, surround sound, etc. may be used to encode the audio mix.
[0159] In some embodiments, a single sound server or a sound server process may be utilized to handle the sound of a single user device. In some embodiments, a single sound server may be used for multiple user devices.
[0160] In some embodiments, a single sound server or a sound server process may be utilized to handle requests from user devices associated with avatars in the same virtual experience. When the avatars are in the same virtual experience, each player hears the same sound source, but the Doppler effects for each avatar based on the distance between the avatar and the sound source, the position of the avatar relative to the sound source, and the relative movement of the avatar and the sound source within the virtual experience are different.
[0161] In some embodiments, computational power can be saved by generating initial audio for each sound source in the same way for each user device (associated with an avatar). Separate loudness effects and Doppler processing can be applied to the generated initial audio to generate two corresponding audio mixes, which can then be encoded and streamed to two different user devices (virtual application clients).
[0162] Even in cases where the virtual application clients require more sounds than a single sound server can generate, computational resource savings can be achieved by using a single sound server to generate sounds for user devices whose avatars are in the same virtual experience. For example, in an illustrative example, a single sound server can generate 6 sounds. In this illustrative example, assume that two user devices have avatars in the same virtual experience and a total of 8 sounds need to be generated. Sharing two sound servers, with each server assigned to generate 4 sounds, may be more computationally efficient than assigning a separate sound server to each user device, with each server generating the final audio mix for each client simultaneously. This approach of using a single sound server to generate sounds for user devices whose avatars are in the same virtual experience can result in fewer servers being used than if the virtual application clients did not share servers.
[0163] For example, in some embodiments, where a single sound server is used to generate all sounds in a single virtual experience, a second request to generate a second plurality of sounds can be received from a second user device whose avatar is participating in the same virtual experience as the user device. Generating the audio segments for the second user device can be based on generating unadjusted audio (sound) segments associated with the sound sources, and then applying respective adjustments for each user device based on their respective virtual application state information (e.g., game state information associated with the avatar associated with each user device).
[0164] For example, a first loudness adjustment can be applied to the audio segment based on the distance between the sound source and the avatar in the virtual experience associated with the user device to generate a first loudness-adjusted audio segment, and a second loudness adjustment can be applied to the audio segment based on the distance between the sound source and the avatar in the virtual experience associated with the second user device to generate a second loudness-adjusted audio segment.
[0165] After generating the loudness-adjusted audio segments, Doppler adjustments can be applied. For example, a first Doppler adjustment can be applied to the first loudness-adjusted audio segment based on the relative velocity (or rate) of the first avatar (associated with the first user device) and the sound source to generate a first loudness and Doppler-adjusted audio segment. A second Doppler adjustment can be applied to the second loudness-adjusted audio segment based on the relative velocity (or rate) of the second avatar (associated with the second user device) and the sound source to generate a second loudness and Doppler-adjusted audio segment.
[0166] Mix multiple first loudness and Doppler-adjusted audio segments, each first loudness and Doppler-adjusted audio segment associated with a respective sound source of a plurality of sounds, to generate a first audio mix of the plurality of sounds, and mix multiple second loudness and Doppler-adjusted audio segments, each second loudness and Doppler-adjusted audio segment associated with a respective sound source of the plurality of sounds, to generate a second audio mix of the plurality of sounds. The first audio mix and the second audio mix can be encoded and transmitted (streamed) to corresponding user devices (virtual application clients).
[0167] In some embodiments, the plurality of sounds can include at least one diegetic sound source and at least one non-diegetic sound source.
[0168] Virtual experiences typically include voice chat between participants. The voices of participants (e.g., gamers) are recorded by microphones on the participants' user devices and can be played back as diegetic sounds (sound sources associated with the virtual world) or as non-diegetic sounds (sounds originating from and / or existing outside the virtual world).
[0169] A remote sound server can be used to decode and mix voice chat streams from other participants. This provides the benefit of offloading processing from the participants' user devices to the remote audio server and enables the remote audio server to provide the same advantages for voice streams from participants as for other virtual application sounds, such as excellent sound fidelity, sound effects, etc.
[0170] To do this, the sound server is used to receive audio streams and transmit (send) audio streams. For diegetic sounds, metadata, e.g., metadata describing the location and velocity (optional) of the sound in the virtual world, is associated with the voice chat stream. After receiving and decoding the stream, the sound server can mix these streams just like all other virtual application sounds.
[0171] In some embodiments, the same codec can be used to generate voice-based sounds and other sounds. In some other embodiments, a separate dedicated voice-optimized codec can be used to process voice-based sounds. The sound server can be configured such that multiple codecs can be used simultaneously for different sounds.
[0172] In some embodiments, a user device may have an upper limit on the number of connectable remote sound servers. Each connection to a remote sound server has a network bandwidth cost from the audio stream, as well as a computational cost associated with decoding and mixing the stream. Thus, the user device (client) may reach the limit on the number of sound servers available for generating a sound collection. For example, in a particular scenario, based on the sound server capacity, 5 remote sound servers may be required to generate the requested number of sounds, but based on a particular limitation, the user device itself may only be able to connect to 1 or 2 remote sound servers.
[0173] In some embodiments, hierarchical mixing can be utilized to reduce the limitation on the number of remote sound servers that can be directly connected to a user device.
[0174] Hierarchical mixing can be achieved by arranging the remote sound servers in a hierarchical tree structure, where layer 0 is the user device (virtual application client). The sound server system is configured such that a sound server in a given layer N can mix up to K sounds and up to L streams (branching factor) received from the higher layer N + 1. Thus, any layer N can be connected to up to L^N servers, which can render up to K * L^N sounds. This configuration can achieve a relatively large number of sounds even with a small number of layers and a small branching factor L. By utilizing additional sound servers that can be called (initialized) based on the current demand of the sound servers, a large total number of sounds can be provided for generation and playback on the user device.
[0175] The priority values of the sound sources of the multiple sounds to be generated are utilized to determine how the sounds are distributed among the hierarchically arranged sound servers. In some embodiments, the hop latency from a first remote sound server to a second remote sound server is measured and the latency is further utilized to adjust the hierarchical mixing. For example, the time to generate a first collection of sounds may be less than the time to generate a second collection of sounds that is larger than the first collection. However, if the latency incurred by sending a request to generate a partial sound to the second server and receiving the generated sound collection increases, it may exceed the processing time cost of generating the second collection of sounds on the first server itself.
[0176] In some embodiments, dynamic adjustment of sound generation on the hierarchically organized servers can be performed based on virtual experience state information and / or network conditions.
[0177] In some embodiments of performing hierarchical mixing, generating an audio mix of multiple sounds may include generating a first portion (set) of the multiple sounds at a first remote sound server. The remaining portions of these sounds may be generated at a second server. A request to generate a second set (remaining portions) of the multiple sounds may be transmitted to the second server. In some embodiments, the hierarchical mixing of the multiple sounds does not include sounds to be generated at the user device itself.
[0178] The request may be received at the second server, which is used to generate the second set of the multiple sounds. In some embodiments, the generated second set of the multiple sounds is directly transmitted to the user device, while in some other embodiments, the generated second set of the multiple sounds may be transmitted back to the first server, where the second set of the multiple sounds may be combined (mixed) with the first set of the multiple sounds generated at the server.
[0179] In some embodiments, depending on the number of sounds in the multiple sounds, a third set of the multiple sounds may be generated at a third server. The third set may include sounds not included in the second set portion of the sounds.
[0180] A request to generate the third set of the multiple sounds may be transmitted to the third server. The request may be received at the third server, and then the third server may generate the third set of the sounds in the multiple sounds. The third set of the sounds may be transmitted back to the second server for combination with the second set of the sounds and then forward-transmitted to the second server in the form of the generation of the third set of the sounds, or transmitted after combination with the second set of the sounds (generated at the second server).
[0181] The generated sounds are received by the user device and played back using the output interface of the user device. In some embodiments, the sounds are played back at the user device using timestamps. In some embodiments, some sounds may be played back in an asynchronous manner. The priority values of the sounds may be utilized to determine which sounds may be played back asynchronously. For example, sounds associated with specific objects important to the virtual experience (e.g., gunshots associated with weapon firing) may have relatively high priority values and may be played synchronously with the actions within the virtual experience, while the sounds of the background crowd noise in a stadium may be assigned relatively low priority values and may be played asynchronously at the user device.
[0182] In some embodiments, the virtual experience state information (e.g., game state information) may further include the head orientation of the user associated with the user device. The head orientation of the user may be utilized to adjust the playback of the sounds at the user device.
[0183] Figure 6is a block diagram of an example computing device 600 that can be used to implement one or more features described herein. In one example, the device 600 can be used to implement a computer device (e.g., Figure 1 102 and / or 110 of) and perform appropriate method implementations described herein. The computing device 600 can be any suitable computer system, server, or other electronic or hardware device. For example, the computing device 600 can be a mainframe computer, a desktop computer, a workstation, a portable computer, or an electronic device (portable device, mobile device, cellular phone, smartphone, tablet computer, television, set-top box, personal digital assistant (PDA), media player, gaming device, wearable device, etc.). In some embodiments, the device 600 includes a processor 602, a memory 604, an input / output (I / O) interface 606, and an audio / video input / output device 614.
[0184] The processor 602 can be one or more processors and / or processing circuits to execute program code and control the basic operations of the device 600. A "processor" includes any suitable hardware and / or software system, mechanism, or component that processes data, signals, or other information. The processor can include a system having a general central processing unit (CPU), multiple processing units, dedicated circuits for implementing functions, or other systems. Processing need not be limited to a particular geographical location nor have a time limit. For example, the processor can perform its functions in a "real-time", "offline", "batch mode", etc. manner. Parts of the processing can be performed by different (or the same) processing systems at different times and different locations. A computer can be any processor that communicates with a memory.
[0185] Memory 604 is typically provided in device 600 for access by processor 602 and can be any suitable processor-readable storage medium, e.g., random access memory (RAM), read-only memory (ROM), electrical erasable read-only memory (EEPROM), flash memory, etc. Memory 604 is suitable for storing instructions for execution by the processor and is separate from and / or integrated with processor 602. Memory 604 can store software operated by processor 602 on server device 600, including operating system 608, one or more applications 610 (e.g., audio spatialization application), and application data 612. In some embodiments, application 610 can include instructions that enable processor 602 to perform (or control) the functions described herein, e.g., with respect to Figure 4 and Figure 5 the portions or all of the methods described.
[0186] For example, application 610 can include an audio spatialization module, which, as described herein, can provide audio spatialization within an online virtual experience server (e.g., 102). Any software in memory 604 can optionally be stored in any other suitable storage location or computer-readable medium. Additionally, memory 604 (and / or other connected storage devices) can store instructions and data used in the features described herein. Memory 604 and any other type of memory (disk, optical disk, magnetic tape, or other tangible medium) can be considered "memory" or "storage device".
[0187] I / O interface 606 can provide functionality that enables server device 600 to interface with other systems and devices. For example, network communication devices, storage devices (e.g., memory and / or data storage area 108), and input / output devices can communicate via interface 606. In some embodiments, the I / O interface can be connected to an interface device that includes input devices (keyboard, pointing device, touch screen, microphone, camera, scanner, etc.) and / or output devices (display device, speaker device, printer, motor, etc.).
[0188] Audio / video input / output device 614 can include user input devices (e.g., mouse, etc.) that can be used to receive user input, display devices (e.g., screen, monitor, etc.) that can be used to provide graphical and / or visual output, and / or combined input and display devices.
[0189] For ease of illustration, Figure 6A box is shown for each of the processor 602, the memory 604, the I / O interface 606, and the software blocks 608 and 610. These boxes may represent one or more processors or processing circuits, an operating system, a memory, an I / O interface, an application, and / or a software engine. In other embodiments, the device 600 may not have all of the components shown and / or may have other types of elements including elements alternative to or in addition to those shown herein. Although the online virtual experience server 102 is described as performing the operations as described in some embodiments herein, any suitable component or combination of components of the online virtual experience server 102 or a similar system, or any suitable one or more processors associated with such a system, may perform the described operations.
[0190] A user device may also implement and / or be used in conjunction with the features described herein. Example user devices may be computer devices that include some components similar to those of the device 600, such as a processor 602, a memory 604, and an I / O interface 606. An operating system, software, and applications suitable for the user device may be provided in the memory and used by the processor. The I / O interface for the user device may be connected to network communication devices and input and output devices, such as a microphone for capturing sound, a camera for capturing images or video, a mouse for capturing user input, a gesture device for recognizing user gestures, a touch screen for detecting user input, an audio speaker device for outputting sound, a display device for outputting images or video, or other output devices. For example, a display device within the audio / video input / output device 614 may be connected to (or included in) the device 600 to display pre-processed and post-processed images as described herein, where such display device may include any suitable display device, such as an LCD, LED, or plasma display screen, a CRT, a television, a monitor, a touch screen, a 3-D display, a projector, or other visual display device. Some embodiments may provide an audio output device, such as a synthetic voice for voice output or reading text.
[0191] One or more methods described herein (e.g., method 400) can be implemented by computer program instructions or code that can be executed on a computer. For example, the code can be implemented by one or more digital processors (e.g., microprocessors or other processing circuits) and can be stored on a computer program product including a non-transitory computer-readable medium (e.g., a storage medium), such as a magnetic, optical, electromagnetic, or semiconductor storage medium, including semiconductor or solid-state memory, magnetic tape, removable computer floppy disk, random access memory (RAM), read-only memory (ROM), flash memory, rigid disk, optical disk, solid-state storage drive, etc. The program instructions can also be embodied in and provided as an electronic signal, e.g., in the form of software as a service (SaaS) delivered from a server (e.g., a distributed system and / or a cloud computing system). Optionally, one or more methods can be implemented in hardware (e.g., logic gates, etc.) or a combination of hardware and software. Example hardware can be a programmable processor (e.g., a field-programmable gate array (FPGA), a complex programmable logic device), a general-purpose processor, a graphics processor, an application-specific integrated circuit (ASIC), etc. One or more methods can be executed as part of or a component of an application running on a system, or as an application or software running together with other applications and an operating system.
[0192] One or more methods described herein can be run in a stand-alone program that can run on any type of computing device, a program running on a web browser, or a mobile application (“app”) running on a mobile computing device (e.g., a phone, a smartphone, a tablet, a wearable device (a watch, an armband, jewelry, a headpiece, goggles, glasses, etc.), a laptop computer, etc.). In one example, a client / server architecture can be used, e.g., a mobile computing device (as a user device) sends user input data to a server device and receives final output data from the server for output (e.g., for display). In another example, all computations can be performed within a mobile application (and / or other applications) on the mobile computing device. In another example, the computations can be split between the mobile computing device and one or more server devices.
[0193] Although described with respect to specific embodiments herein, these specific embodiments are for illustration only and not limitation. The concepts shown in the examples can be applied to other examples and embodiments.
[0194] Note that the functional blocks, operations, features, methods, devices, and systems described in this disclosure can be integrated or divided into different combinations of systems, devices, and functional blocks known to those skilled in the art. Any suitable programming language and programming technique can be used to implement the routines of a particular embodiment. Different programming techniques can be employed, for example, procedural or object-oriented. The routines can be executed on a single processing device or multiple processors. Although steps, operations, or calculations are presented in a particular order, the order can be changed in different particular embodiments. In some embodiments, multiple steps or operations shown as being executed sequentially in this specification can be executed simultaneously.
Claims
1. A computer-implemented method, comprising: receiving, at a server, a first request to generate a plurality of sounds for a user device, wherein the user device is associated with a virtual experience hosted by the server; obtaining, by the server, sound source data of a plurality of sound sources in the virtual experience, each sound source being associated with a specific sound among the plurality of sounds; obtaining, by the server, virtual experience state information, the virtual experience state information including the position of a virtual microphone in the virtual experience and at least one of the following: the speed of the virtual microphone in the virtual experience or the orientation of the virtual microphone in the virtual experience; generating, by the server, an audio mix of the plurality of sounds based on the sound source data and the virtual experience state information; and transmitting the audio mix to the user device.
2. The computer-implemented method according to claim 1, wherein, transmitting the audio mix to the user device includes providing the audio mix encoded in a streaming audio format.
3. The computer-implemented method according to claim 1, wherein, the first request further includes a priority value for at least one sound source among the plurality of sound sources.
4. The computer-implemented method according to claim 3, wherein, the priority value is based on one or more of the following: the loudness of the at least one sound source and the distance between the at least one sound source and the virtual microphone in the virtual experience.
5. The computer-implemented method according to claim 1, wherein, obtaining the virtual experience state information further includes obtaining the head orientation of a user associated with the user device.
6. The computer-implemented method according to claim 1, wherein, generating the audio mix of the plurality of sounds includes: for each sound source among the plurality of sound sources: generating, based on the corresponding sound source data, an audio segment of the sound source; and applying to the audio segment at least one of the following: loudness adjustment based on the distance between the sound source and the virtual microphone in the virtual experience, or Doppler adjustment based on the speed of the virtual microphone in the virtual experience; and after the applying, combining the audio segments of the plurality of sounds to generate the audio mix.
7. The computer-implemented method according to claim 6, further comprising receiving a second request to generate a second plurality of sounds associated with a second virtual microphone participating in the virtual experience, and wherein, generating the audio mix of the plurality of sounds includes: for each sound source: applying to the generated audio segment at least one of the following: a second loudness adjustment based on the distance between the sound source and the second virtual microphone in the virtual experience; and a second Doppler adjustment based on the speed of the second virtual microphone in the virtual experience; and after applying at least one of the second loudness adjustment and the second Doppler adjustment, combining the audio segments of the plurality of sounds to generate a second audio mix.
8. The computer-implemented method according to claim 1, wherein, The plurality of sound sources includes at least one scenario sound source and at least one non-scenario sound source.
9. The computer-implemented method according to claim 1, wherein, generating the audio mix of the plurality of sounds includes: generating a first set of the plurality of sounds at the server; and transmitting a request to a second server to generate a second set of the plurality of sounds, wherein the first set and the second set are mutually exclusive.
10. The computer-implemented method according to claim 9, wherein, generating the first set of sounds includes generating one or more sounds of a sound source associated with a priority value that meets a predetermined priority value threshold.
11. The computer-implemented method according to claim 9, further including: receiving, at the second server, the request to generate the second set of the plurality of sounds; generating, at the second server, a portion of the second set of the plurality of sounds; and transmitting a request to a third server to generate a third set of the plurality of sounds.
12. The computer-implemented method according to claim 1, wherein, obtaining the position of the virtual microphone includes: obtaining the position of a virtual camera placed within the virtual experience; and determining the position of the virtual microphone based on the position of the virtual camera.
13. The computer-implemented method according to claim 1, wherein, the position of the virtual microphone includes: obtaining the position of an avatar within the virtual experience; and determining the position of the virtual microphone based on the position of the avatar.
14. The computer-implemented method according to claim 1, wherein, obtaining the position of the virtual microphone includes: obtaining the position of a virtual camera placed within the virtual experience; obtaining the position of an avatar within the virtual experience; and determining the position of the virtual microphone based on the position of the virtual camera.
15. The computer-implemented method according to claim 14, wherein, determining the position of the virtual microphone such that the virtual microphone is equidistant from the position of the virtual camera and the position of the avatar.
16. A non-transitory computer-readable medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations including the following: receiving, at a server, a first request to generate a plurality of sounds for a user device, wherein, the user device is associated with an avatar participating in a virtual experience hosted by the server; the server obtaining sound source data of a plurality of sound sources associated with the plurality of sounds; the server obtaining virtual experience state information, the virtual experience state information including the position of a virtual microphone in the virtual experience and at least one of the following: the speed of the virtual microphone in the virtual experience or the orientation of the virtual microphone in the virtual experience; the server generating an audio mix of the plurality of sounds based on the sound source data and the virtual experience state information; and transmitting the audio mix to the user device.
17. The non-transitory computer-readable medium according to claim 16, wherein, Transmitting the audio mix to the user device includes providing the audio mix encoded in a streaming audio format.
18. The non-transitory computer-readable medium according to claim 16, wherein, the first request further includes a priority value for at least one of the plurality of sound sources.
19. The non-transitory computer-readable medium according to claim 18, wherein, the priority value is based on one or more of the following: the loudness of the at least one sound source and the distance of the at least one sound source from the virtual microphone in the virtual experience.
20. A system, comprising: a memory storing instructions thereon; and a processing device coupled to the memory, the processing device for accessing the memory and executing the instructions, wherein the instructions cause the processing device to perform operations, the operations including: receiving, at the server, a first request to generate a plurality of sounds for a user device, wherein the user device is associated with a virtual experience hosted by the server; the server obtaining sound source data of a plurality of sound sources associated with the plurality of sounds; the server obtaining virtual experience state information, the virtual experience state information including the position of a virtual microphone in the virtual experience and at least one of the following: the speed of the virtual microphone in the virtual experience or the orientation of the virtual microphone in the virtual experience; the server generating an audio mix of the plurality of sounds based on the sound source data and the virtual experience state information; and transmitting the audio mix to the user device.
21. The system according to claim 20, wherein, obtaining the virtual experience state information further includes obtaining the head orientation of a user associated with the user device.
22. The system according to claim 20, wherein, generating the audio mix of the plurality of sounds includes: for each of the plurality of sound sources: generating an audio segment of the sound source based on the corresponding sound source data; and applying at least one of the following to the audio segment: loudness adjustment based on the distance of the sound source from the virtual microphone in the virtual experience; and Doppler adjustment based on the speed of the virtual microphone in the virtual experience; and after the application, combining the audio segments of the plurality of sounds to generate the audio mix.
23. The system according to claim 20, wherein, the plurality of sound sources includes at least one plot sound source and at least one non-plot sound source.
24. The system according to claim 20, wherein, generating the audio mix of the plurality of sounds includes: generating, at the server, a first set of the plurality of sounds; and transmitting a request to generate a second set of the plurality of sounds to a second server, wherein the first set and the second set are mutually exclusive.