System and method for training model that determines type of environment surrounding user
By using microphones to generate acoustic profiles and integrate machine learning, players can navigate and interact with their environment in multiplayer games, addressing the challenge of visual impairment in HMDs.
Patent Information
- Application Number
- JP2025122163
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-03
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-28
AI Technical Summary
In multiplayer games, players wearing head-mounted displays (HMDs) cannot see their environment, making it difficult to navigate and interact with objects in their surroundings.
Utilizing microphones to detect acoustic characteristics of the environment, generating acoustic profiles, and employing machine learning to identify objects and environments through sound reflections and reverberations, which can be integrated with visual data for a comprehensive understanding of the surroundings.
Enables users to identify and navigate their environment without visual input, providing safety and enhancing interaction with objects, even in virtual reality scenarios.
Smart Images

Figure 2025163066000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to systems and methods for training models that determine the type of environment surrounding a user. [Background technology]
[0002] In a multiplayer game, there are multiple game players. Each player wears a head-mounted display (HMD) to play the game or view the application's environment. While playing the game or running the application, each player may not be able to see the environment in front of them.
[0003] It is in this context that embodiments of the present invention arise. Summary of the Invention
[0004] Embodiments of the present disclosure provide systems and methods for training a model to determine the type of environment surrounding a user.
[0005] In embodiments, one or more microphones are used to detect acoustic characteristics of the current environment. An example of the current environment is a location, such as a room in a house or building where the user is located, or an outdoor environment. Another example of the current environment is a location, such as a room in a building, where the user is playing a video game and making selections to generate input to the video game.
[0006] In one embodiment, an acoustic profile is generated in response to a user request to generate an acoustic profile of an environment. The acoustic profile is generated by capturing sounds in an environment, such as a space or room, including sound reflections and reverberations from various objects located in the environment. The configuration of the environment around the user or around a microphone or microphone array is used to define the current environment.
[0007] In an embodiment, the environment includes sounds generated as background noise. Background noise can be generated by other people talking, music playing, extraneous external sounds, noises made by mechanical objects in or around the environment, etc. By profiling the soundscape of an environment, a machine learning system can be used to identify specific distinguishing characteristics. For example, training can be performed by running an environment profiler application in many different environments. Over time, the learning process using the machine learning system identifies objects within or around the various environments that generate noise or sound. During profiling, the machine learning system also identifies the acoustics of the various environments in which sound is being monitored. Once the machine learning system has sufficiently processed the training, it automatically identifies things, objects, or sounds present in the current environment using a machine learning model trained based on the sound profiles of the various environments. These objects present in the current environment generate and reflect sounds, which are detected and create the acoustics of the current environment. For example, if a user is playing a game in front of one or more monitors, there will be reflections of sounds coming from the monitors during gameplay. These types of reflections are processed using a machine learning system to identify unique characteristics and determine that the user is playing the game in front of another monitor. This type of characteristic identification can be used to detect other objects in the user's current environment.
[0008] In an embodiment, the microphone array may be mounted on a user peripheral device such as eyeglasses, augmented reality (AR) glasses, a head-mounted display (HMD) device, or a handheld controller.
[0009] In one embodiment, during training, if the microphone array is located on glasses or an HMD, the user is asked to turn their head in different environments to capture different sounds. In another embodiment, the configuration of the different environments is done passively, and the audio signals and acoustics of objects in the different environments are tracked over time as the user moves from one of the different environments to another.
[0010] In embodiments, profiling of various environments for acoustic properties provides a type of acoustic vision of the current environment. For example, if a microphone array is located on an HMD or AR glasses or eyeglasses, as a user moves and looks around various environments, a machine learning system can almost instantly identify what is in front of the user based on acoustic reflections and bounceback of signals from the current environment.
[0011] In embodiments, as a person moves within the current environment, the acoustic signals in front of the user change, and based on the profile of the acoustic signals, acoustic vision can be used to identify or verify what may be in front of the user. In one embodiment, acoustic vision of various environments can be blended with data received from a camera to identify or verify the presence of an object based on the acoustic profile of the object in front of the user in the current environment.
[0012] In another embodiment, a virtual profile of a space can be created. For example, if a user wants to appear to be located in a particular location, such as a concert, park, gaming event, studio, etc., a machine learning system can generate sound using a known acoustic profile. The sound generated based on the acoustic profile can be blended with the sound generated by an application program, so that the user appears to be located in a particular location rather than the user's actual location. For example, if a user is publishing a YouTube®™ video, the generated soundscape can be customized for the user based on the type of current environment they want to project or virtually project to third parties watching the YouTube®™ video. For example, if a user wants to provide commentary for a sporting event, a virtual soundscape can be generated behind the commentary to mimic the sporting event, taking into account the acoustic profile that is or may be present at the sporting event.
[0013] In an embodiment, a method for mapping the acoustic properties of materials, surfaces, and geometry of a space is described. The mapping is performed by extracting reverberation information, such as sound reflections and diffusion, detected by one or more microphones. Audio data, once sound-based and captured by one or more microphones, is separated into direct and reverberant components. The direct component is then resynthesized with different reverberant characteristics, effectively replacing or modifying the listener's acoustic environment with some other acoustic profile. The other acoustic profiles are used in combination with visual or geometric mapping of the space, such as with a camera, SLAM (Simultaneous Localization and Mapping), or LiDAR (Light Detection and Ranging), to build a more complete audiovisual mapping. The reverberant component is also used to convey features about the space's geometry and material and surface properties.
[0014] In one embodiment, a method for determining an environment in which a user is located is described. The method includes receiving a plurality of audio data sets based on sounds emitted in the plurality of environments, each of the plurality of environments having a different combination of objects. The method further includes receiving input data regarding the plurality of environments and training an artificial intelligence (AI) model based on the plurality of audio data sets and the input data. The method includes applying the AI model to audio data captured from an environment surrounding a first user to determine a type of environment.
[0015] Some advantages of the systems and methods described herein include helping a user or robot learn objects in an environment without having to acquire images of the objects. For example, a robot learns the identity of objects in an environment, the placement of the objects, and the state of the objects without acquiring images of the objects. Once the robot has learned about the objects, it can be programmed to navigate around the objects and then deployed and used in the environment.
[0016] Further advantages of the systems and methods described herein include providing a visually impaired person with a layout of an environment, which is determined before the visually impaired person visits the environment based on sounds emitted by objects within the environment, making the layout perceptible to the visually impaired person.
[0017] Further advantages of the systems and methods described herein include providing the identity and location of objects in an environment in front of a user when the user is wearing an HMD. When a user wears an HMD, such as in virtual reality (VR) mode, the user may not be able to see the environment in front of them. The systems and methods described herein facilitate providing the identity and location of objects to a user to prevent accidents between the user and the objects.
[0018] Other aspects of the present disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the embodiments described in the present disclosure.
[0019] The various embodiments of the present disclosure can be best understood by referring to the following description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0020] [Figure 1A-1] FIG. 1 is a diagram of an embodiment of an environment in which a user is playing a game. [Figure 1A-2] 1 is a diagram of an embodiment of a system including an input controller, glasses, and a server system. [Figure 1A-3] FIG. 10 is a diagram of an embodiment showing a training session in which a user is asked to turn their head to capture a view of the environment. [Figure 1A-4] FIG. 1 illustrates an embodiment of a list of objects displayed by a graphical processing unit (GPU) to generate an acoustic profile of multiple objects. [Figure 1B] FIG. 1 illustrates an embodiment of an environment in which multiple users are playing a game. [Figure 2] FIG. 1 is a diagram of an embodiment of a system for illustrating processing an input data set and an audio data set to train an artificial intelligence (AI) model. [Figure 3] FIG. 1 is a diagram of an embodiment of a system for illustrating multiple correlations between an audio data set and an environmental system. [Figure 4A] FIG. 1 is a diagram of an embodiment of an object state. [Figure 4B] FIG. 1 is a diagram of an embodiment of object types. [Figure 5A] FIG. 1 is a diagram of an embodiment of an environment to illustrate the use of an AI model to identify objects within an environment, determine the state of objects within the environment, and determine the placement of objects within the environment. [Figure 5B] FIG. 1 is a diagram of an embodiment of a model output from an AI model. [Figure 6A] FIG. 1 is a diagram of an embodiment of an acoustic profile database stored on one or more memory devices. [Figure 6B] 1 is a diagram of an embodiment of a method for illustrating applying audio data to an ambient system to output a corresponding sound. [Figure 7] FIG. 7 is a diagram of an embodiment of a system for illustrating the method of FIG. 6. [Figure 8] FIG. 1 is a diagram of an embodiment of a system for illustrating a microphone. [Figure 9] FIG. 1 is a diagram of an embodiment of a system for illustrating how direct audio data is used to create environmental system effects and reverberant audio data is used to determine the placement of objects within the environmental system, the material type of the objects, and the surface type of the objects. DETAILED DESCRIPTION OF THE INVENTION
[0021] Systems and methods for training a model to determine the type of environment surrounding a user are described. It should be noted that various embodiments of the present disclosure may be practiced without some or all of these specific details. In other instances, well-known process operations have not been described in detail in order to not unnecessarily obscure various embodiments of the present disclosure.
[0022] FIG. 1A-1 is a diagram of an embodiment of an environment 102 in which a user 1 is playing a game G1. An example of the environment 102 is a room in a house 110. The environment 102 includes multiple objects 108A, 108B, 108C, 108D, 108E, 108F, 108G, 108H, 108I, 108J, 108K, 108L, 108M, 108N, and 108O. Object 108A is a chair, object 108B is a side table, object 108C is a display device such as a desktop monitor or television monitor, and object 108D is a window in a wall of the room in the house 110. Object 108D is a window with closed blinds. Object 108E is a desktop table on which a display device is placed, object 108F is a keyboard coupled to the display device, and object 108G is a mouse coupled to the display device. For example, the keyboard is coupled to a central processing unit (CPU) of the display device, and the mouse is also coupled to the CPU.By way of example, object 108E is made of plastic and has a smooth top surface.
[0023] Object 108H is a carpet covering the floor of a room, and object 108I is the floor of the room. Object 108J is a wall on which a window is located. Object 108K is a speaker coupled to a display device, and object 108L is another speaker coupled to the display device. For example, the speaker is coupled to the CPU of the display device. Note that object 108K is in front of object 108C. Also, object 108M is a soda can or container, and object 108N is a handheld controller coupled to game console 112. A soda can is open in environment 102. Game console 112 is coupled to the display device. For example, game console 112 is coupled to the CPU or graphical processing unit (GPU) of the display device. Object 108O is glasses worn by user 1. Examples of glasses described herein include a head-mounted display (HMD), prescription glasses, and augmented reality (AR) glasses.
[0024] Note that object 108P, a vehicle, is located outside of house 110 and passes by house 110 while user 1 is playing game G1. Also note that while user 1 is playing game G1, child 1 and child 2 are playing and talking to each other in separate rooms inside house 110. Each of child 1 and 2 is an example of an object.
[0025] The glasses include a camera C1 and a microphone M1 that converts sound into electrical energy. Examples of camera C1 include a depth camera, a video camera, and a digital camera.
[0026] User 1 accesses game G1 from a game cloud, such as a server system, via a computer network and plays game G1, which renders a virtual scene 114 on a display device. For example, User 1 selects one or more buttons on a handheld controller to provide authentication information, such as a username and password. The handheld controller transmits the authentication information to game console 112, which forwards the authentication information to the game cloud via the computer network. An authentication server in the game cloud determines whether the authentication information is authentic and, if so, provides access to User Account 1 and a game program running on the game cloud. When the game program is executed by one or more processors in the game cloud, image frames of game G1 are generated and encoded, and the encoded image frames are output. The encoded image frames are transmitted to game console 112. Game console 112 decodes the encoded image frames and provides them to a display device to display the virtual scene 114 of game G1, further enabling User 1 to play game G1.
[0027] While playing the game G1, sounds of the game G1 are output from a speaker coupled to the display device. For example, when a virtual object 116A in the virtual scene 114 shoots another virtual object 116B, the sound of the shooting is output. For example, the virtual object 116A is controlled by the user 1 via a handheld controller. Also, while playing the game G1, a vehicle makes sounds such as honking its horn, engine noise, or tire screeching. Also, while playing the game G1, the children 1 and 2 make sounds by talking to each other, arguing with each other, playing with each other, etc. Also, while playing the game G1, when the user 1 opens a can, the user 1 makes the sound of opening a can. Furthermore, while playing the game G1, the user 1 utters words, which are examples of sounds.
[0028] Microphone M1 captures audio data, such as audio frames, generated from sounds associated with, e.g., sounds emanating from or reflected from, one or more of objects 108A-108O located within environment 102. For example, microphone M1 captures audio data generated based on sound emanating from object 108K and received from object 108K via path 106A. In this example, sound path 106A is a direct path from the speaker to microphone M1 and does not impinge on any other surfaces between the speaker and microphone M1. As another example, microphone M1 captures audio data generated based on sound emanating from object 108K and received from object 108K via path 106B. In this example, sound path 106B is an indirect path from object 108K to microphone M1. For purposes of illustration, sound emanating from object 108K impinges on one or more other objects, such as a display device, within environment 102 and is reflected from those one or more other objects toward microphone M1. As yet another example, microphone M1 captures audio data generated based on sound emanating from object 108K and received from object 108K via path 106D. In this example, sound path 106D is an indirect path from the speaker to microphone M1. For purposes of illustration, sound emanating from object 108K impinges on one or more other objects, such as a carpet, in environment 102 and is reflected from those one or more other objects toward microphone M1.
[0029] As another example, microphone M1 captures audio data generated based on sound emanating from object 108L and received from object 108L via path 106D. In this example, sound path 106D is an indirect path from the speaker to microphone M1. For illustrative purposes, sound emanating from object 108L strikes one or more other objects, such as a can, in environment 102 and is reflected from those one or more other objects toward microphone M1. As yet another example, microphone M1 captures audio data generated based on sound emanating from object 108L and received from object 108L via path 106E. In this example, sound path 106E is a direct path from the speaker to microphone M1 and does not strike any other surfaces between the speaker and microphone M1. As yet another example, microphone M1 captures audio data generated based on sound emanating from object 108L and received from object 108L via path 106F. In this example, sound path 106F is an indirect path from the speaker to microphone M1. For purposes of illustration, sound emanating from object 108L impinges on one or more other objects in environment 102, such as a desktop table, and is reflected from those one or more other objects toward microphone M1.
[0030] The glasses' microphone M1 also captures audio data generated based on sounds, such as background noise, emanating from one or more of object 108P, child 1, and child 2, which are located outside but near environment 102. The environment outside environment 102 may be referred to herein as external environment 116. As an example, microphone M1 captures audio data generated based on sounds emanating from object 108P and received from object 108P via path 106G. In this example, path 106G extends through a wall or a door or doorway of environment 102. In this example, if a door of environment 102 is open, the sound spreads through the doorway, and if the door is closed, the sound spreads through the door. As another example, microphone M1 captures audio data generated based on sounds emanating from child 1 or child 2, or both child 1 and child 2, and received from child 1 or child 2 or both child 1 and child 2 via path 106H. In this example, the path 106H extends through a wall or door or doorway of the environment 102.
[0031] Note that an external environment 116 is in proximity to the environment 102 if, for example, a sound emanating from an object in the external environment 116, such as a room next to the environment 102 or a street outside the house 110, can reach and be detected by the microphone M1. For example, a sound emanating from the external environment 116 passes through the wall of the environment 102 and is detected by the microphone M1.
[0032] Microphone M1 captures audio data, such as audio frames generated from sounds associated with objects 108A-108O in environment 102, such as sounds emitted from or reflected by those objects, and background noises emanating from one or more of object 108P, child 1, and child 2 in external environment 116. Encoded audio frames are generated based on the audio frames captured by microphone M1 and provided to one or more processors of the game cloud via a computer network for processing and training artificial intelligence (AI) models.
[0033] In embodiments, sound as used herein includes sound waves.
[0034] In one embodiment, the terms object and item are used interchangeably herein.
[0035] In one embodiment, the display device includes a memory device, and the CPU of the display device is coupled to the memory device.
[0036] In the embodiment, the sound emitted by the speaker is reflected from User 1 and captured by the microphone M1. In this embodiment, User 1 is an example of an object.
[0037] In an embodiment, the virtual scene 114 is not displayed on a display device, but rather on a display screen of glasses worn by the user 1 .
[0038] In this embodiment, User 1 accesses game G1 from the game cloud without needing to use the game console 112. For example, encoded image frames of game G1 are transmitted from the game cloud over a computer network to glasses or a display device placed on a desktop table without being transmitted to the game console 112 for video decoding. The encoded image frames are decoded by the glasses or the display device. In this embodiment, encoded audio frames generated based on audio frames output from microphone M1 are transmitted from the glasses over a computer network to the game cloud without using the game console 112.
[0039] In one embodiment, the virtual scene 114 includes other virtual objects, and based on the movement of these other virtual objects, sounds are output from speakers placed on a desktop table.
[0040] In an embodiment, instead of or in addition to microphone M1, one or more additional microphones, such as standalone microphones, are present to capture sounds emanating from objects in environment 102 and from objects located in external environment 116. For example, a display device located on a desktop table includes an additional microphone. As another example, environment 102 includes one or more standalone microphones. As yet another example, a handheld controller includes an additional microphone.
[0041] In one embodiment, User 1 is not playing game G1. In this embodiment, instead of a game program, one or more processors in the game cloud execute another application program, such as a video conferencing application program or a multimedia program. For example, when the video conferencing application executes, video image frames captured from an additional environment are transmitted over a computer network to a display device or game console or to glasses worn by User 1.
[0042] In one embodiment, the environment 102 is replaced by an outdoor environment such as a concert, a lake, or a park.
[0043] In one embodiment, instead of or in addition to the sound output from object 108K or 108L, sound is output from a speaker integrated within object 108C and detected by microphone M1, and audio data is captured.
[0044] FIG. 1A-2 is a diagram of an embodiment of system 118, including input controller 122, glasses 120, and server system 136. Server system 136 is an example of a game cloud. Glasses 120 are an example of object 108O (FIG. 1A-1). A handheld controller is an example of input controller 112. Glasses 120 include camera C1, video encoder 124, audio encoder 125, network transport device 126, video decoder 128, GPU 130, display screen 132, and CPU 135. An example of video encoder 124 is a circuit that applies a video conversion protocol, such as a video encoding protocol, to image frames to output encoded data, such as encoded image frames, and provides the encoded data to network transport device 126. For illustrative purposes, video encoder 124 generates I-frames, P-frames, or B-frames, which are examples of encoded image frames. Examples of video encoding protocols, such as video compression protocols, include H.262, H.263, and H.264. An example of the audio encoder 125 is a circuit that compresses audio frames into encoded audio frames. For illustrative purposes, the audio encoder applies an audio encoding protocol, such as lossless or lossy compression, to encode the audio frames into encoded audio frames. An example of lossy compression includes a modified discrete cosine transform (MDCT) for converting a sampled waveform in the time domain to the frequency domain. Another example of lossy compression is linear predictive coding (LPC), which analyzes audio data generated based on speech sounds.
[0045] An example of a network transfer device as used herein is a network interface controller such as a network interface card (NIC). Another example of a network transfer device is a wireless access card (WAC). An example of a video decoder as used herein is a circuit that performs H.262, H.263, or H.264 decoding or another type of video decompression and outputs decoded data such as image frames. An example of a display screen 132 is a liquid crystal display (LCD) screen or a light-emitting diode (LED) display screen. An example of a device's communication device is a circuit that communicates with another device or system using a communication protocol, such as a wired communication protocol or a wireless communication protocol. Examples of CPU 135 include a processor, microprocessor, microcontroller, application-specific integrated circuit (ASIC), and programmable logic device (PLD). Camera C1 includes a lens L1 facing an environment, such as environment 102 (FIG. 1A-1). For example, camera C1 is an outward-facing camera.
[0046] The CPU 135 is coupled to other components of the glasses 120. For example, the CPU 135 is coupled to the communication device 134, the camera C1, the network transfer device 126, the video encoder 124, the GPU 130, the video decoder 128, the audio encoder 125, and the microphone M1, and controls other components of the glasses 120. The camera C1 is also coupled to the video encoder 124, which is coupled to the network transfer device 126. The network transfer device 126 is coupled to a computer network 142, the audio encoder 125, and the video decoder 128. The audio encoder 125 is coupled to the microphone M1. The video decoder 128 is coupled to the CPU 130, which is coupled to the display screen 132. Examples of the computer network 142 include a local area network (LAN), a wide area network (WAN), and combinations thereof. For purposes of illustration, the computer network 142 is the Internet or an intranet, or a combination thereof.
[0047] The server system 136 includes a network transport device 138, a video decoder 140, an audio decoder 144, and one or more servers 1-N (N is an integer greater than 0). An example of an audio decoder is a circuit that decompresses encoded audio frames into audio frames. For illustrative purposes, the audio decoder applies an audio decoding protocol (e.g., an audio decompression protocol) to decode the encoded audio frames into audio frames. Each of the servers 1-N includes a processor and a memory device. For example, server 1 includes processor 1 and memory device 1, server 2 includes processor 2 and memory device 2, and server N includes processor N and memory device N. The network transport device 126 is coupled to the computer network 142 and is also coupled to the video decoder 140, which is coupled to one or more of the servers 1-N. The network transport device 126 is coupled to the audio decoder 144, which is coupled to one or more of the servers 1-N. The operation of the system 118 will be described with reference to Figures 1A-3.
[0048] In an embodiment, when glasses 120 are AR glasses, input controller 122 is separate from the handheld controller used to play game G1.
[0049] In one embodiment, glasses 120 include multiple display screens instead of display screen 132. Each display screen has a similar structure and function to display screen 132.
[0050] In an embodiment, the glasses 120 include one or more additional lenses in addition to the lens L1 for capturing an image of an environment, such as the environment 102.
[0051] In one embodiment, glasses 120 include one or more memory devices, such as random access memory (RAM) or read-only memory (ROM), coupled to CPU 135 or GPU 130, or both CPU 135 and GPU 130. For example, CPU 135 includes a memory controller for accessing data from and writing data to the one or more memory devices.
[0052] In an embodiment, the memory controller is a separate device from the CPU 135 .
[0053] FIG. 1A-3 is a diagram of an embodiment illustrating a training session in which user 1 is requested to turn his or her head to capture views of an environment, such as environment 102 (FIG. 1A-1). Once authentication information received from user 1 is authenticated, user 1 is provided with access to user account 1 by an authentication server. User 1 may then access the training session or a game program or other application program. During the training session, before or during execution of the game program or other application program, the training program is executed by one or more processors 1-N of the game cloud to generate a message 121. For example, before or during execution of the game program, server system 136 (FIG. 1A-2) executes the training program to generate image frames having message 121, encodes the image frames to output the encoded image frames having message 121, applies a network transport protocol to the encoded image frames to generate data packets having message 121, and transmits the data packets to glasses 120 (FIG. 1A-2) via computer network 142. In this example, a network transport device 126 (FIG. 1A-2) of the glasses 120 applies a network transport protocol to obtain encoded image frames from data packets and provides the encoded image frames to a video decoder 128 (FIG. 1A-2) of the glasses 120. In this example, the video decoder 128 applies a video conversion protocol, such as a video decoding protocol, to decode the encoded image frames, output image frames having messages 121, and provide the image frames to a GPU 130 (FIG. 1A-2). In this example, the GPU 130 applies a rendering program to display the messages 121 on a display screen 132 (FIG. 1A-2). In this example, the training program is computer software stored on one or more memory devices 1-N of the server system 136. Examples of network transport protocols include Transmission Control Protocol (TCP) over Internet Protocol (IP).
[0054] Upon viewing the message 121, the user 1 turns his / her head to capture a view (e.g., a 360-degree view) of the environment 102. For example, as the user 1 turns his / her head within the environment 102 to view the environment 102, the camera C1 of the glasses 120 captures images of the objects 108A-108O in the environment 102. Referring to FIG. 1A-1 , the camera C1 transmits the images of the objects 108A-108O to the video encoder 124 of the glasses 120. The video encoder 124 applies a video encoding protocol to encode the images received from the camera C1 to output encoded image frames and transmits the encoded image frames to the network forwarding device 126. The network forwarding device 126 applies a network forwarding protocol to generate data packets from the encoded image frames and transmits the data packets to the network forwarding device 138 of the server system 136 via the computer network 142.
[0055] A network transport device 138 of the server system 136 retrieves the data packets from the glasses 120 via a computer network 142 and applies a network transport protocol to the data packets to extract encoded image frames of the objects 108A-108O. The network transport device 138 sends the encoded image frames to a video decoder 140 of the server system 136. The video decoder 140 applies a video decoding protocol to the encoded image frames to output the image frames and provides the image frames to one or more processors 1-N of the game cloud.
[0056] Additionally, microphone M1 generates audio frames based on sounds emanating from user 1 and objects, such as speakers, in environment 102 (FIG. 1A-1) and background noise emanating from object 108P and children 1 and 2 in external environment 116 (FIG. 1A-1). The audio frames are transmitted from microphone M1 to audio encoder 125 of glasses 120. Audio encoder 125 applies an audio encoding protocol to output encoded audio frames. The encoded audio frames are provided from audio encoder 125 to network transport device 126. Network transport device 126 applies a network transport protocol to the encoded audio frames to generate data packets and transmits the data packets to server system 136 via computer network 142.
[0057] A network transport device 138 of the server system 136 receives the data packets from the glasses 120 and applies a network transport protocol to the data packets to output encoded audio frames. The network transport device 128 provides the encoded audio frames to an audio decoder 144. The audio decoder 144 applies an audio decoding protocol to the encoded audio frames to output audio frames and provides the audio frames to one or more processors 1-N for storage in one or more memory devices 1-N.
[0058] It should be noted that although FIGS. 1A-3 are described with reference to a gaming program, FIGS. 1A-3 are equally applicable to other application programs, such as a video conferencing application program.
[0059] In embodiments, although FIGS. 1A-3 are described with reference to user 1, environment 120, and glasses worn by user 1, FIGS. 1A-3 are equally applicable to other users, other environments, and glasses worn by other users.
[0060] 1A-4 is a diagram of an embodiment of a list 150 of objects displayed by GPU 130 (FIG. 1A-2) to generate acoustic profiles for objects 108A-108P (FIG. 1A-1). List 150 includes the state of objects within an environment system. For example, list 150 includes a first entry indicating that there are no blinds covering windows in environment 102 and a second entry indicating that there are blinds covering windows. As another example, list 150 includes a third entry indicating that a can is open and a fourth entry indicating that the can is closed. As yet another example, list 150 includes the placement of objects within environment 102. For illustrative purposes, list 150 includes a fourth entry indicating that the can is closer to microphone M1 than objects 108K or 108L, and a fifth entry indicating that object 108L is closer to microphone M1 than object 108K.
[0061] Instead of, or in addition to, providing message 121, during a training session, before or during execution of the game program, a training program is executed by one or more processors 1-N (FIGS. 1A-2) to generate list 150. For example, before or during execution of the game program, server system 136 executes the training program to generate image frames having list 150, encodes the image frames to output encoded image frames having list 150, applies a network transport protocol to the encoded image frames to generate data packets having list 150, and transmits the data packets over a computer network to glasses 120 (FIG. 1A-2). In this example, network transport device 126 (FIG. 1A-2) of glasses 120 applies the network transport protocol to obtain the encoded image frames from the data packets and provides the encoded image frames to video decoder 128 (FIG. 1A-2) of glasses 120. In this example, video decoder 128 applies a video conversion protocol, such as a video decoding protocol, to decode the encoded image frames, output image frames having list 150, and provide the image frames to GPU 130 (FIG. 1A-2). In this example, GPU 130 applies a rendering program to display list 150 on display screen 132 (FIG. 1A-2). In this example, the training program is computer software stored in one or more of memory devices 1-N of server system 136.
[0062] Upon viewing list 150, user 1 uses input controller 122 to select one or more checkboxes next to one or more items in list 150 to identify objects O108A-O108P and children 1 and 2. The communication device of input controller 122 applies a communication protocol to the selection of one or more checkboxes in list 150 to generate one or more forwarding packets and transmits the forwarding packets to communication device 134 of glasses 120. Communication device 134 applies the communication protocol to the forwarding packets to obtain list 150 from the forwarding packets and transmits the selection of one or more checkboxes in list 150 to CPU 135 of glasses 120. CPU 135 transmits the selection of one or more checkboxes in list 150 to network forwarding device 126 of glasses 120. Network forwarding device 126 applies a network forwarding protocol to the selection of one or more checkboxes in list 150 to generate data packets. Network forwarding device 126 transmits the data packets to one or more processors 1-N of server system 136 via computer network 142. The network forwarding device 138 receives data packets from the computer network 142, applies a network forwarding protocol to the data packets to obtain a selection of one or more checkboxes in the list 150, and provides the selection to one or more processors 1-N of the server system 136.
[0063] In one embodiment, instead of list 150, a list of blank lines is generated by one or more processors 1-N and transmitted to glasses 120 (FIG. 1A-3) over computer network 142 (FIGS. 1A-3). User 1 fills in the blank lines using input controller 122 to provide a list of objects 108A-108P and children 1 and 2.
[0064] FIG. 1B is a diagram of another embodiment of environment 152 in which user 1 and user 2 are playing game G1. Environment 152 is a room in a high-rise building. For example, the room is on the top floor of the building. Environment 152 includes spectators 1 and 2. Each spectator 1 and 2 is an example of an object of environment 152. Environment 152 further includes objects 154A, 154B, 154C, 154D, 154E, 154F, 154G, 154H, 154I, 154J, 154K, 154L, 154M, 154N, 158O, 154P, and 154Q. Object 154A is a banner hanging from the ceiling of environment 152. Object 154B is a display device. An example of a display device is a desktop monitor or a television. Object 154C is a display device, and object 154D is another display device. Each of objects 154C and 154D includes a CPU and a memory device located within the object. User 1 plays game G1, which displays a virtual scene on desktop monitor 154D. User 2 plays game G1, which displays a virtual scene on desktop monitor 154C.
[0065] Object 154E is a speaker coupled to object 154C, and object 154F is a speaker coupled to object 154D. Object 154E is behind object 154C, and object 154F is behind object 154D. Object 154G is a can or container, and object 154H is a table on which objects 154C and 154D are placed. For example, object 154H has a marble top and an uneven surface. In environment 152, the can is closed. Also, object 154I is the floor of environment 152. The floor is not carpeted but bare. For example, the floor has a tiled surface. Object 154J is a chair on which user 1 sits, object 154K is a mouse coupled to object 154D, and object 154L is a keyboard coupled to object 154D. Object 154M is a mouse that is coupled to object 154C, and object 154N is a keyboard that is coupled to object 154C.
[0066] Object 154P is a pair of glasses worn by user 2. Object 154P includes microphone M2 and camera C2. Object 154Q is a cabinet stand on which object 154B is placed. Object 154R, an airplane, is located outside environment 152 and flies above the buildings while users 1 and 2 are playing game G1.
[0067] The environment 152 further includes an object 154S and another object 154T. The object 154S is a window without blinds, and the object 154T is a can or container.
[0068] Note that User 1 moves from one location, such as environment 102 (FIG. 1A-1), to another location, such as environment 152. User 1 accesses game G1 from the game cloud via computer network 142 (FIG. 1A-2) and user account 1 and plays game G1, which causes a virtual scene to be rendered on the display screen of object 154D. For example, User 1 selects one or more buttons on one or more of objects 154L and 154K to provide authentication information, such as a username and password. Objects 154L and 154K transmit the authentication information to object 154D, which forwards the authentication information to the game cloud via computer network 142. An authentication server in the game cloud determines whether the authentication information is authentic and, if so, provides access to User Account 1 and a game program running on the game cloud. When the game program is executed by one or more of processors 1-N in the game cloud, image frames of game G1 are generated and encoded, and the encoded image frames are output. The encoded image frames are transmitted to object 154D. Object 154D decodes the encoded image frames and provides the image frames to the display screen of object 154D in the virtual scene displayed on object 154D, allowing user 1 to play game G1.
[0069] During the play of game G1, sounds of game G1 are output from object 154F. For example, when a virtual object in a virtual scene displayed on object 154D jumps in the virtual scene and lands on the virtual ground, a landing sound is output via object 154F. For example, the virtual object in the virtual scene displayed on object 154D is controlled by user 1 via objects 154L and 154K. During the play of game G1, an airplane flying in environment 152 makes a sound such as a sonic boom. During the play of game G1, spectators 1 and 2 make sounds by talking to each other, arguing with each other, playing with each other, etc. During the play of game G1, user 1 opens a container placed on object 154H, which makes a sound when the container is opened. During the play of game G1, user 1 utters words, which are examples of sounds.
[0070] Both User 1 and User 2 are in the same environment 152 and therefore in the same location. Similarly, User 2 accesses game G1 from the game cloud via computer network 142 and user account 2 and plays game G1, which causes a virtual scene to be rendered on the display screen of object 154C. For example, User 2 selects one or more buttons on objects 154N and 154M to provide authentication information, such as a username and password. Objects 154N and 154M transmit the authentication information to object 154C, which forwards the authentication information to the game cloud via computer network 142. An authentication server in the game cloud determines whether the authentication information is authentic and, if so, provides access to User Account 2 and the game program running on the game cloud. When the game program is executed by one or more of processors 1-N in the game cloud, image frames of game G1 are generated and encoded, and the encoded image frames are output. The encoded image frames are transmitted to object 154C. Object 154C decodes the encoded image frames and provides the image frames to the display screen of object 154C of the virtual scene displayed on object 154C, allowing user 2 to play game G1.
[0071] While playing game G1, sounds of game G1 are output from object 154E. For example, when a virtual object in a virtual scene displayed on object 154C flies within the virtual scene, the sound of the flight is output via object 154E. As an example, the virtual object in the virtual scene displayed on object 154C is controlled by user 2 via objects 154N and 154M. Also, while playing game G1, user 2 opens object 154T placed on object 154H, and opening object 154T makes a sound. Also, while playing game G1, user 2 utters words, which are examples of sounds.
[0072] Each of the microphones M1 and M2 captures audio data, such as audio frames, generated from sounds associated with, such as sounds emanating from or reflected from, one or more of the objects located within the environment 152. For example, the microphones M1 and M2 capture audio data generated based on sounds emanating from the object 154F and received from the object 154F via path 156A. In this example, the sound path 156A is a direct path from the object 154F to the microphones M1 and M2 and does not impinge on any other surfaces between the object 154F and the microphones M1 and M2. As another example, the microphones M1 and M2 capture sounds emanating from the object 154F and received from the object 154F via path 156B. In this example, the sound path 154F is an indirect path from the object 154F to the microphones M1 and M2. To illustrate, sound emanating from object 154F impinges on one or more other objects in environment 152 (eg, object 154D) and is reflected from those one or more other objects towards microphones M1 and M2.
[0073] As yet another example, each microphone M1 and M2 captures audio data generated based on sound emanating from object 154E and received from object 154E via path 156C. In this example, sound path 156C is a direct path from object 154E to microphones M1 and M2 and does not strike any other surfaces between object 154E and microphones M1 and M2. As yet another example, microphones M1 and M2 capture sound emanating from object 154E and received from object 154E via path 156D. In this example, sound path 156D is an indirect path from object 154E to microphones M1 and M2. For purposes of illustration, sound emanating from object 154E strikes one or more other objects (e.g., object 154C) in environment 152 and is reflected from those one or more other objects toward microphones M1 and M2.
[0074] Each microphone M1 and M2 also captures audio data generated based on sounds, such as background noise, emanating from an object 154R located outside the environment 152. The environment outside the environment 152 may be referred to herein as the external environment 158. As an example, microphones M1 and M2 capture sounds emanating from the object 154R and received through the ceiling of the environment 152.
[0075] Microphones M1 and M2 capture sounds associated with objects 154A-154N, 108O, 154P, 154Q, 154S, 154T and one or more of spectators 1 and 2 in environment 152, such as sounds emitted from or reflected therefrom, and background noise emanating from object 154R in external environment 158, to generate audio frames, such as audio data. Encoded audio frames are generated based on the audio frames and provided to one or more of processors 1-N of the game cloud via computer network 142 for processing and AI model training.
[0076] In one embodiment, User 1 plays a different game than game G1, and User 2 plays a different game than game G1.
[0077] In an embodiment, the external environment 158 includes any other number of objects, such as two or three.
[0078] In one embodiment, instead of or in addition to the sound emitted from object 154E, sound is emitted from object 154C and detected by microphone M1 or M2 or a combination thereof, and audio data is captured.
[0079] In an embodiment, instead of or in addition to the sound output from object 154F, sound is output from object 154D and detected by microphone M1 or M2 or a combination thereof, and audio data is captured.
[0080] In an embodiment, instead of or in addition to microphone M2, one or more additional microphones, such as standalone microphones, are present to capture sounds emanating from objects within environment 152 and from objects located in external environment 158. For example, a display device located on a table within environment 152 includes an additional microphone.
[0081] In an embodiment, the virtual scene displayed on object 154D is instead displayed on the display screen of the glasses worn by user 1. Similarly, the virtual scene displayed on object 154C is instead displayed on the display screen of the glasses worn by user 2.
[0082] 2 is a diagram of an embodiment of a system 200 illustrating processing of input datasets 1, 2, and 3 and audio datasets 1, 2, and 3. System 200 includes client device 1, client device 2, and client device 3. System 200 further includes a processor system 202. Examples of client device 1 operated by user 1 include eyeglasses worn by user 1, object 108C (FIG. 1A-1), object 108N, game console 112 (FIG. 1A-1), object 108L (FIG. 1A-1), object 108K (FIG. 1A-1), object 154D (FIG. 1B), object 154L, object 154K, object 154F, and combinations of two or more thereof. Examples of client device 2 operated by user 2 include eyeglasses worn by user 2, object 154C (FIG. 1B), object 154M, object 154N, object 154E, a game console, a handheld controller, and combinations of two or more thereof. Note that glasses 120 (FIGS. 1A-2) are an example of glasses worn by user 2, except that camera C1 is replaced with camera C2 and microphone M1 is replaced with microphone M2. An example of client device 3 is provided below with reference to FIG. 5A. Client device 3 is operated by user 3.
[0083] Processor system 202 is coupled to client devices 1, 2, and 3. For example, processor system 202 is coupled to client devices 1-3 via computer network 142 (FIG. 1B). For purposes of illustration, processor system 202 includes one or more processors 1-N of server system 136 (FIGS. 1A-2) that are coupled to client devices 1-3 via computer network 142.
[0084] An example of audio dataset 1 includes audio data captured by microphone M1 based on sounds associated with environment 102 (FIG. 1A-1) and background noise received from external environment 116 (FIG. 1A-1). An example of audio dataset 2 includes audio data captured by microphone M1 based on sounds associated with environment 152 (FIG. 1B) and background noise received from external environment 158 (FIG. 1B). An example of audio dataset 3 includes audio data captured by microphone M2 based on sounds associated with environment 152 and background noise received from external environment 158.
[0085] An example of input dataset 1 includes one or more images of objects in environment 102 (FIG. 1A-1) captured by camera C1, identifications of objects in environment 102, and identifications of objects in external environment 116 (FIG. 1A-1). For illustrative purposes, example identifications of objects in environment 102 and objects in external environment 116 are received in list 150 (FIG. 1A-4).
[0086] An example of input dataset 2 includes one or more images of objects in environment 152 (FIG. 1B) captured by camera C1, identifications of objects in environment 152 received via user account 1, and identifications of objects in external environment 158 (FIG. 1B) received via user account 1. For purposes of illustration, the identifications of objects in environment 152 and objects in external environment 158 are received in the form of a list via user account 1 from user 1, in the same manner as the identifications of objects in environment 102 and objects in external environment 116 are received via list 150. In this illustration, user 1 selects one or more objects in the list to provide the identifications of objects in environment 152 and objects in external environment 158. The list is generated in the same manner as list 150 is generated and is displayed on object 154D (FIG. 1B) manipulated by user 1.
[0087] An example of input dataset 3 includes one or more images of objects in environment 152 (FIG. 1B) captured by camera C2, identifications of objects in environment 152 received via user account 2, and identifications of objects in external environment 158 (FIG. 1B) received via user account 2. For purposes of illustration, the identifications of objects in environment 152 and objects in external environment 158 are received in the form of a list via user account 2 from user 2 in the same manner as the identifications of objects in environment 102 and objects in external environment 116 (FIGS. 1A-1) are received via list 150. In this illustration, user 2 selects one or more objects in the list to provide the identifications of objects in environment 152 and objects in external environment 158. The list is generated in the same manner as list 150 is generated and is displayed on object 154C (FIG. 1B) manipulated by user 2.
[0088] The processor system 202 includes a game engine and an inference training engine. An example of an engine includes hardware, such as one or more controllers. In this example, each controller includes one or more processors, such as processors 1-N, or one or more processors of the game console 112 (FIG. 1A-1), or a combination thereof. As another example, the engine is software, such as a computer software program. For purposes of illustration, the game engine is a game program. Another example of an engine is a combination of hardware and software.
[0089] The inference training engine includes an AI processor and a memory device 204, which is an example of one of memory devices 1-N. The AI processor is coupled to memory device 204 and is an example of one of processors 1-N. Input data sets 1-3 and audio data sets 1-3 are stored in memory device 204. For example, the AI processor receives input data sets 1-3 and audio data sets 1-3 from client devices 1 and 2 via computer network 142 and stores input data sets 1-3 and audio data sets 1-3 in memory device 204. A game engine is coupled to the inference training engine.
[0090] The AI processor includes a feature extractor, a classifier, and an AI model. For example, the AI processor includes a first integrated circuit that applies the function of the feature extractor, a second integrated circuit that applies the function of the classifier, and a third integrated circuit that applies the function of the AI model. As another example, the AI processor executes a first computer program that applies the function of the feature extractor, a second computer program that applies the function of the classifier, and a third computer program that applies the function of the AI model. The feature extractor is coupled to the classifier, and the classifier is coupled to the AI model.
[0091] The feature extractor extracts, for example, determines, parameters such as one or more amplitudes, one or more frequencies, and one or more sensing directions from the audio datasets 1 to 3. For example, the feature extractor determines the magnitude or peak-to-peak amplitude or zero-to-peak amplitude and frequency of the audio datasets 1 to 3. For illustrative purposes, the feature extractor determines the absolute maximum power of the audio dataset 1 or the absolute minimum power of the audio dataset 1 to determine the magnitude of the audio dataset 1. In this description, the absolute power refers to the magnitude within the entire period during which the audio dataset 1 is generated. In another description, the feature extractor determines a local maximum magnitude and a local minimum magnitude of the audio dataset 1. In this description, the local magnitude refers to the magnitude within a predetermined period, which is shorter than the entire period during which the audio dataset 1 is generated. In this description, multiple local maxima of magnitude and multiple local minima of magnitude are determined from the audio dataset 1, and the feature extractor applies a best fit, average, or median to the local maxima of magnitude and the local minima to determine the maximum magnitude and the minimum magnitude.
[0092] Alternatively, the feature extractor determines a first time when the audio dataset 1 reaches a predetermined magnitude and a second time when the audio dataset 1 reaches the same predetermined magnitude, and calculates the difference between the first and second times to determine a time interval. The feature extractor inverts the time interval to determine an absolute frequency of the audio dataset 1. In this description, the absolute frequency is a frequency within an entire period during which the audio dataset 1 is generated. Alternatively, the feature extractor determines a local frequency of the audio dataset 1. In this description, the local frequency is a frequency within a predetermined period, which is shorter than the entire period during which the audio dataset 1 is generated. In this description, multiple local frequencies are determined from the audio dataset 1, and the feature extractor determines the frequency by applying a best fit, average, or median to the local frequencies. In this description, each local frequency is determined in the same manner as the absolute frequency is determined, except that the local frequency is determined for each predetermined period.
[0093] Alternatively, the feature extractor determines the direction in which the audio data set 1 is sensed. In this example, the microphone M1 includes an array (e.g., a linear array) of transducers arranged in a certain direction. In this example, the array includes a proximal transducer and a distal transducer. In this example, if the proximal transducer outputs a first portion of the audio data 1 and the distal transducer outputs a second portion of the audio data 1, and the first portion has an amplitude greater than the second amplitude, the feature extractor determines that the object 108M (FIG. 1A-1) in the environment 102 (FIG. 1A-1) is closer to the proximal transducer than to the distal transducer. In this example, the feature extractor further determines that the object 108M is in a direction facing the proximal transducer. Alternatively, if the second amplitude is greater than the first amplitude, the feature extractor determines that the object 108M is closer to the distal transducer than to the proximal transducer. In this description, the feature extractor further determines that the object 108M is in a direction pointing towards the distal transducer. The direction in which the audio data set is sensed may be referred to herein as the sensing direction.
[0094] The classifier classifies parameters obtained from sounds associated with environments 102, 116, 152, and 158 based on input datasets 1-3. For example, the classifier determines a combination of objects in an environmental system, such as environment 102 (FIG. 1A-1) or external environment 116 (FIG. 1A-1), or a combination of environments 102 and 116, and establishes a correlation, such as a one-to-one correspondence or a unique relationship, with the parameters determined from audio dataset 1. In this example, the classifier determines the object's location, or the object's state, or a combination of two or more of the object, object's location, and object's state in the environment. For illustrative purposes, the classifier receives, in input dataset 1, the identities of objects 108A-108O in environment 102 and the identities of object 108P, child 1, and child 2 via list 150 and user account 1. In this illustration, the classifier receives, in input dataset 1, the states of objects 108A-108O (FIG. 1A-1). Further, in this description, the classifier determines from the sensing direction that object 108M is located closer to microphone M1 than object 108L and compared to object 108K (FIG. 1A-1). Also, in this description, the classifier determines from the sensing direction that children 1 and 2 are located farther from microphone M1 than objects 108M, 108L, and 108K, determining the locations of children 1 and 2 and objects 108M, 108L, and 108K. Alternatively, the classifier determines from the sensing direction or amplitude or frequency or a combination thereof that object 108P or child 1 or child 2 is located outside environment 102.
[0095] Alternatively, the classifier receives, in input dataset 2, the identities of objects 154A-154N, 108O, 154P, and 154Q in environment 152, as well as the identity of object 154R via a list, such as list 150, and user account 1. In this example, the classifier receives, in input dataset 2, the states of objects 154A-154N, 108O, 154P, and 154Q ( FIG. 1B ). Additionally, in this example, the classifier determines, based on the sensing direction, that object 154D is located closer to microphone M1 than object 154F ( FIG. 1B ). Additionally, in this example, the classifier determines, based on the sensing direction, that object 154R is flying above environment 152, thereby determining the location of object 154R relative to environment 152. Alternatively, the classifier determines, based on the sensing direction, amplitude, or frequency, or a combination thereof, that object 154R is located outside environment 152.
[0096] As yet another illustration, the classifier receives in input dataset 3 the identities of objects 154A-154N, 108O, 154P, and 154Q in environment 152, as well as the identity of object 154R via a list, such as list 150, and user account 2. In this illustration, the classifier receives in input dataset 3 the states of objects 154A-154N, 108O, 154P, and 154Q (FIG. 1B). Further, in this illustration, the classifier determines from the sensing direction that object 154C is located closer to microphone M2 than object 154E (FIG. 1B).
[0097] The AI model is trained based on correlations between parameters associated with environments 102, 116, 152, and 158 (FIGS. 1A-1 and 1B) and input datasets 1-3. For example, referring to FIG. 3 , the AI model is provided by a classifier with an indication of correlations 302 (e.g., links or one-to-one correspondences) between a set including amplitude 1, frequency 1, and sensing direction 1 and a set including a first type of environment, a first combination of objects within the first type of environment, a first arrangement of the objects, and a first state of the objects. In this example, the AI model receives amplitude 1, frequency 1, and sensing direction 1 from the classifier. In this example, amplitude 1, frequency 1, and sensing direction 1 are determined by analyzing audio dataset 1 captured by microphone M1. In this example, the one or more amplitudes determined from audio dataset 1 are referred to herein as amplitude 1. Also, in this example, the one or more frequencies determined from audio dataset 1 are referred to herein as frequency 1, and the direction in which audio dataset 1 is perceived is referred to herein as sensing direction 1. For purposes of illustration, amplitude 1, frequency 1, and sensing direction 1 are example parameters of audio dataset 1. Also in this example, the classifier provides the AI model, via user account 1, with a first type of environment 102, a first combination of objects within the environment 102 and the external environment 116, a first state of the objects, and a first disposition of the objects. For illustrative purposes, the first type of environment includes whether the environment 102 is an indoor environment, such as a room or an open space within a building, or an outdoor environment, such as a park, a concert, or a lake. Alternatively, the first combination of objects includes the identity of the objects 108A-108O, such as a can, a display device, a speaker, or a window. In this example, the first disposition of the objects includes the can being closer to the microphone M1 than the speaker, and the first state includes the can being open or closed. Alternatively, the first combination of objects includes the identity of the objects 108A-108O, such as a window with blinds, a carpet, or a speaker.Also, in this example, the combination of environment 102 and external environment 116 is referred to as a first environmental system.
[0098] As another example, referring to FIG. 3 , an AI model is provided with an indication of a correlation 304 (e.g., a link or one-to-one correspondence) between a set including amplitude 2, frequency 2, and sensing direction 2 and a set including a second type of environment, a second combination of objects within the second type of environment, a second disposition of the objects, and a second state of the objects. In this example, the AI model receives amplitude 2, frequency 2, and sensing direction 2 from a classifier. In this example, amplitude 2, frequency 2, and sensing direction 2 are determined by analyzing audio dataset 2 captured by microphone M1. Also in this example, one or more amplitudes determined from audio dataset 2 are referred to herein as amplitude 2, one or more frequencies determined from audio dataset 2 are referred to herein as frequency 2, and the direction in which audio dataset 2 is perceived is referred to herein as sensing direction 2. For purposes of illustration, amplitude 2, frequency 2, and sensing direction 2 are example parameters of audio dataset 2. Also in this example, the classifier provides the AI model with a second type of environment 152, a second combination of objects within environment 152 and external environment 158, a second state of the objects, and a second disposition of the objects. In this example, the second type, second combination, second state, and second disposition are received via user account 1. For illustrative purposes, the second type of environment includes whether environment 152 is an indoor environment, such as a room or an open space in a building, or an outdoor environment, such as a park, a concert, or a lake. Alternatively, the second combination of objects includes the identities of objects 154A-154N, 154O, 154P-154T, such as cans, displays, speakers, or floors or ceilings. In this example, the second disposition of objects includes object 154G being closer to microphone M1 than object 154F, and the second state includes object 154G being open or closed. Alternatively, the second disposition includes object 154B being farther from microphone M1 than objects 154F or 154D (FIG. 1B).In this illustration, object 154B is a display device emitting sounds from a video game or another application. Also, in this example, the combination of environment 152 and external environment 158 is referred to as a second environment system.
[0099] As yet another example, referring to FIG. 3 , the AI model is provided by the classifier with an indication of a correlation 306 (e.g., a link or one-to-one correspondence) between a set including amplitude 3, frequency 3, and sensing direction 3 and a set including a third type of environment, a third combination of objects within the third type of environment, a third disposition of the objects, and a third state of the objects. In this example, the AI model receives amplitude 3, frequency 3, and sensing direction 3 from the classifier. In this example, amplitude 3, frequency 3, and sensing direction 3 are determined by analyzing audio dataset 3 captured by microphone M2. Also in this example, one or more amplitudes determined from audio dataset 3 are referred to herein as amplitudes 3, one or more frequencies determined from audio dataset 3 are referred to herein as frequencies 3, and the direction in which audio dataset 3 is perceived is referred to herein as sensing direction 3. For illustrative purposes, amplitude 3, frequency 3, and sensing direction 3 are example parameters of audio dataset 3. Also in this example, the classifier provides the AI model with a third type of environment 152, a third combination of objects within environment 152 and external environment 158, a third state of the objects, and a third disposition of the objects. In this example, the third type, third combination, third state, and third configuration are received via user account 2. For illustrative purposes, the third type of environment includes whether environment 152 is an indoor environment, such as a room or an open space within a building, or an outdoor environment, such as a park, a concert, or a lake. Alternatively, the third combination of objects includes the identities of objects 154A-154N, 154O, 154P-154T, such as cans, displays, speakers, floors, or ceilings. In this example, the third configuration of objects includes object 154T being closer to microphone M2 than objects 154E or 154C, and the third state includes object 154T being open or closed. Also in this example, the combination of environment 152 and external environment 158 is referred to as a third environment system.
[0100] In an embodiment, instead of or in addition to receiving identities, such as list 150 (FIGS. 1A-4), of objects in the environmental system, image data, such as image frames, of the environmental system is received from cameras C1 and C2. In this embodiment, a feature extractor identifies objects in the environmental system from the image data. For example, the feature extractor identifies objects 108A-108P from image frames of input dataset 1. In this example, the feature extractor determines that the outline of object 108A matches a pre-stored outline of a chair and determines that object 108A is a chair. In this example, the pre-stored outline is stored in one or more of memory devices 1-N. In this example, the feature extractor compares the size and shape of object 108O with pre-stored sizes and pre-stored shapes of pre-stored eyeglasses and determines that object 108O is eyeglasses 120. In this example, the pre-stored size, pre-stored shape, and pre-stored identification of the eyeglasses are stored in memory device 204. In this example, the pre-stored identification of the eyeglasses includes alphanumeric characters. In this example, the feature extractor provides the AI model with the identities of objects in the environmental system.
[0101] In one embodiment, the feature extractor identifies the location and graphical parameters, such as color, brightness, tone, and texture, of objects within the environmental system from image data received from cameras C1 and C2. For example, the feature extractor determines the positions of objects 108A-108P relative to each other and the graphical parameters of objects 108A-108P. The location and graphical parameters are stored in one or more of memory devices 1-N of server system 136.
[0102] In one embodiment, the processor is replaced by an application specific integrated circuit (ASIC) or a programmable logic device (PLD) or a central processing unit (CPU) or a combination of a CPU and a GPU.
[0103] In an embodiment, the game engine is replaced by the engine of another application, such as a video conferencing application.
[0104] In one embodiment, the classifier receives from the feature extractor the identities of objects 108A-108O in environment 102. The identities of objects 108A-108O are determined by the feature extractor from images captured by camera C1 and user account 1. The images captured by camera C1 are part of input dataset 1. Similarly, the classifier receives from the feature extractor the identities of objects 154A-154N, 108O, 154P-154Q, 154S, and 154T in environment 152. The identities of objects 154A-154N, 108O, 154P-154Q, 154S, and 154T are determined by the feature extractor from images captured by camera C1 and user account 1. The images captured by camera C1 are part of input dataset 2. The classifier also receives from the feature extractor the identities of objects 154A-154N, 108O, 154P-154Q, 154S, and 154T in the environment 152. The identities of objects 154A-154N, 108O, 154P-154Q, 154S, and 154T are determined by the feature extractor from images captured by camera C2 and user account 2. The images captured by camera C2 are part of input dataset 3.
[0105] 3 is a diagram of an embodiment of a system 300 to illustrate correlations 302, 304, and 306. System 300 includes audio data sets 1-3.
[0106] 4A is a diagram of an embodiment of an object state 400. For example, the state of a floor is whether the floor is made of carpet, tile, or concrete. As another example, the state of a window blind is whether the window blind is open or closed. As yet another example, the state of a container is whether the container is open or closed.
[0107] FIG. 4B is a diagram of an embodiment of object type 402. For example, the object type provides the identity of the object. For purposes of illustration, a vehicle is a type of object, a monitor is a type of object, a display device is a type of object, external speakers are a type of object, a human, such as a child or spectator or user, is a type of object, a house is a type of object, a building is a type of object, carpet is a type of object, a bare floor is a type of object, a wall is a type of object, and a window is a type of object. An example of each of objects 108K (FIG. 1A-1), 108L (FIG. 1A-1), 154E (FIG. 1B), and 154F (FIG. 1B) is an external speaker. For purposes of illustration, the external speaker is not integrated within another device, such as a display device, but is located outside the display device.
[0108] 5A is a diagram of an embodiment of an environment 500 to illustrate the use of an AI model to identify objects within the environment 500 and to determine the state of the objects within the environment 500 and the placement of the objects within the environment 500. The environment 500 is a room within a house 504. The environment 500 includes an example object, user 3. The environment 500 further includes object 502A, object 502B, object 502C, object 502D, object 502E, object 502F, object 502G, object 502H, object 502I, object 502J, object 502K, object 502L, object 502M, object 502N, object 502O, object 502P, and object 502Q. Examples of client device 3 (FIG. 2) include eyeglasses worn by user 3, object 502I, object 502E, object 502D, object 502A, object 502G, input controllers coupled to the eyeglasses, and combinations of two or more thereof.
[0109] Object 502A is a display device including a computer. Object 502B is also a display device such as a monitor. Object 502C is a table on which objects 502A, 502B, 502D, 502E, 502F, and 502G are placed. Object 502D is a mouse coupled to object 502A, and object 502E is a keyboard coupled to object 502A. Object 502F is a stapler, and object 502G is a speaker coupled to object 502A. Object 502H is a robotic arm. Object 502I is a handheld controller used by user 3 to play game G1. Object 502J is a chair on which user 3 is sitting. Object 502K is a shelf on which objects 502L, 502M, and 502N are supported. Object 502L is a box, and object 502M is another box. Object 502N is a stack of containers. Object 502O is a pair of glasses worn by user 3. As an example, object 502O has the same structure and functionality as glasses 120 (FIGS. 1A-2). For example, object 502O includes a camera C3 and a microphone M3. In this example, glasses 120, which are an example of object 502O, include a camera C3 and a microphone M3 instead of camera C1 and microphone M1. Object 502P is a window without blinds. Object 502Q is an open soda can.
[0110] User 3 accesses game G1 from the game cloud via computer network 142 (FIGS. 1A-2) and plays game G1, which renders a virtual scene on the display screen of object 502A. For example, user 3 selects one or more buttons on one or more of objects 502I, 502D, and 502E to provide authentication information, such as a username and password. Object 502I or objects 502D and 502E transmit the authentication information to object 502A, which then forwards the authentication information to the game cloud via computer network 142. The game cloud's authentication server determines whether the authentication information is authentic and, if so, provides access to user account 3 and a game program running on the game cloud. When the game program is executed by one or more of processors 1-N of the game cloud, image frames of game G1 are generated and encoded, and the encoded image frames are output. The encoded image frames are transmitted to object 502A. Object 502A decodes the encoded image frames and provides them to its display screen, which applies a rendering program to the image frames to display images of the virtual scene and further enables user 3 to play game G1.
[0111] During the play of game G3, the sound of game G3 is output from object 502G. For example, if the virtual object in the virtual scene displayed on object 502A is a car accelerating on the virtual ground in the virtual scene, the sound of the car running is output via object 502G. For example, the virtual object in the virtual scene displayed on object 502A is controlled by user 3 via object 502I or objects 502D and 502E. Also, during the play of game G1, object 502R, which is a car near house 504, makes sounds such as engine noise. Car 502R is part of external environment 506 located outside house 504.
[0112] Microphone M3 captures audio data, such as audio frames, generated from sounds associated with, e.g., sounds emanating from or reflected from, one or more of the objects located in environment 500. For example, microphone M3 captures audio data generated based on sounds emanating from object 502G and received from object 502G via path 508A. In this example, sound path 508A is a direct path from object 502G to microphone M3 and does not impinge on any other objects between object 502G and microphone M3. As another example, microphone M3 captures sounds emanating from object 502G and received from object 502G via path 508B. In this example, sound path 508B is an indirect path from object 502G to microphone M3. For purposes of illustration, sounds emanating from object 502G impinge on one or more other objects (e.g., objects 502A and 502C) in environment 500 and are reflected from those one or more other objects toward microphone M3.
[0113] Microphone M3 also captures audio data generated based on sounds, such as background noise, emanating from object 502R located outside environment 500. As an example, microphone M3 captures sounds emanating from object 502R and received through a wall or window of environment 500.
[0114] Microphone M3 captures sounds associated with one or more objects in environment 500 (e.g., objects 502A-502Q), such as sounds emanating from or reflected from those objects, and background noise emanating from object 502R in external environment 506, to generate audio frames, such as audio dataset N (N is a positive integer). Encoded audio frames are generated based on the audio frames and provided to one or more of processors 1-N of game cloud via computer network 142 for processing and AI model training. Note, by way of example, that there is no capture of images of environment 500 by camera C3.
[0115] The feature extractor extracts, e.g., determines, parameters such as one or more amplitudes, one or more frequencies, and one or more sensing directions from the audio dataset N in the same manner as the parameters are determined from the audio datasets 1, 2, or 3. For example, the feature extractor determines the magnitude, or the peak-to-peak amplitude, or the zero-to-peak amplitude, and the frequency of the audio dataset N. For illustration, the feature extractor determines the absolute maximum power of the audio dataset N or the absolute minimum power of the audio dataset N to determine the magnitude of the audio dataset N. As another explanation, the feature extractor determines a local maximum magnitude and a local minimum magnitude of the audio dataset N. As another explanation, the feature extractor determines multiple local maxima of magnitude and multiple local minima of magnitude, and the feature extractor applies a best fit or an average or median to the local maxima of magnitude and the local minima to determine the maximum magnitude and the minimum magnitude. Alternatively, the feature extractor determines a first time when the audio data set N reaches a predetermined magnitude and a second time when the audio data set N reaches the same predetermined magnitude, and calculates the difference between the first and second times to determine the time interval. The feature extractor inverts the time interval to determine the absolute frequency of the audio data set N. Alternatively, the feature extractor determines a local frequency of the audio data set N. In this description, the local frequency is a frequency within a predetermined period of time, which is shorter than the entire period over which the audio data set N is generated. In this description, multiple local frequencies are determined from the audio data set N, and the feature extractor applies a best fit, average, or median to the local frequencies to determine the frequency.
[0116] As another example, the feature extractor determines the direction from which audio data set N is perceived. In this example, microphone M3 includes an array (e.g., a linear array) of transducers arranged in a linear direction. The array includes a proximal transducer and a distal transducer. In this example, if the proximal transducer outputs a first portion of audio data N and the distal transducer outputs a second portion of audio data N, with the first portion having an amplitude greater than the second amplitude, the feature extractor determines that object 502G (FIG. 5A) in environment 500 (FIG. 5A) is closer to the proximal transducer than to the distal transducer. In this example, the feature extractor further determines that object 502G is in a direction pointing toward the proximal transducer.
[0117] The one or more amplitudes determined from audio data set N are referred to herein as amplitudes N. Also, the one or more frequencies determined from audio data set N are referred to herein as frequencies N, and the direction in which audio data set N is perceived is referred to herein as perceived direction N.
[0118] In one embodiment, user 3 plays a different game than game G1.
[0119] In an embodiment, the external environment 506 includes any other number of objects, such as two or three.
[0120] In an embodiment, object 502O excludes camera C3.
[0121] In one embodiment, parameters may be referred to herein as features.
[0122] In an embodiment, instead of or in addition to microphone M3, one or more additional microphones, such as standalone microphones, are present to capture sounds emanating from objects within environment 500 and from objects located in external environment 506. For example, a display device located on a table in environment 500 includes an additional microphone.
[0123] 5B is a diagram of an embodiment of model output 550. Model output 550 includes a probability A% that environment 500 includes object combination N, a probability B% that the object has configuration N, a probability C% that the object has state N, and a probability D% that environment 500 is of type N, where A, B, C, and D are positive real numbers. As an example, A, B, C, and D are equal. As another example, one of A, B, C, and D is not equal to at least one of the remaining A, B, C, and D.
[0124] When the AI model is provided with amplitude N, frequency N, and sensing direction N from the feature extractor, the AI model provides model output 500. For example, upon determining that amplitude N is within a predetermined range from amplitude 1 and outside a predetermined range from amplitude 2, the AI model indicates that there is a greater than 50% probability that audio data set N was received from a room in a house, as opposed to a room in a building. In this example, a house is an example of environment type N. For purposes of illustration, a house is an indoor-type environment. As another example, upon determining that frequency N is within a predetermined range from frequency 2 and outside a predetermined range from frequency 1, the AI model indicates that there is a greater than 50% probability that audio data set N was received from a room in a building, as opposed to a room in a house. In this example, a building is an example of environment type N. For purposes of illustration, a building is an indoor-type environment. As another example, a combination of two or more of amplitude N, frequency N, and sensing direction N is used to determine environment 500 of type N.
[0125] As yet another example, upon determining that sensing direction N is within a predetermined range from sensing direction 1 and outside a predetermined range from sensing direction 2, the AI model indicates that there is a greater than 50% probability that audio data set N is output from a speaker behind the display device compared to a speaker in front of the display device of environment 500. Examples of speaker and display placements N are provided depending on whether the speaker is behind or in front of the display device. As another example, a combination of two or more of amplitude N, frequency N, and sensing direction N is used to determine the placement N of an object in environment 500.
[0126] As another example, upon determining that the amplitude N is within a predetermined range from amplitude 2 and outside a predetermined range from amplitude 1, the AI model indicates that there is a greater than 50% probability that a window without blinds exists in environment 500. In this example, the window without blinds is an example of an object combination N or object in environment 500. As another example, a combination of two or more of the amplitude N, frequency N, and sensing direction N is used to determine the object combination N in environment 500.
[0127] As yet another example, upon determining that amplitude N is within a predetermined range from amplitude 1 and outside a predetermined range from amplitude 2, the AI model indicates that there is a greater than 50% probability that a predetermined number of speakers are present in environment 500. In this example, the predetermined number of speakers is an example combination N of objects in environment 500.
[0128] As yet another example, upon determining that amplitude N is within a predetermined range from amplitude 1 and outside a predetermined range from amplitude 2, the AI model indicates that there is a greater than 50% probability that a window in environment 500 has open blinds or that the window does not have blinds. In this example, open or closed blinds are examples of state N of an object in environment 500. As another example, a combination of two or more of amplitude N, frequency N, and sensing direction N is used to determine state N of an object in environment 500.
[0129] As yet another example, upon determining that frequency N is within a predetermined range from frequency 1 and outside a predetermined range from frequency 2, the AI model may indicate that there is a greater than 50% probability that a soda can in environment 500 is open. In this example, an open soda can in environment 500 is an example of state N of an object in environment 500.
[0130] In an embodiment, server system 136 ( FIG. 7 ) receives input via user account 3 and either object 502I or objects 502D and 502E. When user 3 makes one or more selections on object 502I or objects 502D and 502E or an input controller to indicate that user 3 wants to identify an object in environment 500, the input is generated by an input controller coupled to object 502I or objects 502D and 502E or object 502O. In this embodiment, user 3 is wearing an HMD. The input is transmitted from object 502I or objects 502D and 502E or the input controller to server system 136 via object 502O or object 502A and computer network 142. Upon receiving input requesting the identities of objects in environment 500, server system 136 applies AI models to object 502O or object 502A via computer network 142 to obtain probabilities A%, B%, C%, and D% for combination N, arrangement N, state N, and type N of environment 500, as well as probabilities for combination N, arrangement N, state N, and type N. Object 502O or object 502A displays these probabilities, as well as combination N, arrangement N, state N, and type N of objects in environment 500, on the display screen of object 502O or object 502A.
[0131] In one embodiment, camera C3 captures image data of environment 500 and transmits the image data to server system 136 ( FIG. 7 ) via computer network 142. The image data is used to verify the combination N, arrangement N, state N, and type N of environment 500 determined by the AI model. One or more processors 1-N of server system 136 verify the identities of objects in environment 500 determined by the AI model based on the image data. For example, processor N determines that a match exists between a first identity of object 502A, that is, a display device, and a second identity of object 502A, that is, a display device. In this example, the first identity is determined by the AI model, and the second identity is determined based on the image data. On the other hand, if processor N determines that a match does not occur, processor N reapplies the AI model to determine a third identity of object 502A, or decides to ignore the first identity and use the second identity.
[0132] FIG. 6A is a diagram of an embodiment of an acoustic profile database 600 stored in one or more of memory devices 1-N (FIGS. 1A-2). Acoustic profile database 600 includes audio data set 1, audio data set 2, and audio data set N. Acoustic profile database 600 also includes environmental systems 1, 2, and N. Acoustic profile database 600 further includes correspondences between audio data sets 1, 2, and N and environmental systems 1, 2, and N. For example, acoustic profile database 600 includes a first exclusive correlation between audio data set 1 and environmental system 1, a second exclusive correlation between audio data set 2 and environmental system 2, and an Nth exclusive correlation between audio data set N and environmental system N. An example of environmental system 1 is the first environmental system, an example of environmental system 2 is the second environmental system (FIG. 3), and an example of the Nth environmental system is type N of environment 500 or environment 500.
[0133] FIG. 6B is a diagram of an embodiment of a method 650 illustrating the application of audio data to output sounds corresponding to an environmental system. Method 650 is performed by one or more of processors 1-N (FIGS. 1A-2), a processor of game console 112 (FIG. 1A-1), or a combination thereof. Method 650 includes an operation 652 of receiving, from user 3, a selection of an environmental system to be simulated, e.g., presented. For example, user 3 selects environmental system N corresponding to the sounds user 3 wants to hear via object 502I or a combination of objects 502E and 502D. In this example, user 3 is located in an environment different from environment 500. For illustration purposes, object 502P has blinds, or object 502B (FIG. 5A) is removed from environment 500 to create a different environment. In this example, user 3 accesses another application program, such as a game program for game G1 or another application, from the game cloud. To illustrate, when a game program or other application is executed by one or more of processors 1-N, one or more image frames of one or more virtual scenes are generated and displayed on object 502A. Further, in this example, during or before the execution of the game program or other application program, user 3 selects environmental system N. In this example, an indication of the selection of environmental system N is transmitted from object 502A over computer network 142 to server system 136 (FIGS. 1A-2). One or more processors 1-N receive the indication of selection and access database 600 to determine that audio data set N is exclusively associated with environmental system N.
[0134] At operation 654 of method 650, one or more processors 1-N access, e.g., retrieve, audio data set N from database 600. At operation 656 of method 600, one or more processors 1-N apply audio data set N to output sounds corresponding to simulated environmental system N to user 3. For example, one or more processors transmit audio data set N to object 502A over computer network 142. While playing game G1 or running other application programs, object 502A outputs sounds generated based on audio data set N. The sounds are output through a speaker of object 502A or through object 502G. Outputting sounds simulating environmental system N causes user 3 to feel as if he or she is in environmental system N rather than another environment.
[0135] FIG. 7 is a diagram of an embodiment of a system 700 for illustrating method 600 (FIG. 6). System 700 includes an input controller system 702, a display system 704, a server system 136, and a computer network 142. Display system 704 includes a communication device 706, a network transfer device 708, a CPU 712, a sound output system 710, and an audio memory device 714. An example of display system 704 is object 502O (FIG. 5A). Another example of display system 704 is object 502A (FIG. 5A). An example of input controller system 702 is input controller 122 (FIGS. 1A-2). Another example of input controller system 702 is a handheld controller. Yet another example of input controller system 702 is a mouse and keyboard combination.
[0136] The communication device 706 has the same structure as the communication device 134 (FIG. 1A-2) and has similar functions as the communication device 134. The network transfer device 708 has the same structure as the network transfer device 126 (FIG. 1A-2) and has similar functions as the network transfer device 126. The CPU 712 has the same structure as the CPU 135 (FIG. 1A-2) and has similar functions as the CPU 135. An example of the audio memory device 714 is a buffer for storing audio data. The sound output system 710 includes a digital-to-analog converter (DAC), an amplifier, and a speaker.
[0137] The communication device 706 is coupled to the input controller system 702 and a CPU 712. The CPU 712 is coupled to a network transfer device 708, which is coupled to the server system 136 via the computer network 142. The CPU 712 is coupled to a DAC, which is coupled to an amplifier. The amplifier is coupled to a speaker. The CPU 712 is coupled to an audio memory device 714.
[0138] 6 , in operation 652, input controller system 702 generates an input signal, such as an indication, in response to user 3's selection of environmental system N. User 3 uses input controller system 702 to select one or more buttons on input controller system 702. A communication device of input controller 702 applies a communication protocol to the indication of environmental system N selection to generate one or more transport packets and transmits the transport packets to communication device 706. Communication device 706 applies the communication protocol to the data packets to extract the indication of environmental system N selection from the transport packets and transmits the indication to CPU 712.
[0139] The CPU 712 sends an indication of the selection of environmental system N to the network forwarding device 708. The network forwarding device 708 applies a network forwarding protocol to generate data packets including the indication of the selection of environmental system N and transmits the data packets over the computer network 142 to the server system 136. The network forwarding device 138 of the server system 136 applies the network forwarding protocol to the data packets to extract the indication of the selection of environmental system N from the data packets and provides the indication to one or more of the processors 1-N. The one or more processors 1-N perform operation 654 ( FIG. 6B ) to identify and access, e.g., read, audio data set N from one or more of the memory devices 1-N corresponding to the environmental system N and transmit the audio data set N to the network forwarding device 138. By way of example, in addition to performing operation 654, the one or more processors 1-N may access an audio file (including audio information) of a game program or other application program and provide the audio information to the network forwarding device 138. An example of audio information is audio data of a virtual character jumping when the virtual character in game G1 jumps. Another example of audio information is a user providing commentary during a YouTube® video about a sporting event, such as a basketball game, when the other application program is YouTube®. In this example, the sporting event is an example of a context.
[0140] The network transport device 138 applies a network transport protocol to generate data packets from the audio data set N and transmits the data packets to the display system 704 via the computer network 142 for application of the audio data set N in operation 656 (FIG. 6B). For example, the network transport device 138 embeds the audio data set N within data packets and transmits the data packets to the network transport device 708 via the computer network 142. In this example, the network transport device 708 applies a network transport protocol to the data packets to obtain the audio data set N and transmits the audio data set N to the CPU 712. The CPU 712 provides the audio data set N to a DAC, which converts the audio data set N from a digital format to an analog format and outputs an analog audio signal N. The DAC provides the analog audio signal N to an amplifier. The amplifier amplifies, e.g., increases or decreases, the amplitude of the analog audio signal N and outputs the amplified audio signal. The speaker converts the electrical energy of the analog audio signal N into sound energy to output the sound of the environmental system N based on the audio data set N, thereby simulating the environmental system N within the environment 500.
[0141] As another example, the network transport device 138 embeds audio data set N and audio information associated with a game program or other application program in data packets and transmits the data packets to the network transport device 708 over the computer network 142. In this example, the network transport device 708 applies a network transport protocol to the data packets to obtain the audio data set N and the audio information and transmits the audio data set N and the audio information to the CPU 712. The CPU 712 provides the audio data set N and the audio information to a DAC, which converts the audio data set N and the audio information from digital format to analog format and outputs an analog audio signal. The DAC provides the analog audio signal to an amplifier. The amplifier amplifies, e.g., increases or decreases, the amplitude of the analog audio signal and outputs the amplified audio signal. The speaker converts the electrical energy of the analog audio signal into sound energy and outputs a first set of sounds of the game G1 or other application and a second set of sounds of the environmental system N as a background to the first set of sounds. In this example, the first and second sets are blended together when being output simultaneously. In this example, the context of the audio information associated with the game program or other application program and the context of the audio data set N are matched. To illustrate, if a YouTube®™ video includes a commentary of a sporting event, such as a baseball game, then audio dataset N represents the sounds made during the sporting event, such as a baseball game. In this example, the baseball game is an example of a context.
[0142] In an embodiment, one or more of processors 1-N determine not to provide audio information corresponding to a virtual scene of game G1 or a video of another application program. For example, instead of accessing audio information of a virtual character jumping in a virtual scene that is output together with the virtual scene from one or more of memory devices 1-N, one or more of processors 1-N access an audio dataset N from one or more of memory devices 1-N and provide the audio dataset N to be applied together with the virtual scene. As another example, instead of accessing audio information that is output together with a YouTube®™ video from one or more of memory devices 1-N, one or more of processors 1-N access an audio dataset N from one or more of memory devices 1-N and provide the audio dataset N to be applied together with the YouTube®™ video. As yet another example, one or more of processors 1-N stop applying audio information of a virtual character jumping in a virtual scene and instead apply the audio dataset N. As yet another example, one or more of processors 1-N stop applying audio information that is output together with the YouTube®™ video and instead apply the audio dataset N.
[0143] In one embodiment, display system 704 includes additional components, such as a microphone, a display screen, a GPU, an audio encoder, a video encoder, and a display screen. These components have similar structures and functions as the corresponding components of glasses 120. For example, the microphone of display system 704 has the same structure as microphone M1, and if display system 704 is a display device, the display screen of display system 704 is larger than display screen 132. As another example, if display system 704 is glasses, such as glasses 120 (FIGS. 1A-2), the display screen of display system 704 has the same size as display screen 132.
[0144] 8 is a diagram of an embodiment of a system 800 for illustrating a microphone 800. The microphone 800 is an example of microphone M1, M2, or M3. The microphone 800 includes a transducer, a converter, and an analog-to-digital converter (ADC). The transducer is coupled to a sound energy-to-electrical energy converter (SE converter), which is coupled to the ADC. An example of a transducer is a diaphragm. An example of an SE converter is a capacitor or a series of capacitors.
[0145] The transducer detects sound emanating from or reflected from objects in the environment, or both, and outputs vibrations. The vibrations are provided to an SE converter, which modifies the electric field generated within the SE converter to output an electrical audio analog signal. The audio analog signal is provided to an ADC, which converts the audio analog signal from analog format to digital format and outputs audio data such as audio data set 1 or 2 or 3 or N.
[0146] 9 is a diagram of an embodiment of a system 900 illustrating a method for using direct audio data to create effects for an environmental system N and using reverberated audio data to determine the placement of objects within the environmental system N, as well as properties of the objects, such as the object's material type and the object's surface type. System 900 includes an audio data separator, direct audio data, reverberated audio data, a feature extractor, a classifier, and an AI model.
[0147] An example of direct audio data is audio data generated by a microphone based on sound received from a sound source via a direct path. An example of reverberant audio data is audio data generated by a microphone based on sound received from a sound source via an indirect path. For illustrative purposes, the first direct audio data of audio dataset 1 is generated by microphone M1 (FIG. 1A) based on sound received from object 108K (FIG. 1A-1) via path 106A. In this description, the first reverberant audio data of audio dataset 1 is generated by microphone M1 based on sound received from object 108K via path 106B (FIG. 1A-1). Alternatively, reverberant audio data is generated based on sound reflected or diffused from objects in an environmental system.
[0148] The audio data separator may be implemented as hardware or software, or a combination thereof. For example, the audio data separator may be a computer program, and the functions of the computer program may be executed by one or more of processors 1-N (FIGS. 1A-2). As another example, the audio data separator may be an ASIC or PLD.
[0149] The audio data separator is coupled to the feature extractor and audio decoder 144 of the server system 136 (FIG. 7). For example, the audio data separator is coupled between the feature extractor and the audio decoder 144.
[0150] The audio data separator receives audio data sets 1, 2, 3, and N from client devices 1 and 2 (FIG. 2) via computer network 142, network transport device 138, and audio decoder 144 (FIG. 7). The audio data separator determines parameters of audio data sets 1, 2, 3, and N and, based on these parameters, identifies direct audio data and reverberated audio data within each of audio data sets 1, 2, 3, and N. For example, the audio data separator identifies first direct audio data and first reverberated audio data within audio data set 1 and second direct audio data and second reverberated audio data within audio data set 2. For illustrative purposes, the audio data separator determines that a first portion of audio data set 1 has a first amplitude that is greater than a second amplitude of a second portion of audio data set 1. In this description, examples of amplitude include peak-to-peak amplitude and zero-to-peak amplitude. Furthermore, in this description, the audio data separator determines that the first portion is first direct audio data and the second portion is first reverberated audio data. Alternatively, the audio data separator determines that a first portion of audio data set 1 has a first frequency range and a second portion of audio data set 1 has a second frequency range. Further in this description, the audio data separator determines that the first portion is first direct audio data and the second portion is first reverberant audio data. The audio data separator determines parameters for audio data sets 1, 2, 3, and N in the same way that the feature extractor determines parameters for audio data sets 1, 2, 3, and N.
[0151] The direct audio data output from the audio data separator is stored by one or more of processors 1-N in one or more of memory devices 1-N (FIG. 7). Operations similar to operation 654 (FIG. 6) are performed based on the direct audio data. For example, upon receiving a selection of environmental system N to be simulated from user 3, one or more of processors 1-N accesses the direct audio data corresponding to environmental system N from one or more of memory devices 1-N. As another example, the same operations as operation 654 are performed, except that the operations are performed on the direct audio data of audio data set N instead of audio data set N.
[0152] Further, an operation similar to operation 656 (FIG. 6B) is performed to apply the direct audio data to output a sound corresponding to environmental system N. For example, the same operation as operation 656 is performed, except that the direct audio data of audio dataset N is applied instead of audio dataset N. As another example, the direct audio data of audio dataset N is transmitted from server system 136 (FIGS. 1A-2) to object 502A or 502O (FIG. 5A). In this example, object 502A or 502O outputs a sound generated based on the direct audio data of audio dataset N while playing game G1 or running another application program. As yet another example, the direct audio data of audio dataset N is combined with reverberant audio data of audio dataset N to output a sound based on audio dataset N. For illustrative purposes, one or more of processors 1-N combine the direct audio data of audio dataset N with the reverberant audio data of audio dataset N to output audio dataset N. In this illustration, audio dataset N is then applied in the manner described above in operation 656 to simulate environmental system N for user 3. As yet another example, direct audio data of audio dataset N is combined by one or more of processors 1-N with reverberant audio data of another audio dataset, such as audio dataset 1 or 2 or 3, to generate an additional audio dataset. In this example, the additional audio dataset is applied in the same manner as audio dataset N is applied in operation 656. For illustrative purposes, the additional audio dataset is transmitted from server system 136 over computer network 142 to object 502A or 502O, and sound is output based on the additional audio dataset.
[0153] The reverberant audio data output from the audio data separator is sent to the feature extractor, for example, the first reverberant audio data and the second reverberant audio data are sent from the audio data separator to the feature extractor.
[0154] The feature extractor determines parameters of the reverberated audio data of any of audio datasets 1, 2, 3, and N in the same manner as the feature extractor determines parameters of the audio datasets. For example, the feature extractor determines an amplitude 1a of the reverberated audio data of audio dataset 1, an amplitude 2a of the reverberated audio data of audio dataset 2, an amplitude 3a of the reverberated audio data of audio dataset 3, and an amplitude Na of the reverberated audio data of audio dataset N. As another example, the feature extractor determines a frequency 1a of the reverberated audio data of audio dataset 1, a frequency 2a of the reverberated audio data of audio dataset 2, a frequency 3a of the reverberated audio data of audio dataset 3, and a frequency Na of the reverberated audio data of audio dataset N. The feature extractor sends the parameters of the reverberated audio data of audio datasets 1, 2, and 3 to the classifier. The feature extractor also sends the parameters of the reverberated audio data of audio dataset N to the AI model.
[0155] The classifier classifies parameters of the reverberant audio data of the audio datasets 1-3 based on the input datasets 1-3. For example, the classifier determines or identifies a combination of objects in an environmental system, such as the environment 102 (FIG. 1A-1) or the external environment 116 (FIG. 1A-1), or a combination thereof, from the input dataset 1, and establishes a correlation (e.g., a one-to-one correspondence) between the combination of objects and the parameters of the reverberant audio data of the audio dataset 1. In this example, the classifier further determines or identifies the arrangement of the objects relative to each other, or the state of the objects, or the surface type of the objects, or the material type of the objects, or a combination of two or more thereof. In this example, the input dataset 1 includes the material type of the objects and the surface type of the objects. For illustrative purposes, the classifier receives the material type of the objects 108A-108O (FIG. 1A-1) and the surface type of the objects 108A-108O in the input dataset 1.
[0156] Examples of material types include wood, plastic, glass, marble, stainless steel, leather, wool, cloth, cotton, polyester, tile, or granite. For illustrative purposes, list 150 (FIGS. 1A-4) includes an entry indicating that the desktop table is made of wood and another entry indicating that the chair is made of leather. In this illustration, user 1 selects the type of material used to manufacture the desktop table and the type of material used to manufacture the chair in the same way that user 1 selects a chair and a desktop table in list 150. For illustrative purposes, list 150 includes an entry indicating that the desktop table has a textured or smooth surface and an entry indicating that the surface of the chair seat is intact or worn, such as torn or torn. In this illustration, user 1 selects the type of surface of the desktop table and the type of surface of the chair in the same way that user 1 selects a chair and a desktop table in list 150.
[0157] In another description of the classification, the classifier receives in input dataset 2 the identities of objects 154A-154N, 108O, 154P, and 154Q in environment 152, as well as the identity of object 154R via a list, such as list 150, and user account 1. In this description, the classifier receives in input dataset 2 the state, material type, and surface type of objects 154A-154N, 108O, 154P, and 154Q (FIG. 1B). Further, in this description, the classifier determines from a perceived direction of reverberant audio data in audio dataset 2 that object 154D is located closer to microphone M1 than object 154F (FIG. 1B).
[0158] In yet another illustration, the classifier receives, in input dataset 3, the identities of objects 154A-154N, 108O, 154P, and 154Q in environment 152, as well as the identity of object 154R via a list, such as list 150, and user account 2. In this illustration, the classifier receives, in input dataset 3, the state, material type, and surface type of objects 154A-154N, 108O, 154P, and 154Q (FIG. 1B). Further, in this illustration, the classifier determines from the sensing direction that object 154C is located closer to microphone M2 than object 154E (FIG. 1B). In this illustration, the sensing direction is determined from parameters of reverberant audio data in audio dataset 2.
[0159] The AI model is trained based on correlations between the reverberated audio data of audio datasets 1-3 associated with environments 102, 116, 152, and 158 (FIGS. 1A-1 and 1B) and parameters of input datasets 1-3. For example, the AI model is provided, via a classifier, with an indication of a first correlation, such as a link or one-to-one correspondence, between a set including amplitude 1a of the reverberated audio data of audio dataset 1, frequency 1a of the reverberated audio data of audio dataset 1, and perceived direction 1a of the reverberated audio data of audio dataset 1, and a set including a first type of environment, a first combination of objects within the first type of environment, a first arrangement of the objects, a first state of the objects, a first set of object material types, and a first set of object surface types. In this example, the one or more amplitudes determined from the reverberated audio data of audio dataset 1 are referred to herein as amplitude 1a. Also in this example, the one or more frequencies determined from the reverberated audio data of audio dataset 1 are referred to herein as frequency 1a, and the direction in which the reverberated audio data of audio dataset 1 is perceived is referred to herein as perceived direction 1a. For purposes of illustration, amplitude 1a, frequency 1a, and perceived direction 1a are example parameters of reverberant audio data in audio dataset 1. Also in this example, the classifier provides the AI model via user account 1 a first type of environment 102, a first combination of objects in environment 102 and external environment 116, a first state of the objects, a first arrangement of the objects, a first set of object material types, and a first set of object surface types.
[0160] As another example, the AI model is provided with an indication of a second correlation, such as a one-to-one correspondence, between a set including amplitude 2b, frequency 2b, and sensing direction 2a of the reverberated audio data of audio dataset 2 and a set including a second type of environment, a second combination of objects within the second type of environment, a second arrangement of objects, a second state of the objects, a second set of material types of objects, and a second set of surface types of objects. In this example, the AI model receives amplitude 2b, frequency 2b, and sensing direction 2b from the classifier. In this example, amplitude 2b, frequency 2b, and sensing direction 2b are determined by analyzing the reverberated audio data of audio dataset 2 captured by microphone M1. Also in this example, one or more amplitudes determined from the reverberated audio data of audio dataset 2 are referred to herein as amplitude 2b, one or more frequencies determined from the reverberated audio data of audio dataset 2 are referred to herein as frequency 2b, and a direction in which the reverberated audio data of audio dataset 2 is perceived is referred to herein as sensing direction 2b. For illustrative purposes, amplitude 2b, frequency 2b, and sensing direction 2b are examples of parameters of the reverberated audio data of audio dataset 2. Also, in this example, the classifier provides the AI model with a second type of environment 152, a second combination of objects within environment 152 and external environment 158, a second state of the objects, a second arrangement of the objects, a second set of material types of the objects, and a second set of surface types of the objects. In this example, the second type, second combination, second state, and second arrangement, the second set of material types of the objects, and the second set of surface types of the objects are received via user account 1.
[0161] As yet another example, the AI model is provided by the classifier with an indication of a third correlation (e.g., a link) between a set including amplitude 3a, frequency 3a, and perceived direction 3a of the reverberant audio data of the audio dataset 3 and a set including a third type of environment, a third combination of objects within the third type of environment, a third arrangement of objects, a third state of the objects, a third set of object material types, and a third set of object surface types. In this example, the AI model receives amplitude 3c, frequency 3c, and perceived direction 3c from the classifier. In this example, amplitude 3c, frequency 3c, and perceived direction 3c are determined by analyzing the reverberant audio data of the audio dataset 3 captured by microphone M2. Also in this example, one or more amplitudes determined from the reverberant audio data of the audio dataset 3 are referred to herein as amplitude 3c, one or more frequencies determined from the reverberant audio data of the audio dataset 3 are referred to herein as frequency 3c, and a direction in which the reverberant audio data of the audio dataset 3 is perceived is referred to herein as perceived direction 3c. For purposes of illustration, amplitude 3c, frequency 3c, and perceived direction 3c are examples of parameters of the reverberant audio data of the audio dataset 3. Also, in this example, the classifier provides to the AI model a third type of environment 152, a third combination of objects within environment 152 and external environment 158, a third state of the objects, a third arrangement of the objects, a third set of material types of the objects, and a third set of surface types of the objects. In this example, the third type, the third combination, the third state, the third arrangement, the third set of material types of the objects, and the third set of surface types of the objects are received via user account 2.
[0162] When the AI model is provided with the amplitude Na, frequency Na, and sensing direction Na from the feature extractor, the AI model provides a model output 902. For example, upon determining that the amplitude Na is within a predetermined range from amplitude 1a and outside a predetermined range from amplitude 2a, the AI model indicates that there is a greater than 50% probability that the reverberant audio data of audio dataset N is generated based on sound reflected from a plastic table or a table with a smooth top. In this example, the probability that the table is made of plastic or that the table has a smooth top is an example of model output 902. As another example, upon determining that the frequency Na is within a predetermined range from frequency 2a and outside a predetermined range from frequency 1a, the AI model indicates that there is a greater than 50% probability that the reverberant audio data of audio dataset N is generated based on sound reflected from a table with an uneven surface or a table with a marble top. In this example, the probability that the table is made of marble or that the table has an uneven surface is an example of model output 902. As another example, a combination of two or more of the amplitude Na, frequency Na, and sensing direction Na is used to determine the type of material of any of the objects in the environment 500 or the type of surface of any of the objects in the environment 500.
[0163] Note that the type of material an object is made of and the type of surface an object is made of are examples of properties of an object.
[0164] In one embodiment, during operation 654, a visual mapping of a scene of environment N is created on display screen 132 (FIG. 1A-2) based on image data captured by cameras C1 and C2 (FIG. 1B). For example, in addition to directly applying audio data to simulate environment system N, the arrangement of objects in environment system N and the graphical parameters of the objects in environment system N are also applied to simulate environment system N. For illustrative purposes, the arrangement of objects in environment system N and the graphical parameters of the objects in environment system N are accessed by one or more of processors 1-N from one or more of memory devices 1-N of server system 136 and transmitted to glasses 120 (FIG. 1A-2) via computer network 142. Upon receiving the arrangement of objects in environment system N and the graphical parameters of the objects in environment system N, GPU 130 (FIG. 1A-2) displays the graphical parameters of the objects according to the arrangement to simulate environment system N when user 3 is in a different environment.
[0165] It should be noted that in various embodiments, one or more features of some of the embodiments described herein may be combined with one or more features of one or more of the remaining embodiments described herein.
[0166] The embodiments described in this disclosure may be implemented with a variety of computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc. In one implementation, the embodiments described in this disclosure are practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a wired or wireless network.
[0167] With the foregoing embodiments in mind, it should be understood that, in one implementation, the embodiments described herein employ various computer-implemented operations involving data stored in computer systems. These operations are operations requiring physical manipulation of physical quantities. Any of the operations described herein that form part of the embodiments described herein are useful machine operations. Some embodiments described herein also relate to devices or apparatus for performing these operations. The apparatus may be specially constructed for the required purposes, or the apparatus may be a general-purpose computer selectively activated or configured by a computer program stored in the computer. In particular, in one embodiment, various general-purpose machines are used with computer programs written in accordance with the teachings herein. Alternatively, it may be more convenient to construct a more specialized apparatus to perform the required operations.
[0168] In an implementation, some embodiments described in this disclosure are embodied as computer-readable code on a computer-readable medium. The computer-readable medium is any data storage device that stores data which can then be read by a computer system. Examples of computer-readable media include hard drives, network-attached storage (NAS), ROM, RAM, compact disc ROM (CD-ROM), CD-recordable (CD-R), CD-rewritable (CD-RW), magnetic tape, optical data storage devices, non-optical data storage devices, etc. As an example, the computer-readable medium includes computer-readable tangible media distributed over network-coupled computer systems, such that the computer-readable code is stored and executed in a distributed fashion.
[0169] Additionally, although some of the above embodiments are described with respect to a gaming environment, in some embodiments other environments are used instead of a game, such as, for example, a video conferencing environment.
[0170] Although the method operations have been described in a particular order, it should be understood that other housekeeping operations may be performed between operations, or operations may be adjusted to occur at slightly different times, or may be distributed within a system that allows processing operations to occur at various intervals relative to processing, so long as the processing of the overlay operation is performed in the desired manner.
[0171] Although the foregoing embodiments set forth in this disclosure have been described in some detail for clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims. Thus, the present embodiments are to be considered as illustrative and not restrictive, and the present embodiments should not be limited to the details set forth herein but may be modified within the scope of the appended claims and their equivalents.
Claims
1. 1. A method for determining a real-world environment in which a first user is located, comprising: receiving a plurality of image data generated from a plurality of cameras in a plurality of real-world environments; extracting a plurality of features from the plurality of image data; receiving input data relating to the plurality of real-world environments; training an artificial intelligence (AI) model based on a plurality of image datasets generated from the plurality of real-world environments, a plurality of features from the plurality of image data, and input data related to the plurality of real-world environments; applying the AI model to image data captured from a real-world environment surrounding the first user to determine a type of the real-world environment; Including, each of the plurality of real-world environments having a different combination of objects; The step of training the AI model includes: providing the AI model with associations between feature values of the plurality of features and a plurality of types of the plurality of real-world environments; determining, by the AI model, a plurality of probabilities based on associations between feature values of the plurality of features and a plurality of types of the plurality of real-world environments; A method comprising:
2. 2. The method of claim 1, wherein extracting features from the plurality of image data comprises identifying a plurality of positions and a plurality of graphical parameters of objects within the plurality of real-world environments.
3. The method of claim 2 , wherein the graphics parameters include color, intensity, shading, or texture of objects in the multiple real-world environments.
4. The method of claim 2 , wherein the plurality of arrangements of objects includes relative positions of objects in the plurality of real-world environments with respect to other objects in the plurality of real-world environments.
5. providing the AI model with associations between the plurality of features, a plurality of types of the plurality of real-world environments, objects in the plurality of real-world environments, a plurality of arrangements of the objects, and a plurality of graphical parameters of objects in the plurality of real-world environments; determining, by the AI model, a plurality of probabilities based on associations between feature values of the plurality of features, a plurality of types of the plurality of real-world environments, objects in the plurality of real-world environments, a plurality of arrangements of the objects, and a plurality of graphical parameters of objects in the plurality of real-world environments; further comprising 3. The method of claim 2, wherein the plurality of probabilities provides a probability that image data captured from the real-world environment indicates a type of the real-world environment, a probability that the type of real-world environment includes a plurality of items, a probability that the plurality of items have a plurality of graphical parameters, and a probability that the plurality of items have a configuration.
6. receiving an indication of a type of real-world environment to be simulated; accessing image data captured from the real-world environment based on the type of real-world environment being simulated; providing image data captured from the real-world environment to a client device to output a type of the real-world environment; The method of claim 1 further comprising:
7. The method of claim 1 , wherein the input data includes data identifying objects in the plurality of real-world environments.
8. 10. The method of claim 1, wherein the plurality of image data is captured at a location where a plurality of users are located, including a second user and a third user.
9. The method of claim 1 , wherein the plurality of image data comprises a plurality of image frames of the real-world environment.
10. a server for determining an environment in which a first user is located, receiving a plurality of image data generated from a plurality of cameras in a plurality of real-world environments; extracting a plurality of features from the plurality of image data; receiving input data relating to the plurality of real-world environments; training an artificial intelligence (AI) model based on a plurality of image datasets generated from the plurality of real-world environments, a plurality of features from the plurality of image data, and input data related to the plurality of real-world environments; applying the AI model to image data captured from a real-world environment surrounding the first user to determine a type of the real-world environment; a processor that executes a memory device coupled to the processor; Equipped with each of the plurality of real-world environments having a different combination of objects; The step of training the AI model includes: providing the AI model with associations between feature values of the plurality of features and a plurality of types of the plurality of real-world environments; determining, by the AI model, a plurality of probabilities based on associations between feature values of the plurality of features and a plurality of types of the plurality of real-world environments; A server comprising:
11. The server of claim 10, wherein extracting features from the image data includes identifying positions and graphical parameters of objects within the real-world environments.
12. The server of claim 11, wherein the graphics parameters include color, intensity, shading, or texture of objects in the multiple real-world environments.
13. The server of claim 11 , wherein the plurality of placements of objects includes relative positions of objects in the plurality of real-world environments relative to other objects in the plurality of real-world environments.
14. The processor: providing the AI model with associations between the plurality of features, a plurality of types of the plurality of real-world environments, objects in the plurality of real-world environments, a plurality of arrangements of the objects, and a plurality of graphical parameters of objects in the plurality of real-world environments; determining, by the AI model, a plurality of probabilities based on associations between feature values of the plurality of features, a plurality of types of the plurality of real-world environments, objects in the plurality of real-world environments, a plurality of arrangements of the objects, and a plurality of graphical parameters of objects in the plurality of real-world environments; Further execute 11. The server of claim 10, wherein the plurality of probabilities provides a probability that image data captured from the real-world environment indicates a type of the real-world environment, a probability that the type of real-world environment includes a plurality of items, a probability that the plurality of items have a plurality of graphical parameters, and a probability that the plurality of items have a configuration.
15. The processor: receiving an indication of a type of real-world environment to be simulated; accessing image data captured from the real-world environment based on the type of real-world environment being simulated; providing image data captured from the real-world environment to a client device to output a type of the real-world environment; 11. The server of claim 10, further comprising:
16. The server of claim 10 , wherein the input data includes data identifying objects in the plurality of real-world environments.
17. The server of claim 10 , wherein the plurality of image data are captured at a location where a plurality of users, including a second user and a third user, are located.
18. The server of claim 10 , wherein the plurality of image data includes a plurality of image frames of the real-world environment.
19. 1. A system for determining an environment in which a first user is located, comprising: a plurality of client devices; a server coupled to the plurality of client devices; the plurality of client devices, generating from multiple cameras in multiple real-world environments; receiving input data relating to the plurality of real-world environments; Run The server receiving a plurality of image data generated from a plurality of cameras in the plurality of real-world environments; extracting a plurality of features from the plurality of image data; receiving input data relating to the plurality of real-world environments; receiving input data relating to the plurality of real-world environments; training an artificial intelligence (AI) model based on a plurality of image datasets generated from the plurality of real-world environments, a plurality of features from the plurality of image data, and input data related to the plurality of real-world environments; applying the AI model to image data captured from a real-world environment surrounding the first user to determine a type of the real-world environment; providing the AI model with associations between feature values of the plurality of features and a plurality of types of the plurality of real-world environments; determining, by the AI model, a plurality of probabilities based on associations between feature values of the plurality of features and a plurality of types of the plurality of real-world environments; Run A system, wherein each of the plurality of real-world environments has a different combination of objects.
20. The server receiving an indication of a type of real-world environment to be simulated; accessing image data captured from the real-world environment based on the type of real-world environment being simulated; providing image data captured from the real-world environment to the plurality of client devices to output a type of the real-world environment; 20. The system of claim 19, further comprising:
Citation Information
Patent Citations
A codec for processing scenes with almost unlimited detail
JP2021523442A
Room acoustics simulation using deep learning image analysis
WO2020139588A1
A visual object instance descriptor for place recognition
WO2021086422A1