A system and method for training a model to determine the type of environment surrounding a user
The system uses acoustic profiling and machine learning to enable users to identify and interact with their environment in multiplayer games, addressing the issue of visual limitations in head-mounted displays.
Patent Information
- Application Number
- JP2024533065
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-03
- Filing Date
- 2022-11-18
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-11-18
AI Technical Summary
Players in multiplayer games using head-mounted displays cannot see their environment, limiting their interaction and awareness of surroundings.
A system and method using microphones to detect acoustic characteristics of the environment, employing machine learning to identify objects and environments through sound profiling, allowing users to determine their surroundings without visual input.
Enables users to identify and interact with objects in their environment by analyzing acoustic reflections, providing environmental awareness and safety in virtual reality scenarios.
Smart Images

Figure 0007717284000001 
Figure 0007717284000002 
Figure 0007717284000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a system and method for training a model to determine the type of environment surrounding a user.
Background Art
[0002] In a multiplayer game, there are multiple game players. Each player wears a head-mounted display (HMD) to play a game or view the environment of an application. During the play of the game or the execution of the application, each player may not be able to see the environment in front of him or her.
[0003] In this context, embodiments of the present invention occur.
Summary of the Invention
[0004] Embodiments of the present disclosure provide a system and method for training a model to determine the type of environment surrounding a user.
[0005] In an embodiment, one or more microphones are used to detect the acoustic characteristics of the current environment. Examples of the current environment are locations such as a room inside a house or building where the user is located or an outdoor environment. Another example of the current environment is a location such as a room inside a building where the user is playing a video game and making selections to generate an input to the video game.
[0006] In one embodiment, in response to a user's request to generate an acoustic profile of the environment, an acoustic profile is generated. The acoustic profile is generated by capturing the sound in the environment, such as a space or a room, including the reflections and reverberations of sound from various objects located within the environment. The configuration of the environment around the user, or around the microphone or microphone array, is utilized to define the current environment.
[0007] In an embodiment, the environment includes sounds generated as background noise. The background noise is generated by, for example, conversations of other people, music playback, unrelated sounds from outside, noises emitted by mechanical objects within or around the environment, etc. By profiling the soundscape of the environment, it is possible to identify specific identification characteristics using a machine learning system. For example, training can be performed by running an environmental profiling application in many different environments. Over time, through a learning process using the machine learning system, objects within various environments, or objects around various environments that generate noise or sounds, are identified. In profiling, the machine learning system also identifies the acoustics of the various environments where the sound is being monitored. When training using the machine learning system is processed to a sufficient level, a machine learning model trained based on the sound profiles of various environments is used to automatically identify the things, objects, or sounds present in the current environment. These objects are present in the current environment, generate sounds that are reflected, those sounds are detected, and the acoustics of the current environment are generated. For example, when a user is playing a game in front of one or more monitors, there are reflections of the sound flowing from the monitors during gameplay. These types of reflections are processed using the machine learning system to identify unique characteristics and determine that the user is playing a game in front of another monitor. This type of characteristic identification can be used to detect other objects within the current environment where the user is located.
[0008] In an embodiment, the microphone array can be attached on a user peripheral device such as glasses, augmented reality (AR) glasses, a head-mounted display (HMD) device, or a handheld controller.
[0009] In one embodiment, during training, when the microphone array is located on the glass or HMD, the user is required to turn their head in various environments to capture various sounds. In other embodiments, the configuration of different environments is performed passively, and when the user transitions from one different environment to another, the audio signals and acoustic properties of objects in different environments are tracked over time.
[0010] In embodiments, profiling of various environments with respect to acoustic characteristics provides the type of acoustic vision of the current environment. For example, when the microphone array is located on an HMD or AR glasses or glasses, as the user moves and looks around in various environments, the machine learning system can almost instantaneously identify what is in front of the user based on the acoustic reflections and bounces back of the signals from the current environment.
[0011] In embodiments, as a person moves within the current environment, the acoustic signals in front of the user change, and based on the profile of the acoustic signals, it is possible to identify or confirm what may be present in front of the user using acoustic vision. In one embodiment, the acoustic vision of various environments is blended with data received from a camera to identify or verify the presence of an object based on the acoustic profile of the object present in front of the user within the current environment.
[0012] In another embodiment, it is possible to create a virtual profile of a space. For example, if a user wants to appear to be located at a specific location, such as a concert, park, game event, studio, etc., a machine learning system can generate sound using known acoustic profiles. The sound generated based on the acoustic profile can be blended with the sound generated by an application program, so that the user appears to be present at a specific location rather than the user's actual location. For example, when a user is publishing a YouTube (registered trademark) (trademark) video, the generated soundscape can be customized for the user based on the type of current environment that third parties watching the YouTube (registered trademark) (trademark) video want to project onto or virtually project onto. For example, if a user wants to provide commentary on a sports event, a soundscape can be virtually generated behind the commentary considering the acoustic profiles that exist or may exist at the sports event, and the sports event can be mimicked.
[0013] In an embodiment, a method for mapping the acoustic properties of the materials, surfaces, and geometric shapes of a space is described. The mapping is performed by extracting reverberation information such as the reflection and diffusion of sound detected by one or more microphones. When audio data is captured based on sound by one or more microphones, it is separated into a direct component and a reverberant component. Next, the direct component is resynthesized with different reverberation characteristics to effectively replace or change the listener's acoustic environment with some other acoustic profile. The other acoustic profile is used in combination with visual or geometric mapping of the space by, for example, a camera, SLAM (Simultaneous Localization And Mapping), or LiDAR (Light Detection and Ranging) to construct a more complete audiovisual mapping. Also, the reverberant component is used to convey features regarding the geometric shape of the space as well as the properties of the materials and surfaces.
[0014] In one embodiment, a method for determining the environment in which a user is located is described. The method includes receiving a plurality of audio datasets based on sounds emitted in a plurality of environments. Each of the plurality of environments has a different combination of objects. Further, the method includes receiving input data regarding the plurality of environments and training an artificial intelligence (AI) model based on the plurality of audio datasets and the input data. The method includes applying the AI model to audio data captured from the environment surrounding a first user to determine the type of the environment.
[0015] Some advantages of the systems and methods described herein include assisting a user or robot in learning objects in an environment without the need to obtain an image of the objects. For example, a robot can learn the identity of objects in an environment, the placement of the objects, and the state of the objects without obtaining an image of the objects. Once the robot has learned about the objects, it can be programmed to move around the objects and then be deployed into the environment and used within the environment.
[0016] Further advantages of the systems and methods described herein include providing a layout of the environment to a visually impaired person. Before a visually impaired person visits an environment, a layout of the environment is determined based on the sounds emitted by the objects within that environment. By doing so, the visually impaired person can also be made aware of the layout.
[0017] Further advantages of the systems and methods described herein include providing the identity and placement of objects in the environment in front of a user when the user is wearing an HMD. When a user is wearing an HMD, in a virtual reality (VR) mode for example, the user may not be able to see the environment in front of him or her. The systems and methods described herein facilitate providing the user with the identity and placement of the objects to prevent an accident between the user and the objects.
[0018] Other aspects of the present disclosure will become apparent from the following modes for carrying out the invention taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the embodiments described in the present disclosure.
[0019] The various embodiments of the present disclosure can be best understood by reference to the following description in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0020]
Figure 1A-1
Figure 1A-2
Figure 1A-3
Figure 1A-4
Figure 1B
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 6A
Figure 6B
Figure 7
Figure 8
Figure 9
Embodiments for Carrying Out the Invention
[0021] A system and method for training a model for determining the type of environment surrounding a user are described. Note that various embodiments of the present disclosure may be implemented without some or all of these specific details. In other cases, well-known process operations are not described in detail so as not to unnecessarily obscure various embodiments of the present disclosure.
[0022] FIG. 1A-1 is a diagram of an embodiment of an environment 102 in which user 1 is playing game G1. An example of environment 102 is a room within house 110. Environment 102 includes a plurality of objects 108A, 108B, 108C, 108D, 108E, 108F, 108G, 108H, 108I, 108J, 108K, 108L, 108M, 108N, and 108O. Object 108A is a chair, object 108B is a side table, object 108C is a display device such as a desktop monitor or a television monitor, and object 108D is a window in the wall of the room of house 110. Object 108D is a window with the blinds closed. Object 108E is a desktop table on which the display device is placed, object 108F is a keyboard coupled to the display device, and object 108G is a mouse coupled to the display device. For example, the keyboard is coupled to the central processing unit (CPU) of the display device, and the mouse is also coupled to the CPU. As an example, object 108E is made of plastic and has a smooth top surface.
[0023] Object 108H is a carpet covering the floor of the room, and object 108I is the floor of the room. Object 108J is the wall in which the window is located. Object 108K is a speaker coupled to the display device, and object 108L is another speaker coupled to the display device. For example, the speakers are coupled to the CPU of the display device. Note that object 108K is in front of object 108C. Also, object 108M is a soda can or container, and object 108N is a handheld controller coupled to game console 112. The soda can is open in environment 102. Game console 112 is coupled to the display device. For example, game console 112 is coupled to the CPU or the graphical processing unit (GPU) of the display device. Object 108O is glasses worn by user 1. Examples of glasses described herein include head-mounted displays (HMDs), prescription glasses, and augmented reality (AR) glasses.
[0024] The object 108P, which is a vehicle, is located outside the house 110, and it should be noted that it passes by the house 110 while the user 1 is playing the game G1. Further, it should also be noted that while the user 1 is playing the game G1, the child 1 and the child 2 are playing in another room inside the house 110 and talking to each other. Each of the child 1 and the child 2 is an example of an object.
[0025] The glasses include the camera C1 and the microphone M1 that converts sound into electrical energy. Examples of the camera C1 include a depth camera, a video camera, and a digital camera.
[0026] The user 1 accesses the game G1 from a game cloud such as a server system via a computer network and plays the game G1 that causes a virtual scene 114 to be represented on a display device. For example, the user 1 selects one or more buttons on a handheld controller and provides authentication information such as a username and a password. The handheld controller transmits the authentication information to the game console 112, and this game console transfers the authentication information to the game cloud via a computer network. The authentication server of the game cloud determines whether the authentication information is genuine, and if it determines that it is, it provides access to the user account 1 and the game program executed on the game cloud. When the game program is executed by one or more processors of the game cloud, an image frame of the game G1 is generated and encoded, and the encoded image frame is output. The encoded image frame is transmitted to the game console 112. The game console 112 decodes the encoded image frames and provides those image frames to the display device to display the virtual scene 114 of the game G1 and further enables the user 1 to play the game G1.
[0027] During the play of game G1, the sound of game G1 is output from a speaker coupled to the display device. For example, when virtual object 116A within virtual scene 114 shoots another virtual object 116B, the sound of the shooting is output. As an example, virtual object 116A is controlled by user 1 via a handheld controller. Also, during the play of game G1, the vehicle makes a honking sound, or makes sounds such as engine noise or tire screeching sound. Also, during the play of game G1, children 1 and 2 make sounds by talking to each other, arguing with each other, or playing with each other. Also, during the play of game G1, when user 1 opens a can, the sound of opening the can is made. Moreover, during the play of game G1, user 1 utters words which are examples of sounds.
[0028] The microphone M1 captures audio data such as an audio frame generated from sound related to one or more of the objects 108A to 108O located in the environment 102, such as sound emitted or reflected from the objects. For example, the microphone M1 captures audio data generated based on sound emitted from the object 108K and received from the object 108K via the path 106A. In this example, the sound path 106A is a direct path from the speaker to the microphone M1 and does not impinge on any other surface between the speaker and the microphone M1. As another example, the microphone M1 captures audio data generated based on sound emitted from the object 108K and received from the object 108K via the path 106B. In this example, the sound path 106B is an indirect path from the object 108K to the microphone M1. For the sake of explanation, the sound emitted from the object 108K impinges on one or more other objects such as a display device in the environment 102 and is reflected from those one or more other objects towards the microphone M1. As yet another example, the microphone M1 captures audio data generated based on sound emitted from the object 108K and received from the object 108K via the path 106D. In this example, the sound path 106D is an indirect path from the speaker to the microphone M1. For the sake of explanation, the sound emitted from the object 108K impinges on one or more other objects such as a carpet in the environment 102 and is reflected from those one or more other objects towards the microphone M1.
[0029] As another example, the microphone M1 captures audio data generated based on sound emitted from the object 108L and received from the object 108L via the path 106D. In this example, the sound path 106D is an indirect path from the speaker to the microphone M1. For the sake of explanation, the sound emitted from the object 108L impinges on one or more other objects such as a can within the environment 102 and is reflected from those one or more other objects towards the microphone M1. As yet another example, the microphone M1 captures audio data generated based on sound emitted from the object 108L and received from the object 108L via the path 106E. In this example, the sound path 106E is a direct path from the speaker to the microphone M1 and does not impinge on any other surface between the speaker and the microphone M1. As yet another example, the microphone M1 captures audio data generated based on sound emitted from the object 108L and received from the object 108L via the path 106F. In this example, the sound path 106F is an indirect path from the speaker to the microphone M1. For the sake of explanation, the sound emitted from the object 108L impinges on one or more other objects such as a desktop table within the environment 102 and is reflected from those one or more other objects towards the microphone M1.
[0030] The microphone M1 of the glasses also captures audio data generated based on sounds such as background noise emitted from one or more of the object 108P, child 1, and child 2, which are located near but outside the environment 102. The environment outside the environment 102 may be referred to as the external environment 116 in this specification. As an example, the microphone M1 captures audio data generated based on the sound emitted from the object 108P and received via the path 106G from the object 108P. In this example, the path 106G extends through the wall or door or entrance of the environment 102. In this example, when the door of the environment 102 is open, the sound spreads through the entrance, and when the door is closed, the sound spreads through the door. As another example, the microphone M1 captures audio data generated based on the sound emitted by child 1 or child 2 or both child 1 and child 2 and received via the path 106H from child 1 or child 2 or both child 1 and child 2. In this example, the path 106H extends through the wall or door or entrance of the environment 102.
[0031] As an example, it should be noted that the external environment 116 is close to the environment 102 when the sound emitted from an object in the external environment 116, such as a room adjacent to the environment 102 or a street outside the house 110, reaches the microphone M1 and can be detected by the microphone M1. For example, the sound emitted from the external environment 116 passes through the wall of the environment 102 and is detected by the microphone M1.
[0032] Microphone M1 captures audio data such as an audio frame generated from sounds related to those objects, such as sounds emitted from or reflected by objects 108A to 108O in environment 102, and background noise emitted from one or more of objects 108P, child 1, and child 2 in external environment 116. An audio frame encoded based on the audio frame captured by microphone M1 is generated and provided to one or more processors of the game cloud via a computer network for processing and training of an artificial intelligence (AI) model.
[0033] In an embodiment, the sound used in this specification includes sound waves.
[0034] In one embodiment, the terms object and item are used interchangeably in this specification.
[0035] In one embodiment, the display device includes a memory device. The CPU of the display device is coupled to the memory device.
[0036] In an embodiment, the sound emitted by the speaker is reflected from user 1 and captured by microphone M1. In this embodiment, user 1 is an example of an object.
[0037] In an embodiment, virtual scene 114 is not displayed on the display device, but on the display screen of the glasses worn by user 1.
[0038] In an embodiment, user 1 accesses game G1 from the game cloud without the need to use game console 112. For example, the encoded image frames of game G1 are transmitted from the game cloud via a computer network to glasses or a display device placed on a desktop table without being sent to game console 112 for video decoding. The encoded image frames are decoded by the glasses or the display device. In this embodiment, the encoded audio frames generated based on the audio frames output from microphone M1 are transmitted from the glasses via a computer network to the game cloud without using game console 112.
[0039] In one embodiment, virtual scene 114 includes other virtual objects, and based on the movement of these other virtual objects, sound is output from a speaker placed on the desktop table.
[0040] In an embodiment, instead of or in addition to microphone M1, one or more additional microphones, such as a stand-alone microphone, are present to capture the sound emitted from objects within environment 102 and the sound emitted from objects located within external environment 116. For example, a display device located on a desktop table includes an additional microphone. As another example, environment 102 includes one or more stand-alone microphones. As yet another example, a handheld controller includes an additional microphone.
[0041] In one embodiment, user 1 is not playing game G1. In this embodiment, instead of the game program, one or more processors of the game cloud execute another application program, such as a video conferencing application program or a multimedia program. For example, when the video conferencing application is executed, video image frames captured from an additional environment are transferred via a computer network to a display device or a game console or glasses worn by user 1.
[0042] In one embodiment, instead of environment 102, an outdoor environment such as a concert, a lake, or a park is used.
[0043] In one embodiment, instead of or in addition to the sound output from object 108K or 108L, sound is output from a speaker integrated within object 108C and detected by microphone M1, and audio data is captured.
[0044] FIG. 1A-2 is a diagram of an embodiment of system 118 that includes input controller 122, glasses 120, and server system 136. Server system 136 is an example of a game cloud. Glasses 120 are an example of object 108O (FIG. 1A-1). The handheld controller is an example of input controller 112. Glasses 120 include camera C1, video encoder 124, audio encoder 125, network transfer device 126, video decoder 128, GPU 130, display screen 132, and CPU 135. An example of video encoder 124 is a circuit that applies a video conversion protocol, such as a video encoding protocol, to an image frame to output encoded data, such as an encoded image frame, and provides the encoded data to network transfer device 126. For illustration purposes, video encoder 124 generates an example of an encoded image frame, an I-frame, a P-frame, or a B-frame. Examples of video encoding protocols, such as video compression protocols, include H.262, H.263, and H.264. An example of audio encoder 125 is a circuit that compresses an audio frame into an encoded audio frame. For illustration purposes, the audio encoder applies an audio encoding protocol, such as reversible compression or irreversible compression, to encode the audio frame into an encoded audio frame. Examples of irreversible compression include a modified discrete cosine transform (MDCT) for converting a time-domain sampled waveform into a frequency domain. Another example of irreversible compression is linear prediction coding (LPC), which analyzes audio data generated based on spoken sounds.
[0045] Examples of network transfer devices used in this specification are network interface controllers such as network interface cards (NICs). Another example of a network transfer device is a wireless access card (WAC). An example of a video decoder used in this specification is a circuit that performs H.262, H.263, or H.264 decoding, or another type of video extension, to output decoded data such as image frames. An example of the display screen 132 is a liquid crystal display (LCD) screen or a light emitting diode (LED) display screen. An example of a communication device of the device is a circuit that applies a communication protocol such as a wired communication protocol or a wireless communication protocol to communicate with another device or system. Examples of the CPU 135 include a processor, a microprocessor, a microcontroller, an application specific integrated circuit (ASIC), and a programmable logic device (PLD). The camera C1 includes a lens L1 that faces an environment such as the environment 102 (FIG. 1A-1). For example, the camera C1 is an outward-facing camera.
[0046] The CPU 135 is coupled to other components of the glasses 120. For example, the CPU 135 is coupled to the communication device 134, the camera C1, the network transfer device 126, the video encoder 124, the GPU 130, the video decoder 128, the audio encoder 125, and the microphone M1 to control the other components of the glasses 120. Also, the camera C1 is coupled to the video encoder 124, and this video encoder is coupled to the network transfer device 126. The network transfer device 126 is coupled to the computer network 142, the audio encoder 125, and the video decoder 128. The audio encoder 125 is coupled to the microphone M1. The video decoder 128 is coupled to the CPU 130, and this CPU is coupled to the display screen 132. Examples of the computer network 142 include a local area network (LAN), a wide area network (WAN), and combinations thereof. For the sake of explanation, the computer network 142 is the Internet or an intranet or a combination thereof.
[0047] The server system 136 includes a network transfer device 138, a video decoder 140, an audio decoder 144, and one or more servers 1 to N (N is an integer greater than 0). An example of an audio decoder is a circuit that extends an encoded audio frame into an audio frame. For the sake of explanation, the audio decoder applies an audio decoding protocol (such as an audio extension protocol) to decode the encoded audio frame into an audio frame. Each of the servers 1 to N includes a processor and a memory device. For example, server 1 includes processor 1 and memory device 1, server 2 includes processor 2 and memory device 2, and server N includes processor N and memory device N. The network transfer device 126 is coupled to the computer network 142 and also to the video decoder 140, which is coupled to one or more of the servers 1 to N. The network transfer device 126 is coupled to the audio decoder 144, which is coupled to one or more of the servers 1 to N. The operation of the system 118 is described with reference to FIGS. 1A-3.
[0048] In an embodiment, when the glasses 120 are AR glasses, the input controller 122 is different from the handheld controller used to play the game G1.
[0049] In one embodiment, the glasses 120 include a plurality of display screens instead of the display screen 132. Each display screen has a structure and function similar to that of the display screen 132.
[0050] In an embodiment, the glasses 120 include one or more additional lenses in addition to the lens L1 to capture an image of an environment such as the environment 102.
[0051] In one embodiment, the glasses 120 include one or more memory devices, such as a random access memory (RAM) or a read-only memory (ROM). The one or more memory devices are coupled to the CPU 135, the GPU 130, or both the CPU 135 and the GPU 130. For example, the CPU 135 includes a memory controller for accessing data from and writing data to the one or more memory devices.
[0052] In an embodiment, the memory controller is a device separate from the CPU 135.
[0053] Figures 1A-3 are diagrams of an embodiment that describes a training session in which user 1 is required to turn his or her head to capture a view of an environment such as environment 102 (Figure 1A-1). When the authentication information received from user 1 is authenticated, the authentication server provides user 1 with access to user account 1. Thereafter, user 1 can access a training session or a game program or other application program. During the training session, before or during the execution of a game program or other application program, a training program is executed by one or more processors 1 to N of the game cloud, and message 121 is generated. For example, before or during the execution of a game program, server system 136 (Figure 1A-2) executes a training program to generate an image frame having message 121, encodes the image frame to output an encoded image frame having message 121, applies a network transfer protocol to the encoded image frame to generate a data packet having message 121, and transmits the data packet to glasses 120 (Figure 1A-2) via computer network 142. In this example, network transfer device 126 (Figure 1A-2) of glasses 120 applies a network transfer protocol to obtain an encoded image frame from the data packet and provides the encoded image frame to video decoder 128 (Figure 1A-2) of glasses 120. In this example, video decoder 128 applies a video conversion protocol such as a video decoding protocol to decode the encoded image frame and outputs an image frame having message 121, and provides those image frames to GPU 130 (Figure 1A-2). In this example, GPU 130 applies a rendering program to display message 121 on display screen 132 (Figure 1A-2). In this example, the training program is computer software stored in one or more memory devices 1 to N of server system 136. Examples of network transfer protocols include Transmission Control Protocol (TCP) via Internet Protocol (IP).
[0054] When the user 1 visually recognizes the message 121, the user 1 turns his / her head to capture a view of the environment 102 (such as a 360-degree view). For example, when the user 1 turns his / her head within the environment 102 to visually recognize the environment 102, the camera C1 of the glasses 120 captures images of the objects 108A to 108O within the environment 102. Referring to FIG. 1A-1, the camera C1 transmits the images of the objects 108A to 108OP to the video encoder 124 of the glasses 120. The video encoder 124 applies a video encoding protocol to encode the images received from the camera C1 and outputs an encoded image frame, and transmits the encoded image frame to the network transfer device 126. The network transfer device 126 applies a network transfer protocol to generate data packets from the encoded image frame, and transmits those data packets to the network transfer device 138 of the server system 136 via the computer network 142.
[0055] The network transfer device 138 of the server system 136 acquires data packets from the glasses 120 via the computer network 142, applies a network transfer protocol to the data packets to extract the encoded image frames of the objects 108A to 108O. The network transfer device 138 transmits the encoded image frames to the video decoder 140 of the server system 136. The video decoder 140 applies a video decoding protocol to the encoded image frames to output image frames, and provides the image frames to one or more processors 1 to N of the game cloud.
[0056] Also, the microphone M1 generates an audio frame based on sounds emitted from objects such as the user 1 and the speaker in the environment 102 (FIG. 1A-1), and background noise emitted from the object 108P and children 1 and 2 in the external environment 116 (FIG. 1A-1). The audio frame is transmitted from the microphone M1 to the audio encoder 125 of the glasses 120. The audio encoder 125 applies an audio encoding protocol and outputs an encoded audio frame. The encoded audio frame is provided from the audio encoder 125 to the network transfer device 126. The network transfer device 126 applies a network transfer protocol to the encoded audio frame to generate data packets, and transmits those data packets to the server system 136 via the computer network 142.
[0057] The network transfer device 138 of the server system 136 receives the data packets from the glasses 120, applies a network transfer protocol to the data packets, and outputs an encoded audio frame. The network transfer device 128 provides the encoded audio frame to the audio decoder 144. The audio decoder 144 applies an audio decoding protocol to the encoded audio frame to output an audio frame, and provides those audio frames to one or more processors 1 to N for storage in one or more memory devices 1 to N.
[0058] Note that although FIG. 1A-3 is described with reference to a game program, it is equally applicable to other application programs such as a video conferencing application program.
[0059] In the embodiment, FIG. 1A-3 is described with reference to the user 1, the environment 120, and the glasses worn by the user 1, but FIG. 1A-3 is equally applicable to another user, another environment, and the glasses worn by another user.
[0060] FIG. 1A-4 is a diagram of an embodiment of a list 150 of objects displayed by a GPU 130 (FIG. 1A-2) to generate acoustic profiles of objects 108A-108P (FIG. 1A-1). The list 150 includes the state of the objects within the environmental system. For example, the list 150 includes a first entry where there are no blinds covering the window of the environment 102, and a second entry where there are blinds covering the window. As another example, the list 150 includes a third entry where the can is open, and a fourth entry where the can is closed. As yet another example, the list 150 includes the placement of the objects within the environment 102. For illustration purposes, the list 150 includes a fourth entry indicating that the can is closer to microphone M1 compared to object 108K or 108L, and a fifth entry indicating that object 108L is closer to microphone M1 compared to object 108K.
[0061] Instead of or in addition to providing message 121, during a training session, before or during the execution of the game program, a training program is executed by one or more of processors 1 to N (FIG. 1A-2), and list 150 is generated. For example, before or during the execution of the game program, server system 136 executes a training program to generate an image frame having list 150, encodes the image frame to output an encoded image frame having list 150, applies a network transfer protocol to the encoded image frame to generate data packets having list 150, and transmits those data packets to glasses 120 (FIG. 1A-2) via a computer network. In this example, network transfer device 126 (FIG. 1A-2) of glasses 120 applies a network transfer protocol to obtain an encoded image frame from the data packets and provides the encoded image frame to video decoder 128 (FIG. 1A-2) of glasses 120. In this example, video decoder 128 applies a video conversion protocol such as a video decoding protocol to decode the encoded image frame, outputs an image frame having list 150, and provides those image frames to GPU 130 (FIG. 1A-2). In this example, GPU 130 applies a rendering program to display list 150 on display screen 132 (FIG. 1A-2). In this example, the training program is computer software stored in one or more of memory devices 1 to N of server system 136.
[0062] When the user 1 views the list 150, the user 1 uses the input controller 122 to select one or more check boxes adjacent to one or more items in the list 150 to identify the objects O108A to O108P and children 1 and 2. The communication device of the input controller 122 applies a communication protocol to the selection of one or more check boxes in the list 150 to generate one or more transfer packets and transmits those transfer packets to the communication device 134 of the glasses 120. The communication device 134 applies a communication protocol to the transfer packets to obtain the list 150 from the transfer packets and transmits the selection of one or more check boxes in the list 150 to the CPU 135 of the glasses 120. The CPU 135 transmits the selection of one or more check boxes in the list 150 to the network transfer device 126 of the glasses 120. The network transfer device 126 applies a network transfer protocol to the selection of one or more check boxes in the list 150 to generate data packets. The network transfer device 126 transmits those data packets to one or more processors 1 to N of the server system 136 via the computer network 142. The network transfer device 138 receives the data packets from the computer network 142, applies a network transfer protocol to those data packets to obtain the selection of one or more check boxes in the list 150, and provides that selection to one or more processors 1 to N of the server system 136.
[0063] In one embodiment, instead of the list 150, a list of blank lines is generated by one or more processors 1 to N and transmitted to the glasses 120 (FIGS. 1A - 3) via the computer network 142 (FIGS. 1A - 3). The user 1 uses the input controller 122 to fill in the blank lines and provide a list of the objects 108A to 108P and children 1 and 2.
[0064] FIG. 1B is a diagram of an embodiment of another environment 152 in which user 1 and user 2 are playing game G1. Environment 152 is a room within a high-rise building. For example, the room is on the top floor of the building. Environment 152 includes spectators 1 and 2. Each of spectators 1 and 2 is an example of an object in environment 152. Further, environment 152 includes objects 154A, 154B, 154C, 154D, 154E, 154F, 154G, 154H, 154I, 154J, 154K, 154L, 154M, 154N, 108O, 154P, and 154Q. Object 154A is a banner suspended from the ceiling of environment 152. Object 154B is a display device. Examples of display devices are desktop monitors or televisions. Object 154C is a display device and object 154D is another display device. Each of objects 154C and 154D includes a CPU and a memory device located within the object. User 1 plays game G1 that displays a virtual scene on desktop monitor 154D. User 2 plays game G1 that displays a virtual scene on desktop monitor 154C.
[0065] Object 154E is a speaker coupled to Object 154C, and Object 154F is a speaker coupled to Object 154D. Object 154E is behind Object 154C, and Object 154F is behind Object 154D. Object 154G is a can or container, and Object 154H is the table on which Objects 154C and 154D are placed. By way of example, Object 154H has a top surface made of marble and has a textured surface. In Environment 152, the can is closed. Also, Object 154I is the floor of Environment 152. The floor is not carpeted and is bare. For example, the floor has a tiled surface. Object 154J is the chair on which User 1 sits, Object 154K is a mouse coupled to Object 154D, and Object 154L is a keyboard coupled to Object 154D. Object 154M is a mouse coupled to Object 154C, and Object 154N is a keyboard coupled to Object 154C.
[0066] Object 154P is the glasses worn by User 2. Object 154P includes a microphone M2 and a camera C2. Object 154Q is the cabinet stand on which Object 154B is placed. Also, Object 154R, which is an airplane, is located outside Environment 152 and is flying above the building while Users 1 and 2 are playing Game G1.
[0067] Environment 152 further includes Object 154S and another Object 154T. Object 154S is a window without blinds, and Object 154T is a can or container.
[0068] Note that user 1 moves from one location such as environment 102 (Figure 1A-1) to another location such as environment 152. User 1 accesses game G1 from the game cloud via computer network 142 (Figure 1A-2) and user account 1, and plays game G1 that represents a virtual scene on the display screen of object 154D. For example, user 1 selects one or more buttons on one or more of objects 154L and 154K to provide authentication information such as a username and password. Objects 154L and 154K send the authentication information to object 154D, and this object 154D transfers the authentication information to the game cloud via computer network 142. The authentication server of the game cloud determines whether the authentication information is genuine, and if it determines so, provides access to user account 1 and the game program running on the game cloud. When the game program is executed by one or more of processors 1 to N of the game cloud, an image frame of game G1 is generated and encoded, and the encoded image frame is output. The encoded image frame is sent to object 154D. Object 154D decodes the encoded image frame and provides the image frame to the display screen of object 154D of the virtual scene displayed on object 154D, enabling user 1 to play game G1.
[0069] During the play of game G1, the sound of game G1 is output from object 154F. For example, when a virtual object in the virtual scene displayed on object 154D jumps in the virtual scene and lands on the virtual ground, a landing sound is output via object 154F. As an example, the virtual object in the virtual scene displayed on object 154D is controlled by user 1 via objects 154L and 154K. Also, during the play of game G1, an airplane flying on environment 152 makes a sound like a sonic boom. Also, during the play of game G1, spectators 1 and 2 make sounds by talking to each other, arguing with each other, or playing with each other, etc. Also, during the play of game G1, user 1 opens a container placed on object 154H, and opening the container makes a sound. Moreover, during the play of game G1, user 1 utters words which are examples of sounds.
[0070] Since both User 1 and User 2 are within the same environment 152, they are in the same location. Similarly, User 2 accesses Game G1 from the game cloud via computer network 142 and User Account 2, and plays Game G1 that expresses a virtual scene on the display screen of object 154C. For example, User 2 selects one or more buttons on objects 154N and 154M to provide authentication information such as a username and password. Objects 154N and 154M send the authentication information to object 154C, and this object 154C transfers the authentication information to the game cloud via computer network 142. The authentication server of the game cloud determines whether the authentication information is genuine, and if it determines that it is, it provides access to User Account 2 and the game program running on the game cloud. When the game program is executed by one or more of Processors 1 to N of the game cloud, an image frame of Game G1 is generated and encoded, and the encoded image frame is output. The encoded image frame is sent to object 154C. Object 154C decodes the encoded image frame and provides the image frame to the display screen of object 154C of the virtual scene displayed on object 154C, enabling User 2 to play Game G1.
[0071] During the play of Game G1, the sound of Game G1 is output from object 154E. For example, if a virtual object within the virtual scene displayed on object 154C is flying within the virtual scene, a flying sound is output via object 154E. As an example, the virtual objects within the virtual scene displayed on object 154C are controlled by User 2 via objects 154N and 154M. Also, during the play of Game G1, User 2 opens object 154T placed on object 154H, and opening object 154T makes a sound. Also, during the play of Game G1, User 2 utters words, which are an example of sound.
[0072] Each of the microphones M1 and M2 captures audio data such as an audio frame generated from sound related to them, such as sound emitted from or reflected by one or more of the objects located in the environment 152. For example, the microphones M1 and M2 capture audio data generated based on the sound emitted from the object 154F and received from the object 154F via the path 156A. In this example, the sound path 156A is a direct path from the object 154F to the microphones M1 and M2 and does not impinge on any other surface between the object 154F and the microphones M1 and M2. As another example, the microphones M1 and M2 capture the sound emitted from the object 154F and received from the object 154F via the path 156B. In this example, the sound path 154F is an indirect path from the object 154F to the microphones M1 and M2. For the sake of explanation, the sound emitted from the object 154F impinges on one or more other objects (e.g., the object 154D) in the environment 152 and is reflected from those one or more other objects towards the microphones M1 and M2.
[0073] As yet another example, each of the microphones M1 and M2 captures audio data generated based on the sound emitted from the object 154E and received from the object 154E via the path 156C. In this example, the sound path 156C is a direct path from the object 154E to the microphones M1 and M2 and does not impinge on any other surface between the object 154E and the microphones M1 and M2. As yet another example, the microphones M1 and M2 capture the sound emitted from the object 154E and received from the object 154E via the path 156D. In this example, the sound path 156D is an indirect path from the object 154E to the microphones M1 and M2. For the sake of explanation, the sound emitted from the object 154E impinges on one or more other objects (e.g., the object 154C) in the environment 152 and is reflected from those one or more other objects towards the microphones M1 and M2.
[0074] Each of the microphones M1 and M2 also captures audio data generated based on sounds such as background noise emitted from an object 154R located outside the environment 152. The environment outside the environment 152 may be referred to as an external environment 158 in this specification. As an example, the microphones M1 and M2 capture sounds emitted from the object 154R and received through the ceiling of the environment 152.
[0075] The microphones M1 and M2 capture related sounds such as sounds or reflected sounds emitted from one or more of the objects 154A to 154N, 108O, 154P, 154Q, 154S, 154T, and the audience 1 and 2 within the environment 152, and background noise emitted from the object 154R within the external environment 158, and generate an audio frame such as audio data. An audio frame encoded based on the audio frame is generated and provided to one or more of the processors 1 to N of the game cloud via the computer network 142 for processing and training of the AI model.
[0076] In one embodiment, user 1 plays a game different from game G1, and user 2 plays a game different from game G1.
[0077] In an embodiment, the external environment 158 includes any other number such as two or three of the objects.
[0078] In one embodiment, instead of or in addition to the sound output from the object 154E, the sound is output from the object 154C, detected by the microphone M1 or M2 or a combination thereof, and the audio data is captured.
[0079] In an embodiment, instead of or in addition to the sound output from object 154F, the sound is output from object 154D and detected by microphone M1 or M2 or a combination thereof, and audio data is captured.
[0080] In an embodiment, instead of or in addition to microphone M2, one or more additional microphones, such as a stand-alone microphone, are present to capture sounds emitted from objects in environment 152 and sounds emitted from objects located in external environment 158. For example, a display device located on a table in environment 152 includes an additional microphone.
[0081] In an embodiment, the virtual scene displayed on object 154D is instead displayed on the display screen of the glasses worn by user 1. Similarly, the virtual scene displayed on object 154C is instead displayed on the display screen of the glasses worn by user 2.
[0082] Figure 2 is a diagram of an embodiment of a system 200 for explaining the processing of input data sets 1, 2, and 3 and audio data sets 1, 2, and 3. System 200 includes client device 1, client device 2, and client device 3. Further, system 200 includes a processor system 202. Examples of client device 1 operated by user 1 include glasses worn by user 1, object 108C (FIG. 1A-1), object 108N, game console 112 (FIG. 1A-1), object 108L (FIG. 1A-1), object 108K (FIG. 1A-1), object 154D (FIG. 1B), object 154L, object 154K, object 154F, and combinations of two or more thereof. Examples of client device 2 operated by user 2 include glasses worn by user 2, object 154C (FIG. 1B), object 154M, object 154N, object 154E, a game console, a handheld controller, and combinations of two or more thereof. Note that glasses 120 (FIG. 1A-2) are an example of glasses worn by user 2 except that camera C1 is replaced by camera C2 and microphone M1 is replaced by microphone M2. An example of client device 3 is provided below with reference to FIG. 5A. Client device 3 is operated by user 3.
[0083] Processor system 202 is coupled to client devices 1, 2, and 3. For example, processor system 202 is coupled to client devices 1-3 via computer network 142 (FIG. 1B). For the sake of explanation, processor system 202 includes one or more processors 1-N of server system 136 (FIG. 1A-2) coupled to client devices 1-3 via computer network 142.
[0084] An example of the audio dataset 1 includes audio data captured by the microphone M1 based on the sounds related to the environment 102 (FIG. 1A-1), and background noise received from the external environment 116 (FIG. 1A-1). An example of the audio dataset 2 includes audio data captured by the microphone M1 based on the sounds related to the environment 152 (FIG. 1B), and background noise received from the external environment 158 (FIG. 1B). An example of the audio dataset 3 includes audio data captured by the microphone M2 based on the sounds related to the environment 152, and background noise received from the external environment 158.
[0085] An example of the input dataset 1 includes one or more images of the objects in the environment 102 (FIG. 1A-1) captured by the camera C1, the identification of the objects in the environment 102, and the identification of the objects in the external environment 116 (FIG. 1A-1). For the sake of explanation, examples of the identification of the objects in the environment 102 and the objects in the external environment 116 are received within the list 150 (FIG. 1A-4).
[0086] An example of the input dataset 2 includes one or more images of the objects in the environment 152 (FIG. 1B) captured by the camera C1, the identification of the objects in the environment 152 received via the user account 1, and the identification of the objects in the external environment 158 (FIG. 1B) received via the user account 1. For the sake of explanation, the identification of the objects in the environment 152 and the objects in the external environment 158 is received in the form of a list from the user 1 via the user account 1 in the same way as the identification of the objects in the environment 102 and the objects in the external environment 116 is received via the list 150. In this explanation, the user 1 selects one or more objects in the list to provide the identification of the objects in the environment 152 and the objects in the external environment 158. The list is generated in the same way as the list 150 is generated and is displayed on the object 154D (FIG. 1B) operated by the user 1.
[0087] Examples of the input data set 3 include one or more images of objects in the environment 152 (FIG. 1B) captured by the camera C2, the identification of objects in the environment 152 received via the user account 2, and the identification of objects in the external environment ......
[0088] The processor system 202 includes a game engine and an inference training engine. Examples of the engines include hardware such as one or more controllers. In this example, each controller includes one or more processors such as processors 1 to N, or one or more processors of the game console 112 (FIG. 1A-1), or a combination thereof. As another example, the engine is software such as a computer software program. For illustration, the game engine is a game program. Another example of the engine is a combination of hardware and software.
[0089] The inference training engine includes an AI processor and a memory device 204 which is an example of one of the memory devices 1 to N. The AI processor is coupled to the memory device 204 and is an example of one of the processors 1 to N. Input data sets 1 to 3 and audio data sets 1 to 3 are stored in the memory device 204. For example, the AI processor receives the input data sets 1 to 3 and the audio data sets 1 to 3 from the client devices 1 and 2 via the computer network 142 and stores the input data sets 1 to 3 and the audio data sets 1 to 3 in the memory device 204. The game engine is coupled to the inference training engine.
[0090] The AI processor includes a feature extractor, a classifier, and an AI model. For example, the AI processor includes a first integrated circuit that applies the function of the feature extractor, a second integrated circuit that applies the function of the classifier, and a third integrated circuit that applies the function of the AI model. As another example, the AI processor executes a first computer program that applies the function of the feature extractor, a second computer program that applies the function of the classifier, and a third computer program that applies the function of the AI model. The feature extractor is coupled to the classifier, and the classifier is coupled to the AI model.
[0091] The feature extractor extracts, for example determines, parameters such as one or more amplitudes, one or more frequencies, and one or more sensing directions from audio data sets 1 to 3. For example, the feature extractor determines the size or peak-to-peak amplitude or zero-to-peak amplitude of audio data sets 1 to 3 and the frequencies of audio data sets 1 to 3. For the sake of explanation, the feature extractor determines the absolute maximum power or the absolute minimum power of audio data set 1 to determine the size of audio data set 1. In this explanation, the absolute power is the size within the entire period during which audio data set 1 is generated. As another explanation, the feature extractor determines the local maximum value of the size of audio data set 1 and the local minimum value of the size of audio data set 1. In this explanation, the local size is the size within a predetermined period, and the predetermined period is shorter than the entire period during which audio data set 1 is generated. In this explanation, a plurality of local maximum values of the size and a plurality of local minimum values of the size are determined from audio data set 1, and the best fit or average or median is applied to the local maximum value of the size and the local minimum value of the size by the feature extractor to determine the maximum value of the size and the minimum value of the size.
[0092] As another explanation, the feature extractor determines a first time when the audio dataset 1 reaches a predetermined size and a second time when the audio dataset 1 reaches the same predetermined size, and calculates the difference between the first time and the second time to determine a time interval. The feature extractor determines the absolute frequency of the audio dataset 1 by inverting the time interval. In this explanation, the absolute frequency is the frequency within the entire period during which the audio dataset 1 is generated. As yet another explanation, the feature extractor determines the local frequency of the audio dataset 1. In this explanation, the local frequency is the frequency within a predetermined period, and the predetermined period is shorter than the entire period during which the audio dataset 1 is generated. In this explanation, a plurality of local frequencies are determined from the audio dataset 1, and the best fit or average or median is applied to the local frequencies by the feature extractor to determine the frequency. In this explanation, each local frequency is determined in the same way as the method for determining the absolute frequency, except that the local frequency is determined for each predetermined period.
[0093] As another explanation, the feature extractor determines the direction in which the audio data set 1 is sensed. In this explanation, the microphone M1 includes an array of transducers (such as a linear array) arranged in a certain direction. In this explanation, the array includes a proximal transducer and a distal transducer. In this explanation, when the proximal transducer outputs a first portion of the audio data 1 and the distal transducer outputs a second portion of the audio data 1, and the first portion has an amplitude greater than the second amplitude, the feature extractor determines that the object 108M (FIG. 1A-1) in the environment 102 (FIG. 1A-1) is closer to the proximal transducer than to the distal transducer. In this explanation, the feature extractor further determines that the object 108M is in the direction facing the proximal transducer. On the other hand, as another explanation, when the second amplitude is greater than the first amplitude, the feature extractor determines that the object 108M is closer to the distal transducer than to the proximal transducer. In this explanation, the feature extractor further determines that the object 108M is in the direction facing the distal transducer. The direction in which the audio data set is sensed may be referred to herein as the sensing direction.
[0094] The classifier classifies parameters obtained from sounds related to environments 102, 116, 152, and 158 based on input data sets 1 - 3. For example, the classifier determines a combination of objects within the environment system, such as environment 102 (FIG. 1A - 1) or external environment 116 (FIG. 1A - 1) or a combination of environments 102 and 116, and establishes a correlation such as a one - to - one correspondence or unique relationship with the parameters determined from audio data set 1. In this example, the classifier determines the arrangement of objects, or the state of the objects, or a combination of two or more of the objects, object arrangement, and object state within the environment. For the sake of explanation, the classifier receives, within input data set 1, the identities of objects 108A - 108O within environment 102 and, via list 150 and user account 1, the identities of object 108P, child 1, and child 2. In this explanation, the classifier receives, within input data set 1, the states of objects 108A - 108O (FIG. 1A - 1). Further, in this explanation, the classifier determines, from the sensing direction, that object 108M is positioned proximal to microphone M1 as compared to object 108L and also as compared to object 108K (FIG. 1A - 1). Also, in this explanation, the classifier determines, from the sensing direction, that child 1 and 2 are positioned away from microphone M1 as compared to objects 108M, 108L, and 108K, and determines the arrangement of child 1 and 2 and objects 108M, 108L, and 108K. As another explanation, the classifier determines, from the sensing direction or amplitude or frequency or a combination thereof, that object 108P or child 1 or child 2 is located outside environment 102.
[0095] As another explanation, the classifier receives, within input data set 2, the identities of objects 154A - 154N, 108O, 154P, and 154Q within environment 152, as well as the identities of lists such as list 150 and object 154R via user account 1. In this explanation, the classifier receives, within input data set 2, the states of objects 154A - 154N, 108O, 154P, and 154Q (FIG. 1B). Further, in this explanation, the classifier determines, from the sensing direction, that object 154D is positioned proximal to microphone M1 as compared to object 154F (FIG. 1B). Also, in this explanation, the classifier determines, from the sensing direction, that object 154R is flying above environment 152 and determines the placement of object 154R relative to environment 152. As another explanation, the classifier determines, from the sensing direction or amplitude or frequency or a combination thereof, that object 154R is located outside environment 152.
[0096] As yet another explanation, the classifier receives, within input data set 3, the identities of objects 154A - 154N, 108O, 154P, and 154Q within environment 152, as well as the identities of lists such as list 150 and object 154R via user account 2. In this explanation, the classifier receives, within input data set 3, the states of objects 154A - 154N, 108O, 154P, and 154Q (FIG. 1B). Further, in this explanation, the classifier determines, from the sensing direction, that object 154C is positioned proximal to microphone M2 as compared to object 154E (FIG. 1B).
[0097] The AI model is trained based on the parameters related to environments 102, 116, 152, and 158 (FIGS. 1A-1 and 1B) and the correlations among input data sets 1-3. For example, referring to FIG. 3, the AI model is provided by the classifier with an indication of a correlation 302 (such as a link or a one-to-one correspondence) between a set including amplitude 1, frequency 1, and sensing direction 1 and a set including a first type of environment, a first combination of objects within the first type of environment, a first arrangement of the objects, and a first state of the objects. In this example, the AI model receives amplitude 1, frequency 1, and sensing direction 1 from the classifier. In this example, by analyzing audio data set 1 captured by microphone M1, amplitude 1, frequency 1, and sensing direction 1 are determined. In this example, one or more amplitudes determined from audio data set 1 are referred to herein as amplitude 1. Also, in this example, one or more frequencies determined from audio data set 1 are referred to herein as frequency 1, and the direction in which audio data set 1 is sensed is referred to herein as sensing direction 1. For the sake of explanation, amplitude 1, frequency 1, and sensing direction 1 are examples of the parameters of audio data set 1. Also, in this example, the classifier provides the AI model, via user account 1, with a first type of environment 102, a first combination of objects within environment 102 and external environment 116, a first state of the objects, and a first arrangement of the objects. For the sake of explanation, the first type of environment includes whether environment 102 is an indoor environment such as an open space within a room or a building, or an outdoor environment such as a park or a concert or a lake. As another explanation, the first combination of objects includes the identities of objects 108A-108O such as a can or a display device or a speaker or a window. In this explanation, the first arrangement of the objects includes that the can is located closer to microphone M1 than the speaker, and the first state includes that the can is open or closed. As another explanation, the first combination of objects includes the identities of objects 108A-108O such as a window with blinds or a carpet or a speaker.Also, in this example, the combination of the environment 102 and the external environment 116 is called the first environmental system.
[0098] As another example, referring to FIG. 3, the AI model is provided with an indication of a correlation 304 (such as a link or a one-to-one correspondence) between a set including amplitude 2, frequency 2, and sensing direction 2, and a set including a second type of environment, a second combination of objects in the second type of environment, a second arrangement of the objects, and a second state of the objects. In this example, the AI model receives amplitude 2, frequency 2, and sensing direction 2 from a classifier. In this example, amplitude 2, frequency 2, and sensing direction 2 are determined by analyzing an audio dataset 2 captured by microphone M1. Also, in this example, one or more amplitudes determined from audio dataset 2 are referred to herein as amplitude 2, one or more frequencies determined from audio dataset 2 are referred to herein as frequency 2, and the direction in which audio dataset 2 is sensed is referred to herein as sensing direction 2. For the sake of explanation, amplitude 2, frequency 2, and sensing direction 2 are examples of parameters of audio dataset 2. Also, in this example, the classifier provides the AI model with a second type of environment 152, a second combination of objects in environment 152 and the external environment 158, a second state of the objects, and a second arrangement of the objects. In this example, the second type, second combination, second state, and second arrangement are received via user account 1. For the sake of explanation, the second type of environment includes whether environment 152 is an indoor environment such as an open space in a room or a building, or an outdoor environment such as a park or a concert or a lake. As another explanation, the second combination of objects includes the identities of objects 154A - 154N, 108O, 154P - 154T such as cans or display devices or speakers or floors or ceilings. In this explanation, the second arrangement of the objects includes that object 154G is located closer to microphone M1 than object 154F, and the second state includes that object 154G is open or closed. As another explanation, the second arrangement includes that object 154B is farther away from microphone M1 compared to object 154F or 154D (FIG. 1B).In this description, the object 154B is a display device that emits the sound of a video game or another application. Also, in this example, the combination of the environment 152 and the external environment 158 is called a second environmental system.
[0099] As yet another example, referring to FIG. 3, the AI model is provided by the classifier with an indication of a correlation 306 (such as a link or a one-to-one correspondence) between a set including amplitude 3, frequency 3, and sensing direction 3, and a set including a third type of environment, a third combination of objects within the third type of environment, a third arrangement of the objects, and a third state of the objects. In this example, the AI model receives amplitude 3, frequency 3, and sensing direction 3 from the classifier. In this example, amplitude 3, frequency 3, and sensing direction 3 are determined by analyzing an audio data set 3 captured by microphone M2. Also, in this example, one or more amplitudes determined from the audio data set 3 are referred to herein as amplitude 3, one or more frequencies determined from the audio data set 3 are referred to herein as frequency 3, and the direction in which the audio data set 3 is sensed is referred to herein as sensing direction 3. For the sake of explanation, amplitude 3, frequency 3, and sensing direction 3 are examples of parameters of the audio data set 3. Also, in this example, the classifier provides the AI model with a third type of environment 152, a third combination of objects within the environment 152 and the external environment 158, a third state of the objects, as well as a third arrangement of the objects. In this example, the third type, the third combination, the third state, and the third arrangement are received via the user account 2. For the sake of explanation, the third type of environment includes whether the environment 152 is an indoor environment such as an open space within a room or a building, or an outdoor environment such as a park or a concert or a lake. As another explanation, the third combination of objects includes the identities of objects 154A-154N, 108O, 154P-154T such as cans or display devices or speakers or floors or ceilings. In this explanation, the third arrangement of the objects includes that the object 154T is located closer to the microphone M2 than the object 154E or 154C, and the third state includes that the object 154T is open or closed. Also, in this example, the combination of the environment 152 and the external environment 158 is referred to as the third environmental system.
[0100] In an embodiment, instead of or in addition to receiving the identity of an object in the environmental system, such as the list 150 (FIG. 1A-4), image data such as an image frame of the environmental system is received from cameras C1 and C2. In this embodiment, the feature extractor identifies an object in the environmental system from the image data. For example, the feature extractor identifies objects 108A-108P from the image frames of input data set 1. In this example, the feature extractor determines that the outline of object 108A matches the pre-stored outline of a chair, and determines that object 108A is a chair. In this example, the pre-stored outline is stored in one or more of memory devices 1-N. In this example, the feature extractor compares the size and shape of object 108O with the pre-stored size and pre-stored shape of pre-stored glasses, and determines that object 108O is glasses 120. In this example, the pre-stored size, the pre-stored shape, and the pre-stored identification of glasses are stored in memory device 204. In this example, the pre-stored identification of glasses includes alphanumeric characters. In this example, the feature extractor provides the identity of an object in the environmental system to the AI model.
[0101] In one embodiment, the feature extractor identifies the arrangement of objects in the environmental system, as well as graphical parameters such as color, luminance, gradation, and texture, from the image data received from cameras C1 and C2. For example, the feature extractor determines the relative positions of objects 108A-108P and the graphical parameters of objects 108A-108P. The arrangement and graphical parameters are stored in one or more of memory devices 1-N of the server system 136.
[0102] In one embodiment, instead of a processor, an application-specific integrated circuit (ASIC) or a programmable logic device (PLD) or a central processing unit (CPU) or a combination of a CPU and a GPU is used.
[0103] In an embodiment, instead of a game engine, an engine of another application such as a video conferencing application is used.
[0104] In one embodiment, the classifier receives the identities of objects 108A to 108O in environment 102 from the feature extractor. The identities of objects 108A to 108O are determined by the feature extractor from the images captured by camera C1 and user account 1. The images captured by camera C1 are part of input data set 1. Similarly, the classifier receives the identities of objects 154A to 154N, 108O, 154P to 154Q, 154S, and 154T in environment 152 from the feature extractor. The identities of objects 154A to 154N, 108O, 154P to 154Q, 154S, and 154T are determined by the feature extractor from the images captured by camera C1 and user account 1. The images captured by camera C1 are part of input data set 2. Also, the classifier receives the identities of objects 154A to 154N, 108O, 154P to 154Q, 154S, and 154T in environment 152 from the feature extractor. The identities of objects 154A to 154N, 108O, 154P to 154Q, 154S, and 154T are determined by the feature extractor from the images captured by camera C2 and user account 2. The images captured by camera C2 are part of input data set 3.
[0105] FIG. 3 is a diagram of an embodiment of system 300 for explaining correlations 302, 304, and 306. System 300 includes audio data sets 1 to 3.
[0106] FIG. 4A is a diagram of an embodiment of the state 400 of an object. For example, the state of a floor refers to whether the floor is made of carpet, tile, or concrete. As another example, the state of a window blind is whether the window blind is open or closed. As yet another example, the state of a container is whether the container is open or closed.
[0107] FIG. 4B is a diagram of an embodiment of the type 402 of an object. For example, the type of an object provides the identity of the object. For illustration purposes, a vehicle is a type of object, a monitor is a type of object, a display device is a type of object, an external speaker is a type of object, a human such as a child or spectator or user is a type of object, a house is a type of object, a building is a type of object, a carpet is a type of object, a bare floor is a type of object, a wall is a type of object, and a window is a type of object. Examples of each object 108K (FIG. 1A-1), 108L (FIG. 1A-1), 154E (FIG. 1B), and 154F (FIG. 1B) are external speakers. For illustration purposes, the external speakers are not integrated within another device such as a display device and are located external to the display device.
[0108] FIG. 5A is a diagram of an embodiment of environment 500 for explaining the use of an AI model to identify objects within environment 500, determine the state of the objects within environment 500, and determine the placement of the objects within environment 500. Environment 500 is a room within house 504. Environment 500 includes user 3, which is an example of an object. Further, environment 500 includes objects 502A, 502B, 502C, 502D, 502E, 502F, 502G, 502H, 502I, 502J, 502K, 502L, 502M, 502N, 502O, 502P, and 502Q. Examples of client device 3 (FIG. 2) include glasses worn by user 3, object 502I, object 502E, object 502D, object 502A, object 502G, an input controller coupled to the glasses, and combinations of two or more of them.
[0109] Object 502A is a display device including a computer. Object 502B is also a display device such as a monitor. Object 502C is the table on which objects 502A, 502B, 502D, 502E, 502F, and 502G are placed. Object 502D is a mouse coupled to object 502A, and object 502E is a keyboard coupled to object 502A. Object 502F is a stapler, and object 502G is a speaker coupled to object 502A. Object 502H is a robotic arm. Object 502I is a handheld controller used by user 3 to play game G1. Object 502J is the chair on which user 3 is sitting. Object 502K is the shelf that supports objects 502L, 502M, and 502N. Object 502L is a box, and object 502M is another box. Object 502N is a stack of containers. Object 502O is the glasses worn by user 3. As an example, object 502O has the same structure and function as glasses 120 (FIG. 1A-2). For example, object 502O includes camera C3 and microphone M3. In this example, glasses 120, which is an example of object 502O, includes camera C3 and microphone M3 instead of camera C1 and microphone M1. Object 502P is a window without blinds. Object 502Q is an open soda can.
[0110] User 3 accesses game G1 from the game cloud via computer network 142 (FIGS. 1A-2) and plays game G1 that causes a virtual scene to be represented on the display screen of object 502A. For example, User 3 selects one or more buttons on one or more of objects 502I, 502D, and 502E and provides authentication information such as a username and password. Object 502I or objects 502D and 502E send the authentication information to object 502A, and this object 502A transfers the authentication information to the game cloud via computer network 142. The authentication server of the game cloud determines whether the authentication information is authentic, and if it determines that it is, provides access to user account 3 and the game program running on the game cloud. When the game program is executed by one or more of processors 1 to N of the game cloud, an image frame of game G1 is generated and encoded, and the encoded image frame is output. The encoded image frame is sent to object 502A. Object 502A decodes the encoded image frames and provides those image frames to the display screen of object 502A. The display screen of object 502A applies a rendering program to the image frames to display an image of the virtual scene, and further enables User 3 to play game G1.
[0111] During the play of game G3, the sound of game G3 is output from object 502G. For example, when the virtual object in the virtual scene displayed on object 502A is a car accelerating on the virtual ground in the virtual scene, the sound of the car traveling is output via object 502G. As an example, the virtual object in the virtual scene displayed on object 502A is controlled by User 3 via object 502I or objects 502D and 502E. Also, during the play of game G1, object 502R, which is a car near house 504, makes a sound such as engine noise. Car 502R is part of the external environment 506 located outside house 504.
[0112] Microphone M3 captures audio data such as audio frames generated from sounds related to one or more of the objects located within environment 500, such as sounds emitted or reflected from those objects. For example, microphone M3 captures audio data generated based on the sound emitted from object 502G and received from object 502G via path 508A. In this example, the sound path 508A is a direct path from object 502G to microphone M3 and does not impinge on any other object between object 502G and microphone M3. As another example, microphone M3 captures the sound emitted from object 502G and received from object 502G via path 508B. In this example, the sound path 508B is an indirect path from object 502G to microphone M3. For the sake of explanation, the sound emitted from object 502G impinges on one or more other objects (e.g., object 502A and object 502C) within environment 500 and is reflected from those one or more other objects towards microphone M3.
[0113] Microphone M3 also captures audio data generated based on sounds such as background noise emitted from object 502R located outside environment 500. As an example, microphone M3 captures the sound emitted from object 502R and received via the walls or windows of environment 500.
[0114] The microphone M3 captures sounds related to those objects, such as sounds emitted or reflected from one or more objects (such as objects 502A to 502Q) within the environment 500, and background noise emitted from an object 502R within the external environment 506, and generates audio frames such as an audio dataset N (where N is a positive integer). An audio frame encoded based on the audio frame is generated and provided via the computer network 142 to one or more of processors 1 to N of the game cloud for processing and training of the AI model. Note as an example that there is no capture of an image of the environment 500 by the camera C3.
[0115] The feature extractor extracts, for example determines, parameters such as one or more amplitudes, one or more frequencies, and one or more sensing directions from the audio dataset N in the same way as the way the parameters are determined from audio datasets 1, 2, or 3. For example, the feature extractor determines the size of the audio dataset N, or the peak-to-peak amplitude, or the zero-to-peak amplitude, and the frequency of the audio dataset N. For the sake of explanation, the feature extractor determines the absolute maximum power of the audio dataset N, or the absolute minimum power of the audio dataset N to determine the size of the audio dataset N. As another explanation, the feature extractor determines the local maximum value of the size of the audio dataset N, and the local minimum value of the size of the audio dataset N. As another explanation, the feature extractor determines a plurality of local maximum values of the size and a plurality of local minimum values of the size, and the best fit or average or median is applied to the local maximum value of the size and the local minimum value of the size by the feature extractor to determine the maximum value of the size and the minimum value of the size. As another explanation, the feature extractor determines the first time when the audio dataset N reaches a predetermined size and the second time when the audio dataset N reaches the same predetermined size, and calculates the difference between the first time and the second time to determine the time interval. The feature extractor inverts the time interval to determine the absolute frequency of the audio dataset N. As yet another explanation, the feature extractor determines the local frequency of the audio dataset N. In this explanation, the local frequency is the frequency within a predetermined period, and the predetermined period is shorter than the entire period during which the audio dataset N is generated. In this explanation, a plurality of local frequencies are determined from the audio dataset N, and the best fit or average or median is applied to the local frequencies by the feature extractor to determine the frequency.
[0116] As another explanation, the feature extractor determines the direction in which the audio data set N is sensed. In this explanation, the microphone M3 includes an array of transducers (such as a linear array) arranged in a linear direction. This array includes a proximal transducer and a distal transducer. In this explanation, when the proximal transducer outputs a first portion of the audio data N and the distal transducer outputs a second portion of the audio data 1, and the first portion has an amplitude greater than the second amplitude, the feature extractor determines that the object 502G (FIG. 5A) in the environment 500 (FIG. 5A) is closer to the proximal transducer than to the distal transducer. In this explanation, the feature extractor further determines that the object 502G is in the direction facing the proximal transducer.
[0117] One or more amplitudes determined from the audio data set N are referred to herein as amplitude N. Also, one or more frequencies determined from the audio data set N are referred to herein as frequency N, and the direction in which the audio data set N is sensed is referred to herein as the sensing direction N.
[0118] In one embodiment, the user 3 plays a game different from the game G1.
[0119] In an embodiment, the external environment 506 includes any other number, such as two or three of the objects.
[0120] In an embodiment, the object 502O excludes the camera C3.
[0121] In one embodiment, a parameter may be referred to herein as a feature.
[0122] In an embodiment, instead of or in addition to the microphone M3, one or more additional microphones, such as a stand-alone microphone, are present to capture sounds emitted from objects in the environment 500 and sounds emitted from objects located in the external environment 506. For example, a display device located on a table in the environment 500 includes an additional microphone.
[0123] FIG. 5B is a diagram of an embodiment of the model output 550. The model output 550 includes a probability A% that the environment 500 includes a combination N of objects, a probability B% that the object has an arrangement N, a probability C% that the object has a state N, and a probability D% that the environment 500 is of type N, where A, B, C, and D are positive real numbers. As an example, A, B, C, and D are equal. As another example, one of A, B, C, and D is not equal to at least one of the others of A, B, C, and D.
[0124] When the AI model is provided with the amplitude N, the frequency N, and the sensing direction N from the feature extractor, the AI model provides the model output 500. For example, if it is determined that the amplitude N is within a predetermined range from amplitude 1 and outside a predetermined range from amplitude 2, the AI model indicates that there is a probability of more than 50% that the audio dataset N was received from a room in a house and not from a room in a building. In this example, the house is an example of an environment type N. For the sake of explanation, the house is an indoor type of environment. As another example, if it is determined that the frequency N is within a predetermined range from frequency 2 and outside a predetermined range from frequency 1, the AI model indicates that there is a probability of more than 50% that the audio dataset N was received from a room in a building and not from a room in a house. In this example, the building is an example of an environment type N. For the sake of explanation, the building is an indoor type of environment. As another example, a combination of two or more of the amplitude N, the frequency N, and the sensing direction N is used to determine the environment 500 of type N.
[0125] As yet another example, when it is determined that the sensing direction N is within a predetermined range from the sensing direction 1 and outside a predetermined range from the sensing direction 2, the AI model indicates that the probability that the audio dataset N is output from a speaker behind the display device as compared to a speaker in front of the display device of the environment 500 exceeds 50%. Depending on whether the speaker is behind or in front of the display device, an example of the arrangement N of the speaker and the display device is shown. As another example, a combination of two or more of the amplitude N, the frequency N, and the sensing direction N is used to determine the arrangement N of the objects in the environment 500.
[0126] As another example, when it is determined that the amplitude N is within a predetermined range from the amplitude 2 and outside a predetermined range from the amplitude 1, the AI model indicates that the probability that there is a window without a blind in the environment 500 exceeds 50%. In this example, the window without a blind is an example of the combination N of the objects or the objects in the environment 500. As another example, a combination of two or more of the amplitude N, the frequency N, and the sensing direction N is used to determine the combination N of the objects in the environment 500.
[0127] As yet another example, when it is determined that the amplitude N is within a predetermined range from the amplitude 1 and outside a predetermined range from the amplitude 2, the AI model indicates that the probability that a predetermined number of speakers exist in the environment 500 exceeds 50%. In this example, the predetermined number of speakers is an example of the combination N of the objects in the environment 500.
[0128] As yet another example, when it is determined that the amplitude N is within a predetermined range from the amplitude 1 and outside a predetermined range from the amplitude 2, the AI model indicates that the probability that the blind of the window in the environment 500 is open or the probability that the window has no blind exceeds 50%. In this example, the open or closed blind is an example of the state N of the objects in the environment 500. As another example, a combination of two or more of the amplitude N, the frequency N, and the sensing direction N is used to determine the state N of the objects in the environment 500.
[0129] As yet another example, if it is determined that the frequency N is within a predetermined range from frequency 1 and outside a predetermined range from frequency 2, the AI model instructs that the probability that the soda can in the environment 500 is open exceeds 50%. In this example, the fact that the soda can is open in the environment 500 is an example of the state N of the object in the environment 500.
[0130] In an embodiment, the server system 136 (FIG. 7) receives an input via the user account 3 and either the object 502I or the objects 502D and 502E. When the user 3 makes one or more selections on the object 502I or the objects 502D and 502E or the input controller to indicate that the user 3 wants to identify an object in the environment 500, the input is generated by the input controller coupled to the object 502I or the objects 502D and 502E or the object 502O. In this embodiment, the user 3 is wearing an HMD. The input is transmitted from the object 502I or the objects 502D and 502E or the input controller to the server system 136 via the object 502O or the object 502A and the computer network 142. When receiving an input requesting the identity of an object in the environment 500, the server system 136 applies, via the computer network 142, an AI model for obtaining probabilities A%, B%, C%, and D% regarding the combination N, arrangement N, state N, and type N of the environment 500, as well as the probabilities regarding the combination N, arrangement N, state N, and type N, to the object 502O or the object 502A. The object 502O or the object 502A displays these probabilities, as well as the combination N, arrangement N, state N, and type N of the object in the environment 500, on the display screen of the 502O or the object 502A.
[0131] In one embodiment, camera C3 captures image data of environment 500 and transmits the image data to server system 136 (FIG. 7) via computer network 142. Using the image data, combinations N, arrangements N, states N, and types N of environment 500 determined by the AI model are verified. One or more processors 1 to N of server system 136 verify the identities of objects within environment 500 determined by the AI model based on the image data. For example, processor N determines that there is a match between a first identity that object 502A is a display device and a second identity that object 502A is a display device. In this example, the first identity is determined by the AI model and the second identity is determined based on the image data. On the other hand, if it is determined that no match occurs, processor N reapplies the AI model to determine a third identity of object 502A, or determines to ignore the first identity and use the second identity.
[0132] FIG. 6A is a diagram of an embodiment of an acoustic profile database 600 stored in one or more of memory devices 1 to N (FIGS. 1A-2). Acoustic profile database 600 includes audio dataset 1, audio dataset 2, and audio dataset N. Acoustic profile database 600 also includes environment systems 1, 2, and N. Further, acoustic profile database 600 includes correspondences between audio datasets 1, 2, and N and environment systems 1, 2, and N. For example, acoustic profile database 600 includes a first exclusive correlation between audio dataset 1 and environment system 1, a second exclusive correlation between audio dataset 2 and environment system 2, and an Nth exclusive correlation between audio dataset N and environment system N. An example of environment system 1 is the first environment system, an example of environment system 2 is the second environment system (FIG. 3), and an example of the Nth environment system is type N of environment 500 or environment 500.
[0133] FIG. 6B is a diagram of an embodiment of a method 650 for explaining the application of audio data for outputting sounds corresponding to an environmental system. The method 650 is executed by one or more of processors 1 - N (FIGS. 1A - 2), or the processor of the game console 112 (FIG. 1A - 1), or a combination thereof. The method 650 includes an operation 652 of receiving a selection of a simulated environmental system presented to, e.g., the user 3. For example, the user 3 selects an environmental system N corresponding to the sound the user 3 wants to hear via the object 502I or a combination of the objects 502E and 502D. In this example, the user 3 is located in an environment different from the environment 500. For the sake of explanation, the object 502P has a blind or the object 502B (FIG. 5A) is removed from the environment 500 to create a different environment. In this example, the user 3 accesses a game program of game G1 or another application program of another application from the game cloud. For the sake of explanation, when the game program or other application is executed by one or more of the processors 1 - N, one or more image frames of one or more virtual scenes are generated and displayed on the object 502A. Further, in this example, during or before the execution of the game program or other application program, the user 3 selects the environmental system N. In this example, an indication of the selection of the environmental system N is transmitted from the object 502A to the server system 136 (FIG. 1A - 2) via the computer network 142. One or more of the processors 1 - N receive the indication of the selection and access the database 600 to determine that the audio data set N is exclusively associated with the environmental system N.
[0134] In operation 654 of method 650, one or more processors 1 to N access, for example, read from the database 600 to the audio dataset N. In operation 656 of method 600, one or more processors 1 to N apply the audio dataset N and output sound corresponding to the simulated environmental system N to the user 3. For example, one or more processors transmit the audio dataset N to the object 502A via the computer network 142. The object 502A outputs sound generated based on the audio dataset N during the play of the game G1 or during the execution of other application programs. The sound is output via the speaker of the object 502A or via the object 502G. When outputting sound that simulates the environmental system N, the user 3 feels as if he is in the environmental system N rather than in another environment.
[0135] FIG. 7 is a diagram of an embodiment of a system 700 for explaining method 600 (FIG. 6). The system 700 includes an input controller system 702, a display system 704, a server system 136, and a computer network 142. The display system 704 includes a communication device 706, a network transfer device 708, a CPU 712, a sound output system 710, and an audio memory device 714. An example of the display system 704 is the object 502O (FIG. 5A). Another example of the display system 704 is the object 502A (FIG. 5A). An example of the input controller system 702 is the input controller 122 (FIGS. 1A-2). Another example of the input controller system 702 is a handheld controller. Yet another example of the input controller system 702 is a combination of a mouse and a keyboard.
[0136] Communication device 706 has the same structure as communication device 134 (FIGS. 1A-2) and has similar functions to communication device 134. Also, network transfer device 708 has the same structure as network transfer device 126 (FIGS. 1A-2) and has similar functions to network transfer device 126. Further, CPU 712 has the same structure as CPU 135 (FIGS. 1A-2) and has similar functions to CPU 135. An example of audio memory device 714 is a buffer for storing audio data. Sound output system 710 includes a digital-to-analog converter (DAC), an amplifier, and a speaker.
[0137] Communication device 706 is coupled to input controller system 702 and CPU 712. CPU 712 is coupled to network transfer device 708, and this network transfer device is coupled to server system 136 via computer network 142. CPU 712 is coupled to the DAC, and the DAC is coupled to the amplifier. The amplifier is coupled to the speaker. CPU 712 is coupled to audio memory device 714.
[0138] Referring to FIG. 6, in operation 652, input controller system 702 generates an input signal, such as an indication, in response to the selection of environmental system N by user 3. User 3 uses input controller system 702 to select one or more buttons on input controller system 702. The communication device of input controller 702 applies a communication protocol to the indication of the selection of environmental system N to generate one or more transfer packets and transmits those transfer packets to communication device 706. Communication device 706 applies a communication protocol to the data packet to extract the indication of the selection of environmental system N from the transfer packet and transmits that indication to CPU 712.
[0139] The CPU 712 sends an indication of the selection of the environmental system N to the network transfer device 708. The network transfer device 708 applies the network transfer protocol to generate data packets containing the indication of the selection of the environmental system N, and transmits those data packets to the server system 136 via the computer network 142. The network transfer device 138 of the server system 136 applies the network transfer protocol to the data packets to extract the indication of the selection of the environmental system N from the data packets, and provides the indication to one or more of the processors 1 to N. The one or more processors 1 to N identify and access, for example, read, an audio data set N from one or more of the memory devices 1 to N corresponding to the environmental system N, perform an operation 654 (FIG. 6B), and transmit the audio data set N to the network transfer device 138. As an example, in addition to performing the operation 654, the one or more processors 1 to N access an audio file (including audio information) of a game program or other application program, and provide the audio information to the network transfer device 138. Examples of audio information are audio data where a virtual character in the game G1 jumps when the virtual character jumps. Another example of audio information is that when another application program is YouTube (registered trademark) (trademark), the user provides commentary in a YouTube (registered trademark) (trademark) video related to a sports event such as a basketball game. In this example, the sports event is an example of context.
[0140] The network transfer device 138 applies a network transfer protocol to generate data packets from the audio data set N and transmits those data packets via the computer network 142 to the display system 704 for the application of the audio data set N in operation 656 (FIG. 6B). For example, the network transfer device 138 embeds the audio data set N within the data packets and transmits those data packets to the network transfer device 708 via the computer network 142. In this example, the network transfer device 708 applies the network transfer protocol to the data packets to obtain the audio data set N and transmits the audio data set N to the CPU 712. The CPU 712 provides the audio data set N to the DAC, and the DAC converts the audio data set N from digital format to analog format and outputs an analog audio signal N. The DAC provides the analog audio signal N to an amplifier. The amplifier amplifies the amplitude of the analog audio signal N, for example, increases or decreases it, and outputs an amplified audio signal. The speaker converts the electrical energy of the analog audio signal N into sound energy, outputs the sound of the environmental system N based on the audio data set N, and simulates the environmental system N within the environment 500.
[0141] As another example, the network transfer device 138 embeds the audio data set N and the audio information related to the game program or other application program into data packets, and transmits those data packets to the network transfer device 708 via the computer network 142. In this example, the network transfer device 708 applies a network transfer protocol to the data packets to obtain the audio data set N and the audio information, and transmits the audio data set N and the audio information to the CPU 712. The CPU 712 provides the audio data set N and the audio information to the DAC, and the DAC converts the audio data set N and the audio information from digital format to analog format and outputs an analog audio signal. The DAC provides the analog audio signal to the amplifier. The amplifier amplifies the amplitude of the analog audio signal, for example, increases it, or decreases it, and outputs the amplified audio signal. The speaker converts the electrical energy of the analog audio signal into sound energy and outputs the sound of the first set of the game G1 or other application, and the sound of the second set of the environmental system N as the background for the sound of the first set. In this example, the first set and the second set are blended together when output simultaneously. In this example, the context of the audio information related to the game program or other application program matches the context of the audio data set N. For the sake of explanation, if a YouTube (registered trademark) (trademark) video includes commentary on a sports event such as a baseball game, the audio data set N represents the sounds that occur during a sports event such as a baseball game. In this example, the baseball game is an example of context.
[0142] In an embodiment, one or more of Processors 1 to N determine not to provide audio information corresponding to the virtual scene of Game G1 or the video of other application programs. For example, instead of accessing, from one or more of Memory Devices 1 to N, the audio information of a virtual character that jumps within a virtual scene, which is output together with the virtual scene, one or more of Processors 1 to N access Audio Dataset N from one or more of Memory Devices 1 to N and provide Audio Dataset N to be applied together with the virtual scene. As another example, instead of accessing, from one or more of Memory Devices 1 to N, the audio information that is output together with a YouTube (registered trademark) (trademark) video, one or more of Processors 1 to N access Audio Dataset N of one or more of Memory Devices 1 to N and provide Audio Dataset N to be applied together with the YouTube (registered trademark) (trademark) video. As yet another example, one or more of Processors 1 to N stop applying the audio information of a virtual character that jumps within a virtual scene and instead apply Audio Dataset N. As yet another example, one or more of Processors 1 to N stop applying the audio information that is output together with a YouTube (registered trademark) (trademark) video and instead apply Audio Dataset N.
[0143] In one embodiment, display system 704 includes additional components such as a microphone, a display screen, a GPU, an audio encoder, a video encoder, and a display screen. These components have a similar structure and similar functions to the corresponding components of glasses 120. For example, the microphone of display system 704 has the same structure as microphone M1, and when display system 704 is a display device, the display screen of display system 704 is larger than display screen 132. As another example, when display system 704 is glasses such as glasses 120 (FIGS. 1A-2), the display screen of display system 704 has the same size as display screen 132.
[0144] FIG. 8 is a diagram of an embodiment of a system 800 for explaining a microphone 800. The microphone 800 is an example of microphone M1, or M2, or M3. The microphone 800 includes a transducer, a converter, and an analog-to-digital converter (ADC). The transducer is coupled to a sound energy - electrical energy converter (S-E converter), and this converter is coupled to the ADC. An example of the transducer is a diaphragm. Examples of the S-E converter are a capacitor or a series of capacitors.
[0145] The transducer detects sound, either emitted or reflected, from an object in the environment and outputs vibrations. The vibrations are provided to the S-E converter, which changes the electric field generated within the S-E converter to output an audio analog signal, which is an electrical signal. The audio analog signal is provided to the ADC, which converts the audio analog signal from an analog format to a digital format to output audio data such as audio data set 1 or 2 or 3 or N.
[0146] FIG. 9 is a diagram of an embodiment of a system 900 for explaining a method for creating the effect of an environmental system N using direct audio data, determining the placement of objects within the environmental system N using reverberated audio data, and properties of the objects such as the type of material of the object and the type of surface of the object. The system 900 includes an audio data separator, direct audio data, reverberated audio data, a feature extractor, a classifier, and an AI model.
[0147] Examples of direct audio data are audio data generated by a microphone based on sound received via a direct path from a sound source. Also, examples of reverberant audio data are audio data generated by a microphone based on sound received via an indirect path from a sound source. For the sake of explanation, the first direct audio data of audio data set 1 is generated by microphone M1 (FIG. A) based on the sound received from object 108K (FIG. 1A-1) via path 106A. In this explanation, the first reverberant audio data of audio data set 1 is generated by microphone M1 based on the sound received from object 108K via path 106B (FIG. 1A-1). As another explanation, reverberant audio data is generated based on sound reflected or diffused from an object within an environmental system.
[0148] The audio data separator is implemented as hardware or software, or a combination thereof. For example, the audio data separator is a computer program, and the functions of the computer program are executed by one or more of processors 1 to N (FIG. 1A-2). As another example, the audio data separator is an ASIC or a PLD.
[0149] The audio data separator is coupled to the feature extractor and audio decoder 144 of server system 136 (FIG. 7). For example, the audio data separator is coupled between the feature extractor and audio decoder 144.
[0150] The audio data separator receives audio data sets 1, 2, 3, and N from client devices 1 and 2 (Figure 2) via computer network 142, network transfer device 138, and audio decoder 144 (Figure 7). The audio data separator determines the parameters of audio data sets 1, 2, 3, and N, and based on these parameters, identifies the direct audio data and reverberation audio data within each of audio data sets 1, 2, 3, and N. For example, the audio data separator identifies first direct audio data and first reverberation audio data within audio data set 1, and second direct audio data and second reverberation audio data within audio data set 2. For purposes of explanation, the audio data separator determines that the first part of audio data set 1 has a first amplitude that is greater than the second amplitude of the second part of audio data set 1. In this explanation, examples of amplitude include peak-to-peak amplitude and zero-to-peak amplitude. Further, in this explanation, the audio data separator determines that the first part is the first direct audio data and the second part is the first reverberation audio data. As another explanation, the audio data separator determines that the first part of audio data set 1 has a first frequency range and the second part of audio data set 1 has a second frequency range. Further, in this explanation, the audio data separator determines that the first part is the first direct audio data and the second part is the first reverberation audio data. The audio data separator determines the parameters of audio data sets 1, 2, 3, and N in the same way that the feature extractor determines the parameters of audio data sets 1, 2, 3, and N.
[0151] The direct audio data output from the audio data separator is stored in one or more of the memory devices 1 to N (Figure 7) by one or more of the processors 1 to N. Based on the direct audio data, an operation similar to operation 654 (Figure 6) is executed. For example, when receiving a selection of the environmental system N simulated for user 3, one or more of the processors 1 to N access the direct audio data corresponding to the environmental system N from one or more of the memory devices 1 to N. As another example, the same operation as operation 654 is executed except that the operation is performed on the direct audio data of the audio data set N instead of the audio data set N.
[0152] Furthermore, an operation similar to operation 656 (FIG. 6B) is performed to directly apply audio data and output a sound corresponding to the environmental system N. For example, an operation identical to operation 656 is performed, except that instead of the audio data set N, the direct audio data of the audio data set N is applied. As another example, the direct audio data of the audio data set N is transmitted from the server system 136 (FIGS. 1A-2) to the object 502A or 502O (FIG. 5A). In this example, the object 502A or 502O outputs a sound generated based on the direct audio data of the audio data set N during the play of the game G1 or during the execution of other application programs. As yet another example, the direct audio data of the audio data set N is synthesized with the reverberated audio data of the audio data set N, and a sound based on the audio data set N is output. For the sake of explanation, one or more of the processors 1 to N combine the direct audio data of the audio data set N with the reverberated audio data of the audio data set N to output the audio data set N. In this explanation, the audio data set N is then applied to operation 656 in the above-described manner to simulate the environmental system N for the user 3. As yet another example, the direct audio data of the audio data set N is synthesized by one or more of the processors 1 to N with the reverberated audio data of another audio data set such as the audio data set 1 or 2 or 3, and an additional audio data set is generated. In this example, the additional audio data set is applied in the same manner as the audio data set N is applied in operation 656. For the sake of explanation, the additional audio data set is transmitted from the server system 136 to the object 502A or 502O via the computer network 142, and a sound is output based on the additional audio data set.
[0153] The reverberation audio data output from the audio data separator is sent to the feature extractor. For example, the first reverberation audio data and the second reverberation audio data are sent from the audio data separator to the feature extractor.
[0154] The feature extractor determines the parameters of the reverberation audio data of any of the audio data sets 1, 2, 3, and N in the same way as the feature extractor determines the parameters of the audio data set. For example, the feature extractor determines the amplitude 1a of the reverberation audio data of audio data set 1, the amplitude 2a of the reverberation audio data of audio data set 2, the amplitude 3a of the reverberation audio data of audio data set 3, and the amplitude Na of the reverberation audio data of audio data set N. As another example, the feature extractor determines the frequency 1a of the reverberation audio data of audio data set 1, the frequency 2a of the reverberation audio data of audio data set 2, the frequency 3a of the reverberation audio data of audio data set 3, and the frequency Na of the reverberation audio data of audio data set N. The feature extractor sends the parameters of the reverberation audio data of audio data sets 1, 2, and 3 to the classifier. Also, the feature extractor sends the parameters of the reverberation audio data of audio data set N to the AI model.
[0155] The classifier classifies the parameters of the reverberated audio data of the audio data sets 1 to 3 based on the input data sets 1 to 3. For example, the classifier determines, or identifies, a combination of objects within the environment system, such as environment 102 (FIG. 1A-1) or external environment 116 (FIG. 1A-1) or a combination thereof, from input data set 1, and establishes a correlation (such as a one-to-one correspondence) between the combination of objects and the parameters of the reverberated audio data of audio data set 1. In this example, the classifier further determines, or identifies, the arrangement of the objects relative to each other, or the state of the objects, or the type of the surface of the objects, or the type of the material of the objects, or a combination of two or more of them. In this example, input data set 1 includes the type of the material of the objects and the type of the surface of the objects. For the sake of explanation, the classifier receives, within input data set 1, the type of the material of objects 108A to 108O (FIG. 1A-1) and the type of the surface of objects 108A to 108O.
[0156] Examples of the type of material include wood, plastic, glass, marble, stainless steel, leather, wool, cloth, cotton, polyester, tile, granite. For the sake of explanation, list 150 (FIG. 1A-4) includes an entry indicating that the desktop table is made of wood and another entry indicating that the chair is made of leather. In this explanation, in the same way that user 1 selects the chair and the desktop table within list 150, user 1 selects the type of the material used in the manufacture of the desktop table and the type of the material used in the manufacture of the chair. As another explanation, list 150 includes an entry indicating that the desktop table has a rough or smooth surface, and an entry indicating that the surface of the seat of the chair is intact, or worn such as being about to tear or torn. In this explanation, in the same way that user 1 selects the chair and the desktop table within list 150, user 1 selects the type of the surface of the desktop table and the type of the surface of the chair.
[0157] As another description of the classification, the classifier receives, within input data set 2, the identities of objects 154A - 154N, 108O, 154P, and 154Q within environment 152, as well as the identities of lists such as list 150 and object 154R via user account 1. In this description, the classifier receives, within input data set 2, the states, material types, and surface types of objects 154A - 154N, 108O, 154P, and 154Q (FIG. 1B). Further, in this description, the classifier determines, from the sensing direction of the reverberant audio data of audio data set 2, that object 154D is positioned proximal to microphone M1 as compared to object 154F (FIG. 1B).
[0158] As yet another description, the classifier receives, within input data set 3, the identities of objects 154A - 154N, 108O, 154P, and 154Q within environment 152, as well as the identities of lists such as list 150 and object 154R via user account 2. In this description, the classifier receives, within input data set 3, the states, material types, and surface types of objects 154A - 154N, 108O, 154P, and 154Q (FIG. 1B). Further, in this description, the classifier determines, from the sensing direction, that object 154C is positioned proximal to microphone M2 as compared to object 154E (FIG. 1B). In this description, the sensing direction is determined from the parameters of the reverberant audio data of audio data set 2.
[0159] The AI model is trained based on the correlation between the reverberated audio data of audio datasets 1 to 3 and the parameters of input datasets 1 to 3 related to environments 102, 116, 152, and 158 (FIGS. 1A-1 and 1B). For example, the AI model is provided by a classifier with an indication of a first correlation, such as a link or a one-to-one correspondence, between a set including the amplitude 1a of the reverberated audio data of audio dataset 1, the frequency 1a of the reverberated audio data of audio dataset 1, and the perceived direction 1a of the reverberated audio data of audio dataset 1, and a set including a first type of environment, a first combination of objects within the first type of environment, a first arrangement of the objects, a first state of the objects, a first set of types of materials of the objects, and a first set of types of surfaces of the objects. In this example, one or more amplitudes determined from the reverberated audio data of audio dataset 1 are referred to herein as amplitude 1a. Also, in this example, one or more frequencies determined from the reverberated audio data of audio dataset 1 are referred to herein as frequency 1a, and the direction in which the reverberated audio data of audio dataset 1 is perceived is referred to herein as perceived direction 1a. For the sake of explanation, amplitude 1a, frequency 1a, and perceived direction 1a are examples of parameters of the reverberated audio data of audio dataset 1. Also, in this example, the AI model is provided by a classifier with a first type of environment 102, a first combination of objects within environments 102 and external environment 116, a first state of the objects, a first arrangement of the objects, a first set of types of materials of the objects, and a first set of types of surfaces of the objects via user account 1.
[0160] As another example, the AI model is provided with an indication of a second correlation, such as a one-to-one correspondence, between a set including the amplitude 2b, frequency 2b, and sensing direction 2a of the reverberated audio data of the audio dataset 2, and a set including a second type of environment, a second combination of objects in the second type of environment, a second arrangement of the objects, a second state of the objects, a second set of types of materials of the objects, and a second set of types of surfaces of the objects. In this example, the AI model receives the amplitude 2b, frequency 2b, and sensing direction 2b from the classifier. In this example, the amplitude 2b, frequency 2b, and sensing direction 2b are determined by analyzing the reverberated audio data of the audio dataset 2 captured by the microphone M1. Also, in this example, one or more amplitudes determined from the reverberated audio data of the audio dataset 2 are referred to herein as amplitude 2b, one or more frequencies determined from the reverberated audio data of the audio dataset 2 are referred to herein as frequency 2b, and the direction in which the reverberated audio data of the audio dataset 2 is sensed is referred to herein as sensing direction 2b. For the sake of explanation, the amplitude 2b, frequency 2b, and sensing direction 2b are examples of parameters of the reverberated audio data of the audio dataset 2. Also, in this example, by the classifier, the AI model is provided with a second type of the environment 152, a second combination of objects in the environment 152 and the external environment 158, a second state of the objects, a second arrangement of the objects, a second set of types of materials of the objects, and a second set of types of surfaces of the objects. In this example, the second type, second combination, second state, and second arrangement, the second set of types of materials of the objects, as well as the second set of types of surfaces of the objects are received via the user account 1.
[0161] As yet another example, the AI model is provided by the classifier with an indication of a third correlation (such as a link) between a set including the amplitude 3a, frequency 3a, and sensing direction 3a of the reverberated audio data of the audio dataset 3, and a set including a third type of environment, a third combination of objects within the third type of environment, a third arrangement of the objects, a third state of the objects, a third set of types of materials of the objects, and a third set of types of surfaces of the objects. In this example, the AI model receives the amplitude 3c, frequency 3c, and sensing direction 3c from the classifier. In this example, the amplitude 3c, frequency 3c, and sensing direction 3c are determined by analyzing the reverberated audio data of the audio dataset 3 captured by the microphone M2. Also, in this example, one or more amplitudes determined from the reverberated audio data of the audio dataset 3 are referred to herein as amplitude 3c, one or more frequencies determined from the reverberated audio data of the audio dataset 3 are referred to herein as frequency 3c, and the direction in which the reverberated audio data of the audio dataset 3 is sensed is referred to herein as sensing direction 3c. For the sake of explanation, the amplitude 3c, frequency 3c, and sensing direction 3c are examples of parameters of the reverberated audio data of the audio dataset 3. Also, in this example, the classifier provides the AI model with a third type of environment 152, a third combination of objects within the environment 152 and the external environment 158, a third state of the objects, a third arrangement of the objects, a third set of types of materials of the objects, and a third set of types of surfaces of the objects. In this example, the third type, third combination, third state, third arrangement, third set of types of materials of the objects, and third set of types of surfaces of the objects are received via the user account 2.
[0162] When the AI model is provided with the amplitude Na, the frequency Na, and the sensing direction Na from the feature extractor, the AI model provides a model output 902. For example, if it is determined that the amplitude Na is within a predetermined range from amplitude 1a and outside a predetermined range from amplitude 2a, the AI model indicates that the probability that reverberant audio data of the audio dataset N is generated based on the sound reflected from a plastic table or a table with a smooth upper surface exceeds 50%. In this example, the probability that the table is made of plastic or the upper surface of the table is smooth is an example of the model output 902. As another example, if it is determined that the frequency Na is within a predetermined range from frequency 2a and outside a predetermined range from frequency 1a, the AI model indicates that the probability that reverberant audio data of the audio dataset N is generated based on the sound reflected from a table with a rough surface or a table made of marble on its upper surface exceeds 50%. In this example, the probability that the table is made of marble or the table has a rough surface is an example of the model output 902. As another example, a combination of two or more of the amplitude Na, the frequency Na, and the sensing direction Na is used to determine the type of material of any object within the environment 500 or the type of surface of any object within the environment 500.
[0163] Note that the type of material of the object and the type of surface of the object are examples of the properties of the object.
[0164] In one embodiment, during operation 654, a visual mapping of the scene of environment N is created on display screen 132 (FIG. 1A-2) based on the image data captured by cameras C1 and C2 (FIG. 1B). For example, in addition to directly applying audio data to simulate environment system N, the arrangement of objects within environment system N and the graphical parameters of the objects within environment system N are applied to simulate environment system N. For purposes of explanation, the arrangement of objects within environment system N and the graphical parameters of the objects within environment system N are accessed by one or more of processors 1-N from one or more of memory devices 1-N of server system 136 and transmitted to glasses 120 (FIG. 1A-2) via computer network 142. When GPU 130 (FIG. 1A-2) receives the arrangement of objects within environment system N and the graphical parameters of the objects within environment system N, it displays the graphical parameters of the objects according to those arrangements and simulates environment system N when user 3 is in different environments.
[0165] Note that in various embodiments, one or more features of some of the embodiments described herein may be combined with one or more features of one or more of the remaining embodiments described herein.
[0166] The embodiments described in this disclosure can be implemented in various computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, and mainframe computers. In one embodiment, the embodiments described in this disclosure are practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a wired or wireless network.
[0167] With the foregoing embodiments in mind, in one embodiment, it should be understood that the embodiments described in the present disclosure use various computer-implemented operations involving data stored in a computer system. These operations are operations that require physical manipulation of physical quantities. Any of the operations described herein that form part of the embodiments described in the present disclosure are useful mechanical operations. Some embodiments described in the present disclosure also relate to devices or apparatuses for performing these operations. The apparatus is specially constructed for the required purpose, or the apparatus is a general-purpose computer selectively activated or configured by a computer program stored in a computer. Specifically, in one embodiment, various general-purpose machines are used in conjunction with a computer program written in accordance with the teachings herein. Or, it may be more convenient to construct a more specialized apparatus for performing the required operations.
[0168] In some embodiments, some of the embodiments described in the present disclosure are embodied as computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device that stores data that can later be read by a computer system. Examples of computer-readable media include hard drives, network attached storage (NAS), ROM, RAM, compact disc ROM (CD-ROM), CD-recordable (CD-R), CD-rewritable (CD-RW), magnetic tape, optical data storage devices, non-optical data storage devices, and the like. As an example, a computer-readable medium includes computer-readable tangible media distributed on a network-coupled computer system, and the computer-readable code is stored and executed in a distributed manner.
[0169] Furthermore, although some of the above-described embodiments have been described with respect to a game environment, in some embodiments, instead of a game, other environments such as, for example, a video conferencing environment are used.
[0170] Although the method operations have been described in a particular order, other housekeeping operations may be performed during the operations, or the operations may be adjusted to occur at slightly different times, or the processing of the overlay operations may be distributed within the system that allows the processing operations to occur at various intervals as long as the processing is performed in the desired manner. It should be understood that this is the case.
[0171] The foregoing embodiments described in this disclosure have been described in some detail for clarity of understanding, but it will be apparent that certain changes and modifications can be made within the scope of the appended claims. Therefore, these embodiments should be regarded as illustrative rather than limiting, and these embodiments should not be limited to the details described herein, but may be modified within the scope of the appended claims and the scope of equivalents.
Claims
1. A method for determining the environment in which a first user is located, comprising: receiving a plurality of audio datasets based on sounds emitted in a plurality of environments, each of the plurality of environments having a different combination of objects, said receiving; receiving input data regarding the plurality of environments; training an artificial intelligence model (AI model) based on the plurality of audio datasets and the input data; applying the AI model to audio data captured from the environment surrounding the first user to determine the type of the environment; further comprising extracting a plurality of features from the plurality of audio datasets, the plurality of features including a plurality of amplitudes of the plurality of audio datasets, a plurality of frequencies of the plurality of audio datasets, and a plurality of perception directions in which the sound is perceived.
2. further comprising classifying the plurality of features and outputting a correlation between the plurality of features and a plurality of types of the plurality of environments, objects in the plurality of environments, a plurality of arrangements of the objects, and a plurality of states of the objects in the plurality of environments; said training the AI model comprises: providing the AI model with the correlation between the plurality of features and the plurality of types of the plurality of environments, objects in the plurality of environments, the plurality of arrangements of the objects, and the plurality of states of the objects in the plurality of environments; determining, by the AI model, a plurality of probabilities based on the correlation between the plurality of features and the plurality of types of the plurality of environments, objects in the plurality of environments, the plurality of arrangements of the objects, and the plurality of states of the objects in the plurality of environments, the plurality of probabilities providing a probability that the audio data captured from the environment indicates the type of the environment, a probability that the type of the environment includes a plurality of items, a probability that the plurality of items have a plurality of states, and a probability that the plurality of items have an arrangement, said determining; The method according to claim 1, comprising.
3. Receiving an indication of the type of the environment simulated by the AI model Accessing the audio data captured from the environment based on the type of the environment simulated by the AI model; Providing the audio data captured from the environment to a client device to output a sound corresponding to the type of the environment; The method according to claim 1, further comprising the above.
4. The method according to claim 1, wherein the plurality of audio data sets include audio data generated from sounds emitted from one or more of the objects in the plurality of environments and sounds reflected from the remaining of the objects.
5. The method according to claim 1, wherein the input data includes data for identifying the objects in the plurality of environments, or image data captured by a camera in the plurality of environments, or a combination thereof.
6. The method according to claim 1, wherein when a second user in the same environment as the first user moves from one position to another position, the plurality of audio data sets are captured.
7. The method according to claim 1, wherein when there are a plurality of users including a second user and a third user, the plurality of audio data sets are captured.
8. The method according to claim 1, further comprising applying the AI model to the audio data captured from the environment surrounding the first user to identify one or more objects in the environment surrounding the first user, or one or more states of the one or more objects, or the arrangement of the one or more objects, or a combination thereof.
9. A server for determining an environment in which a user is located, configured to receive a plurality of audio data sets based on sounds emitted in a plurality of environments via a computer network, each of the plurality of environments having a different combination of objects; configured to receive input data regarding the plurality of environments via the computer network; configured to train an artificial intelligence (AI) model based on the plurality of audio data sets and the input data; a processor configured to apply the AI model to audio data captured from an environment surrounding a user to determine the type of the environment; A memory device coupled to the processor, comprising The processor is configured to extract a plurality of features from the plurality of audio datasets, the plurality of features including a plurality of amplitudes of the plurality of audio datasets, a plurality of frequencies of the plurality of audio datasets, and a plurality of sensing directions in which the sound is sensed, a server.
10. The processor is configured to classify the plurality of features and output a correlation between the plurality of features and a plurality of types of the plurality of environments, objects in the plurality of environments, a plurality of arrangements of the objects, and a plurality of states of the objects in the plurality of environments. To train the AI model, the processor is configured to provide the AI model with the correlation between the plurality of features and the plurality of types of the plurality of environments, objects in the plurality of environments, the plurality of arrangements of the objects, and the plurality of states of the objects in the plurality of environments. The server according to claim 9, wherein the AI model is used to determine a plurality of probabilities based on the correlation between the plurality of features and the plurality of types of the plurality of environments, objects in the plurality of environments, the plurality of arrangements of the objects, and the plurality of states of the objects in the plurality of environments, the plurality of probabilities including a probability that the audio data captured from the environment indicates the type of the environment, a probability that the type of the environment includes a plurality of items, a probability that the plurality of items have a plurality of states, and a probability that the plurality of items have an arrangement.
11. The processor is configured to receive an indication of a type of an environment simulated by the AI model, configured to access the audio data captured from the environment based on the type of the environment simulated by the AI model, The server according to claim 10, configured to provide the audio data captured from the environment to a client device and output a sound corresponding to the type of the environment.
12. The server according to claim 9, wherein the plurality of audio data sets include audio data generated from sounds emitted from one or more of the objects in the plurality of environments and sounds reflected from the remaining objects.
13. The server according to claim 9, wherein the input data includes data for identifying the objects in the plurality of environments, or image data captured by cameras in the plurality of environments, or a combination thereof.
14. The server according to claim 9, wherein the processor is configured to apply the AI model to the audio data captured from the environment surrounding the user to identify one or more objects in the environment surrounding the user, or one or more states of the one or more objects, or the arrangement of the one or more objects, or a combination thereof.
15. A system for determining an environment in which a user is located, configured to generate a plurality of audio data sets based on sounds emitted in a plurality of environments, each of the plurality of environments having a different combination of objects, a plurality of client devices configured to receive input data regarding the plurality of environments, a server coupled to the plurality of client devices, configured to receive the plurality of audio data sets from the plurality of client devices via a computer network, configured to receive the input data regarding the plurality of environments from the plurality of client devices via the computer network, configured to train an artificial intelligence (AI) model based on the plurality of audio data sets and the input data, the server configured to apply the AI model to audio data captured from an environment surrounding the user to determine the type of the environment, comprising The server is configured to extract a plurality of features from the plurality of audio data sets, the plurality of features including a plurality of amplitudes of the plurality of audio data sets, a plurality of frequencies of the plurality of audio data sets, and a plurality of sensing directions for sensing the sound. system.
16. The server is configured to classify the plurality of features and output a correlation between the plurality of features and a plurality of types of the plurality of environments, objects in the plurality of environments, a plurality of arrangements of the objects, and a plurality of states of the objects in the plurality of environments. To train the AI model, the server is configured to provide the AI model with the correlation between the plurality of features and the plurality of types of the plurality of environments, objects in the plurality of environments, the plurality of arrangements of the objects, and the plurality of states of the objects in the plurality of environments. The system according to claim 15, wherein the AI model is configured to determine a plurality of probabilities based on the correlation between the plurality of features and the plurality of types of the plurality of environments, objects in the plurality of environments, the plurality of arrangements of the objects, and the plurality of states of the objects in the plurality of environments, and the plurality of probabilities provide a probability that the audio data captured from the environment indicates the type of the environment, a probability that the type of the environment includes a plurality of items, a probability that the plurality of items have a plurality of states, and a probability that the plurality of items have an arrangement. [
17. ] The server is configured to receive an indication of a type of an environment simulated by the AI model, is configured to access the audio data captured from the environment based on the type of the environment simulated by the AI model, The system according to claim 16, wherein the server is configured to provide the audio data captured from the environment to a client device and cause the client device to output a sound corresponding to the type of the environment.
Citation Information
Patent Citations
Acoustic program, acoustic device, and acoustic system
EP3799035A1
System and method to convert two-dimensional video into three-dimensional extended reality content
US11551407B1
Stereophonic apparatus for blind and visually-impaired people
US20200258422A1
Room acoustics simulation using deep learning image analysis
WO2020139588A1
Systems and methods for generating audio presentations
WO2021216060A1