System and method for identifying the material type of an object in a real-world environment
Patent Information
- Application Number
- JP2024543581
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-01-31
- Filing Date
- 2023-01-19
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2043-01-19
Smart Images

Figure 0007914220000001 
Figure 0007914220000002 
Figure 0007914220000003
Abstract
Description
[[TECHNICAL FIELD]]
[0001] The present disclosure relates to systems and methods for identifying material types of objects in a real-world environment. [[BACKGROUND ART]]
[0002] In a multiplayer game, there are a plurality of game players. Each player wears a head-mounted display (HMD) to play a game or view an environment generated by executing an application. Several objects are displayed on the HMD during playing the game or executing the application. However, sometimes players cannot recognize these objects in the environment.
[0003] Against this background, embodiments of the present invention have been developed. [[SUMMARY OF THE INVENTION]]
[0004] Embodiments of the present disclosure provide systems and methods for identifying material types of objects in a real-world environment.
[0005] In embodiments, the materials of objects in a real-world environment may have properties that allow them to produce sound when interacted with, such as through compression, contact, and movement. For example, when a user sits on a chair, the chair may react with noise based on the physical properties of the chair's seat material. For instance, if the chair seat is made of plastic, when a user sits on it, it may produce a squelching sound that suggests air is being released. By identifying the properties of an object, this information can be used to extend virtual spaces, such as virtual reality (VR) or augmented reality (AR) environments, to mimic the object being interacted with. For example, if a user sits on a soft chair in a real-world environment, a similar soft chair with similar properties that evoke or relate to the sound detected when the user sits on the chair may be depicted in the AR view or VR space. As another example, if a user sits on a bath seat made of inexpensive plastic and foam, a virtual bath seat with similar properties that evoke or relate to the sound detected when the user sits on it may be depicted in the AR view or VR space. To leverage these properties, a VR reproduction of a certain type of chair may be created, and a virtual user sitting on the same type of chair may be shown. Therefore, when mapping is performed from the real-world environment to the virtual space, and similar virtual objects are placed in the virtual space, the surface, sound, and characteristics of the objects in the real-world environment are mimicked and reproduced using audio cues from the real-world environment. Physical phenomena related to the sounds emitted by real-world objects can be reproduced by virtual objects in the virtual space. Therefore, due to the physical phenomena reproduced in the virtual space, when a virtual user or character in the game sits on a virtual chair, the virtual seat of the virtual chair may appear to indent.
[0006] In one embodiment, when a specific material or object is deformed in the real world, lighting, sound, and other environmental characteristics are detected to determine how that material or object will react in the real world. The reaction of the real-world object is then used in the virtual space to display a virtually changing object, which corresponds to the deformation and change occurring in the real-world environment.
[0007] In one embodiment, a method for identifying the material type of an object in a real-world environment is described. The method includes receiving multiple audio datasets based on sounds received from multiple objects in multiple environments. The method further includes receiving multiple input datasets relating to multiple material types of multiple objects, training and / or inferring from an artificial intelligence (AI) model based on the multiple audio datasets and the multiple input datasets, and applying the AI model to audio datasets captured from a real-world environment to identify the material type of an object in the real-world environment. For example, the input data may include audio data, image data, LiDAR data, or additional input data such as inertial measurement unit (IMU) data, or a combination of audio data, image data, LiDAR data, and additional input data.
[0008] In one embodiment, a server for identifying the material type of an object in a real-world environment is described. The server includes a processor that receives multiple audio datasets based on sounds received from multiple objects in multiple environments. The processor further receives multiple input datasets relating to multiple material types of multiple objects. The processor also trains and / or infers from an AI model based on the multiple audio datasets and the multiple input datasets. For example, the input data may include audio data, image data, LiDAR data, or additional input data such as IMU data, or a combination of audio data, image data, LiDAR data, and additional input data. The processor applies the AI model to the audio datasets captured from the real-world environment to identify the material type of an object in the real-world environment. The server includes a memory device connected to the processor.
[0009] In one embodiment, a system for identifying the material type of an object in a real-world environment is described. The system includes multiple client devices. The multiple client devices generate multiple audio datasets based on sounds received from multiple objects in multiple environments. The client devices also receive multiple input datasets relating to multiple material types of multiple objects. For example, the input data may include audio data, image data, LiDAR data, or additional input data such as IMU data, or a combination of audio data, image data, LiDAR data, and additional input data. The system further includes a server connected to the multiple client devices via a computer network. The server receives multiple audio datasets from the multiple client devices via the computer network, receives multiple input datasets from the multiple client devices via the computer network, and performs training and / or inference on an AI model based on the multiple audio datasets and the multiple input datasets. The server applies the AI model to the audio datasets captured from the real-world environment to identify the material type of an object in the real-world environment.
[0010] Some of the advantages of the systems and methods described herein include providing a way to guide visually impaired persons to where to sit. For example, a visually impaired person wears glasses. The glasses emit a sound to the visually impaired person indicating the material type of the seat in the real-world environment. If a first seat is made of a hard material compared to a second seat made of a soft cushion material, the glasses inform the visually impaired person of this. The visually impaired person can then sit on the second seat. Furthermore, in the example, audio data regarding the sounds emitted from the seat when other people sit down or stand up is received by an AI model. The AI model can be trained on the audio data. The AI model is then applied to determine whether the visually impaired person is about to sit on a hard material or a soft cushion material.
[0011] Further advantages of the systems and methods described herein include providing tools for creating a metaverse that appears real to the user. For example, the systems and methods create a virtual seat surface having properties such as material type and graphic and physical parameters relating to a seat surface in a real-world environment. By using an AI model to identify the material type, the virtual seat surface is displayed photorealistically. For example, by having a physics engine apply physical phenomena relating to the material type, realistic virtual gameplay related to the interaction of virtual objects with the material type may be possible.
[0012] Other aspects of this disclosure will become apparent from the following embodiments for carrying out the invention, in conjunction with the accompanying drawings illustrating the principles of the embodiments described herein.
[0013] Various embodiments of this disclosure are best understood by referring to the following description in conjunction with the accompanying drawings. [Brief explanation of the drawing]
[0014] [Figure 1] This is a diagram illustrating an embodiment of a system that demonstrates a method for reproducing the properties of real-world objects in a virtual world. [Figure 2A] This is a diagram illustrating one of the multiple environments of the system shown in Figure 1, representing an embodiment of the system. [Figure 2B] This is a diagram illustrating an embodiment of a system that shows a method for capturing audio and image data for training an artificial intelligence (AI) model. [Figure 2C] Figure 2A is a diagram illustrating an embodiment of the system, showing a list of material types for the seat and their covers. [Figure 2D] Figure 2B is a diagram illustrating an embodiment of the system, showing a list of material types for the seat and their covers. [Figure 3A]This is a diagram illustrating an embodiment of the system, showing the training of an AI model based on multiple audio datasets, or a combination of multiple image datasets and audio datasets. [Figure 3B] This is a diagram illustrating an embodiment of a table showing the association between a set containing an audio dataset and multiple audio parameters, and a set containing the material type of the seat surface and the material type of the seat cover in the system shown in Figures 2A and 2B. [Figure 3C] This is a diagram illustrating an embodiment of a table showing the association between a set including an audio dataset, audio parameters, an image dataset, multiple physical parameters, and multiple graphic parameters, and a set including the material type of the seat surface and the material type of the seat cover in the system. [Figure 4A] This is a diagram illustrating an embodiment of a system that shows a method for determining the probability that the seat material is of a specific type and the probability that the seat cover material is of a specific type. [Figure 4B] This is a diagram of an embodiment of the system showing an augmented reality (AR) video game played using the method described herein. [Figure 5] This is a diagram illustrating an embodiment of the system, showing communication between eyeglasses and a server system via a computer network. [Figure 6] This is a diagram of an embodiment of eyeglasses showing a sound output system and microphone. [Figure 7] This is a diagram illustrating an embodiment of an input controller that provides haptic feedback. [Modes for carrying out the invention]
[0015] A system and method for identifying the material type of an object in a real-world environment are described. Note that various embodiments of this disclosure may be carried out without some or all of these specific details. In other instances, well-known process behaviors are not described in detail so as not to unnecessarily obscure the various embodiments of this disclosure.
[0016] Figure 1 is a diagram of an embodiment of system 100, illustrating a method for reproducing the characteristics of real-world objects in a virtual world. The system includes an image capture system 102, an audio capture system 104, and a server system 106. An example of the image capture system 102 is one or more cameras, such as a depth camera and a digital camera. For example, one or more cameras may be located in a head-mounted display (HMD) or AR glasses, or a display device or a game console, or they may be standalone cameras. An example of the audio capture system 104 is one or more microphones. For example, one or more microphones may be part of an HMD or AR glasses, or a display device or a game console, or they may be standalone microphones. An example of the server system 106 is multiple servers in a data center, or multiple servers of virtual machines, or a combination of processors and memory devices.
[0017] Examples of processors used herein include application-specific integrated circuits (ASICs), programmable logic devices (PLDs), central processing units (CPUs), and combinations thereof. Examples of memory devices used herein include read-only memory (ROM) devices, random access memory (RAM), and combinations thereof. For example, memory devices may be flash memory devices or redundant arrays of independent disks (RAID).
[0018] The server system 106 includes an inference training engine 108, which includes a material identification system 110. By way of example, an engine as used herein is a computer program executed by one or more of the processors of the server system 106. A computer program is an example of software. As another example, an engine as used herein includes an ASIC, or a PLD, or a combination thereof. It should be noted that ASICs, PLDs, and processors are examples of hardware. Examples of the material identification system 110 include hardware, software, or a combination thereof.
[0019] The server system 106 includes a physical phenomenon imparting system 112, a sound imparting system 114, and a graphics imparting system 116. Any of the physical phenomenon imparting system 112, the sound imparting system 114, and the graphics imparting system 116 may be exemplified by hardware, software, or a combination thereof.
[0020] The image capture system 102 and the sound capture system 104 are connected to the inference training engine 108. Next, the material identification system 110 is connected to the physical phenomenon imparting system 112, the sound imparting system 114, and the graphics imparting system 116.
[0021] The system 100 further includes a plurality of environments 118 and an environment 120. The environment 118 and the environment 120 are real-world environments. By way of example, a real-world environment exists outside of a virtual reality (VR) environment or an augmented reality (AR) environment. Illustratively, a real-world environment cannot be created by a processor.
[0022] Image data is captured from the environment 118 by the image capture system 102 and generated, for example. Audio data is also captured from the environment 118 by the sound capture system 104. The inference training engine 108 is trained according to the audio data, or a combination of audio data and image data.
[0023] The sound capture system 104 captures audio data from the environment 120. The audio data captured from the environment 120 is transmitted from the sound capture system 104 to the material identification system 110. The material identification system 110, which is trained on audio data captured from the environment 118, or on a combination of audio data captured from the environment 118 and image data, identifies one or more materials of one or more real-world objects in the environment 120, outputs one or more identities for one or more materials, and provides one or more identities to the physical phenomenon assignment system 112, the sound assignment system 114, and the graphic assignment system 116. Examples of real-world objects include chair seats, chair seat cushions, sofa seats, sofa seat cushions, chair back cushions, sofa back cushions, chair armrest cushions, sofa armrest cushions, dining tables with wooden tabletops, and dining tables with glass tabletops.
[0024] The physical phenomenon assignment system 112 assigns physical parameters to one or more virtual objects in a virtual environment displayed on a display device placed within the environment 120, such as movement according to the laws of physical phenomena, or changes in position and orientation according to the laws. Examples of virtual environments include virtual scenes such as VR scenes or AR scenes. Examples of display devices include HMDs, AR glasses, computer monitors, and televisions. The laws of physical phenomena are assigned to one or more virtual objects based on one or more types of materials in one or more real-world objects.
[0025] Furthermore, the sound assignment system 114 assigns one or more sound parameters, such as one or more combinations of amplitude and frequency, to one or more virtual objects in the virtual environment. For example, when a first virtual object is displayed in the virtual environment accompanied by the physical phenomena assigned to the first virtual object, a first sound is output, and when a second virtual object is displayed in the virtual environment accompanied by the physical phenomena assigned to the second virtual object, a second sound is output. In this example, the first sound is output by one or more speakers of a display device placed in the environment 120, and the second sound is output by one or more speakers.
[0026] Furthermore, the graphics assignment system 116 assigns one or more sets of graphics parameters, such as intensity, color, and texture, to one or more virtual objects displayed in the virtual environment. For example, when a first virtual object is displayed in the virtual environment with the physical parameters assigned to it, the first virtual object is controlled by one or more processors, such as a CPU or a graphics processing unit (GPU), or a combination thereof, of a display device located in the environment 120, so that it has a first set of graphics parameters. Furthermore, in the example, when a second virtual object is displayed in the virtual environment with the physical parameters assigned to it, the second virtual object is controlled by one or more processors of a display device located in the environment 120 so that it has a second set of graphics parameters.
[0027] In one embodiment, the inference training engine 108 is the same as the material identification system 110.
[0028] In one embodiment, system 100 excludes the image capture system 102.
[0029] In one embodiment, the terms "capture" and "generation" are used interchangeably herein.
[0030] Figure 2A is a diagram of an embodiment of System 200, showing one of the environments 118. Examples of System 200 include a room in a house, or a room in a shed, or a room in a building. System 200 is an example of one of the environments 118. System 200 includes several real-world objects such as an office chair 202, a sofa 204, a curtain behind the sofa 204, a doorway to the left of the sofa 204, a passageway to the kitchen 206, a desk 208, a computer monitor 210, a keyboard, a computer mouse, a handheld controller 212, and a game console 214. The sofa 204 has a seat 204A, on which user 2 sits. For example, the seat 204A is made from a material such as cloth, leather, corduroy, linen, or a combination thereof. The kitchen occupies another room in System 200 and has a dishwasher. A camera 216 is located above the computer monitor 210, which is an example of an image capture system 102 (Figure 1).
[0031] The handheld controller 212 is connected to the game console 214, which in turn is connected to the display device 210 and the server system 106 via a computer network. Examples of computer networks include the internet, an intranet, and combinations thereof. The camera 216 is connected to the game console 214, or to the glasses 218, or to both the glasses 218 and the game console 214. For example, the camera 216 is connected to the glasses 218 via a wireless connection such as Bluetooth®, or via a wired connection. The game console 214 accesses the game from the server system 106 and provides virtual environment data to the display device 210, which then displays the virtual scene 220. The virtual scene 220 includes multiple virtual objects, such as virtual characters, and a virtual background, such as virtual trees and virtual mountain ranges. The glasses 218 are connected to the server system 106 via a computer network. For example, the glasses 218 are connected to the server system 106 via the game console 214 and the computer network. In another example, the glasses 218 are connected directly to the server system 106 via a computer network without using the game console 214. Examples of glasses include HMDs and AR glasses.
[0032] User 1 is seated in an office chair 202 and holds a handheld controller 212 to play a game. For example, the game engine for the game is run by one or more processors in a server system 106 (Figure 1) and generates a virtual scene 220 of the game. The office chair 202 has a seat 202A, which is made from a material such as leather and covered with a plastic cover. User 1 is also wearing glasses 218, which include a microphone M1 and a camera C1. The microphone M1 is an example of a sound capture system 104 (Figure 1). The camera C1 is mounted on the bottom of the glasses 218 and positioned below the bottom. For example, when User 1 is wearing the glasses 218, the lens of camera C1 is oriented downward toward the floor of the system 200. For example, the field of view of camera C1 is directed toward the floor of the room where User 1 is. As another example, the field of view of camera C1 is wide enough to capture images of the movement of the seat 202A and 204A. User 2 is also listening to music on their smartphone. Camera C1 is an example of the image capture system 102 (Figure 1).
[0033] User 1 accesses the game from the server system 106 via a computer network and plays the game with a virtual scene 220 displayed on the display device 210. For example, User 1 selects one or more buttons on the handheld controller 212 to provide authentication information, such as a username and password. The handheld controller 212 sends the authentication information to the game console 214, which then forwards the authentication information to the server system 106 via the computer network. The server system 106 determines whether the authentication information is genuine, and if it is, provides User Account 1 and access to the game engine running on the server system 106. When the game engine is executed by one or more processors of the server system 106, game image frames are generated and encoded, and encoded image frames are output. The encoded image frames are sent to the game console 214. The game console 214 decodes the encoded image frames and provides the image frames to the display device 210 to display the game's virtual scene 220, enabling User 1 to play the game. While User 1 is playing the game, the game's virtual scene 220 is displayed on the display device 210.
[0034] User 1 activates the glasses 218 before or during gameplay. After User 1 activates the glasses 218, the microphone M1 captures an audio dataset 1a related to the seat 202A. For example, User 1 sits on the office chair 202 before playing the game, and the sound emitted by the user's sitting motion is detected by the microphone M1, capturing, outputting, or generating an audio dataset 1a. For example, the microphone M1 captures the creaking sound of the seat 202A when it is compressed, or the sound of air being blown out when the seat 202A is compressed, to generate an audio dataset 1a. For example, when User 1 sits on the seat 202A, the seat 202A is compressed. For another example, the microphone M1 captures the creaking or squeaking sound of the seat 202A when it is depressurized, to generate an audio dataset 1a. In one example, when user 1 stands up from seat 202A, seat 202A is depressurized. In another example, when user 1 jumps on seat 202A during gameplay, causing seat 202A to be compressed and depressurized multiple times, microphone M1 captures multiple creaking sounds from seat 202A. In this example, the sound is detected and audio dataset 1a is captured. In yet another example, microphone M1 captures the sound of a dishwasher operating in the kitchen along with the sound emitted by the movement of seat 202A and outputs audio dataset 1a. Audio dataset 1a is further described below with reference to Figure 3A.
[0035] Furthermore, after user 1 activates the glasses 218, microphone M1 also captures an audio dataset 1b associated with the seat 204A. For example, microphone M1 detects sounds emitted when user 2 sits on or stands up from the seat 204A. In this example, user 2 sits on or stands up from the seat 204A before or during user 1's gameplay. In this example, sound is detected and an audio dataset 1b is captured, for example, generated. In yet another example, microphone M1 captures the sound of a dishwasher operating in the kitchen along with the sounds emitted by the movement of the seat 204A and outputs an audio dataset 1b. The audio dataset 1b is further described below with reference to Figure 3A.
[0036] In the combination of camera C1 and camera 216, either camera C1 or camera 216 detects the movement of the seat surface 202A and other objects within the system 200 and captures an image dataset 1a associated with the seat surface 202A. For example, when user 1 sits on the seat surface 202A, jumps on the seat surface 202A, or stands up from the seat surface 202A, one or more of cameras C1 and cameras 216 capture image data of the seat surface 202A.
[0037] Additionally, in the combination of camera C1 and camera 216, either camera C1 or camera 216 detects the movement of the seat surface 204A and other objects within the system 200 and captures or outputs an image dataset 1b associated with the seat surface 204A. For example, during the period when user 2 is sitting on the seat surface 204A, or standing up from the seat surface 204A, or jumping on the seat surface 204A, one or more of cameras C1 and cameras 216 capture image data of the seat surface 204A.
[0038] To train an AI model, camera 216 transmits image datasets 1a and 1b to server system 106 via a computer network. For example, camera 216 transmits image datasets 1a and 1b to glasses 218 via a wireless connection, and glasses 218 forwards image datasets 1a and 1b to server system 106. In another example, camera 216 transmits image datasets 1a and 1b to server system 106 via game console 214 and a computer network.
[0039] To train an artificial intelligence (AI) model, the glasses 218 transfer image datasets 1a and 1b and audio datasets 1a and 1b to the server system 106 (Figure 1) via a computer network. For example, the glasses 218 send image datasets 1a and 1b and audio datasets 1a and 1b to the server system 106 via the game console 214 and the computer network. As another example, the glasses 218 send image datasets 1a and 1b and audio datasets 1a and 1b to the server system 106 via the computer network without using the game console 214.
[0040] In one embodiment, instead of playing a game, user 1 accesses an application from the server system 106 (Figure 1), and the application's virtual environment is displayed on the display device 210. Examples of applications include video conferencing applications, chat applications, or social networking applications.
[0041] In this embodiment, the glasses 218 do not include the camera C1, and image datasets 1a and 1b are not generated.
[0042] In the embodiments, the AI model is sometimes referred to herein as a machine learning model.
[0043] In the embodiment, the seat surface 202A is manufactured from plastic, or cloth, or vinyl, or mesh, or synthetic leather, or polyurethane, or memory foam, or other material.
[0044] In one embodiment, the backrest of the office chair 202 is manufactured from the same material as the seat 202A, or from a different material than the seat 202A.
[0045] In the embodiment, the office chair 202 includes a headrest, which is manufactured from the same material as the seat surface 202A, or from a different material than the seat surface 202A.
[0046] In one embodiment, the office chair 202 includes two armrests, each armrest being manufactured from the same material as the seat 202A, or from a different material than the seat 202A.
[0047] In this embodiment, the camera C1 is mounted at another location on the eyeglasses 218. For example, the camera C1 is fixed to the side of the eyeglasses 218, and the lens of the camera C1 is oriented downward toward the floor of the system 200.
[0048] In one embodiment, one or more additional cameras are attached to the eyeglasses 218. For example, a second camera is attached to the left edge of the eyeglasses 218. In this example, the second camera also has a lens oriented downward toward the floor of the system 200. Furthermore, in this example, a camera C1 is attached to the right edge of the eyeglasses 218.
[0049] In one embodiment, in addition to user 1, user 2 wears glasses such as AR glasses or an HMD. The glasses capture an audio dataset from sounds detected by system 200 and / or an image dataset of seat movement within system 200, and transmit the audio dataset and / or image dataset to server system 106 via a computer network. The glasses worn by user 2 are connected to server system 106 via the computer network.
[0050] In this embodiment, user 1 logs in to user account 1 using glasses 218 and then accesses user account 1. For example, glasses 218 are connected to an input controller via a wired or wireless connection. User 1 provides authentication information by selecting one or more buttons on the input controller. Upon receiving the authentication information, the input controller generates an input signal according to the authentication information and transmits the input signal to glasses 218 via a wired or wireless connection. Glasses 218 transmit the authentication information to the server system 106 via a computer network, or via both the game console 214 and the computer network. If the server system 106 determines that the authentication information is genuine, it allows user 1 to log in to user account 1.
[0051] In one embodiment, both the seat surface 202A and the seat surface 204B are made from the same material.
[0052] In this embodiment, the covers of the seat surfaces 202A and 204B are made from the same material.
[0053] In one embodiment, an audio dataset is sometimes referred to herein as an audio frame, and an image dataset is sometimes referred to herein as an image frame.
[0054] In the embodiment, a seat in the real-world environment is an example of an object or object type in the real-world environment. For example, the identity of seat 202A indicates that the object type in system 200 is seat 202A, and the identity of seat 204A indicates that the object type in system 200 is seat 204A.
[0055] Figure 2B is a diagram of an embodiment of System 250 showing a method for capturing audio and image data for training an AI model. System 250 includes a real-world environment inside a vehicle such as a bus or van. System 250 is an example of one of the environments in Environment 118 (Figure 1). System 250 includes multiple bus seats 252A, 252B, and 252C. Bus seats 252A and 252B are arranged in a row, and bus seat 252C is positioned behind bus seat 252B. User 3 sits on the seat surface 254B of bus seat 252B, and user 4 sits on the seat surface 254C of bus seat 252C. Also, user 5 is standing inside the vehicle, and user 6 sits on the seat surface 254A of bus seat 252A. For example, each of the seats 254A-254C is manufactured from a material such as wool, plastic, or cloth. As another example, each of the seat surfaces 254A to 254C is made from a material such as cloth and covered with a material such as a vinyl cover or a plastic cover. As yet another example, each of the seat surfaces 254A to 254C is made from a different material than the material that forms the seat surface 202A (Figure 2A). As yet another example, each of the seat surfaces 254A to 254C is manufactured from a different material than the material that forms the seat surface 204A (Figure 2A).
[0056] User 3 is wearing glasses 256, which include a camera C2 and a microphone M2. The glasses 256 are connected to the server system 106 via a computer network. For example, the glasses 256 are connected directly to the server system 106 via a computer network without using a game console (not shown).
[0057] Microphone M2 is an example of the sound capture system 104 (Figure 1). Camera C2 is mounted on the bottom surface of the glasses 256 and positioned below the bottom surface. For example, when user 3 is wearing the glasses 256, the lens of camera C2 is oriented downward toward the floor of system 250. For example, the field of view of camera C2 is directed toward the floor of the vehicle in which user 3 is riding. As another example, the field of view of camera C2 is wide enough to capture images of the movement of the seat surfaces 254A and 254B. System 250 includes a rear window 258 and several side windows 260A and 260B. Camera C2 is an example of the image capture system 102 (Figure 1).
[0058] Microphone M2 detects sounds emitted within the vehicle's real-world environment and outputs audio dataset 2. For example, while the vehicle is driving on a road, microphone M2 detects the sound of user 3 sitting on seat 254B and the sound of seat 254B moving up and down, and outputs audio dataset 2. As another example, M2 detects the sound of user 3 standing up from seat 254B before or while the vehicle is driving on a road and outputs audio dataset 2. As yet another example, while the vehicle is driving on a road, microphone M2 detects the sound of user 6 sitting on seat 254A and the sound of seat 254A moving up and down, and outputs audio dataset 2. As yet another example, M2 detects the sound of user 6 standing up from seat 254A before or while the vehicle is driving on a road and outputs audio dataset 2. As yet another example, microphone M2 detects the sound of user 4 sitting on the seat 254C and the sound of the seat 254C moving up and down while the vehicle is driving on the road, and outputs audio dataset 2. As yet another example, M2 detects the sound of user 4 standing up from the seat 254C before or while the vehicle is driving on the road, and outputs audio dataset 2. Audio dataset 2 is described further below.
[0059] Furthermore, the camera C2 of the glasses 256 detects the movement of one or more of the seat surfaces 254A to 254C and generates image dataset 2. For example, when user 3 sits on seat surface 254B, jumps on seat surface 254B, moves on seat surface 254B due to the movement of the vehicle, or stands up from seat surface 254B, camera C2 detects the movement of seat surface 254B and outputs image dataset 2. As another example, when user 6 sits on seat surface 254A, jumps on seat surface 254A, moves on seat surface 254A due to the movement of the vehicle, or stands up from seat surface 254A, camera C2 detects the movement of seat surface 254A and outputs image dataset 2. As yet another example, when user 4 sits on seat 254C, jumps on seat 254C, moves on seat 254C due to vehicle movement, or stands up from seat 254C, camera C2 detects the movement of seat 254C and outputs image dataset 2. In this example, camera C2 detects the movement of seat 254C when seat 254C is within the camera C2's field of view. For example, when user 3 stands up from seat 254B to talk to user 4, leans against seat 254C, and then turns towards bus seat 252C, camera C2 detects the movement of seat 254C. For training the AI model, audio dataset 2 and image dataset 2 are transmitted from glasses 256 to server system 106 via a computer network.
[0060] Please note that the following are examples of user-seat interaction: the user sitting on the seat, the user jumping on the seat, the user moving on the seat, the user touching the seat, the user using their hands to depressurize or compress the seat, and the user standing up from the seat.
[0061] In one embodiment, one or more additional cameras are attached to the eyeglasses 256. For example, a second camera is attached to the left edge of the eyeglasses 256. In this example, the second camera also has a lens oriented downward toward the floor of the system 250. Furthermore, in this example, a camera C2 is attached to the right edge of the eyeglasses 256.
[0062] In one embodiment, in addition to user 3, one or more of users 4-6 each wear one or more pairs of glasses, such as AR glasses or HMDs. Each pair of glasses captures an audio dataset from sounds detected by system 250 and / or an image dataset of seat movement within system 250, and transmits the audio dataset and / or image dataset to server system 106 via a computer network.
[0063] In this embodiment, the glasses 256 do not include the camera C2, and no image dataset 2 is generated.
[0064] In this embodiment, user 3 logs in to user account 3 using glasses 256 and then accesses user account 3. For example, glasses 256 are connected to an input controller via a wired or wireless connection. User 3 provides authentication information by selecting one or more buttons on the input controller. Upon receiving the authentication information, the input controller generates an input signal based on the authentication information and transmits the input signal to glasses 256 via a wired or wireless connection. Glasses 256 transmits the authentication information to a server system 106 via a computer network. If the server system 106 determines that the authentication information received from glasses 256 is genuine, it allows user 3 to log in to user account 3.
[0065] In one embodiment, the seat surfaces 254A to 254C are made from the same material.
[0066] In this embodiment, the seat surface 254A is made of a different material from the seat surface 254B or 254C.
[0067] In this embodiment, the covers of the seat surfaces 254A to 254C are made from the same material.
[0068] In this embodiment, the cover of the seat surface 254A is made of a different material from the cover material of the seat surface 254B or 254C.
[0069] In one embodiment, where user 3 is in a room with a game console, the glasses 256 are connected to the server system 106 via the game console and a computer network.
[0070] In an embodiment where user 3 is in a room with a game console, the glasses 256 are connected to the server system 106 via a computer network without using the game console.
[0071] Figure 2C is a diagram illustrating an embodiment of List 270 relating to the material types of the seat and their covers in System 200 (Figure 2A). The glasses 218 (Figure 2A) displays List 270 on one or more of the display screens of the glasses 218. For example, in response to User 1 logging into User Account 1, the CPU of the glasses 218 controls the GPU of the glasses 218 to display List 270 on one or more of the display screens of the glasses 218. In this example, the CPU of the glasses 218 receives List 270 from one or more processors of Server System 106 (Figure 1) via a computer network and controls the GPU of the glasses 218 to display List 270 on the glasses 218. In this example, List 270 is generated by one or more processors of Server System 106.
[0072] List 270 includes a blank space to accept the material type of seat 202A (Figure 2A) that user 1 interacts with, and another blank space to accept the material type of seat 204A (Figure 2A) that user 2 interacts with. List 270 further includes a blank space to accept the material type used to cover seat 202A, and another blank space to accept the material type used to cover seat 204A. List 270 also includes blank spaces to accept the identity of real-world objects, such as the name or type of seat 202A of chair 202, seat 204A of chair 204, and the covers for seats 202A and 204A.
[0073] User 1 logs into user account 1 and accesses list 270. User 1 selects one or more buttons on the input controller connected to the glasses 218 to provide identities such as the names of real-world objects, such as the seat 202A of chair 202, the seat 204A of chair 204, and the covers of seats 202A and 204A. User 1 also selects one or more buttons on the input controller connected to the glasses 218 to provide the material types of seats 202A and 204A, the type of material used to cover seat 202A, and the type of material used to cover seat 204A. For example, User 1 selects one or more buttons on the input controller connected to the glasses 218 to spell out the material type used for seat 202A, such as plastic, leather, or vinyl, and selects one or more buttons to spell out the type of material used for the cover of seat 202A. As another example, user 1 selects one or more buttons on an input controller connected to glasses 218 to spell out that seat 202A is the seat of an office chair and seat 204A is the seat of a sofa.
[0074] The glasses 218 receive the identities of real-world objects such as the seat 202A of chair 202 and the seat 204A of chair 204, and upon receiving the material types of the seats 202A and 204A, as well as the material types of the covers of the seats 202A and 204A, it transmits the identities of the real-world objects, the material types of the seats 202A and 204A, and the material types of the covers of the seats 202A and 204A to the server system 106 via a computer network, or via the game console 214 and the computer network, for AI model training. For example, the glasses 218 is equipped with a CPU, which receives the identity of seat 202A, which is the seat of an office chair, and assigns the alphanumeric characters 1a to the identity of seat 202A. The alphanumeric characters 1a assigned to the identity of seat 202A are sometimes referred to herein as identity I1a. As another example, the CPU of the glasses 218 receives the identity of the seat 204A, which is the seat of the sofa, and assigns an alphanumeric character such as 1b to the identity of the seat 204A. The alphanumeric character 1b assigned to the identity of the seat 204A is sometimes referred to herein as identity I1b.
[0075] In one embodiment, instead of using an input controller connected to the glasses 218 to select the material types of the seat surfaces 202A and 204A and the material types of the covers for the seat surfaces 202A and 204A, the selection is made using eye gestures. For example, the glasses 218 include an internal camera directed at the eyes of user 1. User 1 makes eye gestures to identify the seat surface 202A of chair 202, the seat surface 204A of sofa 204, the material type of seat surface 202A, and the material type of the cover for seat surface 202A, which are detected by the internal camera. For example, user 1 makes eye gestures to select the identity of the seat surface 202A of chair 202, the material type of seat surface 202A, and the material type of the cover for seat surface 202A from list 270. The internal camera captures image data including the eye gestures. The CPU of the glasses 218 receives the identities of the seats 202A and 204A selected using eye gestures, the material types of the seats 202A and 204A, and the material types of the covers of the seats 202A and 204A. For training the AI model, the CPU sends a list 270 containing the identities of the seats 202A and 204A, the material types of the seats 202A and 204A, and the material types of the covers of the seats 202A and 204A to the server system 106 via the computer network, or via both the game console 214 and the computer network.
[0076] In this embodiment, image data including eye gestures is transmitted from the glasses 218 to the server system 106 via a computer network. One or more processors in the server system 106 analyze the image data to identify the eye gestures and obtain the identity of the seat surface 202A of the chair 202, the identity of the seat surface 204A of the sofa 204, the material types of the seat surfaces 202A and 204A, and the material types of the covers of the seat surfaces 202A and 204A.
[0077] In this embodiment, a handheld controller 212 is used instead of the input controller used with the eyeglasses 218.
[0078] In one embodiment, List 270 is pre-entered with the identity of seat 202A, which is the seat of chair 202, and the identity of seat 204A, which is the seat of sofa 204. For example, when user 1 logs into user account 1 and camera C1 captures image datasets 1a and 1b, one or more processors in server system 106 generate List 270 and send List 270 to glasses 218 via the computer network. One or more processors in server system 106 identify from image datasets 1a and 1b that system 200 includes office chair 202, sofa 204, seat 202A and 204A, and covers for seat 202A and 204A. For example, one or more processors compare the pre-stored shapes of pre-stored objects, such as an office chair, or a sofa, or a seat, or a seat cover, with the shapes of images of objects such as office chair 202, sofa 204, seat 202A, seat 204A, seat cover 202A, and seat cover 204A. In this example, a comparison is made to determine if the two shapes are similar, and if they are similar, the objects are identified as pre-stored objects. An example of similarity between two shapes is when the two shapes are the same. Another example of similarity between two shapes is when a large portion of the pre-stored shape matches a large portion of the shape of the object's image.
[0079] Figure 2D is a diagram illustrating an embodiment of List 280 relating to the material types of the seat and their covers in System 250 (Figure 2B). The glasses 256 (Figure 2B) display List 280 on one or more of the glasses' display screens. For example, in response to User 3 logging into User Account 3, the CPU of the glasses 256 controls the GPU of the glasses 256 to display List 280 on one or more of the glasses 256's display screens. In this example, the CPU of the glasses 256 receives List 280 from one or more processors of Server System 106 (Figure 1) via a computer network and controls the GPU of the glasses 256 to display List 280 on the glasses 256. In this example, List 280 is generated by one or more processors of Server System 106.
[0080] List 280 includes blanks to accept the material type of seat 254B (Figure 2B) that user 3 interacts with, another blank to accept the material type of seat 254A (Figure 2B) that user 6 interacts with, and yet another blank to accept the material type of seat 254C that user 4 interacts with. List 280 further includes blanks to accept the material type used to cover seat 254B, another blank to accept the material type used to cover seat 254A, and yet another blank to accept the material type used to cover seat 254C. List 280 also includes blanks to accept identities such as the names or types of real-world objects, such as seat 254B of bath chair 252B, seat 254A of bath chair 252A, and seat 254C of bath chair 252C, as well as the covers for seats 254A-254C.
[0081] User 3 logs into user account 3 and accesses list 280. User 3 selects one or more buttons on the input controller connected to the glasses 256 to provide identities such as the names of real-world objects, such as the seat 254B of the bath chair 252B, the seat 254A of the bath chair 252A, the seat 254C of the bath chair 252C, and the covers of seats 254A-254C. User 3 also selects one or more buttons on the input controller connected to the glasses 256 to provide the material types of seats 254A-254C and the types of materials used to cover seats 254A-254C. For example, User 3 selects one or more buttons on the input controller connected to the glasses 256 to spell out the material type used for seat 254B, such as plastic, leather, or vinyl, and selects one or more buttons to spell out the type of material used for the cover of seat 254B. As another example, user 3 selects one or more buttons on an input controller connected to glasses 256 to spell out that seat 254B is the seat of a bath chair.
[0082] The glasses 256 receive the identities of real-world objects such as the seat 254B of the bath chair 252B, the seat 254A of the bath chair 252A, and the seat 254C of the bath chair 252C, and upon receiving the material types of the seat surfaces 254A-254C and the material types of the covers of the seat surfaces 254A-254C, it transmits the identities of the real-world objects, the material types of the seat surfaces 254A-254C, and the material types of the covers of the seat surfaces 254A-254C to the server system 106 via a computer network for AI model training. For example, the glasses 256 is equipped with a CPU, which receives the seat of the bath chair as the identity of seat surface 254B and assigns the alphanumeric code 2 to the identity of seat surface 254B. The alphanumeric code 2 assigned to the identity of seat surface 254B is sometimes referred to herein as Identity I2.
[0083] In one embodiment, instead of using an input controller connected to the glasses 256 to select the material type of the seat surfaces 254A-254C and the material type of the cover of the seat surfaces 254A-254C, the selection is made using eye gestures. For example, the glasses 256 include an internal camera directed at the eyes of user 3. User 3 makes eye gestures to identify the seat surface 254B of the bath chair 252B, the material type of the seat surface 254B, and the material type of the cover of the seat surface 254B, which are detected by the internal camera. For example, user 3 makes eye gestures to select the identity of the seat surface 254B of the bath chair 252B, the material type of the seat surface 254B, and the material type of the cover of the seat surface 254B from list 280. The internal camera captures image data including the eye gestures. The CPU of the glasses 256 receives the identity of the seat surfaces 254A-254C, the material type of the seat surfaces 254A-254C, and the material type of the cover of the seat surfaces 254A-254C, selected using eye gestures. The CPU sends a list 280 containing the identity of the seat surfaces 254A-254C, the material type of the seat surfaces 254A-254C, and the material type of the cover of the seat surfaces 254A-254C to the server system 106 via the computer network for training the AI model.
[0084] In this embodiment, image data including eye gestures is transmitted from the glasses 256 to the server system 106 via a computer network. One or more processors in the server system 106 analyze the image data to identify the eye gestures and obtain the identity of the seat surface 254A of the bath chair 252A, the seat surface 254B of the bath chair 252B, and the seat surface 254C of the bath chair 252C, the material type of the seat surfaces 254A to 254C, and the material type of the cover of the seat surfaces 254A to 254C.
[0085] In this embodiment, a handheld controller is used instead of the input controller used with the eyeglasses 256.
[0086] In one embodiment, List 280 is pre-entered with the identity of seat 254A being the seat of bath chair 252A, the identity of seat 254B being the seat of bath chair 252B, and the identity of seat 254C being the seat of bath chair 252C. For example, when user 3 logs into user account 3 and camera C2 captures image dataset 2, one or more processors in server system 106 generate List 280 and send List 280 to glasses 256 via the computer network. One or more processors in server system 106 identify from image dataset 2 that system 250 includes bath chairs 252A-252C, seat surfaces 254A-254C, and covers for seat surfaces 254A-254C. For example, one or more processors compare the pre-stored shapes of pre-stored objects, such as a bath chair, seat, or seat cover, with the shapes of images of objects such as bath chair 252B, seat 254B, and seat cover 254B. In this example, a comparison is made to determine if the two shapes are similar, and if they are similar, the objects are identified as pre-stored objects.
[0087] Figure 3A is a diagram of an embodiment of system 300 showing the training of an AI model based on audio datasets 1a, 1b, and 2, or a combination of image datasets 1a, 1b, and 2 and audio datasets 1a, 1b, and 2. System 300 includes client device 1, client device 2, client device 3, and server system 106. Examples of client device 1 include glasses 218, or a combination of glasses 218 and camera 216 (Figure 2A), or a combination of glasses 218 and game console 214 (Figure 2A), or a combination of glasses 218, camera 216, and game console 214. An example of client device 2 is glasses 256 (Figure 2B).
[0088] The server system 106 includes a game engine and an inference training engine 108, sometimes referred to herein as an AI processor system. The game engine is used to run the game. For example, the game engine includes game code, which enforces the laws of physics to assign physical parameters in the game, generate the state of virtual objects in the game, or generate graphic parameters for virtual objects. Game code is also executed to apply graphic parameters to one or more virtual objects in the game. The game engine is connected to the inference training engine 108.
[0089] The inference training engine 108 includes an AI processor and a memory device 302. The AI processor is an example of one of the processors of the server system 106, and the memory device 302 is an example of one of the memory devices of the server system 106. The AI processor is connected to the memory device 302. Input datasets 1a, 1b, and 2 are received from glasses 218 and glasses 256 (Figures 2A and 2B) and then stored in the memory device 302. Input dataset 1a includes image dataset 1a, or audio dataset 1a, or a combination thereof. Input dataset 1b includes image dataset 1b, or audio dataset 1b, or a combination thereof, and input dataset 2 includes image dataset 2, or audio dataset 2, or a combination thereof. The AI processor stores image datasets 1a, 1b, and 2, as well as audio datasets 1a, 1b, and 2, in the memory device 302. For example, image dataset 1b is received from glasses 218, or from a combination of camera 216 (Figure 2A) and glasses 218.
[0090] The AI processor includes a feature extractor, a classifier, and an AI model. For example, the AI processor includes a first integrated circuit that applies the functionality of the feature extractor, a second integrated circuit that applies the functionality of the classifier, and a third integrated circuit that applies the functionality of the AI model. In another example, the AI processor runs a first computer program that applies the functionality of the feature extractor, a second computer program that applies the functionality of the classifier, and a third computer program that applies the functionality of the AI model. The feature extractor is connected to the classifier, and the classifier is connected to the AI model. The AI model is an example of a material identification system 110 (Figure 1).
[0091] The feature extractor extracts audio parameters from audio datasets 1a, 1b, and 2, such as one or more amplitudes and one or more frequencies, or combinations thereof, and provides these audio parameters to the classifier. For example, the feature extractor identifies the magnitude, or peak-to-peak amplitude, or zero-to-peak amplitude, of audio datasets 1a, 1b, and 2, and the frequencies of audio datasets 1a, 1b, and 2. For example, the feature extractor identifies the magnitude of audio dataset m by identifying the absolute maximum power or absolute minimum power of audio dataset m, e.g., 1a, 1b, or 2. In this example, the absolute power is the magnitude over the entire period during which audio dataset m was generated. In another example, the feature extractor identifies the local maximum magnitude and local minimum magnitude of audio dataset m. In this example, the local magnitude is the magnitude over a given period, which is shorter than the entire period during which audio dataset m was generated. In the example, multiple local maximums and multiple local minimums are identified from the audio dataset m, and a feature extractor is used to determine the maximum and minimum sizes by applying the best fit, mean, or median to the local maximums and local minimums.
[0092] As another example, the feature extractor identifies a first time when the audio dataset m reaches a predetermined magnitude and a second time when the audio dataset m reaches the same predetermined magnitude, and calculates the difference between the first and second times to identify a time interval. The feature extractor inverts the time interval to identify the absolute frequency of the audio dataset m. In this example, the absolute frequency is the frequency over the entire period during which the audio dataset m was generated. In this example, the feature extractor identifies the absolute frequency as the frequency of the audio dataset m. As yet another example, the feature extractor identifies the local frequencies of the audio dataset m. In this example, the local frequencies are the frequencies over a predetermined period, which is shorter than the entire period during which the audio dataset m was generated. In this example, multiple local frequencies are identified from the audio dataset m, and the feature extractor applies the best fit, mean, or median to the local frequencies to identify the frequency of the audio dataset m. In this example, each local frequency is identified in the same way as the absolute frequency is identified, except that the local frequencies are identified for each predetermined period. As yet another example, the feature extractor identifies the maximum frequency in audio dataset m. The maximum frequency is the highest of all frequencies identified from audio dataset m within a given period. As yet another example, the feature extractor identifies the minimum frequency in audio dataset m. The minimum frequency is the lowest of all frequencies identified from audio dataset m within a given period.
[0093] The classifier receives audio parameters from the feature extractor and classifies the audio parameters obtained from audio datasets 1a, 1b, and 2. After classifying the audio parameters, an association is output between the audio parameters and the material types of the seat surfaces used in systems 200 and 250 and the material types of the seat covers. This association is provided from the classifier to the AI model for training. For example, if one or more processors in the server system 106 determine that the feature extractor has generated audio parameters from audio datasets 1a and 1b, they generate a list 270 (Figure 2C) and send the list 270 to the glasses 218 worn by user 1 via the computer network. For example, after audio parameters are generated from audio datasets 1a and 1b, the list 270 is sent from the server system 106 to the glasses 218 within a predetermined period. In the example, when a material type accepted in List 270 is received from the glasses 218 via the computer network, the classifier establishes an association, such as establishing a one-to-one correspondence, link, or unique relationship, between a first set of material types for seats and seat covers in System 200 and a second set of audio parameters identified from audio datasets 1a and 1b. Note, for example, if seats 202A and 204A do not have covers, List 270 excludes the material types for the covers of seats 202A and 204A. For example, the first set includes the material types for seats 202A and 204A but excludes the material types for the covers of seats 202A and 204A. Note, in another example, that System 200 excludes seat 204A. In the example, the first set includes the material type for seat 202A but excludes the material type for seat 204A. In this example, audio dataset 1a is received without receiving audio dataset 1b, and the audio parameters are identified from audio dataset 1a.Note that in all of the previous examples in this paragraph, the material types of the seat surface 202A and the cover of seat surface 202A of system 200 are examples of types 1ax and 1ay, and the material types of the seat surface 204A and the cover of seat surface 204A of system 200 are examples of types 1bx and 1by. For example, the material type of seat surface 202A is an example of type 1ax, and the material type of the cover of seat surface 202A is an example of type 1ay. As another example, the material type of seat surface 204A is an example of type 1bx, and the material type of the cover of seat surface 204A is an example of type 1by. Each of types 1ax, 1ay, 1bx, and 1by is an example of an output parameter.
[0094] As another example, one or more processors in the server system 106 generate a list 280 (Figure 2D) in response to determining that the feature extractor has generated audio parameters from the audio dataset 2, and transmit the list 280 to the glasses 256 (Figure 2B) worn by user 3 via the computer network. For example, the list 280 is transmitted within a predetermined period after the audio parameters have been generated from the audio dataset 2. In this example, when the material types accepted in the list 280 are received from the glasses 256 via the computer network, the classifier establishes an association, such as a one-to-one correspondence, link, or unique relationship, between a third set of material types for seats and seat covers in system 250 and a fourth set of audio parameters identified from the audio dataset 2. Note, for example, that if seats 254A, 254B, and 254C do not have covers, the list 280 will exclude the material types for the covers of seats 254A, 254B, and 254C. In the example, the third set includes the material types for seat surfaces 254A, 254B, and 254C, but does not include the material types for the covers of seat surfaces 254A, 254B, and 254C. Note that in any of the previous examples in this paragraph, the material types for seat surfaces 254A, 254B, and 254C of system 250, and the material types for the covers of seat surfaces 254A, 254B, and 254C, are examples of types 2x and 2y. For example, the material type for seat surface 254A, 254B, or 254C is an example of type 2x, and the material type for the cover of seat surface 254A, 254B, or 254C is an example of type 2y. Each type 2x and 2y is an example of an output parameter.
[0095] The AI model is trained based on the associations between input datasets 1a, 1b, and 2, such as audio parameters and image datasets 1a, 1b, and 2 related to systems 200 and 250 (Figures 2A and 2B). For example, the AI model is provided with an association between a first set of audio parameters generated from audio datasets 1a and 1b and a second set of material types for the seats 202A and 204A and their covers. As another example, the AI model is provided with an association between a first set of audio parameters generated from audio datasets 1a and 1b and a second set of material types for the seats 202A and 204A. As yet another example, the AI model is provided with an association between a first set of audio parameters generated from audio dataset 1a and a second set of material types for the seats 202A and 204A. As yet another example, the AI model is provided with an association between a first set of audio parameters generated from audio dataset 1a and a second set of material types for the seats 202A and 204A.
[0096] As another example, the AI model is provided with an association between a third set of audio parameters generated from audio dataset 2 and a fourth set of material types for the seat surfaces 254A-254C and the covers of seat surfaces 254A-254C.
[0097] In one embodiment, image datasets 1a, 1b, and 2 are received by a server system 106 from client devices 1 and 2 via a computer network. In the embodiment, image datasets 1a, 1b, and 2 are used in conjunction with audio datasets 1a, 1b, and 2 to facilitate the training of an AI model. For example, a feature classifier identifies real-world objects in systems 200 and 250 from image datasets 1a, 1b, and 2, and further identifies the physical and graphic parameters of the real-world objects. For example, the feature classifier identifies or identifies that the first seat is the seat of an office chair by comparing the size and shape of an image of a first seat, such as a seat 202A or 204A, received in a first image dataset, such as image dataset 1a or 1b, with pre-stored size and shape of a pre-stored seat. In the example, the pre-stored size, pre-stored shape, and pre-stored seat identification are stored in a memory device 302. In the example, the pre-stored seat identification includes alphanumeric characters. In the example, the feature classifier extracts or identifies physical parameters from the first image dataset, such as movement on the first seating surface, i.e., changes in position and orientation. In the example, the feature classifier identifies changes in movement on the first seating surface from a first position to a second position, and changes in movement from a first orientation to a second orientation, or a combination thereof. In the example, changes in the movement of the first seating surface occur when the first seating surface is compressed or decompressed due to interaction between the first user, such as user 1 or 2, and the first seating surface. Furthermore, in the example, as the position and orientation of the first seating surface change, the feature classifier extracts or identifies graphic parameters from the first image dataset, such as the intensity, color, shading, or texture of the first seating surface. In the example, to facilitate the training of the AI model, the feature extractor provides the classifier with the identity of the first seating surface in system 200, the physical parameters of the first seating surface, and the graphic parameters of the first seating surface.In the example, the classifier establishes an association, such as a one-to-one relationship, a unique relationship, or a link, between a set containing the identity of the first seat, audio parameters identified based on the sound generated by the interaction between the first user and the first seat, physical parameters of the first seat, and graphic parameters of the first seat, and a set containing the material type of the first seat and the material type of the cover of the first seat, and outputs the association. In the example, the association is provided from the classifier to the AI model for training the AI model.
[0098] Note that the physical parameters of seat surface 202A from the first image dataset are sometimes referred to herein as physical parameter PP1a, and the physical parameters of seat surface 204A from the first image dataset are sometimes referred to herein as physical parameter PP2a. Also note that the graphic parameters of seat surface 202A from the first image dataset are sometimes referred to herein as graphic parameter GP1a, and the graphic parameters of seat surface 204A from the first image dataset are sometimes referred to herein as graphic parameter GP2a. For example, graphic parameter GP1a includes a first set of graphic parameters when seat surface 202A is in a first position and orientation, and a second set of graphic parameters when seat surface 202A is in a second position and orientation. For example, the first set includes different amounts, intensities, shading, colors, textures, or combinations thereof compared to the amounts, intensities, shading, colors, textures, or combinations thereof of the second set.
[0099] As another example, the feature classifier identifies the second seat as a bus chair seat by comparing the size and shape of an image of a second seat, such as seat 254A, 254B, or 254C, received in a second image dataset, such as image dataset 2, with the size and shape of a pre-stored seat. In this example, the pre-stored size, shape, and seat identification are stored in memory device 302. In this example, the pre-stored seat identification includes alphanumeric characters. Also in this example, the feature classifier extracts or identifies physical parameters such as movement on the second seat, i.e., changes in position and orientation, from the second image dataset. In this example, the feature classifier identifies movement on the second seat from a third position to a fourth position, and movement from a third orientation to a fourth orientation. In the example, changes in the movement of the second seat surface occur when the second seat surface is compressed or decompressed due to interaction between the second user (such as user 3, 4, or 6) and the second seat surface. Furthermore, in the example, when the position and orientation of the second seat surface change due to interaction by the second user, the feature classifier extracts or identifies graphic parameters such as the intensity, color, shading, or texture of the second seat surface from the second image dataset. In the example, for training the AI model, the feature extractor provides the AI model with the identity of the second seat surface in system 250, the physical parameters of the second seat surface, and the graphic parameters of the second seat surface. In the example, the classifier establishes an association, such as a one-to-one relationship, a unique relationship, or a link, between a set containing the identity of the second seat, audio parameters identified based on the sound generated by the interaction between the second user and the second seat, physical parameters of the second seat, and graphic parameters of the second seat, and a set containing the material type of the second seat and the material type of the cover of the second seat, and outputs the association. In the example, the association is provided from the classifier to the AI model for training the AI model.
[0100] Note that the physical parameters of seat surface 254A, 254B, or 254C from the second image dataset are sometimes referred to herein as physical parameters PP2. Also note that the physical parameters of seat surface 254A, 254B, or 254C from the second image dataset are sometimes referred to herein as graphic parameters GP2. Furthermore, note that the identity of seat surface 254A, 254B, or 254C is sometimes referred to herein as identity I2.
[0101] In some embodiments, communication between the server system 106 and client devices 1-3 can be facilitated using wireless technology. Such technology may include, for example, 5G wireless communication technology. 5G is the fifth generation of cellular network technology. A 5G network is a digital cellular network, in which the service area targeted by a provider is divided into small geographical areas called cells. Analog signals representing sound and images are digitized at the client device, converted by an analog-to-digital converter, and transmitted to the cell as a bitstream. All 5G wireless devices within a cell communicate radio waves with local antenna arrays and low-power automatic transceivers (transmitters and receivers) within the cell, and this communication takes place via frequency channels assigned by the transceivers from a pool of frequencies reused by other cells. The local antenna arrays are connected to the cellular network and the internet by high-bandwidth optical fiber or wireless backhaul connections. As with other cell networks, mobile devices moving from one cell to another are automatically transferred to the new cell. It should be understood that 5G networks are merely illustrative types of communication networks, and embodiments of this disclosure may also utilize previous generations of wireless or wired communications, as well as subsequent generations of wired or wireless technologies following 5G.
[0102] In one embodiment, to train the AI model, one or more of the processors 1-P (Figure 5) identify errors between the AI model's output, such as material type, and the actual output parameters, such as material type. One or more of the processors 1-P backpropagate the errors to adjust the AI model. In this embodiment, the output parameters are not used as input to the AI model, but rather as input to identify errors in the AI model. Also in this embodiment, the errors are not used as input to the AI model. Rather, the errors are used by one or more of the processors 1-P to make corrections, such as changing the weights of the AI model.
[0103] Figure 3B is a diagram illustrating an embodiment of Table 350 showing associations 352a, 352b, and 354. Associations 352a, 352b, and 354 are formed between a set including audio datasets 1a, 1b, and 2, and a plurality of audio parameters AP1a, AP1b, and AP2, and a set including identities I1a, I1b, and I2 in systems 200 and 250 (Figures 2A and 2B), material types 1ax, 1bx, and 2x for seat surfaces 202A, 204A, 254A, 254B, and 254C, and material types 1ay, 1by, and 2y for the covers of seat surfaces 202A, 204A, 254A, 254B, and 254C. Audio parameter AP1a represents an audio parameter identified from audio dataset 1a by a feature classifier. Audio dataset 1a is generated based on sounds received from the interaction between user 1 and seat 202A (Figure 2A). Audio parameter AP1b represents the audio parameters identified from audio dataset 1b by the feature classifier. Audio dataset 1b is captured based on sounds received from the interaction between user 2 and seat 204A (Figure 2A). Audio parameter AP2 represents the audio parameters identified from audio dataset 2 by the feature classifier. Audio dataset 2 is generated based on sounds received from the interaction between user 3 and seat 254B (Figure 2B), user 6 and seat 254A (Figure 2B), user 4 and seat 254C (Figure 2B), or a combination thereof.
[0104] Material type 1ax is the material type of seat 202A, and material type 1ay is the material type of the cover of seat 202A. Types 1ax and 1ay are received from user 1 via user account 1 as part of input dataset 1a (Figure 3A). For example, material types 1ax and 1ay are received from user 1 in list 270 (Figure 2C). Similarly, material type 1bx is the material type of seat 204A, and material type 1by is the material type of the cover of seat 204A. Types 1bx and 1by are received from user 1 via user account 1 as part of input dataset 1b (Figure 3A). For example, material types 1bx and 1by are received from user 1 in list 270. Also, material type 2x is the material type of one of seat 254A, 254B, and 254C, and material type 2y is the material type of the cover of one of seat 254A, 254B, and 254C. Material types 2x and 2y are received from user 3 via user account 3 as part of input dataset 1b (Figure 3A). For example, material types 2x and 2y are received from user 3 via user account 3 in list 280 (Figure 2D).
[0105] Association 352a is a unique relationship between the set containing audio dataset 1a and audio parameter AP1a and the set containing identity I1a and material types 1ax and 1ay. Similarly, association 352b is a unique relationship between the set containing audio dataset 1b and audio parameter AP1b and the set containing identity I1b and material types 1bx and 1by. Furthermore, association 354 is a one-to-one relationship between the set containing audio dataset 2 and audio parameter AP2 and the set containing identity I2 and material types 2x and 2y.
[0106] In one embodiment, Table 350 excludes material types 1ay, 1by, and 2y for the seat covers of systems 200 and 250. This occurs when there is no cover on the seat of systems 200 and 250, or when material types 1ay, 1by, and 2y are not received from users 1 and 3 via their respective user accounts 1 and 3.
[0107] Figure 3C is a diagram of an embodiment of Table 360 showing associations 362a, 362b, and 364. Associations 362a, 362b, and 364 are formed between a set including audio datasets 1a, 1b, and 2, multiple audio parameters AP1a, AP1b, and AP2, image datasets 1a, 1b, and 2, physical parameters PP1a, PP1b, and PP2, and graphic parameters GP1a, GP1b, and GP2, and a set including identities I1a, I1b, and I2 in systems 200 and 250 (Figures 2A and 2B), material types 1ax, 1bx, and 2x for seat surfaces 202A, 204A, 254A, 254B, and 254C, and material types 1ay, 1by, and 2y for seat surface covers.
[0108] Association 362a is a unique relationship between a set containing audio dataset 1a, audio parameter AP1a, physical parameter PP1a, and graphic parameter GP1a, and a set containing identity I1a and material types 1ax and 1ay. Similarly, association 362b is a unique relationship between a set containing audio dataset 1b, audio parameter AP1b, physical parameter PP1b, and graphic parameter GP1b, and a set containing identity I1b and material types 1bx and 1by. Furthermore, association 364 is a one-to-one relationship between a set containing audio dataset 2, audio parameter AP2, physical parameter PP2, and graphic parameter GP2, and a set containing identity I2 and material types 2x and 2y.
[0109] In one embodiment, Table 360 excludes material types 1ay, 1by, and 2y for the covers of seat surfaces 202A, 204A, 254A, 254B, and 254C of systems 200 and 250. This occurs when there are no covers for the seat surfaces of systems 200 and 250, or when material types 1ay, 1by, and 2y are not received from users 1 and 3 via their respective user accounts 1 and 3.
[0110] Figure 4A is a diagram of an embodiment of system 400 showing a method for determining the probability N% that the material of the seat surface 402A is of type 1ax and the probability M% that the material of the cover of the seat surface 402A is of type 1ay (M and N are real numbers). Types 1ax and 1ay are examples of output parameters. System 400 is an example of environment 120 (Figure 1). System 400 includes a chair 402 with a seat surface 402A, a user 7, an AI model, a table 404, a virtual chair 408 with a virtual seat surface 408A, and a virtual user 416. User 7 is wearing glasses 410 equipped with a microphone M3 and a camera C3. Glasses 410 is an example of client device 3 (Figure 3A).
[0111] The glasses 410 are connected to the server system 106 via a computer network. For example, the glasses 410 are connected to the server system 106 (Figure 1) via a game console (not shown) and a computer network. In another example, the glasses 410 are connected directly to the server system 106 via a computer network without using a game console (not shown).
[0112] Microphone M3 is an example of the sound capture system 104 (Figure 1). Camera C3 is an example of the image capture system 104 (Figure 1). Camera C3 is mounted on the bottom surface of the rim of the glasses 410. For example, the field of view of camera C3 is oriented downward toward the floor of the real-world environment, such as the room in which the glasses 410 are located. The glasses 410 are connected to the input controller 414 via a wired or wireless connection.
[0113] Microphone M3 detects sounds emitted from the real-world environment where user 7 is located. For example, when user 7 stands up from or sits down on seat 402A, noise is generated, which is detected by microphone M3 and an audio dataset p is captured (where p is an integer). The audio dataset p is transmitted from glasses 410 to server system 106 (Figure 1) via the computer network. Feature extractors in one or more processors of server system 106 extract audio parameters APp from audio dataset p. For example, the feature extractor extracts audio parameters APp from audio dataset p in the same way that audio parameters AP1a are obtained from audio dataset 1a, or audio parameters AP1b are obtained from audio dataset 1b.
[0114] One or more processors of the server system 106 store the audio parameter APp and the audio dataset p in one or more memory devices of the server system 106. The audio parameter APp is provided to the AI model from the feature extractor. Upon receiving the audio parameter APp, the AI model identifies a probability N% that the material of the seat surface 402A is of type 1ax and a probability M% that the material of the cover of the seat surface 402A is of type 1ay. For example, if the AI model identifies that the audio parameter APp is within a predetermined range from the audio parameter AP1a and outside a predetermined range from the audio parameter AP1b or AP2, it indicates that there is a greater than 50% probability that the audio dataset p was generated based on sound reflected from the seat surface 402A, which is made of the same material as the seat surface 202A (Figure 2A), and that there is a greater than 50% probability that the audio dataset p was generated based on sound reflected from the cover of the seat surface 402A, which is made of the same material as the cover of the seat surface 202A (Figure 2A). In this example, the probability that seat surface 402A is made of the same material as seat surface 202A (Figure 2A), and the probability that the cover of seat surface 402A is made of the same material as the cover of seat surface 202A, are examples of model output 412. Model output 412 is the output of the AI model.
[0115] For example, if the AI model identifies that the maximum amplitude of audio parameter APp is within a predetermined range from the maximum amplitude of audio parameter AP1a, but outside the predetermined range from the maximum amplitude of audio parameter AP1b, or outside the predetermined range from the maximum amplitude of audio parameter AP2, it will indicate that there is a greater than 50% probability that seat surface 402A is made of the same material as seat surface 202A, and that there is a greater than 50% probability that the cover of seat surface 402A is made of the same material as the cover of seat surface 202A. As another example, if the AI model identifies that the frequency of audio parameter APp is within a predetermined range from the frequency of audio parameter AP1a, but outside the predetermined range from the frequency of audio parameter AP1b, or outside the predetermined range from the frequency of audio parameter AP2, it will indicate that there is a greater than 50% probability that seat surface 402A is made of the same material as seat surface 202A, and that there is a greater than 50% probability that the cover of seat surface 402A is made of the same material as the cover of seat surface 202A. As another example, the AI model identifies that the maximum amplitude of audio parameter APp is within a first predetermined range from the maximum amplitude of audio parameter AP1a and outside the first predetermined range from the maximum amplitude of audio parameter AP1b, or outside the first predetermined range from the maximum amplitude of audio parameter AP2, and that the frequency of audio parameter APp is within a second predetermined range from the frequency of audio parameter AP1a and outside the second predetermined range from the frequency of audio parameter AP1b, or outside the second predetermined range from the frequency of audio parameter AP2, and then indicates that there is a greater than 50% probability that the seat surface 402A is made of the same material as seat surface 202A, and that there is a greater than 50% probability that the cover of seat surface 402A is made of the same material as the cover of seat surface 202A.
[0116] Furthermore, in this example, user 7 selects one or more buttons on the input controller 414 to instruct the glasses 410 to display a virtual image of a virtual seat 408A similar to the seat 402A, and a virtual image of a virtual user 416. For example, the virtual user 416 is an in-game character or avatar controlled by user 7 during gameplay. For example, the virtual seat 408A is similar to the seat 402A when it has the same material, or the same physical parameters, the same audio parameters, or the same graphic parameters, or a combination thereof.
[0117] In the example, upon receiving an instruction, the glasses 410 transmit the instruction to one or more processors of the server system 106 via the computer network. In the example, upon receiving an instruction from the glasses 410 to display a virtual image of a virtual seat surface 408A similar to the seat surface 402A, the AI model in one or more processors of the server system 106 calculates or identifies, or has already calculated, or has already identified, a model output 412 indicating that there is an N% probability that the seat surface 402A is made of the same material as the seat surface 202A, and an M% probability that the cover of the seat surface 402A is made of the same material as the cover of the seat surface 202A.
[0118] Furthermore, in the example, one or more processors of the server system 106 identify model output 412 and, after receiving an instruction from the glasses 410 to display a virtual image of a virtual seat surface 408A similar to seat surface 402A or 202A, decide to output the audio parameters AP1a of a specific underlying seat surface 202A with probabilities M% and N%. Furthermore, in the example, one or more processors of the server system 106 decide to display the graphic parameters GP1a and physical parameters PP1a of seat surface 202A, and decide to display the graphic parameters GP1a according to the physical parameters PP1a. For example, when physical parameters PP1a indicate that the virtual seat surface 408A should be displayed in a first position and orientation, a first set of graphic parameters is assigned to the virtual seat surface 408A, and when physical parameters PP1a indicate that the virtual seat surface 408A should be displayed in a second position and orientation, a second set of graphic parameters is assigned to the virtual seat surface 408A.
[0119] In the example, one or more processors in the server system generate virtual image data to display an image of virtual user 416 upon receiving an instruction. In the example, one or more processors in the server system determine how the virtual seat 408A moves based on the movements of virtual user 416. For example, when virtual user 416 is shown sitting on seat 408A, the virtual seat 408A is shown being compressed according to the same physical parameters PP1a, the same audio parameters AP1a, and the same graphics parameters GP1a as when user 1 sits on seat 202A. As another example, when virtual user 416 is shown standing up from seat 408A, the virtual seat 408A is shown being depressurized according to the same physical parameters PP1a, the same audio parameters AP1a, and the same graphics parameters GP1a as when user 1 stands up from seat 202A. In this example, one or more processors of the server system 106 access audio parameter AP1a, graphics parameter GP1a, physical parameter PP1a, and virtual image data for displaying an image of the virtual user 416 from one or more memory devices of the server system 106, and transmit audio parameter AP1a, graphics parameter GP1a, physical parameter PP1a, and virtual image data to the glasses 410 via the computer network.
[0120] Furthermore, in the example, the CPU of the glasses 410 receives audio parameter AP1a, graphics parameter GP1a, physical parameter PP1a, and virtual image data for displaying an image of the virtual user 416, and controls the GPU of the glasses 410 to display a virtual chair 408 having a virtual seat surface 408A in which the graphics parameter GP1a changes according to the change in the physical parameter PP1a. Also in the example, the CPU of the glasses controls the GPU of the glasses to display the virtual user 416 sitting on or standing up from the virtual seat surface 408A. Also in the example, the CPU of the glasses 410 controls the sound output system of the glasses 410 to output sound according to the audio parameter AP1a. For example, the CPU of the glasses 410 controls the GPU of the glasses 410 to display the virtual seat surface 408A being compressed or decompressed based on the physical parameter PP1a. For example, the GPU of the glasses 410 displays a virtual seat surface 408A having the intensity, color, and texture indicated by the graphics parameter GP1a. Furthermore, in the example, when the virtual seat surface 408A is shown being compressed and depressurized according to the physical parameter PP1a, the CPU controls the sound output system of the glasses 410 to output sounds of compression and depressurization.
[0121] In one embodiment, each input dataset described herein includes light-detection ranging (LiDAR) data. For example, in system 200, LiDAR data 1a is captured using a LiDAR scanner and transmitted to the server system 106 in the same manner that image data 1a is transmitted from camera C1 to server system 106. For example, the LiDAR scanner is implemented in glasses 218, and LiDAR data 1a includes a LiDAR image of seat surface 202A. In system 250, LiDAR data 1b is captured using a LiDAR scanner, and in system 400, LiDAR data 2 is captured using a LiDAR scanner. For example, as shown in Figure 2B, the LiDAR scanner is implemented in glasses 256, and LiDAR data 1b includes a LiDAR image of one or more seat surfaces 254A, 254B, and 254C. Furthermore, as shown in Figure 4A, the LiDAR scanner is implemented within the glasses 410, and LiDAR data 2 includes a LiDAR image of the seat surface 402A. LiDAR data 1a and 1b are used in the same manner as image data 1a and 1b are used to train the AI model. The AI model then processes LiDAR data 2 to identify probabilities N% and M%.
[0122] In embodiments, each input dataset described herein includes inertial measurement unit (IMU) data. For example, each pair of glasses described herein includes inertial sensors such as a magnetometer, gyroscope, and accelerometer, which detect the movement of the user's head while wearing the glasses. For example, the movement includes the position and orientation of the user's head. Glasses 218 (Figure 2A) detect the movement of user 1's head while user 1 walks on the floor material of system 200 and output an IMU dataset 1a. Furthermore, microphone M1 captures an audio dataset 1a' representing the sound produced by the floor material of system 200 as user 1 walks on it. Similarly, the glasses 256 (Figure 2B) detect the head movement of user 3 while user 3 walks on the floor material of system 250 and output IMU dataset 1b, and the microphone M2 captures audio dataset 1b' showing the sound generated by the floor material of system 250 when user 3 walks on the floor material of system 250. In the same way that input datasets 1a and 1b are used to train the AI model, audio datasets 1a' and 1b', as well as IMU datasets 1a and 1b, are used to train the AI model. In addition, the microphone M3 captures audio dataset 2' while user 7 walks on the floor material of system 400, and the inertial sensor detects the head movement of user 7 while user 7 walks and outputs IMU dataset 2. Audio dataset 2' shows the sound generated by the floor material of system 400 when user 7 walks on the floor of system 400. Audio dataset 2' and IMU dataset 2 are processed by the AI model in the same way that input dataset 2 is processed to identify a probability N%, so that they output the probability of the floor material type for system 400.
[0123] In one embodiment, the AI model identifies a model output 412 before receiving an instruction to display a virtual image of a seat surface 408A similar to the seat surface 402A.
[0124] In one embodiment, the glasses 410 exclude the camera C3.
[0125] In one embodiment, a handheld controller is used instead of the input controller 414. The handheld controller is connected to the glasses 410 via a wired or wireless connection.
[0126] In one embodiment, instead of receiving a selection from one or more buttons on an input controller connected to glasses worn by the user, the user makes a gesture. The glasses' camera captures image data of the gesture. For example, the gesture is used to select the material type of a real-world object in the system, or the material type of the cover of a real-world object. The image data is analyzed by the glasses' CPU or one or more processors of the server system 106 to identify the material type of the real-world object or cover selected by the user.
[0127] Figure 4B is a diagram of an embodiment of System 450 showing an augmented reality (AR) video game using the method described herein. When a handheld controller, including a microphone, is placed on a coffee table, the handheld controller helps provide material estimation for the coffee table. A virtual character then pops out from the handheld controller onto the coffee table and virtually bumps into the coffee table, producing realistic three-dimensional (3D) audio sounds. In this example, the sonic characteristics of the 3D audio produced by the virtual object interacting with the coffee table can be dynamically adjusted using the estimated material, such as wood or glass. The sound produced by the virtual character interacting with the coffee table is modified using the actual audio data of the physical interaction of the handheld controller bumping into the coffee table, so that user 7 (Figure 4A) feels as if the virtual character is actually on the coffee table.
[0128] System 450 includes a view of a living room from the perspective of user 7 wearing glasses 410 (Figure 4A). The living room is an example of System 450 and a real-world environment. The living room includes a coffee table 452, a real video game controller 454, a virtual robot 456, and a display device 458 within user 7's view through the glasses 410. The virtual robot 456 is a video game character. For example, the video game controller 454, or a combination of the video game controller 454 and the display device 458, is an example of client device 3 (Figure 3A). Also, for example, the coffee table 452 has a table surface made of wood, or glass, or another material.
[0129] User 7 uses the video game controller 454 to log in to user account 7, which has been assigned to user 7 by the server system 106 (Figure 3A). Once user 7 is logged in to user account 7, the glasses 410 access games, such as AR video games, from the server system 106. During gameplay, when user 7 places the video game controller 454 on the coffee table 452, a sound is produced, which is represented as an exemplary sound wave 460. For example, the sound wave 460 is produced by the interaction between the video game controller 454 and the coffee table 452. The sound wave 460 is captured by microphone M3, or a combination of microphone M3 and the microphone on the video game controller 454, and an audio dataset q (where q is an integer) is generated. The audio dataset q is transmitted from microphone M3, or the microphone on the video game controller 454, or a combination thereof, to the server system 106 via the internet. The server system 106 processes the audio dataset q in the same way that the audio dataset p in Figure 4A is processed, determining the probability r% that the material type of the coffee table 452 is wood or glass (where r is a positive real number). The audio dataset q is then stored in one or more memory devices of the server system 106 by one or more processors of the server system 106.
[0130] Furthermore, during gameplay, the camera C3 of the glasses 410 captures images of the real-world environment of system 450 and transmits the images to the server system 106 via the internet. The server system 106 determines the positions of objects in the real-world environment of system 450, such as the coffee table 452 and the video game controller 454, from reference points such as the glasses 410 within system 450. In addition, the server system 106 determines the distance between the glasses 410 and any point on the table surface of the coffee table 452 from the images of the real-world environment of system 450.
[0131] During gameplay, the audio dataset q is processed by the server system 106, and after a probability r% is determined, user 7 selects one or more buttons on the video game controller 454. Upon receiving the selection, the video game controller 454 generates an input signal, which is sent from the video game controller 454 to the server system 106 via the internet. Upon receiving the input signal, the server system 106 executes the game code to generate data to display the virtual robot 456 jumping out of the video game controller 454 and landing on the coffee table 452, based on the position of the coffee table 452 and the position of the video game controller 454. Additionally, one or more processors in the server system 106 generate audio data to be output as sound when the virtual robot 456 lands on the coffee table 452. For example, the audio dataset q is accessed from one or more memory devices in the server system 106 to mimic the sound produced when user 7 places the video game controller 454 on the coffee table 452. In the example, an audio dataset q, acquired from the glasses 410 and generated based on the sound when user 7 places the video game controller 454 on the coffee table 452, is accessed from one or more memory devices of the server system 106. Furthermore, the example uses the audio dataset q to generate an audio dataset that is output as sound when the virtual robot 456 lands on the coffee table 452. Also in the example, one or more processors of the server system 106 adjust the audio parameters of the audio dataset q accessed from one or more memory devices of the server system 106 based on the distance between the glasses 410 and the virtual robot 456 displayed on the glasses 410. For example, the further the position where the virtual robot 456 is displayed on the table surface of the coffee table 452 is from the glasses 410, the more one or more processors proportionally reduce the peak-to-peak amplitude of the audio dataset q.Conversely, in the example, the closer the position of the virtual robot 456 on the table surface of the coffee table 452 is to the glasses 410, the more one or more processors proportionally increase the peak-to-peak amplitude of the audio dataset q. The virtual robot 456 is a character controlled by user 7 using a video game controller 454. One or more processors of the server system 106 transmit data for displaying the virtual robot 456 landing on the coffee table 452, and audio data to be output as sound along with the display of the virtual robot 456, to the glasses 410 via the internet.
[0132] Upon receiving data to display the virtual robot, the GPU of the glasses 410 displays the virtual robot 456 on the table surface of the coffee table 452 in the living room view presented by the glasses 410. For example, the GPU of the glasses 410 displays the virtual robot 456 landing on the table surface of the coffee table 452. The CPU of the glasses 410 also receives audio data from the server system 106 via the internet. When the virtual robot 456 lands on the coffee table 452, the CPU of the glasses 410 controls the sound output system of the glasses 410 to generate a sound, represented as a sound wave 462, which is output from the speaker of the glasses 410. The sound is output based on the audio data received from the server system 106. Thus, a simulation is generated in the glasses 410 showing that a sound wave 462 is produced when the virtual robot 456 lands on the coffee table 452. It should be understood that the current technology described herein provides a way for the virtual robot 456 to respond to the user 7 in a perceptually natural manner, such as when the virtual robot 456 interacts with the coffee table 452. Examples of interaction include outputting realistically suitable 3D sounds from the sound output system of the glasses 410 so that the sounds appear to be emanating from the same physical location on the coffee table 452 as the visual representation of the virtual robot 456 presented by the glasses 410, and further adjusting the virtual physical phenomenon attributes applied to the virtual robot 456 so that the landing motion and animation of the virtual robot 456 appear realistic and natural, and are significantly different from the landing motion and animation in which the virtual robot 456 might land and bounce on different physical surfaces such as the seat 204A or 202A shown in Figure 2A. Thus, it should be understood that the systems and methods described herein can be used to modify video game applications, as shown in Figure 4B, to provide more realistic and immersive AR entertainment.
[0133] In one embodiment, when it is determined that the virtual robot 456 will interact with the coffee table 452, haptic feedback is provided to the user 7 via the video game controller 454. For example, when it is determined that the virtual robot 456 will interact with the coffee table 452, one or more of the processors 1-P (Figure 5) generate haptic feedback data and transmit it via the computer network 504 (Figure 5) to a haptic feedback device, such as an eccentric rotating mass (ERM) actuator or a piezoelectric actuator, located within the video game controller 454. Furthermore, one or more processors 1-P transmit instructions to the video game controller 454 via the computer network 504 to output haptic feedback for the duration that the virtual robot 456 is shown interacting with the coffee table 452, such as landing on or moving to the coffee table 452. Upon receiving the haptic feedback data, the haptic feedback device vibrates to provide haptic feedback about the interaction to the user 7 holding the video game controller 454. The video game controller 454 or input controller 414 is an example of a body-connected device.
[0134] In this embodiment, haptic feedback is provided to the user 7 via the glasses 410 or another body-connected device. For example, the glasses 410 include a haptic feedback device. In this example, one or more processors 1-P generate haptic feedback data and transmit it via the computer network 504 to the haptic feedback device of the glasses 410 or another body-connected device. Furthermore, in this example, one or more processors 1-P transmit instructions via the computer network 504 to the glasses 410 or another body-connected device to output haptic feedback during the period in which a virtual robot 456 interacting with the coffee table 452 is displayed. Upon receiving the haptic feedback data, the haptic feedback device of the glasses 410 or another body-connected device vibrates to provide the user 7 with haptic feedback regarding the interaction.
[0135] Figure 5 is a diagram of an embodiment of system 500 showing communication between the glasses 502 and the server system 106 via a computer network 504. System 500 includes the glasses 502, the computer network 504, the server system 106, and an input controller 501. Examples of input controllers used herein include handheld controllers such as game controllers. For example, a game controller is an input device used in video games. In this example, the user makes one or more selections on the game controller to control an object or character in the video game. Examples of glasses 502 include glasses 218 (Figure 2A), glasses 256 (Figure 2B), or glasses 410 (Figure 4A). Examples of glasses 502 include AR glasses or HMDs. Examples of computer network 504 include wide area networks (WANs) such as the Internet, local area networks (LANs) such as intranets, and combinations thereof.
[0136] The glasses 502 include a CPU 504, a GPU 506, a display screen 508, a camera 510, a video encoder 512, a network transfer device 514, an audio encoder 516, a microphone 518, a communication device 507, a video decoder 528, and an audio decoder 530. Examples of the CPU 504 include a processor, an ASIC, and a PLD. Examples of the GPU 506 include a processor, an ASIC, and a PLD. Examples of the display screen 508 include a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. The camera 510 is an example of any of cameras C1, C2, and C3. An example of a video encoder used herein is a circuit that applies a video conversion protocol, such as a video encoding protocol, to an image dataset to output an encoded image dataset, such as an encoded image frame. For example, the video encoder 512 generates an I-frame, a P-frame, or a B-frame, which are examples of encoded image frames. Examples of video encoding protocols include H.262, H.263, and H.264. Microphone 518 is a device that converts sound energy into an audio dataset.
[0137] Examples of audio encoders used herein include circuits that compress an audio dataset into encoded audio datasets, such as encoded audio frames. For example, an audio encoder encodes an audio dataset and generates encoded audio frames by applying an audio encoding protocol, such as lossless or lossy compression. An example of lossy compression is the modified discrete cosine transform (MDCT), which converts a time-domain sampled waveform to the frequency domain. Another example of lossy compression is the linear predictive coding (LPC) protocol, which analyzes an audio dataset.
[0138] Examples of network transmission devices used herein include network interface controllers such as network interface cards (NICs). Another example of a network transmission device is a wireless access card (WAC). Microphone 518 is an example of one of microphones M1, M2, and M3 (Figures 2A, 2B, and 4).
[0139] The server system 106 includes a network transfer device 516, an audio encoder 518, a video decoder 520, a processor system 522, an audio encoder 524, and a video encoder 526. The processor system 522 includes processors 1, 2, 3, ... (and so on) up to processor P (where P is an integer). The processor system 522 also includes memory devices 1, 2, 3, ... (and so on) up to memory device P. For example, the combination of processor 1 and memory device 1 forms the first server, the combination of processor 2 and memory device 2 forms the second server, and so on, and the last combination of processor P and memory device P forms the Pth server. As another example, one or more of processors 1 to P form the AI processor shown in Figure 3A. For example, one or more of a feature extractor, a classifier, and an AI model are implemented in one or more of processors 1 to P, or are executed by one or more of processors 1 to P. Memory device 302 (Figure 3A) is an example of one of memory devices 1 to P.
[0140] Processors 1-P are examples of a sound-adding system 114, a physical phenomenon-adding system 112, and a graphics-adding system 116 (Figure 1). Each processor 1-P of system 522 is connected to the corresponding memory device 1-P of processor system 522. For example, processor P is connected to memory device P. An example of an audio decoder is a circuit that decompresses an encoded audio dataset into audio datasets such as audio frames. For example, the audio decoder decodes an encoded audio frame into an audio frame by applying an audio decoding protocol such as an audio decompression protocol. An example of a video decoder used herein is a circuit that performs a video decoding protocol such as H.262, H.263, or H.264 decoding, or another type of video decompression, on an encoded image frame and outputs decoded data such as an image frame.
[0141] The input controller 501 includes an input control 503 and a communication device 505. Examples of the input control 503 include one or more buttons, a touchpad, a touchscreen, and a joystick, all of which enable the user to make selections. Examples of communication devices used herein include circuits that apply a transfer communication protocol, such as a wired communication protocol or a wireless communication protocol. Examples of wireless communication protocols include the Bluetooth® protocol, the Near Field Wireless Communication Protocol, the Wi-Fi® protocol, and the Radio Frequency (RF) communication protocol.
[0142] The CPU 504 is connected to other components of the glasses 502, such as the GPU 506, display screen 508, camera 510, video encoder 512, microphone 518, audio encoder 516, network transfer device 514, communication device 507, video decoder 528, and audio decoder 530. The GPU 506 is connected to the display screen 508. The camera 510 is connected to the video encoder 512, and the video encoder 512 is connected to the network transfer device 514. The microphone 518 is connected to the audio encoder 516, and the audio encoder 516 is connected to the network transfer device 514. The network transfer device 514 is connected to the computer network 504. The communication device 507 is connected to the network transfer device 514. The network transfer device 514 is connected to the video encoder 516, video decoder 528, and audio decoder 530.
[0143] Furthermore, the network transfer device 516 of the server system 106 is connected to the audio decoder 518, video decoder 520, audio encoder 524, and video encoder 526. The audio decoder 518, video decoder 520, audio encoder 524, and video encoder 526 of the server system 106 are also connected to the processor system 522. For example, each of one or more processors 1 to P is connected to the audio decoder 518, video decoder 520, audio encoder 524, and video encoder 526.
[0144] When microphone 518 detects sound emitted from one or more real-world objects in the real-world environment where microphone 518 is placed, such as seat surfaces 202A or 204A or 254A or 254B or 254C or 402A (Figures 2A, 2B, and 4), it generates an audio dataset n such as audio dataset 1a or 1b or 2 or p (Figures 3A and 4) or a combination thereof (where n is an integer). Microphone 518 transmits audio dataset n to audio encoder 516. Audio encoder 516 applies an audio encoding protocol to audio dataset n and outputs an encoded audio dataset such as encoded audio frames. Audio encoder 516 transmits the encoded audio dataset to network transfer device 514.
[0145] Furthermore, camera 510 captures an image dataset n of one or more real-world objects in the real-world environment where camera 510 is located, such as a seat 202A or 204A or 254A or 254B or 254C or an office chair 202, for example, image dataset 1a or 1b or 2 or a combination thereof. The image dataset n is transmitted from camera 510 to video encoder 512. The video encoder 512 applies a video encoding protocol to the image dataset n to output encoded image frames, such as a combination of I-frames, B-frames, and P-frames, and provides the encoded image frames to network transfer device 514.
[0146] Furthermore, when the user makes one or more selections in the input control 503, the input control 503 generates an input dataset n, such as input dataset 1a, 1b, 2, or a combination thereof. An example of input dataset n is given in Listing 270 or 280 (Figure 2C or Figure 2D). Input dataset n is sent from the input control 503 to the communication device 505 of the input controller 501. The communication device 505 applies a transfer communication protocol to input dataset n to generate a transfer packet. The communication device 505 sends the transfer packet to the communication device 507 of the glasses 502 via a connection such as a wired or wireless connection between the communication device 505 and the communication device 507.
[0147] The communication device 507 applies a forwarding communication protocol to the forwarding packet to extract the input dataset n. Under the control of the CPU 504, the communication device 507 transmits the input dataset n to the network forwarding device 514.
[0148] The network transfer device 514 applies a network transfer protocol such as the Transmission Control Protocol over the Internet Protocol (TCP / IP) to embed encoded image frames received from the video encoder 512, encoded audio frames received from the audio encoder 516, or input dataset n received from the communication device 507, or a combination thereof, into an output data packet. The network transfer device 514 transmits the data packet to the network transfer device 516 of the server system 106 via the computer network 504.
[0149] The network transfer device 516 of the server system 106 applies the network transfer protocol to the data packets received from the network transfer device 514 to extract encoded audio frames, encoded image frames, input dataset n, or a combination thereof from the data packets. The network transfer device 514 sends the encoded audio frames to the audio decoder 518 and the encoded image frames to the video decoder 520. The audio decoder 518 applies the audio decoding protocol to the encoded audio frames to identify the audio dataset n from the encoded audio frames and sends the audio dataset n to the processor system 522. The video decoder 520 also applies the video decoding protocol to the encoded video frames to identify the image dataset n and sends the image dataset n to the processor system 522. Furthermore, the input dataset n is sent from the network transfer device 516 to the processor system 522.
[0150] The processor system 522 applies a feature extractor to extract audio parameters from audio dataset n and graphic and physical parameters from image dataset n. Furthermore, the processor system 522 applies a classifier to identify associations such as associations 352a or 352b or 354 or 362a or 362b or 364 (Figures 3B and 3C). An example of an association is the correspondence between a set containing audio parameters identified from audio datasets 1a, 1b, and 2 and a set containing seat material types and seat cover material types received in input datasets 1a, 1b, and 2. Another example of an association is the unique relationship between a set containing audio parameters identified from audio datasets 1a, 1b, and 2, graphic parameters identified from image datasets 1a, 1b, and 2, and physical parameters identified from image datasets 1a, 1b, and 2 and a set containing seat material types and seat cover material types received in input datasets 1a, 1b, and 2. The processor system 522 applies and trains an AI model based on associations. Upon receiving the audio dataset p, the association-trained AI model analyzes the audio dataset p and outputs the model output 412 (Figure 4A).
[0151] In one embodiment, the input controller 501 is not used with the glasses 502. Rather, in the embodiment, the glasses 502 include an additional internal camera directed at the eyes of the user wearing the glasses 502. The additional internal camera is connected to the CPU 504, the video encoder 512, and the network transfer device 514. The additional internal camera captures image data based on the user's eye gestures. For example, eye gestures are made to select the material type of the seat surface and the material type of the seat cover in the real-world environment where the glasses 502 are placed. For example, eye gestures are made to select the material type of the seat surface and the material type of the seat cover from list 270 or 280. Image data captured by the additional internal camera is an example of input data 1a or 1b or 2 (Figure 3A). Under the control of the CPU 504, the image data is transmitted from the additional internal camera to the video encoder 512. The video encoder 512 applies a video encoding protocol to the image data to output encoded image frames and transmits the encoded image frames to the network transfer device 514. The network transfer device 514 applies a network communication protocol to the encoded audio frame to generate a data packet and transmits the data packet to the network transfer device 516 of the server system 106 via the computer network 504.
[0152] The network transfer device 516 also applies a network communication protocol to the data packets to extract the encoded image frames and transmits the encoded image frames to the video decoder 520. The video decoder 520 applies a video decoding protocol to the encoded image frames to identify the image data captured by the additional internal camera and transmits the image data to the processor system 522. One or more of the processors 1 to P analyze the image data captured by the additional internal camera to identify the eye gestures made by the user, and further identify the material type of the seat surface and the material type of the seat cover in the real-world environment where the glasses 502 are placed.
[0153] Figure 6 is a diagram of an embodiment of the glasses 502, showing the sound output system 602 and the microphone 518. The glasses 502 includes a CPU 504, a sound output system 602, a microphone 518, and an audio memory device 606.
[0154] Microphone 518 includes a transducer, an acoustic energy-to-electrical energy converter (SE converter), an analog-to-digital converter (ADC), and a processor. The transducer is connected to the SE converter, and the SE converter is connected to the ADC. The processor of microphone 518 is connected to the transducer, the SE converter, and the ADC. An example of a transducer is a diaphragm. An example of an SE converter is a capacitor or a series of capacitors. The sound output system 602 includes a digital-to-analog converter (DAC), an amplifier, and a speaker. The DAC is connected to the amplifier, and the amplifier is connected to the speaker. The CPU 504 is connected to the DAC of the sound output system 602, the ADC of microphone 518, and the audio memory device 606.
[0155] The transducer detects sound emitted from or reflected from real-world objects in a real-world environment, such as System 200, 250, or 400 (Figures 2A, 2B, and 4), or both, and outputs vibrations. The vibrations are supplied to the SE transducer, where the electric field generated is modified to output an audio analog signal, which is an electrical signal. The audio analog signal is supplied to the ADC, where it is converted from analog to digital and outputs audio data such as audio dataset 1a, 1b, 2, or p (Figures 3A and 4).
[0156] When it is decided to display the virtual seat 408A (Figure 4A) of the virtual chair 408 and output sound in response to the movement of the virtual seat 408A, one or more processors 1 to P of the server system 106 access the audio parameter AP1a, the graphics parameter GP1a, and the physical parameter PP1a from one or more memory devices 1 to P, for example, by reading them. One or more processors 1 to P generate an image frame from the graphics parameter GP1a and the physical parameter PP1a, and an audio frame from the audio parameter AP1a, and transmit the image frame and audio frame to the glasses 502 via the computer network 504 in order to display the virtual seat 408A and output sound based on the movement of the virtual seat 408A. For example, one or more processors 1 to P receive an instruction from the input controller 501 via the glasses 502, the computer network 504, and the network transfer device 516 to display the virtual user 416 (Figure 4A) sitting on the virtual seat 408A. In the example, once determined in this way, one or more of the processors 1 to P generate a series of image frames having an output sequence of the physical parameters PP1a of the virtual seat surface 408A for showing how the virtual seat surface 408A rises from a decompression position to a compression position, and an output sequence of the graphic parameters GP1a of the virtual seat surface 408A synchronized with the output sequence of the physical parameters PP1a.
[0157] For example, an instance of the virtual seat surface 408A displayed in the first image frame within a series of image frames has a physical parameter PPi that provides the depressurization position Dpi and depressurization orientation DOi of the virtual seat surface 408A, and a graphic parameter GPi that corresponds to the depressurization position Dpi and depressurization orientation DOi of the virtual seat surface 408A. In the example, an instance of the virtual seat surface 408A displayed in the last image frame within a series of image frames has a physical parameter PPf that provides the compression position CPf and compression orientation COf, and a final graphic parameter GPf that corresponds to the compression position CPf and compression orientation COf. In the example, all intermediate image frames between the first and last image frames have physical parameters corresponding to the intermediate position and intermediate orientation, and intermediate graphic parameters. In the example, each image frame, such as the first, intermediate, and last image frames, is time-stamped, and the display order of the image frames is provided to show the virtual seat surface 408A transitioning from a depressurized state to a compressed state. In the example, the timestamp of the image frame is generated by one or more processors 1-P based on the generation time, such as by copying the generation time of the image containing the motion representation of the seat surface 202A from the depressurization position Dpi and depressurization orientation DOi to the compression position CPf and compression orientation COf, or by synchronizing with the generation time. In the example, the generation time of the image containing the motion representation of the seat surface 202A is generated by the camera 510 and received along with the image by one or more processors 1-P via the computer network 504 from the glasses 502.
[0158] In the example, once this is determined, one or more of the processors 1-P generate a series of audio frames having an output order for the audio parameter AP1a such that sound is output when the virtual seat surface 408A rises. Continuing from the previous example, the first audio frame has the first audio parameter that outputs the first sound at the point when an instance of the virtual seat surface 408A is displayed in the first image frame. In the example, the last audio frame has the last audio parameter that outputs the last sound at the point when an instance of the virtual seat surface 408A is displayed in the last image frame. In the example, all intermediate audio frames between the first and last audio frames have audio parameters corresponding to the intermediate position and orientation of the virtual seat surface 408A. In the example, each audio frame, including the first, intermediate, and last audio frames, is timestamped, providing an order in which the audio frames are output as sound as the virtual seat surface 408A transitions from a decompressed state to a compressed state. In the example, the timestamp of the audio frame is generated by one or more processors 1-P based on the generation time, such as by copying the generation time of the audio data generated in synchronization with the movement of the seat surface 202A from the depressurization position Dpi and depressurization orientation DOi to the compression position CPf and compression orientation COf, or in synchronization with the generation time. In the example, the generation time of the audio data generated in synchronization with the movement of the seat surface 202A is generated by the processor of the microphone 518 and is received along with the audio data from the glasses 502 via the computer network 504 by one or more processors 1-P.
[0159] Continuing the example, one or more of processors 1-P send an audio frame generated from audio parameter AP1a to audio encoder 524. Also in the example, one or more of processors 1-P send an image frame generated from graphic parameter GP1a and based on physical parameter PP1a to video encoder 526. In the example, audio encoder 524 applies an audio encoding protocol to the audio frame generated from audio parameter AP1a to output an encoded audio frame and provides the encoded audio frame to network transfer device 516. Furthermore, in the example, video encoder 526 applies a video encoding protocol to the image frame generated from graphic parameter GP1a to output an encoded image frame and sends the encoded image frame to network transfer device 516.
[0160] In the example, network transfer device 516 applies the network transfer protocol to the encoded image frame received from video encoder 526 and the encoded audio frame received from audio encoder 524 to generate a data packet. The example also shows network transfer device 516 sending the data packet to network transfer device 514 via computer network 504. Continuing the example, network transfer device 514 receives the data packet and applies the network transfer protocol to extract the encoded image frame and encoded audio frame. In the example, network transfer device 514 sends the encoded image frame to the video decoder 528 of the glasses 502 and the encoded audio frame to the audio decoder 530 of the glasses 502.
[0161] In the example, the video decoder 528 applies a video decoding protocol to the encoded image frame to output an image frame to show how the virtual seat surface 408A changes from a decompressed state to a compressed state, and provides the image frame to the CPU 504 of the glasses 502. In the example, the audio decoder 530 of the glasses 502 applies an audio decoding protocol to the encoded audio frame to output an audio frame to output sound in synchronization with the display of how the virtual seat surface 408A changes from a decompressed state to a compressed state, and provides the audio frame to the CPU 504. In the example, the CPU 504 controls the GPU 506 and further controls the display screen 508 to show how the virtual seat surface 408A of the virtual chair 408 changes from the initial decompression position DPi and initial decompression orientation DOi to the final compression position DPf and final compression orientation DOf. In the example, the change from the initial decompression position DPi and initial decompression orientation DOi to the final compression position DPf and final compression orientation DOf is an example of the physical parameter PP1a. In the example, GPU 506 displays the virtual seat surface 408A with graphics parameter GP1a synchronized with the movement of the virtual seat surface 408A from the initial decompression position DPi and initial decompression orientation DOi to the final compression position DPf and final compression orientation DOf. Also in the example, CPU 504 controls the sound output system 602 of glasses 502 and outputs sound synchronized with the display of the virtual seat surface 408A being compressed from the initial decompression position DPi and initial decompression orientation DOi to the final compression position DPf and final compression orientation DOf. In the example, sound is output according to the audio frame received by CPU 504.
[0162] Figure 7 shows an embodiment of the input controller 501. The input controller 501 includes a controller 702, a driver system 704, a haptic feedback system 706, and a network transfer device 708. An example of the controller 702 is a combination of a processor and a memory device. For example, the controller 702 includes a microprocessor or is a microcontroller. An example of the driver system is one or more drivers, such as one or more transistors. An example of the haptic feedback system described herein is one or more haptic feedback devices, such as actuators. An example of an actuator is a motor. The controller 702 is connected to the network transfer device 708 and the communication device 505. The controller 702 is also connected to the driver system 704, which is connected to the haptic feedback system 706. The network transfer device 708 is connected to the computer network 504.
[0163] The controller 702 receives haptic feedback data from one or more of the processors 1 to P (Figure 5) via the network transfer device 708 and the computer network 504. Upon receiving the haptic feedback data, the controller 702 sends one or more control signals to the driver system 704 based on the haptic feedback data. In response to receiving the control signals, the driver system 704 generates one or more driver signals, such as current signals, and sends one or more driver signals to the haptic feedback system 706. Upon receiving one or more driver signals, the haptic feedback system 706 vibrates to provide haptic feedback to the user 7 holding the input controller 501. For example, during the same period that the GPU 506 (Figure 5) displays a virtual robot 456 interacting with a coffee table 452, the controller 702 sends one or more control signals to the driver system 704. As another example, the controller 702 sends one or more control signals to the driver system 704 at the same time as it sends control signals from the GPU 506 to the display screen 508 (Figure 5) to show the virtual robot 456 interacting with the coffee table 452.
[0164] In this embodiment, the eyeglasses 502 (Figure 5) or another body-connected device includes a driver system and a haptic feedback system. A CPU 504 (Figure 5) is connected to the driver system of the eyeglasses 502. The driver system is connected to the haptic feedback system of the eyeglasses 502. The CPU 504 receives haptic feedback data from one or more processors 1 to P (Figure 5) via a network transfer device 514 (Figure 5) and a computer network 504. When the CPU 504 receives haptic feedback data for the eyeglasses 502, it transmits one or more control signals to the driver system of the eyeglasses 502 based on the haptic feedback data. In response to receiving the control signals, the driver system generates one or more driver signals and transmits one or more driver signals to the haptic feedback system of the eyeglasses 502. When the haptic feedback system receives one or more driver signals, it vibrates to provide haptic feedback to the user 7 wearing the eyeglasses 502. For example, during the same period that the GPU 506 (Figure 5) displays a virtual robot 456 interacting with the coffee table 452, the CPU 504 sends one or more control signals to the glasses driver system. As another example, while the GPU 506 sends control signals to the display screen 508 (Figure 5) to show the virtual robot 456 interacting with the coffee table 452, the CPU 504 sends one or more control signals to the driver system.
[0165] It should be noted that in various embodiments, one or more features of some embodiments described herein are combined with one or more features of one or more embodiments of the remaining embodiments described herein.
[0166] The embodiments described herein can be implemented in a variety of computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, miniature computers, and mainframe computers. In one embodiment, the embodiments described herein are implemented in a distributed computing environment in which tasks are performed by remote processing devices linked via a wired-based or wireless network.
[0167] With the embodiments described above in mind, it should be understood that in one embodiment, the embodiments described herein utilize various computer operations involving data stored in a computer system. These operations require the physical manipulation of physical quantities. Any of the operations described herein that form part of the embodiments described herein are useful machine operations. Some embodiments described herein also relate to devices or apparatus for performing these operations. The apparatus is either specially constructed for a required purpose, or it is a general-purpose computer selectively started or configured by a computer program stored in the computer. Specifically, in one embodiment, various general-purpose machines are used with computer programs written according to the teachings herein, or it may be more convenient to construct an apparatus more specialized to perform the required operations.
[0168] In some embodiments, several embodiments described in this disclosure are embodied as computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device that stores data to be read later by a computer system. Examples of computer-readable media include hard drives, network-attached storage (NAS), ROM, RAM, compact disk ROM (CD-ROM), CD-recordable (CD-R), CD-rewritable (CD-RW), magnetic tape, optical data storage devices, and non-optical data storage devices. For example, a computer-readable medium may include computer-readable tangible media distributed across a network-attached computer system so that the computer-readable code is stored and executed in a distributed manner.
[0169] Furthermore, while some of the embodiments described above are described in relation to a game environment, in some embodiments other environments, such as a video conferencing environment, are used instead of a game.
[0170] While the method operations have been described in a specific order, it should be understood that, as long as the processing of the overlay operations is performed in the desired manner, other housekeeping operations may be performed between operations, or operations may be coordinated to occur at slightly different times, or operations may be distributed throughout the system to allow processing operations to occur at various processing-related intervals.
[0171] While the embodiments described herein have been described in some detail for clarity, it will be apparent that certain changes and modifications can be made within the scope of the appended claims. Therefore, these embodiments should be considered illustrative rather than restrictive, and should not be limited to the details described herein, but may be modified within the scope of the appended claims and equivalents.
Claims
1. A method for identifying the material type of a real-world object in a real-world environment, The microphone captures an audio dataset based on sounds emitted from the real-world environment, including the real-world objects. Providing an artificial intelligence (AI) model with the audio dataset captured from the real-world environment to identify the material type of the real-world object in the real-world environment, wherein the audio dataset is processed to identify the audio parameters of the audio dataset, and the material type is identified based on a comparison of the audio parameters of the audio dataset with further audio parameters of further audio datasets. Displaying a first virtual object having the material type of the real-world object in a virtual environment, To provide the audio dataset that is output as sound when the first virtual object interacts with the second virtual object in the virtual environment, The method, including the method described above.
2. Receiving multiple audio datasets based on sounds received from multiple real-world objects in multiple real-world environments, Furthermore, the system receives multiple output parameters relating to the multiple real-world objects within the multiple real-world environments, Training the AI model based on the aforementioned multiple audio datasets and the aforementioned multiple output parameters, The method according to claim 1, further comprising:
3. The method according to claim 2, further comprising extracting from the plurality of audio datasets a plurality of features including a plurality of amplitudes and a plurality of frequencies of the plurality of audio datasets, wherein the plurality of audio datasets includes the further audio datasets and the plurality of features includes the further audio parameters.
4. The method further includes classifying the aforementioned multiple features and outputting the association between the aforementioned multiple features and the aforementioned multiple output parameters. Training the AI model includes providing the AI model with the associations between the plurality of features and the plurality of output parameters. Identifying the material type includes, based on the association between the plurality of features and the plurality of output parameters, determining by the AI model the probability that the audio dataset captured from the real-world environment represents the material type of the real-world object. The method according to claim 3.
5. The system further includes receiving multiple input datasets, which include multiple image datasets captured by multiple cameras in the multiple real-world environments, or multiple LiDAR data captured from the multiple real-world environments, or multiple inertial measurement unit (IMU) data captured from the multiple real-world environments, or combinations thereof. Extracting the aforementioned multiple features includes extracting the aforementioned multiple features from the aforementioned multiple input datasets. The aforementioned features include multiple identities of the multiple real-world objects, multiple physical parameters defining the movement of the multiple real-world objects, and multiple graphic parameters of the multiple real-world objects. The method according to claim 3.
6. The plurality of physical parameters include a first physical parameter and a second physical parameter, and the plurality of graphic parameters include a first graphic parameter and a second graphic parameter. The method further includes classifying the plurality of features and outputting the association between the plurality of features and the plurality of output parameters, Training the AI model includes providing the AI model with the associations between the plurality of features and the plurality of output parameters. The method according to claim 5.
7. The method according to claim 6, wherein identifying the material type of the real-world object in the real-world environment includes identifying the probability that the real-world object has the material type, the probability being identified based on the association between the plurality of features and the plurality of output parameters.
8. To generate virtual object data for displaying how the first virtual object interacts with the second virtual object in the virtual environment, When the first virtual object interacts with the second virtual object, the audio dataset output by the first virtual object is generated, When the first virtual object interacts with the second virtual object, haptic feedback data is generated that is output by the body connection device. It further includes, The virtual object data represents the material type, the movement of the real-world object from the first physical parameter to the second physical parameter in the real-world environment, and the change from the first graphic parameter to the second graphic parameter of the real-world object in the real-world environment. The method according to claim 7.
9. The plurality of output parameters include data that identifies the plurality of materials of the plurality of real-world objects in the plurality of real-world environments, The method further includes simulating a virtual interaction between the first virtual object and the second virtual object of a video game based on the material type. The method according to claim 2.
10. This further includes receiving multiple input datasets, The aforementioned multiple input datasets include multiple image datasets captured by multiple cameras in the aforementioned multiple real-world environments, The plurality of audio datasets and the plurality of image datasets are captured when the plurality of real-world objects interact with the plurality of users in the plurality of real-world environments. The method according to claim 2.
11. A server for identifying the material type of real-world objects in a real-world environment, Receiving an audio dataset from a microphone in the aforementioned real-world environment, based on sounds emitted from the aforementioned real-world environment, including the aforementioned real-world objects, Providing the received audio dataset to an artificial intelligence (AI) model to identify the material type of the real-world object in the real-world environment, wherein the audio dataset is processed to identify the audio parameters of the audio dataset, and the material type is identified based on a comparison of the audio parameters of the audio dataset with further audio parameters of further audio datasets. The material type is provided to the client device to display a first virtual object having the material type of the real-world object in the virtual environment. To provide the audio dataset that is output as sound when the first virtual object interacts with the second virtual object in the virtual environment, A processor configured to perform, A memory device connected to the aforementioned processor, The server comprising the above-mentioned equipment.
12. The aforementioned processor, Receiving multiple audio datasets based on sounds received from multiple real-world objects in multiple real-world environments, Furthermore, the system receives multiple output parameters relating to the multiple real-world objects within the multiple real-world environments, Training the AI model based on the aforementioned multiple audio datasets and the aforementioned multiple output parameters, The server according to claim 11, configured to perform the following:
13. The server according to claim 12, wherein the processor is configured to extract from the plurality of audio datasets a plurality of features including a plurality of amplitudes and a plurality of frequencies of the plurality of audio datasets, the plurality of audio datasets including the further audio datasets and the plurality of features including the further audio parameters.
14. The server according to claim 13, wherein the AI model is trained based on the plurality of audio datasets and the plurality of output parameters.
15. The aforementioned processor, The system is configured to classify the aforementioned multiple features and output the association between the aforementioned multiple features and the aforementioned multiple output parameters. To train the AI model, the processor is configured to provide the AI model with the associations between the plurality of features and the plurality of output parameters. To identify the material type, the processor is configured to use the AI model to determine the probability that the received audio dataset represents the material type, based on the association between the plurality of features and the plurality of output parameters. The server according to claim 14.
16. The processor is configured to receive multiple input datasets, The aforementioned plurality of input datasets include a plurality of image datasets captured by a plurality of cameras in the plurality of real-world environments, or light-detection ranging (LiDAR) data captured from the plurality of real-world environments, or inertial measurement unit (IMU) data captured from the plurality of real-world environments, or a combination thereof. In order to extract the aforementioned multiple features, the processor is configured to acquire the aforementioned multiple features from the aforementioned multiple input datasets. The aforementioned features include multiple identities of the multiple real-world objects, multiple physical parameters defining the movement of the multiple real-world objects, and multiple graphic parameters of the multiple real-world objects. The server according to claim 13.
17. The plurality of physical parameters include a first physical parameter and a second physical parameter, and the plurality of graphic parameters include a first graphic parameter and a second graphic parameter. The processor is configured to classify the plurality of features and output the association between the plurality of features and the plurality of output parameters. The processor is configured to train the AI model based on the plurality of audio datasets and the plurality of output parameters. To train the AI model, the processor is configured to provide the AI model with the associations between the plurality of features and the plurality of output parameters. The server according to claim 16.
18. In order to identify the material type of the real-world object in the real-world environment, the processor is configured to determine the probability that the real-world object has the material type. The probability is determined based on the association between the plurality of features and the plurality of output parameters. The server according to claim 17.
19. The aforementioned processor, To generate virtual object data for displaying how the first virtual object interacts with the second virtual object in the virtual environment, When the first virtual object interacts with the second virtual object, the audio dataset output by the first virtual object is generated, When the first virtual object interacts with the second virtual object, haptic feedback data is generated that is output by the body connection device. It is configured to perform, The virtual object data represents the material type, the movement of the real-world object from the first physical parameter to the second physical parameter in the real-world environment, and the change from the first graphic parameter to the second graphic parameter of the real-world object in the real-world environment. The server according to claim 18.
20. A system for identifying the material type of real-world objects in a real-world environment, A client device configured to capture an audio dataset from the real-world environment using a microphone, wherein the audio dataset is captured based on sounds emitted from the real-world environment, including the real-world objects, and the client device, A server connected to the client device via a computer network, The server is equipped with, Receiving the audio dataset from the client device via the aforementioned computer network, Providing an artificial intelligence (AI) model with the audio dataset captured from the real-world environment to identify the material type of the real-world object in the real-world environment, wherein the audio dataset is processed to identify the audio parameters of the audio dataset, and the material type is identified based on a comparison of the audio parameters of the audio dataset with further audio parameters of further audio datasets. To provide the aforementioned material type and to display a first virtual object having the aforementioned material type of the real-world object in a virtual environment, To provide the audio dataset that is output as sound when the first virtual object interacts with the second virtual object in the virtual environment, The system configured to perform the following.
21. The system further comprises multiple client devices, and the multiple client devices are This involves generating multiple audio datasets based on sounds received from multiple real-world objects within multiple real-world environments, and Receiving multiple output parameters relating to multiple material types of the multiple real-world objects, It is configured to perform, The server is connected to the plurality of client devices via the computer network, and the server Receiving the multiple audio datasets from the multiple client devices via the computer network, Receiving the multiple output parameters from the multiple client devices via the computer network, Training the AI model based on the aforementioned multiple audio datasets and the aforementioned multiple output parameters, Extracting from the plurality of audio datasets a plurality of features including a plurality of amplitudes and a plurality of frequencies of the plurality of audio datasets, The system according to claim 20, configured to perform, wherein the plurality of audio datasets include the further audio datasets, and the plurality of features include the further audio parameters.
22. The server is configured to classify the multiple features and output the association between the multiple features and the multiple output parameters. In order to train the AI model, the server is configured to provide the AI model with the associations between the plurality of features and the plurality of output parameters. To identify the material type, the server is configured to use the AI model to determine the probability that the audio dataset captured from the real-world environment represents the material type, based on the association between the plurality of features and the plurality of output parameters. The system according to claim 21.
Citation Information
Patent Citations
Road surface condition measuring device
JP1994138018A
Object identification system
JP1994273394A
Apparatus for converting plastic into oil
JP2002356680A
Apparatus for extracting feature sound, feature sound extraction method, and product evaluation system
JP2005140707A
Tactile information detector and tactile information detection method
JP2010044028A