How to examine the game context to determine the user's voice commands
Context-based automatic speech recognition in video games addresses the challenge of dynamic gameplay contexts by analyzing gameplay elements and offering candidate words for confirmation, improving voice command interpretation and immersion.
Patent Information
- Application Number
- JP2024568839
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-05-28
- Filing Date
- 2023-04-17
- Publication Date
- 2026-03-05
- Estimated Expiration
- 2043-04-17
AI Technical Summary
Existing video game technologies struggle to accurately interpret voice commands due to dynamic gameplay contexts, leading to suboptimal player interaction and immersion.
Implementing context-based automatic speech recognition in video games by analyzing gameplay context, such as scene, avatar attributes, and player movements, to enhance speech recognition accuracy, and providing candidate words for player confirmation when confidence is low.
Improves speech recognition accuracy in dynamic gameplay scenarios, enhancing player interaction and immersion by accurately interpreting voice commands in real-time.
Smart Images

Figure 0007825078000001 
Figure 0007825078000002 
Figure 0007825078000003
Abstract
Description
[Technical Field]
[0001] FIELD OF THE DISCLOSURE This disclosure relates generally to verifying game context to determine a user's voice commands within a video game. [Background technology]
[0002] 2. Description of Related Art The video game industry has undergone many changes over the years. As technology advances, video games continue to achieve greater immersion through sophisticated graphics, realistic sounds, compelling soundtracks, haptic feedback, and more. Players can enjoy immersive gaming experiences that involve and engage with the virtual environment, demanding new ways of interaction.
[0003] It is in this context that embodiments of the present disclosure arise. Summary of the Invention
[0004] Embodiments of the present disclosure include methods, systems, and devices related to verifying game context to determine a user's voice commands in a video game.
[0005] In some embodiments, a method of executing a video game session is provided, the method including the following operations: recording audio of players engaged in gameplay of the video game session; analyzing a game state generated by the execution of the video game session, wherein analyzing the game state identifies a gameplay context; analyzing the recorded audio to identify textual content of the recorded audio using the identified gameplay context and a speech recognition model; and applying the identified textual content as gameplay input for the video game session.
[0006] In some embodiments, the method is performed in substantially real time, such that analyzing the game state is performed in substantially real time as the game state is continually updated as the session is executed, and analyzing the recorded audio to identify textual content is responsive to changes in the game state in substantially real time.
[0007] In some implementations, the identified context of the gameplay includes an identified scene or stage of the video game.
[0008] In some implementations, the identified context of gameplay includes an identified virtual location within a virtual environment defined by the execution of a session of a video game.
[0009] In some implementations, the identified context of the gameplay includes identified attributes of a player's avatar within the video game.
[0010] In some implementations, analyzing the game state includes recognizing activity occurring in game play, the recognized activity at least partially defining an identified context of game play.
[0011] In some implementations, analyzing the game state includes predicting future gameplay activity, the predicted future gameplay at least partially defining the identified context of gameplay.
[0012] In some embodiments, a method of conducting a video game session is provided, the method including: recording audio of a player engaged in gameplay of the video game session; analyzing the recorded audio with a speech recognition model, where the speech recognition model identifies a plurality of candidate words as possible interpretations of the player's audio; presenting the plurality of candidate words to the player during the video game session; receiving a selection input from the player that identifies one of the candidate words as a correct interpretation of the player's audio; and applying, by the video game, the selected one of the candidate words as a gameplay input for the video game.
[0013] In some implementations, presenting the plurality of candidate words is in response to the confidence of recognition by the speech recognition model being below a predetermined threshold.
[0014] In some implementations, presenting the candidate words to the player includes presenting the candidate words in a video generated from the video game session.
[0015] In some implementations, presenting a candidate word pauses game play until a selection input is received.
[0016] In some implementations, applying the selected one of the candidate words includes triggering a command for gameplay of the video game.
[0017] In some implementations, a selected one of the candidate words is used as feedback to refine the speech recognition model.
[0018] In some implementations, receiving the selection input occurs via an input device of the controller.
[0019] Other aspects and advantages of the present disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, illustrating by way of example the principles of the disclosure. [Brief explanation of the drawings]
[0020] The present disclosure may be better understood by reference to the following description taken in conjunction with the accompanying drawings.
[0021] [Figure 1] 1 illustrates a user engaging in gameplay of a video game, according to an embodiment of the present disclosure.
[0022] [Figure 2] 1 conceptually illustrates dynamic adjustment of a speech recognition model based on changing game context, according to an embodiment of the present disclosure.
[0023] [Figure 3] 1 conceptually illustrates a player providing feedback regarding speech recognition applied to their voice during gameplay of a video game, according to an embodiment of the present disclosure.
[0024] [Figure 4] 1 conceptually illustrates game context prediction for improving speech recognition during gameplay of a video game, according to an embodiment of the present disclosure.
[0025] [Figure 5] 1 conceptually illustrates a method for applying voice recognition in a video game, according to an embodiment of the present disclosure.
[0026] [Figure 6] 1 illustrates components of an exemplary device that can be used to implement aspects of various embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0027] The following embodiments of the present disclosure provide methods, systems, and devices for verifying game context to determine a user's voice commands in a video game.
[0028] Broadly speaking, embodiments of the present disclosure are directed to methods and systems for providing contextual automatic speech recognition in video games. As a player plays a given video game, the gameplay context in which the player interacts with the video game's virtual environment is constantly changing. However, the gameplay context may provide important clues about what the player may be saying at any given moment. Therefore, according to embodiments of the present disclosure, the gameplay context is utilized to enhance the player's speech recognition. For example, aspects of the gameplay, such as the scene or setting, the characteristics of the player's avatar, and the player's avatar's movements, are used to improve speech recognition. In some implementations, because speech is used to provide commands for gameplay, achieving high accuracy of speech recognition is important to providing a high-quality gameplay experience for the player. In some implementations, if the system cannot determine with sufficient confidence the correct interpretation of the player's given speech, candidate words or phrases may be presented to the player for selection, indicating which candidate is the correct interpretation of the user's speech.
[0029] With the above summary in mind, the following provides several illustrative figures to facilitate understanding of example embodiments.
[0030] FIG. 1 illustrates a user engaging in gameplay of a video game, according to an embodiment of the present disclosure.
[0031] In the illustrated implementation, a user / player 100 is engaged in interactive gameplay of a video game executed by a computing device 102. By way of example and not limitation, the computing device 102 may be a game console, a personal computer, a laptop, a set-top box, or other general-purpose or special-purpose computer having a processor and memory and capable of executing program instructions for a video game. Furthermore, while a user interacting with a video game is described for purposes of illustrating embodiments of the present disclosure, it will be understood that the principles of the present disclosure are applicable to other types of interactive applications that may be executed by a computing device 102 and with which a user 100 may interact. Furthermore, it will be understood that in other implementations, at least a portion of the functionality attributed to the computing device 102 may be performed in the cloud by remote cloud processing resources 132 accessed via a network 130 (including, for example, the Internet).
[0032] In some implementations, computing device 102 executes video game session 120, which includes generating game state 122 and rendering video and audio of the video game's virtual environment. The video is presented on display 104 (e.g., a television, monitor, LED / LCD display, projector screen, etc.) for viewing by user 100. In some implementations, display 104 takes the form of a head-mounted display (HMD) worn by the user. As user 100 interacts with the video game, user 100 may manipulate a user input device, such as controller 106 in the illustrated implementation, to generate user input for the video game. It will be appreciated that the running video game responds to user input generated as a result of the user's interactivity and continually updates the game state based on that user input to drive the execution of the video game.
[0033] Audio generated by the video game may be presented through an audio device such as speakers (which may be included in the display 104 or, as shown in the illustrated embodiment, in headphones 108 worn by the user 100). The headphones 108 may be wired or wireless and may be directly or indirectly connected to the computing device 102. For example, in some implementations, the headphones 108 are connected to the controller 106, which communicates with the computing device 102 to receive audio that is passed to the headphones 108 for presentation.
[0034] Broadly speaking, embodiments of the present disclosure are directed to methods and systems for providing context-based automatic speech recognition in video games. To that end, the voice of a user 100 during gameplay is captured by a microphone 110. In the illustrated embodiment, the microphone 110 is connected to headphones 108. However, in other embodiments, the microphone may be a separate device, a device integrated into the controller 106, part of another device that senses the local environment (e.g., a device having a camera and microphone for detecting the user and the local environment), part of an HMD worn by the user, or in any other form provided to enable capturing the user's voice during gameplay.
[0035] The user's 100 speech is captured and recorded as audio data. The recorded audio data is analyzed using a speech recognition model 124 to identify the textual content of the user's speech. More specifically, to improve speech recognition, the user's current gameplay context is identified and applied to the speech recognition model. It will be understood that in a given game context, a user is more likely to speak certain words or phrases compared to another game context. For example, in a combat scene, a user is more likely to speak words or phrases specific to combat, such as issuing a command to fire a weapon or launching an attack against an enemy. Therefore, understanding the game context can improve automatic speech recognition of the user's speech.
[0036] In some implementations, the game state 122 is analyzed to identify the context of the user's gameplay. It will be appreciated that gameplay is dynamic and the context may change from moment to moment. Thus, in some implementations, a real-time or substantially real-time understanding of the game context is applied to optimize speech recognition. In this manner, the speech recognition model 124's ability to determine the textual content of the user's (recorded) speech is responsive to real-time changes occurring during gameplay, as reflected in the game state, which is continually updated during interactive gameplay.
[0037] In some implementations, the game state 122 is analyzed using a trained machine learning model configured to recognize or identify gameplay context. In some implementations, such machine learning models can be configured to predict future gameplay activity / context based on the current game state, and this information may also be used to enhance automatic speech recognition. Broadly speaking, the speech recognition model 124 is adjusted based on the determined gameplay context, e.g., to prioritize or weight text decisions that are more likely in light of the gameplay context.
[0038] After the textual content (words or phrases) of the user's speech is determined, it can be applied to the video game session 120. For example, the textual content may be interpreted to trigger one or more commands that are applied to game play. In this way, speech recognition can be improved to better recognize voice commands used during game play of the video game.
[0039] In the illustrated exemplary implementation, the video gaming and speech recognition are performed by computing device 102, however, in other implementations, either or both of the video gaming and speech recognition may be performed / performed remotely by cloud processing resources 132 accessed via network 130. In some implementations, recorded audio of the user's speech is sent to cloud processing resources 132 and analyzed to determine the textual content of the user's speech using speech recognition models and game context.
[0040] FIG. 2 conceptually illustrates dynamic adjustment of a speech recognition model based on changing game context, according to an embodiment of the present disclosure.
[0041] It will be appreciated that the gameplay context of a video game continually changes as a user progresses during a session. For example, in an illustrative implementation, at an earlier point in time, a player may engage in gameplay in a first video game scene 200, in which the player may control a character avatar 202. Meanwhile, at a later point in time in the session, the player may engage in gameplay in a second video game scene 204, in which the player may, by way of example and not limitation, control an airplane 206. It will be appreciated that in such scenes, the types of sounds the player may make will differ due to the different nature of the scene and the different requirements for gameplay activity (e.g., controlling a character avatar versus controlling an airplane / vehicle).
[0042] Accordingly, a dynamic automatic speech recognition (ASR) tuning process 208 is performed, which adjusts the speech recognition model 210 based on the game context occurring at the time. In some implementations, the dynamic ASR process includes analyzing the game state of the video game to determine the current gameplay context. In some implementations, the game state is analyzed using a machine learning model trained to recognize or identify gameplay contexts of the video game. By way of example and not limitation, the game context may include identifying aspects that describe the type of gameplay occurring, such as a setting, scene, stage, section, virtual / spatial location within a virtual environment, progress, timeline / temporal location within a campaign storyline, current objective, character / avatar attributes, abilities, inventory, achievements, skill level, etc., or other types of information about gameplay that may affect the likelihood that a player will speak a particular word or phrase during gameplay.
[0043] For example, during scene 200, the player controls character avatar 202 and is therefore likely to speak words / phrases related to operating character avatar 202. Meanwhile, in scene 204, the player controls airplane 206 and is therefore likely to speak words / phrases related to operating airplane 206. Thus, in each of scenes 200 and 204, dynamic ASR tuning process 208 tunes speech recognition model 210 in a different manner, such as by adjusting / tuning weights or variables of speech recognition model 210, to prioritize words that are more relevant and likely to be spoken.
[0044] In various implementations, the speech recognition model may be based on various recognition models known in the art, including acoustic models and language models, hi some implementations, the speech recognition model is based on a hidden Markov model, an artificial neural network, a deep neural network, an AI / machine learning model, or other type of suitable modeling mechanism.
[0045] FIG. 3 conceptually illustrates a player providing feedback regarding the voice recognition applied to their voice during gameplay of a video game, according to an embodiment of the present disclosure.
[0046] In the illustrated implementation, player 100 is engaged in gameplay of a video game, and gameplay video is being rendered on display 104. As discussed, automatic speech recognition is applied to the voice of player 100. Also, as previously mentioned, gameplay context can be determined and used to improve the results of the speech recognition process. However, in some cases, the voice of player 100 may not be recognized with sufficient confidence.
[0047] Accordingly, in some embodiments, a method is implemented to request additional feedback from the player. For example, at operation 300, a threshold determination is made regarding the confidence of the speech recognition. If the player's speech is recognized by the speech recognition model and the confidence meets or exceeds a predetermined threshold, the recognized textual content of the player's speech is applied as described above. However, if the confidence is below the predetermined threshold, at method operation 302, the system identifies potential candidate words / phrases that are possible interpretations of the player's recorded speech. These candidates are presented to the player 100, such as by displaying a dialog window 308 on the display 104 and asking the player 100 whether any of the candidate words / phrases are intended.
[0048] At operation 304, the player 100 can respond by indicating a selection of one of the candidates as the correct interpretation. For example, the dialog window 308 may indicate that a particular button on the controller 106 operated by the player 100 is associated with the presented candidate word / phrase, allowing the player to indicate selection of the given candidate word / phrase by pressing the corresponding controller button.
[0049] In operation 306, the word / phrase indicated as the correct interpretation of the player's speech is applied to the video game, such as by triggering a command for game play. Additionally, information regarding the correct interpretation of the player's speech may be used as feedback or further training data for a speech recognition model to improve the accuracy of the speech recognition model in determining the textual content of the player's speech.
[0050] FIG. 4 conceptually illustrates game context prediction for improving speech recognition during gameplay of a video game, according to an embodiment of the present disclosure.
[0051] In the illustrated implementation, a player 400 experiences a video game through a VR headset (or head-mounted display) 402. The player can manipulate controllers 404a and 404b and can issue voice commands or other sounds related to the video game. The player 400 views a virtual environment 410, which can be from the perspective of the player's avatar 412 within the virtual environment. It will be appreciated that the VR headset 402 provides the player 400 with an immersive view of the virtual environment 410.
[0052] In some implementations, in addition to determining gameplay context to improve speech recognition, future gameplay activity and context can be predicted based on current game state data, and this future gameplay activity / context can also be used to improve speech recognition. For example, player 400 may be looking in a given direction 414 within virtual environment 410 or moving in direction 414. Based on such gaze direction or movement, a system (e.g., including a machine learning model that analyzes game state as described above) can predict that the player's future game context will include a virtual location within region 416 of virtual environment 410. This data can be used to tune a speech recognition model before the player arrives in region 416, for example, to prioritize words or phrases that are likely to be spoken while in region 416.
[0053] Although predicting future locations within a virtual environment has been described, it will be appreciated that other types of gameplay context information, such as the examples described above, may be predicted and used to enhance a video game player's speech recognition process during gameplay.
[0054] FIG. 5 conceptually illustrates a method for applying voice recognition in a video game, according to an embodiment of the present disclosure.
[0055] A video game session is initiated at method operation 500. At method operation 502, player audio occurring during gameplay of the video game is recorded.
[0056] At method operation 504, a current gameplay context in which the player speech occurs is determined, such as by analyzing the game state of the running session. At method operation 506, the determined gameplay context is applied to a speech recognition process, and speech recognition incorporating the gameplay context is performed on the recorded speech to determine the textual content of the player speech.
[0057] At method operation 508, a confidence level in the speech recognition results is determined. If the confidence level exceeds a predetermined threshold deemed to indicate a sufficiently high degree of confidence in the speech recognition results, then at method operation 510, the textual content of the player's speech as determined by the speech recognition is applied to gameplay of the video game, such as by triggering a command or some other action during the video game session.
[0058] If the confidence does not exceed a predetermined threshold, then at method operation 512, candidate words / phrases determined by the speech recognition process are presented to the player as possible interpretations of the player's speech. At method operation 514, player input is received identifying one of the candidates as the correct interpretation. Once the correct interpretation is identified, the associated text content is applied to the video game at method operation 510, as described above. Further, at method operation 516, the identification of the correct interpretation of the player's speech is used to improve speech recognition, such as by updating / training / tuning a speech recognition model.
[0059] In some implementations described herein, analysis of game context and speech recognition are performed by separate models. However, in other implementations, a single speech recognition model is configured to evaluate game context and perform speech recognition jointly. That is, in some implementations, a speech recognition model may be configured to accept game state data (indicative of game context) and a player's recorded speech and perform contextual speech recognition.
[0060] In other implementations, there may be multiple speech recognition models specific to different game contexts, for example, a first speech recognition model may be trained and configured to recognize speech occurring during a first level / scene of a video game, a second speech recognition model may be trained and configured to recognize speech occurring during a second level / scene of the video game, and so on.
[0061] FIG. 6 illustrates components of an exemplary device 600 that can be used to implement aspects of various embodiments of the present disclosure. The block diagram illustrates device 600, which can incorporate or be a personal computer, video game console, personal digital assistant, server, or other digital device suitable for implementing embodiments of the present disclosure. Device 600 includes a central processing unit (CPU) 602 for executing software applications and optionally an operating system. CPU 602 may be comprised of one or more homogeneous or heterogeneous processing cores. For example, CPU 602 is one or more general-purpose microprocessors having one or more processing cores. Further embodiments can be implemented using one or more CPUs with a microprocessor architecture particularly suited for highly parallel and computationally intensive applications, such as interpreting queries, identifying context-relevant resources, and immediately implementing and rendering context-relevant resources within a video game. Device 600 may be located near players playing game segments (e.g., a game console), or remote from players (e.g., a back-end server processor), or one of many servers using virtualization in a game cloud system for remote streaming of gameplay to clients.
[0062] Memory 604 stores applications and data used by CPU 602. Storage 606 provides non-volatile storage and other computer-readable media for applications and data and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other optical storage devices, as well as signal transmission and storage media. User input device 608 communicates user input from one or more users to device 600; examples of user input device 608 may include a keyboard, mouse, joystick, touchpad, touchscreen, still or video recorder / camera, gesture-recognizing tracking device, and / or microphone. Network interface 614 enables device 600 to communicate with other computer systems over an electronic communications network, which may include wired or wireless communications over a local area network or a wide area network such as the Internet. The audio processor 612 is adapted to generate analog or digital audio output from instructions and / or data provided by the CPU 602, memory 604, and / or storage 606. The components of the device 600, including the CPU 602, memory 604, data storage 606, user input device 608, network interface 610, and audio processor 612, are connected via one or more data buses 622.
[0063] A graphics subsystem 620 is further coupled to the data bus 622 and the components of device 600. Graphics subsystem 620 includes a graphics processing unit (GPU) 616 and a graphics memory 618. The graphics memory 618 includes display memory (e.g., a frame buffer) used to store pixel data for each pixel of an output image. The graphics memory 618 may be integrated into the same device as the GPU 608, connected as a separate device from the GPU 616, and / or incorporated within memory 604. Pixel data may be provided to the graphics memory 618 directly from the CPU 602. Alternatively, the CPU 602 may provide data and / or instructions defining desired output images to the GPU 616, from which the GPU 616 generates pixel data for one or more output images. The data and / or instructions defining the desired output images may be stored in memory 604 and / or the graphics memory 618. In an embodiment, GPU 616 includes 3D rendering functionality for generating pixel data for output images from instructions and data defining scene geometry, lighting, shading, texturing, motion, and / or camera parameters. GPU 616 may further include one or more programmable execution units capable of executing shader programs.
[0064] Graphics subsystem 614 periodically outputs image pixel data from graphics memory 618 for display on display device 610. Display device 610 can be any device capable of displaying visual information in response to signals from device 600, including CRT, LCD, plasma, and OLED displays. Device 600 can provide analog or digital signals to display device 610, for example.
[0065] It should be noted that access services distributed over a wide geographic area, such as providing access to games in the current embodiment, often use cloud computing. Cloud computing is a computing paradigm in which dynamically scalable, often virtualized resources are provided as a service over the Internet. Users do not need to be experts in the technical infrastructure of the "cloud" that supports them. Cloud computing can be categorized into different services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). Cloud computing services often provide common applications, such as video games, online, accessed through a web browser, but the software and data are stored on servers in the cloud. The term cloud is used as a metaphor for the Internet, based on how the Internet is depicted in computer network diagrams, and is an abstract concept that hides a complex infrastructure.
[0066] In some embodiments, a game server may be used to perform the operation of a persistent information platform for video game players. Most video games played over the Internet operate through a connection to a game server. Typically, the game uses a dedicated server application that collects data from players and distributes the collected data to other players. In other embodiments, the video game may be executed by a distributed game engine. In these embodiments, the distributed game engine may run on multiple processing entities (PEs), such that each PE executes a functional segment of the given game engine on which the video game is executed. Each processing entity is viewed by the game engine as simply a computational node. A game engine typically performs a functionally diverse set of operations to execute a video game application along with additional services for the user experience. For example, the game engine implements game logic and performs game calculations, physics, geometry transformations, rendering, lighting, shading, audio, and additional in-game or game-related services. The additional services may include, for example, messaging, social utilities, audio communication, gameplay playback functions, help functions, etc. The game engine may run on an operating system virtualized by a hypervisor on a particular server, but in other embodiments the game engine itself may be distributed across multiple processing entities, each residing on a different server unit in a data center.
[0067] According to this embodiment, each processing entity for performing operations may be a server unit, a virtual machine, or a container, depending on the needs of each game engine segment. For example, if a game engine segment is responsible for camera transformations, that particular game engine segment may be provisioned with a virtual machine associated with a graphics processing unit (GPU) since it will be performing a large number of relatively simple mathematical operations (e.g., matrix transformations). Other game engine segments requiring fewer but more complex operations may be provisioned with processing entities associated with one or more more powerful central processing units (CPUs).
[0068] By distributing the game engine, the game engine has elastic computational characteristics that are not bound by the capabilities of a physical server unit. Instead, the game engine is provisioned with more or fewer computational nodes as needed to meet the demands of the video game. From the perspective of the video game and the video game player, a game engine that is distributed across multiple computational nodes is indistinguishable from a non-distributed game engine running on a single processing entity, as a game engine manager or supervisor distributes the workload and seamlessly integrates the results to provide the video game output component to the end user.
[0069] Users access remote services through client devices that include at least a CPU, a display, and I / O. The client devices may be PCs, mobile phones, netbooks, PDAs, etc. In one embodiment, a network running on the game server recognizes the type of device used by the client and adjusts the communication method to employ. In another example, the client devices access applications on the game server over the Internet using standard communication methods such as HTML. It should be understood that a given video game or game application may be developed for a particular platform and a particular associated controller device. However, when such games are made available through a game cloud system as presented herein, users can access the video game with different controller devices. For example, a game may be developed for a game console and its associated controller, but a user can access a cloud-based version of the game from a personal computer utilizing a keyboard and mouse. In such a scenario, the input parameter configuration can define a mapping from inputs that can be generated by the user's available controller device (in this example, a keyboard and mouse) to inputs that are acceptable for execution of the video game.
[0070] In another example, a user may access the cloud gaming system via a tablet computing device, a touchscreen smartphone, or other touchscreen-driven device. In this case, the client device and the controller device are integrated together within the same device, and input is provided by detected touchscreen input / gestures. For such devices, the input parameter configuration may define specific touchscreen inputs corresponding to game inputs for the video game. For example, buttons, directional pads, or other types of input elements may be displayed or overlaid during the execution of the video game to indicate locations on the touchscreen that the user can touch to generate game input. Gestures such as swipes in specific orientations or specific touch motions may also be detected as game inputs. In one embodiment, to familiarize the user with control operations on the touchscreen, a tutorial showing how to input gameplay inputs via the touchscreen may be provided to the user, for example, before beginning gameplay of the video game.
[0071] In some embodiments, the client device serves as a connection point for the controller device. That is, the controller device communicates with the client device via a wireless or wired connection and transmits inputs from the controller device to the client device. The client device may then process these inputs and then transmit the input data to a cloud gaming server over a network (e.g., a network accessed through a local network device such as a router). However, in other embodiments, the controller itself may be a networked device capable of communicating inputs directly to the cloud gaming server over a network, without the need for such inputs to first be communicated through the client device. For example, the controller may connect to a local network device (such as the router mentioned above) to send and receive data from the cloud gaming server. Thus, while the client device may still be required to receive video output from the cloud-based video game and render it on a local display, input latency can be reduced by the controller transmitting inputs directly to the cloud gaming server over the network, bypassing the client device.
[0072] In one embodiment, the networked controller and client device can be configured to transmit certain types of input directly from the controller to the cloud gaming server and other types of input via the client device. For example, detected inputs that do not rely on any additional hardware or processing, aside from the controller itself, can be transmitted directly from the controller to the cloud gaming server over the network, bypassing the client device. Such inputs can include button inputs, joystick inputs, embedded motion-sensing inputs (e.g., accelerometers, magnetometers, gyroscopes), etc. However, inputs that utilize additional hardware or require processing by the client device can be transmitted by the client device to the cloud gaming server. These can include video or audio captured from the game environment, which can be processed by the client device before transmission to the cloud gaming server. Additionally, inputs from the controller's motion-sensing hardware can be processed by the client device in conjunction with the captured video to detect the position and movement of the controller, which is then communicated by the client device to the cloud gaming server. It should be understood that controller devices according to various embodiments can also receive data (e.g., feedback data) from the client device or directly from the cloud gaming server.
[0073] In one embodiment, various example techniques can be implemented using a virtual environment via a head-mounted display (HMD). An HMD is sometimes referred to as a virtual reality (VR) headset. As used herein, the term "virtual reality" (VR) generally refers to user interaction with a virtual space / environment, including viewing the virtual space through an HMD (or VR headset) in a way that responds in real time to the (user-controlled) movements of the HMD to provide the user with the sensation of being in the virtual space or metaverse. For example, when facing a given direction, the user can see a three-dimensional (3D) view of the virtual space, and when the user turns to the side, thereby similarly changing the orientation of the HMD, a view of that side of the virtual space is rendered on the HMD. The HMD can be worn in a manner similar to glasses, goggles, or a helmet and is configured to display video games or other metaverse content to the user. The HMD can provide a display mechanism in close proximity to the user's eyes, thereby creating a highly immersive experience for the user. Thus, an HMD can provide each of the user's eyes with a display area that occupies most, or even the entire, of the user's field of view, and can also provide viewing with three-dimensional depth and perspective.
[0074] In one embodiment, the HMD can include eye-tracking cameras configured to capture images of the user's eyes while the user interacts with the VR scene. Gaze information captured by the eye-tracking camera(s) can include information related to the user's gaze direction and particular virtual objects and content items within the VR scene that the user is focusing on or interested in interacting with. Thus, based on the user's gaze direction, the system can detect particular virtual objects and content items, such as game characters, game objects, game items, etc., that may be potential focal points for the user if the user is interested in interacting and engaging with them.
[0075] In some embodiments, the HMD may include outward-facing camera(s) configured to capture images of the user's real-world space, including the user's body movements, and images of any real-world objects that may be located in the real-world space. In some embodiments, images captured by the outward-facing cameras may be analyzed to determine the position / orientation of the real-world objects relative to the HMD. Using the known position / orientation of the HMD, inertial sensor data from the real-world objects and the user's gestures and movements may be continuously monitored and tracked while the user interacts with the VR scene. For example, while interacting with an in-game scene, the user may perform various gestures, such as pointing at or walking toward a particular content item in the scene. In one embodiment, the gestures may be tracked and processed by the system to generate a prediction of an interaction with a particular content item in the game scene. In some embodiments, machine learning may be used to facilitate or assist in the above predictions.
[0076] Various types of single-handed and dual-handed controllers may be used during use of the HMD. In some implementations, the controllers themselves can be tracked by tracking lights included on the controllers or by tracking shape, sensor, and inertial data associated with the controllers. These various controllers, or even simple hand gestures performed and captured by one or more cameras, can be used to interface, control, manipulate, interact with, and participate in the virtual reality environment or metaverse rendered on the HMD. In some cases, the HMD can be wirelessly connected to a cloud computing and gaming system via a network. In one embodiment, the cloud computing and gaming system maintains and executes the video game being played by the user. In some embodiments, the cloud computing and gaming system is configured to receive input from the HMD and interface objects over the network. The cloud computing and gaming system is configured to process the input to affect the game state of the running video game. Output, such as video data, audio data, and haptic feedback data, from the running video game is transmitted to the HMD and interface objects. In other implementations, the HMD may communicate with the cloud computing and gaming system wirelessly via alternative mechanisms or channels, such as a cellular network.
[0077] Additionally, while embodiments of the present disclosure may be described with reference to a head-mounted display, it will be understood that in other embodiments, a non-head-mounted display may be used instead, including, but not limited to, a portable device screen (e.g., a tablet, smartphone, laptop, etc.) or any other type of display that may be configured to render video and / or provide a display of an interactive scene or virtual environment in accordance with the present embodiments. It should be understood that the various embodiments defined herein may be combined or assembled into specific implementations using the various features disclosed herein. Thus, the examples provided are merely some of the possible examples and are not limiting of the various embodiments that may be defined by combining various elements. In some examples, an embodiment may include fewer elements without departing from the spirit of the disclosed or equivalent embodiments.
[0078] Embodiments of the present disclosure may be practiced with a variety of computer system configurations including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc. Embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a wire-based or wireless network.
[0079] Although the operations of the method are described in a particular order, it should be understood that other housekeeping operations may be performed between operations, or operations may be coordinated to occur at slightly different times, or operations may be distributed within the system to allow processing operations to occur at various intervals relative to processing, so long as the processing of telemetry and game state data to generate modified game states is performed in the desired manner.
[0080] One or more embodiments may also be fabricated as computer-readable code on a computer-readable medium. The computer-readable medium may be any data storage device that can store data, which can then be read by a computer system. Examples of computer-readable media include hard drives, network-attached storage (NAS), read-only memory, random-access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tape, and other optical and non-optical data storage devices. The computer-readable medium may include tangible computer-readable media distributed over network-coupled computer systems so that the computer-readable code is stored and executed in a distributed fashion.
[0081] In one embodiment, the video game is executed locally on a game console, personal computer, or server. In some cases, the video game is executed by one or more servers in a data center. When the video game is executed, some instances of the video game may be a simulation of the video game. For example, the video game may be executed by an environment or server that generates a simulation of the video game. A simulation, in some embodiments, is an instance of the video game. In other embodiments, the simulation may be generated by an emulator. In either case, when the video game is represented as a simulation, the simulation may be executed to render interactive content that can be interactively streamed, executed, and / or controlled by user input.
[0082] Although the foregoing embodiments have been described in some detail for clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims. Thus, the present embodiments are to be regarded as illustrative and not restrictive, and the present embodiments should not be limited to the details set forth herein but may be modified within the scope of the appended claims and their equivalents.
Claims
1. 1. A method of conducting a video game session, comprising: recording audio of players engaged in gameplay of said session of said video game; analyzing a game state generated by the execution of the session of the video game, wherein analyzing the game state identifies a context for the game play; predicting a future context of the gameplay from the identified context of the gameplay based on the game state; analyzing the recorded speech to identify textual content of the recorded speech using the predicted future context of the game play and a speech recognition model; applying the identified text content as gameplay input for the session of the video game; and Including, wherein analyzing the recorded audio includes dynamically adjusting the speech recognition model using the predicted future context of the game play, and wherein identifying the textual content includes applying the dynamically adjusted speech recognition model to the recorded audio to determine words or phrases in the recorded audio.
2. 2. The method of claim 1, wherein the method is performed in substantially real time, such that analyzing the game state is performed in substantially real time as the game state is continually updated by the execution of the session, and analyzing to identify the textual content of the recorded audio is performed in substantially real time in response to changes in the game state.
3. The method of claim 1 , wherein the identified context of the gameplay comprises an identified scene or stage of the video game.
4. The method of claim 1 , wherein the identified context of the gameplay includes an identified virtual location within a virtual environment defined by the execution of the session of the video game.
5. The method of claim 1 , wherein the identified context of the gameplay includes identified attributes of the player's avatar within the video game.
6. 2. The method of claim 1, wherein analyzing the game state includes recognizing activity occurring in the gameplay, the recognized activity at least partially defining the identified context of the gameplay.
7. 2. The method of claim 1 , wherein analyzing the game state includes predicting future gameplay activity, the predicted future gameplay at least partially defining the identified context of the gameplay.
8. 1. A non-transitory computer readable medium having program instructions embodied thereon, the program instructions, when executed by at least one computing device, causing the at least one computing device to perform a method of conducting a video game session, the method comprising the following operations: recording audio of players engaged in gameplay of said session of said video game; analyzing a game state generated by the execution of the session of the video game, wherein analyzing the game state identifies a context for the game play; predicting a future context of the gameplay from the identified context of the gameplay based on the game state; analyzing the recorded speech to identify textual content of the recorded speech using the predicted future context of the game play and a speech recognition model; applying the identified text content as gameplay input for the session of the video game; and Including, a game system configured to play a game using the gameplay context; a gameplay context configured to play a game using the gameplay context; a game system configured to play a game using the gameplay context; a game system configured to play a game using the gameplay context; a game system configured to play a game using the gameplay context;
9. 9. The non-transitory computer-readable medium of claim 8, wherein the method is performed substantially in real time, such that analyzing the game state is performed substantially in real time as the game state is continuously updated by the execution of the session, and wherein the analyzing to identify the textual content of the recorded audio is responsive to changes in the game state substantially in real time.
10. The non-transitory computer-readable medium of claim 8 , wherein the identified context of the gameplay comprises an identified scene or stage of the video game.
11. 10. The non-transitory computer-readable medium of claim 8, wherein the identified context of the gameplay includes an identified virtual location within a virtual environment defined by the execution of the session of the video game.
12. The non-transitory computer-readable medium of claim 8 , wherein the identified context of the gameplay includes identified attributes of the player's avatar within the video game.
13. 9. The non-transitory computer-readable medium of claim 8, wherein analyzing the game state includes recognizing activity occurring in the gameplay, the recognized activity at least partially defining the identified context of the gameplay.
Citation Information
Patent Citations
Game device, control method of game device, and program
JP2011206471A
Asynchronous audio for network games
JP2011509102A
Command processing program, image command processing apparatus, and image command processing method
JP2019063341A
Game program, character control program, method, and information processing device
JP2019205645A
Game system
JP2020156650A