Providing automated user input to application during disruption
By generating automatic user input based on a user interaction model, the technique addresses the issue of network interruptions in interactive applications, ensuring a seamless user experience during and after disruptions.
Patent Information
- Application Number
- JP2025010591
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-10-22
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-03
AI Technical Summary
Interactive applications, such as video games and streaming services, often experience disruptions due to network interruptions, leading to an inability to receive user input or provide application output, which disrupts the user experience.
The technique involves detecting interruptions and generating automatic user input using a user interaction model that predicts user behavior based on previous inputs and application outputs, allowing the application to continue functioning seamlessly during interruptions.
This approach provides a seamless user experience by allowing interactive applications to continue processing without user input during interruptions, ensuring that the application state is maintained and the user is reintroduced to the experience smoothly when the interruption ends.
Smart Images

Figure 2025084735000001_ABST
Abstract
Description
Background Art
[0001] Background
[0001] In many computer environments, interruptions are a problem. For example, when a user is playing a video game on a network or using another type of interactive application, if a network interruption occurs, the interactive application may become unable to respond to the user's input. Also, interruptions can be essentially non-technical, such as when a family member or friend interrupts the user. There are limited successful examples of automated efforts to address such interruptions.
Summary of the Invention
[0002] Summary
[0002] This summary is provided to introduce selected concepts in a simplified form, which will be further described in the following detailed description. This summary is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0003]
[0003] This description generally relates to techniques for handling interruptions where an application becomes unable to receive user input, the user becomes unable to provide input, and / or the user becomes unable to receive application output. An example includes a method or technique that can be executed on a computing device. The method or technique may include detecting an interruption to an interactive application during a user's interaction with the interactive application. The method or technique may also include generating an automatic user input and providing the automatic user input to the interactive application in response to detecting the interruption.
[0004]
[0004] Another example includes a system having a hardware processing unit and a storage resource storing computer-readable instructions. When executed by the hardware processing unit, the computer-readable instructions can cause the hardware processing unit to detect a network interruption that at least temporarily makes it impossible to receive one or more actual user inputs by a streaming interactive application. Also, the computer-readable instructions can cause the hardware processing unit to generate an automatic user input using a previously received actual user input. Further, the computer-readable instructions can cause the hardware processing unit to use the automatic user input in place of one or more actual user inputs to the streaming interactive application in response to detecting the network interruption.
[0005]
[0005] Another example is a computer-readable storage medium storing computer-readable instructions that, when executed by a hardware processing unit, cause the hardware processing unit to perform operations. The operations can include receiving a video output of an interactive application and receiving an actual user input to the interactive application provided by a user. The operations can also include detecting an interruption that makes it impossible to receive further actual user inputs by the interactive application. The operations can further include providing the video output and the actual user input of the interactive application to a prediction model and obtaining a predicted user input from the prediction model. The operations can also include providing the predicted user input to the interactive application during the interruption.
[0006]
[0006] The examples listed above are intended to provide a simple reference for assisting the reader and are not intended to define the scope of the concepts described herein.
[0007] Brief Description of the Drawings
[0007] The detailed description will be explained with reference to the accompanying drawings. In the figures, the leftmost digit of the reference number identifies the figure in which the reference number first appears. The use of the same reference number in different examples of the description or figures may indicate similar or identical items.
Brief Description of the Drawings
[0008]
Figure 1
[0008] Examples of game environments that conform to some implementations of this concept are shown.
Figure 2
[0009] Examples of timelines that conform to some implementations of this concept are shown.
Figure 3A
[0010] Examples of user experiences that conform to some implementations of this concept are shown.
Figure 3B
[0010] Examples of user experiences that conform to some implementations of this concept are shown.
Figure 3C
[0010] Examples of user experiences that conform to some implementations of this concept are shown.
Figure 3D
[0010] Examples of user experiences that conform to some implementations of this concept are shown.
Figure 3E
[0010] Examples of user experiences that conform to some implementations of this concept are shown.
Figure 4
[0011] Examples of process flows that conform to some implementations of this concept are shown.
Figure 5
[0012] Examples of user interaction models that conform to some implementations of this concept are shown.
Figure 6
[0013] Examples of systems that conform to some implementations of this concept are shown.
Figure 7
[0014] Examples of methods or techniques that conform to some implementations of this concept are shown.
Modes for Carrying Out the Invention
[0009] Detailed Description Summary
[0015] As described, for users of interactive applications (such as video games, augmented reality applications, or other applications) where the user frequently provides input to control the application, interruptions can be a problem. In the case of online applications, due to network interruptions between the user device and the application server, the application server may receive inputs too late and may not be able to effectively use those inputs to control the online application. In addition, due to network interruptions, the user device may receive the output of the application server, such as video or audio output, too late, and the user may not be able to effectively respond. In either case, it disrupts the application experience.
[0010]
[0016] One of the basic approaches to dealing with application interruptions is to continue using the most recent input received before the interruption during the interruption. However, this approach can have an adverse effect on the user's application experience. For example, the most recent input may result in negative or unexpected results in the application because the user did not have the opportunity to adjust their input in response to the application output generated during the interruption time.
[0011]
[0017] A more sophisticated alternative form may involve monitoring the internal application state and adjusting the internal application state during the interruption to provide a more fluid user experience. However, this approach can involve significant development efforts, such as modifying the internal application code to manage the interruption and / or providing hooks to the internal application state to external interruption management software.
[0012]
[0018] The disclosed implementations provide a technique for reducing application interruptions that address the above problems. In the disclosed implementations, during an application interruption, automatic user input is used instead of actual user input. Once the interruption ends, control can be returned to the user.
[0013]
[0019] In some implementations, the automatic user input can be generated without accessing the internal application state. For example, a user interaction model can use information such as application output and previously received user input to generate automatic user input, and the automatic user input generated by the user interaction model can be provided to the application during the interruption. As a result, the disclosed implementations can provide a seamless experience to the user during the interruption without requiring modification of the application code.
[0014]
[0020] In addition, the disclosed implementations can provide a realistic user experience by leveraging automatic user input that accurately reflects how a particular user would interact with a given application in the absence of an interruption. In comparison, as will be further discussed below, approaches that attempt to simulate optimal behavior rather than predict user behavior can lead to unrealistic results.
[0015] Technical terms
[0021] For the purposes of this specification, the term "application" refers to any type of executable software, firmware, or hardware logic for performing a specified function. The term "interactive application" refers to an application that performs processing in response to received user input and iteratively, frequently, or continuously adjusts application output in response to received user input. The term "online application" refers to an application that can be accessed over any type of computer network or communication link by streaming or downloading the application from one device to another. The term "streaming application" refers to an online application that runs on a first device and transmits a stream of application output to one or more other devices over a network or other communication link. The one or more other devices can reproduce the application output, for example, using a display or audio device, and can also provide user input to the streaming application.
[0016]
[0022] The term "interruption" refers to any situation that at least temporarily affects the user interaction with an interactive application or at least temporarily blocks the user interaction with an interactive application. For example, an interruption can be of a technical nature, such as a network interruption or other technical problem that prevents the interactive application from receiving actual user input and / or prevents the user from receiving application output. Common interruptions can include network conditions such as latency, bandwidth limitations, or packet loss. An interruption can also be non-technical in nature, for example, when the user is distracted by a conversation with a friend or family member, an incoming phone call, etc.
[0017]
[0023] The term "user interaction model" refers to any type of machine learning, discovery-based or rule-based approach that can be used to model user interaction with an application, for example, by generating automated user input. The term "actual user input" refers to the input actually provided by the user during the course of the interaction with the interactive application. The term "automated user input" refers to a machine-generated representation that can be used in place of actual user input during lulls. In some cases, the automated user input output by a given model can be used without modification, and in other cases, the automated user input output by the model can be smoothed or combined with previously received actual user input before being provided to the interactive application. Thus, the term "automated user input" encompasses both the unmodified output of the user interaction model and the output of the user interaction model that has been smoothed or combined with actual user input.
[0018]
[0024] The term "machine learning model" refers to any of a wide range of models that can learn to generate automatic user input by observing the characteristics of past interactions between a user and an application. For example, a machine learning model can be a neural network, a support vector machine, a decision tree, a clustering algorithm, etc. In some cases, a machine learning model can be trained using labeled training data, a reward function, or other mechanisms, and in other cases, a machine learning model can learn by analyzing data without explicit labels or rewards. The term "user-specific model" refers to a model that has at least one component that is at least partially trained or constructed for a specific user. Thus, this term encompasses models that are fully trained for a specific user, models that are initialized using multi-user data and adjusted for a specific user, and models that have both general components trained for multiple users and one or more components trained or adjusted for a specific user. Similarly, the term "application-specific model" refers to a model that has at least one component that is at least partially trained or constructed for a specific application.
[0019]
[0025] The term "neural network" refers to a type of machine learning model that uses layers of nodes to perform specific operations. In a neural network, the nodes are connected to each other via one or more edges. A neural network can include an input layer, an output layer, and one or more intermediate layers. Each individual node can process its respective inputs according to a predefined function (e.g., a transfer function such as ReLU or sigmoid) and provide the output of the function to subsequent layers (or, in some cases, previous layers). The inputs to a given node can be multiplied by the corresponding weight values of the edges between the input and the node. In addition, a node can have an individual bias value, which is also used to generate the output. Various training procedures can be applied to learn the edge weights and / or bias values.
[0020]
[0026] A neural network structure can have different layers that perform different specific operations. For example, one or more layers of nodes can collectively perform specific operations such as pooling, encoding, or convolutional operations. For the purposes of this specification, the term "layer" refers to a group of nodes that share inputs and outputs (e.g., from an external source or to / from other layers of the network). The term "operation" refers to a function that can be performed by one or more layers of nodes.
[0021] Example of a video game system
[0027] The following explains some specific examples of how this concept can be adopted in the context of a streaming video game where network interruptions can affect gameplay. However, as discussed elsewhere in this specification, this concept is not limited to video games, not limited to streaming or network-connected applications, and not limited to handling network interruptions. Rather, this concept can be adopted in a wide range of technical environments to address many different types of interruptions for various types of interactive applications.
[0022]
[0028] FIG. 1 shows an exemplary game environment 100 that conforms to the disclosed implementation. For example, FIG. 1 shows exemplary communications between an application server 102, an intermediary server 104, a client device 106, and a video game controller 108. The application server can execute an interactive application such as a streaming video game and generate an output 110, which can include video, audio, and / or haptic output. The application server can send the output to the intermediary server, and the intermediary server can forward the output to the client device 106.
[0023]
[0029] The client device 106 can display the video output on a display and reproduce the audio output through a speaker. In an example where the video game provides haptic output, the client device can transfer the haptic output to the video game controller 108 (not shown in FIG. 1). The video game controller can generate haptic feedback based on the haptic output received from the application server. In addition, the video game controller can generate an actual user input 112 based on user interaction with various input mechanisms of the video game controller. The video game controller can send the actual user input to the client device, and the client device can transfer the actual user input back to the intermediary server 104.
[0024]
[0030] When there is no network interruption, the mediation server 104 can simply operate as a pass-through server and transfer the actual user input 112 to the game server. However, when a network interruption is detected, the mediation server can instead provide the automatic user input 114 to the application server. As will be further described below, with the automatic user input, the application server can continue application processing in a way that reduces or eliminates user awareness of the interruption so that the user is seamlessly reintroduced to the application experience when the interruption ends.
[0025] Example of a timeline
[0031] Figure 2 shows an example of a timeline 200. The timeline 200 encompasses three times, namely, the time 202 that occurs before the network interruption, the time 204 during which there is a network interruption, and the time 206 that occurs after the network interruption. The timeline 200 shows how the disclosed implementations can be used to handle network interruptions, and in Figure 2, the network interruption is shown through the network state representation 208.
[0026]
[0032] At time 202, actual user input 210 is utilized to control the video game. The video game can generate video output 212, audio output 214, and / or haptic output 216 in response to actual controller input and / or the internal game state. These outputs are not necessarily provided at a fixed rate, but note that in some implementations, especially for video and audio, they can be provided at a specified frame rate. In many cases, especially for haptic output and user input, they are not synchronized.
[0027]
[0033] At time 204, the network is interrupted. For the purposes of this example, the interruption affects both the traffic flow to the user and the traffic flow from the user, and thus it is assumed that the user's device does not receive video game output and the application server does not receive actual user input. During the interruption, an automatic user input 218 is provided for use in place of actual user input that is unavailable due to the network interruption. As further discussed below, in some implementations, the automatic user input can be generated by a user interaction model that has access to the video game output during the network interruption. For example, the user interaction model can have local network access to the application server, the network interruption can occur on an external network, or the user interaction model can be executed on the same device as the video game.
[0028]
[0034] At time 206, the network recovers and further actual user input is received and provided to the game. As further discussed below, this allows the video game to seamlessly transition from network interruption to user control without causing confusion to the user. As discussed elsewhere in this specification, some implementations can smooth or combine the automatic user input with the actual user input to further reduce the impact of the network interruption on the user input.
[0029] Examples of user experience
[0035] FIGS. 3A-3E illustrate an exemplary user experience 300 of a user playing a driving video game. In FIG. 3A, a car 302 is shown moving along a road 304. FIG. 3A also shows a direction expression 310 and a trigger expression 320, which represent controller inputs to the driving game. Generally, the direction expression conveys the magnitude of a direction with respect to a direction input mechanism (e.g., a thumbstick for operating a car) of a video game controller. Similarly, the trigger expression 320 conveys the magnitude of a trigger input in a video game controller, e.g., for controlling a car's throttle. The thumbstick and trigger are two examples of “analog” input mechanisms that can be provided on a controller. The term “analog” is used simply to refer to an input mechanism that is not just on / off. For example, by varying the magnitude of an input to an analog input mechanism, a user can cause the analog input mechanism to generate a signal that can be represented digitally using a range of values greater than two.
[0030]
[0036] The direction expression 310 is shown by a received actual direction input 312 shown in black and an automatic direction input 314 shown in white. Similarly, the trigger expression 320 shows a received actual trigger input 322 in black and an automatic trigger input 324 in white. For the purposes below, the automatic inputs are assumed to be generated in the background by a user interaction model as the user plays the driving game.
[0031]
[0037] In FIG. 3A, since the user is gently operating the car 302 to the left at a mid-throttle opening, the car is being controlled via the actual direction input 312 and the actual trigger input 322. The automatic direction input 314 and the automatic trigger input 324 generated by the user interaction model are similar to the actual user inputs but are not being used for controlling the car at the current time.
[0032]
[0038] In FIGS. 3B and 3C, since the user continues to gently navigate the vehicle to the left, vehicle 302 continues to move forward on road 304 as it is, and the actual and automatic user inputs are somewhat different. Since there has been no interruption yet, the vehicle is still under the control of the actual user input.
[0033]
[0039] In FIG. 3D, an interruption has occurred. For the purposes of this example, assume that due to the interruption, the video game cannot receive actual user input, but the video output to the user is not hindered. The user needs to sharply curve vehicle 302 to the left and throttle down in order to correctly navigate the curve of road 304. However, the most recent actual user inputs are old. For example, even if the user adjusts their thumbstick and trigger inputs to sharply curve and throttle down, those actual user inputs are not being received by the video game. As a result, if the old user inputs are used, the magnitude of the leftward direction of the old direction input is too small to correctly navigate the curve, and the vehicle veers off the road as shown by ghost vehicle 330.
[0034]
[0040] At this point, automatic user input can be used instead of actual user input. For example, in FIG. 3D, automatic direction input 314 is used instead of the old most recently received user input, and the automatic direction input sharply curves the vehicle to the left, and automatic trigger input 324 throttles down to decelerate vehicle 302. Therefore, during the interruption, since the vehicle is being controlled by automatic user input, the vehicle continues to move forward on road 304 without veering off the road. Also, FIG. 3D shows the current actual user input 316 (hatched pattern), and the current actual user input 316 is affected by the interruption and is therefore not available for game play control.
[0035]
[0041] Figure 3E shows a video game after recovery from interruption. Since automatic user input was used to control vehicle 302 during the interruption, when control returns to the user, the vehicle is approximately located at the position on road 304 where the user expected it to be when the vehicle is positioned. To indicate that the user correctly made a sharp left turn, the received actual direction input 312 received from the user is moving to the left. Similarly, the received actual trigger input 322 received from the user is throttling back.
[0036]
[0042] The automatic user input used during the interruption was similar to the actual input provided by the user during the interruption, so the user was not very aware of the impact on the gameplay experience. In contrast, if old user input recently received had been used instead of the automatic input, the user might have experienced a collision as indicated by ghost vehicle 330 hitting tree 340.
[0037] Example of a model processing flow
[0043] Figure 4 shows an example of a processing flow 400 that can be employed to selectively provide actual or automatic input to application 402. In processing flow 400, actual input source 404 provides actual user input 406 to user interaction model 408. User interaction model 408 can generate automatic user input 412 using the actual user input and / or application output 410. Also, the application output can be provided to output mechanism 414 for purposes such as displaying an image or video, playing audio, generating haptic feedback, etc.
[0038]
[0044] The input adjudicator 416 can select the actual user input 406 or the automatic user input 412 to provide to the application 402 as the selected input 418. For example, if there is no interruption, the input adjudicator can provide the actual user input to the application as the selected input. If an interruption occurs, the input adjudicator can directly use the automatic user input instead of the actual user input, for example, by outputting the automatic user input as the selected input provided to the application. In other cases, the input adjudicator can smooth or combine the automatic user input provided by the user interaction model with the previously received actual user input and output the smoothed / combined automatic user input to the application.
[0039]
[0045] As further discussed below, the processing flow 400 can be employed in a wide range of technical environments. For example, in some implementations, each part of the processing flow is executed on a single device. In other cases, different parts of the processing flow are executed on different devices. FIG. 1 shows a specific example. For example, the application 402 can be on the application server 102, the actual input source 404 can be the video game controller 108, the output mechanism 414 can be provided by the client device 106, and the user interaction model 408 and the input adjudicator 416 can be provided by the mediation server 104.
[0040]
[0046] Also, it should be noted that the processing flow 400 does not necessarily rely on access to the internal application state to determine the automatic user input. As previously proposed, this can be beneficial because it allows the implementation of concepts disclosed without modifying the application code. However, in other cases, the user interaction model may have access to the internal application state and use the internal application state to determine the automatic user input. Generally, any internal game state can be adopted to determine the automatic user input, from raw data in the CPU or GPU memory to a specific data structure or an intermediate graphics pipeline stage. For example, a driving video game can provide many different simulated driving courses, and the user interaction model may have access to the identifier of the current driving course. As another example, a driving video game can provide many different vehicle models for the user to drive, and the user interaction model may have access to the identifier of the currently selected vehicle model. In an implementation where the user interaction model does not have access to the internal game state, the user interaction model may be able to infer the current driving course or vehicle model from the application output, for example, by analyzing the video output by the application.
[0041] Specific examples of user interaction models
[0047] As mentioned above, one way to implement the user interaction model 408 is to adopt machine learning techniques. FIG. 5 shows a prediction model 500 based on a neural network that can be adopted as a user interaction model consistent with this concept. The following describes a specific implementation of how a neural network-based prediction model can be adopted to predict user input to an interactive application. In the following description, reference numbers starting with "4" refer to elements previously introduced in FIG. 4, and reference numbers starting with "5" refer to elements newly introduced in FIG. 5.
[0042]
[0048] Actual user input 406 and application output 410 can be input for preprocessing 502. For example, when actual user input is provided via a video game controller having buttons and an analog input mechanism, the controller input can be preprocessed by representing the buttons as boolean values and normalizing the values of the analog inputs to a range of values from -1 to 1. Video or audio output can be preprocessed by reducing the resolution of the video and / or audio output, and haptic output can also be normalized, for example, to a range from -1 to 1.
[0043]
[0049] In some implementations, preprocessing 502 maintains respective windows for user input and application output. For example, the following example assumes a time window of one second. The preprocessing can include performing a temporal alignment of the actual user input with the corresponding frames of the video and / or audio data. In addition, haptic output can be temporally aligned with the actual user input. This process results in an input window 504 and an output window 506.
[0044]
[0050] Input window 504 can be input to a fully connected neural network 508, and output window 506 can be input to a convolutional neural network 510. The outputs of the fully connected neural network (e.g., a vector space representation of features extracted from the input window) and the convolutional neural network (e.g., a vector space representation of features extracted from the output window) can be input to a recurrent neural network 512, such as a long short-term memory ("LSTM") network.
[0045]
[0051] The recurrent neural network 512 can output an embedding 514, which can represent user input and application output in a vector space. The embedding can not only be returned to the recurrent neural network (as shown by arrow 513), but also be input into another fully connected network 516. Note that some implementations can employ multiple recurrent neural networks (e.g., a first recurrent neural network for processing the output of the fully connected network 508 and a second recurrent neural network for processing the output of the convolutional neural network 510).
[0046]
[0052] The fully connected network 516 can map from the embedding output by the recurrent neural network to the automatic user input 412. The embedding represents information about the user input and application output used by the fully connected network 512 to determine the automatic user input 412 (e.g., future predicted user input).
[0047]
[0053] In addition, the automatic user input 412 can be input into the preprocessing 502 and preprocessed as described above. Thus, at any given time step, the input to be preprocessed is, for example, the input selected for that time step by the input adjudicator 416. As a result, at any given time, the input window 504 can include all actual user inputs, all automatic user inputs (assuming a break at least as long as the input window), and / or combinations of both actual and automatic user inputs.
[0048] Examples of specific model processing
[0054] As described, the preprocessing 502 can be executed as a series of time steps. For example, a convenient interval can be for using a standardized video frame rate and can have one time step per video frame. At a 60 Hz video frame rate, 60 frames are obtained within a given one-second window. Each input window 504 can have a respective set of 60 actual or automatic user inputs, and each output window 506 can have a respective set of 60 application outputs. As described, the input window and the output window can be time-aligned with each other.
[0049]
[0055] Generally, actual or automatic user inputs and application outputs can be preprocessed to obtain features for input to the fully-connected neural network 508 and the convolutional neural network 510, respectively. Thus, in some cases, the preprocessing 502 can involve extracting features that may be useful for discrimination across the input window 504 and the output window 506 to accurately predict the user input for the next time step. For example, the preprocessing can include extracting uniqueness features from the actual or predicted user input for the current time step, where the uniqueness features indicate whether the user input has changed from the user input of the previous time step.
[0050]
[0056] Some user inputs (such as those provided by a video game controller) can be represented as vectors. For example, certain entries of the vector can represent different button states as boolean on / off values, and other entries of the vector can represent magnitudes for analog input mechanisms such as trigger states and directional pressure applied to thumbsticks. Also, the vector can include an entry indicating whether the value has changed from the state at the previous time. This vector representation can be used for both actual and predicted user inputs.
[0051]
[0057] In an implementation where user input and application output are discretized and time-aligned to a 60 Hz video frame rate using a 1-second window, each of the individual networks of the neural network-based prediction model 500 can process 60 stacked input / output vectors simultaneously, and each vector represents a time of 16.67 milliseconds. Also, the recurrent neural network 512 can maintain an internal recurrent state that can be used to represent the previous output of the recurrent network, and this internal recurrent state can contain information that is not present in the current input / output window. Thus, the recurrent neural network 512 can effectively model a sequence of input / output data that is longer than the duration of each time window provided by the preprocessing 502.
[0052]
[0058] Also, note that in some implementations, automatic user input 412 can be smoothed or combined with previous actual or predicted user input to reduce jitter and / or user perception that the game is not responding to the user's input. Also note that the neural network-based prediction model 500 can predict the final user input value or the delta relative to the last seen user input.
[0053] Model Training and Architecture
[0059] The following discussion provides details on how a model such as the neural network-based prediction model 500 can be trained to generate automatic user input. The following training techniques can be applied to various types of machine learning model types, including other neural network structures and model types other than neural networks.
[0054]
[0060] One way to train a model involves using imitation learning techniques. For example, in S. Ross, G. J. Gordon, and J. Bagnell, “A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,” in AISTATS, 2011, the DAgger algorithm is provided. DAgger presents an iterative approach for learning a policy to imitate a user, where the policy specifies predicted user actions (e.g., user input) as a function of previously received information. As another example, in S. Ross and J. Bagnell, “Reinforcement and Imitation Learning via Interactive No-Regret Learning,” in arXiv:1406.5979, 2014, the AggreVaTe algorithm is provided. AggreVaTe is an extended version of DAgger and uses a “cost-to-go” loss function rather than binary zero or one-class loss to inform training.
[0055]
[0061] In some implementations, separate models can be trained for each user and / or each application. This approach can ultimately provide high-quality models that can accurately predict user input, but generating new models from scratch for each user and each application can be time-consuming. In addition, this approach generally involves significant use of computational resources, such as storage resources for storing separate models and processor time for training separate models for each user. Moreover, this approach relies on obtaining large amounts of training data for each user and / or each application. As a result, users do not benefit from an accurate prediction model until after a large number of interactions with the application have ended. In other words, models developed using this approach do not “provide all the latest information” quickly for new users.
[0056]
[0062] Another high-level approach involves training a model on multiple users and then adapting some or all of that model to a particular user. To achieve this, several different techniques can be employed. One approach is to pre-train the model on many different users and then adapt the entire model to a new user, for example, by using a separate training epoch for the new user to adjust the pre-trained model. In this approach, each user ultimately arrives at a complete user-specific model for themselves, but user data from other users can be used to accelerate the training of the model. Alternatively, the model itself can have some common general components, such as one or more neural network layers trained on multiple users, and one or more user-specific components, such as one or more separate layers trained specifically for each user.
[0057]
[0063] Also, for example, a meta-learning approach can be adopted by starting with a set of weights learned from a large user group and then customizing those weights to fit a new user. Background information on meta-learning approaches can be found in the following literature, namely, Munkhdalai et al., “Rapid Adaptation with Conditionally Shifted Neurons,” in Proceedings of the 35 th International Conference on Machine Learning, pp. 1-12, 2018, Munkhdalai et al., “Meta Networks,” in Proceeding of the 34 th International Conference on Machine Learning, pp. 1-23, 2017, Koch et al., “Siamese Neural Networks for One-shot Image Recognition,” in Proceedings of the 32nd International Conference on Machine Learning, Volume 37, pp. 1-8, 2015, Santoro et al., “Meta-Learning with Memory-Augmented Neural Networks,” in Proceedings of the 33 rd International Conference on Machine Learning, Volume 48, pp. 1-9, 2016, Vinyals et al., “Matching Networks for One Shot Learning,” in Proceedings of the 30 th Conference on Neural Information Processing Systems, pp. 1-9, 2016 and Finn et al., “Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks,” in Proceedings of the 34 th can be found in the International Conference on Machine Learning, pp. 1-13, 2017. These techniques can be used to quickly adapt the disclosed techniques to new users, thus reducing and / or potentially eliminating examples where users are adversely affected by interruptions. In addition, these techniques can quickly adapt to changes in user skill and / or application difficulty, for example, by enabling the prediction model to adapt as the user becomes better at playing a video game and as the video game becomes more difficult in proportion to that.
[0058]
[0064] In a further implementation form, the model can be trained using auxiliary training tasks. For example, the model can be trained to predict which user generated a given input stream or to predict a given user's skill level from an input stream. A reward function can be defined for these goals, and the model obtains a reward for the correct identification of the user or the user's skill level that generated the input stream. By training the model for such auxiliary tasks using input data obtained from multiple users with different characteristics (e.g., skill levels), the model can learn how to distinguish different users. Then, when a new user starts a game play, the model thus trained may have the ability to infer that the new user is similar to one or more other users for whom the model has already been trained.
[0059]
[0065] Model training can use various data sources. For example, a given model can be trained offline based on previous examples of one or more users interacting with its application. In other examples, the model can be trained online while one or more users are interacting with the application. The actual user input received during a given application session can function as labeled training data for the model. Thus, when a given interactive application is executed and the user interacts with the application, the model can generate predicted user inputs, compare them with the actual user inputs, and propagate the error to the model to improve the model. In other words, the model can be trained to mimic the user while the user is interacting with the application.
[0060]
[0066] In some cases, after an interruption, it is possible to receive the actual user input that was intended during the interruption time. For example, packets having the actual user input can be transmitted late (e.g., after automatic user input has been used instead of the actual user input). The actual user input during the interruption time can be discarded for the purpose of controlling the application, but nevertheless, the actual user input during the interruption time can be used for training purposes.
[0061]
[0067] In addition, individual parts of a given model can be trained offline and / or separately from the rest of the model. For example, in some implementations, the convolutional neural network 510 of the prediction model 500 based on a neural network can be trained using video output even in the absence of user input. For example, the convolutional neural network can be trained using an auxiliary training task as described above or using a reconstruction or masking task. A pre-trained convolutional neural network can be inserted into a larger neural network structure together with other layers, and then the entire network structure can be trained collectively using the inputs and outputs obtained while one or more users are interacting with the application. This allows the untrained components of the neural network to learn more quickly than when training the convolutional neural network so that it learns from scratch together with the rest of the model.
[0062]
[0068] Also, the disclosed concepts can be used in a wide range of model architectures. Generally, the more contexts a model grasps, the more accurate the model's prediction of user input becomes. Thus, for example, by using a longer input / output window, generally, the accuracy of the prediction can be improved. On the other hand, by using a longer input / output window, proportionally, the amount of input data for the model increases, thereby increasing the training time and also potentially causing computational latency during execution. Some implementations can adopt alternative model structures to enable the model to grasp more contexts for the purpose of improving efficiency. For example, adopting a stack of dilated causal convolutional layers as discussed in Oord et al., “WaveNet: A Generative Model of Raw Audio,” arXiv:1609.03499 1-15, 2016, and / or adopting a hierarchy of autoregressive models as discussed in Mehri et al., “SampleRNN: An Unconditional End-to-End Neural Audio General Model,” in Proceedings of the 5th International Conference on Learning Representations, pp. 1-11, 2017.
[0063]
[0069] Also, it should be noted that in some implementations, rather than as a separate functionality, input adjudication can be employed within the prediction model. For example, the prediction model can learn when to use predicted user input instead of actual user input using additional features that describe, for example, network conditions (such as latency, bandwidth, packet loss, etc.). Similarly, the prediction model can learn to combine predicted user input with actual user input to reduce the impact of interruptions to the user. For example, the prediction model can be trained using feedback indicating a perceived user interruption and learn to smooth or combine actual user input and predicted user input to minimize or reduce the perceived interruption.
[0064] Input adjudication
[0070] In some implementations, one or more trigger criteria can be used to determine when to employ automatic user input. For example, referring back to FIG. 4, one or more trigger criteria can be used to determine whether input adjudicator 416 selects actual user input 406 or automatic user input 412. For example, one approach is to define a threshold time (e.g., 100 milliseconds) and always select automatic user input if no actual user input is received within that threshold time amount.
[0065]
[0071] Considering a network connection scenario as shown in FIG. 1, due to network interruption, the application server 102 may be unable to receive actual user input. Network interruption can occur due to various underlying reasons, such as the packet reception being too slow, the packet reception order being chaotic, and / or the packet being dropped on the network and not received at all. One way to handle network interruption is to substitute automatic user input when no packet is received from the client device by the mediation server 104 within the threshold time. In some cases, the client can send a heartbeat signal at regular intervals (e.g., 10 milliseconds) so that the mediation server can distinguish between a scenario where the user is simply not handling the controller 108 and network interruption.
[0066]
[0072] In a further implementation form, the threshold time can be adjusted for different users and / or different applications. For example, the user sensitivity to interruption may vary based on user skills or the type of application. Specifically considering video games, experienced users tend to notice shorter interruptions than novice users, or, in relation to this, players of a certain video game may tend to notice shorter interruptions (e.g., 50 milliseconds) that players of another video game do not notice.
[0067]
[0073] In other implementation forms, the threshold is related to the video frame rate. For example, the threshold can be defined as three frames or approximately 50 milliseconds at a frame rate of 60 Hz. A further implementation form can learn the threshold as part of the user interaction model for different applications and / or different users.
[0068]
[0074] Additionally, it should be noted that the term "network interruption" can encompass scenarios where traffic is not completely blocked. For example, network "jitter" or variations in the time delay of data packets can significantly impact the user experience with interactive applications. In some cases, automatic user input can be adopted when jitter exceeds a predefined threshold (e.g., 50 milliseconds). Jitter can be detected using timestamps associated with actual user input.
[0069]
[0075] It should be noted that network interruption can be bidirectional (i.e., it can affect both the flow of actual user input to the application and the flow of application output to the user device). However, in other cases, network interruption may only affect the traffic flow in a specific direction. For example, due to network interruption, actual user input may become unreceivable without affecting the flow of application output to the user. Conversely, due to network interruption, application output may become unreachable to the user without affecting the flow of actual user input to the application.
[0070]
[0076] Automated user input can be used to handle any type of network interruption. For example, while actual user input is not affected by a network interruption, consider a scenario where due to a network interruption, application output cannot reach the user for a threshold amount of time. Even if actual user input is available for this type of network interruption, in some scenarios, there may still be value in using automated user input instead of actual user input received during the interruption. Since the user does not receive application output during the interruption, the user does not have the opportunity to adjust that input to change the application output. In this case, the automated user input may more accurately reflect what the user would have provided if the application output had not been prevented from reaching the user due to the interruption.
[0071]
[0077] In some implementations, the input adjuster 416 can also perform a correction operation after the interruption ends. For example, the input adjuster uses automated user input instead of some previous actual user inputs, and then assumes that those actual user inputs arrive late after the interruption ends. Further, assume that the automated user input steered the vehicle at a steeper angle than indicated by the late-arriving actual user input. In some implementations, this can be corrected by adjusting the later-received actual user input to steer the vehicle more gently after a period of time has passed after the interruption, in order to compensate for the difference between the automated user input and the late-arriving actual user input. By doing so, the position of the vehicle can be brought closer to where the vehicle would be located if there had been no interruption.
[0072] Alternative implementations
[0078] As described above, this concept has been mainly illustrated using examples targeting specific types of interactive applications (e.g., streaming video games), specific types of interruptions (e.g., network interruptions), and specific types of automatic user inputs (e.g., predictive controller inputs generated using machine learning models). However, as will be further discussed below, this concept can be adopted for many different types of interactive application types to handle many different types of interruption types using many different types of automatic user inputs.
[0073]
[0079] Taking the application type first, video games are just one example of interactive applications. Other examples include virtual or augmented reality applications, educational or training applications, etc. For example, consider a multi-user driving simulator where a user learns how to drive a vehicle. Similar to video games, during an interruption of the driving simulator, it may be preferable to predict the actual user behavior rather than the optimal user behavior. This can provide a more realistic experience (e.g., two drivers both making mistakes that lead to a simulated accident). In the event of an interruption, if a model for generating optimal inputs rather than a model for predicting realistic user inputs is used, collisions can be avoided, but it will give the user a false sense of security and reduce the training value of the simulation.
[0074]
[0080] As another example, consider an automated inventory management application that continuously places orders to replenish products. In some cases, due to an interruption, the inventory management application may be unable to learn about user orders for an extended period (e.g., overnight). By using predicted user orders to simulate how the inventory changes during the interruption, such an application can mitigate the impact of the interruption, for example, by ensuring that sufficient parts are available in the factory so that production can continue based on predicted user orders rather than actual user orders.
[0075]
[0081] As yet another example, consider, for example, a semi-autonomous robot for disaster relief or medical applications. Generally, such robots can be controlled by a user located remotely. In some cases, such robots may operate remotely with spotty satellite signal reception and can utilize a local prediction model to take over the role of the remotely located user in the event of a network interruption. In some cases, this type of application can benefit from using a model trained for optimal behavior rather than simulating a particular user.
[0076]
[0082] In addition, the disclosed implementations can be employed to handle other types of interruptions besides network interruptions. For example, some implementations can detect that a user has been distracted from an interactive application by, for example, detecting an incoming message or call, using the gyroscope or accelerometer of the client device, or simply by the lack of input to an input mechanism for a threshold amount of time. Any of these interruptions can be mitigated by providing automatic user input to the interactive application during the interruption.
[0077]
[0083] In addition, the disclosed implementations can be used to enhance the experience for users who may have physical disabilities. Consider users who have difficulty keeping a joystick or other input mechanism in a stable state. Some implementations can smooth or combine the user's actual input with predicted user input to treat tremors or other shaky movements as interruptions and reduce the impact of the user's disability on their application experience.
[0078]
[0084] As another example, consider interruptions within a particular device. For example, consider a single physical server that runs a first virtual machine and a second virtual machine in different time slices. When the first virtual machine is running an interactive application, there may be user input received during a given time slice when the interactive application is not currently active in the processor. With previous techniques, it may have been possible to buffer the actual user input in memory until the context was switched from the second virtual machine to the first virtual machine. By using the disclosed techniques, the actual user input that occurred during the time slice when the first virtual machine was not active can be discarded, an automatic user input can be generated, and provided to the interactive application when the first virtual machine begins to operate on the processor. Thus, since a buffer for holding the actual user input during that time slice is not needed when the first virtual machine is not active, memory can be saved.
[0079]
[0085] As another example, automatic user input can be used for compression purposes. For example, in the example shown in FIG. 1, the client device 106 can discard certain user input (e.g., skip every other user input and not send them to the mediation server 104). The mediation server can use automatic user input instead of every other actual user input. This approach can provide a satisfactory user experience while saving network bandwidth.
[0080]
[0086] As another compression example, instead of completely discarding actual user input, some implementations can discard a portion of certain actual user input. For example, a certain number (e.g., 3) of the most significant bits of the actual user input can be sent to the mediation server, and the remaining bits (e.g., 13 bits for a 16-bit input) can be filled in by the user interaction model 408 at the mediation server. Further, another compression example can employ another example of a prediction model on the user device. The user device can calculate the difference between each actual user input and the corresponding predicted input and send that difference to the mediation server. Since the mediation server is running the same prediction model, the mediation server can derive the actual user input from that difference. Since the difference between the predicted user input and the actual user input is often relatively small, the transmission of this information over the network requires fewer bits than the complete actual user input. The compression techniques described above can be considered a mechanism for intentionally introducing pre-programmed discontinuities to reduce the amount of network traffic.
[0081]
[0087] In some cases, pre-programmed discontinuities can be used to not waste battery life (e.g., on a video game controller). For example, the video game controller can sleep for a specified amount of time (e.g., 50 milliseconds) and then wake up, detecting input and sending a new set of controls each time it wakes up. The mediation server 104 can use automatic input to fill in any gaps during the sleep interval.
[0082]
[0088] As another example, the automatic user input can be used for preemptive scheduling or loading of application code. For example, assume that when a user performs a specific action (such as succeeding in achieving a goal in a video game or crashing a car in a driving simulation), a predetermined application module is loaded into memory. Some implementations can predict whether the user's future input might execute that action. If so, this can be used as a trigger for an interactive application to begin loading that module into memory or scheduling its execution before receiving the user's actual input. In this case, the automatic user input is not necessarily used to control the application, but rather is used to selectively perform certain processing in advance, such as loading code or data from storage to memory or notifying the operating system of scheduling decisions. In some cases, the user interaction model can maintain a future window of predicted user input (e.g., 500 milliseconds) and use the predicted user input within that window to perform selective loading and / or scheduling of application code.
[0083]
[0089] Moreover, the disclosed implementations can be employed in scenarios other than the client / server scenarios described above. For example, a peer-to-peer game can involve communication between two peer devices. When communication between peer devices is disrupted (e.g., when a device temporarily goes out of communication range in the case of a short-range wireless connection such as Bluetooth), automatic user input can be employed to provide a satisfactory user experience. Thus, in this example, the trigger criteria for selecting automatic user input rather than actual user input can be related to characteristics of the short-range wireless connection such as signal strength or throughput. Also, note that both users can be aware of each other's predicted versions during the time the peer-to-peer game is interrupted. This approach can improve the experience of users other than the user whose input is being predicted. For example, if predicted user input is substituted for the first user during an interruption, from the perspective of the second user watching the first user's representation in a peer-to-peer game, the interruption is mitigated.
[0084]
[0090] Also, the disclosed implementations can be employed on a single device. For example, a user interaction model can be provided on a game console or other device. An input adjudicator can be provided on the device to enable the interactive application to select whether to receive actual user input or automatic user input. For example, in some implementations, the input adjudicator can use criteria such as disk throughput, memory usage, or processor usage to recognize instances where there may be a delay in user input to the interactive application and can provide automatic user input to the interactive application in those instances.
[0085]
[0091] Also, as described, while some implementations can use a machine learning user interaction model, other implementations can utilize other types of user interaction models. For example, in a driving game, the user interaction model can provide pre-computed directions or throttle inputs for each location along the driving course, and the pre-computed inputs are static and not user-specific. Those pre-computed inputs can be used in place of actual user inputs during interruptions. Similar to the use of a prediction model, the user's previous inputs can be smoothed or combined with the default user inputs for a given location on the course at runtime.
[0086]
[0092] In addition, some implementations can predict not only user inputs but also application outputs. For example, recall that network interruptions are not always two-way. Previously, it was described that a substitute for automatic user input can be used to mitigate network interruptions where the user is unable to receive application outputs. Alternatively, some implementations can predict application outputs and provide the predicted application outputs to the user during an interruption. In this scenario, during the interruption, those inputs can be provided to the user in response to the predicted application outputs rather than in response to a "frozen" application output, so actual user inputs can be used.
[0087] Examples of systems
[0093] This concept can be implemented in various technical environments and on various devices. FIG. 6 shows an example of a system 600 that can adopt this concept, as will be further discussed below. As shown in FIG. 6, the system 600 includes the application server 102, the mediation server 104, the client device 106, and the video game controller 108, which were introduced above with respect to FIG. 1. Also, the system 600 includes a client device 610. The client device 106 is connected to the video game controller 108 via a local wireless link 606. Both the client device 106 and the client device 610 are connected to the mediation server 104 via a wide area network 620. The mediation server 104 is connected to the application server 102 via a local area network 630.
[0088]
[0094] Certain components of the client devices and servers shown in FIG. 6 may be referred to in this specification by reference numbers in parentheses. For the purposes of the following description, the reference number in parentheses (1) indicates the presence of a predetermined component on the client device 106, (2) indicates the presence of a predetermined component on the client device 610, (3) indicates the presence on the mediation server 104, and (4) indicates the presence on the application server 102. Unless a specific example of a predetermined component is identified, this specification generally refers to components without parentheses.
[0089]
[0095] Generally, the devices shown in FIG. 6 may each have respective processing resources 612 and storage resources 614, which will be discussed in further detail below. Also, the devices may have various modules that function to execute the techniques discussed in this specification using the processing and storage resources, as will be further discussed below.
[0090]
[0096] The video game controller 108 may include a controller circuit 602 and a communication component 604. The controller circuit can digitize inputs received by various controller mechanisms such as buttons or analog input mechanisms. The communication component can transmit the digitized inputs to the client device 106 over a local wireless link 606. The interface module 616 on the client device 106 can obtain the digitized inputs and transmit those inputs to the mediation server 104 over a wide area network 620.
[0091]
[0097] The client device 610 can provide a controller emulator 622. The controller emulator can display a virtual controller on the touch screen of the client device. Inputs received via the virtual controller can be transmitted to the mediation server 104 over a wide area network.
[0092]
[0098] The mediation server 104 can receive packets having actual controller inputs from the client device 106 or the client device 610 over the wide area network 620. The input adjudicator 416 on the mediation server can determine whether to transmit those actual controller inputs to the application server 102 over a local area network 630 or to transmit automatic user inputs generated through the user interaction model 408. For example, the input adjudicator can use automatic user inputs instead of actual user inputs during a network outage and abort the substitution by providing the subsequently received actual user inputs to the application server when the network outage is resolved.
[0093]
[0099] The interactive application 642 can process the received actual or automatic user input and generate corresponding application output. The interactive application can send the output to the mediation server 104 on the local area network 630, and the mediation server 104 can use the output to generate further automatic user input. In addition, the mediation server can transfer the output to the client device 106 or the client device 610 on the wide area network 620.
[0094]
[0100] The client device 106 and / or the client device 610 can display video output and reproduce audio output. When the application output includes haptic feedback, the client device 106 can send a haptic signal to the video game controller 108 on the local wireless link 606, and the video game controller can generate a haptic output based on the received signal. When the output provided to the client device 610 includes a haptic signal, the client device 610 can generate a haptic output by itself through the controller emulator 622.
[0095] Example of a method
[0101] FIG. 7 shows an example of a method 700 that can be used to selectively provide actual or automatic user input to an application consistent with this concept. As discussed elsewhere in this specification, the method 700 can be implemented on many different types of devices, for example, by one or more cloud servers, by client devices such as laptops, tablets or smartphones, or by a combination of one or more servers, client devices, etc.
[0096]
[0102] Method 700 begins at block 702, where actual user input is received. For example, the actual user input can be received via a client device, via a dedicated controller such as a video game or virtual reality controller, or by any other suitable mechanism for providing user input to an application.
[0097]
[0103] Method 700 proceeds to block 704, where application output is received. As described, the application output can include video, audio, and / or haptic feedback generated by an interactive application.
[0098]
[0104] Method 700 proceeds to block 706, where automatic user input is generated. For example, the automatic user input can be generated by a user interaction model when the user is interacting with the application. Alternatively, the automatic user input can be generated prior to user interaction, for example, by pre-computing the automatic user input.
[0099]
[0105] Method 700 proceeds to decision block 708, where a determination is made as to whether an interruption has occurred. As described, an interruption can include a network interruption, an incoming call, the user's inactive time, a pre-programmed interruption, etc.
[0100]
[0106] If no interruption is detected at block 708, method 700 proceeds to block 710, where the actual user input is provided to the application. Method 700 returns to block 702, and another iteration of the method can be executed.
[0101]
[0107] If an interruption is detected at block 708, method 700 proceeds to block 712, where automatic user input is provided to the application. Method 700 returns to block 704 and additional application output is received. Method 700 proceeds to block 706, where additional automatic user input is generated, for example, using the actual user input received before the interruption, the application output received before or during the interruption, and / or the automatic user input generated before or during the interruption. Method 700 can repeat through blocks 704, 706, 708, and 712 until the interruption ends.
[0102]
[0108] Note that in some implementations, block 706 can be executed by a user interaction model that runs in the background during normal operations of the interactive application. In other implementations, the user interaction model can be activated during an interruption and is not active during normal operations.
[0103]
[0109] Also, recall that some interruptions can affect both actual user input and application output, while other interruptions can affect only one of them, not both application output and actual user input. The description of method 700 above corresponds to an interruption that affects actual user input rather than application output. However, as previously described, automatic user input can also be provided to the application during other types of interruptions.
[0104] Device Implementations
[0110] As described above with respect to FIG. 6, system 600 includes several devices including client device 106, client device 610, intermediary server 104, and application server 102. Also, not all device implementations are shown, and from the above and below descriptions, other device implementations should be apparent to those skilled in the art.
[0105]
[0111] As used herein, the terms "device", "computer", "computing device", "client device" and / or "server device" can mean any type of device having some hardware processing capabilities and / or hardware storage / memory capabilities. The processing capabilities can be provided by one or more hardware processors (e.g., hardware processing units / cores) capable of executing computer-readable instructions to provide functionality. The computer-readable instructions and / or data can be stored in storage resources. As used herein, the term "system" can refer to a single device, multiple devices, etc.
[0106]
[0112] The storage resources can be internal or external to their respective associated devices. The storage resources can include any one or more of, among others, volatile or non-volatile memory, hard drives, flash storage devices, and / or optical storage devices (e.g., CDs, DVDs, etc.). In some cases, the modules of system 600 are provided as executable instructions, which are stored in a persistent storage device, loaded into a random access memory device, and read from the random access memory by processing resources for execution.
[0107]
[0113] As used herein, the term "computer-readable medium" can include a signal. In contrast, the term "computer-readable storage medium" excludes a signal. The computer-readable storage medium includes a "computer-readable storage device". Examples of computer-readable storage devices include, among others, volatile storage media such as RAM and non-volatile storage media such as hard drives, optical disks, and flash memory.
[0108]
[0114] In some cases, the device is configured to have a general-purpose hardware processor and storage resources. In other cases, the device may include a system-on-chip (SOC) type of design. In an implementation of the SOC design, the functionality provided by the device can be integrated on a single SOC or multiple combined SOCs. One or more associated processors can be configured to interact with shared resources (such as memory, storage, etc.) and / or one or more dedicated resources (such as hardware blocks configured to perform a particular functionality). Thus, the terms "processor", "hardware processor" or "hardware processing unit", as used herein, can also refer to a central processing unit (CPU), a graphics processing unit (GPU), a controller, a microcontroller, a processor core, or other types of processing devices suitable for implementation in both conventional computing architectures and SOC designs.
[0109]
[0115] Alternatively or in addition thereto, the functionality described herein can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system-on-chip systems (SOC), programmable complex logic devices (CPLD), and the like.
[0110]
[0116] In some configurations, any module / code discussed herein can be implemented in software, hardware, and / or firmware. In any case, the module / code can be provided by the manufacturer of the device or by an intermediary preparing the device for sale to the end user during the manufacture of the device. In other examples, the end user can later install these modules / codes, such as by downloading the executable code and installing the executable code on the corresponding device.
[0111]
[0117] It should also be noted that a device can generally have input and / or output functionality. For example, a computing device can have various input mechanisms such as a keyboard, mouse, touchpad, voice recognition, gesture recognition (e.g., using a depth camera such as a stereo or time-of-flight camera system, an infrared camera system, an RGB camera system, or using an accelerometer / gyroscope, face recognition, etc.). A device can also have various output mechanisms such as a printer, monitor, etc.
[0112]
[0118] It should also be noted that the devices described herein can function stand-alone or in conjunction to implement the techniques described. For example, the methods and functionality described herein can be executed on a single computing device and / or distributed across multiple computing devices communicating on networks 620 and / or 630.
[0113]
[0119] In addition, some implementations can employ any of the disclosed techniques in the context of the Internet of Things (IoT). In such implementations, a household appliance or a vehicle can provide computing resources for implementing the modules of system 600. Also, as mentioned, some implementations can be used for virtual or augmented reality applications by performing some or all of the functionality disclosed on a head-mounted display, a handheld virtual reality controller, or by using techniques such as depth sensors to obtain user input through physical gestures performed by the user.
[0114]
[0120] In the above, examples of various devices have been described. Below, additional examples will be described. One example includes a method executed by a computing device, the method including generating automatic user input for an interactive application, detecting an interruption to the interactive application during a user's interaction with the interactive application, and providing the automatic user input to the interactive application in response to the detection of the interruption.
[0115]
[0121] Another example may include any of the above and / or below examples, and the method further includes generating a user interaction model for a user's interaction with the interactive application and generating automatic user input with the user interaction model.
[0116]
[0122] Another example may include any of the above and / or below examples, and the user interaction model can be a user-specific model for the user.
[0117]
[0123] Another example may include any of the above and / or below examples, and the user interaction model can be an application-specific model for the interactive application.
[0118]
[0124] Another example may include any of the above and / or below examples, and the method further includes obtaining actual user input before interruption, obtaining the output of an interactive application before interruption, and inputting the actual user input and output into a user interaction model.
[0119]
[0125] Another example may include any of the above and / or below examples, and the user interaction model includes a machine learning model.
[0120]
[0126] Another example may include any of the above and / or below examples, and the method further includes training a machine learning model using discarded actual user input, where the discarded actual user input is provided by the user during the interruption and received after the interruption ends.
[0121]
[0127] Another example may include any of the above and / or below examples, and the method is executed without providing the application state inside the interactive application to the user interaction model.
[0122]
[0128] Another example may include any of the above and / or below examples, and the interactive application includes an online application.
[0123]
[0129] Another example may include any of the above and / or below examples, and the interruption includes a network interruption.
[0124]
[0130] Another example may include any of the above and / or below examples, and the method further includes detecting a network interruption when a packet having actual user input is not received during a threshold time.
[0125]
[0131] Another example may include any of the above and / or below examples, and the online application includes a streaming video game.
[0126]
[0132] Another example may include any of the above and / or below examples, and the automatic user input includes predicted controller input for a video game controller.
[0127]
[0133] Another example may include any of the above and / or below examples, and the predicted controller input is used in place of actual user input provided by an analog input mechanism on a video game controller.
[0128]
[0134] Another example includes a system including a hardware processing unit and a storage resource storing computer-readable instructions, which, when executed by the hardware processing unit, cause the hardware processing unit to detect a network interruption that affects the reception of one or more actual user inputs by a streaming interactive application, generate automatic user input using previously received actual user inputs to the streaming interactive application, and use the automatic user input in place of one or more actual user inputs to the streaming interactive application in response to the detection of the network interruption.
[0129]
[0135] Another example may include any of the above and / or below examples, and the computer-readable instructions, when executed by the hardware processing unit, cause the hardware processing unit to transfer actual user inputs received at a computing device running a streaming interactive application when there is no network interruption, and transfer automatic user inputs to the computing device running the streaming interactive application during the network interruption.
[0130]
[0136] Another example may include any of the above and / or below examples, and the computer-readable instructions, when executed by the hardware processing unit, cause the hardware processing unit to generate automatic user input using the output of a streaming interactive application.
[0131]
[0137] Another example may include any of the above and / or below examples, and when the computer-readable instructions are executed by a hardware processing unit, the hardware processing unit is made to detect that the network interruption has been resolved and that further actual user input has been received, and in response to detecting that the network interruption has been resolved, to stop using automatic user input in place of further actual user input.
[0132]
[0138] Another example includes a computer-readable storage medium storing computer-readable instructions that, when executed by a hardware processing unit, cause the hardware processing unit to receive a video output of an interactive application, receive actual user input to the interactive application provided by a user, detect an interruption that affects the receipt of further actual user input by the interactive application, provide the video output and actual user input of the interactive application to a prediction model, obtain predicted user input from the prediction model, and provide the predicted user input to the interactive application during the interruption.
[0133]
[0139] Another example may include any of the above and / or below examples, and the operations further include detecting that the interruption has ended and, in response to detecting that the interruption has ended, providing subsequent received actual user input to the interactive application.
[0134] Conclusion
[0140] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as examples of forms of implementing the claims, and other features and acts that would be recognized by one of ordinary skill in the art are intended to be within the scope of the claims.
Claims
1. A hardware processing unit; a storage resource for storing computer readable instructions; wherein the computer readable instructions, when executed by the hardware processing unit, Detecting a network disruption affecting receipt of one or more actual user inputs by a streaming interactive application; generating automatic user input using previously received actual user input to the streaming interactive application; in response to detecting the network disruption, substituting the one or more actual user inputs to the streaming interactive application. The system further comprises:
2. The computer readable instructions, when executed by the hardware processing unit, forwarding received actual user input to a computing device running the streaming interactive application in the absence of the network disruption; and transferring the automated user input to the computing device running the streaming interactive application during the network disruption; and The system of claim 1 , further comprising:
3. The computer readable instructions, when executed by the hardware processing unit, generating said automated user input using output of said streaming interactive application; The system of claim 1 , further comprising:
4. The computer readable instructions, when executed by the hardware processing unit, detecting that the network outage has been resolved and that further actual user input has been received; in response to detecting that the network disruption has been resolved, ceasing to substitute the automated user input for the further actual user input; The system of claim 1 , further comprising:
5. The system of claim 1 , wherein the streaming interactive application is a video game.
6. The system of claim 5 , wherein the automated user input comprises predictive controller input for a video game controller.
7. The computer readable instructions, when executed by the hardware processing unit, generating said automated user input using a user interaction model of a user interaction with said video game; The system of claim 5 , further comprising:
8. The computer readable instructions, when executed by the hardware processing unit, generating said user interaction model; The system of claim 7 , further comprising:
9. The system of claim 8 , wherein the user interaction model is a user-specific model.
10. The system of claim 9 , wherein the user interaction model is an application-specific model for the video game.
11. The computer readable instructions, when executed by the hardware processing unit, obtaining actual user input prior to said network disruption; obtaining an output of the video game prior to the network disruption; inputting the actual user inputs and outputs into the user interaction model to generate the automatic user inputs; The system of claim 8 , further comprising:
12. the user interaction model is a machine learning model, and the computer readable instructions, when executed by the hardware processing unit, training the machine learning model using discarded actual user input, the discarded actual user input being provided by a user during the network disruption and received after the disruption ends. The system of claim 8 , further comprising:
13. A computer-readable storage medium storing computer-readable instructions that, when executed by the hardware processing unit, receiving a video output of an interactive application; receiving actual user input to the interactive application provided by a user; detecting a disruption affecting receipt of further actual user input by the interactive application; providing the video output and the actual user input of the interactive application to a predictive model; obtaining predicted user inputs from the predictive model; and providing the predictive user input to the interactive application during the disruption; and 23. A computer-readable storage medium for causing the hardware processing unit to perform operations including:
14. The operation, detecting that the disruption has ended; and in response to detecting that the disruption has ended, providing subsequently received actual user input to the interactive application; The computer-readable storage medium of claim 13 , further comprising:
15. The operation, smoothing said subsequently received actual user input with said predicted user input. The computer-readable storage medium of claim 14 , further comprising:
Citation Information
Patent Citations
Method and apparatus for continuing an electronic multiplayer game in the absence of a player
JP2006509548A