Multi-Task Machine Learning for Predicting Touch Interpretation

Through multi-task machine learning systems and machine learning touch explanation prediction models, the limitations of determining touch points in the prior art are solved, and more accurate and customizable touch explanations are achieved, and the responsiveness and interaction efficiency of computing devices are improved.

CN110325949BActive Publication Date: 2025-06-17GOOGLE LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN201780087180.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-12-29
Filing Date
2017-09-28
Publication Date
2025-06-17
Estimated Expiration
2037-09-28

AI Technical Summary

Technical Problem

When determining touch points, existing computing devices use limited fixed processing rules, cannot adapt to new technologies and touch modes of different users, and discard a large amount of original touch sensor data to prevent further processing.

Method used

Using a multi-task machine learning system, the machine learning touch interpretation prediction model receives touch sensor data and outputs multiple predicted touch explanations, including intention touch points, gesture explanations and touch prediction vectors, and uses commonalities and differences to improve prediction accuracy.

Benefits of technology

Achieve more accurate and customizable touch interpretation, reduce touch delay, improve responsiveness and interaction efficiency, and can adapt to the needs of different users and technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110325949B_ABST
    Figure CN110325949B_ABST
Patent Text Reader

Abstract

The present disclosure provides systems and methods for predicting multiple touch interpretations using machine learning. Specifically, the systems and methods of the present disclosure can include and use a machine learning touch interpretation prediction model, where the model has been trained to receive touch sensor data indicative of one or more positions of one or more user input objects relative to a touch sensor at one or more times, and to provide one or more predicted touch interpretation outputs in response to receiving the touch sensor data. Each predicted touch interpretation output corresponds to a different type of predicted touch interpretation based at least in part on the touch sensor data. The predicted touch interpretations can include a set of touch point interpretations, gesture interpretations, and / or touch prediction vectors for one or more future times.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to machine learning. More specifically, this disclosure relates to systems and methods for using multi-task machine learning to determine touch points and other touch interpretations. Background Art

[0002] A user may provide user input to a computing device using a user input object such as, for example, one or more fingers, a stylus operated by the user, or other user input object. Specifically, in one example, a user may use the user input object to touch a touch-sensitive display screen or other touch-sensitive component. The interaction of the user input object with the touch-sensitive display screen enables the user to provide user input to the computing device in the form of raw touch sensor data.

[0003] In some existing computing devices, simple heuristics may be used on a digital signal processor associated with a touch sensor to directly interpret touch sensor data as zero, one, or more "touch points". Traditional analysis used to determine whether touch sensor data results in a touch point determination may limit the types of possible interpretations. In some examples, touch point determination in such existing computing devices uses a limited number of fixed processing rules to analyze touch sensor data. The processing rules sometimes cannot be modified to adapt to new technologies, nor can they be customized for specific touch patterns for different users. Additionally, any additional analysis of the determined touch points involves subsequent use of additional processing rules. Further, touch point determination in existing computing devices discards a large amount of raw touch sensor data after touch point determination, thus precluding the possibility of further processing the raw touch sensor data. Summary of the Invention

[0004] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned by practice of the embodiments.

[0005] One example aspect of the present disclosure is directed to a computing device that determines a touch interpretation from user input objects. The computing device includes at least one processor, a machine learning touch interpretation prediction model, and at least one computer-readable medium (which may be a tangible, non-transitory computer-readable medium or tangible, non-transitory computer-readable media, although the aspect is not limited thereto), the at least one computer-readable medium storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations. The touch interpretation prediction model has been trained to receive touch sensor data indicative of one or more positions of one or more user input objects relative to a touch sensor at one or more times, and to output one or more predicted touch interpretations in response to receiving the touch sensor data. The operations include obtaining a first set of touch sensor data indicative of one or more user input object positions relative to the touch sensor over time. The operations further include inputting the first set of touch sensor data into the machine learning touch interpretation prediction model. The operations further include receiving, as an output of the touch interpretation prediction model, one or more predicted touch interpretations that describe the predicted intent of one or more user input objects. The touch sensor data may be raw touch sensor data, and the first set of touch sensor data may be a first set of raw touch sensor data.

[0006] Another example aspect of the present disclosure is directed to one or more tangible, non-transitory computer-readable media (which may be a tangible, non-transitory computer-readable medium, although the aspect is not limited thereto), the one or more tangible, non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations. The operations include obtaining data describing a machine learning touch interpretation prediction model. The touch interpretation prediction model has been trained to receive touch sensor data indicative of one or more positions of one or more user input objects relative to a touch sensor at one or more times, and to provide a plurality of predicted touch interpretation outputs in response to receiving the touch sensor data. Each predicted touch interpretation output corresponds to a different type of predicted touch interpretation based at least in part on the touch sensor data. The operations further include obtaining a first set of touch sensor data indicative of one or more user input object positions relative to the touch sensor over time. The operations further include inputting the first set of touch sensor data into the machine learning touch interpretation prediction model. The operations further include receiving a plurality of predicted touch interpretations as an output of the touch interpretation prediction model, each predicted touch interpretation describing a different predicted aspect of one or more user input objects. The operations further include performing one or more actions associated with the plurality of predicted touch interpretations. The touch sensor data may be raw touch sensor data, and the first set of touch sensor data may be a first set of raw touch sensor data.

[0007] Another example aspect of the present disclosure is directed to a mobile computing device that determines a touch interpretation from a user input object. The mobile computing device includes a processor and at least one computer-readable medium (which may be a tangible, non-transitory computer-readable medium, although this aspect is not limited thereto), the at least one computer-readable medium storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations. The operations include obtaining data describing a machine learning touch interpretation prediction model that includes a neural network. The touch interpretation prediction model has been trained to receive touch sensor data indicating one or more positions of one or more user input objects relative to a touch sensor at one or more times, and to output two or more predicted touch interpretations in response to receiving the touch sensor data. The operations further include obtaining a first set of touch sensor data associated with one or more user input objects, the first set of touch sensor data describing the positions of the one or more user input objects over time. The operations further include inputting the first set of touch sensor data into the machine learning touch interpretation prediction model. The operations further include receiving two or more predicted touch interpretations, as an output of the touch interpretation prediction model, that describe one or more predicted intents of the one or more user input objects. The two or more predicted touch interpretations include a set of touch point interpretations respectively describing one or more intended touch points and a gesture interpretation characterizing the set of touch point interpretations as a gesture determined from a predefined gesture category. The touch sensor data may be raw touch sensor data, and the first set of touch sensor data may be a first set of raw touch sensor data. A complementary example aspect of the present disclosure is directed to a processor-implemented method that includes obtaining data describing a machine learning touch interpretation prediction model that includes a neural network. The touch interpretation prediction model has been trained to receive touch sensor data indicating one or more positions of one or more user input objects relative to a touch sensor at one or more times, and to output two or more predicted touch interpretations in response to receiving the touch sensor data. The method further includes obtaining a first set of touch sensor data associated with one or more user input objects, the first set of touch sensor data describing the positions of the one or more user input objects over time. The method further includes inputting the first set of touch sensor data into the machine learning touch interpretation prediction model. The method further includes receiving two or more predicted touch interpretations, as an output of the touch interpretation prediction model, that describe one or more predicted intents of the one or more user input objects. The two or more predicted touch interpretations include a set of touch point interpretations respectively describing one or more intended touch points and a gesture interpretation characterizing the set of touch point interpretations as a gesture determined from a predefined gesture category. The touch sensor data may be raw touch sensor data, and the first set of touch sensor data may be a first set of raw touch sensor data.

[0008] Other aspects of the present disclosure are directed to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.

[0009] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] A detailed discussion of embodiments addressed to those of ordinary skill in the art is set forth in the specification with reference to the accompanying drawings, in which:

[0011] Figure 1 A block diagram of an example computing system performing machine learning in accordance with an example embodiment of the present disclosure is depicted;

[0012] Figure 2 A block diagram of a first example computing device performing machine learning in accordance with an example embodiment of the present disclosure is depicted;

[0013] Figure 3 A block diagram of a second example computing device performing machine learning in accordance with an example embodiment of the present disclosure is depicted;

[0014] Figure 4 A first example model arrangement in accordance with an example embodiment of the present disclosure is depicted;

[0015] Figure 5 A second example model arrangement in accordance with an example embodiment of the present disclosure is depicted;

[0016] Figure 6 A first aspect of an example use case in accordance with an example embodiment of the present disclosure is depicted;

[0017] Figure 7 A second aspect of an example use case in accordance with an example embodiment of the present disclosure is depicted;

[0018] Figure 8 A flowchart of an example method of performing machine learning in accordance with an example embodiment of the present disclosure is depicted;

[0019] Figure 9 A flowchart of a first additional aspect of an example method of performing machine learning in accordance with an example embodiment of the present disclosure is depicted;

[0020] Figure 10 A flowchart of a second additional aspect of an example method of performing machine learning in accordance with an example embodiment of the present disclosure is depicted; and

[0021] Figure 11Depicts a flowchart of a training method for a machine learning model according to an example embodiment of the present disclosure.

[0022] Specific implementation

[0023] Overview

[0024] Generally, the present disclosure is directed to systems and methods for implementing touch interpretation using machine learning. Specifically, the systems and methods of the present disclosure may include and use a machine learning touch interpretation prediction model, where the machine learning touch interpretation prediction model has been trained to receive raw touch sensor data (the raw touch sensor data indicating measured touch sensor readings across a grid of points generated in response to the position of one or more user input objects relative to the touch sensor), and, in response to receiving the raw touch sensor data, output multiple predicted touch interpretations. In some examples, the machine learning touch interpretation prediction model has been trained to output at least a first predicted touch interpretation and a second predicted touch interpretation simultaneously. The multiple predicted touch interpretations may include, for example, a set of touch point interpretations respectively describing one or more intended touch points, a gesture interpretation characterizing at least a portion of the raw touch sensor data as a gesture determined from a predefined gesture category, and / or a touch prediction vector describing one or more predicted future positions of one or more user input objects respectively for one or more future times. The touch interpretation prediction model may include a single input layer, but then provide multiple predicted touch interpretations at multiple different and discrete output layers, where the output layers share the same input layer and sometimes share additional layers between the input layer and the output layers. By using multi-task learning (where "multi-task" learning is machine learning that solves multiple learning tasks simultaneously) to predict multiple touch interpretations based on the same set of input data (e.g., raw touch sensor data that is not discarded after initial touch point processing), the commonalities and differences across different touch interpretations can be exploited to provide an improved level of accuracy for the output of the machine learning touch interpretation prediction model. Compared to systems that use multiple sequential data processing steps and potentially multiple unconnected models for prediction, using raw touch sensor data to directly predict user intent can also provide a more efficient and accurate touch sensor assessment. Given the determined predicted touch interpretations, applications or other components that consume data from the touch sensor can have improved responsiveness, reduced latency, and more customizable interactions. For example, mobile devices (e.g., smartphones) or other computing devices that use touch sensor input can benefit from the availability of predicted touch interpretations.

[0025] In one example, a user computing device (e.g., a mobile computing device such as a smartphone) can obtain raw touch sensor data that indicates measured touch sensor readings across a grid of points generated in response to the position of one or more user input objects relative to a touch sensor (e.g., a touch-sensitive display screen or other user input component). The raw touch sensor data can correspond, for example, to voltage levels registered by changes in capacitance across the surface of a capacitive touch-sensitive display screen or changes in resistance across the surface of a resistive touch-sensitive display screen. Example user input objects can include one or more fingers, thumbs, or hands of a user, a stylus operated by the user, or other user input objects. Specifically, in one example, a user can use a user input object to touch the touch-sensitive display screen or other touch-sensitive component of the computing device. The touch and / or movement of the user input object relative to the touch-sensitive display screen can enable the user to provide user input to the computing device.

[0026] In some examples, the raw touch sensor data provided as input to the touch interpretation prediction model can include one or more entries of user input object positions and times. For example, a set of raw touch sensor data can include one or more entries that provide the positions of one or more user input objects in the x-dimension and y-dimension and the timestamps associated with each position. As another example, the raw touch sensor data can include one or more entries that describe the changes in the positions of one or more user input objects in the x-dimension and y-dimension and the timestamps or time changes associated with the changes in each pair of x-values and y-values. In some implementations, when additional touch sensor data is detected, the set of touch sensor data can be iteratively updated, refreshed, or generated.

[0027] In some implementations, the raw touch sensor data provided as input to the touch interpretation prediction model can be provided as a time-stepping sequence of T inputs, each input corresponding to raw touch sensor data obtained at a different time step. For example, the touch sensor can be continuously monitored such that a time-stepping sequence of raw touch sensor data can be iteratively obtained in real time or near real time from the touch sensor. For example, the raw touch sensor data can be provided as a time series of sensor readings Z1, Z2, ..., Z T while. In some examples, the time differences between the T different sampling times (e.g., t1, t2, …, t T ) can be the same or different. Each sensor reading can be an array of points having a generally rectangular shape characterized by a first dimension (e.g., x) and a second dimension (e.g., y). For example, each sensor reading at time t can be represented as Z t = z xyt , x ∈ 1...W, y ∈ 1...H, t ∈ 1...T, where zxyt is the raw sensor measurement of the touch sensor at position (x, y) at time t for a sensor size of WxH. Each sensor reading obtained iteratively at different time steps can iteratively provide a single machine learning model that has been trained to compute multiple touch interpretations in response to a mapped array of the raw touch sensor data.

[0028] In some implementations, the computing device can feed the raw touch sensor data as input to a machine learning touch interpretation prediction model in an online manner. For example, during use in an application, in each instance where an update is received from the relevant touch sensor(s) (e.g., a touch-sensitive display screen), the latest touch sensor data update (e.g., the changed values of x, y, and time) can be fed into the touch interpretation prediction model. Thus, as additional touch sensor data is collected, the raw touch sensor data collection and touch interpretation prediction can be performed iteratively. In this way, one benefit provided by using a machine learning model (e.g., a recurrent neural network) is the ability to maintain context from previous updates when new touch sensor updates are input in the online manner as described above.

[0029] According to one aspect of the present disclosure, a user computing device (e.g., a mobile computing device such as a smartphone) can input the raw touch sensor data into a machine learning touch interpretation prediction model. In some implementations, the machine learning touch interpretation prediction model can include a neural network, and inputting the raw touch sensor data includes inputting the raw touch sensor data into the neural network of the machine learning touch interpretation prediction model. In some implementations, the touch interpretation prediction model can include a convolutional neural network. In some implementations, the touch interpretation prediction model can be a time model that allows the raw touch sensor data to be referenced in a timely manner. In such an instance, the neural network within the touch interpretation prediction model can be a recurrent neural network (e.g., a deep recurrent neural network). In some examples, the neural network within the touch interpretation prediction model is a Long Short-Term Memory (LSTM) neural network, a Gated Recurrent Unit (GRU) neural network, or other forms of recurrent neural networks.

[0030] According to another aspect of the present disclosure, a user computing device (e.g., a mobile computing device such as a smartphone) can receive one or more predicted touch interpretations that describe one or more predicted intents of one or more user input objects as an output of a machine learning touch interpretation prediction model. In some implementations, the machine learning touch interpretation model outputs multiple predicted touch interpretations. In some instances, the machine learning touch interpretation prediction model has been trained to output at least a first predicted touch interpretation and a second predicted touch interpretation simultaneously. In some implementations, multiple predicted touch interpretations can be provided by the touch interpretation prediction model at different and distinct output layers. However, the different output layers can be downstream of at least one shared layer (e.g., the input layer and one or more subsequent shared layers) and include at least one shared layer. Thus, the machine learning touch interpretation prediction model can include at least one shared layer and multiple different and distinct output layers that are structurally located after the at least one shared layer. One or more computing devices can obtain raw touch sensor data and input the raw touch sensor data into the at least one shared layer, while the multiple output layers can be configured to respectively provide multiple predicted touch interpretations.

[0031] In some implementations, the predicted touch interpretation can include a set of touch point interpretations that respectively describe zero (0), one (1), or more intended touch points. Non-intended touch points can also be directly identified or inferred by excluding them from a list of intended touch points. For example, the set of touch point interpretations that describe the intended touch points can be output as a potentially empty set of intended touch points, with each touch point represented as a two-dimensional coordinate pair (x,y) of up to N different touch points (e.g., (x,y)1...(x,y) N )). In some implementations, additional data can accompany each identified touch point in the set of touch point interpretations. For example, in addition to the location of each intended touch point, the set of touch point interpretations can also include the estimated touch pressure at each touch point and / or the touch type that describes the predicted type of user input object associated with each touch point (e.g., which finger, joint, palm, stylus, etc. is predicted to cause the touch), and / or the radius of the user input object on the touch-sensitive display screen, and / or other user input object parameters.

[0032] In some implementations, the set of touch point interpretations includes intended touch points and excludes unintended touch points. Unintended touch points can include, for example, accidental and / or undesired touch locations on a touch-sensitive display screen or component. Unintended touch points can occur under a variety of circumstances, such as touch sensor noise at one or more points, the way of holding a mobile computing device that causes an unintended touch on some parts of the touch sensor surface, pocket dialing, or accidental photographing, etc. For example, when a user holds a mobile device (e.g., a smartphone) with one hand while providing user input to the touch-sensitive display screen with the other hand, unintended touch points may occur. Parts of the first hand (e.g., the user's palm and / or thumb) may sometimes provide input to the touch-sensitive display screen around the edge of the display screen corresponding to the unintended touch points. This situation is particularly common for mobile computing devices with a relatively smaller bezel around the periphery of the mobile computing device.

[0033] In some implementations, the predicted touch interpretation can include characterizing at least a portion of the raw touch sensor data (and additionally or alternatively, a previously predicted set of touch point interpretations and / or a touch prediction vector, etc.) as a gesture interpretation of a gesture determined from a predefined gesture category. In some implementations, the gesture interpretation can be at least partially based on the last δ steps of the raw touch sensor data. Example gestures can include, but are not limited to, non-recognized gestures, click / press gestures (e.g., including hard click / press, soft click / press, short click / press, and / or long click / press), tap gestures, double-tap gestures for selecting or otherwise interacting with one or more items displayed on a user interface, scroll gestures or slide gestures for transitioning the user interface in one or more directions and / or switching the user interface screen from one mode to another, pinch gestures for zooming in or out relative to the user interface, draw gestures for drawing a line or typing a word, and / or other gestures. In some implementations, the predefined gesture category from which the gesture interpretation is determined can be or otherwise include a set of gestures associated with a dedicated accessibility mode (e.g., a visually impaired interface mode through which the user interacts using Braille or other dedicated writing style characters).

[0034] In some implementations, if at least a portion of the raw touch sensor data is characterized as an interpreted gesture, the gesture interpretation can include information that not only identifies the type of the gesture but also additionally or alternatively identifies the location of the gesture. For example, the gesture interpretation can take the form of a three-dimensional data set (e.g., (c, x, y)), where c is a predefined gesture category (e.g., "not a gesture", "tap", "slide", "pinch to zoom", "double-tap slide for zoom", etc.), and x, y are the coordinates in the first dimension (e.g., x) and the second dimension (e.g., y) on the touch-sensitive display where the gesture occurs.

[0035] In some implementations, predictive touch interpretation may include an accessibility mode, where the accessibility mode describes the predicted type of interaction of one or more user input objects determined from a predefined category of accessibility modes including one or more of a standard interface mode and a visually impaired interface mode.

[0036] In some examples, predictive touch interpretation may include a touch prediction vector that describes one or more predicted future positions of one or more user input objects respectively for one or more future times. In some implementations, each future position of the one or more user input objects may be presented as a sensor reading Z depicting an expected pattern on the touch sensor for a future theta number of time steps. t+θ . In some examples, a machine learning touch interpretation prediction model may be configured to generate a touch prediction vector as an output. In some examples, the machine learning touch interpretation prediction model may additionally or alternatively be configured to receive the determined touch prediction vector as an input to assist the machine learning touch interpretation prediction model in making related determinations of other touch interpretations (e.g., gesture interpretation, set of touch point interpretations, etc.).

[0037] The time step theta for which the touch prediction vector is determined can be configured in various ways. In some implementations, the machine learning touch interpretation prediction model may be configured to consistently output predicted future positions that define a predefined set of values of theta for one or more future times. In some implementations, one or more future times theta may be provided as a separate input along with the raw touch sensor data to the machine learning touch interpretation prediction model. For example, a computing device may input a time vector along with the raw touch sensor data into the machine learning touch interpretation prediction model. The time vector may provide a list of the lengths of time (e.g., 10 ms, 20 ms, etc.) that are desired to be predicted by the touch interpretation prediction model. Thus, the time vector may describe one or more future times at which the position of the user input object is to be predicted. In response to receiving the raw touch sensor data and the time vector, the machine learning touch interpretation prediction model may output a touch prediction vector that describes the predicted future positions of the user input object for each time or time length described by the time vector. For example, the touch prediction vector may include a pair of values for the position in the x dimension and the y dimension for each future time, or a pair of values for the change in the x dimension and the y dimension for each future time.

[0038] A computing device can use a touch prediction vector output by a touch interpretation prediction model to reduce or even eliminate touch latency. Specifically, the computing device can perform an operation in response to or otherwise based on the predicted future position of a user input object, thereby eliminating the need to wait to receive and process the remainder of the user input action. To provide an example, a computing device of the present disclosure can input a time-stepped sequence of raw touch sensor data that describes touch positions representing finger movements associated with an initial portion of a user touch gesture (e.g., an initial portion of a left swipe gesture) into a neural network of a touch interpretation prediction model. In response to receiving the raw touch sensor data, the touch interpretation prediction model can predict finger movements / positions associated with the remainder of the user touch gesture (e.g., the remainder of the left swipe gesture). The computing device can perform an action (e.g., render a display screen where a display object has been swiped left) in response to the predicted finger movements / positions. In this way, the computing device does not need to wait and then process the remainder of the user touch gesture. Thus, the computing device can respond more quickly to touch events and reduce touch latency. For example, high-quality gesture prediction can enable touch latency to be reduced to a level imperceptible to a human user.

[0039] By providing a machine learning model that has been trained to output multiple joint variables, improvements in determining some predicted touch interpretations can benefit from improvements in determining other predicted touch interpretations. For example, an improvement in determining a set of touch point interpretations can help improve the determined gesture interpretation. Similarly, an improvement in determining a touch prediction vector can help improve the determination of a set of touch point interpretations. By co-training a machine learning model across multiple desired outputs (e.g., multiple predicted touch interpretations), an effective model can generate different output layers that share the same input layer (e.g., raw touch sensor data).

[0040] According to another aspect of the present disclosure, in some implementations, the touch interpretation prediction model or at least a portion thereof may be used via an Application Programming Interface (API) for one or more applications provided on a computing device. In some instances, a first application uses the API to request access to the machine learning touch interpretation prediction model. The machine learning touch interpretation prediction model may be hosted as part of a second application, or hosted in a dedicated layer, application, or component within the same computing device as the first application, or hosted in a separate computing device. In some implementations, the first application may effectively add a custom output layer to the machine learning touch interpretation prediction model that can be further trained on top of a pre-trained stable model. Such an API would allow for the benefits of maintaining a single machine learning model that runs to predict custom user touch interactions, particularly those targeted by the first application, without any additional cost of implicitly gaining access to the complete set of raw touch sensor data.

[0041] In one example, the first application may be designed to predict user touch interactions when operating in a visually impaired interface mode, such as an application configured to receive braille input via a touch-sensitive display or other touch-sensitive component. This type of interface mode can be substantially different from traditional user interface modes due to relying on a device with a touch sensor capable of capturing multiple touch points simultaneously. In another example, the first application may be designed to process specialized input from a stylus or process input from a different type of touch sensor, which may require adapting the touch interpretation prediction model to a new type of touch sensor data received as input by the computing device or its touch-sensitive component.

[0042] In some examples, to specify an additional training data path for the custom output layer of the touch interpretation prediction model, the application programming interface may allow definition libraries from machine learning tools such as TensorFlow and / or Anao to be flexibly coupled from the first application to the touch interpretation prediction model. In such examples, additional training and post-training operations for the (multiple) custom output layers can be run in parallel with the pre-defined touch interpretation prediction model at very little additional cost.

[0043] According to another aspect of the present disclosure, the touch interpretation prediction model described herein may be trained on ground truth data using a novel loss function. More specifically, a training computing system may use a training data set that includes multiple ground truth data sets to train the touch interpretation prediction model.

[0044] In some implementations, when training a machine learning touch interpretation prediction model to determine a set of touch point interpretations and / or one or more gesture interpretations, a first example training dataset can include touch sensor data and corresponding labels that describe a large number of previously observed touch interpretations.

[0045] In one implementation, a first example training dataset includes a first portion of data corresponding to recorded touch sensor data that indicates the position of one or more user input objects relative to a touch sensor. For example, the recorded touch sensor data can be recorded while a user is operating a computing device having a touch sensor under normal operating conditions. The first example training dataset can also include a second portion of data corresponding to labels of determined touch interpretations applied to the recorded screen content. For example, the recorded screen content can be co-recorded simultaneously with the first portion of data corresponding to the recorded touch sensor data. The labels of the determined touch interpretations can include, for example, a set of touch point interpretations and / or gesture interpretations determined at least in part from the screen content. In some instances, the screen content can be manually labeled. In some instances, traditional heuristics applied to interpret the raw touch sensor data can be used to automatically label the screen content. In some instances, a combination of automatic and manual labeling can be used to label the screen content. For example, manual labeling can be used to best identify cases where a palm is touching the touch-sensitive display screen to prevent the user from using the device in the desired manner, while automatic labeling can be used to potentially identify other touch types.

[0046] In other implementations, a dedicated data collection application that prompts a user to perform certain tasks relative to the touch-sensitive display screen (e.g., asking the user to "tap here" while holding the device in a particular manner) can be used to build the first example training dataset. The raw touch sensor data can be associated with the touch interpretations identified during the operation of the dedicated data collection application to form the first and second portions of the first example training dataset.

[0047] In some implementations, when training a machine learning touch interpretation prediction model to determine a touch prediction vector, a second example training dataset can include a first portion of data corresponding to an initial sequence of touch sensor data observations (e.g., Z1...Z T ) and a second portion of data corresponding to a subsequent sequence of touch sensor data observations (e.g., Z T+1 ...Z T+F ). Because the prediction step can be compared to the observation step when it occurs at a future time, the sequence of sensor data Z1...Z T is automatically suitable for training to predict future sensor data Z T+1 ...Z T+FPredictions. It should be understood that any combination of the above techniques and other techniques can be used to obtain one or more training data sets for training a machine learning touch interpretation prediction model.

[0048] In some implementations, to train a touch interpretation prediction model, a training computing system can input a first portion of a ground truth data set into the touch interpretation prediction model to be trained. In response to receiving such a first portion, the touch interpretation prediction model outputs one or more touch interpretations that predict the remaining portion of the ground truth data set (e.g., the second portion of data). After such a prediction, the training computing system can apply or otherwise determine a loss function, where the loss function compares the one or more touch interpretations output by the touch interpretation prediction model with the remaining portion of the ground truth data that the touch interpretation prediction model is attempting to predict. The training computing system can then backpropagate the loss function through the touch interpretation prediction model to train the touch interpretation prediction model (e.g., by modifying one or more weights associated with the touch interpretation prediction model).

[0049] The systems and methods described herein can provide many technical effects and benefits. For example, the disclosed techniques can improve touch sensor output by predicting user intent with improved accuracy when interfacing with a touch-sensitive display. The accuracy level can be improved due to the ability of machine learning models to implement multi-task learning. By learning how to predict multiple touch interpretations based on the same input data set (e.g., raw touch sensor data that is not discarded after initial touch point processing), the commonalities and differences of different touch interpretations can be exploited. Touch patterns specific to different users can also be identified and used to enhance touch interpretation prediction. More intricate nuances in touch interpretation determination can thus be provided using the disclosed machine learning techniques. When the machine learning model includes a deep neural network as described above, an excellent function approximator that provides a richer predictive ability compared to polynomials can be used to predict touch interpretations. Thus, if properly trained, the touch interpretation prediction model of the present disclosure can provide excellent prediction accuracy.

[0050] Another example technical effect and benefit of the present disclosure is the improved efficiency in determining multiple touch interpretations. Compared to solutions that determine each touch interpretation in a sequential processing manner (e.g., first determining touch points from raw touch sensor data and subsequently separately and sequentially determining gestures from the touch points), a jointly trained touch interpretation model that provides multiple touch interpretations as outputs simultaneously can provide more accurate predictions in a more efficient manner. The sequential processing steps for determining touch interpretations not only take longer but also have the drawback that if an error occurs during one of the earlier steps, the final output cannot be recovered. The disclosed use of a machine learning touch interpretation prediction model that provides multiple touch interpretation outputs simultaneously can avoid this potential problem associated with end-to-end processing.

[0051] Another technical effect and benefit of the present disclosure is to enhance the opportunity to determine specialized touch interpretations in a principled manner. For example, touch point interpretations that also include the predicted touch type of each touch point (e.g., predicting which finger, joint, palm, stylus, etc. caused the touch) can provide useful information for predicting the user's intent with respect to the input to the touch-sensitive display device. The ability to determine touch interpretations in a specialized accessibility mode (e.g., visually impaired mode) greatly broadens the usefulness of the disclosed machine learning models. As new touch-based technologies are developed, the ability to adapt the disclosed machine learning touch interpretation prediction models to include new output layers provides even further advantages.

[0052] In some implementations, such as when the machine learning touch interpretation prediction model is configured to output a touch prediction vector, the disclosed techniques can be used to improve responsiveness during the operation of an application. For example, an application, program, or other component of a computing device (e.g., a handwriting recognition application) can consume or otherwise be provided with multiple touch interpretations that include the touch prediction vector. The application can treat the touch interpretation that includes the touch prediction vector as if the user input object has been moved to such a predicted location. For example, a handwriting recognition application can recognize handwriting based on the predicted future location provided by the touch prediction vector. Thus, the application does not need to wait to receive and process the remainder of the user input action through a large processing stack, thereby reducing latency.

[0053] In addition to reduced touch latency, some implementations of the present disclosure can produce many additional technical benefits, including, for example, smoother finger tracking, improved handwriting recognition, faster and more precise user control of user-operable virtual objects (e.g., objects within a game), and many other benefits in scenarios where user touch input is provided to a computing device.

[0054] Another exemplary technical benefit of the present disclosure is improved scalability. Specifically, by modeling touch sensor data with a neural network or other machine learning model, the research time required is significantly reduced compared to the development of manual touch interpretation algorithms. For example, for a manual touch interpretation algorithm, the designer would need to exhaustively derive heuristic models of how different touch patterns correspond to the user's intent in different scenarios and / or with different users. In contrast, to use the neural network or other machine learning techniques described herein, the machine learning touch interpretation prediction model can be trained on appropriate training data, which can be done on a large scale if the training system allows. Additionally, as new training data becomes available, the machine learning model can be easily modified.

[0055] The systems and methods described herein can also provide the technical effects and benefits of improved computer technology in the form of relatively low memory usage / requirements. Specifically, the neural networks or other machine learning touch interpretation prediction models described herein effectively generalize the training data and compress it into a compact form (e.g., the machine learning model itself). This significantly reduces the amount of memory required to store and implement the touch interpretation prediction algorithm.

[0056] Reference is now made to the accompanying drawings, and example embodiments of the present disclosure will be discussed in further detail.

[0057] Example Devices and Systems

[0058] Figure 1 An example computing system 100 for predicting multiple touch interpretations in accordance with an example embodiment of the present disclosure is depicted. System 100 includes a user computing device 102, a machine learning computing system 130, and a training computing system 150 communicatively coupled via a network 180.

[0059] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0060] The user computing device 102 can include one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple processors operably connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash devices, disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0061] The user computing device 102 may include at least one touch sensor 122. The touch sensor 122 may be, for example, a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch sensor 122 is capable of detecting raw touch sensor data that indicates measured touch sensor readings across a grid of points generated in response to the position of one or more user input objects relative to the touch sensor 122. In some implementations, the touch sensor 122 is associated with capacitive touch-sensitive elements such that the raw touch sensor data corresponds to voltage levels registered by capacitance changes across the surface of the capacitive touch-sensitive elements. In some implementations, the touch sensor 122 is associated with resistive touch-sensitive elements such that the raw touch sensor data corresponds to voltage levels registered by resistance changes across the surface of the resistive touch-sensitive elements. Example user input objects for interfacing with the touch sensor 122 may include one or more fingers, thumbs, or hands of the user, a stylus operated by the user, or other user input objects. Specifically, in one example, the user may use a user input object to touch the touch-sensitive component associated with the touch sensor 122 of the user computing device 102. The touch and / or movement of the user input object relative to the touch-sensitive component may enable the user to provide input to the user computing device 102.

[0062] The user computing device 102 may also include one or more additional user input components 124 that receive user input. For example, the user input component 124 may detect user input based on touchless gestures by analyzing images collected by a camera of the device 102 through a computer vision system or by using radar (e.g., micro radar) to track the movement of the user input object. Thus, the movement of the user input object relative to the user input component 124 enables the user to provide user input to the computing device 102.

[0063] The user computing device may also include one or more user output components 126. The user output components 126 may include, for example, a display device. In some implementations, such a display device may correspond to the touch-sensitive display device associated with the touch sensor 122. The user output component 126 may be configured to display a user interface to the user as part of the normal device operation of the user computing device 102. In some implementations, the user output component 126 may be configured to provide an interface to the user that serves as part of a process for capturing a training data set for training a machine learning touch interpretation prediction model.

[0064] The user computing device 102 may store or include one or more machine learning touch interpretation prediction models 120. In some implementations, one or more touch interpretation prediction models 124 may be received from the machine learning computing system 130 via the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, the user computing device 102 may implement multiple parallel instances of a single touch interpretation prediction model 120 (e.g., to perform parallel touch interpretation predictions for multiple input objects).

[0065] The machine learning computing system 130 may include one or more processors 132 and a memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be a single processor or multiple processors operably connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the machine learning computing system 130 to perform operations.

[0066] In some implementations, the machine learning computing system 130 includes one or more server computing devices or is otherwise implemented by one or more server computing devices. In instances where the machine learning computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0067] The machine learning computing system 130 may store or otherwise include one or more machine learning touch interpretation prediction models 140. For example, the touch interpretation prediction model 140 may be or otherwise include various machine learning models, such as neural networks (e.g., deep recurrent neural networks) or other multi-layer non-linear models, regression-based models, etc. Refer to Figures 4 - 5 for discussion of example touch interpretation prediction models 140.

[0068] The machine learning computing system 130 may train the touch interpretation prediction model 140 via interaction with a training computing system 150 communicatively coupled via the network 180. The training computing system 150 may be separate from the machine learning computing system 130 or may be part of the machine learning computing system 130.

[0069] The training computing system 150 may include one or more processors 152 and a memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be a single processor or multiple processors operably connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash devices, magnetic disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes one or more server computing devices or is otherwise implemented by one or more server computing devices.

[0070] The training computing system 150 may include a model trainer 160 that trains a machine learning model 140 stored at the machine learning computing system 130 using various training or learning techniques (such as, for example, backpropagation). The model trainer 160 may perform a variety of generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the trained model.

[0071] Specifically, the model trainer 160 may train the touch interpretation prediction model 140 based on a set of training data 142. The training data 142 may include ground truth data (e.g., a first data set including recorded touch sensor data and labels of determined touch interpretations applied to co-recorded screen content and / or a second data set including an initial sequence of touch sensor data observations and subsequent sequences of touch sensor data observations). In some implementations, if the user has provided consent, the training data 142 may be provided by the user computing device 102 (e.g., based on touch sensor data detected by the user computing device 102). Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 on user-specific data received from the user computing device 102. In some instances, this process may be referred to as a personalized model.

[0072] The model trainer 160 may include computer logic for providing the desired functionality. The model trainer 160 may be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into the memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium such as RAM, a hard disk, or an optical or magnetic medium.

[0073] Network 180 can be any type of communication network such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over network 180 can be carried via any type of wired and / or wireless connection using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, secure HTTP, SSL).

[0074] Figure 1 An example computing system that can be used to implement the present disclosure is shown. Other computing systems can also be used. For example, in some implementations, user computing device 102 can include a model trainer 160 and a training data set 162. In such implementations, the touch interpretation prediction model can be trained and used locally at the user computing device.

[0075] Figure 2 A block diagram of an example computing device 10 that performs communication assistance according to an example embodiment of the present disclosure is depicted. Computing device 10 can be a user computing device or a server computing device.

[0076] Computing device 10 includes a plurality of applications (e.g., applications 1 to J). Each application contains its own machine learning library and (multiple) machine learning models. For example, each application can include a machine learning communication assistance model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, a virtual reality (VR) application, etc.

[0077] As Figure 2 shown, each application can communicate with a plurality of other components of the computing device (such as, for example, one or more sensors, a context manager, a device status component, and / or additional components). In some implementations, each application can communicate with each device component using an API (e.g., a common API). In some implementations, the API used by each application can be specific to that application.

[0078] Figure 3 A block diagram of an example computing device 50 that performs communication assistance according to an example embodiment of the present disclosure is depicted. Computing device 50 can be a user computing device or a server computing device.

[0079] Computing device 50 includes a plurality of applications (e.g., applications 1 through J). Each application can communicate with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, a virtual reality (VR) application, and the like. In some implementations, each application can communicate with the central intelligence layer (and the (multiple) models stored therein) using an API (e.g., a common API across all applications).

[0080] The central intelligence layer includes a plurality of machine learning models. For example, as Figure 3 shown, a corresponding machine learning model (e.g., a communication assistance model) can be provided for each application, and the corresponding machine learning model is managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model (e.g., a single communication assistance model) for all applications. In some implementations, the central intelligence layer can be included within the operating system of computing device 50 or otherwise implemented by the operating system of computing device 50.

[0081] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a central repository of data for computing device 50. As Figure 3 shown, the central device data layer can communicate with a plurality of other components of the computing device, such as, for example, one or more sensors, a context manager, a device status component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0082] Example Touch Interpretation Prediction Model

[0083] Figure 4 Depicted is a first example touch interpretation prediction model 200 according to an example embodiment of the present disclosure. The touch interpretation prediction model 200 can be a machine learning model. In some implementations, the touch interpretation prediction model 200 can be or otherwise include various machine learning models, such as a neural network (e.g., a deep recurrent neural network) or other multi-layer non-linear models, a regression-based model, and the like. When the touch interpretation prediction model 200 includes a recurrent neural network, this can be a multi-layer long short-term memory (LSTM) neural network, a multi-layer gated recurrent unit (GRU) neural network, or other forms of recurrent neural networks.

[0084] The touch interpretation prediction model 200 can be configured to receive raw touch sensor data 210. In one example, a user computing device (e.g., a mobile computing device) obtains raw touch sensor data 210 including one or more entries of user input object locations and times. For example, the set of raw touch sensor data 210 can include one or more entries, where the one or more entries provide the locations of one or more user input objects in the x and y dimensions (e.g., defined relative to the touch sensor 122) and timestamps associated with each location. As another example, the raw touch sensor data 210 can include one or more entries describing changes in the locations of one or more user input objects in the x and y dimensions and timestamps or time changes associated with changes in each pair of x and y values. In some implementations, the set of raw touch sensor data 210 can be iteratively updated, refreshed, or generated when additional touch sensor data is detected.

[0085] Still referring to Figure 4 , the raw touch sensor data 210 can be provided as input to the machine learning touch interpretation prediction model 200. In some implementations, the touch interpretation prediction model 200 can include one or more shared layers 202 and multiple different and distinct output layers 204 - 208. The multiple output layers 204 - 208 can be structurally located after the one or more shared layers 202 within the touch interpretation prediction model 200. The one or more shared layers 202 can include an input layer and one or more additional shared layers structurally located after the input layer. The input layer is configured to receive the raw touch sensor data 210. The multiple output layers 204 - 208 can be configured to respectively provide multiple predicted touch interpretations as the output of the touch interpretation prediction model 200. For example, the first output layer 204 can be a first touch interpretation output layer configured to provide a first predicted touch interpretation 212 as the output of the touch interpretation prediction model 200. The second output layer 206 can be a second touch interpretation output layer configured to provide a second predicted touch interpretation 214 as the output of the touch interpretation prediction model 200. The third output layer 208 can be a third touch interpretation output layer configured to provide one or more additional predicted touch interpretations 216 as the output of the touch interpretation prediction model 100. Although Figure 4 three output layers 204 - 208 and three model outputs 212 - 216 are shown, any number of multiple output layers and model outputs can be used according to the disclosed multi - task machine learning model.

[0086] In some implementations, the machine learning touch interpretation prediction model 200 can include a neural network, and inputting the raw touch sensor data 210 includes inputting the raw touch sensor data into the neural network of the machine learning touch interpretation prediction model 200. In some implementations, the touch interpretation prediction model 200 can include a convolutional neural network. In some implementations, the touch interpretation prediction model 200 can be a temporal model that allows the raw touch sensor data to be referenced in a timely manner. In such instances, the neural network within the touch interpretation prediction model 200 can be a recurrent neural network (e.g., a deep recurrent neural network). In some examples, the neural network within the touch interpretation prediction model 200 can be a Long Short-Term Memory (LSTM) neural network, a Gated Recurrent Unit (GRU) neural network, or other forms of recurrent neural networks.

[0087] In some implementations, the machine learning touch interpretation prediction model 200 can include many different sizes, numbers of layers, and levels of connectivity. Some layers can correspond to stacked convolutional layers (optionally followed by contrast normalization and max pooling), which are then followed by one or more fully connected layers. For neural networks trained on large data sets, the potential problem of overfitting can be addressed by using dropout to increase the number of layers and layer sizes. In some instances, the neural network can be designed to forego using fully connected upper layers at the top of the network. By forcing the network to perform dimensionality reduction in the middle layers, a fairly deep neural network model can be designed while significantly reducing the number of learning parameters. Additional specific features of exemplary neural networks that can be used according to the disclosed techniques can be found in "Scalable Object Detection using Deep Neural Networks," Dumitru Erhan et al., arXiv:1312.2249 [cs.CV], CVPR2014 and / or "Scalable HighQuality Object Detection," Christian Szegedy et al., arXiv:1412.1441 [cs.CV], Dec. 2015, which are incorporated herein by reference for all purposes.

[0088] In some implementations, the touch interpretation prediction model 200 may be configured to generate multiple predicted touch interpretations. In some examples, the touch interpretation prediction model 200 may also output a learned confidence metric for each of the predicted touch interpretations. For example, the confidence metric for each predicted touch interpretation may be expressed as a confidence metric value within a range (e.g., 0.0 - 1.0 or 0 - 100%) indicating the degree of certainty of the predicted touch interpretation.

[0089] In some implementations, the touch interpretation prediction model 200 may be configured to generate at least a first predicted touch interpretation 212 and a second predicted touch interpretation 214. The first predicted touch interpretation 212 may include a first of a set of touch point interpretations, a gesture interpretation determined from a predefined gesture category, a touch prediction vector for one or more future times, and an accessibility mode. The second predicted touch interpretation 214 may include a second of the set of touch point interpretations, a gesture interpretation determined from a predefined gesture category, a touch prediction vector for one or more future times, and an accessibility mode. The first touch point interpretation 212 may be different from the second touch point interpretation 214. One or more additional predicted touch interpretations 216 may also be generated by the touch interpretation prediction model. The additional predicted touch interpretations 216 may be selected from the same group as described above, or may include specialized touch interpretations as described herein. Refer to Figure 5 Describe more specific aspects of the example touch interpretations described above.

[0090] Figure 5 Depicts a second example touch interpretation prediction model 220 according to an example embodiment of the present disclosure. The touch interpretation prediction model 220 may be similar to Figure 4 the touch interpretation model 200, and the features described with reference to one may be applied to the other, and vice versa. In some implementations, the sensor output prediction model 220 may be a time model that allows sensor data to be referenced in a timely manner. In such an implementation, the raw touch sensor data provided as input to the touch interpretation prediction model 220 may be provided as a time-stepping sequence of T inputs. For example, the raw touch sensor data may be obtained and provided as a sequence of T inputs 224 - 228, each input corresponding to or including raw touch sensor data sampled or obtained at different time steps. For example, a time-stepping sequence of touch sensor data 224 - 228 may be iteratively obtained from the touch sensor in real time or near real time at different sampling times (e.g., t1, t2,..., t T )). In some examples, the T different sampling times (e.g., t1, t2,..., t TThe time difference between every two adjacent sampling times of [[ID=]] can be the same or it can be different. In such an example, the raw touch sensor data can be obtained for each of the T different times. For example, the first touch sensor data set 224 can correspond to the touch sensor data sampled at time t1. The second touch sensor data set 226 can correspond to the touch sensor data sampled at time t2. An additional number of touch sensor data sets can be provided in a sequence of T time-step samples until the last touch sensor data set 228 corresponding to the touch sensor data sampled at time t T is provided. Each of the touch sensor data sets 224 - 228 can be iteratively provided as an input to the touch interpretation prediction model 200 / 220 when being iteratively obtained.

[0091] In some implementations, Figure 5 one or more of the touch sensor data sets 224 - 228 depicted in [[ID=]] can be represented as a time series of sensor readings Z1, Z2,..., Z T . Each sensor reading corresponding to the touch sensor data sets 224 - 228 can be an array of points having a generally rectangular shape characterized by a first dimension (e.g., x) and a second dimension (e.g., y). For example, each sensor reading at time t can be represented as Z t = z xyt , where x ∈ 1...W, y ∈ 1...H, t ∈ 1...T, and z xyt is the raw sensor measurement of the touch sensor at position (x, y) at time t for a sensor size of WxH points.

[0092] In some implementations, the computing device can feed the raw touch sensor data sets 224 - 228 as inputs to the machine learning touch interpretation prediction model in an online manner. For example, during use in an application, in each instance of receiving an update from the relevant touch sensor(s) (e.g., a touch-sensitive display screen), the latest touch sensor data set (e.g., the changing values of x, y, and time) can be fed into the touch interpretation prediction model 220. Thus, as additional touch sensor data sets 224 - 228 are collected, the raw touch sensor data collection and touch interpretation prediction can be performed iteratively.

[0093] In some examples, Figure 5 the touch interpretation prediction model 220 of [[ID=]] can be configured to generate multiple predicted touch interpretations, such as the multiple predicted touch interpretations 212 - 216 shown in [[ID=]] and previously described with reference to [[ID=]] Figure 4 Figure 4 .

[0094] In some implementations, the predicted touch interpretation generated as an output of the touch interpretation prediction model 220 may include a set of touch point interpretations 224 that describe zero (0), one (1), or more intended touch points, respectively. Non-intended touch points may also be directly identified, or may be inferred by excluding non-intended touch points from a list of intended touch points. For example, the set of touch point interpretations 224 that describe intended touch points may be output as a potentially empty set of intended touch points, with each touch point represented as a two-dimensional coordinate pair (x,y) of up to N different touch points (e.g., (x,y)1...(x,y) N )). In some implementations, additional data may accompany each identified touch point in the set of touch point interpretations 224. For example, in addition to the location of each intended touch point, the set of touch point interpretations 224 may also include the estimated touch pressure at each touch point and / or the touch type of a user input object that describes the predicted type associated with each touch point (e.g., which finger, joint, palm, stylus, etc. is predicted to cause the touch), and / or the radius of the user input object on the touch-sensitive display screen, and / or other user input object parameters.

[0095] In some implementations, the set of touch point interpretations 224 is intended to include intended touch points and exclude non-intended touch points. Non-intended touch points may include, for example, accidental and / or undesired touch locations on the touch-sensitive display screen or component. Non-intended touch points may occur in a variety of situations, such as touch sensor noise at one or more points, the way of holding a mobile computing device that causes an unintentional touch on some parts of the touch sensor surface, pocket dials, or accidental photographing. For example, non-intended touch points may occur when a user holds a mobile device (e.g., a smartphone) with one hand while providing user input to the touch-sensitive display screen with the other hand. Parts of the first hand (e.g., the user's palm and / or thumb) may sometimes provide input to the touch-sensitive display screen around the edge of the display screen corresponding to non-intended touch points. This situation is particularly common for mobile computing devices with a relatively smaller bezel around the periphery of the mobile computing device.

[0096] In some implementations, the predicted touch interpretation generated as an output of the touch interpretation prediction model 220 may include a gesture interpretation 226 that characterizes at least a portion of the raw touch sensor data 210 and / or 224 - 228 (and additionally or alternatively, a set of previously predicted touch point interpretations 224 and / or a touch prediction vector 228, etc.) as a gesture determined from a predefined gesture category. In some implementations, the gesture interpretation 226 may be at least partially based on the last δ steps of the raw touch sensor data 210 and / or 224 - 228. Example gestures may include, but are not limited to, a non - recognized gesture, a click / press gesture (e.g., including a hard click / press, a soft click / press, a short click / press, and / or a long click / press), a tap gesture, a double - tap gesture for selecting or otherwise interacting with one or more items displayed on the user interface, a scroll gesture or a swipe gesture for shifting the user interface in one or more directions and / or transitioning the user interface screen from one mode to another, a pinch gesture for zooming in or out relative to the user interface, a drawing gesture for drawing a line or typing a word, and / or other gestures. In some implementations, the predefined gesture category from which the gesture interpretation 226 is determined may be or otherwise include a set of gestures associated with a specialized accessibility mode (e.g., a visually impaired interface mode through which the user interacts using Braille or other specialized writing - style characters).

[0097] In some implementations, if at least a portion of the raw touch sensor data 210 and / or 224 - 228 is characterized as an interpreted gesture, the gesture interpretation 226 may include information that not only identifies the type of the gesture but also additionally or alternatively identifies the location of the gesture. For example, the gesture interpretation 226 may take the form of a three - dimensional data set (e.g., (c, x, y)), where c is a predefined gesture category (e.g., “not a gesture”, “tap”, “swipe”, “pinch to zoom”, “double - tap swipe for zoom”, etc.), and x, y are the coordinates in the first dimension (e.g., x) and the second dimension (e.g., y) on the touch - sensitive display where the gesture occurred.

[0098] In some implementations, the predicted touch interpretation generated as an output of the touch interpretation prediction model 220 may include an accessibility mode, where the accessibility mode describes the predicted type of interaction of one or more user input objects determined from a predefined category of accessibility modes including one or more of a standard interface mode and a visually impaired interface mode.

[0099] In some examples, the predicted touch interpretation generated as the output of the touch interpretation prediction model 220 can include a touch prediction vector 228 that describes one or more predicted future positions of one or more user input objects respectively for one or more future times. In some implementations, each future position of the one or more user input objects can be presented as a sensor reading Z depicting an expected pattern of the touch sensor after θ future time steps. t+θ In some examples, the machine learning touch interpretation prediction model 220 can be configured to generate the touch prediction vector 228 as an output. In some examples, the machine learning touch interpretation prediction model 220 can additionally or alternatively be configured to receive the determined touch prediction vector 228 as an input to assist the machine learning touch interpretation prediction model 220 in making related determinations of other touch interpretations (e.g., gesture interpretation 226, set of touch point interpretations 224, etc.).

[0100] The time step θ for which the touch prediction vector 228 is determined can be configured in various ways. In some implementations, the machine learning touch interpretation prediction model 220 can be configured to consistently output predicted future positions that define a predefined set of values of θ for one or more future times. In some implementations, one or more future times θ can be provided as a separate input to the machine learning touch interpretation prediction model 220 along with the original touch sensor data 224 - 228. For example, a computing device can input a time vector 230 to the machine learning touch interpretation prediction model 220 along with the original touch sensor data 224 - 228. The time vector 230 can provide a list of one or more time lengths (e.g., 10 ms, 20 ms, etc.) at which the touch interpretation prediction model is expected to predict the position of the user input object. Thus, the time vector 230 can describe one or more future times at which the position of the user input object is to be predicted. In response to receiving the original touch sensor data 224 - 228 and the time vector 230, the machine learning touch interpretation prediction model 220 can output a touch prediction vector 228 that describes the predicted future positions of the user input object for each time or time length described by the time vector 230. For example, the touch prediction vector 228 can include a pair of values for the position in the x - dimension and the y - dimension for each future time, or a pair of values for the change in the x - dimension and the y - dimension for each future time identified in the time vector 230.

[0101] Figure 6 and Figure 7 respectively depict a first aspect and a second aspect of an example use case according to an example embodiment of the present disclosure. Specifically, such aspects help to provide context for certain predicted touch interpretations based on the original touch sensor data. Figure 6Depicts a user holding a mobile device 250 (e.g., a smart phone) with a first hand 252 while providing user input to a touch-sensitive display screen 254 with a second hand 256. Desired input to the touch-sensitive display screen 254 of the mobile device 250 can be provided by fingers 258 of the second hand 256. Parts of the first hand 252 (i.e., the user's palm 260 and / or thumb 262) may sometimes provide undesired input to the touch-sensitive display screen 254 of the mobile device 250. Such undesired touch input may be particularly prevalent when the mobile device 250 has a very small bezel 264 around the perimeter. User interaction with the touch-sensitive display screen 254 thus includes touches generated from three different user input objects (e.g., fingers 258, palm 260, and thumb 262).

[0102] Figure 7 Depicts co-recorded screen content and associated overlaid raw touch sensor data associated with a keyboard application in which a user provides input as Figure 6 shown. Figure 7 Depicts potential touch sensor readings during a period when user input corresponds to typing the word "hello" in a screen keyboard provided on the touch-sensitive display screen 254. The desired input corresponds to touch sensor data detected at a first portion 270, while the undesired input corresponds to touch sensor data detected at a second portion 280 and a third portion 290. The desired input detected at the first portion 270 can include touch sensor data representing a first touch sequence 272, a second touch sequence 274, and a third touch sequence 276, where the first touch sequence 272 corresponds to the finger 258 moving from the "H" button to the "E" button in the screen keyboard, the second touch sequence 274 corresponds to the finger 258 moving from the "E" button to the "L" button, and the third touch sequence 276 corresponds to the finger 258 moving from the "L" button to the "O" button. The undesired input detected at the second portion 280 corresponds to the area in the lower left corner of the touch-sensitive display screen 254 where the user's palm 260 is resting. The undesired input detected at the third portion 290 corresponds to the area where the user's thumb 262 is resting.

[0103] Based on the Figure 6 and Figure 7The types of raw touch sensor data received during the user interactions depicted herein, such as the machine learning touch interpretation prediction models disclosed herein, can be trained to predict a set of touch points, where the set of touch points includes intended touch points associated with the raw touch sensor data of the first portion 270 and excludes non-intended touch points associated with the raw touch sensor data of the second portion 280 and the third portion 290. The machine learning touch interpretation prediction models disclosed herein can also be trained to predict a gesture interpretation of the raw touch sensor data of the first portion 270 while typing the word "hello". In some examples, the machine learning touch interpretation prediction models disclosed herein can also be trained to simultaneously generate a touch prediction vector predicting subsequent touch locations. For example, if sensor readings are currently only available for Figure 7 the first touch sequence 272 and the second touch sequence 274 depicted herein, then the touch interpretation prediction model can generate a touch prediction vector predicting the third touch sequence 276. The simultaneous prediction of the set of touch points, the gesture interpretation, and the predicted touch vector can all be generated based on the same set of raw touch sensor data provided as input to the touch interpretation prediction model.

[0104] Example Method

[0105] Figure 8 Depicts a flowchart of an example method 300 for performing machine learning in accordance with an example embodiment of the present disclosure.

[0106] At 302, one or more computing devices can obtain data describing a machine learning touch interpretation prediction model. The touch interpretation prediction model can have been trained to receive raw touch sensor data and generate one or more predicted touch interpretations as output. The touch interpretation prediction model can be or can otherwise include various machine learning models, such as neural networks (e.g., deep recurrent neural networks) or other multi-layer non-linear models, regression-based models, etc. The touch interpretation model can include at least one shared layer and multiple different and distinct output layers structurally located after the at least one shared layer. The touch interpretation prediction model for which data is obtained at 302 can include any of the features described with respect to Figures 4 - 5 the touch interpretation prediction models 200 and 220 or variations thereof.

[0107] At 304, one or more computing devices may obtain a first set of raw touch sensor data. The raw touch sensor data may indicate one or more positions of one or more user input objects relative to the touch sensor over time. At 306, one or more computing devices may input the raw touch sensor data obtained at 304 into a machine learning system of a touch interpretation prediction model. In some implementations, such as when the touch interpretation prediction model is configured to generate at least one touch prediction vector, one or more computing devices may optionally input time information identifying at least one future time into the touch interpretation prediction model at 308. In some implementations, the time information provided as an input at 308 may be in the form of a time vector describing one or more future times. One or more future times may be defined as a length of time relative to the current time and / or the time at which the touch sensor was sampled to obtain the touch sensor data at 304.

[0108] At 310, one or more computing devices may receive two or more predicted touch interpretations that are outputs of a touch interpretation prediction model. The two or more predicted touch interpretations may describe one or more predicted intents of one or more user input objects. In some implementations, the two or more predicted touch interpretations include a first predicted touch interpretation and a second different touch interpretation. In some examples, the two or more predicted touch interpretations received at 310 may include a set of touch point interpretations that respectively describe zero, one, or more touch points. The set of touch point interpretations may also include a touch type that describes the predicted type of user input object associated with each touch point and / or the estimated touch pressure at each touch point. In some examples, the two or more predicted touch interpretations received at 310 may include a gesture interpretation that characterizes at least a portion of a first raw touch sensor data set as a gesture determined from a predefined gesture category. In some examples, the two or more predicted touch interpretations received at 310 may include a touch prediction vector that describes one or more predicted future positions of one or more user input objects respectively for one or more future times. At 312, one or more computing devices may perform one or more actions associated with one or more of the predicted touch interpretations. In one example, the touch sensor data is used as an input to a virtual reality application. In such an instance, performing the one or more actions at 312 may include providing the output of the touch interpretation prediction model to the virtual reality application. In another example, the touch sensor data is used as an input to a mobile computing device (e.g., a smartphone). In such an instance, performing the one or more actions at 312 may include providing the output of the touch interpretation prediction model to an application running on the mobile computing device (e.g., a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.). In yet another example, performing the one or more actions at 312 may include providing one or more predicted touch interpretations to an application via an application programming interface (API).

[0109] Figure 9 FIG. depicts a flowchart of a first additional aspect of an example method 400 of performing machine learning in accordance with an example embodiment of the present disclosure. More specifically, Figure 9Describes the temporal aspects of providing an input to a touch interpretation prediction model and receiving an output therefrom according to an example embodiment of the present disclosure. At 402, one or more computing devices may iteratively obtain a time-stepped sequence of T touch sensor data readings, such that each of the T touch sensor readings includes touch sensor data indicative of one or more positions of one or more user input objects relative to the touch sensor at a given point in time. Each touch sensor reading obtained at 402 may be iteratively input by one or more computing devices at 404 into the touch interpretation prediction model as it is iteratively obtained. At 406, one or more computing devices may iteratively receive a plurality of predicted touch interpretations as an output of the touch interpretation prediction model.

[0110] Figure 10 Depicts a flowchart of a second additional aspect of an example method 500 of performing machine learning according to an example embodiment of the present disclosure. More specifically, Figure 10 Describes providing the output of a touch interpretation prediction model to one or more software applications using an API. At 502, a first application (e.g., a software application running on a computing device) may request access to a machine learning touch interpretation prediction model via an application programming interface (API) such that the touch interpretation prediction model or a portion thereof may be used within the first application. The machine learning touch interpretation prediction model may be hosted as part of a second application, or hosted within a dedicated layer, application, or component within the same computing device as the first application, or hosted in a separate computing device.

[0111] In some implementations, the first application may optionally provide definitions at 504 for one or more additional output layers of the touch interpretation prediction model. Providing such definitions may effectively add custom output layers to the machine learning touch interpretation prediction model that may be further trained on top of a pre-trained stable model. In some examples, an application programming interface may allow for providing a library of definitions at 504 from machine learning tools such as TensorFlow and / or Anao. Such tools may help specify additional training data paths for the specialized output layers of the touch interpretation prediction model. At 506, one or more computing devices may receive one or more predicted touch interpretations from the touch interpretation prediction model in response to a request via the API.

[0112] Figure 11 Depicts a flowchart of a first example training method 600 of a machine learning touch interpretation prediction model according to an example embodiment of the present disclosure. More specifically, at 602, one or more computing devices (e.g., within a training computing system) may obtain one or more training data sets each including a plurality of ground truth data sets.

[0113] For example, one or more training data sets obtained at 602 may include a first example training data set, where the first example training data set includes touch sensor data and corresponding labels describing a large number of previous observed touch interpretations. In one implementation, the first example training data set includes a first portion of data corresponding to recorded touch sensor data indicating the position of one or more user input objects relative to the touch sensor. The recorded touch sensor data may be recorded, for example, when the user is operating a computing device with a touch sensor under normal operating conditions. The first example training data set obtained at 602 may also include a second portion of data corresponding to labels of the determined touch interpretations applied to the recorded screen content. For example, the recorded screen content may be co-recorded simultaneously with the first portion of data corresponding to the recorded touch sensor data. The labels of the determined touch interpretations may include, for example, a set of touch point interpretations and / or gesture interpretations determined at least in part from the screen content. In some instances, the screen content may be manually labeled. In some instances, traditional heuristics applied to interpret the raw touch sensor data may be used to automatically label the screen content. In some instances, a combination of automatic and manual labeling may be used to label the screen content. For example, manual labeling may be used to best identify cases where the palm is touching the touch-sensitive display screen to prevent the user from using the device in the desired manner, while automatic labeling may be used to potentially identify other touch types.

[0114] In other implementations, a dedicated data collection application that prompts the user to perform certain tasks relative to the touch-sensitive display screen (e.g., asking the user to "click here" while holding the device in a specific manner) may be used to construct the first example training data set obtained at 602. The raw touch sensor data may be associated with the touch interpretations identified during the operation of the dedicated data collection application to form the first and second portions of the first example training data set.

[0115] In some implementations, when training a machine learning touch interpretation prediction model to determine a touch prediction vector, one or more training data sets obtained at 602 may include a second example training data set, where the second example training data set may include a first portion of data corresponding to an initial sequence of touch sensor data observations (e.g., Z1...Z T ) and a second portion of data corresponding to a subsequent sequence of touch sensor data observations (e.g., Z T+1 ...Z T+F ). Since the prediction step can be compared with the observation step when it occurs at a future time, the sequence of sensor data Z1...Z T is automatically suitable for training to predict future sensor data Z T+1 ...Z T+Fpredictions because the prediction step size can be compared to the observation step size when the future time occurs. It should be understood that any combination of the above techniques and other techniques can be used to obtain one or more training data sets at 602.

[0116] At 604, one or more computing devices can input a first portion of a training data set of ground truth data into a touch interpretation prediction model. At 606, one or more computing devices can receive, in response to receiving the first portion of the ground truth data, one or more predicted touch interpretations that are the output of the touch interpretation prediction model and that predict the remaining portion of the training data set (e.g., the second portion of the ground truth data).

[0117] At 608, one or more computing systems within the training computing system or otherwise can apply or otherwise determine a loss function, where the loss function compares the one or more predicted touch interpretations generated by the touch interpretation prediction model at 606 to the second portion of the ground truth data (e.g., the remaining portion) that the touch interpretation prediction model is attempting to predict. One or more computing devices can then backpropagate the loss function through the touch interpretation prediction model at 610 to train the touch interpretation prediction model (e.g., by modifying at least one weight of the touch interpretation prediction model). For example, the computing device can perform truncated backpropagation through time to backpropagate the loss function determined at 608 through the touch interpretation prediction model. Optionally, a number of generalization techniques (e.g., weight decay, dropout, etc.) can be performed at 610 to improve the generalization ability of the trained model. In some examples, the training process described in 602 - 610 can be repeated a number of times (e.g., until a target loss function no longer improves) to train the model. After the model has been trained at 610, the model can be provided to and stored at a user computing device for providing predicted touch interpretations at the user computing device. It should be understood that other training methods in addition to the loss function determined by backpropagation can also be used to train neural networks or other machine learning models for determining touch points and other touch interpretations.

[0118] Additional Disclosure

[0119] The techniques discussed herein relate to servers, databases, software applications, and other computer - based systems, as well as the actions taken and the information sent to and received from such systems. The inherent flexibility of computer - based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functions among and between components. For example, the processes discussed herein can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or can be distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0120] Although the subject matter has been described in detail with respect to various specific example embodiments thereof, each example has been provided by way of explaining the disclosure rather than limiting the disclosure. Those skilled in the art can readily generate substitutions, variations, and equivalents of such embodiments after understanding the foregoing. Accordingly, the subject matter disclosure does not exclude including such substitutions, variations, and / or additions of the subject matter, which will be apparent to those of ordinary skill in the art. For example, features shown or described as part of one embodiment can be used with another embodiment to yield yet another embodiment. Accordingly, it is intended that the disclosure cover such substitutions, variations, and equivalents.

[0121] Specifically, although Figures 8 to 11 the steps performed in a particular order have been described separately for purposes of illustration and discussion, the methods of the disclosure are not limited to the particular order or arrangement shown. The various steps of methods 300, 400, 500, and 600 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the disclosure.

Claims

1. A computing device for determining a touch interpretation from a user input object, comprising: At least one processor; A machine learning touch interpretation prediction model, where the machine learning touch interpretation prediction model has been trained based on a second training data set, and the second training data set includes: a first part of data corresponding to an initial sequence of touch sensor data observations, and a second part of data corresponding to a subsequent sequence of touch sensor data observations, where the touch interpretation prediction model has been trained to directly receive raw touch sensor data, the raw touch sensor data indicating measured touch sensor readings of a cross-point grid generated in response to one or more positions of one or more user input objects relative to the touch sensor at one or more times, and in response to receiving the raw touch sensor data, outputting a plurality of simultaneous predicted touch interpretations, each predicted touch interpretation corresponding to a different type of predicted touch interpretation based at least in part on the raw touch sensor data; and At least one tangible, non-transitory computer-readable medium storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including: Obtaining a first set of true first raw touch sensor data indicating one or more positions of one or more user input objects relative to the touch sensor over time; Inputting the first set of true first raw touch sensor data into the machine learning touch interpretation prediction model; Receiving, as an output of the touch interpretation prediction model, a plurality of simultaneous predicted touch interpretations describing the predicted intent of one or more user input objects, each predicted touch interpretation describing a different predicted aspect of the one or more user input objects, where the plurality of simultaneous predicted touch interpretations include: at least a first predicted touch interpretation and a second predicted touch interpretation determined from a group including a set of touch point interpretations respectively describing one or more intended touch points, a gesture interpretation characterizing the set of touch point interpretations as a gesture determined from a predefined gesture category, and a touch prediction vector respectively predicting one or more predicted future positions of a second set of true touch sensor data of the one or more user input objects at one or more future times in response to the first set of true first raw touch sensor data; and Performing one or more actions associated with the plurality of predicted touch interpretations.

2. The computing device according to claim 1, wherein, The plurality of predicted touch interpretations include at least a first predicted touch interpretation and a second predicted touch interpretation, where the first predicted touch interpretation includes a set of touch point interpretations respectively describing one or more intended touch points, and the second predicted touch interpretation includes a gesture interpretation characterizing the set of touch point interpretations as a gesture determined from a predefined gesture category.

3. The computing device according to claim 1, wherein, The machine learning touch interpretation prediction model includes a deep recurrent neural network having a plurality of output layers, each output layer corresponding to a different type of touch interpretation describing one or more predicted intents of one or more user input objects.

4. The computing device according to claim 1, wherein, Obtaining a first set of raw touch sensor data includes obtaining a first set of raw touch sensor data associated with one or more fingers or hand parts of a user or a stylus operated by the user, the first set of raw touch sensor data describing the position of one or more fingers, hand parts or the stylus relative to the touch-sensitive screen.

5. The computing device according to claim 1, wherein, Obtaining a first set of raw touch sensor data includes: obtaining a first set of raw time sensor data, wherein the first set of raw time sensor data provides at least one value describing a change in the position of one or more user input objects in the x dimension, at least one value describing a change in the position of one or more user input objects in the y dimension, and at least one value describing a change in time; or obtaining a first set of raw time sensor data, wherein the first set of raw time sensor data provides at least two values describing at least two positions of one or more user input objects in the x dimension, at least two values describing at least two positions of one or more user input objects in the y dimension, and at least two values describing at least two times.

6. The computing device according to claim 1, wherein, A machine learning touch interpretation prediction model has been trained based on a first training data set, wherein the first training data set includes: a first portion of data corresponding to recorded touch sensor data indicating the position of one or more user input objects relative to the touch sensor, and a second portion of data corresponding to labels of determined touch interpretations applied to the recorded screen content, wherein the first portion of data and the screen content are recorded simultaneously.

7. The computing device according to claim 1, wherein, Multiple predicted touch interpretations include a set of touch point interpretations respectively describing zero, one or more touch points.

8. The computing device according to claim 7, wherein, The set of touch point interpretations further includes one or more of a touch type describing the user input object of the predicted type associated with each touch point and the estimated touch pressure at each touch point.

9. The computing device according to claim 1, wherein, Multiple predicted touch interpretations include a gesture interpretation characterizing at least a portion of the first set of raw touch sensor data as a gesture determined from a predefined gesture category.

10. The computing device according to claim 1, wherein, Multiple predicted touch interpretations include touch prediction vectors describing one or more predicted future positions of one or more user input objects respectively for one or more future times.

11. The computing device according to claim 1, wherein: The machine learning touch interpretation prediction model includes at least one shared layer and a plurality of different and separate output layers structurally located after the at least one shared layer; and Receiving multiple predicted touch interpretations describing the predicted intent of one or more user input objects as an output of the touch interpretation prediction model includes receiving multiple predicted touch interpretations from the plurality of different and separate output layers of the machine learning touch interpretation prediction model.

12. One or more tangible, non-transitory computer-readable media, wherein the one or more tangible, non-transitory computer-readable media store computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: Obtain data describing a machine learning touch interpretation prediction model that has been trained based on a second training data set, where the second training data set includes: a first portion of data corresponding to an initial sequence of touch sensor data observations, and a second portion of data corresponding to a subsequent sequence of touch sensor data observations, where the touch interpretation prediction model has been trained to receive touch sensor data indicating one or more positions of one or more user input objects relative to a touch sensor at one or more times and, in response to receiving the touch sensor data, provide a plurality of simultaneous predicted touch interpretations, each predicted touch interpretation corresponding to a different type of predicted touch interpretation based at least in part on the touch sensor data; Obtain a first set of true first raw touch sensor data indicative of the positions of one or more user input objects relative to the touch sensor over time for a first portion; Directly input the first set of true first raw touch sensor data for the first portion into the machine learning touch interpretation prediction model; Receive a plurality of simultaneous predicted touch interpretations as the output of the touch interpretation prediction model, each predicted touch interpretation describing a different predicted aspect of one or more user input objects, where the plurality of simultaneous predicted touch interpretations includes: at least a first predicted touch interpretation and a second predicted touch interpretation determined from a group of touch point interpretation sets each describing one or more expected touch points, a gesture interpretation characterizing the touch point interpretation sets as a gesture determined according to a predefined gesture category, and a touch prediction vector predicting one or more predicted future positions of a second set of true touch sensor data for one or more user input objects at one or more future times in response to the first set of true first raw touch sensor data; and Perform one or more actions associated with the plurality of predicted touch interpretations.

13. One or more tangible, non-transitory computer-readable media according to claim 12, wherein, The plurality of predicted touch interpretations includes at least a first predicted touch interpretation and a second predicted touch interpretation, the first predicted touch interpretation including the first one of the touch point interpretation sets, the gesture interpretation determined from the predefined gesture category, and the touch prediction vector for one or more future times, and the second predicted touch interpretation including the second one of the touch point interpretation sets, the gesture interpretation determined from the predefined gesture category, and the touch prediction vector for one or more future times, where the second one is different from the first one.

14. One or more tangible, non-transitory computer-readable media according to claim 12, wherein, Performing one or more actions associated with the plurality of predicted touch interpretations includes providing, via an application programming interface (API), one or more of the plurality of touch interpretations to an application.

15. One or more tangible, non-transitory computer-readable media according to claim 12, wherein, Obtaining the first set of raw touch sensor data includes obtaining a first set of raw touch sensor data associated with one or more fingers or hand parts of a user or a stylus operated by the user, the first set of raw touch sensor data describing the positions of one or more fingers, hand parts, or the stylus relative to a touch-sensitive screen.

16. One or more tangible, non-transitory computer-readable media according to claim 12, wherein: Inputting the first set of raw touch sensor data into a machine learning touch interpretation prediction model includes inputting the first set of raw touch sensor data and a time vector describing one or more future times into a neural network associated with the machine learning touch interpretation prediction model; and Receiving a plurality of predicted touch interpretations as the output of the touch interpretation prediction model includes: receiving, as the output of the machine learning touch interpretation prediction model, one or more touch prediction vectors describing one or more predicted future positions of one or more user input objects respectively for one or more future times.

17. A mobile computing device for determining a touch interpretation from a user input object, comprising: Processor; At least one tangible, non-transitory computer-readable medium storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including: Obtaining data describing a machine learning touch interpretation prediction model including a neural network, wherein the machine learning touch interpretation prediction model has been trained based on a second training data set, wherein the second training data set includes: a first portion of data corresponding to an initial sequence of touch sensor data observations, and a second portion of data corresponding to a subsequent sequence of touch sensor data observations, wherein the touch interpretation prediction model has been trained to receive raw touch sensor data indicating one or more positions of one or more user input objects relative to the touch sensor at one or more times, and in response to receiving the raw touch sensor data, output a plurality of simultaneous predicted touch interpretations, each predicted touch interpretation corresponding to a different type of predicted touch interpretation based at least in part on the raw touch sensor data; and Obtaining a first set of true first raw touch sensor data associated with one or more user input objects, the first set of true first raw touch sensor data describing the positions of the one or more user input objects over time; Directly inputting the first portion of the true first raw touch sensor data set into the machine learning touch interpretation prediction model; and Receiving a plurality of simultaneous predicted touch interpretations as the output of the touch interpretation prediction model, the plurality of simultaneous predicted touch interpretations describing different predicted aspects of one or more user input objects, wherein the plurality of simultaneous predicted touch interpretations includes: at least a first predicted touch interpretation and a second predicted touch interpretation determined from a group including a set of touch point interpretations respectively describing one or more expected touch points, a gesture interpretation characterizing the set of touch point interpretations as a gesture determined according to a predefined gesture category, and one or more touch prediction vectors predicting the second portion of the true touch sensor data set of the one or more user input objects at one or more future times respectively in response to the first portion of the true first raw touch sensor data set; and Performing one or more actions associated with the plurality of predicted touch interpretations.

18. The mobile computing device according to claim 17, wherein, Two or more predicted touch interpretations also include touch prediction vectors describing one or more predicted future positions of one or more user input objects respectively for one or more future times.

19. The mobile computing device according to claim 17, wherein, Obtaining a first set of raw touch sensor data associated with one or more user input objects includes: Obtaining a first set of raw touch sensor data, wherein the first set of raw touch sensor data provides at least one value describing a change in the position of one or more user input objects in the x dimension, at least one value describing a change in the position of one or more user input objects in the y dimension, and at least one value describing a change in time; or Obtaining a first set of raw touch sensor data, wherein the first set of raw touch sensor data provides at least two values describing at least two positions of one or more user input objects in the x dimension, at least two values describing at least two positions of one or more user input objects in the y dimension, and at least two values describing at least two times.

Citation Information

Patent Citations

  • Detection of and response to extra-device touch events

    CN105074626A

  • Classification of user input

    CN105339884A

  • Systems and methods for providing response to user input using information about state changes predicting future user input

    CN105556438A

  • Predicting touch input

    US20140204036A1

  • Part and state detection for gesture recognition

    CN105051755A