Spatio-Temporal Deep Learning for Behavioral Biometrics
The spatio-temporal deep learning pipeline addresses the limitations of fixed-dimensional inputs in behavioral biometrics by processing raw data to learn and authenticate user behavior, enabling dynamic profiling and accurate authentication.
Patent Information
- Application Number
- JP2023535704
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-16
- Filing Date
- 2021-12-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-12-10
AI Technical Summary
Existing behavioral biometric systems require fixed-dimensional inputs and preprocessing to remove temporal characteristics, limiting their ability to handle variable-length input features and authenticating based on predefined patterns rather than unique user behavior.
A spatio-temporal deep learning pipeline that processes raw input data without preprocessing, utilizing multiple stages of machine learning models to learn and authenticate user behavior through spatio-temporal features, temporal features, and fixed/categorical features, generating a final output vector for user authentication.
Enables dynamic user profiling and authentication based on unique behavioral patterns, allowing for efficient adaptation to new users and minimizing information loss, while maintaining high authentication accuracy.
Smart Images

Figure 0007710520000001 
Figure 0007710520000002 
Figure 0007710520000003
Abstract
Description
Technical Field
[0001] This application generally relates to improved data processing apparatus and methods, and more particularly to mechanisms for performing spatio-temporal deep learning for behavioral biometrics.
Background Art
[0002] Physical biometrics typically involves the measurement and analysis of unique physical characteristics such as fingerprints, voiceprints, DNA, retinal patterns, etc. used to verify an individual's identifying information. Behavioral biometrics is a field of study that addresses the measurement of unique identification and measurable patterns in human activities. Examples of behavioral biometrics include keystroke dynamics where rhythms and timing patterns occur when a person selects keys on a keyboard device, gait analysis, computer mouse usage characteristics, and signature analysis. Behavioral biometrics is used in many industries for secure authentication, including financial institutions, commerce, government facilities, and retail point of sale (POS) devices.
Summary of the Invention
[0003] This summary is provided to introduce in a simplified form a selection of concepts that are further described herein in the detailed description. This summary is not intended to identify key factors or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0004] In one aspect of the present invention, there is provided a method in a data processing system including at least one processor and at least one memory, the at least one memory including instructions executed by the at least one processor to specifically configure the at least one processor to implement a behavioral biometrics deep learning (BBDL) pipeline including multiple stages of a machine learning computer model operating based on spatio-temporal input data to provide a behavioral biometrics-based authentication mechanism. According to one exemplary embodiment, the method includes receiving spatio-temporal input data corresponding to inputs associated with an entity over a predetermined time window including a plurality of time intervals from one or more sensors, each time interval having a corresponding subset of the spatio-temporal input data. The method also includes, at each of the plurality of time intervals, processing, by one or more machine learning computer models of the corresponding stage in the multiple stages, the subset of the spatio-temporal input data corresponding to the time interval to generate an output vector having a value indicative of an internal representation of spatio-temporal traits of the entity represented in the subset of the spatio-temporal input data. Further, the method includes accumulating the output vectors over the multiple stages of the BBDL pipeline to generate a final output vector including a final output vector value indicative of the spatio-temporal traits of the entity represented in the spatio-temporal input data. Further, the method includes authenticating the entity by the data processing system based on the final output vector.
[0005] The spatio-temporal input data includes input data that is spatially dependent and temporally dependent with respect to a physical area monitored by at least one of one or more sensors. Thus, the BBDL pipeline of an exemplary embodiment does not require preprocessing of the raw spatio-temporal input data to generate a fixed predetermined set of input features and remove temporal characteristics, such as in existing mechanisms. Also, by processing the spatio-temporal input data, the BBDL pipeline can learn the behavioral biometrics of an entity (user) rather than simply authenticating a pattern of input that may be provided by any entity or entities that notice a particular pattern. That is, the spatio-temporal input data represents the unique behavior exhibited by the entity / user, which is presented in the way the entity / user generates the input and is not just a detected pattern.
[0006] Preferably, the present invention provides a method, wherein authenticating an entity includes comparing a final output vector with at least one previously generated output vector generated by a BBDL pipeline stored in a user profile for an authenticated user to determine the probability that the final output vector represents the spatio-temporal characteristics of the authenticated user. Also, authenticating an entity further includes controlling access to protected resources according to the result of the comparison. In this way, a small set of previously generated output vectors may be used to define a user profile representing the behavioral biometrics of an authenticated user, and subsequent user inputs may be authenticated against such behavioral biometric data.
[0007] Preferably, the present invention provides a method comprising: controlling access to a protected resource by permitting a user to access the protected resource in response to determining that the probability that the final output vector represents the user's spatio-temporal characteristics that match the spatio-temporal characteristics of the authenticated user; and updating the user profile to include the final output vector. Thus, the user profile may be dynamically updated to reflect the latest behavioral biometric information about the authenticated user.
[0008] Preferably, the present invention provides a method, wherein each stage of the BBDL pipeline includes an image processing machine learning computer model, and at each stage of the BBDL pipeline, processing a subset of spatio-temporal input data corresponding to a time interval includes: processing the subset of spatio-temporal input data to convert spatio-temporal features in the subset of spatio-temporal input data into an image; and performing image analysis on the image by the image processing machine learning computer model of the stage to generate a first vector output. Thus, the spatio-temporal features can be represented in a format in which image analysis by a machine learning model can be performed to classify spatio-temporal data regarding various spatio-temporal features.
[0009] Preferably, the present invention provides a method, wherein each stage of the BBDL pipeline includes a fully connected neural network machine learning computer model, and at each stage of the BBDL pipeline, processing a subset of spatio-temporal input data corresponding to a time interval includes: processing the subset of spatio-temporal input data to identify temporal features in the subset of spatio-temporal input data; and processing the temporal features by the fully connected neural network machine learning model of the stage to generate a second vector output. Therefore, the temporal features can be identified in the spatio-temporal input data, in which case the temporal features are not influenced by specific spatial elements and can be evaluated separately from the spatio-temporal features before being combined with the results of analyzing the spatio-temporal features.
[0010] Preferably, the present invention provides a method, wherein each stage of the BBDL pipeline includes fixed / category data embedding logic, and at each stage of the BBDL pipeline, processing a subset of spatio-temporal input data corresponding to a time interval includes processing the subset of spatio-temporal input data to identify fixed / category features in the subset of spatio-temporal input data, and processing the fixed / category features by the fixed / category data embedding logic of the stage to generate a third vector output. Thus, the mechanism of the exemplary embodiment can further perform the evaluation of the input data based not only on the fixed / category data, but also on spatio-temporal features and temporal features in the spatio-temporal input data.
[0011] Preferably, the present invention provides a method, wherein each stage of the BBDL pipeline includes combination logic that operates to combine a first output vector, a second output vector, and a third output vector to generate a combined output vector that is input to the main stage machine learning computer model of the stage. Thus, the results of processing spatio-temporal features, temporal features, and fixed / category features in the spatio-temporal input data are combined to generate an internal representation of the spatio-temporal features of the input for use in predicting whether the spatio-temporal input data is associated with an authenticated user.
[0012] Preferably, the present invention provides a method, wherein each main-stage machine learning computer model of each stage of the BBDL pipeline processes the corresponding combined output vector of this corresponding stage together with the input from the previous stage of the BBDL pipeline to generate a stage output vector, and accumulating the output vectors over multiple stages of the BBDL pipeline to generate a final output vector includes, for each stage, outputting the stage output vector as the input to the next stage in the BBDL pipeline, and the stage output vector of the main-stage machine learning computer model of the last stage of the BBDL pipeline is the final output vector. Thus, the processing of spatio-temporal features, temporal features, and fixed / categorical features over multiple time intervals of a time frame is accumulated such that the behavioral biometric pattern is distinguishable within the time frame.
[0013] Preferably, the present invention provides a method, wherein the embedded device information is input into the main-stage machine learning computer model of the first stage of the BBDL pipeline, and the main-stage machine learning computer model processes the combined output vector associated with the first stage together with the embedded device information to generate a stage output vector for the first stage. Thus, the evaluation of the BBDL pipeline can be customized for a specific device associated with the spatio-temporal input data.
[0014] Preferably, the present invention provides a method, wherein one or more sensors include touch sensors of a touch sensitive display device, and the spatio-temporal input data includes sensor data indicating the characteristics of the touch input of an entity. For example, the mechanisms of the exemplary embodiments can be implemented in conjunction with touch sensitive display devices on many modern electronic devices such as smartphones and tablet computers.
[0015] Preferably, the present invention provides a method, which provides a computer program product including a computer-usable or readable medium having a computer-readable program. When the computer-readable program is executed on a computing device, it causes the computing device to perform various ones and combinations of the operations outlined above with respect to the method of the exemplary embodiments.
[0016] According to another aspect of the present invention, a system / apparatus is provided. The system / apparatus may include one or more processors and a memory coupled to the one or more processors. The memory may include instructions that, when executed by the one or more processors, cause the one or more processors to perform various ones and combinations of the operations outlined above with respect to the method of the exemplary embodiments.
[0017] These and other features and advantages of the present invention will become those described in the following detailed description of the exemplary embodiments of the present invention or will be apparent to those skilled in the art upon consideration of the detailed description.
[0018] The present invention, and the preferred modes of this use, as well as further objects and advantages, will be best understood by reference to the following detailed description of the exemplary embodiments when read in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0019]
Figure 1
Figure 1E
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
[0020] Behavioral biometrics is an authentication paradigm that enables continuous authentication as opposed to one-time authentication typically used by physical biometrics and knowledge-based mechanisms such as password-based mechanisms, secret question-based mechanisms, and authentication mechanisms. For example, on a smartphone or tablet device, different users may have different swipe speeds and pinch angles, etc., that can be monitored by a behavioral biometrics mechanism. There may be a mechanism for using behavioral biometrics as an authentication technique, due to the limitations of these existing approaches that cannot handle variable-length input features, as these existing mechanisms only consider pre-defined features with fixed dimensions that eliminate the time dimension.
[0021] That is, existing mechanisms require fixed dimensions to train computer models to make predictions or classifications, while the temporal aspects of behavioral biometrics are variable. Thus, the temporal aspects of raw input data are reduced from the input data processed by a computer model through preprocessing or precomputation on the raw input data, so that high-level features specifically defined by feature engineering to distinguish entities from each other must be generated. For example, a swipe angle feature can be defined that involves calculating the swipe angle and representing it as a boolean value according to a certain threshold. This can also further be used to determine whether the angular rotation or circular rotation in the swipe, for example, is 1 if the swipe is a corner swipe or 0 if it is not a corner swipe, or 1 if it is a circular rotation or 0 if it is not a circular rotation, as represented by the swipe angle. Such features may be able to distinguish users by the type of rotation used when the user performs a swipe gesture, but the features themselves do not contain any temporal aspects or specific spatial information.
[0022] This is preprocessed data that includes features that distinguish high-level entities. The data is generated by converting the raw input data into a fixed-dimensional input with features that distinguish pre-defined entities that have had the variable timing aspects, spatial aspects, or both of the raw input data that has been removed, i.e., processed by an existing behavioral biometric computer model. Such high-level features with fixed dimensions are required due to the limitations of existing behavioral biometric computer models that require fixed-dimensional inputs. Thus, existing mechanisms do not perform these behavioral biometric classification operations based on raw spatio-temporal input data because they cannot take into account the variability of the time dimension regarding behavioral biometrics.
[0023] As an example, the raw input data from a sensor device regarding a user's stroke or swipe on a touch-sensitive screen, etc., can be represented as a spatio-temporal function F(t, x, y) = v (alternatively, F(t) = (x, y, v)) that indicates the value of the perceived characteristics of the stroke / swipe at the coordinates (x, y) at time t, where x is the horizontal coordinate of the screen, y is the vertical coordinate of the screen, t is the time point, and v is the value of the perceived characteristic of the stroke or swipe, such as pressure. Existing mechanisms cannot process such variable inputs because fixed-dimensional computer models have to be trained using fixed-dimensional feature input data. Instead, the input data has to be transformed into fixed-dimensional user-differentiated features through feature extraction transformations performed by preprocessing of the input data, which includes, for example, preprocessing the input such that F(x, y) = 1 or 0 depending on whether the screen was touched at the coordinates (x, y), removing the time dimension of the input, and there are boolean values such as changing F(t, x, y) = v to a boolean input value. That is, through feature extraction performed based on feature engineering, the feature extraction may generate F1(session) = max|F(t, x, y) - avg(F(t, x, y))| - min|F(t, x, y) - avg(F(t, x, y))|, whereby one or more of the dimensions t or x and y or combinations thereof are efficiently excluded. max(F - avg(F)) conveys how large a certain feature during a session can be, such as how much faster a swipe can be made compared to the average of the session, where a large number means a large deviation. This equation can be made more explicit by using the 2-norm instead of the absolute value, for example, F1(session) = max||F(t, x, y) - avg(F(t, x, y))||_2 - min||F(t, x, y) - avg(F(t, x, y))||_2, where "||_2" indicates the 2-norm.
[0024] Such preprocessed input generates fixed-dimensional values representing high-level features of the input, such as "this user forms acute angles", "this user zooms in quickly / slowly", "this user zooms in and out alternately / repeatedly", etc., which are further fed into a computer model that classifies or makes predictions based on the high-level features when spatio-temporal input data has already been removed from the input.
[0025] Exemplary embodiments provide a behavioral biometrics deep learning (BBDL) computer model pipeline (hereinafter referred to as the BBDL pipeline), including a machine learning image analysis computer model, a feature classification machine learning model, and a feature embedding mechanism, etc., that operate to learn patterns of spatio-temporal characteristics of raw spatio-temporal input data. This does not require removing temporal or spatial characteristics or both to generate high-level features as in existing mechanisms. These learned patterns of spatio-temporal characteristics indicate the behavioral biometrics of a specific user and are used to generate a user profile based on the learned machine learning pattern of spatio-temporal features, which can further be used to authenticate subsequent user inputs regarding the behavioral biometrics represented by the spatio-temporal features of the subsequent user inputs. It should be recognized that the raw spatio-temporal input to the BBDL pipeline is raw input data that is not designed to have features that distinguish it as a unique entity in itself, i.e., the raw input data is preprocessed to perform feature extraction based on feature engineering and still contains raw temporal and spatial information. That is, unlike existing mechanisms where the input must operate based on specifically engineered features such that the computer model differentiates unique entities, the input to the BBDL pipeline includes raw, unprocessed spatio-temporal input data from one or more sensor devices over a specific slide time frame for any raw input data that occurs for a particular embodiment.
[0026] The time frames during which raw input data is collected and input into the BBDL pipeline are sliding or moving time frames where each individual time interval can be overlapping or non - overlapping. For example, if the time frame includes time points T1, …, T12, the individual time intervals (or time slices) can be (T1, T2), (T2, T3), …(T11, T12) or (T1, T3), (T2, T4), (T3, T5), …(T10, T12) depending on whether overlapping time intervals are desired for a particular implementation. (Assumed to be a human user for illustrative purposes, but not so limited) Observable input events (detectable by one or more sensors) from an entity can occur at any time during T1 to T12 in this example. That is, if T1 = 12:00:10 (hour:minute:second), T2 = 12:00:15 (based on actual elapsed time), and a swipe event can occur at 12:00:12 which can be associated with the first time interval (T1, T2). By overlapping time intervals, the possibility of losing information is minimized. For example, two events at 12:00:11 and 12:00:13 can be considered to fall into the same time interval without overlapping, but two events at 12:00:14 and 12:00:16 can be considered to fall into two different time intervals despite having the same time difference.
[0027] Therefore, the BBDL pipeline operates based on raw spatio-temporal input data that has no direct correlation with specific high-level features of entity behavior. As a result, the BBDL pipeline mechanism, as will be described in more detail later, can be dynamically trained for new classifications such as new users based on only a very small set of examples, such as embedding two swipe sessions and comparing the similarities of those embeddings. In contrast, existing mechanisms require that a computer model be trained against a fixed set of classifications or that separate pre-trained computer models be trained against different classifications, and if a new classification is desired, the entire computer model must be retrained with the new classification or a new computer model must be generated.
[0028] According to one exemplary embodiment, the BBDL pipeline includes multiple stages of computer logic, each stage implementing logic for a set of machine learning models (or "sub-models" if the entire BBDL pipeline is considered a computer "model") that evaluate spatio-temporal features, temporal features, and fixed / categorical features at specific corresponding time intervals including time points t, t+1, t+2, up to t+n, where n indicates the size of the time frame in which spatio-temporal feature patterns are evaluated to determine behavioral biometrics for a particular user. Thus, the BBDL pipeline will have n stages, each stage having multiple machine learning (ML) computer models (or "sub-models" of the overall BBDL pipeline model) that evaluate the spatio-temporal features, temporal features, and fixed / categorical features of user input to a computing device to which the user is interface-connected, such that the features can be generated by one or more sensors of a user input device associated with the computing device, or by the computing device itself, or both. For example, the multiple sensors can be provided in association with one or more user-operated devices such as smartphones, tablet computers, touch screen devices, computer mice, trackball input devices, image capture devices, biometric reader devices, or any other source of biometric or behavioral biometric input data.
[0029] For example, in some exemplary embodiments, these sensors may include touch sensors such as a touch-sensitive screen, an accelerometer associated with the device itself, an ambient light sensor, and a camera device whose input can be expressed in terms of F(t, x, y). Examples of spatio-temporal features sensed by such sensors and provided as data for input to the BBDL pipeline include, but are not limited to, finger swipe data on the touch screen including coordinate data for specific touch screen dimensions, such as swipe size, swipe speed, swipe acceleration, swipe direction, etc., pressure information regarding touch pressure on the touch screen device, mouse over information, detected gesture information from an imaging device (such as a camera), the touch screen, or any other gesture information source, voiceprint information associated with speech input, and acceleration information representing the movement of a particular device by the user. In one exemplary embodiment, the spatio-temporal features used include Boolean touch (see subsequent description), touch size (e.g., how large is the region around the point (x, y) when the value of F(t, x, y) is 1), touch pressure, (x, y) touch speed (judged from a sequence of points {(t1, x1, y1), (t2, x2, y2),...} where the speed is {(t1, (x2 - x1) / (t2 - t1), (y2 - y1) / (t2 - t1),...} and F(t, x, y) and F’(t, x, y) are constructed from these values), (x, y) touch acceleration, scalar touch speed, and scalar touch acceleration. Although Cartesian coordinates are used in the examples herein, exemplary embodiments are not limited thereto and any other coordinate system may be utilized without departing from the scope of the invention. For example, it should be recognized that F(t, x, y) may be transformed to F(t, phi, r) to represent similar information (touch size, speed, acceleration, etc.) in a polar coordinate system.
[0030] Generally, spatio-temporal features are a set of features that vary with respect to the device's screen dimensions or graphical user interface dimensions relative to the movement of the device's input range, for example, a touch screen device, or other user input devices such as a computer mouse, trackball, or graphical user interface, representing user input, and that vary over time. For example, boolean spatio-temporal touch features are a set of features representing which parts of the device screen are touched at a particular time step. Formally, this can be represented as $f_{Boolean touch}(x, y, t)$, where the value is 1 if (x, y) on the device screen is touched at time step t on the device screen, and 0 otherwise. Thus, for each value of (x, y) at time step t, the value for the spatio-temporal feature is either 1 or 0, thereby generating a matrix or bitmap representation of the user input. In the case of the function F(t, x, y) = v, the value for v may not be limited to boolean values, but instead may be within a predefined range for the value of v depending on the perceived input, for example, v can be from max(v) to min(v) (e.g., 1 to -1).
[0031] The use of spatio-temporal features according to an exemplary embodiment ignores the location with respect to the input range of the device and is significantly different from previous gesture-based authentication mechanisms where the gesture of the user action changes over time (previous gesture-based authentication mechanisms relate only to high-level features that exclude spatial or temporal characteristics or both, as discussed previously). Previous gesture-based authentication mechanisms store a fixed set of gestures and authenticate based on whether the same or a different set of gestures has been input, such that if the input has exactly the same set of gestures, the user is authenticated, and if not, the user is not authenticated. Therefore, anyone who knows the stored fixed pattern can succeed in authentication. In contrast, the exemplary embodiment uses raw input data and operates based on the subtle differences in user (entity) behavior that are not associated with a fixed pattern. That is, the subtle differences in user behavior provide an expression of the spatio-temporal characteristics of a particular user that the user uses with various inputs of the user, such as the drawing speed of a curve, as well as spatio-temporal user input characteristics such as zoom angle and tendency. These characteristics are captured with respect to temporal information and are uniquely identifiable to a particular user, rather than a fixed pattern of non-spatio-temporal features that can be replicated by any user who knows the fixed pattern of the input.
[0032] In an exemplary embodiment, time characteristics that are independent of the input range associated with a particular device for which spatio-temporal information or non-spatio-temporal information associated with user input can be aggregated are considered. For example, changes in an accelerometer are independent of the specific input range of the device, e.g., the dimensions of the touch screen of the device where the user provides input, but simply proportional to the previous acceleration measurement. That is, the change in acceleration is the same regardless of the input range of the particular device. Formally, $f_{accelerometer-x}(t)$ is the change in the accelerometer along the x-axis at time step t. Such time characteristics include, but are not limited to, the values of a 3D accelerometer in Euclidean / polar coordinates, the first derivative of acceleration, the second derivative of acceleration, the aggregated values of spatio-temporal characteristics, and context embeddings where components of, e.g., a graphical user interface or a touch screen or both, are touched.
[0033] In addition to the spatio-temporal feature input data and the temporal feature input data, the exemplary embodiments further obtain data from information sources that do not depend on biometric data, behavioral biometric data, or other data associated with the user's input. For example, other features are identification information of the user operation device itself, such as the manufacturing and model numbers (e.g., Apple's iPhone(R) X, Samsung's Galaxy(R) S20, etc.), the configuration of the device (e.g., the type of sensors and their characteristics, etc.), the dimensions of the device or the limits of operation (e.g., the size of the screen, the size of the touch panel, the size of the graphical user interface, etc.), and such user input specific data that does not depend on a specific application to which the entity / user connects the interface to provide the input. These features are features that do not change over time, that is, they are fixed features, but nevertheless are useful in combination with spatio-temporal features and temporal features to authenticate the user's identification information. Further examples of fixed features that can be used include configuration information indicating a range of measurable values. For example, different devices may measure inputs using different ranges of values or different units, etc.
[0034] These features may be embedded as one or more vector representations of the device itself. Any feature by category can be embedded, and the embedded vectors are provided as inputs to the rest of the stage network or the entire BBDL pipeline or both. Examples of other features that can be embedded in this way include, but are not limited to, volume, speakerphone on / off, camera state, camera flash state, active / touched application, touched user interface element, device name / type, operating system version, etc.
[0035] Sources of feature data, such as spatio-temporal and temporal features, may continuously or periodically collect or generate or both feature data at multiple time points or time intervals, and provide the feature data to a preprocessor that operates to convert the raw feature data into a form usable by various machine learning models at stages of a BBDL pipeline, such as a time×2D feature matrix or map for spatio-temporal features, or a time×1D matrix or map for temporal features. Further, for any fixed device features or category type features, the embedding of these features may be performed by the preprocessor or other embedding components, and the embedded vector representation may be provided together with the output of the neural network for downstream processing as will be discussed hereinafter. Embedding is a process of generating a vector representation of the input values. In a neural network mechanism, embedding is a process for generating a low-dimensional learned continuous vector representation of discrete variables. The process of embedding an input into a vector representation is generally known in the art.
[0036] In some exemplary embodiments, each stage of the BBDL pipeline performs image analysis or computer-vision-based analysis of spatio-temporal feature data received from various sources and preprocessed by a preprocessor into a time×2D feature matrix or map representation through a machine learning process trained to be one or more first ML computer models (or "sub-models"). The one or more first ML computer models may include a single ML computer model trained to classify or categorize multiple different spatio-temporal features, may be a single ML computer model trained to classify or categorize a single spatio-temporal feature, or may be a combination of multiple ML computer models each trained to perform classification or categorization for different spatio-temporal features. In the case of multiple ML computer models operating on spatio-temporal feature data, the results of each ML computer model may be aggregated into a final vector output representation through merge logic that outputs the final vector output representation to another neural network or machine learning model that generates an output to the next stage of the BBDL pipeline, as will be described later.
[0037] Since the training of the ML computer models at each stage of the BBDL pipeline can be done as part of the overall training operation of the BBDL pipeline, it is not necessary for each individual model (or "sub-model") at each stage to be pre-trained, although it should be recognized that this can be done in some exemplary embodiments. Thus, the BBDL pipeline (or "model") can be trained, including training the computer models (or "sub-models") at each stage, as part of the overall training operation of the BBDL pipeline. In some embodiments, one or more of the stage ML computer models or "sub-models" may be pre-trained or trained as part of a separate machine learning-based training operation, if desired, and it should be recognized that training data for training the individual ML computer models or "sub-models" is available. For ease of explanation, the ML computer models at each stage of the BBDL pipeline will hereinafter be referred to as ML computer models rather than "sub-models", and the BBDL pipeline will be referred to as the BBDL pipeline rather than the BBDL pipeline ML computer model. However, it should be recognized that the BBDL pipeline is an overall computer model having various stages where each stage includes one or more "sub-models" of the overall BBDL pipeline ML computer model.
[0038] Spatio-temporal feature data may be represented as a time x 2D feature matrix or map representation that essentially provides an image processable by image analysis, or a computer vision neural network, or other image / vision analysis machine learning models, and generates an output vector representing the categorization of the input image. Examples of image / vision analysis machine learning models or "sub-models" at each stage of the overall BBDL pipeline computer model include AllConvNet, ResNet, Inception, and Xception. The image / vision analysis machine learning model of the exemplary embodiments is specifically trained through a machine learning process to classify or categorize spatio-temporal matrices or map representations related to specific behavioral biometrics categories. Various categorizations for various spatio-temporal features for various time x 2d feature matrices or maps may be combined to generate a vector representation of the categorization of spatio-temporal features present in the input at a particular point in time or time interval.
[0039] In some exemplary embodiments of the present invention, the final logit layer of the image / vision analysis machine learning model may not be used. That is, in a neural network, a layer further causes features to be processed by a dense layer so as to cause a vector to be produced with the same dimension as the number of output classes. This is referred to as the logit layer, and the final activation is applied in addition to the output generated by a logic layer, for example, a softmax activation. Since these final two layers, namely, the logit and activation layers, are not used in the image / vision analysis of the exemplary embodiments in some exemplary embodiments, the input to the logit layer is used as the internal representation of the image for the next stage in the BBDL pipeline.
[0040] The temporal features may be represented as a time × 1D matrix or map that is input into one or more second neural networks or trained machine learning computer models that generate a vector output representing a categorization of the temporal features. In some exemplary embodiments, the one or more second neural networks or trained machine learning computer models can be one or more trained fully connected neural networks or other dense layer neural networks. The one or more second neural networks or trained machine learning computer models output one or more vector representations of the classification of the input of the temporal features at time intervals, which are combined with a machine learning model that processes spatio-temporal features and other embedded features for fixed or categorical, i.e., non-time-dependent features, to generate a final vector output that is input into a neural network or machine learning computer model (hereinafter referred to as the main stage machine learning (ML) model) that generates an output to the next stage of the BBDL pipeline.
[0041] It should be recognized that the output of the models at each stage, and thus the internal representations learned as part of training the overall BBDL pipeline by each main-stage ML model. That is, since the BBDL pipeline attempts to classify whether two input sessions are from the same user, the models at each stage will attempt to extract features that best represent the user's input image. Since these outputs of the computer models at that stage are internal representations, these outputs are not explicitly / separately trained to generate a particular feature, but rather are trained as part of the overall machine learning-based training of the entire BBDL pipeline. The dimension of the output of each main-stage ML model is a certain spatio-temporal characteristic of the entity (user). For example, assume that dimension 1 represents the spatio-temporal characteristic of a user who performs several consecutive and rapid zooms in and out to check the details of something on the smartphone screen (as opposed to a single zoom in), while dimension 2 represents the characteristic of a single zoom in. In both dimensions, 1 means that the executor of this input has that characteristic, and 0 means that the executor does not have it. The BBDL pipeline accumulates the detected characteristics of the main-stage ML models over a time frame. For example, the first main-stage ML model may detect a zoom in, the second main-stage ML model may detect a zoom out, the third main-stage ML model may detect another zoom in, and the fourth main-stage ML model may detect a zoom out. Since these detections by the various stages of the BBDL pipeline are accumulated over a time frame, the BBDL pipeline determines that dimension 1 is 1 (not a single zoom-in characteristic).
[0042] In one exemplary embodiment, the main-stage ML model for the first stage of the BBDL pipeline receives, as input, embeddings of device information, or other fixed or categorical information, or combinations thereof, not provided at each time step. In subsequent stages of the BBDL pipeline, the main-stage ML model of the previous stage is combined with the main-stage ML model of the next stage and provides its output as input to the main-stage ML model of the next stage. In one exemplary embodiment, the main-stage ML model for each stage of the BBDL pipeline may be a recurrent neural network (RNN), such as a long short-term memory (LSTM) RNN.
[0043] At each stage along the BBDL pipeline, the ML model of the corresponding stage performs a similar function as above, except for specific spatio-temporal features, temporal features, and other fixed / categorical features at the corresponding time point or time interval. The resulting output vector of the ML model of that stage is combined in the current stage's ML model with the input from the main-stage ML model of the previous stage as feature inputs to be post-processed by the main-stage ML model of that stage to generate the vector output that is input to the main-stage ML model of the next stage. This process continues until the final stage of the BBDL pipeline, which then generates an output vector K that represents the user input over the time frame for the spatio-temporal input, temporal input, and fixed / categorical input. The value of "K" represents the number of values included in the output vector that represents the characteristics of the user input. Since this output vector K is unique to the behavioral biometric characteristics of a particular user, it can be used to authenticate the user.
[0044] In one exemplary embodiment, the training of the machine learning models for the stages of the BBDL pipeline may be performed using any suitable machine learning process, such as a supervised or unsupervised machine learning process. It should be recognized that the BBDL pipeline machine learning training is performed with respect to the entire BBDL pipeline, which includes multiple machine learning models at various stages.
[0045] Generally, machine learning is related to the design and development of techniques that employ input empirical data (such as network statistical data and performance evaluation metrics), and recognizes complex patterns in these data. One pattern among machine learning techniques is the use of a base model M, where, given input data, the parameters are optimized to minimize a cost function associated with M. For example, in the context of binary classification, the model M could be a straight line that separates data into two classes (e.g., labels), such as M = a * x + b * y + c, and the cost function could be the number of misclassified points. The learning process further operates by adjusting the parameters a, b, c such that the number of misclassified points is minimized. After this optimization phase (or learning phase), the model M is available for classifying new data points. Of course, this is a simple example for binary classification, and other models for classification into more than two classes will use other similar techniques.
[0046] According to an exemplary embodiment, the machine learning process generally involves learning how to represent the behavioral biometrics of a particular user with respect to the spatio-temporal features, temporal features, and fixed / categorical features of user input. The machine learning process generally involves, in each iteration of each epoch of machine learning training, obtaining a pair of user-session information representing user input over a particular period or time frame. The pair of user-session information involves a first set of session information corresponding to the target user. The second session information in this pair can be from the same user or a different user. Corresponding label information is obtained for the second session information indicating whether the user is the same user or a different user. The first session information and the second session information are each processed separately through the BBDL pipeline to generate corresponding K-dimensional output vectors. The K-dimensional output vector for the first session information is compared with the K-dimensional output vector for the second session information to determine whether the session information is classified as being from the same user or a different user by the BBDL pipeline, i.e., to determine whether the session information for both sessions is similar enough to represent the same user.
[0047] This resulting determination can then be compared with the label for the second session information to determine whether the second session information is actually from the same user or a different user. If the similar result is inaccurate based on the label of the second session information, the operating parameters of the machine learning model are adjusted to reduce the error at the value of K generated by the BBDL model. That is, if the similar comparison indicates that the users for the two sessions are the same user and they are not the same as indicated by the label of the second session information, the operating parameters of the machine learning model are adjusted such that the generated value of K is different for different users. If the similar comparison indicates that the users for the two sessions are not the same user and they are the same as indicated by the label of the second session information, the operating parameters of the machine learning model are adjusted such that the generated value of K is more similar for different sessions by the same user. If the similar comparison indicates that the users for the two sessions are the same user and they are the same as indicated in the label of the second session information, no adjustment is necessary. It should be recognized that this process may be repeated for different second session information for the same user or different users or both, for example, in the training data provided for training the machine learning model of the BBDL pipeline.
[0048] Once trained, the BBDL pipeline may be executed against one or more user sessions of a target user to generate an authenticated user profile that stores the output of one or more K values of the BBDL pipeline that represent user input of an authenticated user for a particular device. For example, session history information of a user may be obtained for a plurality of user sessions, each of which may be processed by the BBDL pipeline to generate a K-dimensional output vector for each session stored in the user profile. This group of user sessions, i.e., the user session history, may be a continuously updated set of user session histories, e.g., the last 10 sessions may be maintained, and the corresponding user profile may be dynamically updated as new sessions occur by the user. That is, for each subsequent user session, the BBDL pipeline is executed against the input data from the user session to generate the corresponding K-dimensional output vector, which is then used to replace the oldest K-dimensional output vector entry in the user profile. Thus, a dynamic user profile is maintained for the authenticated user.
[0049] Whether the new user input information for a new session is actually from an authenticated user can be determined in a manner similar to that previously described for the training of the BBDL pipeline by determining whether the K-dimensional output vector for the new user input information is sufficiently similar to the K-dimensional output vector information stored in the authenticated user's profile. In this case, "sufficiently similar" may be determined with respect to a predetermined threshold indicating an acceptable amount of difference, while still indicating that the two K-dimensional output vectors represent the same user (entity). In response to a determination that the comparison between the user's profile and the K-dimensional output vector for the new user input information is not sufficiently similar, the user's access to computing resources, physical facilities, or any other resources protected by the behavioral biometrics authentication mechanism of the exemplary embodiment is denied. Further, during runtime operation, a determination that the new user information for a new session is from an authenticated user may cause the update engine to update the user's profile with the K-dimensional output vector for this new session, in addition to granting access to the protected resources, such as by replacing the oldest K-dimensional output vector entry in the user's profile if a predetermined number of K-dimensional output vector entries exist in the user's profile.
[0050] Therefore, the exemplary embodiments provide a mechanism for implementing a spatio-temporal deep learning mechanism for behavioral biometrics. The exemplary embodiments provide a pipeline of stages that include a set of machine learning models that classify or categorize input data at corresponding points in time or time intervals with respect to spatio-temporal features, temporal features, and fixed / categorical features. In some exemplary embodiments, the spatio-temporal features are rendered as image data and then processed by a trained image analysis or computer vision analysis machine learning model, and the temporal features are processed through fully connected or dense layer-based or both machine learning models. The stages of the pipeline operate based not only on the user input data at their corresponding points in time or time intervals but also on the output from previous stages in the pipeline to ultimately generate a K-dimensional output vector that represents the behavioral biometrics of the user input over the time frame or period being processed by the pipeline. This can be used to generate a user profile, which can be used in subsequent user sessions to authenticate the user.
[0051] Before continuing to discuss various aspects of the exemplary embodiments and the improved computer operations performed by the exemplary embodiments, it should first be recognized that throughout this description, the term "mechanism" is used to refer to elements of the present invention that perform various operations and functions, etc. "Mechanism", as the term is used herein, can be an implementation form of the function or aspect of an exemplary embodiment in the form of an apparatus, procedure, or computer program product. In the case of a procedure, the procedure is implemented by one or more devices, apparatuses, computers, or data processing systems, etc. In the case of a computer program product, the logic represented by computer code or instructions embodied in or on the computer program product is executed by one or more hardware devices to implement the functionality associated with the specific "mechanism" or to perform operations. Thus, the mechanisms described herein may be implemented as dedicated hardware, implemented as software executed on hardware to configure the hardware to implement the specific functionality of the present invention that otherwise cannot be executed by the hardware, implemented as software instructions stored on a medium such that the instructions are easily executable by the hardware to specifically configure the hardware to execute the described functionality and the specific computer operations, functions, procedures, or methods described herein, or implemented as any combination of the above.
[0052] In the specification and claims of the present invention, the terms "a", "at least one of", and "one or more of" may be utilized with respect to specific features and elements of exemplary embodiments. It should be recognized that these terms and phrases are intended to indicate that there is at least one of the specific features or elements in a particular exemplary embodiment, but that there may also be more than one. That is, these terms / phrases are not intended to limit the present specification or claims to a single such feature / element, nor are they intended to require the presence of multiple such features / elements. Instead, these terms / phrases only require that there be at least a single such feature / element, with the possibility that multiple such features / elements are within the present specification and claims.
[0053] Also, when the term "engine" is used herein in connection with describing embodiments and features of the present invention, it should be recognized that it is not intended to limit any particular implementation for achieving or performing or both of actions, steps, processes, etc. that may be attributable to or performed by or both of an engine. An engine may be, but is not limited to, software, hardware, or firmware or combinations thereof, or, but is not limited to, a general-purpose processor or a dedicated processor or both, in combination with appropriate software loaded or stored in a machine-readable memory and executed by a processor, for performing the specified function of any of these combinations. Further, any name associated with a particular engine is, unless otherwise specified, for convenience of reference and is not intended to be limited to a particular implementation. Additionally, any functionality attributable to an engine may be equally performed by multiple engines, incorporated into or combined with or both the functionality of another engine of the same type or a different type, or may be distributed across one or more engines of various configurations.
[0054] Furthermore, in the following description, it should be recognized that multiple different examples are used for various elements of the exemplary embodiments in order to further illustrate the exemplary implementations of the exemplary embodiments and to aid in the understanding of the mechanisms of the exemplary embodiments. These examples are non-limiting and are not intended to cover all the various possibilities of implementing the mechanisms of the exemplary embodiments. From the perspective of this specification, it will be apparent to those skilled in the art that there are many other alternative implementations for these various elements that can be used in addition to or instead of the examples provided herein without departing from the scope of the present invention.
[0055] The present invention may be a system, a method, or a computer program product or a combination thereof. The computer program product may include a computer-readable storage medium (s) having computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0056] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or raised structures in grooves in which instructions are recorded, and any suitable combination of the foregoing. A computer-readable storage medium, as used herein, should not be construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0057] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network or a wireless network, or a combination thereof. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.
[0058] Computer-readable program instructions for carrying out the operation of the present invention may be written in any combination of one or more program languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Java®, Smalltalk®, or C++, and conventional procedural programming languages such as the “C” programming language or similar programming languages, either in source code or object code. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to carry out aspects of the present invention.
[0059] Aspects of the present invention will be described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0060] These computer readable program instructions are provided to a computer processor or other programmable data processing apparatus to produce a machine, such that the instructions executed via the computer processor or other programmable data processing apparatus implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, programmable data processing apparatus, or other device to function in a particular manner, such that the medium storing the instructions comprises an article of manufacture including instructions which implement the function / act specified in one or more blocks of the flowchart and / or block diagram.
[0061] The computer readable program instructions may also be loaded onto a computer, other programmable apparatus, or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram by causing a series of operational steps to be performed on the computer, other programmable apparatus, or other device.
[0062] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions that include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may be performed out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order depending upon the functionality involved. It should also be noted that each block of the block diagrams or flowchart diagrams, or combinations of blocks in the block diagrams or flowchart diagrams or both, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or by a combination of dedicated hardware and computer instructions.
[0063] As described above, the mechanisms of the exemplary embodiments provide a Behavioral Biometric Deep Learning (BBDL) pipeline that includes multiple stages of a machine learning computer model that processes input data regarding spatio-temporal features, temporal features, and fixed / categorical features at various time steps or time intervals of a time frame to characterize that input data as a vector representation that can be used to generate a user profile for authenticating a user, compare to a user profile, or both. The BBDL pipeline includes a data preprocessor that takes raw behavioral biometric data and converts that data into a format usable by the machine learning models of the stages of the BBDL pipeline. For example, the preprocessor converts spatio-temporal features present in the raw input data into a time x 2D feature matrix or map. The preprocessor also converts temporal features of the raw input data into a time x 1D feature matrix or map. Further, the preprocessor may, by itself or in conjunction with a separate embedding engine, embed the fixed / categorical input data into a vector representation of embedded values for use in the output of other machine learning models at corresponding stages of the BBDL pipeline logic.
[0064] Figures 1(A) - 1(D) show exemplary plots of spatio-temporal features that can be used for user authentication according to one exemplary embodiment. The plots of these spatio-temporal features in these examples are referred to as time x 2D feature plots and represent two-dimensional behavioral biometric features of user input over time. Each time x 2D feature plot in Figures 1(A) - 1(D) includes a time dimension along a horizontal axis along which spatio-temporal features are plotted as points along a vertical axis representing a first characteristic and 2D characteristics of other characteristics of the two-dimensional (2D) characteristics as a hatch or color. The particular 2D characteristics plotted are dependent on the particular spatio-temporal features being plotted. As can be seen from Figures 1(A) - 1(D), the plots of the spatio-temporal features result in an image representing the spatio-temporal features over time.
[0065] In the time×2D features shown in FIGS. 1(A) to 1(D), they represent spatio-temporal features associated with user input of a swipe on the touch screen of a user operation device such as a smartphone or a tablet computing device. FIG. 1(A) is a time series of a time×2D feature plot for user input of a swipe. The x-axis and y-axis correspond to the (x, y) coordinates on the touch screen of the user operation device, thereby constituting a 2D plot. Each plot represents the user's input at a specific point in time or time interval. The shading of the pixels in the 2D feature plot represents the value of the feature at a specific coordinate. In this case, the specific value and the meaning of the value depend on the specific feature being represented. For example, a boolean touch may be 1 or 0 (represented by two different colors) at a specific coordinate, and the speed may be a real number represented by different shadings on a spectrum such as a possible shading from white to black.
[0066] The time×2D feature plot in FIG. 1(B) represents the 2D characteristics of the size of the plotted swipe versus time. In the plot of FIG. 1(B), the 2D characteristic may be a value indicating the pixel region or area around a certain pixel touched by the user at a specific point in time or time interval.
[0067] The time×2D feature plot in FIG. 1(C) represents the 2D characteristics of the plotted swipe speed versus time. In the plot of FIG. 1(C), the 2D characteristic may include the swipe speed which can be the x-component of the speed. Another plot (not shown) may represent the y-component of the speed. The shading in the plot may represent positive / negative values where the symbol is for a specific axis. In the shown figure, the shading of the pixels indicates that the left half of the movement is going from left to right (positive direction with respect to the x-axis) and the right half is going from right to left (negative direction with respect to the x-axis). Thus, the shown movement is a pinch gesture such as when the user is trying to zoom out on the touch screen. Similarly, as described above, the time×2D feature plot in FIG. 1(D) represents the 2D characteristics of the plotted swipe acceleration versus time.
[0068] Raw sensor data is provided as raw input data to a preprocessor that performs operations to transform the spatio-temporal features into a time × 2D matrix or map that is an image representing the individual spatio-temporal features. The image data generated by the preprocessor as a result of the generation of this matrix or map of time × 2D feature plots can then be input into the corresponding stage of the BBDL pipeline's image analysis / computer vision machine learning model. There may be a single image analysis / computer vision ML model for each spatio-temporal feature, or there may be a single image analysis / computer vision ML model that addresses image analysis / computer vision operations for multiple spatio-temporal feature matrices / maps.
[0069] Furthermore, as previously described, the preprocessor may also transform the temporal features in the input data into a time × 1D matrix or map. These are features in the input data that do not vary across the input range of the device, e.g., across the location of a touch screen, but vary over time. Thus, the time × 1D feature matrix or map may be used to plot the time and one-dimensional features of these temporal features, i.e., the values of the corresponding temporal features. FIG. 1E shows an exemplary plot of temporal features that can be used for user authentication according to one exemplary embodiment. Further, the horizontal axis represents time and the vertical axis represents the value of a temporal feature such as the x-axis of an accelerometer recording. The dashed box represents a sliding time window for sampling the temporal features. The time × 1D matrix or map may be input into a fully connected or dense layer-based machine learning model that generates a vector output representing the temporal features. The output of the fully connected / dense layer-based machine learning model may be combined with the output of the image analysis / computer vision machine learning model and the embedding of the fixed / categorical input data to generate a combined vector output representation of the user input at a point in time or time interval, which can also be provided as input for processing to the main stage ML model.
[0070] FIG. 2 is an exemplary block diagram of an action biometrics deep learning (BBDL) pipeline according to one exemplary embodiment. As shown in FIG. 2 and previously described, BBDL pipeline 200 includes a plurality of stages 210-230, each stage 210-230 including a set of machine learning models 212-218, 222-228, and 232-238 that evaluate spatio-temporal features, temporal features, and fixed / categorical features at a specific corresponding time interval including one or more time points, e.g., t, t+1, t+2, up to t+n, where n indicates the size of the time frame over which the spatio-temporal feature pattern is evaluated to determine the action biometrics for a particular user. Thus, BBDL pipeline 200 will have n stages, and each stage 210-230 has a plurality of trained machine learning (ML) computer models 212-218, 222-228, and 232-238 that evaluate spatio-temporal features, temporal features, and fixed / categorical features of user input received via an input device associated with a computing device to which the user is interface-connected.
[0071] User input may be captured or collected by one or more sensors associated with a user input device associated with a computing device, or may be generated by the computing device itself, or both. That is, according to one or more exemplary embodiments, a plurality of sensors 210-214 are provided in connection with one or more user operation devices (not shown) such as a smartphone, a tablet computer, a touch screen device, a computer mouse, a trackball input device, an image capture device, a biometrics reader device, or any other source of biometrics or behavioral biometrics input data. For example, in some exemplary embodiments, these sensors 201-204 may include touch sensors in a touch-sensitive screen and accelerometers associated with the device itself. Examples of spatio-temporal features sensed by such sensors 201-204 and provided as part of input data 205-208 for input to the BBDL pipeline 200 include, but are not limited to, finger swipe data on a touch screen including coordinate data for a particular touch screen dimension, e.g., swipe size, swipe speed, swipe acceleration, swipe direction, etc., pressure information regarding touch pressure on a touch screen device, mouse over information, detected gesture information from an imaging device (such as a camera), a touch screen, or any other gesture information source, voiceprint information associated with speech input, and acceleration information representing the movement of a particular device by the user. In one exemplary embodiment, the spatio-temporal features used include boolean-type touches (see the following description), touch size, (x, y) touch acceleration, scalar-type touch speed, and scalar-type touch acceleration.
[0072] Generally, spatio-temporal features change with the movement of a user input device, such as the input ranges of user interface devices 292-294 or computing device 290 or both, e.g., a touch screen device, or other user input devices representing user input, such as a computer mouse, trackball, or graphical user interface, and change with time, and are a set of features that change with the screen dimensions of the device or the dimensions of the graphical user interface with respect to the movement of the user input device. Further, as described above, an exemplary spatio-temporal feature is a set of features representing which part of the device screen was touched at a particular time step, and may be a boolean spatio-temporal touch feature that is an actual value feature having various values along a range of potential values, e.g., pressure values. However, it should be recognized that spatio-temporal features are not limited to features that are only perceivable by sensors of a user input device or a computing device or both. In some exemplary embodiments, other spatio-temporal features may also be evaluated, such as any spatio-temporal feature corresponding to a physical area monitored by a sensor from which spatio-temporal data can be obtained. For example, in the case of ambient light values, the sensor may detect light levels over a time frame and over a particular physical area, and this light level data may represent spatio-temporal input data.
[0073] In exemplary embodiments, temporal features that do not depend on the input ranges associated with particular user input devices 292-295 or computing device 290 are also considered. As described above, these temporal features may be aggregated spatio-temporal information or non-spatio-temporal information associated with user input. Such temporal features include, but are not limited to, values of a three-dimensional accelerometer in Euclidean coordinates / polar coordinates, the first derivative of acceleration, the second derivative of acceleration, aggregated values of spatio-temporal features, and context embeddings in which components of, e.g., a graphical user interface or a touch screen or both, are touched.
[0074] Furthermore, for any fixed device feature or category type feature, the embedding of these features may be performed by preprocessors 280, 282, 284 corresponding to specific stages 210 - 230, or other embedded components 216, 226, 236 of stages 210 - 230. The embedded components 216, 226, and 236 may be machine learning models trained to embed the corresponding fixed / category features into a vector representation that represents these fixed / category features.
[0075] It should be recognized that sensors 201 - 204 provide raw input data 205 - 208 including spatio-temporal features, temporal features, and fixed / category features for each of a plurality of time points or time intervals. These time points or time intervals together represent the time frame of the user input. The time frame of the user input may be a moving or sliding frame such that sensors 201 - 204 continue to monitor the user input during the user session. Thus, at time point / interval t + n + 1, the time frame processing is initiated in the processing of stage 210, i.e., time point / interval t + n + 1 becomes time point / interval t for the new time frame. The time intervals may be overlapping or non - overlapping as previously described.
[0076] Each raw input data 205 - 208 is provided to preprocessors 280, 282, 294 corresponding to specific stages 210 - 230 of the logic. The preprocessors 280, 282, 284 generate a time × 2D feature matrix / map, a time × 1D matrix / map, and fixed / category features, which are provided to the corresponding machine learning models of stages 210 - 230 of the logic as previously explained.
[0077] In addition to user input data, device data 250 may be obtained from and embedded in the computing device 290, the user interface device, or both. For example, other device characteristics that do not depend on such user input specific data, such as identification information of the user operation device itself, such as the manufacturing and model numbers (e.g., Apple's iPhone(R) X, Samsung's Galaxy(R) S20, etc.), the configuration of the device (e.g., the type of sensors and their characteristics, etc.), and the dimensions or operating limits of the device (e.g., the size of the screen, the size of the touch panel, the size of the graphical user interface, etc.), are embedded as a vector representation (embedding) 260 and provided as input to the main stage ML model 218 of the first stage 210 of the BBDL pipeline 200.
[0078] The sensors 201 - 204 of the user input devices 292 - 294, computing device 290, or both, may continuously or periodically collect, generate, or both, raw input data 205 - 208 including spatio - temporal features, temporal features, and fixed / categorical features at multiple points in time or time intervals, and provide the feature data to the corresponding pre - processors 280 - 284 associated with stages 210 - 230 of the BBDL pipeline 200. The pre - processors 280 - 282 operate to convert the raw user input data into a form usable by the various machine learning models 212 - 218, 222 - 228, and 232 - 238 of the BBDL pipeline 200 at stages 210 - 230, e.g., a time × 2D feature matrix or map for spatio - temporal features, a time × 1D matrix or map for temporal features, and fixed / categorical features provided for embedding. The pre - processors 280 - 282 are configured to identify different types of features present in the raw input data 205 - 208. That is, the pre - processors 280 - 284 may parse the raw input data 205 - 208 and, based on its structure, categorize portions of the raw input data 205 - 208 with respect to whether those portions represent spatio - temporal features, temporal features, or fixed / categorical features. Based on the categorization of portions of the raw input data 205 - 208, the corresponding processing by the processors 280 - 284 is performed to generate a time × 2D matrix / map, a time × 1D matrix / map, and identification information for fixed / categorical features, which can be input for embedding. The pre - processed user input data 240 - 244 is then provided to the logic of the corresponding stages 210 - 230 for processing by the machine learning models of the corresponding stages 210 - 230.
[0079] As shown in FIG. 2, each stage 210-230 of the BBDL pipeline 200 includes one or more first ML computer models 212, 222, 232 trained through a machine learning process to perform image analysis or computer vision-based analysis of the spatio-temporal feature data present in the pre-processed user input data 240-244 for the corresponding time points / slices. That is, the one or more first ML computer models 212, 222, 232 may include a single ML computer model trained to perform classification or categorization of multiple different spatio-temporal features, or may be a single ML computer model trained to perform classification or categorization of a single spatio-temporal feature, or may be a combination of multiple ML computer models each trained to perform classification or categorization regarding different spatio-temporal features. In the case of multiple ML computer models operating based on spatio-temporal feature data, the results of each ML computer model may be combined by the combination logic 217, 227, 237 of stages 210-230 after concatenating or aggregating the outputs of the various ML computer models 212, and may be aggregated into a single final vector output representation to the main stage ML model 218 of the BBDL pipeline 200 through a merge logic (not shown) that can generate a single final vector output representation.
[0080] As discussed above, the spatio-temporal feature data may be represented as a time × 2D feature matrix or map representation that basically provides an image that can be processed by image analysis, or a computer vision neural network, or other image / vision analysis machine learning models, and generates an output vector representing the categorization of the input image. The image / vision analysis machine learning models of the exemplary embodiments are specifically trained through a machine learning process to classify or categorize the spatio-temporal matrix or map representation regarding specific behavioral biometric categories. The various categorizations for the various time × 2D feature matrices / maps for the various spatio-temporal features may be combined or aggregated to generate a vector representation of the categorization of the spatio-temporal features present in the input at a specific point in time or time interval.
[0081] The temporal characteristics in the preprocessed user inputs 240 - 244 may be represented as a time x 1D matrix or map that can be input into one or more second neural networks or trained machine learning computer models 214, 224, 234 that generate a vector output representing the categorization of the temporal characteristics. In some exemplary embodiments, the one or more second neural networks or trained machine learning computer models 214, 224, 234 may be one or more trained fully-connected neural networks or other dense layer neural networks. The one or more second neural networks or trained machine learning computer models 214, 224, 234 output one or more vector representations of the classification of the temporal characteristic inputs at time intervals, which are combined by combination logic 217, 227, 237 with the output from machine learning models 212, 222, 232 that process spatio-temporal characteristics, and with other embedded features generated by embedding logic 216, 226, 236 for fixed or category-specific, i.e., non-time-dependent, features, to generate a final vector output that is input into main stage machine learning model (ML) models 218, 228, 238 that generate an output to the next stage of the BBDL pipeline 200.
[0082] For the primary stage ML model 218 of the first stage 210 of the BBDL pipeline 200, as input, it receives an embedding 268 of device information or other fixed or categorical information 250, or a combination of both, which is not provided at each time step. In subsequent stages 220-230 of the BBDL pipeline 200, the primary stage ML model of the previous stage is combined with the primary stage ML model of the next stage and provides its output as input to the primary stage ML model of the next stage. For example, the output of the primary stage ML model 218 is input to the primary stage ML model 228 of stage 220, and the output of the primary stage ML model 228 is provided as input to the primary stage ML model 238. In one exemplary embodiment, each of the primary stage ML models 218, 228, 238 of stages 210-230 of the BBDL pipeline 200 may be a recurrent neural network (RNN), such as a long short-term memory (LSTM) RNN.
[0083] At each stage 210-230 along the BBDL pipeline 200, the ML models 212-218, 222-228, and 232-238 of the corresponding stages perform similar functions as above, except for specific spatio-temporal features, temporal features, and other fixed / categorical features at the corresponding time points or time intervals. The output vector resulting from the ML model of that stage is combined, in the current stage's ML model, with the input from the primary stage ML model of the previous stage as feature input to be post-processed by the primary stage ML model of that stage, to generate a vector output that is input to the primary stage ML model of the next stage. This process continues until the final stage of the BBDL pipeline 200, for example, stage 230, which then generates an output vector K270 representing the user input over a time frame related to the spatio-temporal input, temporal input, and fixed / categorical input. In this way, the BBDL pipeline 200 accumulates the characteristics of the entity (user) represented in the input at the time points represented in the time frame. The resulting output vector K270 is unique to the specific user's behavioral biometric features (or user characteristics) and can thus be used to authenticate the user.
[0084] The training of the machine learning model at the stage of the BBDL pipeline 200 may be performed using any suitable machine learning process, e.g., a supervised or unsupervised machine learning process. FIG. 3 is an exemplary block diagram showing a high-level representation of a training process for training the BBDL pipeline 200 according to one exemplary embodiment. In the exemplary embodiment, the machine learning process generally involves learning how to represent the behavioral biometrics of a particular user with respect to the spatio-temporal features, temporal features, and fixed / categorical features of those user inputs. The machine learning process generally involves obtaining, for each iteration of each epoch of machine learning training, a pair of user session information 310 and 320 representing user inputs over a particular period or time frame. The pair of user session information 310 and 329 is accompanied by a first set of session information 310 corresponding to the target user, such as raw input data 205-208 obtained from various sensors or computing devices or both, and device information 250 about the user input device / computing device. The second session information 320 in this pair may be from the same user or a different user and may include other raw input data 305-308 and corresponding device information 250. Corresponding label information 322 is obtained for the second session information 320 indicating whether the user is the same user or a different user. In some embodiments, each of the user session information 310 and 320 may have label data 312, 322 specifying the identification information of the user providing the user session information 310 and 320, so that by comparing the user identification information, it is possible to determine whether the users are the same or different for training purposes.
[0085] The first user session information 310 and the second user session information 320 are each processed separately through a model 340 including a BBDL pipeline 200 after preprocessing by a preprocessor 330 that may include preprocessors 280-284 for generating corresponding K-dimensional output vectors 342, 344 such as the K-dimensional output vector 270 in FIG. 2. The K-dimensional output vector 342 for the first session information is compared with the K-dimensional output vector 344 for the second session information when determining whether the session information is classified by the model 340 including the BBDL pipeline 200 as being from the same user or a different user, i.e., when determining whether the session information for both sessions is similar enough to represent the same user.
[0086] The K-dimensional output vector K1 342 and the K-dimensional output vector K2 344 are compared by comparison logic 350 to determine whether the K1 output vector 342 and the K2 output vector 344 are similar enough (the difference in the individual values of the K1 and K2 output vectors is below a predetermined or learned threshold) to conclude that the user providing the user session information 310 is the same user or a different user as the user providing the user session information 320. Depending on the result of the comparison, a probability value is generated indicating the likelihood that the user providing the user session information 310 is the same as the user providing the user session information 320, which can then be compared with a threshold probability. If the calculated probability meets or exceeds the threshold probability, it is predicted that the user providing the user session information 310 is the same user as the one who provided the user session information 320.
[0087] Once a prediction is obtained, the accuracy of the prediction is then checked by the training logic 360 to determine whether the BBDL model 340, including the BBDL pipeline 200, has generated accurate results for the user session information 310, 320. For example, the labels 312, 322 for the user session information 310, 320 can be compared, or in an embodiment where only the second user session information 320 is labeled, the label 322 can be compared to the identification information of the certified user for which the model 340 is being trained. The comparison made by the training logic 360 determines whether the second session information is actually from the same user or a different user. If the similar predictions are inaccurate based on the label 312 or 322 or both of the user session information 310, 320, the training logic 360 adjusts the operating parameters of the machine learning model of the BBDL pipeline 200 of the BBDL model 340 to reduce the error (loss or cost) at the value of K generated by the BBDL model 340.
[0088] It should be recognized that the training of the BBDL model 340 and the BBDL pipeline 200 depends on the user in that the BBDL pipeline 200 and the BBDL model 340 are trained to accurately generate K-vector outputs for instances of the same user and different users. In training, the BBDL pipeline 200 and the BBDL model 340 correctly generate K-vector outputs that are sufficiently the same for the same user or sufficiently different for different users to determine whether the prediction of user authentication is accurately made based on those behavioral biometric data present in the user input. Instead of having to train separate machine learning models for different users, the mechanism of the exemplary embodiment can provide a single BBDL pipeline 200 and BBDL model 340 that can generate a profile of an authenticated user by operating based on the input of the authenticated user. Thereafter, the BBDL pipeline 200 and the BBDL model 340 can operate based on new user input to generate a K-vector output that can then be compared with the profile of the authenticated user stored to authenticate whether the user input is from the authenticated user.
[0089] That is, the BBDL model 340 including the BBDL pipeline 200, when trained, may be executed against one or more user sessions of a target user to generate a profile of the authenticated user that stores the output of one or more K values of the BBDL pipeline representing the user input of the authenticated user for a particular device. FIG. 4 is an exemplary block diagram showing a high-level representation of the runtime operation of the BBDL pipeline according to one exemplary embodiment. As shown in FIG. 4, user session history information 410 may be obtained for a plurality of user sessions, each of which may be processed by the BBDL pipeline 200 of the corresponding preprocessor 420 and model 430 to generate a K-dimensional output vector for each session stored in the user profile 440. This group of user sessions, i.e., the user session history, may be a continuously updated set of user session history 410. For example, the last 10 sessions may be maintained, and the corresponding user profile 440 may be dynamically updated when a new session by the user occurs. That is, for each subsequent user session, the model 430 including the BBDL pipeline 200 is executed against the input data from the user session to generate the corresponding K-dimensional output vector, which is then used to replace the oldest K-dimensional output vector entry in the user profile 440. Thus, a dynamic user profile 440 is maintained for the authenticated user.
[0090] Whether the new user input information for a new session is actually from an authenticated user may be determined in the same manner as previously described for the training of model 340 and BBDL pipeline 200. As one exception to the training operation, during runtime operation after training, if it is determined that the user for a subsequent session is not the authenticated user itself, or the session information does not match the registered authenticated user of the specific device from which the session information is obtained, the permission to access protected resources may be denied. That is, the new user input information for a new session 450, processed through the preprocessor 420 and model 430 including BBDL pipeline 200, may result in a K-dimensional output vector 460 that can be compared with the K-dimensional output vectors stored in the user profile 440. To determine whether the K-dimensional output vector 460 sufficiently matches any of the K-dimensional output vectors stored in the user profile 440, the K-dimensional output vector 460 may be compared with each of the respective K-dimensional output vectors. If the K-dimensional output vector 460 sufficiently matches any one of these stored K-dimensional output vectors in the user profile 440, it is determined that the user providing the new session information 450 is an authenticated user and access to protected resources may be permitted. If the K-dimensional output vector 460 does not sufficiently match any of the K-dimensional output vectors stored in the user profile, the user providing the new session information 450 may not be an authenticated user, and thus access to protected resources may be denied. The determination that the new session information 450 is from an authenticated user, in addition to the fact that access to protected resources will be permitted, may cause the update engine 480 to update the user profile 440 with the K-dimensional output vector 460 for this new session information 450, such as by replacing the oldest K-dimensional output vector entry in the user profile 440 if a predetermined number of K-dimensional output vectors exist in the user profile 440, or by adding a new entry if the maximum number of entries has not been reached.
[0091] Thus, as previously described, the exemplary embodiments provide a mechanism for implementing a spatio-temporal deep learning mechanism for behavioral biometrics. The exemplary embodiments provide a pipeline of stages that include a set of machine learning models that classify or categorize input data at corresponding points in time or time intervals with respect to spatio-temporal features, temporal features, and fixed / categorical features. In some exemplary embodiments, the spatio-temporal features are rendered as image data and then processed by a trained image analysis or computer vision analysis machine learning model, and the temporal features are processed through a fully connected or dense layer-based or both machine learning models. The stages of the pipeline operate based not only on the user input data at their corresponding points in time or time intervals but also on the output from previous stages in the pipeline to ultimately generate a K-dimensional output vector representing the behavioral biometrics of the user input over the time frame or period being processed by the pipeline. This can be used to generate a user profile, which can be used in subsequent user sessions to authenticate the user.
[0092] FIG. 5 is a flowchart outlining an exemplary operation of a BBDL pipeline according to one exemplary embodiment. The operations outlined in FIG. 5 may be performed by the BBDL pipeline as part of its runtime operation, assuming that the BBDL pipeline has been trained through a machine learning process as previously described above.
[0093] As shown in FIG. 5, the operation starts by receiving user input for a new session (step 510). The user input is preprocessed (step 520) to identify and preprocess the spatio-temporal features, temporal features, and fixed / categorical features present in the received user input. This preprocessing may include not only the generation of a time×2D feature matrix / map and a time×1D feature matrix / map, but also identifying the fixed / categorical features present in the received user input. The user input is processed by the corresponding stage of the BBDL pipeline at each point in the time frame represented by the user input, and the results of that stage are combined (step 530) to generate a K-vector output representation of the behavioral biometrics of the user input.
[0094] The identification information of the authorized user for the device is determined, and the corresponding user profile for the authorized user is retrieved (step 540). The identification information of the user may be specified in the user identifier during a user input, such as a logon process that can be used to retrieve the user profile for the identified user. In other cases, the identification information of the device may be used to perform a lookup operation in the registry of authorized users for a particular device.
[0095] A comparison of the K-vector entries in the user profile with the K-vector output representation of the user input is performed (step 550) to determine whether the vectors are similar enough to indicate the same user or not similar enough to represent different users (step 560). If the user is predicted to be the same user, access to the protected resource is permitted (step 570) and the user profile is updated with the K-vector output for the user input (step 580). If the user is predicted to be different, access to the protected resource is denied (step 590). The operation then ends.
[0096] Exemplary embodiments may be utilized to protect various types of computing resources or physical facilities or devices, etc. based on behavioral biometrics based on the analysis of spatio-temporal features, temporal features, and fixed / categorical features of behavioral biometric inputs in many different types of data processing environments. The data processing environment may include a single computing device or a distributed data processing environment or both. FIGS. 6 and 7 are provided hereinafter as exemplary environments in which aspects of the exemplary embodiments may be implemented. It should be recognized that FIGS. 6 and 7 are merely examples and are not intended to imply or claim any limitation with respect to the environments in which aspects or embodiments of the present invention may be implemented. Many modifications to the shown embodiments may be made without departing from the scope of the present invention.
[0097] FIG. 6 shows a graphical representation of an exemplary distributed data processing system in which aspects of the exemplary embodiments may be implemented. The distributed data processing system 600 may include a network of computers in which aspects of the exemplary embodiments may be implemented. The distributed data processing system 600 contains at least one network 602 that is a medium used to provide communication links between various devices and computers connected together within the distributed data processing system 600. The network 602 may include connections such as wired communication links, wireless communication links, or fiber optic cables.
[0098] In the example shown, servers 604 and 606 are connected to network 602 along with storage unit 608. Further, clients 610, 612, and 614 are also connected to network 602. These clients 610, 612, and 614 may be, for example, personal computers or network computers. In the example shown, server 604 provides data such as boot files, operating system images, and applications to clients 610, 612, and 614. Clients 610, 612, and 614 are clients with respect to server 604 in the example shown. Distributed data processing system 600 may include additional servers, clients, and other devices not shown.
[0099] In the example shown, distributed data processing system 600 is the Internet, which represents a worldwide collection of networks and gateways where network 602 uses the Transmission Control Protocol / Internet Protocol (TCP / IP) suite to communicate with each other. At the core of the Internet is a backbone of high-speed data communication lines between major nodes or host computers, which consists of thousands of commercial, government, educational, and other computer systems that route data and messages. Of course, distributed data processing environment 600 may also be implemented to include several different types of networks, such as, for example, an intranet, a local area network (LAN), or a wide area network (WAN). As described above, FIG. 6 is intended to be an example and not an architectural limitation for different embodiments of the present invention. Accordingly, the specific elements shown in FIG. 6 should not be regarded as limitations regarding the environment in which exemplary embodiments of the present invention may be implemented.
[0100] As shown in FIG. 6, one or more computing devices, such as server 604, may be configured to implement an Action Biometrics Deep Learning (BBDL) model that includes a BBDL pipeline, such as BBDL model 340 of FIG. 3, specifically including BBDL pipeline 200 of FIG. 2. Configuring the computing device may include providing application-specific hardware or firmware, etc., to facilitate the performance of the operations and the generation of the output described herein with respect to the exemplary embodiments. Configuring the computing device may further or alternatively include providing a software application stored in one or more storage devices and loaded into the memory of the computing device, such as server 604, to configure the one or more hardware processors of the computing device to perform the operations and generate the output described herein with respect to the exemplary embodiments by executing the software application. Also, any combination of application-specific hardware, firmware, or software applications executed on the hardware may be used without departing from the scope of the exemplary embodiments.
[0101] When a computing device is configured in one of these ways, it should be recognized that the computing device becomes a dedicated computing device specifically configured to implement the mechanisms of the exemplary embodiments and is not a general-purpose computing device. Further, as described herein, implementation of the mechanisms of the exemplary embodiments improves the functionality of the computing device and provides useful and specific results that facilitate the evaluation of user inputs such as user inputs 205-208 from user input devices of client computing device 610 such as user input devices 205-208. The user input represents behavioral biometrics related to spatio-temporal features, temporal features, and fixed / categorical features. BBDL model 340 and BBDL pipeline 200 characterize the behavioral biometrics of the user input and generate a user profile that can then be stored in user profile registry 630 for access to protected resources of the authenticated user. The user profile in user profile registry 630 may authenticate the user in subsequent sessions by being compared to the characterization of the behavioral biometrics of subsequent user inputs. Authentication logic 620, such as comparison logic 470 and update logic 480 in FIG. 4, compares, for example, the characterization of behavioral biometrics to authenticate the user, and if the user is not the true user, access to the protected resources may be denied. If the user is an authenticated user, the user may be permitted access to the protected resources, and in some embodiments, the profile of the authenticated user may be dynamically updated by being updated by the characterization of the user input of the session, for example, by authentication logic 620 including update logic 470.
[0102] In some embodiments, user input may be received from sensors of user interface devices 292-295, which may be part of client computing system 610, and other data may characterize client computing system 610 itself. The user input or device information or both may be provided to server 604 that implements BBDL pipeline 200 via network 602, and then the user input may be processed to authenticate the user. Server 604 may further also return an output that authorizes or denies access to protected resources. Although a client / server configuration is shown in FIG. 6, it should be recognized that exemplary embodiments are not so limited. Further, client computing devices 610-614 may take many different forms such as a client computer, a tablet computer, a smartphone, other smart devices, and Internet of Things (IoT) devices. Briefly, any computing device may have a user input interface through which behavior biometrics-based user input may be received and provided to the BBDL pipeline, and the BBDL model may be used without departing from the scope of the present invention.
[0103] As described above, the mechanisms of the exemplary embodiments utilize a specifically configured computing device or data processing system to implement the BBDL pipeline 200 and perform operations for user input behavioral biometrics deep learning-based authentication. These computing devices or data processing systems may include various hardware elements specifically configured to implement one or more of the systems / subsystems described herein, either by a hardware configuration, a software configuration, or a combination of a hardware configuration and a software configuration. FIG. 7 is a block diagram of merely one exemplary data processing system in which aspects of the exemplary embodiments may be implemented. The data processing system 700 is an example of a computer such as the server 604 in FIG. 6, where computer-usable code or instructions for implementing the processes and aspects of the exemplary embodiments of the present invention are located or executed or both, so as to realize the operations, outputs, and external effects of the exemplary embodiments described herein.
[0104] In the example shown, the data processing system 700 uses a hub architecture that includes a North Bridge and Memory Controller Hub (NB / MCH) 702 and a South Bridge and Input / Output (I / O) Controller Hub (SB / ICH) 704. The processing unit 706, main memory 708, and graphic processor 710 are connected to the NB / MCH 702. The graphic processor 710 may be connected to the NB / MCH 702 through an Accelerated Graphic Port (AGP).
[0105] In the example shown, the local area network (LAN) adapter 712 is connected to the SB / ICH704. The audio adapter 716, the keyboard and mouse adapter 720, the modem 722, the read only memory (ROM) 724, the hard disk drive (HDD) 726, the CD-ROM drive 730, the universal serial bus (USB) ports and other communication ports 732, and the PCI / PCIe devices 734 are connected to the SB / ICH704 through the bus 738 and the bus 740. The PCI / PCIe devices may include, for example, an Ethernet (R) adapter, an add-in card, and a PC card for a notebook computer. A card bus controller is used in PCI but not in PCIe. The ROM 724 may be, for example, a flash basic input output system (BIOS).
[0106] The HDD 726 and the CD-ROM drive 730 are connected to the SB / ICH704 through the bus 740. The HDD 726 and the CD-ROM drive 730 may use, for example, an Integrated drive electronics (IDE) or a serial advanced technology attachment (SATA) interface. The super I / O (SIO) device 736 may be connected to the SB / ICH704.
[0107] The operating system runs on the processing unit 706. The operating system coordinates and provides control of the various components within the data processing system 700 in FIG. 7. As a client, the operating system may be a commercially available operating system such as Microsoft (R) Windows10 (R). An object-oriented programming system such as the Java (TM) programming system may run in conjunction with the operating system and provide calls to the operating system from Java (TM) programs or applications running on the data processing system 700.
[0108] As a server, the data processing system 700 may be, for example, an IBM(R) eServer(TM) System p(R) computer system or a Power(TM) processor-based computer system that executes an Advanced Interactive Executive (AIX(R)) operating system or a LINUX(R) operating system. The data processing system 700 may be a symmetric multi-processor (SMP) system that includes a plurality of processors in a processing unit 706. Alternatively, a single processor system may be used.
[0109] Instructions for an operating system, an object-oriented programming system, and an application or program may be located on a storage device such as HDD 726 and loaded into main memory 708 for execution by processing unit 706. The processes for exemplary embodiments of the present invention may be executed by processing unit 706 using, for example, computer-usable program code located in main memory 708, memory such as ROM 724, or, for example, one or more peripheral devices 726 and 730.
[0110] A bus system such as bus 738 or bus 740 shown in FIG. 7 may be composed of one or more buses. Of course, the bus system may be implemented using any type of communication fabric or architecture that provides for the transfer of data between different components or devices attached to the fabric or architecture. The communication unit, such as modem 722 or network adapter 712 in FIG. 7, may include one or more devices used to transmit and receive data. The memory may be, for example, main memory 708, ROM 724, or cache found, for example, in NB / MCH 702 in FIG. 7.
[0111] As described above, in some exemplary embodiments, the mechanisms of the exemplary embodiments may be implemented as application software stored on a storage device such as HDD 726 and loaded into a memory such as main memory 708 for execution by one or more hardware processors such as application specific hardware or firmware, or processing unit 706. As such, the computing device shown in FIG. 7 is specifically configured to implement the mechanisms of the exemplary embodiments and to perform the operations and generate the outputs described herein with respect to the BBDL pipeline 200 in FIG. 2 and the authentication mechanism based on the characteristics of the user input generated by the BBDL pipeline 200.
[0112] One of ordinary skill in the art will recognize that the hardware in FIGS. 6 and 7 may vary depending on the implementation. Other internal hardware or peripheral devices such as flash memory, equivalent non-volatile memory, or optical disk drives may be used in addition to or instead of the hardware shown in FIGS. 6 and 7. Also, the processes of this exemplary embodiment may be applied to multiprocessor data processing systems other than the previously described SMP systems without departing from the scope of the present invention.
[0113] Furthermore, data processing system 700 may take the form of any of several different data processing systems, including a client computing device, a server computing device, a tablet computer, a laptop computer, a telephone or other communication device, or a personal digital assistant (PDA), among others. In some illustrative examples, data processing system 700 may be a portable computing device configured with flash memory that provides non-volatile memory for storing, for example, operating system files or user-generated data or both. Basically, data processing system 700 may be any data processing system known or later developed without architectural limitations.
[0114] As noted above, it should be recognized that exemplary embodiments may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment containing both hardware elements and software elements. In one exemplary embodiment, the mechanisms of the exemplary embodiments are implemented in software or program code, including but not limited to firmware, resident software, microcode, etc.
[0115] A data processing system suitable for storing or executing or both program code will include at least one processor directly or indirectly coupled to a memory element through a communication bus, such as a system bus. The memory element may include local memory used during actual execution of the program code, a mass storage area, and a cache memory that provides at least some temporary storage of at least some program code to reduce the number of times the code must be retrieved from the mass storage area during execution. The memory may be of various types including, but not limited to, ROM, PROM, EPROM, EEPROM, DRAM, SRAM, flash memory, and solid state memory.
[0116] Input / output devices, i.e., I / O devices (including but not limited to keyboards, displays, pointing devices, etc.), can be coupled to the system either directly or through intervening wired or wireless I / O interfaces or controllers or both. The I / O devices can take many different forms other than conventional keyboards, displays, and pointing devices, such as communication devices coupled through a wired or wireless connection, including but not limited to smartphones, tablet computers, touchscreen devices, and voice recognition devices. Any I / O device known or later developed is intended to be within the scope of the exemplary embodiments.
[0117] Network adapters can also be coupled to the system to enable the data processing system to be coupled to other data processing systems or remote printers or storage devices through an intervening private or public network. Modems, cable modems, and Ethernet cards are just a few of the commercially available types of network adapters for wired communication. Network adapters based on wireless communication, including but not limited to 802.11a / b / g / n wireless communication adapters and Bluetooth(R) wireless adapters, may also be utilized. Any network adapter known or later developed is intended to be within the scope of the present invention.
[0118] The description of the present invention is presented for purposes of illustration and description, and is not intended to be exhaustive or limited to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, and to enable others skilled in the art to understand the invention for various embodiments with various modifications that are suited to the particular use contemplated. The terminology used herein was chosen in order to best explain the principles of the embodiments, the practical application or technical improvement over the technology found in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
Claim 1 A method in a data processing system including at least one processor and at least one memory, wherein the at least one memory includes instructions executed by the at least one processor to specifically configure the at least one processor to implement an action biometrics deep learning (BBDL) pipeline including a plurality of stages of a machine learning computer model operative to implement the method, the method comprising: Receiving spatio-temporal input data corresponding to inputs associated with an entity over a predetermined time frame including a plurality of time intervals from one or more sensors, each time interval having a corresponding subset of the spatio-temporal input data, the spatio-temporal input data including spatially and temporally dependent input data regarding a physical area monitored by at least one of the one or more sensors; Processing, at each time interval of the plurality of time intervals, by one or more machine learning computer models of corresponding stages in the plurality of stages, a subset of the spatio-temporal input data corresponding to the time interval to generate an output vector having a value indicative of an internal representation of spatio-temporal features of the entity represented in the subset of the spatio-temporal input data; Accumulating the output vectors over the plurality of stages of the BBDL pipeline to generate a final output vector including a final output vector value indicative of the spatio-temporal features of the entity represented in the spatio-temporal input data; Authenticating the entity by the data processing system based on the final output vector. Claim 2 Authenticating the entity comprises: Comparing the final output vector with at least one previously generated output vector generated by the BBDL pipeline stored in a user profile for an authenticated user to determine a probability that the final output vector represents spatio-temporal features of a user that match the spatio-temporal features of the authenticated user; Controlling access to a protected resource in response to a result of the comparison. The method of claim 1 Claim 3 Controlling access to the protected resource, in response to determining that the probability indicates that the final output vector represents the user's spatio-temporal characteristics that match the spatio-temporal characteristics of the authenticated user, permitting the user to access the protected resource; updating the user profile to include the final output vector, the method according to claim 2. **Claim 4** Each stage of the BB-DL pipeline includes an image processing machine learning computer model, and in each stage of the BB-DL pipeline, processing the subset of the spatio-temporal input data corresponding to the time interval processing the subset of the spatio-temporal input data to convert spatio-temporal features in the subset of the spatio-temporal input data into an image; performing image analysis on the image by the image processing machine learning computer model of the stage to generate a first vector output, the method according to claim 1. **Claim 5** Each stage of the BB-DL pipeline includes a fully connected neural network machine learning computer model, and in each stage of the BB-DL pipeline, processing the subset of the spatio-temporal input data corresponding to the time interval processing the subset of the spatio-temporal input data to identify temporal features in the subset of the spatio-temporal input data; processing the temporal features by the fully connected neural network machine learning model of the stage to generate a second vector output, the method according to claim 4. **Claim 6** Each stage of the BB-DL pipeline includes fixed / categorical data embedding logic, and in each stage of the BB-DL pipeline, processing the subset of the spatio-temporal input data corresponding to the time interval processing the subset of the spatio-temporal input data to identify fixed / categorical features in the subset of the spatio-temporal input data; processing the fixed / categorical features by the fixed / categorical data embedding logic of the stage to generate a third vector output, the method according to claim 5. **Claim 7** Each stage of the BBDL pipeline operates to generate a combined output vector that combines the first vector output, the second vector output, and the third vector output and is input to the main stage machine learning computer model of that stage, the method of claim 6.
8. Each main stage machine learning computer model of each stage of the BBDL pipeline processes the corresponding combined output vector of this corresponding stage together with the input from the previous stage of the BBDL pipeline to generate a stage output vector, and accumulating the output vectors over the plurality of stages of the BBDL pipeline to generate a final output vector includes, for each stage, outputting the stage output vector as an input to the next stage in the BBDL pipeline, and the stage output vector of the main stage machine learning computer model of the last stage of the BBDL pipeline is the final output vector, the method of claim 7.
9. The embedded device information is input to the main stage machine learning computer model of the first stage of the BBDL pipeline, and the main stage machine learning computer model processes the combined output vector associated with the first stage together with the embedded device information to generate a stage output vector for the first stage, the method of claim 8.
10. The one or more sensors include touch sensors of a touch-sensitive display device, and the spatio-temporal input data includes sensor data indicating the characteristics of the touch input of the entity, the method of claim 1.
11. A computer program for causing a computer to execute the method according to any one of claims 1 to 10.
12. An apparatus, at least one processor, at least one memory coupled to the at least one processor, wherein the at least one memory, when executed by the at least one processor, includes instructions that cause the at least one processor to implement an action biometrics deep learning (BBDL) pipeline that includes multiple stages of a machine learning computer model, and the machine learning computer model receiving spatio-temporal input data corresponding to inputs associated with an entity over a predetermined time frame including a plurality of time intervals from one or more sensors, each time interval having a corresponding subset of the spatio-temporal input data, the spatio-temporal input data including spatially and temporally dependent input data regarding a physical area monitored by at least one of the one or more sensors, the receiving of the spatio-temporal input data processing, by one or more machine learning computer models of a corresponding stage in the plurality of stages, a subset of the spatio-temporal input data corresponding to the time interval at each time interval in the plurality of time intervals to generate an output vector having a value indicative of an internal representation of spatio-temporal features of the entity represented in the subset of the spatio-temporal input data accumulating the output vectors over the multiple stages of the BBDL pipeline to generate a final output vector including a final output vector value indicative of the spatio-temporal features of the entity represented in the spatio-temporal input data authenticating the entity by the apparatus based on the final output vector A device that operates to perform the above.
13. Authenticating the entity comprises comparing the final output vector with at least one previously generated output vector generated by the BBDL pipeline stored in a user profile for an authenticated user to determine a probability that the final output vector represents the spatio-temporal features of a user that match the spatio-temporal features of the authenticated user controlling access to protected resources according to the result of the comparison, the device according to claim 12.
14. Controlling access to the protected resource is, in response to determining that the probability indicates that the final output vector represents the user's spatio-temporal characteristics that match the spatio-temporal characteristics of the authenticated user, permitting the user to access the protected resource, updating the user profile to include the final output vector, the apparatus according to claim 13. **Claim 15** Each stage of the BBDL pipeline includes an image processing machine learning computer model, and in each stage of the BBDL pipeline, processing the subset of the spatio-temporal input data corresponding to the time interval is processing the subset of the spatio-temporal input data to convert spatio-temporal features in the subset of the spatio-temporal input data into an image, performing image analysis on the image by the image processing machine learning computer model of the stage to generate a first vector output, the apparatus according to claim 12. **Claim 16** Each stage of the BBDL pipeline includes a fully connected neural network machine learning computer model, and in each stage of the BBDL pipeline, processing the subset of the spatio-temporal input data corresponding to the time interval is processing the subset of the spatio-temporal input data to identify temporal features in the subset of the spatio-temporal input data, processing the temporal features by the fully connected neural network machine learning model of the stage to generate a second vector output, the apparatus according to claim 15. **Claim 17** Each stage of the BBDL pipeline includes fixed / category data embedding logic, and in each stage of the BBDL pipeline, processing the subset of the spatio-temporal input data corresponding to the time interval is processing the subset of the spatio-temporal input data to identify fixed / category features in the subset of the spatio-temporal input data, processing the fixed / category features by the fixed / category data embedding logic of the stage to generate a third vector output, the apparatus according to claim 16. **Claim 18** Each stage of the BBDL pipeline includes combination logic that operates to combine the first vector output, the second vector output, and the third vector output to generate a combined output vector that is input to the main stage machine learning computer model of that stage, the apparatus of claim 17.
19. Each main stage machine learning computer model of each stage of the BBDL pipeline processes the corresponding combined output vector of this corresponding stage along with the input from the previous stage of the BBDL pipeline to generate a stage output vector, and accumulating the output vectors over the plurality of stages of the BBDL pipeline to generate a final output vector includes, for each stage, outputting the stage output vector as an input to the next stage in the BBDL pipeline, and the stage output vector of the main stage machine learning computer model of the last stage of the BBDL pipeline is the final output vector, the apparatus of claim 18.
20. Embedded device information is input to the main stage machine learning computer model of the first stage of the BBDL pipeline, and the main stage machine learning computer model processes the combined output vector associated with the first stage along with the embedded device information to generate a stage output vector for the first stage, the apparatus of claim 19.
21. The one or more sensors include touch sensors of a touch-sensitive display device, and the spatio-temporal input data includes sensor data indicating the characteristics of the touch input of the entity, the apparatus of claim 12.
Citation Information
Patent Citations
Continuous authentication by mobile device
JP2017515178A