Systems and methods for differentiating user, action, and device specific features recorded in motion sensor data
By dividing the motion signal into fragments and using the Cycle-GAN algorithm to convert the features, the problem that the user identification system in the prior art is difficult to distinguish between the user and the device's action characteristics, and more efficient user identification and authentication are achieved.
Patent Information
- Application Number
- CN202180013376.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-06
- Filing Date
- 2021-01-06
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-01-06
AI Technical Summary
The existing smartphone user identification system based on motion sensor data is difficult to distinguish between the discriminant characteristics of the user from the discriminant characteristics of the device and the action when the attacker tries to authenticate, resulting in a decrease in performance and accuracy.
By dividing the motion signal into fragments and using trained transformation algorithms, such as a Cycle-GAN, the fragments are converted into transformed fragments, eliminating the discriminant features of the device and the action, thereby extracting the discriminant features of the user.
It realizes more reliable and effective user identification, improves the performance and accuracy of the system, and reduces the false alarm rate and missed alarm rate.
Smart Images

Figure CN115087973B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 62 / 957,653, filed on January 6, 2020, entitled "System and Method for Disentangling Features Specific to Users, Actions and Devices Recorded in Motion Sensor Data", the content of which is incorporated herein by reference as if set forth in its entirety herein. Technical field
[0003] This application relates to systems and methods for extracting features of a user, and particularly to systems and methods for extracting discriminative features of a user of a device from motion sensor data related to the user. Background art
[0004] When an attacker attempts to authenticate on an owner's smartphone, standard machine - learning (ML) systems designed for smartphone user (owner) identification and authentication based on motion sensor data suffer a significant performance and accuracy degradation (e.g., higher than 10%). This problem naturally occurs when the ML system fails to disentangle the discriminative features of the user from the discriminative features of the smartphone device or the discriminative features of common actions (e.g., picking up the phone from a table, answering a call, etc.). The problem is caused by the fact that the signals recorded by motion sensors (e.g., accelerometers and gyroscopes) during an authentication session contain all these features (representing the user, the action, and the device) simultaneously.
[0005] One way to solve this problem is to collect additional motion signals during user registration by: (a) asking the user to authenticate on multiple devices while performing different actions (e.g., sitting on a chair, standing, changing hands, etc.), or (b) asking the smartphone owner to let another person perform some authentication (to simulate a potential attack). However, both of these options are inconvenient for the user.
[0006] Therefore, there is a need for a more reliable and efficient way to extract discriminative features of a user of a mobile device. Summary of the invention
[0007] In a first aspect, there is provided a computer-implemented method for distinguishing discriminative features of a user of a device from motion signals captured by a mobile device. The mobile device has one or more motion sensors, a storage medium, instructions stored on the storage medium, and a processor configured by executing the instructions. In the method, the processor divides each captured motion signal into segments. Then, the processor uses one or more trained transformation algorithms to transform the segments into transformed segments. Then, the processor provides the segments and the transformed segments to a machine learning system. Then, the processor uses the machine learning system that applies one or more feature extraction algorithms to extract discriminative features of the user from the segments and the transformed segments.
[0008] In another aspect, the discriminative features of the user are used to identify the user when the user uses the device in the future. In another aspect, one or more of the motion sensors include at least one of a gyroscope and an accelerometer. In another aspect, one or more of the motion signals correspond to one or more interactions between the user and the mobile device.
[0009] In another aspect, the motion signals include discriminative features of the user, discriminative features of actions performed by the user, and discriminative features of the mobile device. In a further aspect, the step of dividing one or more of the captured motion signals into segments eliminates the discriminative features of actions performed by the user. In a further aspect, the step of transforming the segments into transformed segments eliminates the discriminative features of the mobile device.
[0010] In another aspect, one or more of the trained transformation algorithms include one or more Cycle-Consistent Generative Adversarial Networks (Cycle-GANs), and the transformed segments include synthetic motion signals that simulate motion signals originating from another device.
[0011] In another aspect, the step of dividing one or more of the captured motion signals into segments includes dividing each motion signal into a fixed number of segments, where each segment has a fixed length.
[0012] In a second aspect, there is provided a computer-implemented method for authenticating a user on a mobile device from motion signals captured by the mobile device. The mobile device has one or more motion sensors, a storage medium, instructions stored on the storage medium, and a processor configured by executing the instructions. In the method, the processor divides one or more captured motion signals into segments. The processor transforms the segments into transformed segments using one or more trained transformation algorithms. The processor provides the segments and the transformed segments to a machine learning system. Then, by assigning a score to each of the segments and the transformed segments, the processor classifies the segments and the transformed segments as belonging to an authorized user or an unauthorized user. Then, the processor applies a voting scheme or a meta-learning model to the scores assigned to the segments and the transformed segments. Then, the processor determines whether the user is an authorized user based on the voting scheme or the meta-learning model.
[0013] In another aspect, the classifying step includes: comparing the segments and the transformed segments with the features of an authorized user extracted from sample segments provided by the authorized user during an enrollment process, wherein the features are stored on the storage medium; and assigning a score to each segment based on a classification model.
[0014] In another aspect, one or more of the motion sensors include at least one of a gyroscope and an accelerometer. In another aspect, the step of dividing one or more captured motion signals into segments includes dividing each motion signal into a fixed number of segments, and each segment has a fixed length. In another aspect, at least a portion of the segments overlap.
[0015] In another aspect, one or more of the trained transformation algorithms include one or more Cycle-GANs, and the transforming step includes: transforming the segments into transformed segments via a first generator that mimics segments generated on another device; and re-transforming the transformed segments via a second generator to mimic segments generated on the mobile device.
[0016] In another aspect, the transformed segments include synthetic motion signals that simulate motion signals originating from another device. In another aspect, the providing step includes: using the processor to extract features from the segments and the transformed segments using one or more feature extraction techniques to form feature vectors; and applying a learned classification model to the feature vectors corresponding to the segments and the transformed segments.
[0017] In a third aspect, a system is provided for differentiating discriminative features of a user of a device from motion signals captured on a mobile device and authenticating the user on the mobile device, where the mobile device has at least one motion sensor. The system includes a network communication interface, a computer-readable storage medium, and a processor configured to interact with the network communication interface and the computer-readable storage medium and execute one or more software modules stored on the storage medium. The software modules include:
[0018] A segmentation module that, when executed, configures the processor to divide each captured motion signal into segments;
[0019] A transformation module that, when executed, configures the processor to transform the segments into transformed segments using one or more trained Cycle-Consistent Generative Adversarial Networks (Cycle-GANs);
[0020] A feature extraction module that, when executed, configures the processor to extract the extracted discriminative features of the user from the segments and the transformed segments, where the processor uses a machine learning system;
[0021] A classification module that, when executed, configures the processor to assign scores to the segments and the transformed segments and determine whether the segments and the transformed segments belong to an authorized user or an unauthorized user based on the respective scores of these segments;
[0022] A meta-learning module that, when executed, configures the processor to apply a voting scheme or a meta-learning model to the scores assigned to the segments and the transformed segments based on the stored segments corresponding to the user.
[0023] In another aspect, at least one motion sensor includes at least one of a gyroscope and an accelerometer.
[0024] In another aspect, the transformation module configures the processor to: transform the segments into transformed segments via a first generator that mimics segments generated on another device; and re-transform the transformed segments via a second generator to mimic segments generated on the mobile device.
[0025] In another aspect, the feature extraction module is further configured to employ a learned classification model for the extracted features corresponding to the segments and the transformed segments. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1A A high-level diagram of a system for differentiating discriminative features of a user of a device from motion sensor data and authenticating the user from motion sensor data in accordance with at least one embodiment disclosed herein is disclosed;
[0027] Figure 1Bis a block diagram of a computer system for distinguishing a user of a device from motion sensor data and authenticating the user from the motion sensor data according to at least one embodiment disclosed herein;
[0028] Figure 1C is a block diagram of a software module for distinguishing a user of a device from motion sensor data and authenticating the user from the motion sensor data according to at least one embodiment disclosed herein;
[0029] Figure 1D is a block diagram of a computer system for distinguishing a user of a device from motion sensor data and authenticating the user from the motion sensor data according to at least one embodiment disclosed herein;
[0030] Figure 2 is a diagram showing an exemplary machine learning-based standard smartphone user identification system and flowchart according to one or more embodiments;
[0031] Figure 3 is a diagram showing an exemplary mobile device system and flowchart for identifying a user by removing features that distinguish actions based on machine learning according to one or more embodiments;
[0032] Figure 4 is a diagram of an exemplary Cycle-GAN (Cycle-Generative Adversarial Network) for signal-to-signal transformation according to one or more embodiments;
[0033] Figures 5A - 5D shows a system and flowchart for distinguishing discriminative features of a user of a mobile device from motion sensor data and authenticating the user from the motion sensor data according to one or more embodiments;
[0034] Figure 6A discloses a schematic block diagram showing a computational process for distinguishing discriminative features of a user of a device from motion sensor data according to at least one embodiment disclosed herein; and
[0035] Figure 6B discloses a schematic block diagram showing a computational process for authenticating a user on a mobile device from motion sensor data according to at least one embodiment disclosed herein. Detailed Description
[0036] In an overview and introductory manner, this document discloses exemplary systems and methods for extracting or differentiating discriminative features of a smartphone user from motion signals and separating them from features that can be used to differentiate actions and devices without the user having to perform any additional authentication during registration. According to one or more embodiments of the method, in a first stage, the features differentiating actions are eliminated by cutting the signal into smaller chunks and independently applying a machine learning system to each of these chunks. It should be understood that the term "eliminating" does not necessarily mean that the signal features related to an action (or device) are removed from the resulting (one or more) motion data signals. Instead, the process confounds those discriminative features through the independent processing of the signal chunks, effectively eliminating the discriminative features of the action. In other words, since the system cannot reconstruct the entire signal from the smaller chunks, the actions (e.g., finger swipes, gestures) of the user corresponding to the motion signal can no longer be recovered and recognized, and are thus effectively eliminated. In a second stage, the features differentiating devices are effectively eliminated by using a generative model (e.g., a generative adversarial network) to simulate authentication sessions on a predefined set of devices. The generative model is trained to take signal chunks from a device as input and provide similar signal chunks as output, where these similar signal chunks replace the features of the input device with the discriminative features of other devices in the predefined set. After injecting the features from different devices into the signals of a certain user, the machine learning system can learn from the original signal chunks and the simulated signal chunks which features do not change across devices. These are the features that are useful in differentiating the relevant user.
[0037] For example, the methods and systems described herein can be inserted into any machine learning system designed for identifying and authenticating smartphone users (owners). The described methods and systems focus on the discriminative features of the user while eliminating the features differentiating devices and actions. These methods and systems have been confirmed in a set of experiments, and the benefits have also been empirically demonstrated.
[0038] Figure 1A A schematic diagram of the present system 100 for differentiating the discriminative features of a user of a device from motion sensor data and authenticating the user according to at least one embodiment is disclosed. The method can be implemented using one or more aspects of the present system 100, as described in further detail below. In some implementations, the system 100 includes a cloud-based system server platform that communicates with fixed PCs, servers, and devices (such as smartphones, tablets, and laptops) operated by the user.
[0039] In one arrangement, system 100 includes a system server (backend server) 105 and one or more user devices including one or more mobile devices 101. Although the present system and method are generally described as being performed partially with mobile devices (e.g., smartphones), in at least one embodiment, the present system and method can be implemented on other types of computing devices such as workstations, personal computers, laptop computers, access control devices, or other suitable digital computers. For example, although user-facing devices such as mobile device 101 typically capture motion signals, one or more of the exemplary processing operations are intended to distinguish discriminative features of the user from the motion signals, and authentication / identification can be performed by system server 105. System 100 may also include one or more remote computing devices 102.
[0040] System server 105 can actually be any computing device or data processing apparatus capable of communicating with user devices and remote computing devices and receiving, transmitting, and storing electronic information and processing requests, as further described herein. Similarly, remote computing device 102 can actually be any computing device or data processing apparatus capable of communicating with the system server or user devices and receiving, transmitting, and storing electronic information and processing requests, as further described herein. It should also be understood that the system server or remote computing device can be any number of networked or cloud-based computing devices.
[0041] In one or more embodiments, the (one or more) user devices, the (one or more) mobile devices 101 can be configured to communicate with each other, communicate with system server 105 or remote computing device 102, transmit electronic information thereto and receive electronic information therefrom. The user devices can be configured to capture and process motion signals from the user, e.g., corresponding to one or more gestures (interactions) from user 124.
[0042] (One or more) mobile devices 101 can be any mobile computing device or data processing apparatus capable of embodying the systems and methods described herein, including but not limited to personal computers, tablet computers, personal digital assistants, mobile electronic devices, cellular phones, or smartphone devices, etc.
[0043] It should be noted that although Figure 1A system 100 is depicted for distinguishing discriminative features of the user and user authentication with respect to (one or more) mobile devices 101 and remote computing devices 102, any number of such devices can interact with the system in the manner described herein. It should also be noted that although Figure 1A system 100 is depicted for distinguishing discriminative features of the user and authentication with respect to user 124, any number of users can interact with the system in the manner described herein.
[0044] It should be further understood that although the various computing devices and machines referred to herein (including but not limited to the (one or more) mobile devices 101 and the system server 105 and the remote computing device 102) are referred to herein as individual or single devices and machines, in some embodiments, the referred devices and machines and their associated or attendant operations, features, and functions may be combined or arranged across multiple such devices or machines or otherwise employed, such as via a network connection or a wired connection, as known to those skilled in the art.
[0045] It should also be understood that the exemplary systems and methods described herein in the context of the (one or more) mobile devices 101 (also referred to as smartphones) are not specifically limited to mobile devices and may be implemented using other enabled computing devices.
[0046] Now referring Figure 1B , the mobile device 101 of the system 100 includes various hardware and software components for enabling the system to operate, including one or more processors 110, a memory 120, a microphone 125, a display 140, a camera 145, an audio output 155, a storage device 190, and a communication interface 150. The processor 110 is used to execute client applications in the form of software instructions that can be loaded into the memory 120. The processor 110 can be any number of processors, a central processing unit CPU, a graphics processing unit GPU, a multi-processor core, or any other type of processor, depending on the embodiment.
[0047] Preferably, the memory 120 and / or the storage device 190 can be accessed by the processor 110, enabling the processor to receive and execute instructions encoded in the memory and / or on the storage device to cause the mobile device and its various hardware components to perform operations of aspects of the systems and methods described in more detail below. For example, the memory can be a random access memory (RAM) or any other suitable volatile or non-volatile computer-readable storage medium. Additionally, the memory can be fixed or removable. The storage device 190 can take various forms, depending on the embodiment. For example, the storage device can include one or more components or devices, such as a hard disk drive, a flash memory, a rewritable optical disc, a rewritable magnetic tape, or some combination of the above. The storage device can also be fixed or removable.
[0048] One or more software modules 130 can be encoded in the storage device 190 and / or the memory 120. The software modules 130 can include one or more software programs or applications having computer program code or a set of instructions that are executed in the processor 110. In an exemplary embodiment, as Figure 1CDepicted, preferably included in software module 130, are user interface module 170, feature extraction module 172, segmentation module 173, classification module 174, meta - learning module 175, database module 176, communication module 177, and transformation module 178, which are executed by processor 110. Such computer program code or instructions configure processor 110 to perform the operations of the systems and methods disclosed herein and may be written in any combination of one or more programming languages.
[0049] Specifically, user interface module 170 may include one or more algorithms for performing steps related to capturing motion signals from a user and authenticating the user's identity. Feature extraction module 172 may include one or more feature extraction algorithms (e.g., machine - learning algorithms) for performing steps related to extracting discriminative features of the user from segments and transformed segments of the user's motion signals. Segmentation module 173 may include one or more algorithms for performing steps related to dividing the captured motion signals into segments. Transformation module 178 includes one or more transformation algorithms for performing steps related to transforming segments of the motion signals into transformed segments. Classification module 174 includes one or more algorithms for performing steps related to scoring the segments and transformed segments (e.g., assigning class probabilities to them). Meta - learning module 175 includes one or more algorithms (e.g., voting schemes or meta - learning models) for performing steps related to integrating the scores assigned to the segments and transformed segments in order to identify or reject the user attempting authentication. Database module 176 includes one or more algorithms for storing or saving data related to motion signals, segments, or transformed segments to database 185 or storage device 190. Communication module 177 includes one or more algorithms for transmitting and receiving signals between computing devices 101, 102, and / or 105 of system 100.
[0050] The program code may execute entirely on mobile device 101 as a stand - alone software package, partially on the mobile device and partially on system server 105, or entirely on the system server or another remote computer or device. In the latter case, the remote computer may be connected to mobile device 101 via any type of network connection, which includes local area network (LAN) or wide area network (WAN), mobile communication network, cellular network, or the connection may be for an external computer (e.g., using an Internet service provider over the Internet).
[0051] In one or more embodiments, the program code of software module 130 and one or more computer - readable storage devices (such as memory 120 and / or storage device 190) form a computer program product that may be manufactured and / or distributed in accordance with the present invention, as is known to those of ordinary skill in the art.
[0052] In some exemplary embodiments, one or more of the software modules 130 may be downloaded from another device or system via the communication interface 150 over a network to the storage device 190 for use within the system 100. Additionally, it should be noted that other information and / or data related to the operation of the present system and method (such as the database 185) may also be stored on the storage device. Preferably, such information is stored on an encrypted data repository that is specifically allocated for the secure storage of information collected or generated by a processor executing a security authentication application. Preferably, encryption measures are used to locally store the information on the mobile device storage device and to transmit the information to the system server 105. For example, a 1024-bit polymorphic cipher or (depending on export controls) the AES256-bit encryption method may be used to encrypt such data. Additionally, encryption may be performed using a remote key (seed) or a local key (seed). As will be understood by those skilled in the art, alternative encryption methods may be used, such as SHA256.
[0053] Additionally, the user's motion sensor data or mobile device information may be used as an encryption key to encrypt data stored on the mobile device(s) 101 and / or the system server 105. In some implementations, the foregoing combination may be used to create a complex unique key for the user, which may be encrypted on the mobile device using elliptic curve cryptography (preferably with a length of at least 384 bits). Additionally, this key may be used to protect the user data stored on the mobile device or the system server.
[0054] Furthermore, in one or more embodiments, the database 185 is stored on the storage device 190. As will be described in more detail below, the database 185 contains or maintains various data items and elements used throughout the various operations of the system 100 and method for differentiating users and user authentication. The information stored in the database may include, but is not limited to, user motion sensor data templates and profile information, as will be described in more detail herein. It should be noted that while the database is depicted as being locally configured to the mobile device 101, in certain implementations, the database or various data elements stored therein may additionally or alternatively be remotely located (such as on a remote device 102 or the system server 105—not shown) and connected to the mobile device via a network in a manner known to those of ordinary skill in the art.
[0055] A user interface 115 is also operably connected to the processor. The interface may be one or more input or output devices, such as switch(es), button(s), key(s), touch screen, microphone, etc., as will be understood in the art of electronic computing devices. The user interface 115 is used to facilitate the capture of commands from the user, such as switch commands or user information and settings for user identification related to the operation of the system 100. For example, in at least one embodiment, the interface 115 may be used to facilitate the capture of certain information from the mobile device(s) 101, such as personal user information for registration with the system, in order to create a user profile.
[0056] The mobile device 101 may also include a display 140, which is also operably connected to the processor 110. The display includes a screen or any other such presentation device that enables the system to instruct or otherwise provide feedback to the user regarding the operation of the system 100 for distinguishing between the discriminative features of the users and user authentication. By way of example, the display may be a digital display, such as a dot matrix display or other two-dimensional display.
[0057] By way of further example, the interface and display may be integrated into a touch screen display, as is common in smartphones such as mobile device 101. Thus, the display is also used to show a graphical user interface that can display various data and provide a "form" including fields that allow a user to enter information. Touching the touch screen at a location corresponding to the display of the graphical user interface allows a person to interact with the device to enter data, change settings, control functions, etc. Thus, when the touch screen is touched, the user interface communicates the change to the processor, and may change settings, or may capture information entered by the user and store it in memory.
[0058] The mobile device 101 may also include a camera 145 capable of capturing digital images. The mobile device 101 or camera 145 may also include one or more light or signal emitters (e.g., LEDs, not shown), such as visible light emitters or infrared light emitters, etc. The camera may be integrated into the mobile device, such as a front camera or a rear camera containing a sensor (e.g., and not limited to a CCD or CMOS sensor). As will be understood by those skilled in the art, the camera 145 may also include additional hardware, such as a lens, a light meter (e.g., an illuminance meter), and other conventional hardware and software features that may be used to adjust image capture settings, such as zoom, focus, aperture, exposure, shutter speed, etc. Alternatively, the camera may be external to the mobile device 101. Possible variations of the camera and light emitters will be understood by those skilled in the art. In addition, as will be understood by those skilled in the art, the mobile device may also include one or more microphones 125 for capturing audio recordings.
[0059] The audio output 155 is also operatively connected to the processor 110. The audio output can be any type of speaker system configured to play electronic audio files, as would be understood by those skilled in the art. The audio output can be integrated into the mobile device 101 or external to the mobile device 101.
[0060] A variety of hardware devices or sensors 160 can also be operatively connected to the processor. For example, the sensors 160 can include: an on-board clock for tracking the time of day, etc.; a GPS-enabled device for determining the location of the mobile device; a gravity magnetometer for detecting the Earth's magnetic field to determine the three-dimensional orientation of the mobile device; a proximity sensor for detecting the distance between the mobile device and other objects; an RF radiation sensor for detecting the level of RF radiation; and other such devices as would be understood by those skilled in the art.
[0061] (One or more) mobile devices 101 also include an accelerometer 135 and / or a gyroscope 136, which are configured to capture motion signals from the user 124. In at least one embodiment, the accelerometer can also be configured to track the orientation and acceleration of the mobile device. The mobile device 101 can be set (configured) to provide accelerometer and gyroscope values to the processor 110 that executes various software modules 130, including, for example, a feature extraction module 172, a classification module 174, and a meta-learning module 175.
[0062] The communication interface 150 is also operatively connected to the processor 110 and can be any interface capable of communicating between the mobile device 101 and external devices, machines, and / or elements, including the system server 105. Preferably, the communication interface includes, but is not limited to, a modem, a network interface card (NIC), an integrated network interface, a radio frequency transmitter / receiver (such as Bluetooth, cellular, NFC), a satellite communication transmitter / receiver, an infrared port, a USB connection, and / or any other such interface for connecting the mobile device to other computing devices and / or communication networks (such as private networks and the Internet). Such connections can include wired connections or wireless connections (e.g., using the 802.11 standard), but it should be understood that the communication interface can actually be any interface capable of communicating to and from the mobile device.
[0063] At various points during the operation of the system 100 for differentiating discriminative features of a user and user authentication, the mobile device 101 can communicate with one or more computing devices, such as the system server 105 and / or a remote computing device 102. Such computing devices transmit data to and / or receive data from the mobile device 101, thereby preferably initiating, maintaining, and / or enhancing the operation of the system 100, as will be described in more detail below.
[0064] Figure 1DFIG. 0 is a block diagram illustrating an exemplary configuration of system server 105. System server 105 may include a processor 210 operably connected to various hardware and software components for enabling system 100 to operate to distinguish discriminative features of a user and user discrimination. Processor 210 is for executing instructions to perform various operations related to user discrimination, as will be described in more detail below. Processor 210 may be multiple processors, multi-processor cores, or some other type of processor, depending on the particular implementation.
[0065] In some embodiments, processor 210 may access memory 220 and / or storage device 290, enabling processor 210 to receive and execute instructions stored on memory 220 and / or storage device 290. Memory 220 may be, for example, random access memory (RAM) or any other suitable volatile or non-volatile computer-readable storage medium. Additionally, memory 220 may be fixed or removable. Storage device 290 may take various forms, depending on the particular implementation. For example, storage device 290 may comprise one or more components or devices such as a hard disk drive, flash memory, rewritable optical disk, rewritable magnetic tape, or some combination of the above. Storage device 290 may also be fixed or removable.
[0066] One or more software modules 230 are encoded in storage device 290 and / or memory 220. One or more of software modules 230 may include one or more software programs or applications having computer program code or a set of instructions that execute in processor 210. In one embodiment, software modules 230 may include one or more of software modules 130. Such computer program code or instructions for performing operations of aspects of the systems and methods disclosed herein may be written in any combination of one or more programming languages, as will be understood by those skilled in the art. The program code may execute entirely on system server 105 as a stand-alone software package, partially on system server 105 and partially on a remote computing device (such as remote computing device 102 and / or one or more mobile devices 101), or entirely on such remote computing devices. In one or more embodiments, as Figure 1B depicted, preferably included in software module 230 are a feature extraction module 172, a segmentation module 173, a classification module 174, a meta-learning module 175, a database module 176, a communication module 177, and a transformation module 178 that may be executed by processor 210 of the system server.
[0067] In addition, preferably, stored on the storage device 290 is the database 280. As will be described in more detail below, the database 280 contains or maintains various data items and elements used throughout the operations of the system 100, including but not limited to, user profiles, as will be described in more detail herein. It should be noted that although the database 280 is depicted as being locally configured to the computing device 105, in some embodiments, the database 280 or various data elements stored therein may be stored on a computer-readable memory or storage medium that is remotely located and connected to the system server 105 via a network (not shown) in a manner known to those of ordinary skill in the art.
[0068] The communication interface 250 is also operatively connected to the processor 210. The communication interface 250 can be any interface capable of communicating between the system server 105 and external devices, machines, or elements. In some embodiments, the communication interface 250 includes but is not limited to a modem, a network interface card (NIC), an integrated network interface, a radio frequency transmitter / receiver (e.g., Bluetooth, cellular, NFC), a satellite communication transmitter / receiver, an infrared port, a USB connection, or any other such interface for connecting the computing device 105 to other computing devices or communication networks (such as private and public networks and the Internet). Such connections can include wired or wireless connections (e.g., using the 802.11 standard), but it should be understood that the communication interface 250 can actually be any interface capable of communicating to and from the processor 210.
[0069] The operation of the system 100 and its various elements and components can be further understood with reference to, for example, the methods for using motion sensor data to distinguish discriminative features of a user and user authentication described in Figure 2 , Figure 3 , Figure 4 , Figures 5A - 5D , Figures 6A - 6B . The processes depicted herein are shown from the perspective of the mobile device 101 and / or the system server 105. However, it should be understood that these processes can be performed, in whole or in part, by one or more of the mobile device 101, the system server 105, and / or other computing devices (e.g., the remote computing device 102) or any combination of the foregoing. It should be understood that more or fewer operations than those shown in the figures and described herein may be performed. These operations may also be performed in an order different from that described herein. It should also be understood that one or more of the steps may be performed by the mobile device 101 and / or on other computing devices (e.g., the system server 105 and the remote computing device 102).
[0070] In recent literature, several user behavior authentication (UBA) systems for smartphone users based on motion sensors have been proposed. Modern systems achieving top accuracy rely on machine learning (ML) and deep learning principles to learn discriminative models, such as "N. Neverova, C. Wolf, G. Lacey, L. Fridman, D. Chandra, B. Barbello, G. Taylor. Learning Human Identity from Motion Patterns. IEEE Access, vol. 4, pp. 1810 - 1820, 2016". However, such models are evaluated on datasets collected from smartphone users, with each device having one user (the legitimate owner).
[0071] Figure 2 Shows a hybrid system and process flow diagram illustrating a machine - learning - based standard smartphone user identification system according to one or more embodiments. Specifically, Figure 2 Shows the execution flow of a typical UBA system 300 that can be implemented using a smartphone. As Figure 2 shown, in a typical UBA system 300, accelerometer or gyroscope signals are recorded during each authentication or registration session. The signals are then provided as input to a machine - learning system. The system 300 extracts features and learns a model on a set of training samples collected during user registration. During user authentication, the trained model classifies the signals as belonging to the owner (authorized session) or to a different user (rejected session). The specific example steps of the UBA system 300 are as follows:
[0072] 1. During user registration or authentication, capture signals from built - in sensors such as accelerometers and gyroscopes.
[0073] 2. Inside the ML system, employ a set of feature extraction techniques to extract relevant features from the signals.
[0074] 3. During registration, inside the ML system, train a classification model on the feature vectors corresponding to the signals recorded during user registration.
[0075] 4. During authentication, inside the ML system, apply the learned classification model to the corresponding feature vectors to distinguish the legitimate user (smartphone owner) from potential attackers.
[0076] In this setup, the decision boundary of the ML system can be influenced by discriminative features of the device or the action, rather than by discriminative features of the user. Experiments were conducted to test this hypothesis. The empirical results show that it is actually easier to rely on features corresponding to the device than on features corresponding to the user. This problem has not been pointed out or addressed in previous literature. However, this problem arises when attackers can possess the smartphone owned by a legitimate user and they attempt to authenticate to an application protected by the UBA system. Additionally, attackers can impersonate legitimate users after thoroughly analyzing and mimicking the movements performed during authentication. If the ML system relies on more prominent features characterizing the device or the action, it is likely to grant attackers access to the application. Consequently, the ML system will have an increased false positive rate and will not be able to reject such attacks.
[0077] To address this problem, system 300 can be configured to require the user to authenticate using different movements on multiple devices (e.g., authenticate with the left hand, right hand, while sitting, while standing, etc.), and to obtain negative samples from the same device, system 300 can prompt the user to let someone else perform several authentication sessions on his own device. However, all of these are impractical solutions that lead to a cumbersome registration process.
[0078] Accordingly, this paper provides methods and systems that offer a practical two-stage solution to the problems inherent in motion sensor-based UBA systems, namely, by distinguishing the discriminative features of the user from the discriminative features of the device and the action. In one or more embodiments of the present application, the disclosed method consists of the following two main processing stages:
[0079] 1. To eliminate features that distinguish the movements or actions of the user (e.g., finger swipes, gestures), the motion signal is cut into very small chunks, and the ML system is applied to each of the individual chunks. Since the entire signal cannot be reconstructed back from these chunks, the movement (the general action performed by the user) can no longer be recognized. However, these small chunks of the original signal still contain the discriminative features of the user and the device.
[0080] 2. To eliminate features that distinguish the device, a set of transformation algorithms (e.g., generative models, such as cycle-consistent generative adversarial networks for short Cycle-GAN) are applied to simulate authentication sessions on a set of predefined devices. The generative model is trained to take the signal chunks as input and provide similar chunks as output, which include features from the devices in our predefined group. After injecting features from different devices into the signal of a particular user, the ML system can learn which features do not change across devices. These are the features that help distinguish the corresponding user.
[0081] In a second stage, user behavior transformation from smartphone to smartphone is achieved by using one or more transformation algorithms, such as Cycle-GAN, which helps to obscure the characteristics describing smartphone sensors and reveal the characteristics shaping user behavior. The positive consequence of clarifying legitimate user behavior is that the UBA system becomes more vigilant and harder to deceive, i.e., the false positive rate is reduced. In one or more embodiments, the disclosed systems (e.g., system 100) and methods provide an enhanced UBA system or, alternatively, can be incorporated into a conventional UBA system to enhance an existing UBA system.
[0082] Conventionally, common mobile device authentication mechanisms (such as PINs, graphical passwords, and fingerprint scans) provide limited security. These mechanisms are vulnerable to guessing (or spoofing in the case of fingerprint scans) and are susceptible to side-channel attacks (such as smudge, reflection, and video capture attacks). As a result, continuous authentication methods based on behavioral biometric signals have received attention in both academia and industry.
[0083] The first research article analyzing accelerometer data to discern the gait of mobile device users is "E. Vildjiounaite, S.-M. Make la, M. Lindholm, R. Riihimaki, V. Kyllonen, J. Mantyjarvi, H. Ailisto. Unobtrusive multimodal biometrics for ensuring privacy and information security with personal devices. In: Proceedings of International Conference on Pervasive Computing, 2006".
[0084] Subsequently, the research community has proposed various UBA systems, such as: "N. Clarke, S. Furnell. Advanced user authentication for mobile devices. Computers & Security, vol. 26, no. 2, 2007" and "P. Campisi, E. Maiorana, M. Lo Bosco, A. Neri. User authentication using keystroke dynamics for cellular phones. Signal Processing, IET, vol. 3, no. 4, 2009", which focus on keystroke dynamics, and "C. Shen, T. Yu, S. Yuan, S., Y. Li, X. Guan. Performance analysis of motion-sensor behavior for user authentication on smartphones. Sensors, vol. 16, no. 3, pp. 345-365, 2016", "A. Buriro, B. Crispo, F. Del Frari, K. Wrona. Hold&Sign: A Novel Behavioral Biometrics for Smartphone User Authentication. In: Proceedings of Security and Privacy Workshops, 2016", "G. Canfora, P. di Notte F. Mercaldo, C. A. Visaggio. A Methodology for Silent and Continuous Authentication in Mobile Environment. In: Proceedings of International Conference on E-Business and Telecommunications, pp. 241-265, 2016", "N. Neverova, C. Wolf, G. Lacey, L. Fridman, D. Chandra, B. Barbello, G. Taylor. Learning Human Identity from Motion Patterns. IEEE Access, vol. 4, pp. 1810-1820, 2016", which focus on machine or deep learning techniques.
[0085] The methods proposed in the research literature or patents do not address the problem of motion signals that collectively contain features specific to the user, action, and device. High performance levels have been reported in recent work when users perform authentication on their own devices respectively. In such a setup, it is not clear whether the high accuracy of the machine learning model is due to the model's ability to distinguish users or distinguish devices. This problem arises because each user performs authentication on his own device, and the devices are not shared among users.
[0086] Experiments conducted on a group of users performing authentication on a group of devices such that each user authenticates on each device have revealed that it is actually easier to distinguish devices (accuracy of approximately 98%) than to distinguish users (accuracy of approximately 93%). This implies that the UBA systems proposed in the research literature and patents are more likely to perform well because they rely on device-specific functions rather than user-specific features. When an attacker performs authentication on a device stolen from the owner, such systems tend to have a high false positive rate (the attacker is authorized to enter the system).
[0087] Since this problem has not been discussed in the literature, at least in the context of user identification based on smartphone sensor motion signals, the disclosed systems and methods are the first to address the task of distinguishing features specific to the user, action, and device. In at least one embodiment, the method includes one or more of two phases / paths, one phase / path that distinguishes features specific to the user and device from features specific to the action, and one phase / path that distinguishes features specific to the user from features specific to the device.
[0088] The latter approach is inspired by recent research on image style transfer based on generative adversarial networks. In “I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, Y. Bengio. Generative Adversarial Nets. In: Proceedings of Advances in Neural Information Processing Systems, pp. 2672-2680, 2014”, the authors introduced generative adversarial networks (GANs), a model consisting of two neural networks, a generator and a discriminator, which generate new (realistic) images by learning the distribution of training samples by minimizing the Kullback-Leibler divergence. Several other GAN-based approaches have been proposed for mapping the distribution of a set of source images to a set of target images (i.e., performing style transfer). For example, methods such as “J. Y. Zhu, T. Park, P. Isola, A. A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In: Proceedings of IEEE International Conference on Computer Vision, pp. 2223-2232, 2017” and “J. Kim, M. Kim, H. Kang, K Lee. U-GAT-IT: Unsupervised generative attentional networks with adaptive layer-instance normalization for image-to-image translation. arXiv preprint arXiv:1907.10830, 2019” add a cycle-consistency loss between the target distribution and the source distribution. However, these existing methods are specific to image style transfer.
[0089] In one or more embodiments of the present application, the deep generative neural network adapts to motion signal data by focusing on signal transfer features from motion sensor data signals in the time domain. More specifically, the convolutional layer and pooling layer, which are the central components of the deep generative neural network, are modified to perform convolution and pooling operations only in the time domain, respectively. Additionally, in at least one embodiment, CycleGAN is used to transfer signals in a multi-domain setting (e.g., a multi-device setting that can include 3 or more devices). In contrast, existing GANs operate in the image domain and are applied to transfer styles only between two domains (e.g., between natural images and paintings).
[0090] According to one or more embodiments, and as described above, to distinguish the discriminative features of the user from the discriminative features of the device and the action, a two-stage method is provided.
[0091] The first stage is mainly based on preprocessed signals, which are then provided as inputs to the ML system. The preprocessing includes dividing the signals into several blocks. The number of blocks, their length, and the interval for selecting the blocks are parameters of the proposed preprocessing stage, and they can have a direct impact on accuracy and time. For example, a fixed number of blocks and a fixed length for each block are set. In this case, the interval for extracting the blocks is calculated from the signal by considering the signal length. Faster authentication will result in shorter signals and the blocks may overlap, while slower authentication will result in longer signals and the blocks may not cover the entire signal (some parts will be lost). Another example is to fix the interval and the signal length, so as to obtain different numbers of blocks according to the input signal length. Yet another example is to fix the number of blocks and calculate the interval and the corresponding length to cover the entire input signal. All the exemplified cases (and further similar cases) are covered by the preprocessing approach of the present application implemented by the disclosed system. As Figure 3 illustrated, the resulting blocks are further subjected to the feature extraction and classification pipeline of the exemplary UBA system.
[0092] Figure 3 Shows a diagram of an exemplary mobile device system and process flow chart 305 for user identification by removing discriminative features of actions based on machine learning, according to one or more embodiments. In one embodiment, one or more elements of system 100 (e.g., mobile device 100, alone or in combination with system server 105) can be used to implement the process. As Figure 3As shown, the accelerometer or gyroscope signal (motion signal) 310 is recorded by the mobile device 101 during each authentication session. The signal 310 is then divided by the system (e.g., the processor of the mobile device 100 and / or the system server 105) into N smaller (possibly overlapping) blocks or segments 315. The blocks 315 are considered individual samples and are provided as input to the machine learning system. Then, the machine learning system extracts features and learns a model on a set of training samples collected during user registration. During user authentication, the machine learning system classifies the blocks 315 as belonging to the owner or to a different user. A meta-learning approach or (weighted) majority voting can be applied to the labels or scores assigned to the blocks (extracted from the signals recorded during the authentication session) to authorize or reject the session.
[0093] Exemplary systems and methods disclosed herein may be implemented using motion sensor data captured during implicit and / or explicit authentication sessions. Explicit authentication refers to an authentication session in which the user is prompted to perform a prescribed action using a mobile device. In contrast, implicit authentication refers to a user authentication session that is performed without explicitly prompting the user to perform any action. The purpose of partitioning the motion signal 310 into blocks 315 (i.e., applying the above preprocessing) is to eliminate the features that distinguish actions. During explicit or implicit authentication, the user may perform different operations, although explicit authentication may exhibit lower variability. In fact, in an explicit authentication session, the user is likely to always perform the same steps, e.g., scanning a QR code, pointing the smartphone at the face, placing a finger on a fingerprint scanner, etc. However, these steps may be performed using one hand (left or right) or both hands, and the user may perform them while sitting on a chair, standing, or walking. In each of these cases, the recorded motion signal will be different. If the user performs the authentication steps with one hand during registration and with the other hand during authentication, problems may arise because the trained ML model will not generalize well to such changes. In such a case, a conventional system will reject legitimate users and thus have a high false positive rate. The same situation is more prevalent in implicit user authentication, i.e., when the user interacts with some data-sensitive applications (such as a banking application). In this setting, for example, not only the way the hand moves or the user's posture can be different, but also the actions performed by the user (tapping gestures at different positions on the screen, swiping gestures in different directions, etc.) can be different. The most straightforward way to separate (classify) such actions is to look at how the motion signal changes over time from the start to the end of the recording. However, the purpose here is to eliminate the system's ability to separate actions. If the motion signal 310 is separated into small blocks 315 and these blocks are processed independently, as implemented by one or more embodiments of the disclosed system for discriminative features that distinguish smartphone users from motion signals, the ML system will no longer have the opportunity to consider the entire recorded signal as a whole. Since it does not know which block goes where, the ML system will not be able to discern the action, which can only be identified by looking at the entire recording. This occurs because putting the blocks back together in a different order will correspond to different actions. It has been observed that when the user performs a set of actions during training of the ML system and a different set of actions during testing, the signal preprocessing stage of the present method improves the discrimination accuracy by 4%. Since the ML system makes decisions for each block (e.g., generates a label or calculates a score), the system can apply a voting scheme or a meta-learning model to a set of decisions corresponding to authentication in order to determine whether the user performing the authentication is legitimate.
[0094] As is well known, due to built-in defects, hardware sensors can be easily identified by looking at the outputs generated by these sensors, as detailed in "N.Khanna,A.K.Mikkilineni,A.F.Martone,G.N.Ali,G.T.C.Chiu,J.P.Allebach,E.J.Delp.A survey of forensic characterization methods forphysical devices.Digital Investigation,vol.3,pp.17-28,2006". For example, in "K.R.Akshatha,A.K.Karunakar,H.Anitha,U.Raghavendra,D.Shetty.Digital cameraidentification using PRNU:A feature based approach.Digital Investigation,vol.19,pp.69-77,2016", the authors describe a method for identifying smartphone cameras by analyzing captured photos, while in "A.Ferreira,L.C.Navarro,G.Pinheiro,J.A.dos Santos,A.Rocha.Laserprinter attribution:Exploring new features and beyond.Forensic ScienceInternational,vol.247,pp.105-125,2015", the authors propose a way to identify laser printing devices by analyzing printed pages. In a similar way, accelerometer and gyroscope sensors can be uniquely identified by analyzing the generated motion signals. This means that the motion signals recorded during user registration and authentication will inherently contain discriminative features of the device. Based on data from previous systems, it has been determined that it is much easier to identify a smartphone (with an accuracy of 98%) based on the recorded motion signals than to identify a user (with an accuracy of 93%) or an action (with an accuracy of 92%).
[0095] While dividing the signal into blocks eliminates the features that distinguish actions, it does not alleviate the problems caused by the features that distinguish devices. The main problem is that the discriminative features of the device and the discriminative features of the user (the smartphone owner) are entangled within the motion signals recorded during user registration and within the blocks obtained after our preprocessing step.
[0096] According to a further prominent aspect, the systems and methods disclosed herein for separating discriminative features of a smartphone user from motion signals provide a solution to the problem that does not require the user (smartphone owner) to perform additional steps other than standard registration using a single device. The disclosed methods and systems for separating discriminative features of a device's user from motion sensor data and authenticating the user from motion sensor data are partly inspired by the success of using style transfer in cycle-consistent generative adversarial networks for image-to-image transformation. As shown in "J.Y. Zhu, T. Park, P. Isola, A.A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In: Proceedings of IEEE International Conference on Computer Vision, pp. 2223-2232, 2017", Cycle-GAN can replace the style of an image with a different style while preserving its content. In a similar manner, the disclosed systems and methods replace the device used to record motion signals with a different device while preserving the discriminative features of the user. However, as described above, existing methods are specific to image style transfer, while the approach disclosed herein is specifically directed to transferring features from motion sensor data signals of a particular device to signals simulated for another device in the time domain.
[0097] Thus, in one or more embodiments, the system 100 implements at least one transformation algorithm (such as Cycle-GAN) for signal-to-signal transformation, as Figure 4 illustrated. Specifically, Figure 4 is an exemplary cycle-consistent generative adversarial network (Cycle-GAN) 400 for signal-to-signal transformation according to one or more embodiments. As Figure 4 shown and according to at least one embodiment, the input signal x recorded on device X is transformed using generator G to make it appear as if it were recorded on a different device Y. The signal is transformed back to the original device X using generator F. Discriminator D Y discriminates between the signal recorded on device Y and the signal generated by G. Generator G is optimized to deceive discriminator D Y , while discriminator D Y is optimized to separate samples in an adversarial manner. Additionally, the generative adversarial network (comprising generators G and F and discriminator D YThe (formed) is optimized to reduce the reconstruction error calculated after transforming the signal x back to the original device X. In at least one embodiment, the optimization is performed using stochastic gradient descent (or one of its many variants), which is an algorithm commonly used to optimize neural networks, as will be understood by those skilled in the art. The gradient is calculated with respect to the loss function and backpropagated through the neural network using the chain rule, as will be understood by those skilled in the art. In one or more embodiments, an evolutionary algorithm can be used to perform the optimization of the neural network. However, it should be understood that the disclosed method is not limited to optimization by gradient descent or evolutionary algorithms.
[0098] Adding the reconstruction error to the overall loss function ensures cycle consistency. In addition to transforming from device X to device Y, the Cycle-GAN is also trained to transfer from device Y to device X. Thus, ultimately, the disclosed system and method transform signals in both directions. In one or more embodiments, the loss function used to train the Cycle-GAN to perform signal-to-signal transformation in both directions is:
[0099]
[0100] , where:
[0101] · G and F are generators;
[0102] · D X and D Y are discriminators;
[0103] · x and y are motion signals (blocks) from device X and device Y, respectively;
[0104] · λ is a parameter that controls the importance of cycle consistency relative to the GAN loss;
[0105] · is the cross-entropy loss corresponding to the transformation from device X to device Y, where E[■] is the expected value and Pdata(■) is the probability distribution of the data samples;
[0106] · is the cross-entropy loss corresponding to the transformation from device Y to device X;
[0107] · is the sum of the cycle consistency losses of the two transformations, where ||■|| 1 is the l 1 norm.
[0108] As an addition to or an alternative for Cycle GAN, the U-GAT-IT model introduced in "J. Kim, M. Kim, H. Kang, K Lee. U-GAT-IT: Unsupervised generative attentional networks with adaptive layer-instance normalization for image-to-image translation. arXiv preprint arXiv:1907.10830, 2019" can be incorporated into the current systems and methods for signal-to-signal transformation. The U-GAT-IT model incorporates attention modules in both the generator and the discriminator, and incorporates a normalization function (adaptive layer instance normalization) with the aim of improving the transformation from the source domain to the target domain. The attention map is obtained using an auxiliary classifier, and the parameters of the normalization function are learned during training. In at least one embodiment, the loss function for training U-GAT-IT is:
[0109]
[0110] where:
[0111] · G and F are generators;
[0112] · D X and D Y are discriminators;
[0113] · x and y are motion signals (blocks) from device X and device Y respectively;
[0114] · λ 1 、λ 2 、λ 3 and λ 4 are parameters that control the importance of various loss components;
[0115] · is the least squares loss corresponding to the transformation from device X to device Y;
[0116] · is the least squares loss corresponding to the transformation from device Y to device X;
[0117] · is the sum of the cycle consistency losses of the two transformations and ||■|| 1 is the l 1 norm;
[0118] · is the sum of the identity losses that ensure the amplitude distributions of the input and output signals are similar;
[0119] · is the sum of the least squares losses that introduce the attention map.
[0120] In one or more embodiments, to obtain transfer results that generalize across multiple devices, several Cycle-GAN (or U-GAT-IT) models are used to transfer signals between several pairs of devices. In at least one embodiment, a fixed number of smartphones T are set and motion signals are collected from each of the T devices. In one or more embodiments, a fixed number of users who perform registration on each of the T devices are set. Data collection is done before the UBA system is deployed to production, i.e., the end user is never required to perform registration on the T devices, which would be infeasible. The number of trained Cycle-GANs is T such that each Cycle-GAN learns to transform signals from a specific device to all other devices and vice versa, as Figures 5A - 5D illustrated, which will be discussed in further detail below.
[0121] An objective of the systems and methods disclosed herein is to obtain a general Cycle-GAN model that is capable of transforming signals captured on a certain original device into one of the T devices in the group, regardless of the original device. In at least one embodiment, the same scope can be achieved by using different GAN architectures, various network depths, learning rates, or optimization algorithms. Generalization ability is ensured by learning to transform from multiple devices to a single device rather than just learning to transform from one device to another. Thus, during the training of the transformation algorithm (e.g., Cycle-GAN model), the transformation algorithm (e.g., Cycle-GAN) of the disclosed embodiments can be applied to transform signals captured on a certain user's device into one of the T devices in the group without the need to know the user or have information about the device he / she is using. Each time the original signal is transformed into one of the T devices, the characteristics of the owner's device are replaced with the characteristics of a certain device from the group of T devices while maintaining user-specific characteristics. According to one or more embodiments of the system for discriminative features of the user that distinguish devices from motion sensor data, by feeding both the original signal and the transformed signal into the ML system, the ML system can no longer learn the discriminative features specific to the original device. This occurs because the ML system is configured to place (classify) the original signal and the transformed signal in the same class, and the most relevant way to obtain such a decision boundary is by looking at the features that distinguish the user.
[0122] Exemplary Embodiment
[0123] Exemplary embodiments of the present system and method are described below with reference to Figures 5A - 5D 、Figures 6A - 6B and Figures 1A - 1D will be discussed in further detail along with the actual application of the technology and other actual scenarios in which systems and methods can be applied that can distinguish the user of the device based on motion signals captured by motion sensors of the mobile device and authenticate.
[0124] In one or more embodiments, the method disclosed herein provides a modified UBA execution pipeline. The modified UBA execution pipeline is capable of distinguishing the discriminative features of the user from the discriminative features of the actions and the device, as Figures 5A - 5D shown. Figures 5A - 5D FIG. shows a system and flowchart for distinguishing the discriminative features of a user of a mobile device from motion sensor data and authenticating the user according to one or more embodiments. Figures 5A - 5D The steps in the flowchart of can be performed using an exemplary system for distinguishing the discriminative features of the user of the device from motion sensor data of a mobile device 101 and / or a system server 105, such as system 100. Figures 5A - 5D The steps in the flowchart of are detailed as follows:
[0125] 1. During user registration or authentication, the mobile device 101 captures motion signals 310 from built-in sensors such as accelerometers and gyroscopes.
[0126] 2. The motion signals 310 are divided into smaller blocks 315.
[0127] 3. Using a set of trained Cycle-GANs 400, new signals 500 (transformed signals) are generated to simulate the authentication session of the user on other devices.
[0128] 4. Inside the ML system 505, by adopting a set of feature extraction techniques, relevant features are extracted from the original signal blocks and the transformed signal blocks to form feature vectors.
[0129] 5. During registration, inside the ML system 505, a classification model is trained on the feature vectors corresponding to the recorded and transformed signal blocks 500.
[0130] 6. During authentication, inside the ML system 505, the learned classification model is adopted for the feature vectors corresponding to the recorded and transformed signal blocks.
[0131] 7. During authentication, in order to distinguish the legitimate user (smartphone owner) from potential attackers, a voting scheme or a meta-learning model is adopted for the scores or labels obtained for the recorded and transformed blocks corresponding to the authentication session.
[0132] These steps and other steps are in Figures 6A - 6Bis further described and illustrated in the following exemplary methods.
[0133] According to one or more embodiments, Figure 6A a schematic block diagram is disclosed that shows a computational flow for distinguishing discriminative features of a user of a device from motion sensor data. In one or more embodiments, Figure 6A the method of Figure 1A can be performed by the present system, such as Figure 1A exemplary system 100. Although many of the following steps are described as being performed by mobile device 101 (
[0134] Referring to Figure 5A and Figures 1A - 1D , the process begins at step S105, where the processor of mobile device 101 is configured to cause at least one motion sensor of the mobile device (e.g., accelerometer 135, gyroscope 136) to capture data from the user in the form of one or more motion signals by executing one or more software modules (e.g., user interface module 170). In one or more embodiments, the motion signal is a multi-axis signal corresponding to the physical movement or interaction of the user with the device collected from at least one motion sensor of the mobile device during a specified time domain. The physical movement of the user can be in the form of a "gesture" (e.g., finger tap or finger swipe) or other physical interaction with the mobile device (e.g., picking up the mobile device). For example, the motion sensor can collect or capture motion signals corresponding to the user writing their signature in the air ("explicit" interaction) or the user tapping their phone ("implicit" interaction). Thus, the motion signal contains features specific to (distinguishing) the user, the action performed by the user (e.g., gesture), and the particular mobile device.
[0135] In one or more embodiments, the collection or capture of motion signals by a motion sensor of a mobile device may be performed during one or more predetermined time windows, which are preferably short time windows. For example, in at least one embodiment, the time window may be about 2 seconds. In at least one embodiment, such as during a user's enrollment, the mobile device may be configured to collect (capture) motion signals from the user by prompting the user to make a specific gesture or explicit interaction. Additionally, in at least one embodiment, the mobile device may be configured to collect motion signals without prompting the user, such that the collected motion signals represent implicit gestures or interactions of the user with the mobile device. In one or more embodiments, a processor of the mobile device 101 may be configured to save the captured motion signals in a database 185 of the mobile device by executing one or more software modules (e.g., database module 176), or alternatively, the captured motion signals may be saved in a database 280 of the backend server 105.
[0136] At step S110, a processor of the mobile device is configured to divide one or more captured motion signals into segments by executing one or more software modules (e.g., segmentation module 173). Specifically, the captured motion signals are divided into N smaller blocks or segments. As previously discussed, the disclosed systems and methods are configured to partially address the problem of distinguishing features specific to or exclusive to a particular user in a motion signal from features that distinguish an action performed by the user. Thus, by dividing the motion signal into segments, the discriminative features of the action performed by the user are eliminated. In one or more embodiments, the step of dividing one or more captured motion signals into segments includes dividing each motion signal into a fixed number of segments, where each segment has a fixed length. In one or more embodiments, at least some of the segments or blocks are overlapping segments or blocks.
[0137] Continuing to refer Figure 5A , at step S115, a processor of the mobile device is configured to transform the segments into transformed segments using one or more trained transformation algorithms (such as Cycle-Consistent Generative Adversarial Network (Cycle-GAN)) by executing one or more software modules 130 (e.g., transformation module 178). As previously discussed, the present systems and methods are configured to partially address the problem of distinguishing features specific to or exclusive to a particular user in a motion signal from features that distinguish the mobile device of the user. By using Cycle-GAN to transform the segments into transformed segments, the method effectively eliminates the discriminative features of the mobile device from the (one or more) collected motion data signals provided to the ML model.
[0138] More specifically, according to one or more embodiments, the chunks / fragments are considered as individual samples and these chunks / fragments are provided as input to one or more transformation algorithms (e.g., Cycle-GAN). In one or more embodiments utilizing Cycle-GAN, each Cycle-GAN is trained offline on a separate dataset of motion signals collected on a predefined set of T different devices. Given a fragment as input, Cycle-GAN transforms each fragment into a transformed fragment such that the transformed fragment corresponds to a fragment that would be recorded on a device different from the user's mobile device from the predefined set of T devices. In other words, the transformed fragment includes one or more synthetic motion signals that simulate motion signals originating from another device. For example, an input fragment of the motion signal x recorded on device X is transformed using Cycle-GAN (generator) G to make it look like it was recorded on a different device Y.
[0139] Each time a raw fragment is transformed using one of the Cycle-GANs, the characteristics of the user's own device are replaced with those of a certain device from the predefined set of T devices while maintaining user-specific characteristics. As previously mentioned, one purpose of this process is to separate the user-specific (discriminative) characteristics in the motion signal from those that are specific (discriminative) only to the device utilized by the user. Thus, according to at least one embodiment, in order to obtain user-specific results such that the user can be identified across multiple devices, several Cycle-GAN models (or U-GAT-IT or other transformation algorithms) are used to transform fragments between several devices. In one or more embodiments, a fixed number T of devices (e.g., smartphones) are set and fragments of motion signals are collected from each of the T devices. A fixed number of users who perform registration on each of the T devices are also set. Thus, a certain number of Cycle-GANs are trained such that each Cycle-GAN learns to transform fragments from a specific device to all other devices in the group and vice versa, as Figure 5A illustrated. Thus, according to at least one embodiment, the present system produces a general Cycle-GAN model that is capable of transforming signals captured on a certain original device into another device from a predefined set of T devices, regardless of the original device.
[0140] Note that, in at least one embodiment, this step can be implemented by using different GAN architectures, various network depths, learning rates, or optimization algorithms. Generalization ability is ensured by learning to transform from multiple devices to a single device rather than just learning to transform from one device to another. Thus, during the training of Cycle-GAN, Cycle-GAN can be applied to transform the signals captured on the user's device to another device in the set of T devices without the need to know the user or have information about the device he or she is using. Each time the original segment is transformed to one of the T devices in the set, the features of the user's device are replaced with the features of another device in the set of T devices while maintaining the user-specific features.
[0141] In one or more embodiments of step S115, once the input segment is transformed, the processor of the mobile device is configured to transform the segment back by executing one or more software modules (e.g., transformation module 178). For example, as described in the previous example, the input segment of the motion signal x recorded on device X is transformed using Cycle-GAN (generator) G to make it look like it was recorded on a different device Y. Then the transformed segment is transformed back to the original device X using Cycle-GAN (generator) F. According to at least one embodiment, discriminator D Y then discriminates between the signal recorded on device Y and the signal generated by Cycle-GAN (generator) G. Generator G is optimized to deceive discriminator D Y , while discriminator D Y is optimized to separate the samples in an adversarial manner. Additionally, in at least one embodiment, the entire system is optimized to reduce the reconstruction error calculated after transforming the signal x back to the original device X. Adding the reconstruction error to the overall loss function ensures cycle consistency.
[0142] Continuing to refer Figure 5A , at step S120, the processor of the mobile device is configured to provide the segment and the transformed segment to the machine learning system by executing one or more software modules (e.g., feature extraction module 172). The segment and the transformed segment provided as input to the machine learning system are treated as individual samples.
[0143] At step S125, the processor of the mobile device is then configured to extract discriminative features of the user from the segments and the transformed segments by executing one or more software modules using a machine learning system that applies one or more feature extraction algorithms. For example, in one or more embodiments during the user registration process, the processor is configured to extract relevant features from the segments and the transformed segments using one or more feature extraction techniques to form a feature vector. At step S127, the processor is configured to employ and train a learned classification model on the feature vectors corresponding to the segments and the transformed segments. In at least one embodiment, the machine learning system can be an end-to-end deep neural network, including the feature extraction (S125) and classification (S127) steps. In one or more embodiments, the machine learning system can be formed by two components (a feature extractor and a classifier) or three components (a feature extractor, a feature selection method—not shown—and a classifier). In any embodiment, there is a trainable component, i.e., a deep neural network or a classifier. The trainable component is typically trained on a dataset of samples (motion signal blocks) and corresponding labels (user identifiers) by applying an optimization algorithm (such as gradient descent) with respect to a loss function that expresses how well the trainable component can predict the correct label for the training data samples (the original or transformed segments collected during user registration), as will be understood by those skilled in the art. The purpose of the optimization algorithm is to minimize the loss function, i.e., to improve the prediction ability of the trainable component. At step S130, the method ends.
[0144] Figure 6B A schematic block diagram is disclosed showing a computational flow for authenticating a user on a mobile device from motion sensor data according to at least one embodiment. Now referring to Figure 5B , the method begins at step S105. As Figure 6B shown, steps S105 - S120 are the same steps as described above for the method shown in Figure 6A . Specifically, at step S105, the mobile device captures one or more motion signals from the user, and at step S110, the one or more captured motion signals are divided into segments. At step S115, the segments are transformed into transformed segments using one or more trained transformation algorithms (e.g., Cycle - GAN), and at step S120, the segments and the transformed segments are provided to the machine learning system.
[0145] Continuing to refer to Figure 5B and Figures 1A - 1D, after step S120, at step S135, the processor of the mobile device is configured to classify (e.g., score, assign class probabilities to) the segments and transformed segments by executing one or more software modules (e.g., classification module 174). More specifically, in one or more embodiments, the feature vectors representing the segments and transformed segments are analyzed and scored. In one or more embodiments, the segments and transformed segments may correspond to an authentication session such that the segments and transformed segments are scored according to a machine learning model (e.g., a classifier or deep neural network previously trained on data collected during user registration at step S127 in Figure 6A .
[0146] At step S140, the mobile device is configured to apply a voting scheme or a meta-learning model to the scores (e.g., class probabilities) assigned to the segments and transformed segments obtained during the authentication session by executing one or more software modules 130 (e.g., meta-learning module 175). By applying the voting scheme or the meta-learning model, the mobile device is configured to provide a consistent decision to authorize or reject the user. In one or more embodiments, the consistent decision for the session is based on the voting scheme or the meta-learning model applied to the scores of all segments and transformed segments.
[0147] Finally, at step S145, the mobile device is configured to determine whether the user is an authorized user based on the voting or meta-learning step by executing one or more software modules (e.g., meta-learning module 175). As previously mentioned, the segments and transformed segments are scored according to the class probabilities given as output by a machine learning model trained during registration to identify known users. Thus, based on the voting scheme or meta-learner applied to the scores of the corresponding segments and transformed segments, the processor is configured to determine whether the segments and transformed segments belong to a particular (authorized) user and thus ultimately determine whether the user is an authorized user or an unauthorized user. The authentication process disclosed herein refers to one-to-one authentication (user verification). At step S150, the method for authenticating a user based on motion sensor data ends.
[0148] Experimental results
[0149] In this section, experimental results obtained with the discrimination method disclosed herein are presented according to one or more embodiments. Two different datasets were used in the following experiments. The first dataset (hereinafter referred to as the 5x5 database) consists of signals recorded by accelerometers and gyroscopes from 5 smartphones when 5 individuals perform authentication using 5 smartphones. These individuals were required to change positions during authentication (i.e., stand up, or sit down and use the right or left hand), thus performing different general actions. In each position, signals from 50 authentications were captured, which means that each person performed a total of 1000 authentications on each of the 5 smartphones, i.e., the total number of sessions was 5000. The second dataset (hereinafter referred to as the 3x3 database) was formed with the same position changes as the first but with 3 different individuals and 3 different smartphones.
[0150] Signals from each database were divided into 25 blocks during the preprocessing phase. Signal blocks originating from the 5x5 dataset were used to train 5 Cycle - GANs with the aim of transforming signals back and forth across 5 devices. Then, each signal block obtained from the 3x3 database was fed into the trained GANs, thus obtaining a completely new set of signals in this way. These new (transformed) signals were further divided into two subsets: one for training and the other for testing the ML system for user identification based on motion signals.
[0151] As is well known, GAN training is highly unstable and difficult because GANs need to find the Nash equilibrium of a non - convex min - max game with high - dimensional parameters. However, in these experiments, it was observed that the loss function of the generative network decreased monotonically, supporting the idea of reaching an equilibrium point.
[0152] By dividing the signals into blocks, user - discriminative features were captured and the importance of movement (actions) was minimized. Clear evidence is that by feeding signal blocks (without applying GAN) into the ML system, a 4% increase in the accuracy of user discrimination was observed. Transforming user behavior across different mobile devices (by applying GAN) led to an additional 3% increase in user discrimination accuracy (relative to the accuracy obtained using only signal blocks). Therefore, it can be concluded that by simulating features from various devices, a more robust and invariant ML system with respect to device features was obtained.
[0153] As the number of mobile devices has grown, the frequency of attacks has increased significantly. As a result, various user behavior analysis algorithms have been proposed in recent years, including systems based on signals captured by motion sensors during user authentication. Currently, the main problem faced by algorithms based on motion sensor data is that it is difficult to distinguish user-specific features from action- and device-specific features. The system and method of the present application solve this problem by dividing motion signals into blocks and transforming the signals into other devices using a transformation algorithm (e.g., Cycle-GAN) without any changes to the user registration process.
[0154] In fact, dividing the signal into smaller blocks helps reduce the influence of the action (performed by the user) on the decision boundary of the ML system, thereby increasing the user identification accuracy by 4%, as shown in the example above. Additionally, it is well known that sensors from different devices vary due to manufacturing processes, even if the devices are of the same brand and model and from the same production line. This fact leads to a significant impact of mobile device sensors on the decision boundary of the ML system for user identification. According to one or more embodiments, the systems and methods disclosed herein that utilize GAN simulate authentication from multiple devices based on recorded motion signals from any device, reduce the influence of the mobile device on the ML system, and increase the influence of user-specific features. The disclosed systems and methods further improve the ML system accuracy by approximately 3%. Thus, overall, in one or more embodiments, the present system and method can improve performance by 7%, thereby reducing the false positive rate and the false negative rate.
[0155] Exemplary systems and methods for distinguishing discriminative features of a user from a device in a motion signal and for authenticating a user on a mobile device from the motion signal are set forth in the following items:
[0156] Item 1. A computer-implemented method for distinguishing discriminative features of a user from a device in a motion signal captured by a mobile device, the mobile device having one or more motion sensors, a storage medium, instructions stored on the storage medium, and a processor configured by executing the instructions, comprising:
[0157] Dividing, by the processor, each captured motion signal into segments;
[0158] Transforming, by the processor, the segments into transformed segments using one or more trained transformation algorithms;
[0159] Providing, by the processor, the segments and the transformed segments to a machine learning system; and
[0160] Using the processor, the machine learning system that applies one or more feature extraction algorithms extracts discriminative features of the user from the fragments and the transformed fragments.
[0161] Item 2. The method according to item 1, wherein the discriminative features of the user are used to identify the user when the user uses the device in the future.
[0162] Item 3. The method according to the preceding item, wherein the one or more motion sensors include at least one of a gyroscope and an accelerometer.
[0163] Item 4. The method according to the preceding item, wherein the one or more motion signals correspond to one or more interactions between the user and the mobile device.
[0164] Item 5. The method according to the preceding item, wherein the motion signals include discriminative features of the user, discriminative features of an action performed by the user, and discriminative features of the mobile device.
[0165] Item 6. The method according to item 5, wherein the step of dividing the one or more captured motion signals into the fragments eliminates the discriminative features of the action performed by the user.
[0166] Item 7. The method according to item 5, wherein the step of converting the fragments into the transformed fragments eliminates the discriminative features of the mobile device.
[0167] Item 8. The method according to the preceding item, wherein the one or more trained transformation algorithms include one or more Cycle-Consistent Generative Adversarial Networks (Cycle-GANs), and wherein the transformed fragments include synthetic motion signals that simulate motion signals originating from another device.
[0168] Item 9. The method according to the preceding item, wherein the step of dividing the one or more captured motion signals into fragments includes dividing each motion signal into a fixed number of fragments, wherein each fragment has a fixed length.
[0169] Item 10. A computer-implemented method for authenticating a user on a mobile device from motion signals captured by the mobile device, the mobile device having one or more motion sensors, a storage medium, instructions stored on the storage medium, and a processor configured by executing the instructions, comprising:
[0170] Using the processor to divide the one or more captured motion signals into fragments;
[0171] Use the processor to transform the segment into a transformed segment using one or more trained transformation algorithms;
[0172] Use the processor to provide the segment and the transformed segment to a machine learning system; and
[0173] Classify the segment and the transformed segment as belonging to an authorized user or an unauthorized user by assigning scores to each of the segment and the transformed segment using the processor; and
[0174] Apply a voting scheme or a meta-learning model to the scores assigned to the segment and the transformed segment using the processor; and
[0175] Use the processor to determine whether the user is an authorized user based on the voting scheme or the meta-learning model.
[0176] Item 11. The method according to item 10, wherein the step of scoring includes:
[0177] Compare the segment and the transformed segment with the features of the authorized user extracted from the sample segments provided by the authorized user during the registration process, wherein the features are stored on the storage medium; and
[0178] Assign scores to each segment based on a classification model.
[0179] Item 12. The method according to item 10 or 11, wherein the one or more motion sensors include at least one of a gyroscope and an accelerometer.
[0180] Item 13. The method according to items 10-12, wherein the step of dividing the one or more captured motion signals into segments includes dividing each motion signal into a fixed number of segments, wherein each segment has a fixed length.
[0181] Item 14. The method according to items 10-13, wherein at least a part of the segments overlap.
[0182] Item 15. The method according to items 10-14, wherein the one or more trained transformation algorithms include one or more Cycle-GANs, and wherein the step of transformation includes:
[0183] Transform the segment into the transformed segment via a first generator, which mimics the segments generated on another device; and
[0184] Re-transform the transformed segment via a second generator to mimic the segments generated on the mobile device.
[0185] Item 16. The method according to Items 10-15, wherein the transformed segment includes a synthetic motion signal that simulates a motion signal from another device.
[0186] Item 17. The method according to Items 10-16, wherein the providing step includes:
[0187] extracting features from the segment and the transformed segment using one or more feature extraction techniques by the processing to form a feature vector; and
[0188] applying a learned classification model to the feature vectors corresponding to the segment and the transformed segment.
[0189] Item 18. A system for distinguishing discriminative features of a user of a device from motion signals captured on a mobile device having at least one motion sensor and authenticating the user on the mobile device, the system including:
[0190] a network communication interface;
[0191] a computer-readable storage medium;
[0192] a processor configured to interact with the network communication interface and the computer-readable storage medium and execute one or more software modules stored on the storage medium, including:
[0193] a segmentation module that, when executed, configures the processor to divide each captured motion signal into segments;
[0194] a conversion module that, when executed, configures the processor to use one or more trained Cycle-Consistent Generative Adversarial Networks (Cycle-GANs) to transform the segments into transformed segments;
[0195] a feature extraction module that, when executed, configures the processor to extract the extracted discriminative features of the user from the segments and the transformed segments, wherein the processor uses a machine learning system;
[0196] a classification module that, when executed, configures the processor to assign scores to the segments and the transformed segments and determine whether the segments and the transformed segments belong to an authorized user or an unauthorized user based on the corresponding scores of these segments;
[0197] a meta-learning module that, when executed, configures the processor to apply a voting scheme or a meta-learning model based on the stored segments corresponding to the user based on the scores assigned to the segments and the transformed segments.
[0198] Item 19. The system according to Item 18, wherein the at least one motion sensor includes at least one of a gyroscope and an accelerometer.
[0199] Item 20. The system according to Items 18-19, wherein the conversion module configures the processor to: transform the segment into the transformed segment via a first generator, which mimics a segment generated on another device; and re-transform the transformed segment via a second generator to mimic a segment generated on the mobile device.
[0200] Item 21. The system according to Items 18-20, wherein the feature extraction module is further configured to apply a learned classification model to the extracted features corresponding to the segment and the transformed segment.
[0201] At this time, it should be noted that although most of the foregoing description is directed to systems and methods for using motion sensor data to distinguish discriminative features and user authentication of users, the systems and methods disclosed herein can be similarly deployed and / or implemented in scenarios, situations, and settings outside of the reference scenario.
[0202] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any implementation or what can be claimed, but rather as descriptions of features specific to particular embodiments that can be specific to a particular implementation. Certain features described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from the claimed combination can be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variant of the sub-combination.
[0203] Similarly, although operations are depicted in the figures in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or sequentially, or that all illustrated operations be performed to obtain a desired result. In some cases, multitasking and parallel processing can be advantageous. Additionally, the separation of the various system components in the foregoing embodiments should not be understood as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0204] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the present invention. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be noted that the use of ordinal terms such as "first", "second", "third", etc. in the claims to modify the claim elements themselves does not imply any priority, precedence or order of one claim element over another or the temporal order of performing the acts of the method, but is only used as a label to distinguish one claim element having a particular name from another element having the same name (were it not for the use of the ordinal term) to distinguish the claim elements. Furthermore, the language and terminology used herein are for descriptive purposes and should not be regarded as limiting. The use of "comprises", "comprising" or "has", "containing", "involves" and variations thereof herein is intended to cover the items listed thereafter and their equivalents as well as additional items. It should be understood that like numerals in the figures represent like elements in several figures, and not all embodiments or arrangements require all of the components and / or steps described and illustrated with reference to the figures.
[0205] Accordingly, exemplary embodiments and arrangements of the disclosed systems and methods provide a computer-implemented method, a computer system, and a computer program product for user authentication using motion sensor data. The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments and arrangements. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions.
[0206] The foregoing subject matter is provided by way of illustration only and should not be construed as limiting. Various modifications and changes can be made to the subject matter described herein without following the example embodiments and applications illustrated and described, and without departing from the true spirit and scope of the invention set forth in the appended claims.
Claims
1. A computer-implemented method for distinguishing discriminative features of a user of a device from motion signals captured by a mobile device having one or more motion sensors, comprising: providing, at a processor, one or more captured motion signals captured by the mobile device; dividing, by the processor, each captured motion signal into segments; transforming, by the processor, the segments into transformed segments using one or more trained transformation algorithms, wherein the one or more trained transformation algorithms include one or more cycle-consistent generative adversarial networks, and wherein the transformed segments include synthetic motion signals simulating motion signals originating from another device; providing, by the processor, the segments and the transformed segments as inputs to a machine learning system; and extracting, by the processor, the discriminative features of the user from the segments and the transformed segments using the machine learning system applying one or more feature extraction algorithms, wherein the motion signals include discriminative features of the user, discriminative features of actions performed by the user, and discriminative features of the mobile device, and the step of dividing the one or more captured motion signals into the segments eliminates the discriminative features of the actions performed by the user.
2. The method according to claim 1, wherein, the discriminative features of the user are used to identify the user when the user uses the device in the future.
3. The method according to claim 1, wherein, the one or more motion sensors include at least one of a gyroscope and an accelerometer.
4. The method according to claim 1, wherein, the one or more captured motion signals correspond to one or more interactions between the user and the mobile device.
5. The method according to claim 1, wherein, the steps of transforming the segments into the transformed segments and providing the segments and the transformed segments to the machine learning system effectively eliminate the discriminative features of the mobile device.
6. The method according to claim 1, wherein, the step of dividing the one or more captured motion signals into segments includes dividing each motion signal into a fixed number of segments, wherein each segment has a fixed length.
7. A computer-implemented method for authenticating a user of a mobile device from motion signals captured by the mobile device having one or more motion sensors, comprising: providing, at a processor, one or more captured motion signals; dividing, by the processor, the one or more captured motion signals into segments; transforming, by the processor, the segments into transformed segments using one or more trained transformation algorithms, wherein the one or more trained transformation algorithms include one or more cycle-consistent generative adversarial networks, and wherein the transformed segments include synthetic motion signals simulating motion signals originating from another device; providing, by the processor, the segments and the transformed segments as inputs to a machine learning system; and Using the processor and the machine learning system, classify the segment and the transformed segment as belonging to an authorized user or an unauthorized user by assigning a score to each of the segment and the transformed segment; Apply a voting scheme or a meta-learning model to the scores assigned to the segment and the transformed segment by the processor; and Based on the voting scheme or the meta-learning model, determine by the processor whether the user is an authorized user, The motion signal includes discriminative features of the user, discriminative features of an action performed by the user, and discriminative features of the mobile device, and the steps of converting the segment into the transformed segment and providing the segment and the transformed segment to the machine learning system effectively eliminate the discriminative features of the mobile device.
8. The method according to claim 7, wherein, The classifying step includes: Comparing the segment and the transformed segment with the features of the authorized user extracted from the sample segments of the authorized user during the registration process, wherein the features are stored on a storage medium; and Assigning a score to each segment based on a classification model.
9. The method according to claim 7, wherein, The one or more motion sensors include at least one of a gyroscope and an accelerometer.
10. The method according to claim 7, wherein, The step of dividing the one or more captured motion signals into segments includes dividing each motion signal into a fixed number of segments, wherein each segment has a fixed length.
11. The method according to claim 7, wherein, At least a portion of the segments are overlapping.
12. The method according to claim 7, wherein, The converting step includes: Transforming the segment into the transformed segment via a first generator that mimics segments generated on another device; and Re-transforming the transformed segment via a second generator to mimic segments generated on the mobile device.
13. The method according to claim 7, wherein, The providing step includes: Using the processor to extract features from the segment and the transformed segment using one or more feature extraction techniques to form feature vectors; and Applying a learned classification model to the feature vectors corresponding to the segment and the transformed segment.
14. A system for differentiating discriminative features of a user of a device from motion signals captured on a mobile device having at least one motion sensor and authenticating the user of the mobile device, the system comprises: A network communication interface; A computer-readable storage medium; A processor configured to interact with the network communication interface and the computer-readable storage medium and execute one or more software modules stored on the storage medium, the software modules including: A segmentation module that, when executed, configures the processor to divide each captured motion signal into segments; A conversion module that, when executed, configures the processor to use one or more trained cycle-consistent generative adversarial networks to convert the segment into a transformed segment, where the transformed segment includes a synthetic motion signal that mimics a motion signal originating from another device; A feature extraction module that, when executed, configures the processor to extract the user's extracted discriminative features from the segment and the transformed segment, where the processor uses a machine learning system; A classification module that, when executed, configures the processor to assign scores to the segment and the transformed segment and determine whether the segment and the transformed segment belong to an authorized user or an unauthorized user based on the respective scores of these segments; A meta-learning module that, when executed, configures the processor to apply a voting scheme or a meta-learning model to the scores assigned to the segment and the transformed segment based on the stored segments corresponding to the user; The motion signal includes the discriminative features of the user, the discriminative features of the action performed by the user, and the discriminative features of the mobile device, and dividing one or more captured motion signals into the segments eliminates the discriminative features of the action performed by the user.
15. The system according to claim 14, wherein, the at least one motion sensor includes at least one of a gyroscope and an accelerometer.
16. The system according to claim 14, wherein, the conversion module configures the processor to: transform the segment into the transformed segment via a first generator that mimics a segment generated on another device; and re-transform the transformed segment via a second generator to mimic a segment generated on the mobile device.
17. The system according to claim 14, wherein, the feature extraction module is further configured to employ a learned classification model for the extracted features corresponding to the segment and the transformed segment.
Citation Information
Patent Citations
Method and device for authenticating user using user's behavior pattern
KR1020190099156A
System and method for user recognition using motion sensor data
US20190286242A1