Methods and devices for user recognition
A method and system using machine learning on user device data from multiple sensors to verify user behavior patterns improve authentication reliability by detecting fraudulent impersonation, addressing the limitations of single-sensor biometric methods.
Patent Information
- Application Number
- JP2023530649
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-20
- Filing Date
- 2021-11-19
- Publication Date
- 2026-08-26
- Estimated Expiration
- 2041-11-19
AI Technical Summary
Existing user authentication methods, such as face and fingerprint recognition, are not 100% reliable, allowing for the possibility of fraudulent impersonation or unauthorized access.
A computer-implemented method and system that processes data from multiple sensors on a user device to generate user verification data, using machine learning models to determine the probability that a user interacting within a specific time interval is the same user who interacted at other times, combining data from inertial, camera, and radio sensors to verify user behavior patterns.
Enhances the reliability of user authentication by providing a secondary verification layer that can detect fraudulent impersonation, even when primary biometric methods fail, by analyzing user behavior patterns across multiple sensors.
Smart Images

Figure 0007911541000001 
Figure 0007911541000002 
Figure 0007911541000003
Abstract
Description
Technical Field
[0001] Technical Field The present invention relates to a method and apparatus for user recognition, and more particularly but not exclusively, to a computer-implemented method for enabling a computer to recognize whether a user interacting with a user device within an identified time interval is the same user as a user who interacted with the device at other times.
Background Art
[0002] Background Many interactions between a user and a user device may require authentication. For example, the interaction may involve using application software on a user device, such as a mobile phone or a computer, to verify the user's identification information in order to approve entry into a building or vehicle or to perform any process that requires user recognition. Conventionally, a user's interaction with a user device is authenticated at the time of the interaction. For example, face recognition and / or fingerprint recognition can be used to verify the user's identification information, and if the verification fails, the interaction may be rejected. If the verification is successful, the authenticated interaction may continue. However, the reliability of the verification is typically not 100%, and the user during a given interaction may not be the intended user.
Summary of the Invention
[0003] Summary According to a first aspect of the present invention, there is provided a computer-implemented method for enabling a computer to recognize a user interacting with a user device within an identified time interval by processing data in the user device to generate user verification data for use within an interaction verification system, the method comprising: deriving first user behavior data by processing a first plurality of sets of data, each generated by a respective one of a plurality of different elements of a user device including at least one sensor, each representing that the user is interacting with the user device; Identifying at least a first time interval related to the user's interaction with the user device, Deriving second user behavior data by processing a second set of data, each generated by multiple different elements of a device including at least one sensor, each representing that the user is interacting with the user device during at least a first time interval, and The user device transmits user confirmation data based on the first user behavior data and the second user behavior data from the user device to the interaction confirmation system. A method is provided that includes this.
[0004] This method allows the dialogue verification system to process verification data and determine the probability that the user interacting with the user device within a first time interval is the same user who interacted with the device at the time associated with a first set of data.
[0005] In one example, this method involves identifying a first time interval as the time interval at which the dialogue occurs.
[0006] This makes it possible for a second set of user behavior data to influence the user's behavior during interactions.
[0007] In one example, this method identifies a second time interval as a time interval in which the dialogue occurs before it, and / or Identifying a third time interval as the time interval in which the dialogue takes place afterward. Includes, The second set of data represents the user's interaction with the device during the first time interval and the second and / or third time intervals, respectively.
[0008] This allows the second user behavior data to better represent user behavior, under the assumption that the user is the same for the first, second, and third time intervals.
[0009] In one example, this method involves collecting second user behavior data in response to receiving an indication that a dialogue is in progress.
[0010] This makes it possible to collect appropriate data regarding the duration of the dialogue.
[0011] In one example, this method is To store second user behavior data in the storage system on the user device, Receiving timing data indicating the first time interval from the dialogue confirmation system, and To obtain second user behavior data from the memory system based on timing data. Includes.
[0012] This makes it possible to identify and retrieve data related to the problematic dialogue for use when processing by the dialogue verification system.
[0013] One example involves using a hardware abstraction function module configured to derive first and second user behavior data, which is configured to convert data generated by multiple different elements of the user device into converted element data having a normalized format.
[0014] Generating converted element data in a normalized format allows the interactive verification system to process the data regardless of the characteristics of the specific user device.
[0015] For example, deriving first and second user behavior data involves using a data processing function module configured to perform summarization, aggregation, and joining functions on the transformed element data in order to generate processed element data.
[0016] This makes it possible to convert the collected raw data into typically summarized and compressed processed data for use within the user behavior function module.
[0017] In one example, deriving the first user behavior data includes using a user behavior function module configured to extract information regarding typical behavior of a user from processed element data related to the first plurality of sets of data.
[0018] This enables extraction of information from the processed data regarding how a user typically operates a device used to conduct an interaction.
[0019] In one example, deriving the second user behavior data includes using a behavior function module configured to extract information regarding the behavior of a user from processed element data related to the second plurality of sets of data. This enables extraction of user behavior information related to a certain period before, during, and after an interaction is typically conducted.
[0020] In one example, the user confirmation data includes an output from a machine learning model.
[0021] In one example, parameters for the machine learning model are received from a verification system.
[0022] In one example, the input to the machine learning model includes the first user behavior data and the second user behavior data, and the user confirmation data includes an output of the machine learning model.
[0023] In one example, the output of the machine learning model includes the probability that a user within a first time interval is different from the user corresponding to the first user behavior data. The machine learning model can be a deep neural network (DNN) trained to detect abnormal time intervals within a series of time intervals. For example, the machine learning model can be trained by supervised learning using a sequence of sets of user behavior data where most of them are known to be from a given user and one or more of them are known to be from a different user.
[0024] Alternatively, the input to the machine learning model may include at least a first set of data and a second set of data, and the output of the machine learning model may include first user behavior data and second user behavior data. In one example, the machine learning model is trained by using unsupervised learning to sort interactions in trial data into clusters. The machine learning model can process individual time intervals to estimate which cluster a time interval belongs to, and user confirmation data may include estimation of which cluster a time interval belongs to. This allows the interaction confirmation system to compare which cluster a first time interval belongs to with a cluster among several clusters to which the time interval corresponding to the first data belongs.
[0025] According to a second aspect of the present invention, a user device is provided which includes a processor configured to perform a method for enabling a computer to recognize a user interacting with the user device within an identified time interval by processing data in the user device and generating user recognition data for use in an interaction recognition system.
[0026] According to a third aspect of the present invention, a computer program is provided which, when executed on a computer, causes a computer to perform steps of a computer-implemented method for enabling the computer to recognize a user interacting with a user device within an identified time interval by processing data in a user device and generating user recognition data for use in an interaction recognition system.
[0027] A fourth aspect of the present invention provides a non-temporary computer-readable storage medium that holds instructions for causing one or more processors to perform steps of a computer-implemented method for enabling a computer to recognize a user interacting with a user device within an identified time interval by processing data in a user device and generating user recognition data for use in an interaction recognition system.
[0028] A fifth aspect of the present invention provides a system for confirming a dialogue after it has taken place, including a user device and a dialogue confirmation system. Typically, the dialogue confirmation system is configured to process user confirmation data in order to confirm a given dialogue.
[0029] One example includes a dialogue verification system that includes customer and end-user profile modules configured to store user verification data.
[0030] This allows the interaction confirmation system to confirm the interaction even when there is no current connection to the user device.
[0031] In one example, the dialogue verification system includes a dialogue verification module configured to process user verification data to give an estimate of the probability that a given dialogue contains a given user.
[0032] This means that the dialogue verification system may be able to verify, with a certain degree of confidence, whether or not it is actually an instance of fraudulent authentication in question.
[0033] One example includes a data processing module configured to determine data processing rules applied by the user device and to transmit data indicating the data processing rules to the user device.
[0034] This allows the determination of data processing rules, which typically require significant data processing power, to be performed within a processor outside the user device, and furthermore, the rules can be developed using, for example, artificial intelligence techniques based on data from multiple devices. The interaction confirmation system may include a machine learning model to be used in determining parameters for use in the corresponding machine learning model for the user device.
[0035] Further features and advantages of the present invention will become apparent from the following description of examples of the present invention, which will be made with reference to the accompanying drawings.
[0036] Brief explanation of the drawing To make the present invention easier to understand, an example of the present invention will now be described with reference to the accompanying drawings. [Brief explanation of the drawing]
[0037] [Figure 1] This is a schematic diagram showing multiple user devices communicating with the dialogue confirmation system. [Figure 2] This is a schematic diagram showing a user device configured to process data in order to generate user verification data to be sent to a dialogue verification system. [Figure 3] This is a schematic diagram showing an interactive confirmation system for processing user confirmation data received from at least one user device. [Figure 4] This is a schematic diagram showing the system, including user equipment and an interaction confirmation system. [Figure 5] This is a schematic diagram showing a system including a user device and an interaction confirmation system that uses a machine learning model within the user device to generate first and second user behavior data. [Figure 6] This is a schematic diagram showing a system that includes user equipment and an interaction verification system, which uses a machine learning model within the user's device to generate user verification data. [Figure 7] This demonstrates the training of a machine learning model in a dialogue confirmation system. [Figure 8] This indicates that the dialogue verification system will send machine learning model parameters to the machine learning models on each user's device. [Figure 9] This shows the signal flow within a machine learning model, including a deep neural network. [Figure 10] This document presents an example of a training method for deep neural networks during the training and deployment phases. [Figure 11] This is a collaboration diagram illustrating the activation of new users. [Figure 12]This is a collaboration diagram illustrating the execution of further user initial setup functions. [Figure 13] This is a collaboration diagram related to the problematic dialogue. [Figure 14] This is a block diagram showing the configuration of the backend system to be used for the customer. [Figure 15] This block diagram shows the configuration of a backend system shared among multiple customers. [Figure 16] This flowchart illustrates how to process data within a user device to generate user verification data for use within a dialogue verification system. [Modes for carrying out the invention]
[0038] Detailed explanation Examples of the present invention are described in the context of a system for verifying an interaction after it has taken place. An example of an interaction with a biometric identification system for accessing a building or vehicle is described, but it will be understood that the present invention is not limited to these examples and may relate to verification of other interactions, such as authentication of identification information for verifying financial transactions, or any interaction with a user device where it is necessary to verify whether the user interacting with the user device within an identified time interval is the same user who interacted with the device at other times. User verification data is generated in the user device from sensors and other elements of the user device, represents the user's behavior during the interaction and at other times, and is transmitted to the interaction verification system. In this example, the interaction verification system provides a second level of identification in addition to the existing biometric identification system, so that if a first level of identification event is at issue, the second level provided by the system can be used to verify whether the first level of identification was wrong or abnormal, such as in the case of impersonation, simulated impersonation, or forced authentication.
[0039] In one example, a user authentication format using fingerprint recognition is verified, and verification is based on user behavior data derived from inertial sensors within the user device. When using a fingerprint sensor, the user may perform a characteristic series of movements that can be detected using inertial sensors or accelerometers. Movements measured in three or more axes (for example, acceleration and angular velocity can be considered as a six-axis inertial frame) can be recorded before, during, and after the interaction with the fingerprint sensor. User behavior data for identified interactions can be compared with user behavior data for interactions recorded at other times. In addition to data from inertial sensors, user behavior data may be based on data from one or more cameras and / or radio sensors. Cameras and radio sensors provide further data representing the background environment as part of the user behavior, such as ambient light conditions and color and typical radio frequency signal levels. Inertial sensors, cameras, and radio sensors can also be used to generate user behavior data to verify, for example, face detection and voice detection.
[0040] As shown in Figure 1, the system includes one or more user devices 1a, 1b, 1c, such as mobile phones or computers, configured to generate user verification data, and a dialogue verification system 2, typically implemented by data processing outside the user devices, which may be called "backend" data processing. The backend processing can be implemented within a data processor at a central station, or it can be implemented by distributed or cloud processing. One or more user devices 1a, 1b, and 1c are shown to communicate with the dialogue verification system 2 via a data network 3. The data network may include a cellular radio network and other data connections.
[0041] Figure 2 is a schematic diagram showing a user device 1 configured to process data to generate user confirmation data for use within a dialogue confirmation system. As shown in Figure 2, the user device has multiple different elements 4 used to generate multiple sets of data from which first user behavior data is derived. Each set of data represents a user interacting with the user device.
[0042] These elements can be sensors in a user device, such as a camera, microphone, inertial sensor, temperature sensor, fingerprint sensor, keyboard, touchpad, and mouse. One or more of these elements may include wireless interfaces of the device, such as a WiFi interface, GPS / GNSS interface, Bluetooth interface, cellular wireless interface, and NFC interface. These elements may include wired interfaces, such as a USB interface. These elements may also include a screen interface, touchscreen interface, speaker or earphone interface, operating system, and timer. In each case, this allows for the derivation of data representing user interaction involving one or more of these elements. Peripheral interfaces used for identification, such as a keyboard if identification requires the input of a username and password, may be considered a specific type of sensor.
[0043] User device 1 is configured to derive first user behavior data from a first set of data, each of which is generated by at least some of several different elements 4 of the user device. In the first example, user behavior data is derived from the outputs of a fingerprint sensor and an inertial sensor. In the second example, user behavior data is derived from the outputs of an inertial sensor and a camera. In the third example, user behavior data is derived as a function of time from the outputs of an inertial sensor, a camera, and a keyboard.
[0044] The user device is configured to identify at least a first time interval related to an interaction in which the user of the user device is involved, and to derive second user behavior data from a second set of data, each generated by several different elements of the device, each representing that the user is interacting with the user device during at least the first time interval. Identifying at least the first time interval may include receiving instructions from the interaction confirmation system to identify the at least first time interval. For example, these instructions may be instructions for the time of the interaction being queried or the interaction in question. The first time interval may be, for example, the time during which fingerprint recognition or facial recognition authentication takes place.
[0045] As shown in Figure 2, the user device includes a hardware abstraction function module 5, a data processing module 6, a user behavior module 10, and an interaction behavior module 11. The hardware abstraction function module 5, the data processing module 6, and the user behavior module 10 are used to derive first user behavior data, and the hardware abstraction function module 5, the data processing module 6, and the interaction behavior module 11 are used to derive second user behavior data.
[0046] The hardware abstraction function module 5 is used to derive first and second user behavior data by converting data generated by multiple different elements 4 of the user device into converted element data having a normalized format for the interaction confirmation system. This allows the interaction confirmation system to process the data regardless of the characteristics of a particular user device. The data processing function module 6 has summarization 7, aggregation 8, and joining 9 function blocks. These act on the converted element data to generate processed element data. This allows the collected raw data to be converted into typically summarized processed data for use within the user behavior function module.
[0047] Hardware abstraction module 5 converts data from user device elements into a common normalized format that conforms to user devices that enable interactive verification. For example, various user devices may have different camera resolution specifications, and the hardware abstraction module is responsible for converting data from the camera to provide data that conforms to other functional modules that need to collect, process, and store data, regardless of the specific user device.
[0048] The data processing module 6 transforms the collected raw data into multi-level processed data based on the data received from the hardware abstraction module 5. The data processing functions of the data processing module 6 can be divided into three main classes: summarization, aggregation, and merging. These functions can be performed by programmed computational algorithms as well as artificial intelligence functions such as machine learning models. This module also functions on the backend side, i.e., on the user device side of the corresponding module 17 in the dialogue confirmation system 2. The data processing module 6 can be called a data processing and artificial intelligence module.
[0049] The user behavior module 10 extracts information about how a user typically operates the user device based on the data provided by the data processing module 6. This information, conveyed within the first user behavior data, may relate to the device identifier, how often and when the device is used during a day and a week. This information may also relate to behaviors such as pressing or swiping keys with one or both hands to input data. This information may also relate to the most frequently used applications or, for example, the location where the user device is used.
[0050] The dialogue behavior module 11 extracts detailed user behavior information for a certain period before, during, and / or after an interaction, based on the data provided by the data processing module 6. Its purpose is similar to that of the user behavior module 10, but the dialogue behavior module 11 focuses particularly on how the user operates the device during customer-related interactions. The dialogue behavior module 11 provides second user behavior data.
[0051] For example, to record the conversation in detail, the conversation recording module may record conversation-related data such as screenshots, keystrokes, video, audio, and fingerprint authentication. Recording can be activated at various levels of detail, and the level of detail of the raw data or the data processed by the data processing module 6 may be determined by technical and legal constraints, such as privacy constraints.
[0052] The storage and data protection module 13 stores the collected data on the user device's memory, taking into account any technical and / or legal constraints that may limit the amount and / or type of data that can be retained, such as privacy. The storage and data protection module 13 also aims to protect the data from corruption or deletion that an end user or unauthorized user may attempt to perform in the event of simulated or unsimulated fraudulent impersonation.
[0053] The backend communication module 12 enables communication with the backend system, i.e., the dialogue confirmation system. The backend communication module 12 can also manage technical and / or legal constraints that limit the amount and / or type of data that can be transferred from the user device to the backend. The backend communication module 12 transmits user confirmation data 14, based on first user behavior data and second user behavior data, from the device to the dialogue confirmation system. The user confirmation data may include first user behavior data and second user behavior data. Alternatively, the user confirmation data may include data derived by processing first user behavior data and second user behavior data. For example, the user confirmation data may be the output of a machine learning model whose inputs are first user behavior data and second user behavior data.
[0054] Figure 3 is a schematic diagram showing an example of the dialogue confirmation system 2. The dialogue confirmation system 2 is configured to confirm a given dialogue by processing user confirmation data which may include first user behavior data and second user behavior data received from the user device 1.
[0055] As can be seen in Figure 3, the dialogue confirmation system includes a user device communication module 15 for enabling the reception of user confirmation data from the user device, and customer and end-user profile modules 16 configured to store the user confirmation data.
[0056] The dialogue confirmation system 2 includes a data processing module 17, which includes modules for data summarization 18, aggregation 19, and combination 20, and a data processing module 21 for determining data processing rules, such as parameters for a machine learning model that can be determined as part of an artificial intelligence system. The data processing module 17 is configured to determine data processing rules to be applied by the user device 1 and to transmit data indicating the data processing rules to the user device 1 via the user device communication module 15.
[0057] The dialogue verification system includes a dialogue verification module 22 configured to process first and second user behavior data to give an estimate of the probability that a given dialogue includes a given user.
[0058] The user device communication module 15 mirrors the communication module on the user device to the backend, thereby protecting communication with the user device.
[0059] The customer and end-user profile module 16 stores information about the customer and end-user in accordance with the confirmation of the interaction, such as the amount and / or type of data that can be stored and / or transferred to the backend in accordance with the privacy agreement accepted by the end-user.
[0060] The storage and data protection module 24 stores the collected data on the backend after it has been transmitted to the backend by the user device, taking into account any technical and / or legal constraints that may necessitate limiting the amount and / or type of data that can be retained.
[0061] A data processing module 17, which may include artificial intelligence functions and may be called a data processing and artificial intelligence module, can perform the same functions as the corresponding module on the user device, namely summarization, aggregation, and merging, on the backend side whenever, for example, the necessary data on the user device is no longer available, while a copy of that data remains available on the backend side. However, this module on the backend side determines data processing rules and / or AI rules, such as parameters for machine learning modules that the corresponding module on the user device must apply. These rules are determined centrally, and the actual application of the rules is left to the user device.
[0062] The dialogue verification module 22 is a module that can verify with a certain degree of confidence whether the dialogue in question was an actual simulated or unsimulated instance of fraudulent authentication.
[0063] The customer communication module 23 implements an interface between the customer's IT system and the dialogue verification backend. The customer can request that dialogue verification be performed for the dialogue in question and receive the verification result provided by the dialogue verification module. The customer is the entity requesting dialogue verification. The dialogue can be a dialogue conducted by a user via a user device and may include user authentication.
[0064] Figure 4 is a schematic diagram showing an example of a system including a user device 1a and a dialogue verification system 2. As shown in Figure 4, data from a hardware element 4 including at least one sensor is processed by a hardware abstraction module 5 to generate first abstraction data 31 relating to a time other than a first time interval related to a given dialogue, and second abstraction data 32 relating to a first time interval related to a given dialogue. The first abstraction data 31 is processed by a user behavior module 10 to generate first user behavior data 33, and the second abstraction data 32 is processed by a dialogue behavior module 11 to generate second user behavior data 34. In this example, user verification data including the first user behavior data 33 and the second user behavior data 34 is sent to the dialogue verification system 2. In the dialogue verification system 2, the user verification data is processed by a customer and end-user profile module 16, a data processing module 17, and a dialogue verification module 22. This results in an output 35 which may indicate the probability that a user interacting with the user device within an identified time interval is the same user who interacted with the device at other times.
[0065] Figure 5 is a schematic diagram showing a system including a user device 1a and an interaction confirmation system 2 that uses a machine learning model 37 running on a processor 36 in the user device to generate first user behavior data 33 and second user behavior data 34. The first user behavior data 33 and second user behavior data 34 may be processed within the interaction confirmation system to generate an output 35 which may indicate the probability that a user interacting with the user device within an identified time interval is the same user who interacted with the device at another time.
[0066] Figure 6 is a schematic diagram showing a system including a user device and an interaction confirmation system, similar to that in Figure 5, except that the machine learning model 37 in the user device generates user confirmation data 14 that does not include the first and second user behavior data in this example. In this example, the first and second user behavior data are input to the machine learning model.
[0067] Figure 7 shows the training of the machine learning model 38 in the dialogue verification system. The machine learning model 38 in the dialogue verification system has similar characteristics to the machine learning model 37 in the user device, and therefore the parameters learned for the machine learning model 38 in the dialogue verification system can be used for the machine learning model 37 in the user device. Figure 8 shows the transmission of machine learning model parameters from the dialogue verification system to the machine learning models in each user device.
[0068] As shown in Figure 8, there may be machine learning models (such as neural network models, conventionally referred to as deep neural networks (DNNs)) in each user device 1a, 1b, and 1c, and copies in the backend system 2. The machine learning models in the backend system can be trained using trial data before deployment to the live system. Then, parameters for the machine learning model (such as weights in the case of a DNN) resulting from the training can be sent to load onto the machine learning models in the user devices. This makes it possible to obtain trial data in situations where privacy issues may not be a major constraint. If privacy requirements permit, it may be possible to train the machine learning model using data from the live system. Updated weights may be periodically sent to the user devices for use within the machine learning model in the user devices. Typically, there are training and deployment phases. In the training phase, the machine learning models in the backend are trained using sample data from a number of users, not necessarily including the final end-users of the deployment system.
[0069] In the first example, the DNN is trained by supplying batch-based data, each batch containing data representing several interactions, most of which are from a given user for that particular batch, but may include one or more interactions from different users. These different users represent fraudulent interactions. A large number of these batches need to be assembled for the training phase, using various combinations of data from trial participants. For each batch, one trial participant is designated as the legitimate user, and any other trial participants whose interactions are included in the batch are designated as fraudulent users.
[0070] The DNN generates a probability for each dialogue in a batch that the dialogue is anomalous, i.e., from a different user. During training, known anomalous dialogues are labeled with a probability of 1, and known legitimate dialogues are labeled with a probability of 0. The DNN is trained by a supervised learning process used to update the DNN's parameters to minimize the loss function property of the mismatch between labeled probabilities and predicted probabilities. In this way, the DNN is trained to accept data corresponding to a set of dialogues and to generate a probability that each of them is anomalous, i.e., fraudulent. The DNN can be configured to accept data corresponding to a set of dialogues simultaneously (i.e., different inputs to the DNN receive data about different dialogues) or to receive data corresponding to a set of dialogues serially (e.g., using a recurrent neural network (RNN) architecture or a bidirectional RNN architecture). The DNN's parameters (weights) are not user-specific. The DNN is trained to detect anomalousness in a set of dialogues, and the accuracy of detection should improve as training progresses across a large number of datasets. The data is appropriately formatted and processed data from various sensors of the device. The user dialogue data 14 transmitted to the dialogue verification system 2 may include data on the probability that each dialogue is abnormal. If applying the second user behavior data to a machine learning model indicates that there is a high probability that the dialogue represented by the data is abnormal, and if applying the first user behavior data to a machine learning model indicates that there is a low probability that the dialogue represented by the data is abnormal, this is indicated by the user dialogue data. The user verification system processes the user dialogue data to determine whether the user interacting with the user device within an identified time interval is the same user who interacted with the device at other times, as indicated by the first user behavior data.
[0071] The input data for "dialogue" can, in some cases, be background data over a certain period of time rather than actual dialogue. However, in some cases, such as a combination of a fingerprint sensor and an accelerometer, the appropriate data is related to the actual dialogue, including the fingerprint sensor.
[0072] Figures 9 and 10 show the first example. Input batches of training data T1-T n This is shown as follows, where T1 is data related to user 1's interaction, and so on. Each bus T of the training data n This may include sets of data derived from the outputs of several elements of the device, such as a fingerprint or facial recognition device, an inertial sensor and a camera, and / or a radio wave sensor.
[0073] For each dialogue, generate a probability that the dialogue is abnormal. Output P1~P n This is shown as follows, where P1 is a probability, an anomaly in user 1, and so on. The DNN has multiple layers, DNN1 is the input layer, and (DNN2...DNN(N)) are hidden layers. Solid arrows represent forward path data flowing during training and within the introduced system. Dashed arrows represent backpropagation only during training.
[0074] Each bus T in the training data n This may include a set of data derived as a function of time from the outputs of several elements of the device, such as the outputs of an inertial sensor and a camera.
[0075] In the second example, a machine learning model is trained using unsupervised learning to sort the dialogues in the trial data into categories. These categories can be so-called clusters, and the machine learning model may implement a clustering algorithm, such as k-means clustering, Gaussian mixture clustering, or DNN-based clustering. This process allows the machine learning model to learn to identify clusters of similar dialogues without being told in advance what the categories should be. The machine learning process can divide the dialogues into, for example, k clusters, where k can be predetermined or learned from the data. The number of clusters is typically much smaller than the number of users.
[0076] Once implemented, the machine learning model can process individual conversations and estimate which cluster a conversation belongs to. The machine learning model can directly determine which cluster each conversation on a user device belongs to, or it can generate the probability that each conversation on a user device is in each cluster. From this, it can be determined that one of the conversations is in a different cluster or has a high probability of being in a different cluster, and is therefore abnormal. Alternatively, it can be determined that all conversations are in the same cluster, or that each has a similar probability of being in each cluster, which may indicate that the same user was involved in each conversation. The user conversation data 14 transmitted to the conversation confirmation system 2 may include data-related conversations to clusters.
[0077] In the first and second examples, a machine learning model can be implemented using standard software libraries, and the code is compiled in the backend before training. Once the machine learning model in the backend is trained, the parameters for the machine learning model are sent to the user device for loading onto the machine learning model on the user device. The deployed machine learning model may be implemented using firmware or software that runs, for example, a standard GPU processor or other processing means.
[0078] In other examples, digital signal processing may be used to implement data processing modules that do not necessarily need to be trained by machine learning. Digital signal processing may include summarization 7, aggregation 8, and joining 9 functions, which operate as follows: The summarization function transforms raw data, typically in a normalized form, into a summary that retains some key elements that may be required as input by other function modules. For example, a face recognition function can capture raw data produced by a camera and determine whether the face of the person using the device corresponds to person "A" and not person "B". Another example is a function that can determine whether certain text has been entered on the device by typing with one hand, both hands, or only certain fingers.
[0079] Aggregation function 8 performs statistical analysis on raw or already summarized data to later identify typical ways in which the device is used. For example, one function could evaluate the average length of text messages typed on a messaging system, since some users tend to break long messages into shorter ones, while others type a single long message instead. Another example is whether facial recognition consistently identifies the same person in front of the device (who is likely to be the device's regular user), or whether the device is frequently used by a variety of people.
[0080] The function of the combining function 9 enables the combination of raw or already summarized or aggregated data from multiple elements of a user device, such as sensors, into a new type of data that can then be further summarized, aggregated, or combined. For example, information related to keyboard use, such as whether multiple fingers are used for typing, swiping, etc., can be combined with video data in a way that user recognition utilizes both pieces of information together.
[0081] Generally speaking, there may be various levels of data processing, aggregation, and combination that can be further combined to best provide useful information to other modules.
[0082] The above functions can be implemented using two different methods that are not mutually exclusive but can be combined: programmed algorithms and artificial intelligence (AI) algorithms. Using programmed algorithms, the output, i.e., processed data, is calculated by applying an automated sequence of statements, formulas, conditional expressions, etc., that essentially correspond to the functions given by the computer programming language. Using artificial intelligence algorithms, the output is the result of applying rules derived from the experience the computing system gains by processing existing datasets collected in the past. Machine learning, where experience from existing training datasets is translated into data processing rules applicable to future datasets, can be a component of AI algorithms.
[0083] Both methods utilize rules, which are expressed in different ways: in programmed algorithms, they are expressed by statements, mathematical formulas, etc., while in the context of AI, they are expressed by neural networks with appropriate weights on the connections between nodes of a given topology. Such rules are applied on the user device side. Instead, the corresponding data processing and artificial intelligence modules on the backend side are primarily responsible for determining the rules that should be applied on the user device side.
[0084] The user behavior module is an additional data processing / AI module that conceptually performs further aggregation, but is specifically designed to identify typical ways in which the device is used on a daily and weekly basis. User behavior indicators are determined, including the identification of device usage, how much the device is used (e.g., turned off, idle, charging, messaging, phone calls, internet browsing, reading emails, etc.), and when and on which days this occurs, typical lighting and background noise conditions, typical visited locations revealed using GNSS and other means (e.g., surrounding WiFi, SSID, Bluetooth devices, etc.), and typical keyboard and mouse usage (pressing or swiping keys with one hand or both hands and / or specific fingers for typing, etc.).
[0085] The dialogue behavior module 11 performs similar functions to the user behavior module 10, but its functions are particularly based on data collected before, during, and over a certain period after the dialogue takes place. Since the dialogue can be started at any time and requires the availability of data over a certain period before the dialogue starts, circular memory is used as a buffer to store the data required at the start of the dialogue. The purpose of this module is to determine user behavior during dialogues that involve customers in particular.
[0086] In the first scenario, all collected and stored data is sent to the backend as soon as a communication channel to the backend becomes available. This communication configuration is ideal for ensuring maximum data availability to the backend for interaction confirmation, even if the user device is destroyed by accident or intentional sabotage, or if data security is compromised. However, this configuration may not be usable due to legal (e.g., privacy) constraints.
[0087] In the second scenario, the data remains stored within the user's device, and only a minimal set of data is transmitted to the backend when the interaction in question occurs. This communication configuration is the most secure from a privacy standpoint, but the most vulnerable to device damage / sabotage.
[0088] A compromise between the two extreme measures described above for the first and second scenarios is implemented by this module and controlled along with all other configuration settings by the customer and end-user profile modules located in the backend.
[0089] This module also communicates locally (i.e., on the user's device) with the customer's application for interaction, i.e., the software running on the user's device. A unique interaction ID is assigned and shared between the customer's application and the interaction verification system. This module is also responsible for logging the user into the backend system using a single sign-on (SSO) procedure, i.e., a single login that works for both the customer's application and the background interaction verification function.
[0090] Examples of authentication verification based on fingerprint recognition, facial recognition, and speaker recognition, which utilize data from other sensors that are collected and processed, are as follows:
[0091] Authentication verification based on fingerprint recognition Several fingerprint spoofing techniques are known that allow for the creation of artificial fingerprints of a genuine person and their submission to a fingerprint recognition system in order to commit fraud. These fingerprint spoofing techniques have demonstrated that they can deceive fingerprint recognition systems by submitting artificial replicas of fingerprints made from various materials, such as silicon and gelatin, to an electronic capture device. These images are then processed as "real" fingerprints, thus causing a viable fraudulent impersonation.
[0092] As a result of the above, algorithms have been developed that aim to verify the authenticity of submitted fingerprints. For example, one proposed technique, named "activity detection," attempts to measure activity from the properties of the fingerprint image itself by applying an image processing algorithm. Other techniques have also been proposed. While these techniques help reduce the likelihood that fingerprint spoofing attempts, when present, could lead to successful fraudulent authentication, these types of algorithms are still not 100% reliable. Therefore, an unresolved possibility remains that fraudulent authentication may occur even when the most sophisticated authenticity verification algorithms are used.
[0093] A limitation of such authenticity verification algorithms is that they only consider the characteristics of the fingerprint image to determine the authenticity of a submitted fingerprint, without using other sensors that could help more accurately assess whether the fingerprint was submitted by the actual person whose fingerprint is being recognized or by someone else using a spoofed fingerprint. The term "single-sensor fingerprint recognition" may be used to describe systems and algorithms that perform fingerprint recognition based on data from a fingerprint sensor, possibly complemented by an authenticity verification algorithm.
[0094] A key characteristic of this example is that whenever the authenticity of fingerprint recognition becomes an issue, data from a single-sensor fingerprint recognition system is combined with data from other sensors.
[0095] As an example, data from a single-sensor fingerprint recognition system can be combined with data from inertial sensors, which may be called accelerometers and / or gyroscopes, found in most modern smartphones and tablets. The algorithm designed to combine data from both sensors is as follows:
[0096] Raw data (3-axis acceleration and / or 3-axis angular velocity) from the inertial sensor is continuously collected and stored in a circular memory large enough to store several seconds of raw inertial data.
[0097] Each time fingerprint recognition is performed, raw inertial sensor data recorded for several seconds before, during, and after the recognition is saved and stored in local persistent memory. The term "persistent memory" is used to identify a data storage area within the device that can retain available data even when the device is turned off and until the data is transmitted to the backend. Therefore, persistent memory is actually temporary storage, and is "persistent" in the sense defined above, namely in the sense that data is retained throughout the power off / on cycle.
[0098] When an activity detection algorithm or other single-sensor authenticity testing algorithm is applied on a smartphone / tablet, the results of such algorithms are also stored in local persistent memory.
[0099] Data that is formally tagged by a timestamp reference and stored in local persistent memory is sent to the backend system immediately or later.
[0100] Next, the movement of the smartphone / tablet that occurred for several seconds prior to, during, and after fingerprint recognition is reconstructed from the collected raw inertial data using any trajectory reconstruction algorithm available in the literature. This reconstruction is a form of data summarization and aggregation that can be performed locally on the smartphone / tablet (the reconstructed trajectory is then saved and sent to a backend system) or on the backend system (based on the raw inertial sensor data received from the smartphone / tablet).
[0101] To handle any issues that may arise with fingerprint recognition later, the backend system stores the above data in a database for each fingerprint recognition performed by the user.
[0102] Whenever fingerprint recognition is required, the backend system retrieves data from a database collected for each fingerprint recognition performed on the same user using the same smartphone / tablet. The collected smartphone trajectories are compared with each other and with the trajectories recorded for the recognition in question. Using any technique that allows for evaluation of whether such trajectories, considered to be motor functions, are similar (e.g., cross-correlation, pattern recognition), the similarity between the trajectory recorded for the recognition in question and all the trajectories recorded for other recognitions is calculated. Artificial intelligence algorithms for pattern recognition can also be used for this purpose.
[0103] The results of the algorithm described above are an assessment of how the smartphone / tablet was moved when fingerprint recognition occurred. When a legitimate user performs fingerprint recognition, they are likely to move the smartphone / tablet in a specific way (e.g., slightly rotate the smartphone to the left or right) to facilitate the presentation of their fingerprint to the fingerprint sensor. If a spoofed fingerprint is used, the movements performed are likely to be different, and therefore the abnormal movements may be recognized based on their lack of similarity to other (presumably non-spoofed) recognitions. To obtain a more accurate assessment of the presumed authenticity of the recognition, the results of activity detection or other single-sensor authenticity verification algorithms (running from the beginning on the smartphone or calculated / recalculated on a backend system) can also be combined with the results of the smartphone / tablet trajectory similarity calculation.
[0104] As a further example within the scope of the present invention, data from a single-sensor fingerprint recognition system can be combined with data provided by an RF interface within any smartphone / tablet / laptop. An example of an algorithm for combining data from both sensors is shown below.
[0105] Each time fingerprint recognition is performed, the RF interface of the smartphone / tablet / laptop is activated to collect data representing the current RF environment surrounding the device, such as the ID of the received GSM cell, the SSID of the received WiFi network, the Bluetooth address of any nearby Bluetooth devices in advertising mode, and, if any, the GNSS location, or the last known GNSS location. This "RF snapshot information" related to the current RF environment is stored in the device's local persistent memory.
[0106] If an activity detection algorithm or other single-sensor authenticity testing algorithm is applied on the device, the results of such algorithms are also stored in local persistent memory.
[0107] Data that is formally tagged by a timestamp reference and stored in local persistent memory is sent to the backend system immediately or later.
[0108] To handle any issues that may arise during fingerprint recognition, the backend system stores the above data in a database for each fingerprint recognition performed by the user.
[0109] Whenever fingerprint recognition is required, the backend system retrieves the data collected for each instance from a database, and fingerprint recognition is performed on the same user using the same device. The collected RF snapshot information is compared with the RF snapshots recorded for the recognition in question and for the recognition in question, and the similarity between the RF snapshots recorded for the recognition in question and all RF snapshots recorded for other recognitions is calculated using any technique that allows for evaluation of whether such RF snapshots are similar.
[0110] The results of the above algorithm are an assessment of the RF environment surrounding the device at the time of fingerprint recognition, for evaluating whether the RF environment surrounding the device is reliable compared to other RF environments typically experienced by the user and the device. To obtain a more accurate assessment of the presumed authenticity of the recognition, the results of activity detection or other single-sensor authenticity verification algorithms (running from the beginning on the smartphone or calculated / recalculated on the backend system) can also be combined with the calculation results of RF snapshot similarity.
[0111] As a further example, data from a single-sensor fingerprint recognition system can be combined with data provided by the (front and / or rear) cameras found in all modern smartphones and tablets. The algorithm envisioned for combining data from both sensors is as follows:
[0112] Raw image data from the camera is continuously collected and stored in a circular memory large enough to store several seconds of raw data.
[0113] Each time fingerprint recognition is performed, raw image data recorded for a few seconds before, during, and after the recognition is saved and stored in local persistent memory.
[0114] When an activity detection algorithm or other single-sensor authenticity testing algorithm is applied on a smartphone / tablet, the results of such algorithms are also stored in local persistent memory.
[0115] Data that is formally tagged by a timestamp reference and stored in local persistent memory is sent to the backend system immediately or later.
[0116] The collected raw image data is processed to reconstruct both the movement of the smartphone / tablet that occurred during, in the process of, and after fingerprint recognition (similar to the case of inertial sensors, using apparent movement on the image instead of inertial data), and to identify surrounding visual elements (objects, faces, background features, etc.). These reconstructions and identifications are in the form of data summarization and aggregation, which can be performed locally on the smartphone / tablet (the reconstructed trajectory and identified elements are then stored and sent to the backend system) or on the backend system (based on the raw image data received from the smartphone / tablet).
[0117] To handle any issues that may arise with certain fingerprint recognition, the backend system stores the above data in a database for each and all fingerprint recognition performed by the user.
[0118] Whenever fingerprint recognition is required, the backend system retrieves all data collected from the database for each and every fingerprint recognition performed on the same user using the same smartphone / tablet. The trajectory of all collected smartphones / tablets and all identified visual elements are compared with the data recorded for the recognition in question and for any similarity between the data and all other recognitions.
[0119] In this case as well, the results of the algorithm described above are an assessment of how the smartphone / tablet was moved and what visual elements were present when the fingerprint recognition was performed, compared to corresponding data collected during other (presumably unspoofed) recognition processes.
[0120] Verification of authentication based on facial recognition Similar to fingerprint recognition, several techniques are known for submitting artificial / approximate reconstructions of real people's faces to a facial recognition system, so that the face is recognized as if it were the real person's face. As with fingerprint recognition, algorithms have been developed to reduce the likelihood of facial recognition spoofing attempts leading to successful fraudulent authentication (including certain activity detection algorithms based on examining movements such as blinking or pupil dilation in response to changes in light intensity). However, these types of algorithms are still not 100% reliable, and the possibility of fraudulent authentication remains even when the most sophisticated authenticity verification algorithms are used.
[0121] As already mentioned regarding fingerprint recognition, the key feature of this example is that data from a single-sensor facial recognition system (i.e., a conventional system using one or more cameras) is combined with data from other sensors to improve the overall authenticity check performance whenever the authenticity of facial recognition becomes an issue.
[0122] As an example, data from a single-sensor facial recognition system can be combined with data from inertial sensors (accelerometers and / or gyroscopes) found in most modern smartphones and tablets. The algorithm designed to combine data from both sensors is as follows:
[0123] Raw data (3-axis acceleration and / or 3-axis angular velocity) from the inertial sensor is continuously collected and stored in a circular memory large enough to store several seconds of raw inertial data.
[0124] Each time face recognition is performed, raw inertial sensor data recorded for several seconds before, during, and after face recognition is saved and stored in local persistent memory.
[0125] When an activity detection algorithm or other single-sensor authenticity testing algorithm is applied on a smartphone / tablet, the results of such algorithms are also stored in local persistent memory.
[0126] Data that is formally tagged by a timestamp reference and stored in local persistent memory is sent to the backend system immediately or later.
[0127] Next, the movement of the smartphone / tablet that occurred for several seconds prior to, during, and after face recognition is reconstructed from the collected raw inertial data using any trajectory reconstruction algorithm available in the literature. This reconstruction is a form of data summarization and aggregation that can be performed locally on the smartphone / tablet (the reconstructed trajectory is then saved and sent to a backend system) or on the backend system (based on the raw inertial sensor data received from the smartphone / tablet).
[0128] To handle any issues that may arise during facial recognition, the backend system stores the above data in a database for each facial recognition performed by the user.
[0129] Whenever face recognition is required, the backend system retrieves all data collected from a database for each and every face recognition performed on the same user using the same smartphone / tablet. The trajectories of all collected smartphones are compared with those recorded for the recognition in question, and any technique (e.g., cross-correlation, pattern recognition) that allows for evaluation of whether such trajectories, considered to be motor functions, are similar is used to calculate the similarity between the trajectory recorded for the recognition in question and all trajectories recorded for other recognitions. Artificial intelligence algorithms for pattern recognition can also be used for this purpose.
[0130] The results of the algorithm described above are an evaluation of how the smartphone / tablet was moved when face recognition occurred. When a legitimate user performs face recognition, they are likely to move the smartphone / tablet in a specific way (e.g., slightly rotate the smartphone to the left or right) to facilitate the presentation of their face to the camera. If a spoofed face image is used, the movements performed are likely to be different, and therefore the abnormal movements may be recognized based on a lack of similarity to other (presumably unspoofed) recognitions. To obtain a more accurate assessment of the presumed authenticity of the recognition, the results of other single-sensor authenticity verification algorithms (either run initially on the smartphone or calculated / recalculated on a backend system) can also be combined with the results of the smartphone / tablet trajectory similarity calculation.
[0131] As a further example, data from a single-sensor facial recognition system can be combined with data provided by the RF interface in any smartphone / tablet / laptop. The algorithm intended to combine data from both sensors is as follows:
[0132] Each time face recognition is performed, the RF interface of the smartphone / tablet / laptop is activated to collect data representing the current RF environment surrounding the device, such as the ID of the received GSM cell, the SSID of the received WiFi network, the Bluetooth address of any nearby Bluetooth devices in advertising mode, and, if any, the GNSS location, or the last known GNSS location. This "RF snapshot information" related to the current RF environment is stored in the device's local persistent memory.
[0133] If an activity detection algorithm or other single-sensor authenticity testing algorithm is applied on the device, the results of such algorithms are also stored in local persistent memory.
[0134] Data that is formally tagged by a timestamp reference and stored in local persistent memory is sent to the backend system immediately or later.
[0135] To handle any issues that may arise during facial recognition, the backend system stores the above data in a database for each facial recognition performed by the user.
[0136] Whenever face recognition is required, the backend system retrieves all data collected for each and all face recognitions performed on the same user using the same device from a database. All collected RF snapshot information is compared with the RF snapshots recorded for the recognition in question, and the similarity between the RF snapshots recorded for the recognition in question and the RF snapshots recorded for the other recognitions is calculated using any technique that allows for evaluation of whether such RF snapshots are similar.
[0137] The results of the above algorithm are an assessment of the RF environment surrounding the device when face recognition is performed, to evaluate whether the RF environment surrounding the device is reliable compared to other RF environments typically experienced by the user and the device. To obtain a more accurate assessment of the estimated authenticity of the recognition, the results of activity detection or other single-sensor authenticity verification algorithms (running from the beginning on the smartphone or calculated / recalculated on the backend system) can also be combined with the calculation results of RF snapshot similarity.
[0138] Verification of authentication based on speaker recognition In speaker recognition, several techniques are known to enable the recognition of voices (imitated, synthesized, or recorded) as if they were the voice of a specific person. Consequently, algorithms have been developed to reduce the likelihood of speaker recognition spoofing attempts leading to successful fraudulent authentication (such as activity detection algorithms that require answering various questions). However, these types of algorithms are still not 100% reliable, and the unresolved possibility of fraudulent authentication remains, even when the most sophisticated authenticity verification algorithms are used.
[0139] As already mentioned with respect to fingerprint and facial recognition, the key feature of this example is that data from the speaker recognition system (i.e., a conventional system using one or more microphones) is combined with data from other sensors to improve the overall authenticity check performance whenever the authenticity of speaker recognition becomes an issue.
[0140] For example, data from a speaker recognition system can be combined with data from inertial sensors (accelerometers and / or gyroscopes) found in most modern smartphones and tablets. The algorithm designed to combine data from both sensors is as follows:
[0141] Raw data (3-axis acceleration and / or 3-axis angular velocity) from the inertial sensor is continuously collected and stored in a circular memory large enough to store several seconds of raw inertial data.
[0142] Each time speaker recognition is performed, raw inertial sensor data recorded for several seconds before, during, and after speaker recognition is saved and stored in local persistent memory.
[0143] When an activity detection algorithm or other single-sensor authenticity testing algorithm is applied on a smartphone / tablet, the results of such algorithms are also stored in local persistent memory.
[0144] Data that is formally tagged by a timestamp reference and stored in local persistent memory is sent to the backend system immediately or later.
[0145] Next, the movement of the smartphone / tablet during, and for several seconds after, prior to speaker recognition, is reconstructed from the collected raw inertial data using any trajectory reconstruction algorithm available in the literature. This reconstruction is a form of data summarization and aggregation that can be performed locally on the smartphone / tablet (the reconstructed trajectory is then saved and sent to a backend system) or on the backend system (based on the raw inertial sensor data received from the smartphone / tablet).
[0146] To handle any issues that arise with speaker recognition later, the backend system stores the above data in a database for each and all speaker recognition performed by the user.
[0147] Whenever speaker recognition is required, the backend system retrieves all data collected from a database for each and every speaker recognition performed on the same user using the same smartphone / tablet. The trajectories of all collected smartphones are compared with those recorded for the recognition in question, and any technique (e.g., cross-correlation, pattern recognition) that allows for evaluation of whether such trajectories, considered as motor functions, are similar is used to calculate the similarity between the trajectory recorded for the recognition in question and all trajectories recorded for other recognitions. Artificial intelligence algorithms for pattern recognition can also be used for this purpose.
[0148] The results of the algorithm described above are an evaluation of how the smartphone / tablet was moved when speaker recognition was performed. When a legitimate user performs speaker recognition, they are likely to move the smartphone / tablet in a specific way (e.g., slightly rotate the smartphone left or right, or up or down) to facilitate the presentation of their voice to the microphone. If a spoofed voice is used, the movements performed are likely to be different, and therefore the abnormal movements may be recognized based on their lack of similarity to other (presumably non-spoofed) recognitions. To obtain a more accurate assessment of the presumed authenticity of the recognition, the results of other authenticity checking algorithms (either run initially on the smartphone or calculated / recalculated on a backend system) can also be combined with the results of the smartphone / tablet trajectory similarity calculation.
[0149] As a further example, data from a speaker recognition system can be combined with data provided by the RF interface in any smartphone / tablet / laptop. The algorithm intended to combine data from both sensors is as follows:
[0150] Each time speaker recognition occurs, the RF interface of the smartphone / tablet / laptop is activated to collect data representing the current RF environment surrounding the device, such as the ID of the received GSM cell, the SSID of the received WiFi network, the Bluetooth address of any nearby Bluetooth devices in advertising mode, the GNSS location if any, or the last known GNSS location if any. This "RF snapshot information" related to the current RF environment is stored in the device's local persistent memory.
[0151] If an activity detection algorithm or other single-sensor authenticity testing algorithm is applied on the device, the results of such algorithms are also stored in local persistent memory.
[0152] Data that is formally tagged by a timestamp reference and stored in local persistent memory is sent to the backend system immediately or later.
[0153] To handle any issues that arise with speaker recognition later, the backend system stores the above data in a database for each and all speaker recognition performed by the user.
[0154] Whenever speaker recognition is an issue, the backend system retrieves all data collected for each and all speaker recognitions performed on the same user using the same device from a database. All collected RF snapshot information is compared with each other and with RF snapshots recorded for the recognition in question, and the similarity between the RF snapshot recorded for the recognition in question and all RF snapshots recorded for other recognitions is calculated using any technique that allows for evaluation of whether such RF snapshots are similar.
[0155] The results of the above algorithm are an assessment of the RF environment surrounding the device at the time of speaker recognition, for evaluating whether the RF environment surrounding the device is reliable compared to other RF environments typically experienced by the user and the device. To obtain a more accurate assessment of the presumed authenticity of the recognition, the results of activity detection or other authenticity check algorithms (running from the beginning on the smartphone or calculated / recalculated on the backend system) can also be combined with the calculation results of RF snapshot similarity.
[0156] As a further example within the scope of the present invention, data generated from a speaker recognition system can be combined with data provided by (front and / or rear) cameras found in all modern smartphones and tablets. The algorithm envisioned for combining data from both sensors is as follows:
[0157] Raw image data from the camera is continuously collected and stored in a circular memory large enough to store several seconds of raw data. Each time speaker recognition is performed, raw image data recorded for several seconds before, during, and after speaker recognition is saved and stored in local persistent memory.
[0158] When an activity detection algorithm or other single-sensor authenticity testing algorithm is applied on a smartphone / tablet, the results of such algorithms are also stored in local persistent memory.
[0159] Data that is formally tagged by a timestamp reference and stored in local persistent memory is sent to the backend system immediately or later.
[0160] The collected raw image data is processed to reconstruct both the movement of the smartphone / tablet that occurred during, in the process of, and after speaker recognition (similar to the case of inertial sensors, using apparent motion on the image instead of inertial data), and to identify surrounding visual elements (objects, faces, background features, etc.). These reconstructions and identifications are in the form of data summarization and aggregation, which can be performed locally on the smartphone / tablet (the reconstructed trajectory and identified elements are then stored and sent to the backend system) or on the backend system (based on the raw image data received from the smartphone / tablet).
[0161] To handle any issues that arise with speaker recognition later, the backend system stores the above data in a database for each and all speaker recognition performed by the user.
[0162] Whenever speaker recognition is required, the backend system retrieves all data collected from the database for each and all speaker recognitions performed on the same user using the same smartphone / tablet. The trajectory of all collected smartphones / tablets and all identified visual elements are compared with the data recorded for the recognition in question and for any similarity between the data and all other recognitions.
[0163] In this case as well, the results of the algorithm described above are an evaluation of how the smartphone / tablet was moved and what visual elements were present when speaker recognition was performed, compared to corresponding data collected during other (presumably unspoofed) recognition processes.
[0164] Collection and transmission of data for authentication verification. The process of collecting data from a user device (smartphone, tablet, computer, etc.) and sending that data to a backend system for authentication can be invoked within a mobile application and executed using a so-called Software Development Kit (SDK) that is responsible for timely data collection and transmission to the backend. Before being sent to the backend in batches at the appropriate time by the SDK's software function, commonly known as dispatching, the collected data is first stored locally in the mobile device's memory. In one example, dispatching may occur every 30 minutes, or in another, every 2 minutes. The frequency can be customized depending on the specific application.
[0165] In the case where the user device is a personal computer, a snippet of JavaScript code can collect and send data to a backend system, and the data sent is critical for the intended authentication verification purpose (e.g., inertial sensor data, RF sensor data, camera data, etc.). In the case where the user device is a mobile phone, a specific SDK may be used and developed for this particular application that collects and sends the relevant types of data. One possible way to prevent authentication verification from being performed in order to possibly turn off or possibly damage / destroy the mobile device before the data constituting fraudulent authentication is sent to the backend is to send the sendout to verify the authentication in question fairly frequently, for example, every two minutes or more. However, the inability to obtain the data constituting fraudulent authentication in question due to the device being turned off or destroyed may itself be evidence that fraudulent authentication has occurred.
[0166] Example of signal flow between modules Figure 11 shows a collaboration diagram related to the activation of new users. A user or end-user is typically the person who will interact with the user device. This person may be a customer of, for example, a bank or credit card organization or another organization, and may use interaction verification methods to validate the interaction in question.
[0167] Customers are typically banks, credit card organizations, or other organizations (e.g., service providers that offer dialogue verification services to other organizations) that use dialogue verification methods to possibly verify the dialogue in question.
[0168] As shown in Figure 11, when a new end-user (e.g., a new bank account or credit card holder) is activated by the customer, a first communication occurs between the customer's IT system 42 and the customer communication module 23. The customer's IT system informs the interaction confirmation system via the communication interface that a new end-user must be added, and all relevant information (data collection by applications, data transmission, privacy agreement, configuration settings for parameters for SSO) is provided to the customer communication module. The customer communication module communicates with the customer and end-user profile module 16 so that a user profile is created for the target user and permanently stored. The system then confirms to the caller that the operation is complete.
[0169] Next, the user downloads the customer's application (or equivalent customer application software that runs on the user's device). The user device module for interaction verification is integrated as an SDK within the customer application software. The end user logs into the customer's application, and SSO identifies and logs the end user for interaction verification purposes as well. When this step occurs, further user initialization functions are performed as shown in Figure 12.
[0170] Upon initial login, a communication channel is established between the customer and end-user profile module 16 in the backend and the storage and data protection module on the user device, with the involvement of the user device communication module 15 and the backend communication module 12. This program the user device to collect and transmit data according to defined rules (including data processing rules defined by the data processing and artificial intelligence module on the backend). This involves a handshake between the two customer and end-user profile modules (one in the backend and one on the user device) (indicated by the larger dashed arrow in the diagram above), so that information specific to the user device (e.g., which sensors are on the device and which are not, what are the characteristics of the sensors, etc.) is added to the user profile, and the most appropriate data processing rules are selected. Then, on the user device side, the customer and end-user profile module 16 instructs the data processing and artificial intelligence (AI) module 6 (on the user device side) regarding the data processing rules to be applied. If anything changes over time regarding the end-user profile, including data processing rules (for example, some improvements to the data processing rules may be introduced based on the data being collected), all changes are propagated from the backend to the end-user device or vice versa by the same handshake mechanism.
[0171] Once the user device is fully initialized, all modules begin collecting, processing, and potentially sending data to the backend, according to a defined user profile that includes the data processing rules required by their respective functions. Each interaction is handled on demand, and a unique interaction ID is assigned to the interaction so that it can be tracked later.
[0172] Figure 13 shows a collaboration diagram related to a problematic dialogue. When a problematic dialogue occurs, the customer's IT system 25 submits a request to the customer communication module 23 via the communication interface to verify a certain dialogue ID. The customer communication module 23 activates the dialogue verification module 22, which activates the data processing and artificial intelligence (AI) module (backend side) 17. The data processing and AI module 17 then retrieves the required data from the storage and data protection modules 24 and 13 (data already transmitted to the backend is from the backend, and data not yet transmitted to the backend is from the user device). Since the user device may be off or not connected, data retrieval from the user device does not need to be immediate; therefore, requests for data transmitted by the user device are queued to be fulfilled as soon as a connection to the end-user device can be established. Once the data is available and the dialogue verification module is ready to respond, the result is transmitted to the customer's IT system by the customer communication module.
[0173] This collaboration diagram does not include cases where a single backend is shared among multiple customers, such as when a dialogue validation service is provided to multiple customers (various banks, credit card organizations, online payment providers, etc.) by an independent entity (i.e., a dialogue validation service provider - TVSP). Sharing many end users from multiple customers can be valuable because it provides a larger dataset for testing and fine-tuning data processing systems, and in the case of artificial intelligence systems, it provides a larger dataset for training and testing AI algorithms.
[0174] Backend system configuration Figures 14 and 15 show two backend system configuration examples (dedicated backend and shared backend).
[0175] In the case of a dedicated backend, especially when it is installed and physically integrated with the customer's IT system, the backend itself 26 can be logically considered as part of the customer's IT system 25, as shown in Figure 14.
[0176] In the case of a shared backend, regardless of whether they are physically jointly installed or even share the same cloud server, the logical distinction between the backend 27 and the IT systems of various customers is important in the example in Figure 15. In this example, the customer communication module is logically and physically connected to the IT systems 28, 29, and 30 of multiple customers, and is ready to receive interaction verification requests from each of them. The customer communication module provides relevant responses that maintain the necessary logical distinctions between requests arising from the various customers.
[0177] Figure 16 is a flowchart illustrating how data in the user device is processed in steps S16.1, S16.2, S16.3, and S16.4 to generate user confirmation data for use within the dialogue confirmation system.
[0178] In the examples already described, the interaction can be a transaction, such as a financial transaction, and the interaction confirmation system may be called a transaction confirmation system. The examples may relate to a system for confirming transactions, a method for processing data in a user device to generate user confirmation data for use within a transaction confirmation system, and a system for confirming transactions after they have been made. The recognition of the user by the computer may be an inspection of identification information. One example provides a method for processing data in a user device to generate user confirmation data for use within a transaction confirmation system, which includes deriving first user behavior data from a first set of data, each generated by several different elements of the user device and each representing that a user is interacting with the user device; identifying at least a first time interval relating to a transaction involving the user of the user device; deriving second user behavior data from a second set of data, each generated by several different elements of the device and each representing that a user is interacting with the user device during at least a first time interval; and transmitting user confirmation data, including the first and second user behavior data, from the device to the transaction confirmation system. This method allows the verification system to process first user behavior data and second behavior data, for example, to investigate a problematic transaction and determine whether the transaction involved interaction with a given user's device.
[0179] Any feature described in any one example may be used alone or in combination with other features described, in combination with one or more features of any other example, or in any combination of any other example. Furthermore, equivalents and modifications not described above may be used without departing from the scope of the invention as defined in the attached claims.
Claims
1. A computer-implemented method for enabling a computer to recognize a user interacting with a user device within an identified time interval, by processing data in the user device and generating user confirmation data for use in a dialogue confirmation system to verify the dialogue in question after the dialogue has taken place, The method involves deriving first user behavior data by processing a first set of data, each of which is generated by a plurality of different elements of the user device, including at least one sensor, and represents that the user is interacting with the user device. Identifying at least a first time interval related to the interaction between the user of the user device and the user device, The method involves deriving second user behavior data by processing a second set of data, each of which is generated by the multiple different elements of the device, including at least one sensor, and represents that the user is interacting with the user device during at least the first time interval. The second user behavior data is stored in the storage system on the user device, Receiving timing data indicating the first time interval from the dialogue confirmation system, Based on the timing data, the second user behavior data is obtained from the storage system, The user device transmits the first user behavior data and user confirmation data based on the second user behavior data to the dialogue confirmation system. Includes, This method enables the identification and acquisition of data related to a problematic dialogue for use when processing by the dialogue confirmation system.
2. The method according to claim 1, comprising identifying the first time interval as the time interval at which the dialogue occurs.
3. Identifying a second time interval as the time interval before the aforementioned dialogue occurs, and / or Identifying a third time interval as the time interval after the aforementioned dialogue occurred. The second set of data includes, respectively, that the user interacts with the device during the first time interval and the second and / or third time intervals. The method according to claim 1.
4. The method according to claim 1, comprising collecting second user behavior data in response to receiving an instruction that a dialogue is in progress.
5. The method according to claim 1, wherein deriving the first and second user behavior data includes using a hardware abstraction function module configured to convert data generated by the plurality of different elements of the user device into converted element data having a normalized format.
6. The method according to claim 5, wherein deriving the first and second user behavior data includes using a data processing function module configured to perform summarization, aggregation, and joining functions on the transformed element data in order to generate processed element data.
7. The method according to claim 6, wherein deriving the first user behavior data includes using a user behavior function module configured to extract information about typical user behavior from processed element data relating to the first set of data.
8. The method according to claim 6, wherein deriving the second user behavior data includes using a behavior function module configured to extract information about user behavior from processed element data relating to the second set of data.
9. The method according to claim 1, wherein the user verification data includes output from a machine learning model.
10. The method according to claim 9, wherein parameters for the machine learning model are received from a validation system.
11. The method according to claim 9, wherein the input to the machine learning model includes the first user behavior data and the second user behavior data, and the user verification data includes the output of the machine learning model.
12. The method according to claim 11, wherein the output of the machine learning model includes the probability that a user in the first time interval is different from a user corresponding to the first user behavior data.
13. The method according to claim 12, wherein the machine learning model is a deep neural network (DNN).
14. The method according to claim 13, wherein the deep neural network is trained to detect anomaly time intervals within a series of time intervals.
15. The method according to claim 9, wherein the input to the machine learning model includes at least the first set of data and the second set of data, and the output of the machine learning model includes the first user behavior data and the second user behavior data.
16. The method according to claim 15, wherein the machine learning model is trained by using unsupervised learning to sort the dialogues in the trial data into clusters.
17. The method according to claim 16, wherein the machine learning model processes individual time intervals to estimate which cluster each time interval belongs to.
18. The method according to claim 17, wherein the user verification data includes estimating which cluster the time interval belongs to.
19. A user device comprising one or more processors configured to perform the method described in any one of claims 1 to 18.
20. A computer program including instructions that, when the program is executed on a computer, cause the computer to perform a step of the method according to any one of claims 1 to 18.
21. A non-temporary computer-readable storage medium for holding instructions for causing one or more processors to perform steps of the method according to any one of claims 1 to 18.
22. A system for confirming a dialogue after it has taken place, comprising a user device configured to perform the method according to any one of claims 1 to 18 and a dialogue confirmation system.
23. The system according to claim 22, wherein the dialogue confirmation system is configured to process the user confirmation data in order to confirm a given dialogue.
24. The system according to claim 22, wherein the dialogue confirmation system includes customer and end-user profile modules configured to store the user confirmation data.
25. The system according to claim 22, wherein the dialogue confirmation system includes a verification module configured to process the user confirmation data to give an estimate of the probability that a given dialogue includes a given user.
26. The system according to claim 22, wherein the dialogue confirmation system includes a data processing module configured to determine data processing rules to be applied by the user device and to transmit data indicating the data processing rules to the user device.
27. The system according to claim 22, wherein the dialogue confirmation system includes a machine learning model for use in determining parameters for use in a corresponding machine learning model for a user device.
Citation Information
Patent Citations
Methods of enrolling and authenticating user in authentication system, facial authentication system, and methods of authenticating user in authentication system
JP2016051482A
Information processing apparatus
JP2017142614A
Information processing system, information processing method, information processing program, and information processing apparatus
JP2019101566A
Methods and processes for utilizing information collected for enhanced verification
US20200274860A1