Method for identifying a person, and system for identifying a person

DYNAMIDE addresses the challenges of legacy authentication technologies by using spatio-temporal trajectory analysis of anatomical landmarks to uniquely identify individuals based on their action performance, achieving efficient and robust real-time authentication.

JP7695730B2Active Publication Date: 2025-06-19RAMOT AT TEL AVIV UNIVERSITY LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024029663
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-20
Filing Date
2024-02-29
Publication Date
2025-06-19
Estimated Expiration
2041-07-30

AI Technical Summary

Technical Problem

Legacy authentication technologies face challenges in facilitating use and providing authentication quality as required by the increasing complexity and interdependence of modern citizens' actions, particularly in meeting strong customer authentication (SCA) specifications of the revised European Payment Services Directive (PSD2).

Method used

The Dynamic Identification (DYNAMIDE) method, which identifies individuals based on the uniqueness of how they perform actions, by identifying anatomical landmarks and processing their spatio-temporal trajectories to determine features that distinguish individual actions and identify the specific individual performing them.

Benefits of technology

DYNAMIDE provides a robust and efficient biometric authentication method that can identify individuals in real-time, even in complex scenarios, without degrading the quality of experience, thus addressing the limitations of legacy authentication technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695730000001
    Figure 0007695730000001
  • Figure 0007695730000002
    Figure 0007695730000002
  • Figure 0007695730000003
    Figure 0007695730000003
Patent Text Reader

Abstract

To provide a method and a device for identifying a person based on idiosyncrasies of a manner in which the person performs an activity.SOLUTION: A method for identifying a person comprises the steps of: acquiring spatiotemporal data for each of a plurality of anatomical landmarks associated with an activity engaged in by a person that defines a spatiotemporal trajectory of the anatomical landmark during the activity; modeling the acquired spatiotemporal data as a spatiotemporal graph (ST-Graph); and processing the ST-Graph using at least one non-local graph convolution neural network (NLGCN) to provide an identity for the person.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] 〔Related Applications〕 This application claims the benefit of U.S. Provisional Application No. 63 / 067,890, filed Aug. 20, 2020, under 35 U.S.C. 119(e), the disclosure of which is incorporated herein by reference.

[0002] 〔Technical Field〕 Embodiments of the present disclosure relate to methods and apparatuses for providing biometric authentication of a person's identity.

Background Art

[0003] The ever-growing list of services requires an authentication procedure, customarily called a multi-factor authentication procedure (MFA), to authenticate and authorize user access to the services. In an MFA procedure, a user is required to provide an appropriate response to each of a plurality of categories of challenges. The challenge categories are called "authentication factors". General MFA is called two-factor authentication (2FA), and a user is challenged to correctly respond to at least two of three authentication factors: a knowledge factor, a possession factor, and an inherence factor. The knowledge factor tests something the user should know, such as a password. The possession factor requires the presentation of something the user is expected to have, such as a credit card or a smartphone. The inherence factor requires the presentation of a biometric characteristic that characterizes the user, such as a fingerprint, a voiceprint, or an iris scan.

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, legacy authentication technologies appear to face difficulties in facilitating use and providing authentication quality as required by the increasing complexity and interdependence of the matrix of actions regularly performed by modern citizens. For example, legacy MFA configurations appear to be severely pressured to meet the strong customer authentication (SCA) specifications of the revised European Payment Services Directive (PSD2) promulgated to integrate consumers, banks, and third-party providers (TPPs) in the open banking initiative. The implementation of SCA is two years behind schedule. The system, originally scheduled to start in September 2019, was postponed until March 14, 2021, and then postponed until the current deadline of September 14, 2021.

Means for Solving the Problems

[0005] One aspect of an embodiment of the present disclosure relates to providing a method. This method may be a dynamic identification (DYNAMIDE) method, or simply referred to as DYNAMIDE. This method identifies a person based on the uniqueness of the way the person performs an action. According to one embodiment of the present disclosure, DYNAMIDE includes identifying anatomical landmarks during an action performed by people and identifying the spatio-temporal trajectory of the anatomical landmarks while the action is being performed. The anatomical landmarks are optionally referred to as activity fiducials (AFID) and indicate various degrees of movement or the lack thereof. The spatio-temporal trajectory can be used to identify the action. DYNAMIDE includes processing the trajectory to determine features of the trajectory that are advantageous for distinguishing the actions performed by a specific individual and for identifying the specific individual performing the action.

[0006] The characteristics of an action that can distinguish the action by the individual performing the action can be very subtle. The AFID trajectories associated with the action can exhibit substantially slight and nonintuitive crosstalk. As a result, a characteristic of one spatio-temporal trajectory of an action that can intuitively appear independently of the characteristics of another spatio-temporal trajectory of the action may actually be unique to the individual performing the action and can provide a criterion for identifying the individual. According to one embodiment of the present disclosure, the provision of spatio-temporal determination that is advantageous for discovering and using the uniqueness indicated by the trajectory for the identification of the trajectory and for spatial and / or temporal processing may be non-local and a number of a priori processing constraints. The a priori processing constraints are assumed for the movement indicated by the advantageously limited AFID trajectories.

[0007] According to one embodiment, identifying a particular individual based on a given action that an individual can perform includes obtaining a series of images of the individual performing the given action and identifying an AFID associated with the given action within the image. The image can be processed to determine a spatio-temporal trajectory indicated by the identified AFID and a trajectory processed to identify the particular individual who performed the action from among a plurality of individuals who can perform the action. Optionally, processing the AFID trajectory includes determining local and non-local spatio-temporal correlations indicated by the AFID during the execution of the given action and using the correlations to determine the identity of a particular individual. The spatio-temporal correlations can include correlations based on spatial parameters, temporal parameters, or both spatial and temporal parameters that characterize the spatio-temporal trajectory or trajectories in one or more AFIDs.

[0008] In embodiments of the present disclosure, an AFID associated with a given action can be an anatomical landmark of any body part such as a limb, face, or head that exhibits a spatiotemporal trajectory in the execution of the given action suitable for use in identifying the person performing the action. For example, the AFID can be a joint of a limb or a bone of the skeleton that exhibits a spatiotemporal trajectory that conforms during an action such as walking, hitting a golf ball, or typing a password at an ATM. For typing, the AFIDS can include a plurality of joints to which the bones of the hand are connected. The AFID can be a facial landmark such as the eyebrows, eyes, and corners of the lips whose movement is used to define an action unit (AU) of the Facial Action Coding System (FACS). The Facial Action Coding System is used to classify expressions and micro-expressions. The AFID can also be a detailed pair of features of the fingerprints of a plurality of fingers of the hand, imaged non-contact with sufficient spatial resolution to enable identification of the pair of features.

[0009] According to one embodiment, DYNAMIDE uses at least one neural network for processing an image of an action to identify the individual performing the action. In one embodiment, the at least one neural network is trained to detect body parts or regions of interest (BROI) within the image and identify the AFIDs they may contain. The spatial and temporal progression of the identified AFIDs during the execution of the action is represented by a spatiotemporal graph (ST-Graph). In the spatiotemporal graph, the AFID is a node connected by spatial and temporal edges that define the spatiotemporal AFID trajectory of the action. The at least one neural network can comprise at least one graph convolutional network (GCN) for processing the trajectory and classifying the action according to the individual performing the action.

[0010] In one embodiment, at least one GCN comprises a nonlocal neural network (NLGCN) having at least one nonlocal neural network block for processing AFID spatio-temporal trajectories. The at least one nonlocal neural network block may comprise at least one spatial nonlocal neural network block and / or at least one temporal nonlocal neural network block. Optionally, the NLGCN is configured as a multi-stream GCN comprising a plurality of component NLGCNs operative to process a set of data characterized by independent degrees of freedom based on AFID trajectories. In one embodiment, the output of the multi-stream GCN may comprise a weighted average of the outputs of each of the component GCNS.

[0011] As an example, in DYNAMIDE configured to identify an individual by the way the individual types, the AFIDs that are hand joints are characterized by degrees of freedom of movement (e.g., distances between different finger joints) independent of the degrees of freedom of movement obtained for the AFIDs that are the hand bones connecting the joints. Thus, in one embodiment, DYNAMIDE may include a two-stream 2s-NLGCN multi-stream GCN having two component NLGCNs. One of the two component NLGCNs processes joint AFIDs and the other component NLGCN processes bone AFIDs. In one embodiment, the joint NLGCN comprises at least one learnable “adaptive” adjacency matrix that is driven by substantial data to reduce the number of a priori constraints that can be used to configure the 2s-NLGCN. The 2s-NLGCN with an adaptive adjacency matrix according to one embodiment of the present disclosure may be referred to as an adaptive 2s-NLGCN (2s-ANLGCN). The outputs of the joint and bone NLGCNs in the typing DYNAMIDE, in the 2s-NLGCN or 2s-ANLGCN, may be fused to identify the individual.

[0012] According to one embodiment, identifying a specific individual is performed in real time. The real-time identification according to one embodiment refers to the identification of an individual while the individual is performing an action, or the identification of an individual within a time frame in which the quality of experience (QoE) of the service for which the identification is performed is not substantially degraded by the identification process.

[0013] This summary is provided to introduce, in a simplified form, a selection of concepts that are further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

Brief Description of the Drawings

[0014] Non-limiting examples of embodiments of the present invention are described below with reference to the drawings attached hereto, which are listed after this paragraph. Identical features that appear in two or more figures are generally labeled with the same label in all figures in which the feature appears. In the figures, labels that label icons representing a given feature of an embodiment of the present invention may be used to refer to the given feature. The dimensions of the features shown in the figures are selected for convenience of presentation and clarity and are not necessarily shown to scale.

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 6A

Figure 6B

Mode for Carrying Out the Invention

[0015] In the discussion, unless otherwise specified, adjectives such as "substantially" and "about" that modify the relationship of one or more characteristic points of the embodiments of the present disclosure are understood to mean that the state or characteristic is defined within an acceptable tolerance range for the steps of the desired embodiments in the specification. Whenever a general term in the present disclosure is explained by reference to an example or a list of examples, the example(s) mentioned are for the purpose of non-limiting illustration of the general term. Also, the general term is not intended to be limited to the specific example(s) mentioned. The phrase "in an embodiment" is used to introduce illustrative study materials, regardless of whether it is related to permissiveness such as "possible", "optionally", or "for the purpose of illustration". However, this phrase does not necessarily introduce the configurations required in the possible embodiments of the present disclosure. Unless otherwise expressly stated, the term "or" in the specification and claims is considered to be inclusive rather than exclusive, indicating at least one of the multiple items to be combined, or any combination thereof.

[0016] FIG. 1 shows a high-level flow diagram 20 of a process according to an embodiment of the present disclosure, optionally also referred to by the number 20, based on which DYNAMIDE can operate to identify a person in response to an action performed by the person.

[0017] In block 22, DYNAMIDE according to an embodiment of the present disclosure optionally obtains a series of video frames of a person involved in an action. DYNAMIDE is configured to perform processing to determine the identity of the person involved in the action. In block 24, DYNAMIDE processes the video frames to identify an image of a body region of interest (BROI) in the video frame that visualizes at least one AFID related to the action. Identifying the BROI in the video frame optionally includes determining at least one bounding box in the frame that includes the image of the BROI. In block 26, DYNAMIDE processes each of the bounding boxes determined for the video frame to identify an image of at least one AFID in each of the bounding boxes. Identifying the image of the AFID within the bounding box of the video frame optionally includes associating with the image a spatiotemporal ID (ST-ID), "AFID ST-ID", which includes an identification label of the AFID. Here, the "AFID ST-ID" labels all the identified images of the same AFID within the video frame and is used to determine the spatiotemporal coordinates of the image. The spatiotemporal coordinates include a time stamp and at least two spatial coordinates. The time stamp identifies the time position at which the video frame including the bounding box in which the AFID is located was acquired relative to the time at which other video frames in the series of video frames were acquired. The at least two spatial coordinates correspond to the spatial position of the AFID at the time indicated by the time stamp. Optionally, the AFID ST-ID for a given identified AFID includes a standard deviation (sd) for each spatial coordinate and the probability that the AFID-ID label associated with the AFID ST-ID is correct. The oldest and latest time stamps and the extreme spatial coordinates determined for the AFID ST-ID determine a spatiotemporal volume.The spatio-temporal volume may be referred to as a spatio-temporal AFID hull (ST-Hull) that includes the spatio-temporal coordinates of all instances of the AFID imaged and identified in a series of video frames.

[0018] In block 28, DYNAMIDE uses the ST-IDs of the AFIDs to construct the identified instances of the AFID as nodes of an AFID spatio-temporal graph (ST-graph) connected by spatial and temporal edges. The spatial edges connect the ST-graph nodes. The ST-graph nodes represent the imaged instances of the AFID identified by the same timestamp, i.e., the instances of the AFID imaged within the same video frame, and the spatial constraints imposed on the AFID by the structure of the human body. The configuration of the nodes connected by the spatial edges representing the spatial relationships of the instances of the AFID imaged at the same given frame and given time may be referred to as the spatial graph (S-graph) of the AFID at a given time. The temporal edges connect the temporally adjacent nodes in the ST-graph representing the images of the same AFID in two consecutively acquired video frames in a series of video frames. The temporal edges represent the elapsed time between two consecutive timestamps. The ST-graph may be regarded as including the S-graph corresponding to the AFIDs connected by the temporal edges.

[0019] In one embodiment, in block 30, DYNAMIDE processes the AFID ST-graph using an adaptively adaptable adaptive non-local graph convolutional neural network, ANLGCN. Thereby, DYNAMIDE optionally determines in real time which of a plurality of persons trained to be recognized by ANLGCN is involved in or attempting to be involved in the action. In one embodiment, ANLGCN is configured to span the AFID ST-hull and be weighted by weights learned from data associated with the imaged instances of AFID at any spatio-temporal position within the hull. Also, ANLGCN is configured to contribute to the convolution by ANLGCN, which is performed on spatio-temporal positions at any other location within the hull. Optionally, NLGCN is configured as a multi-stream GCN comprising a plurality of component NLGCNs that operate to process a set of AFID data characterized by independent degrees of freedom.

[0020] FIG. 2 schematically shows a DYNAMIDE system 100 according to an embodiment of the present disclosure. The DYNAMIDE system 100 is configured to execute the process shown in FIG. 1 and identify the person involved in the action based on the way the person performs the action. The DYNAMIDE system 100 may optionally include a cloud-based processing hub 120 and an imaging system 110 having a field of view (FOV) indicated by the dashed line 111. As an example, in this figure, the action is the action of typing on the keypad 62 in which the person 50 is involved at the ATM 60.

[0021] Imaging system 110 is operable to provide an array 114 of videos consisting of a plurality (``N'' number) of 2D and / or 3D video frames 114n of the hand 52 of person 50 typing on keypad 62. Here, 1 ≦ n ≦ N. Imaging system 110 is connected to hub 120 by at least one wired and / or wireless communication channel 113, through which imaging system 110 transmits the acquired video frames to the hub. Hub 120 is configured to process the received video frames 114n to identify person 50. Person 50 is the person whose hand 52 is imaged within the video frame. The hub comprises and / or has access to data and / or executable instructions and any of various electronic and / or optical physical and / or virtual processors, memories, and / or wired or wireless communication interfaces. These may be required to support the functions provided by the hub. The data and / or executable instructions are hereinafter also referred to as software. Also, the processor, memory, and / or communication interface are hereinafter also referred to as hardware.

[0022] As an example, hub 120 comprises software and hardware that support object detection module 130, AFID identifier module 140, and classifier module 150. Object detector module 130 is operable to detect the ROI within video frame 114n. AFID identifier module 140 identifies the AFIDs within the detected ROI and provides an ST-ID for each of the identified AFIDs. Classifier module 150 comprises a non-local classifier operable to process the set of ST-IDs as a spatio-temporal graph to identify person 50.

[0023] In one embodiment, the object BROI detector module 130 comprises a fast object detector such as a YOLO (You Look Only Once) detector that can detect related BROIs in real time. The AFID identifier module 140 may comprise a convolutional pose machine (CPM) for identifying AFIDs in the detected BROIs. The classifier module 150 comprises an optionally adaptable non-local graph convolutional network as described above and discussed below. In FIG. 2, the classifier module 150 is schematically shown to provide an output of probabilities represented by a histogram 152. The histogram gives, for each of a plurality of persons, the probability that the given person is the person in which the typing hand 52 is imaged within the video frame. DYNAMIDE 100 is trained to recognize that a given person is the person in which the typing hand 52 is imaged within the video frame. DYNAMIDE 100 is schematically shown as successfully identifying person 50 as the person in which the typing hand 52 is imaged in video frame 114n.

[0024] In one embodiment, the AFIDs that DYNAMIDE 100 uses to identify a person's typing are the joints of the typing hand (finger joints and / or wrist joints) and the finger bones (phalanges). FIG. 3A schematically shows an image of a hand 200 having finger joints (also called knuckles) and wrist joints that are optionally used as AFIDs by DYNAMIDE 100 to process a video image of the typing hand, according to an embodiment of the present disclosure. The joints have an arrangement on the hand 200 indicated by the plus symbol "+", and as shown in the figure, can be generically referred to by the label "JH" of the wrist joint, and can be individually distinguished by the numerical labels J0, J1, …, J20. A given phalanx that DYNAMIDE 100 can use as an AFID for typing behavior is identified when the given phalanx is referred to by alphanumeric labels indicating two connecting finger joints. For example, in FIG. 3A, the finger bone connecting joints J5 and J6 is schematically shown in FIG. 3A by a dashed line labeled B5-6, and the phalanx B18-19 connects finger joints J18 and J19. The finger bones can be generically referred to by the label BH.

[0025] FIG. 3B schematically shows a spatial graph (S-graph 200) that can be used to represent the spatial relationship of the AFIDs at a given time, according to an embodiment of the present disclosure. As an example, the spatial graph is shown by the hand 200 at a given time when the hand 200 is imaged. In the spatial S-graph 200, the finger joint AFIDs JH shown in FIG. 3A are represented by nodes generically referred to by the label JN. The nodes JN are individually distinguished by the alphanumeric labels JN0, JN1, …, JN20 corresponding to the homologous finger joints J0, J1, …, J20 shown in FIG. 3A, respectively. The edges of the S-graph 200 connecting the nodes JN represent the finger bones, i.e., the bone AFIDs connecting the finger joints. As shown in FIG. 3B, the edges can be generically referred to by the label BE and are individually referred to by reference labels corresponding to the homologous finger bones in the hand 200. For example, the edge BE5-6 in FIG. 3B corresponds to the bone B5-6 in FIG. 3A.

[0026] Figure 4A schematically shows an enlarged image of the nth video frame 114n in the array 114 (Figure 2) of video frames acquired by the imaging system 110 at the acquisition time tn and transmitted to the DYNAMIDE hub 120 for processing. The video frame 114n images the hand 52 typing on the keypad 62, as well as features in the environment surrounding the hand that may be located within the FOV 111 (Figure 2) of the imaging system 110. The surrounding features schematically shown in Figure 4A may include, for example, parts of the structure of the ATM 60 such as the counter 64 and the side wall 66, as well as the mobile phone 55 placed by the person 50 on the counter 64.

[0027] As described above, in the processing of the array 114 of video frames 114n, the object detection module 130 may determine a bounding box that locates the image of the hand 52 in the frame as an object having an articulated AFID that the AFID detector 140 identifies and that DYNAMIDE 100 uses to identify the person 50. The bounding box determined by the object detector module 130 for the hand 52 within the video frame 114n is shown by the dashed rectangle 116. The finger joint AFIDs detected and identified by the AFID detector 140 within the bounding box 116 are shown by the general-purpose AFID label JH (Figure 3A). Figure 4B schematically shows the spatial S-graph-52(tn) that models the hand 52 as a graph based on the image of the hand within the video frame 114n acquired at the acquisition time tn. The finger joint nodes within the S-graph-52(tn) are shown by appropriate finger joint node labels JN0, JN1, …, JN20, and an argument indicating the acquisition time tn associated with the S-graph-52(tn) to which the node belongs may be added. For example, the nodes JN0, JN1, ···, JN20 of the S-graph-52(tn) may be referred to as JN0(tn), JN1(tn), …, JN20(tn).

[0028] FIG. 5A schematically shows an enlarged image of a video array 114 shown in FIG. 2, including video frames 1141, 1142, 1143, ..., 114N that image a hand 52 typing on an ATM 60 at respective times t1, t2, t3, ..., tN. FIG. 5B schematically shows an ST-graph 52 that models the spatio-temporal progression of typing behavior based on images of the hand 52 within video frames 1141 to 114N, according to an embodiment of the present disclosure. The ST graph 52 includes a spatial S-graph -52(tn) corresponding to an image of the hand 52 in video frames 1141, ..., 114N. Here, 1 ≦ n ≦ N. Homologous nodes JN in adjacent S-graphs, S-graph -52(tn) and S-graph -52(tn+1), are connected by a time edge representing the elapsed time between respective acquisition times tn and tn+1. All time edges between adjacent S-graphs -52(tn) and S-graph -52(tn+1) have the same temporal length and are labeled TEn,n+1. Some of the time edges in FIG. 5B are labeled by their respective labels.

[0029] The node data associated with ST-graph-52 provides a set of spatio-temporal input features that are processed by the classifier module 150 of DYNAMIDE hub 120 to determine the identity of the person 50 typing on the keypad 62 of ATM60. The set of input features can be modeled as an input spatio-temporal feature tensor 300, as schematically shown in FIG. 6A, and the input spatio-temporal feature tensor 300 has AFID, time, and channel axes that indicate positions within the tensor by row, column, and depth. In ST-graph-52, the AFID axis is calibrated with node numbers indicating specific joints in the hand 52, and the time axis is calibrated with consecutive frame numbers or frame acquisition times. As an example, the channel axis of the spatio-temporal feature tensor 300 schematically shows four channels, but it should be noted that the spatio-temporal feature tensor according to one embodiment may have more or fewer than four channels. For example, the entry along the channel axis corresponding to a given node and a given time, each shown along the AFID and time axes, may provide two or three spatial coordinates that determine a spatial position corresponding to the given node at the given time. The channel entry may also provide an error estimate of the accuracy of the coordinates and the probability that a given node is correctly identified.

[0030] In one embodiment, the classifier module 150 according to one embodiment of the present disclosure has a classifier that includes at least one non-local graph convolutional network (NLGCN) for processing data within the tensor 300 and may provide the identity of the person 50. Optionally, the at least one NLGCN comprises at least one adaptive ANLGCN that includes an adaptive adjacency matrix in addition to the non-local GCN layer. The adaptive adjacency matrix operates to improve the classifier recognition of spatio-temporal motion at the hand joints that are related to each other. The spatio-temporal motion is not affected by the spatial structure and is a motion specific to the way a person performs typing.

[0031] As an example, FIG. 6B shows a schema of a classifier 320 that can be used for the DYNAMIDE hub 120 to process data within the tensor 300, according to an embodiment of the present disclosure. The classifier 320 optionally includes convolutional neural network blocks 322, 324, and 326 that forward data to a fully connected network FCN328. The fully connected network FCN328 provides, for each of a plurality of persons, a probability as to whether the person is the person whose hand 52 was imaged in the video sequence 114 (FIG. 2) of typing on the keypad 62. Block 322 optionally includes a GCN that forwards data to a time convolutional network (TCN). Block 324 includes an ANL-GCN that forwards data to the TCN, and block 326 includes a GCN that forwards data to the ANL-TCN.

[0032] Accordingly, according to one embodiment of the present disclosure, a method for identifying a person is provided. The method includes obtaining spatio-temporal data for each of a plurality of anatomical landmarks related to an action in which a person is involved, the data providing at least one spatio-temporal trajectory of the anatomical landmark during the period of the action; modeling the obtained spatio-temporal data as a spatio-temporal graph (ST-graph); and processing the ST-graph using at least one non-local graph convolutional neural network (NLGCN) to provide an identity corresponding to the person. Optionally, the at least one NLGCN comprises at least one adaptive NLGCN (ANLGCN) including an adaptive adjacency matrix learned in response to data regarding anatomical landmarks consisting of the plurality of anatomical landmarks, the anatomical landmarks consisting of the plurality of anatomical landmarks not being determined solely by the body structure of the person. Additionally or alternatively, processing the ST-graph includes segmenting the plurality of anatomical landmarks into a plurality of sets of anatomical landmarks, each set being characterized by a configuration with a different degree of freedom of movement. Optionally, the method includes modeling the obtained spatio-temporal data related to the anatomical landmarks within each set as an ST-graph. The processing step includes processing the ST-graph modeled for each set in the plurality of sets of anatomical landmarks using an NLGCN consisting of the at least one NLGCN to determine data indicating the identity of the person, the determination being optionally independent of processing the other sets among the plurality of sets. The method optionally includes fusing the determined data from all the sets to provide the identity for the person.

[0033] In one embodiment, obtaining the spatio-temporal data includes obtaining a series of video frames that image the person involved in the action, each video frame including an image of at least one body region of interest (BROI) that images an anatomical landmark consisting of the plurality of anatomical landmarks. Optionally, the method includes processing the video frame to detect the at least one ROI in each video frame. Additionally, or alternatively, the method optionally includes identifying, in each of the at least one detected ROI, an image of an anatomical landmark consisting of the plurality of anatomical landmarks. Optionally, the method includes processing the identified image of the anatomical landmark to determine the data that defines the spatio-temporal trajectory.

[0034] In one embodiment, the plurality of anatomical landmarks includes joints. Optionally, the plurality of anatomical landmarks includes bones that connect the joints. Additionally, or alternatively, the joints include finger joints. Optionally, the action includes a series of finger movements. The finger movements may include movements involved in operating a keyboard.

[0035] In one embodiment, the joints include joints of the large extremities. Optionally, the action is a sport. Optionally, the sport is soccer. Optionally, the sport is golf.

[0036] In one embodiment, the plurality of anatomical landmarks includes facial landmarks. Optionally, the facial landmarks include facial landmarks whose movements are used to define action units (AUs) of the Facial Action Coding System (FACS) for classifying expressions and micro-expressions. In one embodiment, the plurality of anatomical landmarks includes detailed pairs of features of fingerprints of multiple fingers of the hand.

[0037] Furthermore, according to an embodiment of the present disclosure, a system for identifying a person is provided. The system includes an imaging system operable to acquire an image having video frames that image a person involved in an action, and software usable to process the video frames according to any of the claims to provide an identity corresponding to the person.

[0038] The description of embodiments of the invention in this application is provided by way of example and is not intended to limit the scope of the invention. The described embodiments include different features, and not all of them are required in all embodiments. Some embodiments utilize only some of the features, or possible combinations of features. Variations of the described embodiments of the invention, and embodiments including different combinations of the features described in the described embodiments, will occur to those skilled in the art. The scope of the invention is limited only by the claims.

Claims

1. 1. A method for identifying a person, comprising: storing data and executable instructions in a memory; and executing the executable instructions in a processor to perform the steps of: The process comprises: acquiring spatiotemporal data for each of a plurality of anatomical landmarks associated with an activity involving a human being, the data providing data defining a spatiotemporal trajectory of at least one of the anatomical landmarks during the activity; modelling the acquired spatio-temporal data as a spatio-temporal graph (ST-graph); processing the ST-graph using a multi-stream graph convolutional neural network (GCN) to provide an identity corresponding to the person; Including, The multi-stream GCN comprises a plurality of non-local graph convolutional networks (NLGCNs); the multi-configuration NLGCN is configured to process a set of anatomical landmark data characterized by independent degrees of freedom of motion; A method according to claim 1, wherein an output of the multi-stream GCN is used to provide the identity.

2. The method described in claim 1, wherein at least one of the NLGCNs comprises at least one adaptive NLGCN (ANLGCN) including an adaptive adjacency matrix that is learned in response to data regarding anatomical landmarks consisting of the plurality of anatomical landmarks, and the anatomical landmarks consisting of the plurality of anatomical landmarks are not determined solely by the person's physical structure.

3. 2. The method of claim 1, wherein acquiring the spatiotemporal data includes acquiring a series of video frames imaging the person engaged in the activity, each video frame including an image of at least one body region of interest (BROI) imaging an anatomical landmark of the plurality of anatomical landmarks.

4. The method of claim 3 , comprising processing the video frames to detect the at least one BROI in each video frame.

5. The method of claim 3 , further comprising identifying, in each of the at least one detected BROI, an image of an anatomical landmark of the plurality of anatomical landmarks.

6. The method of claim 5 , comprising processing images of the identified anatomical landmarks to determine the data defining the spatiotemporal trajectory.

7. The method of claim 3 , wherein the plurality of anatomical landmarks includes joints.

8. The method of claim 7 , wherein the plurality of anatomical landmarks includes bones connecting the joints.

9. The method of claim 7 , wherein the joints include finger joints.

10. The method of claim 9 , wherein the actions include a sequence of finger movements.

11. The method of claim 10 , wherein the finger movements include movements involved in operating a keyboard.

12. 1. A system for identifying a person, comprising: an imaging system operable to capture a video having video frames capturing the person engaged in an activity; acquiring spatiotemporal data for each of a plurality of anatomical landmarks associated with the activity involving the person, the anatomical landmarks providing data determining a spatiotemporal trajectory of at least one of the anatomical landmarks during the activity; Modeling the acquired spatio-temporal data as a spatio-temporal graph (ST-graph); processing the ST-graph with a multi-stream non-local graph convolutional neural network (GCN) to provide an identity corresponding to the person; and software operable to process the video frames and provide an identity corresponding to the person; processing the ST-graph using a multi-stream graph convolutional neural network (GCN) to provide an identity corresponding to the person; Including, The multi-stream GCN comprises a multi-component non-local graph convolutional network (NLGCN); the multi-configuration NLGCN is configured to process a set of anatomical landmark data characterized by independent degrees of freedom of motion; The output of the multi-stream GCN is used to provide the identity.

Citation Information

Patent Citations

  • Identity authentication method based on dynamic gestures

    CN111444488A

  • Authentication device, crime prevention system, authentication method, and program

    JP2017049867A