Dynamic identity authentication
By utilizing the spatiotemporal trajectory features of anatomical landmarks and nonlocal neural networks, the dynamic identification method (DYNAMIDE) solves the problem that traditional multi-factor authentication technology struggles to achieve high-quality authentication in complex activity matrices, enabling real-time individual identification and efficient user authentication.
Patent Information
- Application Number
- CN202180050653.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-20
- Filing Date
- 2021-07-30
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2041-07-30
AI Technical Summary
Traditional multi-factor authentication technologies struggle to meet the complexity and interdependence of modern citizen activity matrices, especially under the stringent customer authentication requirements of the European Payment Services Directive, making it difficult to provide high-quality user authentication.
The Dynamic Recognition Method (DYNAMIDE) is employed to identify specific individuals by recognizing the spatiotemporal trajectory features of anatomical landmarks during an individual's performance of an activity. This is achieved by using nonlocal neural networks to process image sequences, including graph convolutional networks and adaptive nonlocal neural networks to process spatiotemporal graphs of the activity reference.
It enables real-time individual identification during the execution of activities, improving the accuracy and efficiency of user authentication and meeting the requirements of strong customer authentication specifications.
Smart Images

Figure CN116635910B_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims the benefit under 35 U.S.C. 119(e) of U.S. provisional application number 63 / 067,890, filed August 20, 2020, the disclosure of which is incorporated herein by reference. TECHNICAL FIELD
[0003] Embodiments of the present application relate to methods and devices for providing biometric authentication of personal identity. BACKGROUND
[0004] An increasing list of services requires a procedure to authenticate and authorize user access to the service, commonly referred to as a multi-factor authentication procedure (MFA). In an MFA procedure, the user is required to provide a suitable response to each of a plurality of challenges. The challenge categories are referred to as “authentication factors”. A common MFA is referred to as two-factor authentication (2FA), in which the user is required to correctly respond to at least two of three authentication factors: a knowledge factor, which tests for something the user should know, such as a password; a possession factor, which requires the user to present something they should have, such as a credit card or smartphone; and an inherent factor, which requires the user to present something that represents a characteristic of the user, such as a biometric like a fingerprint, voice print, or iris scan.
[0005] However, conventional authentication techniques appear to struggle with being easy to use and providing high quality authentication, as required by the complexity and interdependence of the matrix of activities that modern citizens often engage in. For example, conventional MFA configurations appear to struggle with meeting the Strong Customer Authentication (SCA) specifications of the revised European Payment Services Directive (PSD2), which brings consumers, banks, and third party providers (TPPs) into the Open Banking initiative. Implementation of SCA has been delayed twice. The regime originally scheduled to be in place in September 2019 was delayed to March 14, 2021, and then again to September 14, 2021 - the current planned deadline. SUMMARY
[0006] One aspect of the embodiments of the present application refers to providing a method, which can be referred to as a dynamic identification method, or simply DYNAMIDE, for identifying a person based on characteristics of the way the person performs an activity. According to one embodiment of the present application, DYNAMIDE comprises identifying anatomical landmarks, optionally referred to as activity fiducials (AFIDs), which exhibit various degrees of motion or lack of motion during performance of an activity by a person, and whose spatiotemporal trajectories during performance of the activity can be used to identify the activity. DYNAMIDE comprises processing said trajectories to determine characteristics of said trajectories, which are conducive to distinguishing between activities performed by individuals performing said activities, and to identifying a specific individual performing said activities.
[0007] For an activity performed by an individual performing the activity, the activity characteristics that distinguish said activity can be extremely subtle, and the AFID trajectories associated with said activity can exhibit a large amount of subtle and non-intuitive cross-talk. As a result, a characteristic of one spatiotemporal trajectory of an activity can appear to be unrelated to a characteristic of another spatiotemporal trajectory of the activity, but in fact it can be specific to the individual performing the activity, and provides a basis for identifying the individual. One embodiment of the present application provides a spatiotemporal method that is conducive to discovering and using characteristics exhibited by trajectories for identification, whose spatial and / or temporal processing can be non-local, and advantageously limits many of the a priori processing constraints assumed by motion exhibited by AFID trajectories.
[0008] According to one embodiment, a scheme for identifying a specific individual based on a given activity that the individual can perform comprises acquiring a sequence of images of the individual performing the given activity, and identifying in the images AFIDs associated with the given activity. The spatiotemporal trajectories exhibited by the identified AFIDs can be determined by processing the images, and said trajectories can be processed to identify from a plurality of individuals who can have performed said activity, a specific individual who performed said activity. Optionally, processing said AFID trajectories comprises determining local and non-local spatiotemporal correlations exhibited by the activity fiducials during performance of said given activity, and using said correlations to determine the identity of said specific individual. The spatiotemporal correlations can comprise correlations based on spatial parameters, temporal parameters, or both temporal and spatial parameters, which characterize one or more spatiotemporal trajectories of one or more AFIDs.
[0009] According to one embodiment of the present application, the activity reference associated with a given activity can be an anatomical landmark of any body part, such as a limb, a face or a head, which exhibits a spatiotemporal trajectory suitable for identifying the person performing the activity when performing the given activity. For example, the activity reference can be a joint or a bone of a limb which exhibits a suitable spatiotemporal trajectory during the performance of an activity such as walking, playing golf or entering a password on an ATM machine. For a typing activity, the activity reference can include a plurality of joints of the hand skeleton. The activity reference can be a facial landmark, such as eyebrows, eyes and corners of the lips, whose movements are used to define action units (AUs) of the Facial Action Coding System (FACS) used to classify facial expressions and microexpressions. The activity reference can also be a minutiae pair of fingerprints of a plurality of fingers of a hand, non-contact imaged with sufficient optical resolution to enable the identification of the minutiae pair.
[0010] According to one embodiment, DYNAMIDE uses at least one neural network to process images of an activity to identify the individual performing the activity. In one embodiment, the at least one neural network is trained to detect target body parts or regions of interest (BROI) in an image and to identify the activity reference they can include. The spatial and temporal progression of the activity reference identified during the performance of the activity is represented by a spatiotemporal graph (ST-Graph) in which the activity reference is a plurality of nodes connected by spatial and temporal boundaries which define the activity reference spatiotemporal trajectory of said activity. The at least one neural network can include at least one graph convolutional network (GCN) for processing said trajectory and classifying said activity according to the individual performing said activity.
[0011] In one embodiment, the at least one GCN includes a nonlocal neural network (NLGCN) having at least one nonlocal neural network block for processing the activity reference spatiotemporal trajectory. The at least one nonlocal neural network block can include at least one spatial nonlocal neural network block and / or at least one temporal nonlocal neural network block. Optionally, the NLGCN is configured as a multi-stream graph convolutional network including a plurality of component nonlocal neural networks for processing a plurality of sets of data characterized by independent degrees of freedom based on said activity reference trajectory. In one embodiment, the output of the multi-stream graph convolutional network can include a weighted average of the outputs of each component graph convolutional network.
[0012] For example, when DYNAMIDE is configured to recognize an individual by way of individual typing, the articulation reference of the hand joints is characterized by degrees of freedom of motion (e.g. distances between the joints of different fingers) that are independent of the degrees of freedom of motion available through the articulation reference of the hand bones that are connected joints. Thus, in one embodiment, DYNAMIDE can include a dual-stream 2s-NLGCN multi-stream graph convolutional network with two component non-local neural networks. One of the two component non-local neural networks processes the articulation reference, and the other processes the bone articulation reference. In one embodiment, the articulation non-local neural network includes at least one learnable “adaptive” adjacency matrix that is essentially data-driven to reduce some a priori constraints that are available to configure the 2s-non-local neural network. In accordance with an embodiment of the present application, a 2s-non-local neural network that includes an adaptive adjacency matrix can be referred to as an adaptive 2s-non-local neural network (2s-ANLGCN). The outputs of the 2s-non-local neural network or the articulation and bone non-local neural networks of the adaptive 2s-non-local neural network of a typing-type DYNAMIDE can be fused to recognize the individual.
[0013] In accordance with one embodiment, recognizing a particular individual is done in real time. The real-time recognition referred to by the embodiments refers to recognizing the individual while the individual is performing an activity, or within a time frame in which the quality of experience (QoE) of the service performing the recognition is not substantially degraded by the recognition process.
[0014] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS
[0015] Non-limiting examples of embodiments of the present application are described below with reference to the accompanying drawings, in which:
[0016] Figure 1 A flowchart of a process by which DYNAMIDE can process a sequence of video frames of a person performing an activity to recognize the person, in accordance with an embodiment of the present application, is shown.
[0017] Figure 2The illustration schematically depicts the process by which the DYNAMIDE system, according to one embodiment of this application, processes a sequence of video frames as a person types on an ATM (Automatic Teller Machine) keypad to identify that person;
[0018] Figure 3A The illustration shows an image of a person's hand typing on a keyboard, along with a reference point for the hand's movement, which DYNAMIDE can use to identify the person, according to one embodiment of this application.
[0019] Figure 3B This illustration schematically shows one embodiment according to the present application, regarding... Figure 3A The spatial graph after modeling the hand in the image is called the S-Graph.
[0020] Figure 4A The diagram schematically shows an enlarged view of a video frame according to an embodiment of this application, from... Figure 2 The image shows a sequence of video frames captured from a person typing on an ATM.
[0021] Figure 4B The illustration schematically depicts one embodiment according to this application, regarding... Figure 4A The S-graph is a model of the hand imaged in the video frame shown.
[0022] Figure 5A An embodiment according to this application is illustrated schematically. Figure 2 A magnified image of the video frame sequence shown;
[0023] Figure 5B This illustration schematically shows an embodiment of the present application, with Figure 5A The spatiotemporal graph (ST-Graph) corresponding to the images of the video frame sequence shown;
[0024] Figure 6A An illustrative spatiotemporal feature tensor according to an embodiment of this application is shown, which includes and Figure 5B The data related to the nodes in the spatiotemporal graph shown;
[0025] Figure 6B This application demonstrates a pattern of a nonlocal neural network according to an embodiment of the present application, which Dynamide can use to process... Figure 6A The data in the tensor shown. Detailed Implementation
[0026] In the discussion, unless otherwise stated, adjectives such as "substantially" and "approximately" that modify another adjective or a noun mean that the captioned term is "within X percent of" the stated term, where X is up to 10% or 5% or 1% of the stated term. When used in the discussion, general terms appear anywhere in one or more illustrative examples, the one or more examples are non-limiting examples of the general term, and the general term is not limited to the one or more specific examples mentioned. The phrase "in an embodiment," whether or not associated with a permission such as "may," "optionally," or "as an example," is used to introduce an example of an optional embodiment of the application, but is not necessarily a required configuration. Unless otherwise stated, the word "or" is considered to be inclusive of "and" in the description and claims, and means it connects at least one of one or more conjuncts.
[0027] According to embodiments of the application, Figure 1 A high-level flowchart 20, optionally also denoted by the numeral 20, is shown by which DYNAMIDE can respond to activities performed by a person, thereby identifying the person.
[0028] According to embodiments of the present application, in block 22, DYNAMIDE optionally acquires a sequence of video frames of a person participating in an activity, DYNAMIDE is configured to determine the identity of the person participating in the activity by processing the sequence. In block 24, DYNAMIDE processes the video frames to identify images of body regions of interest (BROI) in the video frames, at least one activity reference related to the activity is imaged therein. Identifying the BROI in the video frames optionally includes determining at least one bounding box of frames comprising images of the BROI. In block 26, DYNAMIDE processes each bounding box determined from the video frames to identify images of at least one activity reference in each bounding box. Identifying images of the activity reference in the bounding box of the video frames optionally includes associating a spatiotemporal ID (ST-ID) with the image, the spatiotemporal ID comprises an identification tag of the activity reference, “AFID ST-ID”, which is used to mark all identified images of the same activity reference in the video frames, and determines spatiotemporal coordinates of the image. The spatiotemporal coordinates comprise a time stamp and at least two spatial coordinates. The time stamp identifies a time (a position in time) at which the video frame comprising the bounding box in which the activity reference is located is obtained, among other times in which other video frames in the sequence of video frames are obtained. The at least two spatial coordinates correspond to a spatial position of the activity reference at the time indicated by the time stamp. Optionally, a given AFID ST-ID of an identified activity reference comprises a standard deviation (SD) of each spatial coordinate and a probability of correctness of the activity reference ID tag associated with the AFID ST-ID. The earliest and latest time stamps and extreme spatial coordinates determined by the AFID ST-ID determine a spatiotemporal volume, which can be referred to as a spatiotemporal AFID hull (ST-Hull), which contains spatiotemporal coordinates of all activity reference instances imaged and identified in the sequence of video frames.
[0029] In block 28, DYNAMIDE uses the spatiotemporal IDs of the activity fiducials to configure the identified instances of the activity fiducials as nodes of an activity fiducial spatiotemporal graph (ST-Graph) connected by spatial and temporal boundaries. The spatial boundaries connect nodes of the spatiotemporal graph that represent imaging instances of the activity fiducials identified by the same timestamp, i.e. imaging instances of the activity fiducials imaged in the same video frame, and by the human structure representing the spatial constraints imposed on the activity fiducials. The configuration of nodes connected by spatial boundaries and representing the spatial relationship of the activity fiducial instances imaged in the same given frame and at a given time can be referred to as a spatial graph (S-Graph) of the activity fiducials at the given time. The temporal boundaries connect temporally adjacent nodes in the spatiotemporal graph, the spatiotemporal graph representing the images of the same activity fiducials in two consecutively acquired video frames in a sequence of video frames. The temporal boundaries represent the elapsed time between two consecutive timestamps. The spatiotemporal graph can be considered to comprise the spatial graph of the activity fiducials connected by temporal boundaries.
[0030] In block 30 of one embodiment, DYNAMIDE uses a non-local graph convolutional neural network (optionally adaptive, i.e. ANLGCN) to process the activity fiducial spatiotemporal graph to determine (optionally in real-time) the person among a plurality of persons participating or engaged in the activity for which the adaptive non-local graph convolutional neural network is trained to identify. In one embodiment, the adaptive non-local graph convolutional neural network is configured to span the activity fiducial spatiotemporal hull and enable data associated with imaging instances of the activity fiducials at any spatiotemporal location in the hull to be weighted by learned weights and contribute to the convolution performed by the adaptive non-local graph convolutional neural network for any spatiotemporal location in the hull. Optionally, the NLGCN is configured as a multi-stream graph convolutional neural network comprising a plurality of component non-local graph convolutional neural networks for processing a plurality of groups of activity fiducial data characterized by independent degrees of freedom.
[0031] According to embodiments of the present application, Figure 2 The DYNAMIDE system 100 is schematically illustrated as configured to perform Figure 1 The illustrated process, and the way in which the person performing the activity is identified. The DYNAMIDE system 100 can comprise a processing center 120 (optionally cloud-based) and an imaging system 110 having a field of view (FOV) indicated by dashed line 111. As an example, in this figure the activity is the entering of a PIN on keypad 62 by person 50 at ATM 60.
[0032] The imaging system 110 is used to provide the 2D and / or 3D “N” frame video frames 114 of the one or more hands 52 of the plurality of persons 50 entering the PIN on keypad 62n (1 < n < N) of a video sequence 114. The imaging system 110 is connected to a processing center 120 through at least one wired and / or wireless communication channel 113, through which the imaging system 110 transmits the video frames it acquires to the processing center. The processing center 120 is configured to process the received video frames 114 n to identify the person 50 whose hand 52 is imaged in the video frames. The processing center comprises and / or has access to data and / or executable instructions (hereinafter also referred to as software), and to any of various electronic and / or optical-physical and / or virtual processors, memories and / or wired or wireless communication interfaces (hereinafter also referred to as hardware), which can be required to support the functions provided by the processing center.
[0033] By way of example, the processing center 120 comprises software and hardware supporting an object detection module 130 usable to detect relevant body regions in the video frames 114 n , an activity landmark identification module 140 usable to identify activity landmarks within the detected relevant body regions and to provide a spatio-temporal ID of each identified activity landmark, and a classifier module 150 comprising a non-local classifier usable to process the set of spatio-temporal IDs into a spatio-temporal graph to identify the person 50.
[0034] In one embodiment, the object relevant body region detection module 130 comprises a fast object detector, such as a YOLO (You Look Only Once) detector capable of detecting relevant body regions in real time. The activity landmark identification module 140 can comprise a convolutional pose machine (CPM) for identifying activity landmarks within the detected relevant body regions. The classifier module 150 comprises a non-local graph convolutional network (optionally adaptive), mentioned above and discussed below. The classifier module 150 is schematically shown in Figure 2 , providing an output of probabilities represented by a histogram 152. The histogram gives the probability of each given person out of a plurality of persons being identified when the DYNAMIDE 100 is trained to identify the given person whose hand 52 is imaged in the video frames and is typing. The Dynamide 100 is schematically represented as being able to successfully identify the owner 50 of the hand when the imaged hand 52 is typing in the video frames 114 n .
[0035] In one embodiment, the DYNAMIDE100 uses the activity references of the person typing as the starting point to identify the joints (finger and / or wrist joints) and finger bones (phalanges) of the hand typing. According to one embodiment of this application, Figure 3A An image of a hand 200 with finger joints (also called knuckles) and a wrist joint is schematically shown. The wrist joint is optionally used by the DYNAMIDE 100 as a reference point for processing video images of the typing hand. The position of the joint on the hand 200 is indicated by a plus sign "+" and, as shown, is typically indicated by a hand joint label "JH" and individually distinguished by alphanumeric labels J0, J1, ..., J20. When the alphanumeric label is referenced, a given finger bone used as a reference point for typing activity by the DYNAMIDE 100 can be identified; the alphanumeric label indicates the two knuckles connected to the given finger bone. For example, in... Figure 3A In the diagram, the phalanges connecting joints J5 and J6 are schematically shown by dashed lines marked B5-6, and phalanges B18-19 connect joints J18 and J19. Phalanges are generally referenced by the label BH.
[0036] According to one embodiment of this application, Figure 3B A spatial diagram 200 is schematically shown, which can be used to represent the spatial relationships of an activity reference at a given time, and the spatial relationships of a hand 200 imaged at a given time are illustrated as an example. Figure 3A As shown in spatial diagram 200, the hand joint motion reference JH is represented by nodes typically referenced by the label JN. Nodes JN are distinguished by alphanumeric labels JN0, JN1, ..., JN20, and correspond to... Figure 3A The homologous phalanges J0, J1, ..., J20 are shown. The edges of the spatial diagram 200 connecting nodes JN represent phalanges, i.e., the bone reference points for movement, which connect to the phalanges. Figure 3B As shown, edges can generally be referenced via the label BE, and separately referenced via reference labels corresponding to the homologous finger bones in hand 200. For example, Figure 3B The edge BE5-6 in the middle corresponds to Figure 3A B5-6 of the bone.
[0037] According to one embodiment of this application, Figure 4A The illustration schematically shows sequence 114 in the video frame. Figure 2 The nth video frame 114 n A magnified image of the video frame obtained by the imaging system 110 at time t. n The video frame 114 is acquired and transmitted to the DYNAMIDE processing center 120 for processing. nThe hand 52 typing on the keypad 62 and the surrounding environment of the hand (the hand may be located in the FOV 111 of the imaging system 110) Figure 2 Imaging is performed using features within the image. Surrounding features are also included. Figure 4A The illustration may include, for example, a part of the structure of the ATM 60, such as the counter 64 and side wall 66, and a mobile phone 55 that a person 50 has placed on the counter 64.
[0038] As mentioned above, in processing video frame 114 n In sequence 114, the object detection module 130 can determine the bounding box on the frame that locates the image of the hand 52 as an object, which includes the joint joint motion reference identified by the motion reference detector 140 and the object used by the DYNAMIDE 100 to identify the person 50. This is achieved by the object detector module 130 in video frame 114. n The bounding box defined by the middle opponent 52 is represented by dashed rectangle 116. The active reference detector 140 detects and identifies the knuckle active reference within the bounding box 116 using a generic active reference label JH( Figure 3A )express. Figure 4B The diagram schematically illustrates the space map-52(t) n In this figure, based on the time t obtained n 114 video frames obtained n The image of the hand will be used to model hand 52. Spatial diagram - 52(t) n In the diagram, knuckle nodes can be represented by appropriate knuckle node labels JN0, JN1, ..., JN20, with an added parameter to indicate the spatial graph to which the node belongs. n The relevant acquisition time t n For example, spatial diagram -52(t) n The nodes JN0, JN1, ..., JN20 in ) can be referenced from JN0(t n ),JN1(t n ),…,JN20(t n ).
[0039] Figure 5A schematically shown Figure 2 The displayed magnified image of video sequence 114 includes the corresponding times t1, t2, t3, ..., t N The video frames 1141, 1142, 1143, ..., 114 of the hand typing on the ATM60 are captured. N According to one embodiment of this application, Figure 5B The spatiotemporal diagram 52 is schematically shown, on which the data is based on video frames 1141-114... NThe spatio-temporal progression of the typing activity performed by the images of the inner hand 52 is modeled. The spatio-temporal graph 52 comprises spatial graphs - 52(t n ) for 1 < n < N, corresponding to the video frames 1141-114 N of the images of the inner hand 52. Between adjacent spatial graphs, spatial graph-52(t n ) and spatial graph-52(t n+1 ), homologous nodes JN are connected by a time edge representing the elapsed time between their respective acquisition times t n and t n+1 . All time edges between adjacent spatial graphs - 52(t n ) and spatial graph-52(t n+1 ) have the same time length and are labeled TE n,n+1 . Figure 5B Some of the time edges are labeled by their respective labels.
[0040] The node data associated with the spatio-temporal graph - 52 provides a set of spatio-temporal input features, which are processed by the classifier module 150 of the DYNAMIDE processing hub 120 to determine the identity of the person 50 typing on the keypad 62 of the ATM 60. As schematically illustrated in Figure 6A , this set of input features can be modeled into an input spatio-temporal feature tensor 300, which has activity, time and channel axes indicating the position in the tensor by row, column and depth. For the spatio-temporal graph - 52, the activity axis is calibrated by the node number, which represents a particular joint in the hand 52, and the time axis is calibrated by the sequential frame number or acquisition time of the frames. It should be noted, for example, that although the channel axis of the spatio-temporal feature tensor 300 schematically illustrates four channels, according to embodiments, the spatio-temporal feature tensor can have more or less than four channels. For example, a channel axis entry indicating a given node and a given time along the activity and time axes, respectively, can provide two or three spatial coordinates for determining the spatial position of the given node at the given time. The channel entry can also provide an error estimate for the precision of the coordinates and the probability that the given node is correctly identified.
[0041] In one embodiment, the classifier module 150 can have a classifier comprising at least one non-local graph convolutional network (NLGCN) to process the data in the tensor 300 and provide an identity for the person 50 in accordance with embodiments of the present application. Optionally, the at least one non-local graph convolutional neural network comprises at least one adaptive non-local graph convolutional neural network comprising an adaptive adjacency matrix in addition to non-graph convolutional network layers. The adaptive adjacency matrix is used to improve classifier recognition of the spatio-temporal motion of the hand joints relative to each other, which is not restricted by the spatial structure and is specific to the way the person types.
[0042] In accordance with embodiments of the present application, as an example, Figure 6B A schema of the classifier 320 is shown, which can be used by the DYNAMIDE processing center 120 to process the data in the tensor 300. The classifier 320 optionally comprises convolutional neural network blocks 322, 324 and 326 that feed data forward into a fully connected network 328 (FCN) that provides a probability for each of the plurality of persons as to whether this person is the person that is typing on the keypad 62 and whose hands 52 are imaged in the video frames 114 Figure 2 ). Block 322 optionally comprises a graph convolutional network that feeds data forward into a temporal convolutional network (TCN). Block 324 comprises an adaptive non-local graph convolutional network (ANL-GCN) that feeds data forward into a temporal convolutional network, and block 326 comprises a graph convolutional network that feeds data forward into an adaptive non-local graph convolutional network.
[0043] Accordingly, one embodiment of the present application provides a method of identifying a person, the method comprising: obtaining spatio-temporal data for each of a plurality of anatomical landmarks associated with an activity in which the person is engaged, the spatio-temporal data providing data defining at least one spatio-temporal trajectory of the anatomical landmark during the activity; modeling the obtained spatio-temporal data as a spatio-temporal graph (ST-Graph); and processing the spatio-temporal graph using at least one non-local graph convolutional neural network (NLGCN) to provide an identity of the person. Optionally, the at least one non-local neural network comprises at least one adaptive non-local neural network (ANLGCN) comprising an adaptive adjacency matrix trained to be responsive to data associated with anatomical landmarks of the plurality of anatomical landmarks that are not solely determined by a body structure of the person. Additionally or optionally, processing the spatio-temporal graph comprises segmenting the plurality of anatomical landmarks into a plurality of groups of anatomical landmarks, each group of anatomical landmarks characterized by a different configuration of degrees of freedom of motion. Optionally, the method comprises modeling the obtained spatio-temporal data associated with the anatomical landmarks in each group as a spatio-temporal graph. The processing can comprise processing the spatio-temporal graph modeling each of the plurality of groups of anatomical landmarks by at least one of the non-local networks of the at least one non-local graph convolutional neural network independently of processing the other groups of the plurality of groups to determine data indicative of the identity of the person. The method optionally comprises fusing the determined data from all groups to provide the identity of the person.
[0044] In one embodiment, obtaining the spatio-temporal data comprises obtaining a sequence of video frames imaging the person engaged in the activity, each video frame comprising an image of at least one body region of interest (BROI) imaging one of the plurality of anatomical landmarks. Optionally, the method comprises processing the video frames to detect the at least one body region of interest in each video frame. Additionally or optionally, the method comprises identifying an image of one of the plurality of anatomical landmarks in each of the at least one detected body region of interest. Optionally, the method comprises processing the identified image of the anatomical landmark to determine data defining the spatio-temporal trajectory.
[0045] In one embodiment, the plurality of anatomical landmarks comprises a joint. Optionally, the plurality of anatomical landmarks comprises a bone connecting the joint. Additionally or optionally, the joint comprises a finger joint. Optionally, the activity comprises a sequence of finger operations. The finger operations can comprise operations of engaging a keyboard.
[0046] In one embodiment, the joint comprises a joint of an upper extremity. Optionally, the activity is a sport. Optionally, the sport is football. Optionally, the sport is golf.
[0047] In one embodiment, the plurality of anatomical landmarks includes facial landmarks. Optionally, the facial landmarks include facial landmarks whose motion is used to define action units (AUs) of a facial action coding system (FACS) that is used to classify facial expressions and microexpressions. In one embodiment, the plurality of anatomical landmarks includes minutiae of fingerprints of a plurality of fingers of a hand.
[0048] According to one embodiment, there is also provided a system for identifying a person, the system comprising: an imaging system for acquiring a video having video frames that image a person attending an event; and software operable to process the video frames to provide an identity of the person in accordance with any of the preceding statements.
[0049] The description of embodiments of the application in this application is provided to enable any person skilled in the art to make or use the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without the use of the inventive faculty. Nothing in the application should be construed as requiring the inclusion of any particular feature in embodiments of the application. Various features of the described embodiments can be combined, substituted, or deleted, according to the application. The scope of the application is defined by the claims.
Claims
1. A method for identifying a person, characterized in that, The method includes: Spatiotemporal data of each of a plurality of anatomical landmarks associated with a human-participated activity are acquired, the spatiotemporal data providing data defining at least one spatiotemporal trajectory of the anatomical landmark during the activity; The acquired spatiotemporal data is modeled as a spatiotemporal graph (ST-Graph); and The spatiotemporal graph is processed using at least one nonlocal graph convolutional neural network (NLGCN) to provide the identity of the person, wherein processing the spatiotemporal graph includes segmenting the plurality of anatomical landmarks into multiple sets of anatomical landmarks, each set of anatomical landmarks being characterized by a different configuration of motion degrees of freedom.
2. The method as described in claim 1, characterized in that, The at least one nonlocal graph convolutional neural network includes at least one adaptive nonlocal graph convolutional neural network (ANLGCN), which includes an adaptive adjacency matrix trained to respond to data of relevant anatomical landmarks of the plurality of anatomical landmarks, the adaptive adjacency matrix being determined not only by the human body structure.
3. The method as described in claim 1, characterized in that, The method includes modeling the acquired spatiotemporal data associated with the anatomical landmarks in each group into a spatiotemporal graph.
4. The method as described in claim 3, characterized in that, The process includes: The spatiotemporal graph modeled for each of the plurality of anatomical landmarks is processed by one of the at least one nonlocal graph convolutional neural networks, independent of processing the other groups of the plurality of anatomical landmarks, to determine data representing the identity of the person.
5. The method as described in claim 4, characterized in that, The method includes fusing the determined data from all groups to provide the identity of the person.
6. The method according to any one of claims 1-5, characterized in that, Acquiring the spatiotemporal data includes acquiring a sequence of video frames that image people participating in the activity, each video frame including an image of at least one target body part (BROI) imaged on one of the plurality of anatomical landmarks.
7. The method as described in claim 6, characterized in that, The method includes processing the video frames to detect the at least one relevant body part in each video frame.
8. The method as described in claim 6, characterized in that, The method includes identifying an image of one of the plurality of anatomical landmarks in each of the at least one relevant body part being detected.
9. The method as described in claim 8, characterized in that, The method includes processing the images of the identified anatomical landmarks to determine the data that defines the spatiotemporal trajectory.
10. The method as described in claim 6, characterized in that, The multiple anatomical landmarks include joints.
11. The method as described in claim 10, characterized in that, The plurality of anatomical landmarks include bones that connect the joints.
12. The method as described in claim 10, characterized in that, The joints include finger joints.
13. The method as described in claim 12, characterized in that, The activity involves a series of finger movements.
14. The method as described in claim 13, characterized in that, The finger operations include operations involving keyboard manipulation.
15. The method as described in claim 10, characterized in that, The joints include those of the large appendages.
16. The method as described in claim 15, characterized in that, The activity in question is exercise.
17. The method as described in claim 16, characterized in that, The sport in question is football.
18. The method as described in claim 17, characterized in that, The sport in question is golf.
19. A method for identifying a person, characterized in that, The method includes: Spatiotemporal data of each of a plurality of anatomical landmarks associated with an activity in which a person participates are acquired, the spatiotemporal data providing data defining at least one spatiotemporal trajectory of the anatomical landmark during the activity, wherein the plurality of anatomical landmarks include facial landmarks; The acquired spatiotemporal data is modeled as a spatiotemporal graph (ST-Graph); and The spatiotemporal graph is processed using at least one nonlocal graph convolutional neural network (NLGCN) to provide the identity of the person.
20. The method as described in claim 19, characterized in that, The facial landmarks include facial landmarks whose motion is used to define action units (AUs) of a facial action coding system (FACS) for classifying facial expressions and micro-expressions.
21. The method as described in any one of the preceding claims, characterized in that, The multiple anatomical landmarks include detailed pairs of features of fingerprints from multiple fingers of the hand.
Citation Information
Patent Citations
Vision-based fitness guidance system
CN110909621A
Identity authentication method based on dynamic gestures
CN111444488A