Method and computer readable medium for identifying persors in video
The system uses a fully adaptive graph convolutional network to integrate spatial, temporal, and emotional data from video frames, addressing the challenges of diverse facial expressions and environments in video identity recognition, enhancing accuracy and reliability.
Patent Information
- Application Number
- CN202510498893.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-01
- Filing Date
- 2022-11-30
- Publication Date
- 2025-07-15
AI Technical Summary
When existing facial recognition technology recognizes people in videos, it is difficult to accurately identify individuals when facing complex environments and diverse facial expressions, especially in different ages, natural and virtual environments, and is susceptible to fake videos generated by artificial intelligence.
Using the STEM-IDR method, facial anatomical markers and emotional features in video frames are processed using a fully adaptive graph convolutional network (FAGC) to generate a unique human representation vector to identify people in the video.
It improves the accuracy and robustness of identifying individuals in videos, can recognize nuances and emotional changes in facial expressions, enhances the ability to recognize different environments, and resists the threat of fake videos.
Smart Images

Figure CN120318887A_ABST
Abstract
Description
[0001] This application is a divisional application of the application with the application number 202280078979.7, the application date of November 30, 2022, and the invention title "System and method for identifying a person in a video". Technical Field
[0002] Embodiments of the present disclosure relate to methods for processing facial images of a person to identify the person. Background Art
[0003] As human communication has shifted from being mediated through face-to-face interactions to being mediated through technology, more and more people may communicate and interact with other people and computers through video presentations. To support the functional operation of such communication and interaction, the ability to accurately identify the people appearing in the video is crucial, not only to know who is communicating and interacting with whom or with what computer application, but also to protect the communication and interaction from malicious impersonation and distortion. With the globalization and promotion of interactions between populations of different cultures and ethnicities, the extension of lifespan, and the sophistication and capabilities of communication and data processing technologies increasing, the complexity and difficulty involved in providing reliable and accurate identity recognition may intensify. Modern communication technologies and applications need to sustain an ever-increasing volume of communication, where people with different facial expressions can communicate via video, the same person can communicate via video at very different ages throughout his or her life, and people can appear in videos in different natural, artificial, and virtual environments. In addition, with the proliferation of artificial intelligence technologies with rapidly enhanced data processing and capabilities, reliable and trustworthy communication is threatened and may be used to generate videos in which impostors may communicate with targeted people and communities, disrupting and / or subconsciously influencing the behavior of targeted people and groups.
[0004] The environments, backgrounds, and conditions in which people may appear in videos are very different and are increasing in variety, which may challenge the ability of current facial recognition technologies to accurately identify these people. Summary of the Invention
[0005] Aspects disclosed herein relate to providing a system and method for extracting and melding spatial, temporal, and emotional data (STEM data) from video for identifying facial or other features of a person appearing in the video, and a fully adaptive graph convolutional network (FAGC) for processing data to identify a person. The method for extracting, configuring, and using STEM data to identify a person in a video may be referred to herein as a STEM identification method or STEM-IDR. STEM-IDR may be implemented as instructions that, when executed by at least one processor, cause the at least one processor to perform the extraction, configuration, and use of STEM data to identify a person in the video and / or generate a person representation vector that is unique to the person in the video.
[0006] According to one embodiment of the present disclosure, STEM-IDR may include extracting, from each of a plurality of consecutive frames in a video, features that identify and characterize anatomical facial landmarks of a person appearing in the frame, and features that characterize the emotional / psychological state that the person can be inferred to exhibit in the frame. In some embodiments, STEM-IDR may use the features that characterize anatomical facial landmarks to define a landmark feature vector (LF) of the frame, and use the features that characterize the emotional state to define an emotional feature vector (EF) of the frame. Although anatomical facial landmarks of a human face can generally be considered relatively prominent features of the human face, such as but not limited to the commissure of the lips or eyes, the endpoints of the eyebrows, the edges of the nostrils, and the bridge and tip of the nose, the implementation of embodiments of the present disclosure is not limited to such prominent features. For example, facial landmarks according to an embodiment may include features that characterize regions of facial skin texture, and transient features that appear, disappear, or change shape during a particular facial pose, such as wrinkles and dimples. In one embodiment, the landmarks may include vertices and / or edges of a polygonal facial mesh (e.g., a triangular mesh) that is constructed to model the surface of the face.
[0007] In some embodiments, STEM-IDR is operable to define a basis set “S” of emotion basis vectors B(s), 1 ≤ s ≤ S, for the frames of a video, which substantially spans the space of the emotion state vectors EF. Each emotion basis vector B(s) may be considered to represent a basis emotion state in which the person can be inferred to be in each of the consecutive frames of the video. For each frame, the emotion feature vector EF determined for the frame may be projected onto the emotion basis vectors to determine an emotion projection vector PB. Each component of the emotion projection vector may provide the probability that the person in the video is in a different emotion state among the S basis emotion states represented by the emotion basis vectors in the frame.
[0008] In some embodiments, STEM-IDR can use the determined probabilities to determine, for each first basic emotional state in the first frame of a video and each second basic emotional state in a second subsequent frame (which may not be consecutive frames) of the video, the transition probability of a person transitioning from the first basic emotional state to the second basic emotional state.
[0009] According to some embodiments, STEM-IDR can use the transition probabilities to generate a spatio-temporal emotional STEM data profile (STEM-DC) for the video, where the STEM-DC can be associated with each ordered pair of a plurality of basic emotional states defined as a global flag feature vector (GLFV). The global flag feature vector GLFV associated with a given ordered pair of basic emotional states can include a set of components, where each component is a function of at least one flag feature vector LF determined for a frame, and the at least one flag feature vector LF is weighted by the transition probability associated with the ordered pair of basic emotional states.
[0010] According to some embodiments, the STEM-DC can be processed into a graph by a fully adaptive graph convolutional network (FAGC). The nodes of the graph can be cells of the STEM-DC, each cell being related to a state transition, and the node features can be global flag vectors. For each of the multiple layers in the FAGC, the adjacency matrix of the nodes can be learned by the FAGC from data associated with each sample of the person that the FAGC intends to identify.
[0011] Consistent with some disclosed embodiments, a system, method, and computer-readable medium for identifying a person in a video are disclosed herein. The system, method, device, and non-transitory computer-readable medium can include at least one processor, which can be configured to generate a spatio-temporal emotional data profile (STEM-DC) from the video; and use a deep fully adaptive graph convolutional network (FAGC) to process the STEM-DC to determine a first person representation vector representing the person in the video.
[0012] In some embodiments, the at least one processor can also be configured to compare the first person representation vector with a subsequent second person representation vector determined from a subsequent video and a subsequent STEM-DC, and thus identify the person as appearing in the subsequent video when the first person representation vector and the second person representation vector are substantially similar.
[0013] In some embodiments, the first person representation vector and the second person representation vector can be based on recognizable traits of the person in the video. In some embodiments, the recognizable traits can include at least one of face, emotion, gait, body, limb, or typing style.
[0014] In some embodiments, the generation of STEM-DC may include generating an Iterative Feature Vector (IFV). In some embodiments, the generation of the IFV may include iterating a weighted series of landmark feature vectors by a function of a plurality of transition probabilities between a plurality of underlying emotional states of the person detected in a plurality of subsequent frames of the video. In some embodiments, the function of the plurality of transition probabilities may be represented by a Weighted Sum Matrix (WSUM).
[0015] In some embodiments, each of the plurality of underlying emotional states may be determined by projecting an emotional feature vector onto a series of emotion basis vectors. In some embodiments, each of the series of landmark feature vectors for a given facial image may include L landmarks characterized by P features. In some embodiments, the series of landmark feature vectors may be determined by processing a plurality of facial images extracted from the video using a pre-trained Facial Landmark Extraction Network (FLEN) to identify the L facial landmarks, where each facial landmark is characterized by P features.
[0016] In some embodiments, the plurality of facial images may be extracted from the video by locating and rectifying a plurality of images of the face of the person located in the video. In some embodiments, the FAGC may include a feature extraction module and a data merging module including a plurality of convolutional blocks. In some embodiments, for videos with a higher frame rate, the resolution of the underlying emotional states may be increased.
[0017] Consistent with some disclosed embodiments, a system, method, device, and non-transitory computer-readable medium may include at least one processor configured to identify a person in a video by: obtaining a video of video frames having a plurality of occurrences of the face of the person; processing each frame to determine an emotional feature vector (EF) and a landmark feature vector (LF) of the face in each frame; projecting the EF onto each emotion basis vector of a set of emotion basis vectors spanning the space of the EF vector to determine the probability that the person in the frame exhibits an underlying emotional state represented by the emotion basis vector; using the probabilities to determine a transition probability that the person transitions from a first underlying emotional state to a second underlying emotional state for each first underlying emotional state for which the person has a determined probability of being exhibited in a first frame of the video and each second underlying emotional state for which the person has a determined probability of being exhibited in a second consecutive frame of the video;
[0018] Use the LF vectors and the transition probabilities to determine a STEM data profile associated with each ordered pair of the base emotional states, where each component of a set of components of the base emotional states is a function of at least one signature feature vector LF determined for the video frame, and the at least one signature feature vector LF is weighted by the transition probabilities associated with the ordered pair;
[0019] And process the data profile into a graph using a fully adaptive graph convolutional network and generate a human representation vector.
[0020] This summary of the invention is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. It is understood that this summary of the invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The following describes non-limiting examples of the embodiments disclosed herein with reference to the drawings listed after this paragraph. The drawings and the description are intended to explain and clarify the embodiments disclosed herein and should not be considered limiting in any way. The same elements in different drawings may be represented by the same numbers. The elements in the drawings are not necessarily drawn to scale.
[0022] Figure 1 is a block diagram of an exemplary computing device for extracting and fusing STEM data from a video to identify the face or other recognizable aspects of a person appearing in the video according to some implementations;
[0023] Figure 2A A flowchart of a process for generating a set of emotion basis vectors for processing a video to determine the emotional state of a person in the video is shown according to some implementations, and the emotion basis vectors may be beneficial for identifying a person;
[0024] Figure 2B According to some implementations, it shows Figure 2A a schematic visual illustration of the flowchart shown in;
[0025] Figure 3A A flowchart of a process for processing a video to generate a STEM data profile is shown according to some implementations;
[0026] Figure 3B According to some implementations, it shows Figure 3A a schematic visual illustration of the actions for determining the emotion projection vector and the signature feature vector for each frame in the video in the flowchart of;
[0027] Figure 3CA schematic visual illustration of an array of transition probabilities determined for transitions between basic emotional states is shown according to some implementations;
[0028] Figure 3D A schematic visual illustration is shown according to some implementations Figure 3A of the actions in the flowchart shown, the actions being for iteratively flagging feature vectors of temporally adjacent frames using transition probabilities;
[0029] Figure 3E A STEM data profile configured to identify a three-dimensional tensor of a person is schematically shown according to some implementations;
[0030] Figure 4A A flowchart of a process for processing STEM data to determine a latent vector representing a person in a video using a fully adaptive graph convolutional network (FAGC) is shown according to some implementations; and
[0031] Figure 4B The processing of STEM data using a fully adaptive graph convolutional network is schematically shown according to some implementations. DETAILED DESCRIPTION
[0032] Non-limiting examples of implementations for identifying a person in a video shown in the accompanying drawings will now be described in detail. These examples are described below with reference to the accompanying drawings, in which like reference numerals refer to like elements. When similar reference numerals are shown, the corresponding description is not repeated, and the interested reader is referred to the previously discussed figures for the description of similar elements.
[0033] Aspects of the present disclosure can provide technical solutions to the challenging technical problem of identifying a person in a video and can relate to a system for extracting STEM data from a video for identifying a person appearing in the video, where the system has at least one processor (e.g., the processor, processing circuit, or other processing structure described herein), including methods, systems, devices, and computer-readable media. For ease of discussion, an exemplary method is described below, while understanding that aspects of the exemplary method are equally applicable to systems, devices, and computer-readable media. For example, some aspects of such a method can be implemented by a computing device or software running thereon. The computing device can include at least one processor (e.g., a CPU, GPU, DSP, FPGA, ASIC, or any circuit for performing logical operations on input data) for executing the exemplary method. Other aspects of such a method can be implemented via a network (e.g., a wired network, a wireless network, or both).
[0034] As another example, some aspects of such a method can be implemented as operations or program code in a non-transitory computer-readable medium. The operations or program code can be executed by at least one processor. As described herein, a non-transitory computer-readable medium can be implemented as any combination of hardware, firmware, software, or any medium capable of storing data that can be read by any computing device having a processor for executing a method or operation represented by the stored data. In the broadest sense, the exemplary method is not limited to a particular physical or electronic tool, but can be implemented using many different tools.
[0035] Figure 1 is a block diagram of an exemplary computing device 100 for extracting and fusing STEM data from a video for identifying the face or other recognizable aspects of a person appearing in the video, according to some implementations. The computing device 100 can include processing circuitry 110, such as a central processing unit (CPU). In some embodiments, the processing circuitry 110 can include or can be a component of a larger processing unit implemented with one or more processors. The one or more processors can be implemented with any combination of a general-purpose microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic device (PLD), a controller, a state machine, gated logic, discrete hardware components, a dedicated hardware finite state machine, or any other suitable entity capable of performing computational or other operations on information. The processing circuitry, such as processing circuitry 110, can be coupled to a memory 120 via a bus 105.
[0036] The memory 120 can further include a memory portion 122 that can contain instructions that, when executed by the processing circuitry 110, can perform the processes and methods described in more detail herein. The memory 120 can further be used as a working scratch pad, temporary memory, and other memory for the processing circuitry 110, as appropriate. The memory 120 can be a volatile memory, such as but not limited to random access memory (RAM), or a non-volatile memory (NVM), such as but not limited to flash memory. The processing circuitry 110 can further be connected to a network device 140, such as a network interface card, for providing a connection between the computing device 100 and a network (e.g., network 150). The processing circuitry 110 can further be coupled to a storage device 130. The storage device 130 can be used to store videos, video frames, data structures, and / or data elements associated with the data structures, or any other data structure. Although shown as a single device in Figure 1 it should be understood that the storage device 130 can include multiple collocated or distributed devices.
[0037] The processing circuit 110 and / or the memory 120 may also include a machine-readable medium for storing software. As used herein, "software" broadly refers to any type of instruction, whether referring to software, firmware, middleware, microcode, hardware description language, or others. The instructions may include code (e.g., source code format, binary code format, executable code format, or any other suitable code format). When the instructions are executed by one or more processors, they can cause the processing system to perform various functions described in further detail herein.
[0038] In some embodiments, the system 100 may include a human-machine and hardware interface 145. The human-machine interface device may include a screen, a keyboard, a touch interface, and / or a mouse. Non-limiting examples of the use of the interface 145 may include enabling a user of the system 100 to input data or parameters, or to indicate a video to be processed or a person to be identified, or to display the processed video, the identified person, or the processing status of STEM-IDR by the system 100, including the processes (200, 300, 400) described herein. The hardware interface may provide for transferring files such as a video to be processed to other components of the system 100. Video files may also be received or retrieved via the network 150.
[0039] Although the system 100 is presented herein as having specific components and modules, those skilled in the art should understand that the illustrated architectural configuration of the system 100 may be merely one possible configuration, and other configurations with more or fewer components are also possible. As described herein, the "components" of the system 100 may include Figure 1 one or more of the modules or services included within the system 100 as shown therein. Where it is said herein that STEM-IDR or the system 100 provides a specific function or performs an action, it should be understood that the function or action may be performed by the processing circuit 110 running software that may call other components of the system 100 and / or an external system (not shown).
[0040] Figure 2A A flowchart of a process 200 is shown according to an embodiment of the present disclosure, which can determine a set of emotion basis vectors for characterizing the emotional state of a person appearing in a video and identifying this person. Figure 2B A graphical illustration of the operation of the process 200 is shown according to an embodiment of the present disclosure, and the features and entities referred to in the flowchart are schematically shown.
[0041] A non - transitory computer - readable medium may include instructions that, when executed by at least one processor, cause the at least one processor to determine a set of emotion - based vectors for characterizing the emotional state of a person appearing in a video and for identifying the person, e.g., in process 200. Process 200 may be part of STEM - IDR. Process 200 may be executed, for example, by system 100 as described above. The non - transitory computer - readable medium may include instructions that perform the operations described at each step that is part of process 200 when executed by at least one processor. The at least one processor may correspond to system 100 and / or processing circuit 110.
[0042] In step 202, a deep neural network (DNN) Facial Emotion Classifier (FEMC) may be pre - trained to classify the emotional states inferable from facial images, so as to classify the emotional states expressed in N facial training images of different people. In some embodiments, the DNN may be used to classify such features based on images associated with other identifying features of different people. Non - limiting examples of such features may include gait or typing style or body identifier, etc. The embodiments and examples provided herein relate to facial recognition, but this should not be considered restrictive, and it should be understood that the systems described herein may be applied to identify other features associated with an individual for the purpose of identifying the individual. The FEMC is schematically shown in Figure 2B as classifier 222, which operates to classify the emotional state of the 4th image among the N facial training images 224 that the classifier can be used to classify.
[0043] In step 204, a set of H emotion features may be extracted from the same hidden layer as layer 226 of FEMC 222 (schematically shown in Figure 2B ) for each processed training image, and the training image may be used, according to embodiments, to define an embedded emotion feature vector EF (dimension ). For the nth (1 ≤ n ≤ N) facial expression training image, the extracted embedded emotion feature vector may be described by the following equation (1),
[0044]
[0045] Figure 2B Schematically shows the emotion feature vector labeled by the number 228 which may be extracted from layer 226 for the training image 4 processed by FEMC 222.
[0046] In some embodiments, in step 206, the N feature vectors EF(n) can be clustered in a clustering space 229, which is schematically represented by Figure 2B the small asterisks 231 in, which determines the N face training images as multiple optionally predetermined numbers S of clusters 230.
[0047] In step 208, the centroid vector 232 of dimension can optionally be determined for each cluster. It should be noted that the dimension G of the centroid vector 232 can be equal to or less than the dimension H of the embedded emotion feature vector EF(n). For example, if it is found that the centroid vector lies on a hyperplane of the space including the embedded emotion feature vector, then G can be less than H. Or it can be determined in response to principal component analysis. The number G of components less than H can advantageously be used to characterize the embedded emotion feature vector EF(n).
[0048] In some embodiments, in step 210, the S centroid vectors 232 can be defined as emotion basis vectors B(s), 1 ≤ s ≤ S, represented by Figure 2B the label 234 in, which has components b(s) g , 1 ≤ g ≤ G:
[0049]
[0050] According to embodiments of the present disclosure, by using features from the hidden layer to distinguish a person's emotional state and determine the emotion basis vectors B(s) across emotional states, STEM-IDR can have enhanced sensitivity to the nuances of emotional states, which can provide STEM-IDR with improved resolution for distinguishing the ways different people display the same or similar emotional states. The enhanced sensitivity and resolution can provide advantageous performance in identifying people in videos where people appear.
[0051] Figure 3A FIG. shows a flowchart of a process 300 for extracting STEM data from a video of a person to be identified and generating a STEM data compendium according to embodiments of the present disclosure. Figures 3B to 3E FIG. provides a schematic diagram of the operations of the process 300 according to one embodiment and schematically shows the features and entities involved in the process 300.
[0052] A non-transitory computer-readable medium may include instructions that, when executed by at least one processor, cause the at least one processor to define a STEM data profile from a video of a person to be identified (e.g., in process 300). Process 300 may be part of STEM-IDR. Process 300 may be performed, for example, by system 100 as described above. The non-transitory computer-readable medium may include instructions that perform the operations described at each step as part of process 300. The at least one processor may correspond to system 100 and / or processing circuitry 110.
[0053] In step 302, a video 320 including N frames F(n), 1 ≤ n ≤ N, may be received, as schematically shown in Figure 3B for processing to determine the identity of a person appearing in the video according to an embodiment of the present disclosure. In some embodiments, STEM-IDR including processes 200, 300, and 400 may identify a person using facial features or other identifying features. Non-limiting examples of such features may include gait or typing style or body identifiers, etc. The embodiments and examples provided herein relate to facial recognition, but this should not be considered limiting, and it should be understood that the systems described herein may be applied to identify other identifiable features associated with an individual to effect identification of the individual.
[0054] In some embodiments, a video 320 having a frame rate (fps) higher than 60 may be provided, and this higher number of frames per period may enable detection of micro-expressions (defined herein as facial expressions) for a very short period of less than half a second. Thus, in some embodiments, for videos including a higher frame rate, the resolution of the determined underlying emotional state may optionally be increased.
[0055] In step 304, an image of a person's face may be located in each frame and then the image may be corrected to a corrected face image FI(n) such that the person's face appears in all of the corrected face images FI(n) in a substantially same desired pose and size. In some embodiments, the FI(n) images may be corrected such that the person's face is in a full-face frontal pose and the images may be cropped to the same standard size.
[0056] In step 306, a pre-trained facial landmark extraction net (FLEN) (e.g., MobileNet-V1) can be optionally used to process each of the facial images FI(n) to identify L facial landmarks in FI(n), where each facial landmark is characterized by P features from the facial image. In addition to the P features extracted by the FLEN for each landmark identified by the FLEN, each facial image FI(n) of a given video frame F(n) can also be associated with "V" additional spatio-temporal features that the FLEN may not extract. In one embodiment, the additional features can include features determined in response to the localization and correction processes mentioned in step 304. For example, the V additional features can include features from the hidden layer of the FLEN and / or features in response to the position and pose of the face in the frame and / or the change in pose and / or position relative to temporally adjacent frames. For a given facial image FI(n), a landmark feature vector LF(n) having L landmarks characterized by P features can have a dimension where F = LxP + V. Assume lf(n) f represents the f-th feature of LF(n), and LF(n) can be explicitly represented as:
[0057] LF(n) = {lf(n)1, lf(n)2,..., lf(n) F} = {lf(n) f | 1 ≤ f ≤ F}. (3)
[0058] For the sake of narrative convenience and efficiency, the FLEN is considered to provide V and P features for characterizing one landmark.
[0059] Optionally, in step 308, the FEMC for defining the emotion basis vector B(s) ( Figure 2A referred to in Figure 2B and schematically shown as FEMC 222 in Figure 2B vector 234 in can be used to extract an emotion feature vector EF(n) for each facial image FI(n)
[0060] where H EF(n) = {ef(n)1, ef(n)2,..., ef(n) h} = {ef(n)
[0061] | 1 ≤ h ≤ H}. (4) Figure 3B For example, Figure 3B schematically shows that the deep neural network DNN-FEMC 222 extracts the embedded emotion vector EF(1) 322 from the facial image in frame F(1).Embodiments according to the present disclosure also show the DNN FLEN 324 that uses FLEN to extract the landmark feature vector LF(1) 326 from the source image F(1). In Figure 3B , the landmark feature vectors LF(n) 331 and LF(n+1) 334 respectively extracted from the facial images FI(n) and FI(n+1) by FLEN are also schematically shown.
[0062] Optionally, in step 310, each embedded emotion vector EF(n) extracted from the facial image FI(n) can be projected onto S basis vectors B(s) ( Figure 2A , Figure 2B ) to generate an emotion projection vector PB(n) with components pB(n) s for each facial image, where 1 ≤ s ≤ S. The emotion projection vector can be explicitly written as:
[0063] PB(n) = {pB(n)1, pB(n)2,..., pB(n) S} = pB(n) s |1 ≤ s ≤ S}. (5)
[0064] In some embodiments, the projection component pB(n) s can be determined by the Euclidean distance d(EF(n), B(s)) between the embedded emotion feature vector EF(n) and the emotion basis vector B(s). Let d(EF(n), B(s)) denote the Euclidean distance between EF(n) and B(s), and pB(n) s can be determined according to Equation (6),
[0065] pB(n) s = [1 / d(EF(n), B(s))] / ∑ s 1 / (d(EF(n), B(s)). (6)
[0066] For the case where the dimension G of the emotion basis vector B(s) is less than the dimension of the embedded emotion feature vector EF(n), the components of the vector EF(n) can be restricted to the components located in the space of the emotion basis vector B(s).
[0067] Figure 3B The emotion projection vectors PB(1) 329 of PB(n) and PB(n+1) generated from EF(1) by the projection operator 325 and generated for the facial images FI(n) and FI(n+1) respectively provided by frames F(n) and F(n+1) are schematically shown. It should be noted that the projection component pB(n) sis normalized such that the sum Σ s [pB(n) s = 1, and the given component pB(n) s determined by the embedded emotion feature vector EF(n) can be considered as the probability that the person in the facial image FI(n) is in the basic emotion state represented by the emotion basis vector B(s). For the sake of description, the basic emotion state represented by the emotion basis vector B(s), 1 ≤ s ≤ S, can be referred to as the basic emotion state BE(s).
[0068] In step 312, the emotion state transition probability can be determined for each facial image FI(n) (1 ≤ n ≤ N - 1), where the emotion state transition probability is for a person with probability pB(n) j being in the basic emotion state BE(j) (1 ≤ j ≤ S) to transition to the subsequent basic emotion state BE(k) in the subsequent facial image FI(n + 1), where the person has probability pB(n + 1) k being in the basic emotion state BE(k) in the subsequent facial image FI(n + 1). In some embodiments, let the transition probability be represented by the transition weight w(n, n + 1) j,k . In one embodiment, the transition probability w(n, n + 1) j,k can be determined according to Equation (7),
[0069] w(n, n + 1) j,k = pB(n) j · pB(n + 1) k . (7)
[0070] If a set of transition probabilities w(n, n + 1) j,k for the facial images FI(n) and FI(n + 1) is represented by the transition weight matrix W(n, n + 1), then
[0071] W(n, n + 1) = {w(n, n + 1) j,k |1 ≤ j, k ≤ S}. (8)
[0072] It should be noted that generally W(n, n + 1) is not a symmetric matrix, and w(n, n + 1) j,k ≠ w(n, n + 1) k,j .
[0073] Figure 3C According to an embodiment of the present disclosure, the transition weight matrix W(n, n + 1)340 determined by the emotion projection vectors PB(n) and PB(n + 1) is schematically shown.
[0074] Optionally, in step 314, the transition weight sum matrix WSUM(n + 1) can be determined according to Equation (9):
[0075]
[0076] where
[0077]
[0078] In step 316, the landmark feature vector LF(n) can be iterated over n from n = 1 to n = (N - 1) to generate an iterative feature vector IFV(N) for each transition from a given basic emotional state BE(j) (1 ≤ j ≤ S) to a given basic emotional state BE(k) (1 ≤ k ≤ S) for the video j,k,f . The iterative feature vector IFV(N) for an ordered pair of indices (1 ≤ j, k ≤ S) j,k,f can be a vector of dimension , that is, the sum of the vector LF(n) weighted by a function of the transition probability w(n, n + 1) j,k . In one embodiment, for the nth iteration (n = 1 → N - 1) and given indices j, k,
[0079]
[0080] where
[0081]
[0082] For the last (N - 1)th iteration,
[0083]
[0084] and
[0085] IFV(N) j,k,f = (wsum(N - 1) j,k / wsum(N) j,k ) IFV(N - 1) j,k,f + (w(N - 1, N) j,k / wsum(N) j,k ).lf(N) f . (14)
[0086] Note that IFV(1) j,k,f is not defined by Equations (11)-(14). In one embodiment, IFV(1) j,k,fIt can optionally be determined according to the following formula
[0087] IFV(1) j,k,f =[wsum(N) j,k / (N - 1)].lf(1) f or = lf(1) f (15)
[0088] IFV(N) j,k can be referred to as the Global Flag Feature Vector (GLFV), which is used for the transition of a video (e.g., Figure 3B video 320) from the basic emotional state BE(j) to the basic emotional state BE(k). Figure 3D Schematically shows the iterative feature vector IFV(n + 1) 350 determined by the nth iteration of the ordered index pair j,k.
[0089] In some embodiments, in step 318, the spatio - temporal emotion STEM data profile can be defined as a three - dimensional tensor STEM - DC:
[0090]
[0091] Figure 3E Schematically shows the three - dimensional STEM - DC 360 according to an embodiment of the present disclosure. STEM - DC 360 can also be referred to as the encoded state transition matrix herein.
[0092] Figure 4A Shows a flowchart of a process 400 for using a fully adaptive graph convolutional network (FAGC) to process STEM - DC to determine a latent vector representing a person in a video according to some implementations. Figure 4B Schematically shows a process 400 including the processing of STEM - DC using a fully adaptive graph convolutional network FAGC according to some implementations. A non - transitory computer - readable medium may contain instructions that, when executed by at least one processor, cause the at least one processor to process STEM - DC using a deep FAGC (e.g., in process 400). Process 400 can be part of STEM - IDR. Process 400 can be executed, for example, by the system 100 as described above. A non - transitory computer - readable medium may contain instructions that perform the operations described at each step as part of process 400 when executed by at least one processor. The at least one processor may correspond to system 100 and / or processing circuit 110.
[0093] In step 402, a STEM-DC can be generated, for example, by using the above-described process 300. In step 402, according to an embodiment of the present disclosure, the STEM-DC 360 can be processed by a deep FAGC, herein referred to as the STEM-DEEP network 500, or simply as STEM-DEEP 500, to determine a latent vector that represents the person in the video 320 of the generated STEM-DC ( Figure 3B ). In some embodiments, the STEM-DEEP 500 may include a feature extraction module 502, followed by a data merging module 504, which can generate a latent vector, optionally referred to as a STEM person representation vector (STEM-PR), and is schematically shown as PR 510 in Figure 4B . In some embodiments, in step 404, a first PR 510 determined for a specific individual from the provided video may be compared with a subsequent second PR 510 determined from a subsequent video (by repeating steps 402 and 404), so as to confirm the identity of the individual appearing in the subsequent video when the subsequent second PR 510 is substantially similar to the determined first PR510.
[0094] In one embodiment, the feature extraction module 502 may include a plurality (optionally five) of FAGC blocks 503. The data merging module 504 can provide an output to a fully connected (FC) layer 506 that generates the PR 510. In some embodiments, the data merging module 504 may include a plurality of convolutional blocks 505, and the convolutional blocks 505 may be two-dimensional (2D), three-dimensional (3D), or other suitable structures. In one embodiment, each FAGC block 503 may include a data driven attention tensor A att (not shown) and a learned adjacency matrix A adj (not shown), each having a dimension If the input feature to the FAGC block 503 is represented by InF in and the output feature provided by the FAGC block 503 in response to InF in is represented by OtF out , the operation of a given FAGC block 503 can be represented by the following formula,
[0095] OtF out = W(OtF out (A att + A adj )) (17)
[0096] where W represents the learned weight of the FAGC block 503.
[0097] In one embodiment, STEM-DEEP 500 can be trained on a video training set including multiple videos obtained for each of multiple different people using a triplet margin loss (TML) function and an appropriate distance metric. For each video in the video training set, STEM-DEEP 500 can be used to generate and process STEM-DC 360 to generate a STEM-PR 510 person representation vector. The TML operation provides a favorable distance based on the metric between the STEM-PR 510 vectors of different people in the training set.
[0098] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The materials, methods, and examples provided herein are illustrative only and not restrictive.
[0099] As used herein, the terms "convolutional network", "neural network", "machine learning", or "artificial intelligence" refer to using algorithms on a computing device to parse data, learn from the data, and then make a determination or generate data, where the determined or generated data is not deterministically replicated (e.g., using deterministic software known in the art).
[0100] Implementations of the methods and systems of the present disclosure may involve performing or completing certain selected tasks or steps manually, automatically, or a combination thereof. Additionally, according to the preferred embodiments of the methods and systems of the present disclosure, actual instruments and devices may implement several selected steps through hardware (HW) or software (SW) on any operating system of any firmware, or a combination thereof. For example, as hardware, selected steps of the present disclosure may be implemented as a processor chip or circuit. As software or an algorithm, selected steps of the present disclosure may be implemented as multiple software instructions executed by a computer / processor using any suitable operating system. In any case, selected steps of the methods and systems of the present disclosure may be described as being executed by a data processor, such as a computing device for executing multiple instructions.
[0101] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuits, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which can be special purpose or general purpose, coupled to receive data and instructions from, and to send data and instructions to, a storage system, at least one input device, and at least one output device.
[0102] Although the present disclosure is described in terms of "computing device", "computer", or "mobile device", it should be noted that optionally, any device having a data processor and the ability to execute one or more instructions can be described as a computing device, including but not limited to any type of personal computer (PC), server, distributed server, virtual server, cloud computing platform, cellular phone, IP phone, smart phone, smart watch, or PDA (personal digital assistant). Any two or more such devices that communicate with each other can optionally form a "network" or "computer network".
[0103] To provide interaction with a user, the systems and techniques described herein can be implemented on a computing device having a display (indicator / monitor / screen / array) (e.g., LED (light emitting diode), OLED (organic LED), LCD (liquid crystal display), or other display technology) for presenting information to the user, as well as a keyboard and pointing device (e.g., mouse, joystick, or trackball) or separate buttons / knobs / levers (e.g., drive wheel buttons / shift levers) by which the user can provide input to the computing device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including analysis of acoustic, speech, user head position and / or eye movement, or tactile input.
[0104] It should be understood that the above methods and devices can vary in many ways, including omitting or adding steps, changing the order of steps, and the types of devices used. It should be understood that different features can be combined in different ways. In particular, in each embodiment or implementation of the present disclosure, not all of the above features in a particular embodiment or implementation are necessary. Further combinations of the above features and implementations are also considered to be within the scope of some embodiments or implementations of the present disclosure.
[0105] Although certain features of the described embodiments have been illustrated as described herein, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. It should be understood that they are presented by way of example only and not limitation, and various changes can be made in form and detail. Any part of the apparatus and / or method described herein can be combined in any combination, except mutually exclusive combinations. The implementations described herein can include various combinations and / or sub - combinations of the functions, components, and / or features of the different implementations and embodiments described.
Claims
1. A method for identifying a person in a video, characterized in that: The method includes: via a computing device, generating a spatio-temporal emotion data summary (STEM-DC) from the video, wherein generating the STEM-DC includes generating an iterative feature vector (IFV), and generating the IFV includes iterating over a series of landmark feature vectors, the series of landmark feature vectors being weighted by a function of a plurality of transition probabilities between a plurality of basic emotion states of the person detected in a plurality of subsequent frames of the video; and using a deep fully adaptive graph convolutional network (FAGC) to process the STEM-DC to determine a first person representation vector representing the person in the video.
2. The method according to claim 1, characterized in that: The method further includes comparing the first person representation vector with a subsequent second person representation vector determined from a subsequent video and subsequent STEM-DC, thereby identifying the person as appearing in the subsequent video when the first person representation vector and the second person representation vector are substantially similar.
3. The method according to claim 2, wherein: The first person representation vector and the second person representation vector are based on recognizable characteristics of the person in the video.
4. The method according to claim 3, wherein: The recognizable characteristics include at least one of face, emotion, gait, body, limb, or typing style.
5. The method according to claim 1, wherein: The function of the plurality of transition probabilities is represented by a transition weight and a matrix.
6. The method according to claim 1, wherein: Each of the plurality of basic emotion states is determined by projecting an emotion feature vector onto a series of emotion basis vectors.
7. The method according to claim 1, wherein: Each of the series of landmark feature vectors for a given facial image includes L landmarks characterized by P features.
8. The method according to claim 1, characterized in that: The FAGC includes a feature extraction module and a data merging module including a plurality of convolutional blocks.
9. The method according to claim 1, characterized in that: For videos with a higher frame rate, increase the resolution of the basic emotion states.
10. The method according to claim 9, wherein: The series of landmark feature vectors are determined by processing a plurality of facial images extracted from the video using a pre-trained facial landmark extraction network (FLEN) to identify L facial landmarks, where each facial landmark is characterized by P features.
11. The method according to claim 10, characterized in that: The plurality of facial images are extracted from the video by locating and rectifying a plurality of images of the face of the person located in the video.
12. A non-transitory computer-readable medium comprising a plurality of instructions that, when executed by at least one processor, cause the at least one processor to perform a plurality of operations for identifying a person in a video, characterized in that: The plurality of operations include: generating a spatio-temporal emotion data summary (STEM-DC) from the video, wherein generating the STEM-DC includes generating an iterative feature vector (IFV), and generating the IFV includes iterating over a series of landmark feature vectors, the series of landmark feature vectors being weighted by a function of a plurality of transition probabilities between a plurality of basic emotion states of the person detected in a plurality of subsequent frames of the video; and using a deep fully adaptive graph convolutional network (FAGC) to process the STEM-DC to determine a first person representation vector representing the person in the video.
13. The non-transitory computer-readable medium according to claim 12, wherein: The multiple operations include: comparing the first person representation vector with a subsequent second person representation vector determined from a subsequent video and a subsequent STEM-DC, thereby identifying the person as appearing in the subsequent video when the first person representation vector and the second person representation vector are substantially similar.
14. The non-transitory computer-readable medium according to claim 13, wherein: The first person representation vector and the second person representation vector are based on recognizable characteristics of the person in the video.
15. The non-transitory computer-readable medium according to claim 14, wherein: The recognizable characteristics include at least one of face, emotion, gait, body, limb, or typing style.
16. The non-transitory computer-readable medium according to claim 12, wherein: The function of the multiple transition probabilities is represented by a transition weight and matrix (WSUM).
17. The non-transitory computer-readable medium according to claim 12, wherein: Each of the multiple basic emotional states is determined by projecting an emotion feature vector onto a series of emotion basis vectors.
18. The non-transitory computer-readable medium according to claim 12, wherein: Each of the series of landmark feature vectors for a given facial image includes L landmarks characterized by P features.
19. The non-transitory computer-readable medium according to claim 12, wherein: The FAGC includes a feature extraction module and a data merging module including a plurality of convolutional blocks.
20. The non-transitory computer-readable medium according to claim 12, wherein: For videos with a higher frame rate, increase the resolution of the basic emotional state.