System and method for identifying people in a video
Through the STEM-IDR system and method, STEM data and FAGC network are used to solve the complexity of face recognition in videos, and accurate identification and identity identification in a diverse environment are achieved, reducing the risk of counterfeiting.
Patent Information
- Application Number
- CN202280078979.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-01
- Filing Date
- 2022-11-30
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-11-30
AI Technical Summary
The prior art is difficult to accurately identify faces in videos, especially in situations where environment, background and mood changes vary, and face threats of fakes and identity forgery.
Using a system and method, called the STEM Identity Identification Method (STEM-IDR), this data is processed by extracting and fusing spatial, temporal, and emotional data (STEM data) from the video and using a fully adaptive graph convolutional network (FAGC) to generate a unique vector representing the face or other features of a person in the video.
Accurate facial recognition in complex environments and diverse situations is achieved, the ability to identify people's identities is enhanced, and the risks of counterfeiting and identity forgery are reduced.
Smart Images

Figure CN118318236B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 284,643, filed on December 1, 2021, which is expressly incorporated herein by reference in its entirety. Technical Field
[0003] Embodiments of the present disclosure relate to methods for processing facial images of persons to recognize persons. Background Art
[0004] As human communication shifts from being mediated through face-to-face interactions to being mediated through technology, more and more people may communicate and interact with other people and computers through video presentations. To support the functional operation of such communications and interactions, the ability to accurately identify people appearing in videos is critical, not only to know who is communicating and interacting with whom or with what computer application, but also to protect communications and interactions from malicious impersonations and distortions. The complexity and difficulties involved in providing reliable and accurate identification are likely to increase as globalization integrates and promotes interactions between populations of different cultures and ethnicities, life spans increase, and the sophistication and capabilities of communications and data processing technologies increase. Modern communications technologies and applications are needed to sustain an ever-increasing volume of communications, where people with different facial expressions can communicate via video, the same person can communicate via video at very different ages throughout his or her lifetime, and people can appear in videos in different natural, artificial, and virtual environments. Furthermore, with the proliferation of artificial intelligence technologies with rapidly increasing data processing and capabilities, reliable and trustworthy communications are threatened and may be used to generate videos in which impersonators may communicate with, disrupt, and / or subconsciously influence the behavior of targeted people and communities.
[0005] The circumstances, backgrounds, and conditions in which people may appear in videos are very different and increasing, which may challenge the ability of current facial recognition technology to accurately identify these people. Summary of the invention
[0006] Aspects disclosed herein relate to providing a system and method for extracting and melding spatial, temporal, and emotional data (STEM data) from a video for use in identifying facial or other features of a person appearing in the video, and a fully adaptive graph convolutional network (FAGC) for processing the data to identify a person. The method for extracting, configuring, and using STEM data to identify a person in a video may be referred to herein as a STEM identity recognition method or STEM-IDR. STEM-IDR may be implemented as instructions that, when executed by at least one processor, cause the at least one processor to perform extraction, configuration, and use of STEM data to identify a person in a video and / or generate a person representation vector that is unique to the person in the video.
[0007] According to one embodiment of the present disclosure, STEM-IDR may include extracting features of anatomical facial landmarks that identify and characterize a person appearing in the frame and features that characterize an emotional / psychological state that the person may be inferred to exhibit in the frame from each of a plurality of consecutive frames in a video. In some embodiments, STEM-IDR may define a landmark feature vector (LF) of a frame using features that characterize anatomical facial landmarks, and define an emotional feature vector (EF) of a frame using features that characterize emotional states. Although anatomical facial landmarks of a human face may generally be considered relatively significant features of a human face, such as, but not limited to, the commissure of lips or eyes, the endpoints of eyebrows, the edges of nostrils, and the bridge and tip of the nose, implementation of embodiments of the present disclosure is not limited to such significant features. For example, facial landmarks according to an embodiment may include features that characterize areas of facial skin texture, and transient features that appear, disappear, or change shape during specific facial gestures, such as wrinkles and dimples. In one embodiment, a landmark may include vertices and / or edges of a polygonal facial mesh (e.g., a triangular mesh) that is constructed to model the surface of a face.
[0008] In some embodiments, STEM-IDR is operable to define a basis set "S" of emotion basis vectors B(s) for frames of a video, 1≤s≤S, which substantially spans the space of emotion state vectors EF. Each emotion basis vector B(s) can be considered to represent a basic emotional state that a person can be inferred to be in each of the consecutive frames of the video. For each frame, the emotion feature vector EF determined for the frame can be projected onto the emotion basis vector to determine an emotion projection vector PB. Each component of the emotion projection vector can provide the probability that a person in the video is in a different emotional state in the frame among the S basic emotional states represented by the emotion basis vector.
[0009] In some embodiments, STEM-IDR can use the determined probabilities to determine, for each first base emotional state in a first frame of a video and each second base emotional state in a second subsequent frame (which may not be consecutive frames) of the video, a transition probability of a person switching from a first base emotional state to a second base emotional state.
[0010] According to some embodiments, STEM-IDR may use transition probabilities to generate a spatiotemporal emotional STEM data summary (STEM-DC) for a video, which may be associated with each ordered pair of multiple basis emotional states defined as a global landmark feature vector (GLFV). The global landmark feature vector GLFV associated with a given ordered pair of basis emotional states may include a set of components, where each component is a function of at least one landmark feature vector LF determined for a frame, wherein the at least one landmark feature vector LF is weighted by the transition probabilities associated with the ordered pair of basis emotional states.
[0011] According to some embodiments, STEM-DC can be processed into a graph by a fully adaptive graph convolutional network (FAGC). The nodes of the graph can be cells of STEM-DC, each cell is related to a state transition, and the node features can be global flag vectors. For each of the multiple layers in the FAGC, the adjacency matrix of the node can be learned by the FAGC from the data associated with each sample of the person that the FAGC intends to recognize.
[0012] In accordance with some disclosed embodiments, systems, methods, and computer-readable media for identifying people in a video are disclosed herein. The systems, methods, devices, and non-transitory computer-readable media may include at least one processor, which may be configured to generate a spatiotemporal emotion data summary (STEM-DC) from a video; and process the STEM-DC using a deep fully adaptive graph convolutional network (FAGC) to determine a first person representation vector representing a person in the video.
[0013] In some embodiments, the at least one processor may be further configured to compare the first person representation vector with a subsequent second person representation vector determined from a subsequent video and a subsequent STEM-DC, thereby identifying the person as appearing in the subsequent video when the first person representation vector and the second person representation vector are substantially similar.
[0014] In some embodiments, the first person representation vector and the second person representation vector may be based on identifiable traits of people in the video. In some embodiments, the identifiable traits may include at least one of a face, an emotion, a gait, a body, a limb, or a typing style.
[0015] In some embodiments, generating the STEM-DC may include generating an iterated feature vector (IFV). In some embodiments, generating the IFV may include iterating a weighted series of signature feature vectors by a function of a plurality of transition probabilities between a plurality of base emotional states of the person detected in a plurality of subsequent frames of the video. In some embodiments, the function of the plurality of transition probabilities may be represented by a transition weight sum matrix (WSUM).
[0016] In some embodiments, each of the plurality of basis emotional states may be determined by projecting the emotion feature vector onto a series of emotion basis vectors. In some embodiments, each of the series of landmark feature vectors for a given facial image may include L landmarks represented by P features. In some embodiments, the series of landmark feature vectors may be determined by processing a plurality of facial images extracted from the video using a pre-trained facial landmark extraction network (FLEN) to identify the L facial landmarks, each of which is represented by P features.
[0017] In some embodiments, the plurality of facial images may be extracted from the video by locating and correcting a plurality of images of the faces of people in the video. In some embodiments, the FAGC may include a feature extraction module and a data merging module including a plurality of convolution blocks. In some embodiments, for videos with a higher frame rate, the resolution of the basic emotional state may be increased.
[0018] Consistent with some disclosed embodiments, systems, methods, devices, and non-transitory computer-readable media may include at least one processor that may be configured to identify a person in a video by: acquiring a video having a plurality of video frames in which a person's face appears; processing each frame to determine an emotion feature vector (EF) and a facial landmark feature vector (LF) for the face in each frame; projecting the EF onto each emotion basis vector in a set of emotion basis vectors spanning a space of the EF vectors to determine a probability that the person in the frame displays a basis emotion state represented by the emotion basis vector; using the probabilities to determine, for each first basis emotion state that the person has a determined probability of being displayed in a first frame of the video and each second basis emotion state that the person has a determined probability of being displayed in a second consecutive frame of the video, a transition probability of the person transitioning from the first basis emotion state to the second basis emotion state;
[0019] determining a STEM data profile associated with each ordered pair of the base emotional states using the LF vectors and the transition probabilities, each component of a set of components of the base emotional states being a function of at least one landmark feature vector LF determined for the video frame, the at least one landmark feature vector LF being weighted by the transition probabilities associated with the ordered pairs;
[0020] And use a fully adaptive graph convolutional network to process the data summary into a graph and produce a person representation vector.
[0021] The provision of this summary is to introduce in a simplified form a selection of concepts further described in the detailed description below. It is understood that this summary is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. The details of one or more embodiments are set forth in the accompanying drawings and the following description. Other features will be apparent from the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Non-limiting examples of the embodiments disclosed herein are described below with reference to the drawings listed after this paragraph. The drawings and descriptions are intended to explain and illustrate the embodiments disclosed herein and should not be considered to be limiting in any way. The same elements in different drawings may be represented by the same numbers. The elements in the drawings are not necessarily drawn to scale.
[0023] Figure 1 is a block diagram of an exemplary computing device for extracting and fusing STEM data from a video to recognize faces or other identifiable aspects of people appearing in the video, according to some implementations;
[0024] Figure 2A A flowchart showing a process for generating a set of emotion basis vectors for processing a video to determine the emotional states of a person in the video, which emotion basis vectors may be useful for identifying a person, according to some implementations;
[0025] Figure 2B According to some implementations, Figure 2A A schematic visual representation of the flow chart shown in ;
[0026] Figure 3A A flowchart showing a process for processing a video to generate a STEM data summary according to some implementations;
[0027] Figure 3B According to some implementations, Figure 3A A schematic visual illustration of the actions for determining the emotion projection vector and the signature feature vector for each frame in the video in the flowchart of;
[0028] Figure 3CA schematic visual illustration of an array of transition probabilities determined for transitions between underlying emotional states is shown according to some implementations;
[0029] Figure 3D According to some implementations, Figure 3A A schematic visual illustration of the actions in the flowchart shown for iterating landmark feature vectors of temporally adjacent frames using transition probabilities;
[0030] Figure 3E schematically illustrates a STEM data summary of a three-dimensional tensor configured for identifying a person according to some implementations;
[0031] Figure 4A A flowchart illustrating a process for processing STEM data summaries using a fully adaptive graph convolutional network (FAGC) to determine latent vectors representing people in a video according to some implementations; and
[0032] Figure 4B Processing of STEM data summarization using a fully adaptive graph convolutional network is schematically illustrated according to some implementations. DETAILED DESCRIPTION
[0033] Reference will now be made in detail to non-limiting examples of implementations for identifying people in videos shown in the accompanying drawings. These examples are described below with reference to the accompanying drawings, wherein like reference numerals refer to like elements. When similar reference numerals are shown, the corresponding description is not repeated, and the interested reader is referred to the previously discussed figures for descriptions of similar elements.
[0034] Aspects of the present disclosure may provide technical solutions to the challenging technical problems of identifying people in videos, and may relate to a system for extracting STEM data from a video for identifying people appearing in the video, wherein the system has at least one processor (e.g., a processor, processing circuit, or other processing structure described herein), including methods, systems, devices, and computer-readable media. For ease of discussion, an exemplary method is described below, with the understanding that aspects of the exemplary method are also applicable to systems, devices, and computer-readable media. For example, some aspects of this method may be implemented by a computing device or software running thereon. The computing device may include at least one processor for performing the exemplary method (e.g., a CPU, GPU, DSP, FPGA, ASIC, or any circuit for performing logical operations on input data). Other aspects of this method may be implemented over a network (e.g., a wired network, a wireless network, or both).
[0035] As another example, some aspects of such methods may be implemented as operations or program codes in a non-transitory computer-readable medium. The operations or program codes may be performed by at least one processor. As described herein, non-transitory computer-readable media may be implemented as any combination of hardware, firmware, software, or any medium capable of storing data, which may be read by any computing device having a processor for performing the method or operation represented by the stored data. In the broadest sense, the exemplary methods are not limited to specific physical or electronic tools, but may be implemented using many different tools.
[0036] Figure 1 1 is a block diagram of an exemplary computing device 100 for extracting and fusing STEM data from a video for identifying faces or other identifiable aspects of people appearing in the video, according to some implementations. The computing device 100 may include a processing circuit 110, such as a central processing unit (CPU). In some embodiments, the processing circuit 110 may include or may be a component of a larger processing unit implemented with one or more processors. The one or more processors may be implemented with any combination of a general-purpose microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a state machine, a gated logic, a discrete hardware component, a dedicated hardware finite state machine, or any other suitable entity that can perform calculations or other operations on information. A processing circuit such as the processing circuit 110 may be coupled to a memory 120 via a bus 105.
[0037] Memory 120 may further include a memory portion 122, which may contain instructions that, when executed by processing circuit 110, may perform the processes and methods described in more detail herein. Memory 120 may further serve as a working scratch pad, temporary storage, and other storage for processing circuit 110, as appropriate. Memory 120 may be a volatile memory, such as, but not limited to, random access memory (RAM) or non-volatile memory (NVM), such as, but not limited to, flash memory. Processing circuit 110 may further be connected to a network device 140, such as a network interface card, for providing a connection between computing device 100 and a network (e.g., network 150). Processing circuit 110 may be further coupled to storage device 130. Storage device 130 may be used to store video, video frames, data structures and / or data elements associated with data structures, or any other data structure. Although in Figure 1 1 is shown as a single device, but it should be understood that storage device 130 may include multiple devices that are collocated or distributed.
[0038] Processing circuitry 110 and / or memory 120 may also include a machine-readable medium for storing software. As used herein, "software" refers broadly to any type of instructions, whether software, firmware, middleware, microcode, hardware description language, or otherwise. Instructions may include code (e.g., in source code format, binary code format, executable code format, or any other suitable code format). When the instructions are executed by one or more processors, the processing system may be caused to perform various functions described in further detail herein.
[0039] In some embodiments, the system 100 may include a human-machine and hardware interface 145. Human-machine interface devices may include a screen, a keyboard, a touch interface, and / or a mouse. Non-limiting examples of the use of the interface 145 may include being used to enable a user of the system 100 to input data or parameters, or to indicate a video to be processed or a person to be identified, or to display the processing status of a processed video, an identified person, or a STEM-IDR by the system 100, including the processes (200, 300, 400) described herein. The hardware interface may provide for transferring files such as videos to be processed to other components of the system 100. Video files may also be received or retrieved via the network 150.
[0040] Although system 100 is presented herein as having specific components and modules, those skilled in the art will appreciate that the illustrated architectural configuration of system 100 may be only one possible configuration, and that other configurations with more or fewer components are possible. As described herein, “components” of system 100 may include Figure 1 1 and 12, and may be included in the system 100. Where it is said that the STEM-IDR or system 100 provides a particular function or performs an action, it should be understood that the function or action may be performed by the processing circuit 110 running software that may call other components of the system 100 and / or an external system (not shown).
[0041] Figure 2A A flow chart of a process 200 that can determine a set of emotion basis vectors for characterizing the emotional state of a person appearing in a video and identifying the person is shown according to an embodiment of the present disclosure. Figure 2B A graphical illustration of the operation of process 200 is shown in accordance with an embodiment of the present disclosure, and the features and entities referred to in the flowchart are schematically illustrated.
[0042] A non-transitory computer-readable medium may contain instructions that, when executed by at least one processor, cause the at least one processor to determine a set of emotion basis vectors for characterizing the emotional state of a person appearing in a video and identifying the person, such as in process 200. Process 200 may be part of STEM-IDR. Process 200 may be performed, for example, by system 100 as described above. A non-transitory computer-readable medium may contain instructions that, when executed by at least one processor, perform the operations described at each step as part of process 200. The at least one processor may correspond to system 100 and / or processing circuit 110.
[0043] In step 202, a deep neural network (DNN) facial emotion classifier (FEMC) may be pre-trained to classify emotional states that may be inferred from facial images, thereby classifying the emotional states expressed in N facial training images of different persons. In some embodiments, the DNN may be used to classify other identifying features of different persons based on images associated with those features. Non-limiting examples of such features may include gait or typing style or body identifiers, etc. The embodiments and examples provided herein relate to facial recognition, but this should not be considered limiting, and it should be understood that the systems described herein may be applied to identify other features associated with an individual to enable identification of the individual. The FEMC in Figure 2B Schematically shown in FIG. 1 is a classifier 222 that operates to classify the emotional state of a 4th image among N facial training images 224 that the classifier can use to classify.
[0044] In step 204, the same hidden layer as layer 226 of FEMC 222 may be used (in Figure 2B A set of H emotion features are extracted for each processed training image, and the training image can be used to define an embedded emotion feature vector EF (with a dimension of ). For the nth (1≤n≤N) facial expression training image, the extracted embedded emotion feature vector can be described by the following formula (1):
[0045]
[0046] Figure 2B The emotional feature vector marked by the number 228 is schematically shown It may be extracted from layer 226 for training image 4 processed by FEMC 222 .
[0047] In some embodiments, in step 206, the N feature vectors EF(n) may be clustered in a clustering space 229, which is schematically represented by Figure 2B The small asterisk 231 in denoted that the N facial training images are determined into a plurality of clusters 230 of an optionally predetermined number S.
[0048] In step 208, the dimension is The centroid vector 232 may optionally be determined for each cluster. It is noted that the dimension G of the centroid vector 232 may be equal to or less than the dimension H of the embedded emotion feature vector EF(n). For example, if the centroid vector is found to be located on a hyperplane of the space including the embedded emotion feature vectors, G may be less than H. Alternatively, it may be determined in response to a principal component analysis, and the number of components G less than H may be advantageously used to characterize the embedded emotion feature vector EF(n).
[0049] In some embodiments, in step 210, S centroid vectors 232 may be defined as emotion basis vectors B(s), 1≤s≤S, which are obtained by Figure 2B denoted by reference numeral 234 in , which has a component b(s)g, 1≤g≤G:
[0050]
[0051] According to an embodiment of the present disclosure, by using features from a hidden layer to distinguish a person's emotional state and determining an emotional basis vector B(s) across the emotional state, STEM-IDR can have enhanced sensitivity to subtle differences in emotional state, which can provide STEM-IDR with improved resolution for distinguishing how different people display the same or similar emotional states. The enhanced sensitivity and resolution can provide advantageous performance in identifying people in videos where people appear.
[0052] Figure 3A A flow chart of a process 300 for extracting STEM data from a video of a person to be identified and generating a STEM data compendium is shown in accordance with an embodiment of the present disclosure. FIG. 3B to FIG. 3E A schematic diagram of the operation of process 300 according to one embodiment is provided, and features and entities involved in process 300 are schematically illustrated.
[0053] A non-transitory computer-readable medium may include instructions that, when executed by at least one processor, cause the at least one processor to define a STEM data profile from a video of a person to be identified (e.g., in process 300). Process 300 may be part of a STEM-IDR. Process 300 may be performed, for example, by system 100 as described above. A non-transitory computer-readable medium may include instructions that, when executed by at least one processor, perform the operations described at each step as part of process 300. The at least one processor may correspond to system 100 and / or processing circuit 110.
[0054] In step 302, a video 320 including N frames F(n) may be received, 1≤n≤N, by Figure 3B Schematically illustrated in , a process for determining the identity of a person appearing in a video according to an embodiment of the present disclosure. In some embodiments, STEM-IDR including processes 200, 300, and 400 may use facial features or other identifying features to identify a person. Non-limiting examples of such features may include gait or typing style or body identifiers, etc. The embodiments and examples provided herein relate to facial recognition, but this should not be considered limiting, and it should be understood that the system described herein can be applied to identify other identifiable features associated with an individual to enable identification of the individual.
[0055] In some embodiments, video 320 may be provided at a frame rate (fps) higher than 60, and this higher number of frames per time period may enable detection of micro-expressions (defined herein as facial expressions) that last for very short time periods of less than half a second. Thus, in some embodiments, the resolution of the determined underlying emotional state may be optionally increased for videos including higher frame rates.
[0056] In step 304, the image of the person's face can be positioned in each frame, and then the image can be corrected into a corrected facial image FI(n) so that the person's face in all corrected facial images FI(n) appears in substantially the same desired pose and size. In some embodiments, the FI(n) images can be corrected so that the person's face is in a full-face frontal pose, and the images can be cropped to the same standard size.
[0057] In step 306, each of the facial images FI(n) may optionally be processed using a pre-trained facial landmark extraction net (FLEN) (e.g., MobileNet-V1) to identify L facial landmarks in FI(n), each facial landmark being represented by P features from the facial image. In addition to the P features extracted by FLEN for each landmark identified by FLEN, each facial image FI(n) of a given video frame F(n) may also be associated with "V" additional spatiotemporal features that may not be extracted by FLEN. In one embodiment, the additional features may include features determined in response to the localization and correction process mentioned in step 304. For example, the V additional features may include features from a hidden layer of FLEN and / or features responsive to the position and pose of the face in the frame and / or changes in pose and / or position relative to temporally adjacent frames. For a given facial image FI(n), a landmark feature vector LF(n) having L landmarks represented by P features may have dimensions Where F = LxP + V. Assuming lf(n)f represents the fth feature of LF(n), LF(n) can be explicitly expressed as:
[0058] LF(n)={lf(n)1, lf(n)2,..., lf(n)F}={lf(n)f|1≤f≤F}. (3)
[0059] For ease and efficiency of description, FLEN is considered to provide V and P features that characterize a logo.
[0060] Optionally, in step 308, Figure 2A Cited in and Figure 2B Schematically shown in FIG. 2 is a method for defining the emotion basis vector B(s) ( Figure 2B The FEMC of the vector 234 in FIG. 234 can be used to extract the emotion feature vector EF(n) of each facial image FI(n) in
[0061] EF(n)={ef(n)1, ef(n)2,...ef(n)H}={ef(n)h|1≤h≤H}. (4)
[0062] For example, Figure 3B The deep neural network DNN-FEMC 222 is schematically shown extracting an embedded emotion vector EF(1) 322 from the facial image in frame F(1). Figure 3B The embodiment according to the present disclosure also shows that FLEN is used as a DNN FLEN 324 for extracting a landmark feature vector LF(1) 326 from the image F(1). Figure 3B , the landmark feature vectors LF(n) 331 and LF(n+1) 334 extracted from the facial images FI(n) and FI(n+1) respectively by FLEN are also schematically shown.
[0063] Optionally, in step 310, each embedded emotion vector EF(n) extracted from the facial image FI(n) may be projected onto S basis vectors B(s) ( Figure 2A , Figure 2B ) to generate an emotion projection vector PB(n) with components pB(n)s for each facial image, 1≤s≤S. The emotion projection vector can be explicitly written as:
[0064] PB(n)={pB(n)1, pB(n)2,...pB(n)S}={pB(n)S|1≤s≤S}. (5)
[0065] In some embodiments, the projected component pB(n)s can be obtained by embedding the emotion feature vector EF(n) The Euclidean distance between EF(n) and B(s) is used to determine the pB(n)s. Let d(EF(n),B(s)) represent the Euclidean distance between EF(n) and B(s). pB(n)s can be determined according to formula (6):
[0066] PB(n)S=[1 / d(EF(n), B(s))] / ∑S1 / (d(EF(n), B(s)). (6)
[0067] For the case where the dimension G of the emotion basis vector B(s) is smaller than the dimension of the embedded emotion feature vector EF(n), the components of the vector EF(n) may be restricted to the components located in the space of the emotion basis vector B(s).
[0068] Figure 3B The emotion projection vector PB(1) 329 generated by the projection operator 325 from EF(1) and PB(n) and PB(n+1) generated for the facial images FI(n) and FI(n+1) provided by the frames F(n) and F(n+1), respectively, is schematically shown. It should be noted that the projected components pB(n)s are normalized so that the sum Σs[pB(n)s]=1, and a given component pB(n)s determined for the embedded emotion feature vector EF(n) can be considered as the probability that the person in the facial image FI(n) is in the basic emotional state represented by the emotion basis vector B(s). For ease of description, the basic emotional state represented by the emotion basis vector B(s), 1≤s≤S, can be referred to as the basic emotional state BE(s).
[0069] In step 312, an emotional state transition probability may be determined for each facial image FI(n) (1≤n≤N-1), the emotional state transition probability being the probability that a person (1≤j≤S) with probability pB(n)j being in a base emotional state BE(j) transitions to a subsequent base emotional state BE(k) in a subsequent facial image FI(n+1), wherein the person has probability pB(n+1)k being in the base emotional state BE(k) in the subsequent facial image FI(n+1). In some embodiments, let the transition probability be represented by a transition weight w(n,n+1)j,k. In one embodiment, the transition probability w(n,n+1)j,k may be determined according to equation (7),
[0070] w(n,n+1)j,k=pB(n)j·pB(n+1)k. (7)
[0071] If a set of transition probabilities w(n,n+1)j,k for facial images FI(n) and FI(n+1) is represented by a transition weight matrix W(n,n+1), then
[0072] W(n, n+1)={w(n, n+1)j, k|1≤j, k≤S}. (8)
[0073] It should be noted that W(n,n+1) is usually not a symmetric matrix and w(n,n+1)j,k≠w(n,n+1)k,j.
[0074] Figure 3C According to an embodiment of the present disclosure, a conversion weight matrix W(n, n+1) 340 determined by the emotion projection vectors PB(n) and PB(n+1) is schematically shown.
[0075] Optionally, in step 314, the transition weight sum matrix WSUM(n+1) may be determined according to equation (9):
[0076]
[0077] in
[0078]
[0079] In step 316, the landmark feature vector LF(n) may be iterated over n from n=1 to n=(N-1) to generate an iterated feature vector IFV(N)j,k,f for each transition from a given basic emotional state BE(j) (1≤j≤S) to a given basic emotional state BE(k) (1≤k≤S) for the video. The iterated feature vector IFV(N)j,k,f given an ordered pair of indices (1≤j,k≤S) may be of dimension , that is, the sum of vectors LF(n) weighted by a function of the transition probabilities w(n,n+1)j,k. In one embodiment, for the nth iteration (n=1→N-1) and a given indicator j,k,
[0080]
[0081] in
[0082] IFV(n+1) j,k,f =(wsum(n) j,k / wsum(n+1) j,k ).IFV(n) j,k,f
[0083] +(w(n,n+1) j,k / waum(n+1) j,k ).lf(n+1)f. (12)
[0084] For the last (N-1)th iteration,
[0085]
[0086] and
[0087] IFV(N) j,k,f =(wsum(N-1) j,k / wsum(N) j,k )IFV(N-1) j,k,f
[0088] +(w(N-1,N) j,k / wsum(N)j,k).lf(N)f}. (14)
[0089] It should be noted that IFV(1) j,k,f is not defined by equations (11)-(14). In one embodiment, IFV(1) j,k,f can be optionally determined according to the following formula,
[0090] IFV(1) j,kf =[wsum(N)j,k / (N-1)].lf(1) f or=lf(1) f (15)
[0091] IFV(N) j,k It can be called a global landmark feature vector (GLFV), which is used for video (e.g., Figure 3B The video 320) shows the transition from the basic emotional state BE(j) to the basic emotional state BE(k). Figure 3D The iterative feature vector IFV((n+1) 350 determined by the nth iteration of the ordered index pair j, k is schematically shown.
[0092] In some embodiments, in step 318, the spatiotemporal emotion STEM data summary may be defined as a three-dimensional tensor STEM-DC:
[0093]
[0094] Figure 3E A three-dimensional STEM-DC 360 is schematically shown according to an embodiment of the present disclosure. The STEM-DC 360 may also be referred to herein as an encoding state transition matrix.
[0095] Figure 4A A flow chart of a process 400 for processing STEM-DC using a fully adaptive graph convolutional network (FAGC) to determine a latent vector representing a person in a video is shown according to some implementations.
[0096] Figure 4B A process 400 including processing of STEM-DC using a fully adaptive graph convolutional network FAGC is schematically illustrated according to some implementations. A non-transitory computer-readable medium may include instructions that, when executed by at least one processor, cause the at least one processor to process STEM-DC using a deep FAGC (e.g., in process 400). Process 400 may be part of a STEM-IDR. Process 400 may be performed, for example, by the system 100 as described above. A non-transitory computer-readable medium may include instructions that, when executed by at least one processor, perform the operations described at each step as part of process 400. The at least one processor may correspond to system 100 and / or processing circuit 110.
[0097] In step 402, a STEM-DC may be generated, for example, by using the above-described process 300. In step 402, according to an embodiment of the present disclosure, the STEM-DC 360 may be processed by a deep FAGC, referred to herein as a STEM-DEEP network 500, or simply STEM-DEEP 500, to determine a latent vector representing a video 320 ( Figure 3B ). In some embodiments, STEM-DEEP 500 may include a feature extraction module 502, followed by a data merging module 504, which may generate a latent vector, optionally referred to as a STEM person representation vector (STEM-PR), and in Figure 4B Schematically shown as PR 510. In some embodiments, in step 404, a first PR 510 determined for a particular individual from the provided video may be compared with a subsequent second PR 510 determined from a subsequent video (by repeating steps 402 and 404), thereby confirming the identity of the individual appearing in the subsequent video if the subsequent second PR 510 is substantially similar to the determined first PR 510.
[0098] In one embodiment, the feature extraction module 502 may include multiple (optionally five) FAGC blocks 503. The data merging module 504 may provide output to a fully connected (FC) layer 506 that generates the PR 510. In some embodiments, the data merging module 504 may include multiple convolution blocks 505, which may be two-dimensional (2D), three-dimensional (3D) or other suitable structures. In one embodiment, each FAGC block 503 may include a data driven attention tensor A att (not shown) and the learned adjacency matrix A adj (not shown), each having dimensions If the input feature to the FAGC block 503 is InF in Indicates that the FAGC block 503 responds to InF in The output features provided by OtF out It means that the operation of a given FAGC block 503 can be expressed as follows,
[0099] OtF out =W(OtF out (A att +A adj )), (17)
[0100] Where W represents the learning weight of FAGC block 503.
[0101] In one embodiment, STEM-DEEP 500 can be trained on a video training set including multiple videos acquired for each of multiple different persons using a triplet margin loss (TML) function and an appropriate distance metric. For each video in the video training set, STEM-DC 360 can be generated and processed using STEM-DEEP 500 to generate a STEM-PR 510 person representation vector. TML operates to provide a favorable distance based on the metric between the STEM-PR 510 vectors of different persons in the training set.
[0102] Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.The materials, methods, and examples provided herein are illustrative only and not limiting.
[0103] As used herein, the terms "convolutional network," "neural network," "machine learning," or "artificial intelligence" refer to the use of algorithms on a computing device to parse data, learn from the data, and then make determinations or generate data, where the determinations or generated data are not deterministically replicable (e.g., using determination-oriented software known in the art).
[0104] The implementation of the method and system of the present disclosure may involve performing or completing certain selected tasks or steps manually, automatically, or a combination thereof. In addition, according to the actual instruments and equipment of the preferred embodiments of the method and system of the present disclosure, several selected steps can be implemented by hardware (HW) or software (SW) on any operating system of any firmware, or by a combination thereof. For example, as hardware, the selected steps of the present disclosure can be implemented as a processor chip or circuit. As software or algorithms, the selected steps of the present disclosure can be implemented as multiple software instructions executed by a computer / processor using any suitable operating system. In any case, the selected steps of the method and system of the present disclosure can be described as being executed by a data processor, such as a computing device for executing multiple instructions.
[0105] Various implementations of the systems and techniques described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs executable and / or interpretable on a programmable system comprising at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to send data and instructions to the storage system, the at least one input device, and the output device.
[0106] Although the present disclosure is described with respect to a "computing device", "computer" or "mobile device", it should be noted that, alternatively, any device with a data processor and the ability to execute one or more instructions may be described as a computing device, including but not limited to any type of personal computer (PC), server, distributed server, virtual server, cloud computing platform, cellular phone, IP phone, smart phone, smart watch or PDA (personal digital assistant). Any two or more such devices in communication with each other may selectively form a "network" or "computer network".
[0107] To provide interaction with a user, the systems and techniques described herein may be implemented on a computing device having a display (indicator / monitor / screen / array) for displaying information to the user (e.g., LED (light emitting diode), OLED (organic LED), LCD (liquid crystal display), or other display technology) and a keyboard and pointing device (e.g., a mouse, joystick, or trackball) or separate buttons / knobs / sticks (e.g., drive wheel buttons / signal sticks) through which the user can provide input to the computing device. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form, including acoustic, voice, analysis of the user's head position and / or eye movements, or tactile input.
[0108] It should be understood that the above methods and devices can be varied in many ways, including omitting or adding steps, changing the order of steps and the type of equipment used. It should be understood that different features can be combined in different ways. In particular, in each embodiment or implementation of the present disclosure, not all of the above features in a particular embodiment or implementation are necessary. Further combinations of the above features and implementations are also considered to be within the scope of some embodiments or implementations of the present disclosure.
[0109] Although certain features of the described embodiments have been described as described herein, those skilled in the art will now make many modifications, substitutions, changes and equivalents. It should be understood that they are presented by way of example only and not limitation, and that various changes in form and detail may be made. Any portion of the apparatus and / or method described herein may be combined in any combination, except mutually exclusive combinations. The implementations described herein may include various combinations and / or sub-combinations of the functions, components and / or features of the different implementations and embodiments described.
Claims
1. A system for identifying a person in a video, comprising: a computing device configured to generate a spatiotemporal emotion data summary (STEM-DC) from the video; and being configured to process the STEM-DC using a deep fully adaptive graph convolutional network (FAGC) to determine a first person representation vector representing the person in the video, wherein generating the STEM-DC comprises generating an iterated feature vector (IFV), wherein generating the IFV comprises iterating over a series of landmark feature vectors, wherein the series of landmark feature vectors are weighted by a function of a plurality of transition probabilities between a plurality of underlying emotional states of the person detected in a plurality of subsequent frames of the video.
2. The system of claim 1 , further configured to compare the first person representation vector with a subsequent second person representation vector determined from a subsequent video and a subsequent STEM-DC, thereby identifying the person as appearing in the subsequent video when the first person representation vector and the second person representation vector are substantially similar.
3. The system according to claim 2, wherein: The first person representation vector and the second person representation vector are based on identifiable features of the person in the video.
4. The system of claim 3, wherein the identifiable characteristics include at least one of face, emotion, gait, body, limbs, or typing style.
5. The system according to claim 1, wherein: The function of the plurality of transition probabilities is represented by a transition weight and a matrix.
6. The system according to claim 1, wherein: Each of the plurality of basis emotional states is determined by projecting an emotion feature vector onto a series of emotion basis vectors.
7. The system according to claim 1, wherein: Each of the series of landmark feature vectors for a given facial image includes L landmarks characterized by P features.
8. The system according to claim 7, wherein: The series of landmark feature vectors are determined by processing a plurality of facial images extracted from the video using a pre-trained facial landmark extraction network (FLEN) to identify the L facial landmarks, wherein each facial landmark is represented by P features.
9. The system according to claim 8, wherein: The plurality of facial images are extracted from the video by locating and correcting a plurality of images of faces of persons located in the video.
10. The system according to claim 1, wherein: The FAGC includes a feature extraction module and a data merging module including a plurality of convolution blocks.
11. The system according to claim 1, wherein: For videos with higher frame rates, the resolution of the underlying emotional state is increased.
Citation Information
Patent Citations
Video multi-target fast tracking method based on joint probability data association
CN101783020A
Facial expression motion unit identification method based on space-time diagram convolutional network
CN112633153A