Access Control Using Facial Recognition and Heterogeneous Information

By combining multimodal facial attributes and auxiliary attributes, such as time attributes, in the facial recognition system, constructing new feature representations and using joint classification models, the high error recognition rate and recognition bias problems of facial recognition systems in the prior art are solved, and the facial recognition performance is significantly improved.

CN116569226BActive Publication Date: 2025-06-13ASSA ABLOY AB

Patent Information

Application Number
CN202080106094.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-12
Publication Date
2025-06-13
Estimated Expiration
2040-10-12

AI Technical Summary

Technical Problem

In existing facial recognition systems, the trained image classification neural network has a high rate of error recognition, especially when dealing with different ethnic groups, age groups and genders, and large facial pose changes can lead to verification failure.

Method used

The combination of multimodal facial attributes, including age, gender, race and head posture, is used in combination with visual information to improve the verification performance of the facial recognition system by constructing new feature representations. In addition, use auxiliary attributes such as temporal attributes and visual information to generate more accurate facial recognition results through joint classification models.

Benefits of technology

Through the combination of multimodal facial attributes and auxiliary attributes, the verification performance of the facial recognition system is significantly improved, the error recognition rate is reduced, the recognition accuracy of different ethnic groups, age groups and genders is enhanced, and different facial posture changes are adapted to different facial posture changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116569226B_ABST
    Figure CN116569226B_ABST
Patent Text Reader

Abstract

Describes the use of multimodal facial attributes in a facial recognition system. Additionally, the use of one or more auxiliary attributes, such as temporal attributes, can be combined with visual information to improve the facial recognition performance of the facial recognition system. In some examples, the use of multimodal facial attributes in a facial recognition system can be combined with the use of one or more auxiliary attributes, such as temporal attributes. Each of these techniques can improve the verification performance of the facial recognition system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates generally, and not by way of limitation, to facial recognition systems and methods. Background Art

[0002] Image classification techniques (e.g., face classification techniques) have multiple use cases. Example use cases include allowing authorized personnel (and prohibiting unauthorized personnel) to enter secure physical locations, authenticating users of electronic devices, or identifying objects (e.g., such as chairs, tables, etc.). One problem with some implementations of image classification techniques is the high false recognition rate of trained image classification neural networks. Brief Description of the Drawings

[0003] In the drawings, which are not necessarily to scale, like reference numerals may describe similar components in different views. Like reference numerals with different letter suffixes may represent different instances of similar components. The drawings generally illustrate, by way of example and not by way of limitation, the various embodiments discussed in this document.

[0004] Figure 1 is a conceptual diagram showing an example of a facial recognition system using multi-modal facial attributes.

[0005] Figure 2 is a conceptual diagram showing an example of a facial recognition system using one or more auxiliary attributes, which can be used in combination with visual information to improve facial recognition performance.

[0006] Figure 3 is a flowchart showing an example of a computer-implemented method using one or more auxiliary attributes, which can be used in combination with visual information to improve the facial recognition performance of a facial recognition system, as described above with respect to Figure 2 described.

[0007] Figure 4 is a flowchart showing an example of computer-implemented method 100, which uses multi-modal facial attributes in a facial recognition system to improve the facial recognition performance of the facial recognition system, as described above with respect to Figure 1 described.

[0008] Figure 5 shows an example machine learning module 200 according to some examples of the present disclosure.

[0009] Figure 6 shows a block diagram of an example machine on which any one or more of the techniques (e.g., methods) discussed herein can be executed. Summary of the Invention

[0010] The present disclosure describes the use of multimodal facial attributes in a facial recognition system. Additionally, the use of one or more auxiliary attributes such as temporal attributes can be used in combination with visual information to improve the facial recognition performance of the facial recognition system. In some examples, the use of multimodal facial attributes in a facial recognition system can be combined with the use of one or more auxiliary attributes such as temporal attributes. Each of these techniques can improve the verification performance of the facial recognition system.

[0011] In some aspects, the present disclosure relates to a computer-implemented method of using a facial recognition system to identify a person, the method comprising: extracting an attribute of a person by applying a first representation of a first image of the person to a pre-trained attribute classifier machine learning model to generate an attribute classifier output; applying the attribute classifier output and a distance measurement output generated using a second representation of a second image of the person to a pre-trained fusion verification machine learning model; using the pre-trained fusion verification machine learning model to generate a facial recognition system output; and using the facial recognition system output to control access to a secure asset.

[0012] In some aspects, the present disclosure relates to a computer-implemented method of using a facial recognition system to identify a person based on an image of the person, the method comprising: performing a facial embedding; applying the facial embedding to a classification pipeline to identify similar images; generating a classification pipeline output based on the identified images; applying the classification pipeline output to a joint classification model; applying an auxiliary attribute to the joint classification model; using the joint classification model to generate a joint classification output; and using the joint classification output to control access to a secure asset. Detailed Description

[0013] Facial verification is an important aspect in almost any modern facial recognition system. Facial verification can compute the one-to-one similarity between a probe and a gallery to determine whether the two images have the same object, where the gallery information can be obtained through a face classification pipeline (one-to-many).

[0014] Due to the fact that there are too many variations and nuisances in the original facial image, verification cannot be performed at the pixel level between the probe and the gallery. Instead, high-level features are extracted from the facial image by traditional methods such as HOG, SIFT, etc. or more advanced and data-driven neural network methods such as Dlib, Arcface, etc. Verification can then be performed among the facial feature vectors using a similarity metric such as Euclidean distance or cosine similarity.

[0015] Although advanced facial feature embeddings can provide a simplified way to perform face verification, the inventors have recognized that there are still limitations associated with the embedding vectors, which can sometimes lead to verification failures. The limitations mainly come from the following aspects.

[0016] The neural network that can be used for feature extraction may heavily rely on the quality of the generated face embeddings, and the generated face embeddings manage the performance of verification. At the same time, the performance of the neural network may be affected by many factors such as the training dataset, network design, and loss function.

[0017] In terms of the training dataset, most publicly available face image datasets have many problems in terms of demographic bias. For example, the ethnic groups across training identities are usually unbalanced, with most identities belonging to one ethnic group. For example, face image datasets such as VGGFace, MS-celeb-1M, and Youtube Face. The imbalance of ethnic groups may cause the performance of the neural network to be biased towards one ethnic group, thus resulting in a negative impact on the verification performance of other ethnic groups.

[0018] In addition to race, age variations in the training dataset may also introduce bias towards a specific age group, depending on how the collection process is carried out. Age bias may lead to a bias in the performance of the trained neural network.

[0019] In addition to race and age, gender imbalance is also a problem. When constructing a list of identities for training based on public information, gender imbalance can usually be controlled, but it may not be perfect if the evaluation is not carefully performed during collection.

[0020] Finally, large facial pose variations may be another reason for verification failures. Comparing two profile face images for verification purposes is a more difficult task, even from a human perspective.

[0021] This disclosure describes various techniques for overcoming the above problems. The inventors have recognized the need to construct new feature representations from multi-modal facial attributes. The new feature representation can be a combination of one or more facial attributes such as age, gender, and race, as will be described in more detail below.

[0022] In addition to describing the use of multi-modal facial attributes in a face recognition system, this disclosure describes the use of one or more auxiliary attributes such as time attributes, which can be used in combination with visual information to improve the face recognition performance of the face recognition system. In some examples, the use of multi-modal facial attributes in a face recognition system can be combined with the use of one or more auxiliary attributes such as time attributes. Each of these techniques can improve the verification performance of the face recognition system.

[0023] Modern face recognition systems in unconstrained environments attempt to accomplish two tasks: 1) face classification (1-to-N), which attempts to match a query face image to the most similar face in an existing gallery database collected through enrollment; and 2) face verification (1-to-1), which attempts to determine whether the query image and the most similar gallery image are of the same person. Both processes rely heavily on face embeddings obtained through a face feature extraction pipeline that ranges from traditional hand-designed image feature extractors to neural networks.

[0024] As the process indicates, the effectiveness of face verification depends, to some extent, on the performance of face classification. To achieve better matching, machine learning methods such as linear classifiers like SVM or logistic regression or decision trees like XGboost or LightGBM can be utilized. The idea is to train a classification model on face embeddings extracted from an enrolled face database to better understand the target population. In some examples, each identity will enroll a small number of its face images with high quality and controlled variations.

[0025] Figure 1 FIG. 7 is a conceptual diagram showing an example of a face recognition system 10 using multi-modal face attributes. To recognize person 12, a face or facial recognition network 14 can generate a face embedding 16 based on a first representation (e.g., a first template, a first vector, or a first image) of a first image of the person. The face embedding 16 can include a numerical vector that can represent each detected face in the representation of the image (e.g., in a 1024-dimensional space).

[0026] In some examples, a machine learning model or a neural network can be used to generate the face embedding. For example, a neural network such as ArcFace can be used to generate the face embedding. For example, the pre-trained weights of the neural network can be converted from the open-source deepinsight / insightface git repository, which can be trained using the LResNet50E-IR network architecture and the ArcFace loss function.

[0027] Using the various techniques of the present disclosure, a computer-implemented method can construct a new feature representation from multi-modal face attributes. The new feature representation can be used to train a binary classifier for face verification purposes. The new feature representation can be a combination of multiple face attributes such as age, gender, and ethnicity.

[0028] For example, a computer-implemented method can extract one or more attributes of a person 12 by applying a first representation of a first image of the person to a pre-trained attribute classifier machine learning model. Attributes can include, for example, the age, gender, race, and / or head pose of the person. For example, a face embedding can generate a vector that is applied to one or more pre-trained attribute classifier machine learning models, and then each model can extract its specific corresponding attribute. In another example, one or more pre-trained attribute classifier machine learning models can each receive the raw image as input, generate features, and then extract the corresponding attribute.

[0029] By way of non-limiting specific examples, the method can use inferred age scores in 5 categories, such as 15 to 20 years old, 25 to 32 years old, 38 to 43 years old, 48 to 53 years old, and over 60 years old; inferred gender scores in 2 categories: male, female; inferred race scores in 5 categories; and the pitch, yaw, and roll angles of the head for the estimation of the head pose of the person.

[0030] The pre-trained attribute classifier machine learning models can form part of the verification engine 18. The age classifier machine learning model 20 can extract the age of the person 12 from a representation of an image of the person. Similarly, the gender classifier machine learning model 22 can extract the gender of the person 12. The race classifier machine learning model 24 can extract the race of the person 12. The head pose classifier machine learning model 26 can extract the head pose of the person 12, such as pitch, yaw, and roll angles.

[0031] Additionally, a computer-implemented method can use a second representation (e.g., a second template, a second vector, or a second image) of a second image of the person to obtain a distance measurement. In some examples, the first template can be the same as the second template. In some examples, the first template can be different from the second template. In some examples, the first image can be the same as the second image. In some examples, the first image can be different from the second image, e.g., taken at a different angle.

[0032] The distance measurement classifier 28 of the verification engine 18 can compare two vectors from two embeddings, e.g., templates. For example, the distance measurement classifier 28 can compare the vector from the face embedding presented to the reader with the vector from the face embedding of the person 12 obtained, for example, by a camera device at the scene. These two vectors are represented by the bottom inputs 30 of the distance measurement classifier 28.

[0033] In some examples, the distance measurement classifier 28 may alternatively compare a vector from a face embedding presented to the reader or imaging device with, for example, a vector from a face embedding of a person found during a search of images in a database. For example, a classification pipeline 32 such as a pre-trained machine learning model or algorithm may perform a search for the most similar template in the template database. The output of the classification pipeline 32 may include the most similar template. The computer-implemented method may then use, for example, centroid lookup 34 to perform verification of the image.

[0034] The distance measurement classifier 28 of the verification engine 18 may compare a vector from a face embedding presented to the reader or imaging device with, for example, a vector from a face embedding of the most similar image found during use of the classification pipeline 32 and centroid lookup 34. These two vectors are represented by the top inputs 36 of the distance measurement classifier 28. In some examples, the distance measurement classifier 28 may use two inputs 30, 36.

[0035] The distance measurement output 38 of the distance measurement classifier 28 may be a mathematical value generated by the comparison of the two vectors. In some examples, the distance measurement classifier 28 may be a pre-trained machine learning model. In other examples, the distance measurement classifier 28 may be an algorithm.

[0036] Each of the pre-trained attribute classifier machine learning models 20 to 26 may generate corresponding attribute classifier outputs 40 to 46. Using the techniques of the present disclosure, one or more of the attribute classifier outputs 40 to 46 may be applied together with the distance measurement output 38 to a pre-trained fusion verification machine learning model 48. In some examples, the pre-trained fusion verification machine learning model 48 may be a binary XGBoost classifier constructed from a decision tree.

[0037] By fusing or combining the distance measurement output 38 with one or more of the one or more attribute classifier outputs 40 to 46, the pre-trained fusion verification machine learning model 48 may improve the verification performance of the face recognition system.

[0038] The computer-implemented method may use the pre-trained fusion verification machine learning model 48 to generate a face recognition system output 50, which may, for example, control access to a secure asset such as a room, building, or computer system.

[0039] The verification engine 18 may be a computer program. The verification engine 18, the pre-trained machine learning model, and any algorithms described in the present disclosure may use Figure 6implemented by the machine 300. For example, the verification engine 18, one or more of the machine learning models 20-26, and the fusion verification machine learning model 48 can be implemented using instructions 324 executed by the processor 302. The verification engine and the machine learning models can be co-located on one machine or located in different machines.

[0040] The present inventors evaluated Figure 1 the techniques, and the results are shown in Table 1 below. Table 1 shows the verification results to demonstrate the performance of the new feature representation from multi-modal face attribute fusion:

[0041] Table 1

[0042]

[0043]

[0044] The present inventors constructed an exhaustive list of combinations for all the new feature categories, as shown in Table 1, where "1" indicates that the feature (age, gender, race, and yaw angle) was used, and "0" indicates that the feature was not used. For each combination, an XGBoost classifier was trained using the VGGFace2 test dataset, where half of the identities were designated as genuine and the other half as impostors. The results shown in Table 1 indicate that the verification performance can be improved by using one or more of the available attributes when constructing the new feature representation.

[0045] As mentioned above, in addition to using multi-modal face attributes in a face recognition system, the present disclosure describes the use of one or more auxiliary attributes such as time attributes, which can be used in combination with visual information to improve the face recognition performance of the face recognition system. The use of auxiliary features is shown and described below with respect to Figure 2 illustrates and describes the use of auxiliary features.

[0046] Several factors can limit the performance of visual classifiers. For example, face embedding generation neural networks are not perfect. Since face feature extraction methods play an important role in the entire face recognition pipeline, research is underway to develop state-of-the-art face feature extraction methods. Although the performance of neural networks has reached a very high level, there is still no solution for achieving perfect face recognition on large-scale and customized datasets. In addition, in an unconstrained imaging environment, the variation of query images also affects the matching performance. Some of the most common nuances are described above, including face pose angle, age variation, race, or query image quality.

[0047] In view of the foregoing, the present inventors have recognized a need for a new architecture for combining visual and auxiliary attributes such as temporal information for improved face recognition performance, the new architecture being shown and described below in Figure 2 as shown and described.

[0048] Figure 2 is a conceptual diagram illustrating an example of a face recognition system 60 that uses one or more auxiliary attributes that can be used in combination with visual information to improve face recognition performance. Figure 2 Conceptually depicts a computer-implemented method of identifying a person based on an image of the person using a face recognition system. As described below, when a visual classifier predicts multiple candidates with a close confidence level, auxiliary attributes such as temporal information can be used to re-rank the predictions for better identification. Figure 2 The techniques of can improve the accuracy of face classification (1-to-N).

[0049] To identify person 62 (ID_a), a face recognition network 64 can generate or perform a face embedding based on a first representation of a first image of the person (e.g., a first template, a first vector, or a first image). The face embedding can include a numerical vector that can represent each detected face in the representation of the image (e.g., in a 1024-dimensional space).

[0050] The face embedding (e.g., vector) generated by the face recognition network can then be applied to a classification pipeline 66. For example, the classification pipeline 66, such as a pre-trained machine learning model or algorithm, can perform a search for the most similar template in a template database. The output of the classification pipeline 66 can include the most similar template. In Figure 2 the example shown, the most similar template found by the classification pipeline is template ID_b, shown at 68. Figure 2 The classification pipeline 66 in determines that for person 62 (ID_a), it has a higher prediction confidence in template ID_b than in template ID_a.

[0051] The classification pipeline 66 can generate a classification pipeline output 70 based on the recognized image. In some examples, the classification pipeline output 70 can include a distance measurement. In other examples, the classification pipeline output 70 can include a score or match probability for each enrolled identity in the system.

[0052] For example, the classification pipeline 66 can include a distance measurement classifier, such as the one described above with respect to Figure 1As described, a distance measurement classifier can compare two vectors from two embeddings, such as templates. In some examples, the distance measurement classifier can compare a vector from a face embedding presented to a reader or imaging device with a vector from a face embedding of, for example, the most similar image found during classification pipeline and centroid finding.

[0053] The output of the distance measurement classifier (which can also be the output of classification pipeline 66) can be a mathematical value resulting from the comparison of the two vectors. In some examples, the distance measurement classifier can be a pre-trained machine learning model. In other examples, the distance measurement classifier can be an algorithm.

[0054] Using the various techniques of the present disclosure, auxiliary attributes 72 can be used in combination with visual information obtained via classification pipeline 66 to improve the face recognition performance of face recognition system 60. Auxiliary attributes 72 can include time information, the height of person 62, and / or the social pool of person 62, etc.

[0055] The time information can include, for example, the time range during which person 62 typically arrives at the face recognition system, such as at a secure entrance to a building, such as a door. The machine learning model can learn this time range of person 62 over time. In some examples, if there are multiple secure entrances, the machine learning model can learn the specific secure entrance of the building used by person 62, such as a door. Then, using the various techniques of the present disclosure, the time information or other auxiliary attributes can be applied to joint classification model 74 together with classification pipeline output 70.

[0056] For illustrative purposes, by way of non-limiting specific examples, the machine learning model of face recognition system 60 pre-determines that person 62 regularly arrives at the secure entrance of the building between 6:00 am and 6:15 am. This time information or other auxiliary attributes can be applied to joint classification model 74 together with classification pipeline output 70 to improve the face recognition performance of face recognition system 60. Classification pipeline 66 initially determines a higher prediction confidence in template ID_b than in template ID_a for person 62 (ID_a), as shown at 68 in Figure 2 However, the person associated with template ID_b regularly arrives at the secure entrance between 9:00 am and 9:15 am, and the current time is 6:10 am. Thus, person 62 is more likely to be associated with template ID_a rather than template ID_b.

[0057] By using time information, for example, face recognition system 60, and more specifically, joint classification model 74 can use joint classification model 74 to generate a joint classification output 76, which determines a higher prediction confidence in template ID_a than in template ID_b for person 62 (ID_a), asFigure 2 as shown at 78 locations in

[0058] The time information is an auxiliary attribute that can be used. In other examples, a person's height can be used. In other examples, a person's clothing can be used. In other examples, a person's gait can be used. A machine learning model can be trained to classify a person's height, clothing, and gait.

[0059] In some examples, the social pool of person 62 can be used. The social pool can include information about the network of people associated with person 62. For example, the social pool of person 62 can include the people who leave work, arrive at work, go to lunch, etc. with person 62. The machine learning model can learn the identities of these individuals over time to determine the social pool of person 62. For example, when person 62 arrives at a secure entrance, a pre-trained machine learning model can determine the prediction confidence of one or more individuals with respect to person 62, compare the prediction confidence with the social pool of person 62, and apply the social and classification pipeline output 70 to the joint classification model 74 to improve the face recognition performance of the face recognition system 60.

[0060] In some examples, the face recognition system 60 can generate a joint classification output 76 by using the joint classification model 74 to generate a joint probability. For example, Bayes' theorem can be used to determine the joint probability based on two inputs of the joint classification model 74:

[0061]

[0062] The above equation indicates that the matching probability (Pr(query = gallery n |v, t)) (Equation 1) of query face and each face in the enrollment gallery that jointly uses visual and time information is proportional to the product of the matching probability (Pr(query = gallery n|v)) from a visual classifier trained on visual features and the probability (Pr(t|query = gallery n)) that the predicted person arrives at a gateway or entrance at a timestamp (Equation 2).

[0063] In this way, a combination of visual information and auxiliary attributes can be used to improve the face recognition performance of the face recognition system.

[0064] Figure 3 is a flowchart of an example of a computer-implemented method 80 that uses one or more auxiliary attributes that can be combined with visual information to improve the face recognition performance of a face recognition system, as described above with respect to Figure 2 described.

[0065] At block 82, the computer-implemented method 80 may perform a face embedding. For example, a face recognition network such as Figure 2 the face recognition network 64 may generate or perform a face embedding based on a first representation (such as a first template, a first vector, or a first image) of a first image of a person.

[0066] At block 84, the computer-implemented method 80 may apply the face embedding to a classification pipeline to identify similar images. For example, the face embedding (e.g., a vector) generated by the face recognition network may be applied to a classification pipeline such as Figure 2 the classification pipeline 66.

[0067] At block 86, the computer-implemented method 80 may generate a classification pipeline output based on the identified images. For example, Figure 2 the classification pipeline 66, such as a pre-trained machine learning model or algorithm, may perform a search for the most similar template in a template database and output the most similar template. In some examples, Figure 2 the classification pipeline output 70 may be a distance measurement.

[0068] At block 88, the computer-implemented method 80 may apply the classification pipeline output to a joint classification model. For example, the classification pipeline output such as Figure 2 the classification pipeline output 70 may be applied to a joint classification model such as Figure 2 the joint classification model 74, which may calculate a joint probability using Bayesian theory.

[0069] At block 90, the computer-implemented method 80 may apply auxiliary attributes to the joint classification model. For example, time information or other auxiliary attributes may be applied to a joint classification model such as Figure 2 the joint classification model 74.

[0070] At block 92, the computer-implemented method 80 may use the joint classification model to generate a joint classification output.

[0071] At block 94, the computer-implemented method 80 may use the joint classification output to control access to a secure asset. For example, Figure 2 the joint classification output 76 may be used to control access to a room, a building, or a computer system, such as by providing a control signal to a control system communicating with the secure asset.

[0072] Figure 4 is a flowchart of an example of a computer-implemented method 100 for using multi-modal face attributes to improve the face recognition performance of a face recognition system, as described above with respect to Figure 1 .

[0073] At block 102, the computer-implemented method 100 can extract attributes of a person, such as age, gender, race, and head pose, by applying a first representation of a first image of the person, such as a first template, a first vector, or a first image, to a pre-trained attribute classifier machine learning model to generate an attribute classifier output. For example, the corresponding pre-trained attribute classifier machine learning models 20 to 26 of Figure 1 can be used to extract one or more attributes of a person, such as age, gender, race, and / or head pose.

[0074] In some examples, a face embedding can generate a vector that is applied to one or more pre-trained attribute classifier machine learning models, and then each model can extract its specific corresponding attribute. In other examples, one or more pre-trained attribute classifier machine learning models can each receive the original image as input, generate features, and then extract the corresponding attributes.

[0075] At block 104, the computer-implemented method 100 can apply the attribute classifier output and a distance measurement output generated using a second representation of a second image of the person, such as a second template, a second vector, or a second image, to a pre-trained fusion verification machine learning model. For example, Figure 1 one or more of the attribute classifier outputs 40 to 46 of Figure 1 and the distance measurement output 38 of the distance measurement classifier 28 of Figure 1 can be applied to the pre-trained fusion verification machine learning model 48 of

[0076] At block 106, the computer-implemented method 100 can use the pre-trained fusion verification machine learning model to generate a face recognition system output. For example, Figure 1 the pre-trained fusion verification machine learning model 48 of

[0077] In some examples, the use of multimodal face attributes in a face recognition system can be combined with the use of one or more auxiliary attributes, such as temporal attributes. Each of these techniques can improve the verification performance of the face recognition system. That is, in some examples, Figure 2 the techniques of Figure 1 can optionally be combined with the techniques of Figure 4 These optional techniques are shown at blocks 108 and 110 in

[0078] At block 108, the computer-implemented method 100 can optionally apply the face recognition system output and the auxiliary attribute to a joint classification model. For example, Figure 1The output 50 of the face recognition system can be applied together with auxiliary attributes such as time information to Figure 2 the joint classification model 74.

[0079] At block 110, the computer-implemented method 100 can optionally use a joint classification model such as Figure 2 the joint classification model 74 to generate a joint classification output.

[0080] At block 112, the computer-implemented method 100 can use the joint classification output or, if no auxiliary attributes are used, the output 50 of the face recognition system to control access to a secure asset. For example, the joint classification output 76 can be used to control access to a room, building, or computer system, such as by providing a control signal to a control system communicating with the secure asset.

[0081] Figure 5 An example machine learning module 200 in accordance with some examples of the present disclosure is shown. The machine learning module 200 can be implemented in whole or in part by one or more computing devices. In some examples, the training module 202 can be implemented by a different device than the prediction module 204. In these examples, a model 214 can be created on a first machine and then sent to a second machine.

[0082] The machine learning module 200 utilizes a training module 202 and a prediction module 204. The training module 202 can use machine learning in a face recognition system to implement a computerized method of a training processing circuit (such as Figure 6 the processor 302) to identify a person. The training module 202 inputs training data 206 into a selector module 208.

[0083] The training data 206 can include, for example, images or feature vectors generated by a feature extractor network. Using the training data 206, a machine learning algorithm 212 can obtain, for each sample, an image or feature vector with a desired label (such as age, gender, etc.) and pass it through a machine learning algorithm such as SVM, XBoost.

[0084] In some examples, the training data 206 can be labeled. In other examples, the training data can be unlabeled, and feedback data (such as through reinforcement learning methods) can be used to train the model.

[0085] Selector module 208 selects training vectors 210 from training data 206. The selected data can populate the training vectors 210 and include a set of training data determined to predict material classification. The information selected for inclusion in the training vectors 210 can be all of the training data 206, or in some examples, can be a subset of all of the training data 206. Training learning algorithm 212 can utilize the training vectors 210 (along with any applicable labels) to produce a model 214 (a trained machine learning model). In some examples, other data structures besides vectors can be used. The machine learning algorithm 212 can learn one or more layers of the model.

[0086] Example layers can include convolutional layers, dropout layers, pooling / upsampling layers, SoftMax layers, etc. An example model can be a neural network, where each layer includes a plurality of neurons that take a plurality of inputs, weight the inputs, input the weighted inputs into an activation function to produce an output, which can then be sent to another layer. Example activation functions can include rectified linear unit (ReLu), etc. The layers of the model can be fully or partially connected.

[0087] In prediction module 204, data 216 can be input to selector module 218. The data 216 can include an acoustic imaging dataset, such as an S matrix. The selector module 218 can operate the same or differently from the selector module 208 of the training module 202. In some examples, selector modules 208 and 218 are the same module or different instances of the same module. The selector module 218 produces a vector 220 that is input into the model 214 to generate an image of the material depicting the defect probability for each pixel or voxel, thereby producing an image 222.

[0088] For example, by applying the vector 220 to the first layer of the model 214 to produce an input to the second layer of the model 214, and so on until the output material classification, the weights and / or network structure, etc., learned by the training module 202 can be performed on the vector 220. As previously mentioned, other data structures besides vectors (e.g., matrices) can be used.

[0089] The training module 202 can operate in an offline manner to train the model 214. However, the prediction module 204 can be designed to operate in an online manner. It should be noted that the model 214 can be periodically updated via additional training and / or user feedback. For example, when a user provides feedback on age, gender, race, head pose, etc., and / or auxiliary attributes such as time information, additional training data 206 can be collected. The feedback, along with the data 216 corresponding to the feedback, can be used to refine the model by the training module 202.

[0090] The machine learning algorithm 212 can be selected from among many different potential supervised or unsupervised machine learning algorithms. Examples of learning algorithms include artificial neural networks, convolutional neural networks, Bayesian networks, instance-based learning, support vector machines, decision trees (e.g., Iterative Dichotomiser 3, C4.5, Classification and Regression Trees (CART), Chi-squared Automatic Interaction Detector (CHAID), etc.), random forests, linear classifiers, quadratic classifiers, k-nearest neighbors, linear regression, logistic regression, region-based CNNs, full CNNs (for semantic segmentation), Mask R-CNN algorithms for instance segmentation, and hidden Markov models. Examples of unsupervised learning algorithms include the expectation maximization algorithm, vector quantization, and information bottleneck methods.

[0091] In this way, according to the present disclosure, Figure 5 the machine learning module 200 can assist in implementing a computerized method of using a face recognition system to identify a person.

[0092] The techniques shown and described in this document can be performed using a part or the whole of the machine 300, as discussed below with respect to Figure 6 what is discussed.

[0093] Figure 6 A block diagram of an example of a machine is shown on which any one or more of the techniques (e.g., methods) discussed herein can be performed. In an alternative embodiment, the machine 300 can operate as a stand-alone device or be connected (e.g., networked) to other machines. In a networked deployment, the machine 300 can operate in a server-client network environment as a server machine, a client machine, or both. In an example, the machine 300 can act as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. The machine 300 is a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), mobile phone, smartphone, network device, network router, switch, or bridge, server computer, database, conferencing device, or any machine capable of executing instructions (sequentially or otherwise) specifying actions to be taken by that machine. In various embodiments, the machine 300 can perform one or more of the above-described processes. Additionally, although only a single machine is shown, the term "machine" should also be regarded as including any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein, such as cloud computing, software as a service (SaaS), other computer cluster configurations.

[0094] As described herein, an example can include logic or multiple components, modules, or mechanisms (all of which are hereinafter referred to as "modules") or can operate on logic or multiple components, modules, or mechanisms. A module is a tangible entity (e.g., hardware) capable of performing specified operations and is configured or arranged in some manner. In an example, a circuit can be arranged as a module in a specified manner (e.g., internally or with respect to an external entity such as another circuit). In an example, all or part of one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware processors are configured by firmware or software (e.g., instructions, an application portion, or an application) to operate to perform specified operations as a module. In an example, the software can reside on a non-transitory computer-readable storage medium or other machine-readable medium. In an example, when the software is executed by the underlying hardware of the module, it causes the hardware to perform the specified operations.

[0095] Accordingly, the term "module" is understood to encompass a tangible entity that is physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., transiently) configured (e.g., programmed) to operate in a specified manner or to perform some or all of the operations described herein. Considering examples in which a module is configured transiently, each module within the module need not be instantiated at any given time. For example, in the case where a module includes a general hardware processor configured using software, the general hardware processor is configured as various different modules at different times. Thus, the software can configure the hardware processor to constitute a particular module at one instance of time and a different module at a different instance of time.

[0096] A machine (e.g., a computer system) 300 can include a processor (e.g., a hardware processor) 302 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 304, and a static memory 306, some or all of which can communicate with each other via an interconnecting link 308 (e.g., a bus). The machine 300 can also include a display unit 310, an alphanumeric input device 312 (e.g., a keyboard), and a user interface (UI) navigation device 314 (e.g., a mouse). In an example, the display unit 310, the input device 312, and the UI navigation device 314 are touchscreen displays. The machine 300 can additionally include a storage device (e.g., a drive unit) 316, a signal generation device 318 (e.g., a speaker), a network interface device 320, and one or more sensors 321, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or other sensors. The machine 300 can include an output controller 328, such as a serial (e.g., a universal serial bus (USB), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection, to communicate with or control one or more peripheral devices (e.g., a printer, a card reader, etc.).

[0097] The storage device 316 can include a machine-readable medium 322 on which is stored a set or more sets of data structures or instructions 324 (e.g., software) that implement any one or more of the techniques or functions described herein or are utilized by any one or more of the techniques or functions described herein. The instructions 324 can also reside, in whole or at least in part, within the main memory 304, within the static memory 306, or within the processor (e.g., a hardware processor) 302 during execution thereof by the machine 300. In an example, one or any combination of the processor (e.g., a hardware processor) 302, the main memory 304, the static memory 306, or the storage device 316 can constitute a machine-readable medium.

[0098] Although the machine-readable medium 322 is shown as a single medium, the term "machine-readable medium" can include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) configured to store one or more instructions 324.

[0099] The term "machine-readable medium" can include any medium that can store, encode, or carry instructions for execution by machine 300 and cause machine 300 to perform any one or more of the techniques in the technology of the present disclosure, or any medium that can store, encode, or carry data structures used by or associated with such instructions. Non-limiting examples of machine-readable media can include solid-state memories and optical and magnetic media. Specific examples of machine-readable media can include: non-volatile memories, such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)), and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; random access memory (RAM); solid-state drives (SSD); and CD-ROM and DVD-ROM disks. In some examples, the machine-readable medium can include a non-transitory machine-readable medium. In some examples, the machine-readable medium can include a machine-readable medium that is not a transitory propagated signal.

[0100] Instructions 324 can also be sent or received over communication network 326 via network interface device 320 using a transmission medium. Machine 300 can communicate with one or more other machines using any one of a plurality of transmission protocols (e.g., Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc.). Example communication networks can include local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), mobile telephone networks (e.g., cellular networks), plain old telephone (POTS) networks, and wireless data networks (e.g., the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard family, known as ), the IEEE 802.16 standard family, known as ), the IEEE 802.15.4 standard family, the Long Term Evolution (LTE) standard family, the Universal Mobile Telecommunications System (UMTS) standard family, peer-to-peer (P2P) networks, etc. In an example, network interface device 320 can include one or more physical jacks (e.g., Ethernet, coaxial, or telephone jacks) or one or more antennas to connect to communication network 326. In an example, network interface device 320 can include multiple antennas to perform wireless communication using at least one of single-input multiple-output (SIMO) technology, multiple-input multiple-output (MIMO) technology, or multiple-input single-output (MISO) technology. In some examples, network interface device 320 can perform wireless communication using multi-user MIMO technology.

[0101] As described herein, an example may include logic or multiple components, modules, or mechanisms, or an example may operate on logic or multiple components, modules, or mechanisms. A module is a tangible entity (e.g., hardware) capable of performing specified operations and configured or arranged in some manner. In an example, a circuit may be arranged as a module in a specified manner (e.g., internally or relative to external entities such as other circuits). In an example, all or part of one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware processors are configured by firmware or software (e.g., instructions, an application portion, or an application) to operate to perform specified operations as a module. In an example, the software may reside on a machine-readable medium. In an example, when the software is executed by the underlying hardware of the module, it causes the hardware to perform the specified operations.

[0102] Accordingly, the term "module" is understood to encompass a tangible entity that is physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., transiently) configured (e.g., programmed) to operate in a specified manner or to perform some or all of the operations described herein. Considering examples where a module is transiently configured, each module in the module need not be instantiated at any given time. For example, in the case where a module includes a general hardware processor configured by software, the general hardware processor is configured as various different modules at different times. Thus, the software can configure the hardware processor to, for example, constitute a particular module at one time instance and a different module at a different time instance.

[0103] Various embodiments are implemented, in whole or in part, in software and / or firmware. The software and / or firmware may take the form of instructions included in or on a non-transitory computer-readable storage medium. Those instructions can then be read and executed by one or more processors to enable performance of the operations described herein. The instructions are in any suitable form, such as but not limited to source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. Such a computer-readable medium may include any tangible non-transitory medium for storing information in a form readable by one or more computers, such as but not limited to read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory, etc.

[0104] Various annotations

[0105] The above specific embodiments include references to the accompanying drawings, which form a part of the specific embodiments. By way of illustration, the accompanying drawings show specific embodiments in which the present invention may be practiced. Such embodiments are also referred to herein as "examples". Such examples may include elements other than those shown or described. However, the inventors also contemplate examples in which only those elements shown or described are provided. In addition, the inventors also contemplate examples using any combination or permutation of those elements (or one or more aspects of those elements) shown or described with respect to a particular example (or one or more aspects of a particular example) or with respect to other examples (or one or more aspects of other examples) shown or described herein.

[0106] The above description is intended to be illustrative, not restrictive. For example, the above examples (or one or more aspects of the above examples) may be used in combination with each other. After reviewing the above description, other embodiments may be used by, for example, a person of ordinary skill in the art. Additionally, in the above specific embodiments, various features may be combined together to simplify the present disclosure. This should not be construed as intending that the disclosed features not claimed are necessary for any claim. Rather, the inventive subject matter may lie in less than all of the features of a particular disclosed embodiment.

Claims

1. A computer-implemented method for identifying a person using a facial recognition system, the method comprises: extracting an attribute of the person by applying a first representation of a first image of the person to a pre-trained attribute classifier machine learning model to generate an attribute classifier output; applying the attribute classifier output and a distance measurement output generated using a second representation of a second image of the person to a pre-trained fusion verification machine learning model, wherein the fusion verification machine learning model fuses the attribute classifier output and the distance measurement output to obtain a fused feature representation; using the pre-trained fusion verification machine learning model to generate a facial recognition system output; applying the facial recognition system output to a joint classification model; applying an auxiliary attribute to the joint classification model; using the joint classification model to generate a joint classification output; and using the joint classification output to control access to a secure asset.

2. The method according to claim 1, wherein, the attribute is a first attribute, wherein the pre-trained attribute classifier machine learning model is a pre-trained first attribute classifier machine learning model, and wherein the attribute classifier output is a first attribute classifier output, the method comprises: extracting a second attribute by applying the first representation of the first image to a pre-trained second attribute classifier machine learning model to generate a second attribute classifier output; and applying the second attribute classifier output to the pre-trained fusion verification machine learning model.

3. The method according to claim 1, wherein, the attribute comprises at least one of age, gender, race or head angle.

4. The method according to claim 1, wherein, the first representation of the first image comprises at least one vector.

5. The method according to any one of claims 1 to 4, wherein, the distance measurement output is generated by applying a first distance measurement to a distance measurement classifier, the method comprises: performing a face embedding; applying the face embedding to a classification pipeline to identify similar images; generating a second distance measurement based on the identified images; and applying the second distance measurement to the distance measurement classifier to generate the distance measurement output.

6. The method according to claim 1, wherein, the auxiliary attribute comprises a time attribute.

7. The method according to claim 1, wherein, the auxiliary attribute comprises the social pool of the person.

8. The method according to claim 1, wherein, the auxiliary attribute comprises the height of the person.

9. The method according to claim 1, wherein, the first representation of the first image is the same as the second representation of the second image.

10. The method according to claim 1, wherein, the first image is the same as the second image.

11. The method according to claim 1, wherein, using the joint classification output to control access to a secure asset comprises: controlling access to a secure entrance of a building.

12. The method according to claim 1, wherein, generating the joint classification output comprises: Use the joint classification model to generate joint probabilities.

13. The method according to claim 5, wherein, the auxiliary attribute includes a time attribute.

14. The method according to claim 5, wherein, the auxiliary attribute includes the social pool of the person.

15. The method according to claim 5, wherein, the auxiliary attribute includes the height of the person.

16. The method according to claim 1, wherein, the joint classification model includes a pre-trained machine learning model.

17. The method according to claim 5, wherein, using the joint classification output to control access to secure assets includes: controlling access to secure entrances of a building.

Citation Information

Patent Citations

  • Face authentication method and system, authentication equipment and nonvolatile storage medium

    CN108875514A

  • Face recognition method, device, electronic device and computer-readable medium

    CN109117808A

  • Similar face retrieval method and device and storage medium

    CN109993102A

  • Human face image retrieval method and apparatus

    WO2017107957A1

Cited By

  • Access control with face recognition and heterogeneous information

    US12518562B2