Systems and methods for utilizing person identifiability across a network of devices

By using machine learning's recognizability model to determine high-quality reference files in the device network, the cumbersome biometric identification registration process caused by the increase in the number of devices is solved, and fast and simple identity recognition and resource conservation are achieved.

CN114127801BActive Publication Date: 2025-09-02GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980098069.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-08-14
Publication Date
2025-09-02
Estimated Expiration
2039-08-14

AI Technical Summary

Technical Problem

As the number of intelligent computing devices grows, the registration process of biometric identification becomes time-consuming and cumbersome when executed on each device, and users want to simply expand identity identification in the device network without repeated registration.

Method used

By capturing user's reference files on one device and determining the recognizability score using machine learning recognizability models, selecting high-quality reference files for storage, and other devices compare them with reference information to achieve identity identification, avoiding repeated registrations on each device.

Benefits of technology

It realizes rapid and simple expansion of identity identification in the device network, reduces the redundancy consumption of computing resources, saves time and computing resources, and reduces the possibility of false positives and false negatives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114127801B_ABST
    Figure CN114127801B_ABST
Patent Text Reader

Abstract

The present disclosure relates to computer-implemented systems and methods for performing recognition on a network of devices. Generally, the systems and methods implement a machine-learned recognizability model that can process information such as a person's voice, facial features, or similar information to determine a recognizability score without having to generate or store biometric information that can be used to identify the person. The recognizability score can serve as a proxy for the quality of the information, serving as a reference for biometric recognition that can be performed on other devices in the device network. Thus, a single device can be used to register a person in the network (e.g., by capturing multiple photos of the person). Thereafter, the connection of other devices can utilize sensors (e.g., cameras) on the other devices to compare features of the reference information with input received by the sensors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to machine learning and more particularly to an enrollment process (e.g., using a machine learning model) that enables user identification to occur across a network of devices while limiting biometric analysis to certain trusted devices. Background Art

[0002] Biometric recognition, such as facial recognition, fingerprint recognition, and voice recognition, has been implemented in a variety of devices, including smartphones and personal home assistants. Typically, these recognition methods are used as a form of authentication to control access to the device or certain features of the device.

[0003] As the number of computing devices grows, particularly network-connected devices, which may be generally referred to as "smart" devices and / or the Internet of Things (IoT), there is a corresponding need to define access permissions on a per-device basis.

[0004] Typically, to implement biometric recognition, a user may participate in an enrollment process that may include generating one or more reference profiles of the user (e.g., a reference image, fingerprint scan, voice sample, etc.). However, as the number of intelligent computing devices grows, redundant execution of this enrollment process for each separate device may become time-consuming, cumbersome, or frustrating for the user. Consequently, when a user adds a new device to her device network, she may wish to simply extend the ability to recognize her identity to the new device without having to perform the enrollment process again.

[0005] What is needed in the art are methods and systems that can advantageously manage biometric recognition across a network of devices. Summary of the Invention

[0006] The present disclosure relates to computer-implemented systems and methods for performing recognition on a network of devices. Generally, the systems and methods implement a machine-learned recognizability model that can process information such as a person's voice, facial features, or similar information to determine a recognizability score without having to generate or store biometric information that can be used to identify a person. The recognizability score can be used as a proxy for information quality as a reference for biometric recognition that can be performed on other devices in the device network. Thus, a single device can be used to register a person in the network (e.g., by capturing multiple photos of the person). Thereafter, connections to other devices can utilize sensors (e.g., cameras) on the other devices to compare features of this reference information with input received by the sensors. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] A detailed discussion of embodiments for those skilled in the art is set forth in the specification with reference to the accompanying drawings, in which:

[0008] Figure 1A Depicted is a block diagram of an example computing system that performs recognition across a network of devices according to example embodiments of the present disclosure.

[0009] Figure 1B Depicted is a block diagram of an example computing device that may be used to implement identification and / or registration in identification according to example embodiments of the present disclosure.

[0010] Figure 1C Depicted is a block diagram of an example computing device that may be used to implement identification and / or registration in identification according to example embodiments of the present disclosure.

[0011] Figure 2 Depicted is a diagram of an example network of devices according to an example embodiment of the present disclosure.

[0012] Figure 3 Depicted is a block diagram of an example network of devices according to example embodiments of the present disclosure.

[0013] Figure 4 Depicted is a flow diagram of an example method for performing registration in a network of devices according to an example embodiment of the present disclosure.

[0014] Figure 5 Depicted is a block diagram showing an example process for training a recognizability model, according to an example embodiment of the present disclosure.

[0015] Repeated reference numbers across multiple drawings are intended to identify like features in the various embodiments. DETAILED DESCRIPTION

[0016] In general, the present disclosure relates to computer-implemented systems and methods for performing identification on a network of devices. In particular, as described above, when a user adds a new device to her network of devices, she may wish to simply extend the ability to recognize her identity to such new device without having to perform the registration process again. Various aspects of the present disclosure enable such a process by capturing and storing a user's reference files (e.g., a gallery of reference images) at one or more first devices (e.g., user devices such as smartphones and / or server computing systems). Thereafter, when the user wishes to extend identity recognition to a second device (e.g., a new home assistant device), the user is able to simply instruct the first device to share the reference files with the second device. In this way, the user can quickly and easily register the new device (e.g., enable the new device to perform the recognition process to recognize her) without having to perform the registration process of collecting reference files again. In addition, other aspects of the present disclosure relate to using machine learning models to facilitate the registration and identification processes. Specifically, aspects of the present disclosure may include using machine learning identifiability models (e.g., at or used by a first device, such as a user device and / or a server device), which enables the curation of high-quality reference files without computing biometric or other personally identifiable information about the user.

[0017] More specifically, according to one aspect of the present disclosure, one or more devices participating in a network may include and employ a machine-learned recognizability model that processes information such as a person's voice, facial features, or the like to determine a recognizability score without having to generate or store biometric information that can be used to identify the person. Generally, the recognizability score can be used as a proxy for information quality, serving as a reference for biometric recognition that can be performed on other devices in the device network.

[0018] Without agreeing on any one definition of quality or identifiability, generally these terms are used to indicate the condition of identifying data (image or sound) that displays sufficient detail to distinguish an individual. For example, the more information that is contained in an image or audio file about the individual performing the registration, the higher the quality of the file, in general. For example, an image file showing only the upper half of a face will be of lower quality than an image file showing the entire face. As another example, an audio file containing a voice recording obtained in a quiet room will be of higher quality than a voice recording obtained outdoors or in a crowded environment. Thus, in general, identifiability can be related to both the amount of data and properties of the data, such as low background relative to identifying features. For example, low identifiability may be associated with a lower amount of data and / or files displaying high background features.

[0019] Other definitions of identifiability can be associated with a query. As an example, a high identifiability score can be used to indicate that for a query signal with high identifiability and an unknown identity, when provided with a gallery of signals (images) of known identities, there is a greater probability (e.g., 75% or greater) that the identity can be accurately determined. This example can also be used to define an example of low identifiability. Thus, an identifiability score can be used to indicate the probability that an identity can be accurately determined from an image or other file.

[0020] Thus, in some embodiments, a newly captured reference document (e.g., an image captured by a user's device as part of an initial enrollment process) can be evaluated by a machine-learned recognizability model to determine a recognizability score that indicates the extent to which such a document (e.g., an image) is useful for identifying the individual depicted or referenced by the document. However, the recognizability score itself does not contain biometric information or other information capable of identifying an individual. Instead, the recognizability score simply indicates whether the document is useful for performing recognition via a separate recognition process, which can be performed by a different device (e.g., a "secondary" device to which the user later requests their identity be extended).

[0021] Based on the corresponding recognizability scores, certain newly captured reference files can be selected for inclusion in a reference file set, which will be used as reference files for subsequent recognition of the user. As an example, newly captured images (e.g., images captured by a user device as part of an initial registration process) can be evaluated by a machine-learned recognizability model to determine a recognizability score for each image. Images that receive a recognizability score that meets a certain threshold score (e.g., are determined to have high "recognizability") can be selected (e.g., by the user device and / or server device) and stored (e.g., by the user device and / or server device) in an image gallery associated with the user. However, importantly, while recognizability analysis can be used to establish a reference file set (e.g., to produce a high-quality reference set that only includes reference files that are useful for performing the recognition process), actual calculation of biometric information does not necessarily occur to generate the reference file set. Therefore, a high-quality reference set can be established even in situations where the first device (e.g., the user's device) is prohibited from calculating or storing biometric information (e.g., due to policy constraints, permissions, or other reasons).

[0022] When the user requests to do so, the gallery can then be shared with or made accessible to a new auxiliary device (e.g., a home assistant device) whose recognition capabilities the user wishes to expand. In particular, in some embodiments, the auxiliary device can include and / or employ a machine-learned recognition model to recognize the user based at least in part on a reference file (e.g., a gallery).

[0023] More specifically, another aspect of the present disclosure relates to using a machine-learned recognition model (separate from a recognizability model) that operates to identify an individual (e.g., through calculation or analysis of biometric information). Specifically, the auxiliary device may include one or more sensors (e.g., a camera, a microphone, a fingerprint sensor, etc.) that capture additional files (e.g., images, audio, etc.) that depict or otherwise represent a person. The auxiliary device may employ the machine-learned recognition model to analyze the additional files and reference files to determine whether the person represented by the additional files can be identified as a user. As an example, the machine-learned recognition model may be a neural network that has been trained (e.g., via a triplet training technique) to generate embeddings (e.g., in the final layer and / or in one or more hidden layers) that facilitate recognition. For example, a triplet training scheme may be used to train the machine-learned recognition model to generate corresponding embeddings for corresponding inputs, where the distance (e.g., L2 distance) between a pair of embeddings represents the probability that the corresponding pair of inputs (e.g., images) depict or otherwise reference the same person. Thus, in some implementations, a machine-learned recognition model can generate embeddings for the attached file and the reference file, and the corresponding embeddings can be compared to determine whether the person represented by the attached file can be identified as a user.

[0024] Another aspect of the present disclosure, described in further detail elsewhere herein, relates to training a machine-learned identifiability model based on a machine-learned recognition model using a distillation training technique. In particular, the distillation training technique exploits the fact that hidden layer outputs from one or more hidden layers of a machine-learned recognition model contain information about the identifiability of the input in addition to biometric information about the input. Furthermore, the calculation of a metric (e.g., a norm or other cumulative statistic) associated with the hidden layer outputs can remove or destroy biometric or personally identifiable information while preserving the identifiability information. Thus, in some embodiments, a machine-learned identifiability model can be trained to predict the norm or other metric of one or more hidden layer outputs from one or more hidden layers of a machine-learned recognition model. In this manner, a machine-learned identifiability model can be trained to produce an identifiability score that is indicative of identifiability, but does not include or contain biometric data or other personally identifiable information.

[0025] Thus, in some example embodiments, a single device can be used to register a person in a network (e.g., by capturing multiple photos of the person). Thereafter, other connected devices can utilize sensors (e.g., cameras) on the other devices to compare features of the reference information with input received by the sensors to perform recognition of the person.

[0026] Embodiments of the present disclosure may provide advantages for defining device access policies across a network of connected devices. This may be particularly useful as the number of Internet of Things (IoT) devices continues to increase and defining permissions on a per-device basis becomes increasingly cumbersome. Rather than registering each device with voice, facial, fingerprint, or other biometric recognition, a single registration may be performed to determine high-quality information to select as a reference. A person attempting to access one of the devices in the network may then undergo a recognition analysis (e.g., using a trained machine learning recognition model) that compares newly captured data obtained by such additional device to a reference file. In this manner, a user may avoid redundant execution of the registration process for multiple different devices. Eliminating redundant execution of the registration process may conserve computing resources (e.g., process usage, memory usage, network bandwidth, etc.) because the process is only performed once, rather than multiple times.

[0027] As an example for illustrative purposes, a person who wants to build a smart home that includes features such as a home assistant, keyless entry, and / or additional devices that utilize biometric features (e.g., fingerprint, eyes, face, voice, etc.) may want to set facial recognition as an access policy for interacting with each device or for accessing certain capabilities of the device. To complete the registration process on a device network, an individual can capture one or more images using a personal computing device (e.g., a smartphone) that includes software or hardware that implements the methods of the present disclosure. The personal computing device can apply a recognizability model to determine which of the one or more images (if any) to transmit as a reference file to a server or other centralized computing system (e.g., a cloud network). Generally, the centralized computing system can communicate with each device so that data can be transmitted between each device and the centralized computing system over a network (e.g., the Internet, Bluetooth, a local area network, etc.). Thereafter, access to each device can be performed according to the policy of each device. For example, accessing a device can include using a recognition model included in the device to compare input data received by a device sensor such as a camera in the case of facial recognition with one or more reference files.

[0028] Example embodiments of the present disclosure may include a method for registering a personal identity across a network of devices. Generally, the method includes obtaining a dataset that includes one or more files representing a person (e.g., fingerprints, images of eyes, faces, or similar information and / or voice recordings). Based on these one or more files, a machine-learned identifiability model (e.g., a distillation model) may determine an identifiability score for each of the one or more files by providing the files to the machine-learned identifiability model. Based at least in part on the identifiability scores, a portion of the dataset may be selected as a reference file stored on one or more devices. On this basis, attempting to access one of the devices included in the network may include an identification step. As an example, implementing the identification step may include obtaining sensor information describing the person attempting to access the device (e.g., using a camera or microphone). The sensor information may be compared to the reference file to determine whether the biometric information indicates a match, which match will allow access to the device, an application on the device, or a combination of both.

[0029] Aspects of a method for registering a personal identity may include obtaining a dataset comprising one or more files representing a person using a first device included in a network of devices. In some embodiments, the first device may include a personal computing device, such as a smartphone or personal computer, which may include built-in components, such as a camera or other image capture device and / or a microphone. Additional features of the first device may include an image processor that may be configured to detect the presence of one or more persons in an image. For brevity, embodiments of the present disclosure are discussed using a person as an example use case; however, this does not limit these or other embodiments to registering only a single person or images containing a single person. Image filters or other image processing accessible by one or more devices may be used to segment the image into personal identities (separate detected persons) for performing the registration.

[0030] Another aspect of registering a personal identity includes determining a recognizability score for each of the one or more files. In an example embodiment, the recognizability score may be determined using a recognizability model that has been trained using distillation and may be referred to as a distilled model. As an example, a recognizability model according to the present disclosure may include a distilled model trained from one or more outputs of one or more other neural networks. A distilled model may provide advantages such as lower computational cost, which may allow the distilled model to be executed on a personal computing device such as a laptop or smartphone.

[0031] Training a distillation model can include obtaining a neural network and / or one or more outputs of the neural network. By providing an input (e.g., a facial image) to the neural network, the neural network can be used to generate an output comprising one or more hidden layers. Since each hidden layer can include one or more features, a metric (e.g., a norm) can be calculated from the one or more hidden layers. Training the distillation model can then include optimizing an objective function for predicting the metric calculated from the one or more hidden layers determined for a given input.

[0032] For example, an example method for training a distillation model may include: obtaining a neural network configured to determine a series of hidden layers; determining a plurality of outputs by providing a plurality of inputs to the neural network, wherein each output is associated with a corresponding input, and wherein each output comprises a portion of the series of hidden layers; calculating a metric for at least one hidden layer included in the portion of the series of hidden layers; and training the distillation model to predict the metric based at least in part on receiving the corresponding inputs.

[0033] Aspects of the neural network may include a network configuration that describes the number of hidden layers that the neural network is configured to determine. For example, the neural network may be configured to determine at least three layers, such as at least 5 hidden layers, at least 7 hidden layers, at least 10 hidden layers, at least 20 hidden layers, and so on. Typically, the at least one hidden layer used to calculate the metric does not include the first or last layer of layers. Therefore, to train a distillation model, typically, an intermediate layer of the neural network may be selected to calculate the metric. As an example for illustration, the second-to-last layer (i.e., the layer next to the last) may be selected as the hidden layer for calculating the metric. Additionally, in some cases, the neural network may be configured to limit the outputs determined. For example, because an intermediate layer of the neural network may be selected to calculate the metric, subsequent layers of the neural network do not need to be calculated, and the neural network may be configured to stop determining other hidden layers or other outputs of the neural network.

[0034] Using a distilled model can provide certain advantages because the distilled model can perform identifiability analysis without generating biometric information that could otherwise be used to identify a person. This can provide advantages to users because they do not need to be familiar with the policies or capabilities of each device included in the device network. Instead, the user can allow each device to operate according to its own policies. Furthermore, the distilled model can provide a more lightweight implementation that can enable faster identification and / or selection of reference files on the user device.

[0035] Another example aspect of embodiments of the present disclosure may include selecting a portion of a dataset to store as a reference file based at least in part on a recognizability score. According to certain embodiments, the reference file can be accessed as a proxy for comparison with a person attempting to access one of the devices included in the network. Thus, in some cases, the selection can be optimized to reduce false positives (e.g., allowing a person to access a device when the person is not registered), false negatives (e.g., preventing a person from accessing a device when the person is already registered), or a combination of both. For example, embodiments of the present disclosure may provide advantages for reducing false negatives that may be caused by built-in image or voice comparison models present on the device the person is attempting to access. The recognizability model can determine or otherwise identify high-quality information representing the person during the registration process and, in some cases, may even indicate to the user attempting to register that none of the files included in the dataset meet a recognizability criterion or threshold. As another example, embodiments of the present disclosure may provide advantages for reducing false positives by selecting only high-quality images. For example, if a person registers a blurry image, the identifying information may be blurred, making it easier for a different person to access the device. Generally, the blurrier the image, the fewer identifying features it contains, resulting in a higher likelihood of false positives.

[0036] In some embodiments, the threshold value can be determined by a metric, such as a percentile, minimum, maximum, or other similar composite measure determined based on the identifiability scores of one or more files. Additionally or alternatively, the threshold value can include a preset value, and all or a set number of files that meet or exceed the value can be selected as part of the data set to be stored as reference files. Including a preset value can provide advantages in situations where the files captured during registration include low-quality data and a comparison between the identifiability scores of each file and the threshold value indicates that no score meets or exceeds the threshold value. In these situations, the device performing the registration can provide a prompt to the user, such as displaying a message on the device that the registration should be repeated or that additional files need to be included in the data set. Another exemplary advantage of performing the registration on the first device can include saving and / or reducing network traffic, as the first device can determine which (if any) files meet the threshold value for selection. Then, only those selected files can be transferred (e.g., to a second device in the device network), rather than transferring all the files obtained. For example, there may be a situation where no files meet the threshold value, and therefore no files need to be transferred to other devices included in the network.

[0037] For files with identifiability scores that meet or exceed a threshold, these can be transferred to a second device for storage as reference files. In some embodiments, the second device can include a server, cloud computing device, or similar device accessible by every device in the device network. Having such a centralized reference can provide advantages, such as easier registration updates for authorized device access personnel and / or reduced data storage.

[0038] As an example embodiment, a person attempting to access a device included in a device network and / or an operation / application executed by the device may undergo biometric analysis on the device. Biometric analysis may include accessing sensors included on the device to obtain a signal (e.g., video from a camera, audio from a microphone, etc.) containing information about the person attempting to access the device. This signal may be processed by a biometric analyzer, such as a machine learning recognition model trained to determine a set of features associated with the person (e.g., facial characteristics). The same biometric analyzer or a similarly trained biometric analyzer may process a reference file to determine a reference feature set. The two feature sets may then be compared, and based on the comparison, a response may be provided to the person attempting to access the device. For example, if the person attempting to access the device has already registered with the device network, the response may include opening the device's home screen or executing an operation / application included on the device. Alternatively, if the person attempting to access the device is not registered with the device network, the response may include prompting the person to register, providing an error message to the person, and / or sending a notification to the person who has already registered.

[0039] In general, a biometric analyzer may be included in one or more devices included in the device network and may be configured to perform biometric analysis according to a policy of the device. For example, a third device included in the device network may include a computer assistant, such as Google Home, or other similar device configured to receive natural language input and generate output based on the input. Each of these devices may include its own model (e.g., a machine learning recognition model) for performing biometric recognition. For example, a machine learning model may implement a neural network to generate an embedding that describes a representation of a person attempting to access the device. The devices may also include one or more sensors for obtaining a signal that includes information describing the person attempting to access the device.

[0040] As an example of technical effects and benefits, methods and systems for performing identification across a network of devices can provide greater control and reduce computing resources for managing and updating access policies. For example, time and computing resources can be saved by performing only one registration, rather than individually updating each device included in the network. In addition, a single registration can determine high-quality information, thereby reducing the need for re-registration or the possibility of false negatives or false positives. Similarly, in addition to during registration, the identifiability analysis described herein can also be performed at the time of identification (for example, by an auxiliary device such as a home assistant device). Using identifiability analysis at the time of identification can save computing resources by preventing the recognition analysis from being performed on low-quality files (for example, images) with low identifiability.

[0041] In general, embodiments of the present disclosure may include or otherwise access an identifiability model to perform identifiability analysis. For certain embodiments, the identifiability model may be trained using distillation and may be referred to as a distilled model. For example, an identifiability model according to the present disclosure may include a distilled model trained based on outputs from one or more neural networks. A distilled model may provide advantages such as lower computational cost, which may allow the distilled model to be executed on a personal computing device such as a laptop or smartphone. In particular, the distilled models described herein may be very fast and lightweight specialized models, thereby conserving computational resources such as processor and memory usage.

[0042] Referring now to the accompanying drawings, example embodiments of the present disclosure will be discussed in further detail.

[0043] Example devices and systems

[0044] Figure 1A A block diagram of an example computing system 100 capable of performing registration in a device network according to an example embodiment of the present disclosure is depicted. System 100 includes a user computing device 102, a server computing system 130, a training computing system 150, and an auxiliary computing device 170 communicatively coupled via a network 180.

[0045] The user computing device 102 can be any type of computing device, such as a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or game controller, a wearable computing device, an embedded computing device, a home assistant (e.g., Google Home or Amazon Alexa), or any other type of computing device.

[0046] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be one processor or multiple processors operatively connected. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0047] In some implementations, the user computing device 102 may store or include one or more identifiability models 120. For example, the identifiability model 120 may be or include various machine learning models, such as a neural network (e.g., a deep neural network) or other types of machine learning models, including nonlinear models and / or linear models. The neural network may include a feedforward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks.

[0048] In some implementations, one or more recognizability models 120 may be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 may implement multiple parallel instances of a single recognizability model 120 (e.g., perform parallel registration and / or determine recognizability scores across multiple instances of the recognizability model 120).

[0049] More specifically, the recognizability model can include a machine learning model that has been trained using distillation techniques to process recognition information, such as pixels of a person or face and / or a voice signal, to determine whether the information is recognizable. Generally, the person recognizability analyzer can be configured not to calculate or store any biometric information, such as facial embeddings, voice embeddings, facial landmarks (such as eyes or nose), or voice features (such as accents). This aspect of the recognizability model can be achieved by training the recognizability model to output a recognizability score corresponding to the quality of the input information.

[0050] Additionally or alternatively, one or more identifiability models 140 may be included in, or stored and implemented by, a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the identifiability models 140 may be implemented as part of a web service by the server computing system 140. Thus, one or more models 120 may be stored and implemented at the user computing device 102, and / or one or more models 140 may be stored and implemented at the server computing system 130.

[0051] In certain embodiments, the user computing device may also include a recognition model 124. The recognition model 124 may include a machine learning model (e.g., a trained neural network) for performing biometric recognition. Generally, the recognition model 124 differs from the identifiability model 120 in that the recognition model 124 may generate and / or store biometric information that can be used to identify an individual (e.g., facial features such as pupil distance). In some embodiments, the recognition model 124 may not be included as part of the user computing device 102. Instead, the user computing device 102 may access the recognition model 144 stored as part of another computing system, such as the server computing system 130.

[0052] The user computing device 102 may also include one or more user input components 122 for receiving user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other examples of user input components include a camera, a microphone, a traditional keyboard, or other means by which a user can provide user input.

[0053] The server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one or more processors operatively connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 executed by the processor 132 to cause the server computing system 130 to perform operations.

[0054] In some implementations, server computing system 130 includes or is implemented by one or more server computing devices. In instances where server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0055] As described above, the server computing system 130 may store or otherwise include one or more machine-learned identifiability models 140. For example, the model 140 may be or include various machine-learned models. Example machine-learned models include neural networks or other multi-layer nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks.

[0056] Additionally, in certain embodiments, the server computing system 130 may store or otherwise include one or more machine-learned recognition models 144. As described above, the recognizability model 140 and the recognition model 144 may be distinguished by their ability to store or generate biometric information. Generally, the recognizability model 140 may be used as a filter to determine whether the information provided to the model includes sufficient detail or quality to perform biometric recognition (e.g., using the recognition model 144).

[0057] The user computing device 102 and / or the server computing system 130 are communicatively coupled to the training computing system 150 via the network 180, and the user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 via interaction with the training computing system 150. The training computing system 150 can be separate from the server computing system 130, or can be part of the server computing system 130.

[0058] The auxiliary computing device 170 can be any type of computing device, such as a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or game controller, a wearable computing device, an embedded computing device, a home assistant (e.g., Google Home or Amazon Alexa), or any other type of computing device. Generally, the auxiliary computing device can include one or more processors 172, a memory 174, a recognition model 182, and a user input component 184. In an exemplary embodiment, the auxiliary computing device 170 can be an IoT device, which can include an AI assistant, such as Google Home. Furthermore, although shown as a single auxiliary computing device 170, the auxiliary computing device 170 can represent one or more connected devices that include a recognition model 182 for performing biometric recognition (e.g., facial recognition, voice recognition, fingerprint recognition, etc.). One aspect of the auxiliary computing device 170 is that this device does not need to include the identifiability model 120 or 140 for determining the identifiability score. Instead, secondary computing device 170 may access reference files (e.g., data 136 stored on server computing system 130 or data 116 stored on the user computing device) that are selected based at least in part on the identifiability scores determined by identifiability models 120 and / or 140 included in user computing device 102 and / or server computing system 130. In this manner, a user attempting to access secondary computing device 170 does not need to perform registration for each secondary computing device 170.

[0059] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one or more processors operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 can store instructions 158 and data 156 executed by the processor 152 to enable the training computing system 150 to perform operations. In some embodiments, the training computing system 150 includes or is implemented by one or more server computing devices.

[0060] The training computing system 150 may include a model trainer 160 that trains the machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as back propagation of errors. In some embodiments, performing back propagation of errors may include performing truncated back propagation through time. The model trainer 160 may perform various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the trained model.

[0061] In particular, model trainer 160 can train recognizability models 120 and / or 140 based on a set of training data 162. Training data 162 can include, for example, outputs from one or more machine learning models, such as models configured to perform facial or voice recognition. These one or more machine learning models can include a neural network configured to generate three or more hidden layers. In an example embodiment, recognizability models 120 and / or 140 can be trained using features of the hidden layers generated by one or more neural networks rather than the outputs of the neural networks. Additionally, in some cases, a metric (e.g., a norm) can be used to summarize the features of the hidden layers, and the recognizability models 120 and / or 140 can be trained using training data 162 including the metric. For example, a distilled model for facial recognition can be learned by taking a small thumbnail image as input and regressing directly to a network that determines a metric (e.g., an L2 norm value) from the penultimate hidden layer.

[0062] In some embodiments, if the user has provided consent, the training examples may be provided by the user computing device 102. Thus, in such embodiments, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some instances, this process may be referred to as personalized modeling.

[0063] The model trainer 160 includes computer logic for providing the desired functionality. The model trainer 160 can be implemented using hardware, firmware, and / or software that controls a general-purpose processor. For example, in some embodiments, the model trainer 160 includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In other embodiments, the model trainer 160 includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium, such as a RAM hard disk or optical or magnetic media.

[0064] Network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. In general, communications on network 180 can be carried via any type of wired and / or wireless connection, using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0065] Figure 1A An example computing system that can be used to implement the present disclosure is shown. Other computing systems may also be used. For example, in some embodiments, the user computing device 102 may include a model trainer 160 and a training dataset 162. In such embodiments, the model 120 can be trained and used locally at the user computing device 102. In some embodiments, the user computing device 102 may implement the model trainer 160 to personalize the model 120 based on user-specific data.

[0066] Figure 1B Depicted is a block diagram of an example computing device 10 capable of performing registration across a network of devices according to an example embodiment of the present disclosure. Computing device 10 may be a user computing device or a server computing device.

[0067] Computing device 10 may include multiple applications (e.g., applications 1 to N). Each application can include its own machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include text messaging applications, personal assistant applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.

[0068] like Figure 1B As shown in , each application can communicate with multiple other components of the computing device, such as one or more sensors, context managers, device state components, and / or additional components. In some embodiments, each application can use an API (e.g., a public API) to communicate with each device component. In some embodiments, the API used by each application is specific to that application.

[0069] Figure 1C Depicted is a block diagram of an example computing device 50 performing in accordance with an example embodiment of the present disclosure. Computing device 50 may be a user computing device or a server computing device.

[0070] The computing device 50 includes a plurality of applications (e.g., applications 1 through N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like. In some embodiments, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a public API across all applications).

[0071] The central intelligence layer includes many machine learning models. For example, Figure 1C As shown in , a corresponding machine learning model (e.g., model) can be provided to each application and managed by the central intelligence layer. In other embodiments, two or more applications can share a single machine learning model. For example, in some embodiments, the central intelligence layer can provide a single model (e.g., a single model for all applications). In some embodiments, the central intelligence layer is included in or implemented by the operating system of the computing device 50.

[0072] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data repository for computing devices 50. Figure 1C As shown in , the central device data layer can communicate with many other components of the computing device, such as one or more sensors, context managers, device state components, and / or additional components. In some embodiments, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0073] Example model layout

[0074] Figure 2A diagram of an example device network according to an example embodiment of the present disclosure is depicted. As shown, the device network can include at least three devices, such as a mobile computing device 202, a cloud or server computing device 203, and an auxiliary or secondary device 205, such as a computer assistant device. The secondary device 205 can also include a sensor 206, such as a camera or microphone, for acquiring information (e.g., new files, such as new images). In an example embodiment, a person 201 performing registration in the device network can use the mobile computing device 202 to obtain a data set including one or more files representing the person 201. For example, these files can include pictures, sounds, or other identifying information. At the mobile computing device 202 or the cloud computing device 203, an identifiability model can be used to determine which files, if any, should be transmitted over the communication network 204 to be stored as reference files on the cloud computing device 203. After registration, when person 201 requests to register another device included in the network, such as computer assistant device 205, computer assistant device 205 can access or receive reference files from mobile computing device 202 and / or cloud computing device 203 to perform biometric analysis (e.g., using a machine learning recognition model).

[0075] Figure 3 Depicted is a block diagram of an example network of devices according to example embodiments of the present disclosure. Figure 3 Provided Figure 2 In the example case, each of at least three devices is shown to include certain components or perform certain operations. Figure 3 , a mobile computing device 300 is shown as including an image capture device 301 for obtaining images 302 representing people performing registration in a device network. These images 302 can be provided to an image processor 303 to identify or group the images 302 into detected people 304 if the images 302 contain more than one person. For example, the image processor 303 can apply an object detection model or process to detect people in the images 302.

[0076] The grouping of detected persons 304 may then be provided to a person recognizability analyzer 305, such as a machine learning distillation model or recognizability model as described herein. Based at least in part on the recognizability scores determined by the person recognizability analyzer 305, a person image selector 306 may separately determine images and selected persons to transmit to a cloud computing device 320 as reference images 322 to be included in a gallery 321 that may be created for a particular user or person. Figure 3306 , but the person recognizability analyzer 305 and the person image selector 306 can be implemented as a single operation of the recognizability model and the logic associated therewith. Similarly, although components 303-306 are shown at the mobile computing device 300, some or all of these components can alternatively be included at or executed at the cloud computing device 320.

[0077] Figure 3 Also depicted is a third device, shown as a computer assistant device 310. The device 310 is shown to include an image capture device 311 that can be used to obtain an additional image 312 representing a person attempting to access the device 310 or an application executed by the device 310. The device 310 also includes a person biometric analyzer 315 that can perform biometric analysis on the images (e.g., image 312 and / or image 322) to analyze biometric information associated with the images. For example, the person biometric analyzer 315 can include or employ a machine learning recognition model as described herein. An example recognition model is FaceNet, its variants, and similar models. See FaceNet: A Unified Embedding for Face Recognition and Clustering by Schroff et al. (https: / / arxiv.org / ABS / 1503.03832), which provides an example triplet training procedure that can be used to train a recognition model to produce pairs of embeddings for pairs of inputs, where the distance directly corresponds to a measure of facial similarity in the inputs.

[0078] While the computer assistant device 310 is shown as including an image processor 313 to detect one or more persons 314, these elements need not be present, and the image 312 captured by the image capture device 311 may be directly input to a person biometric analyzer 315 to determine person appearance biometrics, such as the embeddings, measurements, or locations of unique features, etc. The same or a different biometric analyzer 315 may be used to process a user reference image 322 to determine biometric information 316 from a gallery of user images 321, which may be compared to person appearance biometrics 317 using, for example, a person appearance identifier (e.g., which may compare corresponding embeddings (e.g., distances therebetween), corresponding features, etc.), to generate a confidence score for identifying whether certain persons depicted in the image 312 are also included in the gallery of user images 321.

[0079] Example Method

[0080] Figure 4Depicted is a flow chart of an example method performed according to an example embodiment of the present disclosure. Although for purposes of illustration and discussion, Figure 4 The steps are depicted as being performed in a particular order, but the method of the present disclosure is not limited to the specific order or arrangement shown. The various steps of method 400 can be omitted, rearranged, combined and / or adapted in various ways without departing from the scope of the present disclosure.

[0081] At 402, a computing system may obtain a dataset comprising one or more files representing people on a first device. The first device may comprise a personal computing device, such as a smartphone or a personal computer, having built-in components, such as a camera or other image capture device and / or a microphone. Additional features of the first device may include an image processor configured to detect the presence of one or more people in an image.

[0082] At 404, the computing system may determine a recognizability score for each file by providing each of the one or more files to a distillation model that has been trained using metrics calculated from one or more hidden layers of a neural network. Generally, the recognizability score may be calculated before transferring the file to the second device. Thus, the recognizability model may be implemented on the first device, or accessed by the first device, to determine the recognizability score. While preferably minimizing storage and computational costs, the cloud service may automatically upload any files generated on the first device to the second device (e.g., a server). Thus, in some embodiments, determining the recognizability score may be performed on the second device.

[0083] At 406, the computing system may select a portion of the dataset to store as a reference file based at least in part on the recognizability score. Generally, selecting a portion of the dataset to store as a reference file may include transferring the reference file to the second device. Alternatively or additionally, the selection may include specifying a reference location for storing the reference file, such as a gallery or record of user images accessible to other devices included in the network. In this manner, files uploaded directly to the second device may be filtered so that only the specified reference file is accessible during biometric identification when a person attempts to access a device included in the network.

[0084] Figure 5

[0014] Example aspects of certain methods and systems according to the present disclosure are shown.For some implementations, the methods and systems may include a trained identifiability model and / or a trained identifiability model. Figure 5 A block flow diagram showing an example method for training a recognizability model 500 according to the present disclosure is shown. Figure 5A plurality of inputs 502 are shown being provided to a recognition model 506, which is configured as a neural network including a plurality of hidden layers 508. The recognition model 506 may generate the plurality of hidden layers 508 based in part on one of the inputs 504 provided to the recognition model 506. One or more hidden layers (e.g., hidden layer N 508) may then be extracted to determine a metric 512, such as the norm of the features included in the hidden layer 508. Continuing this process for each input 504 included in the plurality of inputs 502 may generate a calculated metric for each input. The set 514 of inputs and calculated metrics may then be used to train a recognizability model using a distillation technique. In this manner, the recognizability model may be trained to determine the calculated metric 512 based at least in part on the corresponding input received for determining the metric 512. For some embodiments, the recognition model 506 may be configured to not determine any further hidden layers 508 or outputs 510 after generating the hidden layer 508 used to generate the metric 512. Therefore, the recognition model 506 used during training of the identifiability model 500 does not need to be the same as Figure 1A The identification models included in the device network shown are identical.

[0085] Additional Disclosure

[0086] The technology discussed herein relates to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and received from these systems. The inherent flexibility of computer-based systems allows for a variety of possible configurations, combinations, and divisions of tasks and functions between and among components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0087] Although the present invention has been described in detail with reference to various specific example embodiments of the present invention, each example is provided as an illustration, not as a limitation of the present invention. Those skilled in the art, after understanding the foregoing, can easily make changes, modifications, and equivalents to these embodiments. Therefore, the present invention discloses that such modifications, variations, and / or additions to the present invention are not excluded, and it is clear to those skilled in the art. For example, a feature shown or described as part of one embodiment can be used together with another embodiment to produce yet another embodiment. Therefore, the present invention is intended to cover these changes, variations, and equivalents.

Claims

1. A computing system comprising: A registration device comprising one or more non-transitory computer-readable media and one or more processors collectively storing instructions that, when executed by the one or more processors, configure the registration device to: obtaining a plurality of images depicting a user going through a registration process; processing each of the plurality of images using the machine-learned recognizability model to determine a corresponding recognizability score for each image as an output of the machine-learned recognizability model, wherein the recognizability score for each image indicates the recognizability of a user depicted in the image and does not include biometric information associated with the user; selecting at least one image from the plurality of images for inclusion in a gallery associated with the user based at least in part on respective identifiability scores of the plurality of images; and The gallery is transmitted directly or indirectly to one or more secondary computing devices for use in identifying the user by the one or more secondary computing devices.

2. The computing system of claim 1 , further comprising: The one or more auxiliary computing devices are configured to: Receive and store the gallery; obtain additional images depicting people; as well as The attached image is compared to the gallery to determine whether the person depicted in the attached image is the user.

3. A computing system as claimed in any preceding claim, wherein: The one or more auxiliary computing devices include a server computing device.

4. The computing system of claim 1 or 2, wherein: The one or more auxiliary computing devices include a computer assistant device.

5. The computing system of claim 1 or 2, wherein: The one or more auxiliary computing devices include a server computing device configured to: Receive gallery from registered devices; as well as The gallery is selectively forwarded to the one or more additional devices in response to a request from the user to register the one or more additional devices with a user account associated with the user.

6. The computing system of claim 1 or 2, wherein: The registered device includes a user device associated with a user.

7. The computing system of claim 1 or 2, wherein: The registration device includes a server computing device, and wherein the server computing device obtains the plurality of images from a user device that captured the plurality of images and is associated with a user.

8. The computing system of claim 1 or 2, wherein: Each of the one or more auxiliary computing devices is configured to process each image included in the gallery using a machine-learned facial recognition model that obtains a facial embedding of the image, the facial embedding including biometric information associated with the user.

9. The computing system of claim 1 or 2, wherein: The machine-learned recognizability model has been learned via a distillation training technique, wherein the machine-learned recognizability model is trained to predict the norm of hidden layer outputs generated by a hidden layer of a machine-learned facial recognition model configured to produce a facial embedding of an input image.

10. A computer-implemented method for registering a personal identity across a network of devices, the method comprising: obtaining, by one or more computing devices, a dataset comprising one or more files representing a person on a first device; determining, by one or more computing devices, a recognizability score for each of the one or more files by providing each file to a machine-learned distillation model, wherein the distillation model has been trained using a metric computed from one or more hidden layers of a neural network, wherein the recognizability score for each file indicates recognizability of a person and does not include biometric information associated with the person; and A portion of the dataset is selected, by one or more computing devices, to be stored as a reference profile for the person based at least in part on the recognizability score.

11. The computer-implemented method of claim 10, wherein: Selecting a portion of the dataset to store as a reference file includes: comparing, by the one or more computing devices, the identifiability score of each of the one or more files to a threshold value; and When none of the identifiability scores meet the threshold: providing, by one or more computing devices, a prompt on the first device requesting the character to generate an additional file; When the identifiability score of one or more files included in the dataset meets the threshold: The file is transferred to the second device by one or more computing devices.

12. The computer-implemented method of claim 11 , wherein: The second device includes a cloud computing device or a server computing device, and wherein the second device communicates with at least one other device included in the device network via the communication network.

13. The computer-implemented method of any one of claims 10-12, further comprising: Attempting access by one or more computing devices to one of the devices included in the device network, an operation performed by one of the devices, or both, wherein the attempted access includes performing, by the one or more computing devices, a biometric analysis, the biometric analysis comprising: obtaining, by one or more computing devices, a signal including information representing a person; accessing the reference file by one or more computing devices; comparing, by one or more computing devices, the reference file to the signal; and A response is provided by one or more computing devices allowing or denying the access attempt based at least in part on the comparison of the reference file to the signal.

14. The computer-implemented method of claim 13, wherein: Obtaining, by the one or more computing devices, a signal including information representing a persona includes obtaining, by a third device, a signal including information representing a persona.

15. The computer-implemented method of claim 14, wherein: The third device includes a computer assistant configured to receive input including at least one of visual, audio, or textual input; and provide output based at least in part on the input.

16. The computer-implemented method of claim 13, wherein: Comparison of the reference file with the signal includes: A set of biometric information is determined by one or more computing devices by providing a reference document to a machine learning model.

17. The computer-implemented method of claim 16, wherein: The machine-learned model includes a neural network, and the set of biometric information includes embeddings generated by the neural network.

18. The computer-implemented method of claim 10, wherein: The first device comprises a mobile computing device.

19. The computer-implemented method of claim 10, wherein: The first device includes a computer assistant configured to receive input including at least one of visual, audio, or text; and provide output based at least in part on the input.

20. The computer-implemented method of claim 10, wherein: The one or more files include audio, video, photos, or a combination thereof.

21. The computer-implemented method of claim 10, wherein: The first device is prohibited from computing the biometric identifier.

22. The computer-implemented method of claim 21, wherein: The biometric identifier includes an embedding generated by a recognition neural network.

23. The computer-implemented method of claim 10, wherein: The distillation model is trained using a training method, the training method comprising: Obtaining, by one or more computing devices, a recognition neural network trained to compute a series of hidden layers upon receiving an input; determining, by one or more computing devices, a plurality of outputs by providing a plurality of inputs to a recognition neural network, wherein each output of the plurality of outputs is associated with a corresponding input, and wherein each output comprises at least one intermediate output from at least one hidden layer in a series of hidden layers; computing, by one or more computing devices, a metric of at least one intermediate output from at least one hidden layer in the series of hidden layers for each output; and A distillation model is trained, by one or more computing devices, to predict a metric based at least in part on receiving input for determining at least one intermediate output for computing the metric.

24. The computer-implemented method of claim 23, wherein: The metric comprises a norm of at least one intermediate output.

25. The computer-implemented method of claim 23 or 24, wherein: The recognition neural network is configured to determine three or more hidden layers, and wherein at least one hidden layer used to calculate the metric does not include a first layer or a last layer of the three or more hidden layers.

26. The computer-implemented method of claim 23 or 24, wherein: The recognition neural network is configured to determine that there are no additional hidden layers following the at least one hidden layer used to calculate the metric.

27. A computer system comprising: Memory; as well as A processor configured to execute the method according to any one of claims 10 to 26.

28. A computer-implemented method comprising performing any of the operations described in any of claims 1-9.

29. One or more non-transitory computer-readable media storing instructions for performing any of the operations described in any of claims 1-26.