System for biometric identity enrollment

By enabling enrollment through a generic device and matching pre-enrollment representations with biometric data, the system addresses the inconvenience of traditional enrollment processes, enhancing user experience and operational usability of biometric systems.

JP2026503066AInactive Publication Date: 2026-01-27AMAZON TECH INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025540200
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-10
Filing Date
2024-01-04
Publication Date
2026-01-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional biometric identification systems require users to undergo inconvenient enrollment processes that involve specific interactions at designated locations, limiting system usage and user experience.

Method used

A system that allows users to initiate the enrollment process using a generic input device, such as a smartphone, to capture biometric data in multiple modalities, which is then processed to create a pre-enrollment representation that can be matched with biometric input data from specialized devices, enabling enrollment without additional steps at the time of initial use.

Benefits of technology

This approach simplifies the enrollment process, improving user convenience and allowing biometric systems to be used operationally during initial interactions, reducing the need for separate enrollment steps and enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026503066000001_ABST
    Figure 2026503066000001_ABST
Patent Text Reader

Abstract

User enrollment in a biometric identification system begins with a pre-enrollment process on a selected generic input device (GID), such as a smartphone. The user enters identifying data, such as their name, and may use the GID's camera to capture first image data, such as of their hand. The first image data is processed to determine a first representation. Upon presenting the hand at the biometric input device, second image data is captured. The second image data is processed to determine a second representation. If the second representation is deemed associated with the first representation, the enrollment process may be completed by saving the second representation for later use.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. patent application Ser. No. 18 / 152,403, filed Jan. 10, 2023, entitled "System for Biometric Identification Enrollment," the contents of which are incorporated herein by reference. [Background technology]

[0002] Biometric input data can be used to recognize and assert a user's identity.

[0003] The detailed description will be set forth with reference to the accompanying drawings. In the drawings, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. Use of the same reference number in different figures indicates similar or identical items or features. The figures are not necessarily drawn to scale, and in some figures, proportions or other aspects may be exaggerated to facilitate an understanding of particular aspects. [Brief explanation of the drawings]

[0004] [Figure 1] 1 illustrates a biometric identification system that provides a pre-registration process, according to some embodiments. [Figure 2] 1 illustrates supplemental data used by the system, according to some embodiments. [Figure 3] 1 illustrates labeled training data for training a cross-processing module that may be used for pre-registration, according to some embodiments. [Figure 4] 1 illustrates a block diagram of processing modules, including an intersection processing module, during training, according to some embodiments. [Figure 5] FIG. 1 illustrates a block diagram of a loss function used during training of the intersection processing module, according to some embodiments. [Figure 6]1 illustrates a block diagram of an intersection processing module during inference, according to some embodiments. [Figure 7] FIG. 1 is a block diagram of a cross-comparison module that may be used for pre-registration, according to some embodiments. [Figure 8] 10 illustrates processing training data to determine transformer training data, according to some embodiments. [Figure 9] 1 illustrates a transformer module during training, according to some embodiments. [Figure 10] FIG. 10 is a block diagram of using a transformer module and a comparison module for pre-registration, according to some embodiments. [Figure 11] FIG. 1 is a block diagram of a computing device for implementing a system according to some embodiments. [Figure 12A] 1 illustrates a process flow diagram for performing a pre-registration process, according to some embodiments. [Figure 12B] 1 illustrates a process flow diagram for performing a pre-registration process, according to some embodiments. [Figure 12C] 1 illustrates a process flow diagram for performing a pre-registration process, according to some embodiments. [Figure 13] 1 is a flow diagram of a process for performing a pre-registration process using cross-comparison, according to some embodiments. [Figure 14] 1 is a flow diagram of a process for performing a pre-registration process using cross-comparison, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0005] Although embodiments are described herein by way of example, those skilled in the art will recognize that the embodiments are not limited to the examples or figures described. The drawings and their detailed description are not intended to limit the embodiments to the particular forms disclosed; on the contrary, it should be understood that the invention is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope as defined by the appended claims. The headings used herein are for organizational purposes only and are not meant to be used to limit the scope of the description and claims. As used throughout this application, the term "may" is used in the permissive sense (i.e., meaning having the potential to), rather than the essential sense (i.e., meaning required). Similarly, the terms "include," "including," and "includes" mean including, but not limited to,

[0006] Biometric identification offers several advantages in a wide variety of use cases: for example, biometric identification can be used to facilitate payments at the point of sale, to provide access to facilities, etc.

[0007] Biometric identification may use various types of biometric input data as input. In one embodiment, the biometric input data may include an image of a hand. The biometric input data may be acquired using one or more modalities. For example, a first modality may include an image of the epidermis of a user's palm, and a second modality may include an image of subcutaneous features, such as veins, in the user's palm. Using input data including two or more modalities provides several advantages to biometric identification. One advantage is the potential for decorrelation between modalities, which may improve overall accuracy. For example, each modality may provide different features, providing clearer information that may better distinguish one person from another and determine the presence of artifacts, such as realistic-looking face masks or fake hands. Some features may be common to different modalities, while other features vary. For example, the overall shape of the hand may be apparent across modalities, but information such as friction ridges, veins, etc. may differ between images of different modalities acquired from the same hand.

[0008] During operation, the biometric identification system acquires input data using these different modalities. During the enrollment process, a user opts in to participate in using the system. The user may provide identification data such as their name, account credentials, phone number, etc. The user provides biometric input data using a biometric input device. For example, the user may present their hand to the biometric input device, which can acquire biometric input image data. This data includes image data acquired using various modalities. The biometric input image data is processed to determine an enrollment representation. The enrollment representation is stored for later use. Identification data is also stored and associated with the enrollment representation. A match of a query expression to the enrollment representation can later be used to assert the identity represented by the associated identification data.

[0009] In addition to capturing biometric input image data, biometric input devices may include other features such as robust tamper-resistant elements, liveness detection hardware for determining whether the object being imaged is a living human, etc. Deployment of biometric input devices may be limited to businesses, institutions, or other organizations. In comparison, general-purpose input devices such as mobile phones, tablet computers, and consumer electronic devices such as smart displays with cameras are widely available.

[0010] Traditionally, the enrollment process in biometric identification systems requires specific interactions that can be inconvenient for users. For example, a user may be required to go to a designated location where a biometric input device is located, submit identification data, and capture biometric input image data using the biometric input device. In another example, a user may provide identification data using a separate device, receive a quick response (QR) code or other token, and present it at a designated location. Continuing the example, this reduces the need to capture identification data at a designated location, but still requires the submission and capture of biometric input image data as a separate process. The designated location and time associated with completing the enrollment process can be inconvenient for users and can result in an adverse user experience. This can limit system usage.

[0011] Described in this disclosure are techniques and systems for biometric identity enrollment. Initially, a user initiates the pre-enrollment process using a generic input device. For example, the generic input device may be used to acquire pre-enrollment identification data from the user and generic input image data for the user. The generic input image data may be acquired using the same or a different modality as the biometric input device. For example, the generic input device may include a visible light camera that acquires red-green-blue (RGB) images depicting a hand in visible light. Continuing the example, the biometric input device may include an infrared (IR) camera that acquires infrared images depicting a hand in infrared light.

[0012] The generic input image data is processed to determine a pre-registered representation. The pre-registered representation is stored and associated with the pre-registered identification data. In some implementations, the pre-registered data, or portions thereof, may expire after a specific time or event. For example, a pre-registered representation may be deleted if not used within 10 days.

[0013] At a later time, the user presents their hand to the biometric input device. For example, the user may wish to make a payment at a point of sale or be granted access to a facility. The biometric input device obtains biometric input image data, which is processed to determine a second representation. The second representation may be compared to previously stored enrollment representation data. If no match is found, the second representation is compared to the pre-enrollment representation.

[0014] The pre-enrollment representation data and the representation data associated with the biometric input device may be represented in different embedding spaces. This may be due to differences in modality, processing techniques used to determine the representation, etc. Several techniques may be used to perform the comparison between the second representation and the pre-enrollment representation.

[0015] In one embodiment, the trained machine learning network accepts a generic input image and a biometric input image as input. A first intersection representation is determined based on the generic input image, and a second intersection representation is determined based on the biometric input image. The intersection representations are in the same embedding space and can be compared. If the two intersections are less than a threshold distance in the embedding space, they can be considered to represent the same hand. As a result, the biometric input image can be associated with the generic input image. The second representation can then be saved as an enrollment representation, associated with identification data, and later used for identification.

[0016] In another embodiment, the trained machine learning network accepts a generic input image as input and determines transformed representation data that is in the same embedding space as the second representation. The intersection representation is then in the same embedding space and can be compared. If the two intersections are less than a threshold distance in the embedding space, they can be considered to represent the same hand. As a result, the biometric input image can be associated with the generic input image. The second representation can then be saved as an enrollment representation, associated with identification data, and later used for identification.

[0017] In some situations, during this initial use of the biometric input device, the user may be prompted to provide additional information. For example, the user may be asked to enter a phone number, enter an account name, provide an EMV or other card to be electronically read (or "dipped"), enter data sent to a related but different device, such as a code sent to a mobile phone, answer authentication questions, etc.

[0018] By using these techniques and systems, from the user's perspective, interactions related to the registration process are limited to operations on a general input device. When a pre-registered user proceeds to initial use of a biometric input device, that use will obtain a biometric input image and other data, such as liveness detection, without requiring a separate series of steps or operations for the user. Instead of requiring the user to go to a designated location, such as a specific kiosk or customer service area, the user can use any participating biometric input device in the biometric identification system to complete the registration process with little or no additional effort.

[0019] The user experience and utilization of the biometric input device is also improved because the user can utilize the system in an operational capacity during the initial use of the biometric input device. For example, a transaction associated with the user, such as a purchase or opening a door, may be completed during that initial use of the biometric input device. From the user's perspective, the user can present their hand and the intended transaction, such as a purchase or opening a door, is completed.

[0020] As noted above, in some situations during initial use, the user may be prompted to provide additional information at the biometric input device. This interaction may be minimal and easily performed by the user. For example, the user may be accustomed to presenting or "dipping" an EMV card at the point of sale. Later during use of the system, the user may complete a transaction simply by presenting their hand.

[0021] Exemplary System 1 illustrates a biometric identification system 100 that provides a pre-registration process, according to some embodiments. System 100 is described as being used as part of a biometric identification system that determines the identity of a user. However, the systems and techniques described herein may be used in other contexts.

[0022] At a first time, the user's hand 102 is shown positioned over the generic input device 170. During the first time, the user may perform a pre-registration process described below.

[0023] The generic input device 170 may comprise a home electronic device such as a mobile phone, a tablet computer, a smart display, a doorbell security camera, a robot, or the like. The generic input device 170 may include a computing device and a sensor, such as a camera. The generic input device 170 is used to acquire generic input image data 172 or other input data. The generic input image data 172 may include visible image data 176 representing a visible light image. The generic input image data 172 may include a single image, a series of images, video data, or other image data. The camera may be configured to acquire images in a first modality. For example, the camera of the generic input device 170 may acquire a red-green-blue (RGB) image using visible light. In operation, as described in more detail below, the generic input image data 172 is processed to determine pre-registered representation data 182. The generic input image data 172 is illustrated for ease of explanation, but not necessarily by way of limitation, as having a single visible image data 176. In some implementations, the generic input image data 172 may include other modalities. For example, general input image data 172 or other input data may include LIDAR data of hand 102, point cloud data of hand 102, different images captured while illuminated with different colored lights, data captured from a fingerprint sensor, etc. In another example, other input data captured by general input device 170 may include other types of data, such as audio data containing speech from a user. In some implementations, other input data may include other images of the user, such as an image of the user's face.

[0024] Pre-registration identification data 184 may also be determined. Pre-registration identification data 184 may include one or more of a name, account credentials, a phone number, etc. In some implementations, pre-registration identification data 184 may include data obtained during operation of general input device 170. For example, a user may enter account login credentials into an application executing on general input device 170. Pre-registration identification data 184 may be determined based at least in part on a successful login.

[0025] Generic input device 170 may also determine supplemental data 174(1), which may include one or more of system metadata or user data. Supplemental data 174 is discussed in further detail with respect to FIG. 2.

[0026] At a second time, such as after the user completes the pre-registration process described below, the user can proceed to use one or more biometric input devices 104. The biometric input device 104 may include a computing device 106 and a camera 108. The camera 108 has a field of view (FOV). During operation of the biometric input device 104, the camera 108 acquires images of an object within the FOV, such as the hand 102, using one or more modalities to provide biometric input image data 112(1). The images may include a single image, a series of images, video data, or other image data. The biometric input device 104 may include other components not shown. For example, the biometric input device 104 may include a light to illuminate objects within the FOV, a communication interface, etc.

[0027] As shown for a second time, the user utilizes biometric input device 104(1). The user's hand 102 is shown positioned above biometric input device 104(1). In other implementations, other configurations may be used. For example, camera 108 of biometric input device 104(2) has a FOV that widens downward, and the user places hand 102 within the field of view below biometric input device 104(2).

[0028] In some implementations, the biometric input device 104 can capture other input data. For example, the biometric input device 104 can capture an image of the user's face. In another example, the biometric input device 104 can capture audio data including speech from the user.

[0029] During operation, biometric input device 104 may also determine supplemental data 174(2). For example, supplemental data 174(2) may include additional information such as results from liveness detection, data obtained from electronically reading an EMV card provided by a user, etc.

[0030] Once the user completes the enrollment process at the second time, at the third time the user may continue to use one or more biometric input devices 104 of the system 100 .

[0031] In one implementation, the biometric input device 104 is configured to acquire images of the hand 102 illuminated using infrared light having two or more particular polarizations, such as with different illumination patterns. For example, during a motion, the user may present the hand 102 with the palm or palmar region facing the biometric input device 104. As a result, the biometric input image data 112 provides an image of the front of the hand 102. In other implementations, the biometric input image data 112 may include the back of the hand 102. Separate images may be acquired using different combinations of polarizations provided by the infrared light.

[0032] Depending on the polarization used, the image generated by the biometric input device 104 may be of a first modality feature or a second modality feature. The first modality may utilize an image acquired by the camera 108 in which the hand 102 is illuminated with light having a first polarization and a polarizer passes light also having the first polarization to the camera 108. The first modality feature may include features near or on the surface of the user's hand 102. For example, the first modality feature may include surface features such as creases, wrinkles, scars, dermal papillae, etc. in at least the epidermis of the hand 102. The image acquired using the first modality may be associated with one or more surface features.

[0033] The second modality features include features below the epidermis. The second modality may utilize images acquired by a camera 108 in which the hand 102 is illuminated with light having a second polarization and a polarizer that passes light having a first polarization through the camera 108. For example, the second modality features may include subcutaneous anatomical structures such as veins, bones, and soft tissue. Some features may be visible in both the first modality image and the second modality image. For example, a palmar crease may include superficial first modality features as well as deeper second modality features within the palm. Images acquired using the second modality may be associated with one or more subcutaneous features.

[0034] Separate images of the first and second modalities may be acquired using different combinations of polarization provided by the infrared light. In this example, the biometric input image data 112 includes nth modality image data 116. In some implementations, the biometric input image data 112 may include images acquired using multiple modalities. For example, the biometric input image data 112 may include a first image acquired using a first modality, a second image acquired using a second modality, etc. Multiple images of the same object may be acquired in rapid succession with respect to each other. For example, the camera 108 may operate at 60 frames per second to acquire individual frames of the nth modality image data 116.

[0035] In operation, the generic input device 170 may acquire generic input image data 172 using the same or a different modality than that used by the biometric input device 104 to acquire the biometric input image data 112. For example, the generic input image data 172 may include an RGB image of the hand 102 having a first resolution and using illumination from a single source (such as the flash of the generic input device 170). Continuing the example, the biometric input image data 112 may include an image of the hand 102 having a second resolution, using infrared illumination provided by multiple sources, and using differential polarization to image veins and other subcutaneous features. In this example, the generic input image data 172 and the biometric input image data 112 are different modalities.

[0036] Shown is a computing device 118. One or more computing device(s) 118 may execute one or more of the following modules:

[0037] During "training time," training data 120 is used to train one or more processing module(s) 130 to determine representation data 132. In one implementation, training data 120 may include a plurality of labeled first and second modality images. For example, label data may indicate sample identifiers, identity labels, modality labels, etc. Training data 120 is discussed in further detail with respect to FIG. 3.

[0038] The processing module(s) 130 may include a machine learning network having several distinct portions. As part of training, the processing module(s) 130, or portions thereof, determine the trained model data that is used to operate the processing module 130 during inference. The machine learning network and training process are discussed in more detail with respect to Figures 3-5.

[0039] Once trained, the processing module(s) 130, or portions thereof, may be used during inference to process inputs, such as biometric input image data 112, and provide as output representation data 132. The operation of the trained processing module(s) 130 is discussed in further detail with respect to FIG.

[0040] In some implementations, the processing module(s) 130 includes a machine learning network including several portions, including one or more backbones, a first embedding portion, an intersection portion, and an XOR portion. The intersection portion facilitates training to generate representation data representing features present in two or more modalities. The XOR portion facilitates training to generate representation data representing features that are distinct or exclusive between input modalities. Examples of features common to two or more modalities include the overall contour of the hand, deep wrinkles on the palm and joints, etc. Such features would be intersection features that appear in multiple modalities. In comparison, features that appear in one modality but not another may be considered distinct or exclusive.

[0041] The training and use of this portion of the system 100 is discussed in more detail with respect to FIGS.

[0042] In some implementations, processing module(s) 130 include a machine learning network trained to accept first representation data associated with a first embedding space and provide transformed representation data associated with a second embedding space. The training and use of this portion of system 100 is discussed in more detail with respect to Figures 8-10 and 14.

[0043] The pre-registration module 134 may be used to determine whether the second representation data 132 based on the biometric input image data 112 corresponds to the first representation data 132 based on the general input image data 172. For example, if the representations are in a common embedding space and within a threshold distance of each other, they may be considered associated with each other. The operation of the pre-registration module 134 is discussed in more detail with respect to Figures 7, 10, and 12-14.

[0044] During "enrollment time," a user may utilize system 100 by performing an enrollment process. The enrollment process may be subdivided into a first time during which the user performs a pre-enrollment portion of the process using general input device 170 and a second time during which enrollment is completed using biometric input device 104. Enrollment module 140 may coordinate the enrollment process. Enrollment may associate biometric information, such as expression data 132, with specific identification data 144, including information such as name, account number, etc.

[0045] During a first time in the enrollment process, the user opts in and presents their hand 102 to the generic input device 170. The generic input device 170 acquires generic input image data 172, such as visible image data 176. The visible image data 176 is then provided to a computing device 118 executing trained processing module(s) 130. The trained processing module(s) 130 accept the generic input image data 172 as input and provide pre-registered expression data 182 as output. The pre-registered expression data 182 represents at least some of the features depicted in the generic input image data 172. In some implementations, the pre-registered expression data 182 may include one or more vector values ​​in one or more embedding spaces. The pre-registered expression data 182 may include data associated with one or more intermediate or final layers of the processing module(s) 130. In some implementations, the intermediate layers may include the initial input layer.

[0046] The pre-registered expression data 182 is associated with the above-described pre-registered identification data 184. The pre-registered identification data 184 is associated with the pre-registered expression data 182. For example, both may be acquired using the same general input device 170. In another example, a first general input device 170(1), such as a tablet computer, may be used to acquire a portion of the pre-registered identification data 184, while a second general input device 170(2), such as a mobile phone, is used to acquire a second portion of the pre-registered identification data 184.

[0047] The pre-registration data 136 is stored, including the pre-registration expression data 182 and the pre-registration identification data 184. An individual instance of the pre-registration expression data 182 is associated with each instance of the pre-registration identification data 184. The pre-registration data 136, or portions thereof, may expire after one or more specified events or times. For example, the pre-registration expression data 182 may be deleted if it has not been used for more than a specified number of days. In another example, the pre-registration expression data 182 and the associated pre-registration identification data 184 may be deleted after registration is complete. Now that the pre-registration data 136 is available, the system 100 is ready for the user to complete the registration process a second time.

[0048] During a second time in the enrollment process, the user presents their hand 102 to the biometric input device 104. The biometric input device 104 provides the biometric input image data 112 to a computing device 118 executing trained processing module(s) 130. The trained processing module(s) 130 accept the biometric input image data 112 as input and provide second representation data 132 as output. The second representation data 132 represents at least some of the features depicted in the biometric input image data 112. In some implementations, the second representation data 132 may include one or more vector values ​​in one or more embedding spaces. The second representation data 132 may include data associated with one or more intermediate or final layers of the processing module(s) 130. In some implementations, the intermediate layers may include the initial input layer.

[0049] During the registration process, the second expression data 132 may be checked using the identification module 150 to determine whether the user has been previously registered. Successful registration may include storing enrollment user data 142, including identification data 144, such as name, phone number, account number, etc., and storing one or more of the expression data 132 or data based thereon, as enrollment expression data 146. In some implementations, the enrollment user data 146 may include additional information associated with the processing of the biometric input image data 112 by the processing module(s) 130. For example, the enrollment expression data 146 may include data associated with one or more intermediate layers of the processing module(s) 130, such as values ​​of the penultimate layer of one or more portions of the processing module(s) 130.

[0050] If the second expression data 132 is deemed not to correspond to a previously registered user, the pre-registration module 134 determines whether the second expression data 132 corresponds to the pre-registered expression data 182. The second expression data 132 and the pre-registered expression data 182 may be associated with different representation or embedding spaces. As a result, a direct comparison between the second expression data 132 and the pre-registered expression data 182 may be infeasible. Several implementations are described that may be used to determine correspondence between the second expression data 132 and the pre-registered expression data 182.

[0051] The determination as to whether the second expression data 132 corresponds to the pre-registered expression data 182 can utilize various techniques. These techniques can include determining a mapping between the expressions. In one implementation, a processing module 130 that determines intersection features can be used to determine whether the second expression data 132 corresponds to the pre-registered expression data 182. In another implementation, a transformer module can be used to use the pre-registered expression data 182 as input to determine transformed expression data that is in the same representation space as the second expression data 132. The transformed expression data can then be compared to the second expression data 132 to determine a correspondence between the two.

[0052] If the pre-registration module 134 determines that the second expression data 132 is associated with the pre-registered expression data 182, the registration process may be completed. For example, the pre-registration identification data 184 may be stored as the identification data 144, and the second expression data 132 may be stored as the registered expression data 146. Once registration is complete, the corresponding pre-registered expression data 182 and pre-registered identification data 184 may be deleted from the pre-registration data 136.

[0053] In some implementations, other data may be used by pre-enrollment module 134. In some implementations, other biometric input data 112 may be used, such as an image of the user's face, audio data representing the user's speech, etc. For example, in addition to comparing second expression data 132 to pre-enrollment expression data 182, a comparison may be made between a first image of the user's face acquired by generic input device 170 and a second image of the user's face acquired by biometric input device 104. In another example, in addition to comparing second expression data 132 to pre-enrollment expression data 182, a comparison may be made between first audio data of the user's speech acquired by generic input device 170 and second audio data of the user's speech acquired by biometric input device 104. In yet another example, speech and video data may be acquired by both generic input device 170 and biometric input device 104 and later compared.

[0054] In some implementations, the user may be prompted to provide additional information to complete the registration process. For example, during the second time, the user may be prompted to provide an EMV or other card for electronic reading by the biometric input device 104. In other examples, the user may be asked to enter a code sent to another device such as a mobile phone, to provide input using an application running on another device such as a mobile phone, etc.

[0055] From the user's perspective, during this initial use of the biometric input device 104 during the second time period, minimal or no interaction related to the enrollment process occurs, resulting in a significant improvement in user convenience. This also reduces the time associated with user enrollment, allowing the biometric input device 104 to be used for non-enrollment purposes. From the system 100's perspective, during the second time period, not only is the second expression data 132 provided by the biometric input device 104 acquired, but other information, such as liveness verification, authentication information, etc., is also acquired.

[0056] During "identification time," a (not yet identified) user presents their hand 102 at the biometric input device 104. The resulting query biometric input image data 112 can be processed by the (now trained) processing module(s) 130 to determine expression data 132. In some implementations, the computing device 106 can execute the trained processing module(s) 130. The computing device 106 can perform other functions, such as encryption and transmission of the biometric input image data 112 or data based thereon (e.g., expression data 132).

[0057] An identification module 150 executing on the computing device(s) 118 may accept as input input expression data 132 associated with biometric input image data 112 acquired by the biometric input device 104. The input expression data 132 is compared to previously stored data, such as enrollment expression data 146, to determine asserted identification data 152. In one implementation, the asserted identification data 152 may include a user identifier associated with the previously stored enrollment expression data 146 that is closest in embedding space to the input expression data 132 associated with the user who presented the hand 102 during the identification time. The identification module 150 may utilize other considerations, such as requiring that the query expression data 132 be less than or equal to a maximum distance in embedding space from a particular user's enrollment expression data 146, before determining the asserted identification data 152.

[0058] The asserted identification data 152 may then be used by subsequent systems or modules. For example, the asserted identification data 152 or information based thereon may be provided to a facility management module 160.

[0059] Facility management module 160 can use asserted identity data 152 to associate an identity with a user as the user moves through the facility. For example, facility management module 160 may use data from cameras or other sensors in the environment to determine the user's location. Given the user's known path from an entrance utilizing biometric input device 104, the user identity indicated in identification data 144 can be associated with the user as the user uses the facility. For example, the identified user may go to a shelf, retrieve an item, and exit the facility. Facility management module 160 may determine that conversation data indicating the retrieval of the item is associated with the user identifier specified in asserted identity data 152 and charge the account associated with the user identifier. In another embodiment, facility management module 160 may include a point-of-sale system. The user may present hand 102 at checkout to assert their identity and make payment using a payment account associated with their identity.

[0060] The above systems and techniques are discussed with respect to images of human hands. These systems and techniques may also be used with respect to other forms of data, other types of objects, etc. For example, these techniques may be used in facial recognition systems, object recognition systems, etc.

[0061] 2 illustrates at 200 supplemental data 174 used by system 100, according to some implementations. Supplemental data 174 may include one or more of system metadata 202 or user data 204. Supplemental data 174 may be associated with or otherwise indicative of an action, such as obtaining general input image data 172 or other data, biometric input image data 112, etc.

[0062] System metadata 202 includes information related to devices and their respective operations. For example, system metadata 202 may include one or more of a device identifier indicating a particular device, device location data indicating the location of the device, timestamp data indicating the date and time, a network address associated with the device's operation, the software version used by the device, liveness detection data indicating whether an input image is associated with a live user or an artifact, etc.

[0063] User data 204 includes information related to a particular user. User data 204 may be based on input from a user or may be associated with a user. For example, user data 204 may include a phone number associated with a user, a payment account number, EMV card data, user account data obtained by an application running on the device, an authentication code, or other information.

[0064] 3 illustrates labeled training data 120 at 300 for training processing module(s) 130, according to some implementations. The training data 120 includes multiple images representing multiple training identities 302(1), 302(2), ..., 302(N). Each training identity 302 is considered unique with respect to other training identities 302.

[0065] The information associated with each training identity 302 may include actual image data obtained from users who opted in to provide information for training, generated synthetic input data, or a combination thereof. In one implementation, training data 120 may exclude individuals who registered to use system 100 for identification. For example, registered users with identification data 144 may be excluded from inclusion in training data 120. In another implementation, some registered users may opt in to explicitly allow biometric input image data 112 obtained during enrollment to be saved as training data 120.

[0066] The synthetic input data may include synthetic data that matches expected biometric input image data 112. For example, the synthetic input data may include output from a generative adversarial network (GAN) trained to generate synthetic images of a user's hands. In some implementations, the synthetic input data may be based on real input data. In other implementations, other techniques may be used to determine the synthetic data.

[0067] Each training identity 302(1)-302(N) includes modality image data and associated label data 340. The label data 340 may include information such as a sample identifier (ID) 342 and a modality label 344. The sample ID 342 indicates a particular training identity. The sample ID 342 may be used to distinguish one training identity 302 from another. In embodiments in which actual input data is used as part of the training data 120, the sample ID 342 may be assigned independently of the actual identification data 144 associated with that user. For example, the sample ID 342 may have a value of "User 4791" and not have the actual identity of "Bob Patel." The modality label 344 indicates whether the associated input data represents a first modality, a second modality, etc.

[0068] In this example, each training identity 302(1)-302(N) includes general input image data 172(1) and its associated sample ID 342(1) and modality label 344(1), and biometric input image data 112(1) and its associated sample ID 342(2) and modality label 344(2).

[0069] In embodiments in which additional modalities are used, the training data 120 for the training identity 302 may also include an Mth modality image data 306(1) and an associated sample ID 342(M) and modality label 344(M).

[0070] FIG. 4 shows a block diagram 400 of processing module(s) 130, including an intersection processing module 402, during training, according to some implementations.

[0071] During training, training data 120 is provided as input to processing module(s) 130. A machine learning network is used to implement processing module(s) 130. The machine learning network may include several portions or branches. In the illustrated embodiment, a portion of the training 440 is trained as specified and described below. The remainder of the portions may have been previously trained for their respective functions. In other embodiments, one or more portions of the entire machine learning network may be trained during training.

[0072] During training, processing module(s) 130 may include a first backbone module 404(1), a second backbone module 404(2), a first processing module 408(1), an intersection processing module 450, and an XOR processing module 460. In some implementations, during training, processing module(s) 130 may also include a second processing module 408(2).

[0073] The backbone module(s) 404 comprise the backbone architecture of the artificial neural network. The backbone module 404 accepts training data 120 as input and generates intermediate representation data 406. In the illustrated implementation, a first backbone module 404(1) accepts general input image data 172 as input and provides first intermediate representation data 406(1) as output. Also shown in FIG. 4, a second backbone module 404(2) accepts biometric input image data 112 as input and provides second intermediate representation data 406(2) as output. In some implementations, a single backbone module 404 can be used to process training data 120 and determine intermediate representation data 406. For example, the same backbone module 404 can be used at different times to determine intermediate representation data 406 for each input.

[0074] In one embodiment, the backbone module(s) 404 may utilize a neural network having at least one layer that utilizes inverted residuals with a linear bottleneck. For example, MobileNetV2 implements this architecture. (See, for example, “MobileNetV2: Inverted Residuals and Linear Bottlenecks,” Sandler, M. et al., 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition, 18-23 June 2018.)

[0075] First processing module 408(1) accepts first intermediate representation data 406(1) as input and determines first representation data 420. For example, first representation data 420 may represent one or more features present in general input image data 172. First representation data 420 may include data associated with one or more intermediate or final layers of first processing module 408(1) or other portions of processing module(s) 130. In implementations in which first backbone module 404(1) accepts general input image data 172 as input, first representation data 420 may represent one or more surface features.

[0076] In some implementations, the network may include a multi-head network. For example, different heads may be trained to determine or utilize specific features in the training data 120. For example, the first processing module 408(1) may include or work in conjunction with another module that provides additional "heads" to determine specific features, such as feature points that represent specific characteristics of friction ridges present on human skin. In another implementation, this portion may be trained to receive and utilize feature point data determined using another system, such as a feature point determination algorithm.

[0077] The second processing module 408(2) accepts the second intermediate representation data 406(2) as input and determines the second representation data 422. For example, the second representation data 442 may represent one or more features present in the biometric input image data 112. The second representation data 442 may include data associated with one or more intermediate or final layers of the second processing module 408(2) or other portions of the processing module(s) 130. In implementations in which the second backbone module 404(2) accepts the biometric input image data 112 as input, the second representation data 442 may represent one or more of the subcutaneous features.

[0078] The machine learning network of processing module(s) 130 includes a crossover branch implemented by crossover processing module 450 and an XOR branch implemented by XOR processing module 460. In some implementations, this provides a joint model training framework. During training, crossover processing module 450 and XOR processing module 460 utilize their respective loss functions to determine loss values. Based on these loss values, each portion determines trained model data. For example, crossover processing module 450 determines trained model data 452, while XOR processing module 460 determines trained model data 462. Some implementations of loss functions that may be used are discussed with respect to FIG. 5.

[0079] The cross-representation branch implemented by the cross-processing module 450 processes the first intermediate representation data 406(1) and the second intermediate representation data 406(2) to determine the cross-representation data 454 and the loss value. For example, the cross-representation branch is trained such that a first modality image and a second modality image having the same sample ID 342 value belong to the same class. As a result, after training, the cross-representation data 454 represents features depicted in the first modality image and the second modality image that are labeled as having the same identity. The cross-representation data 454 may include data associated with one or more intermediate or final layers of the cross-processing module 450 or other portions of the processing module(s) 130.

[0080] The XOR branch implemented by the XOR processing module 460 processes the first intermediate representation data 406(1) and the second intermediate representation data 406(2) to determine XOR representation data 464 for each modality being trained. Each modality of sample ID 342 may be assigned a different class label for training in the XOR branch. The XOR processing module 460 is trained such that first and second modality images having the same sample ID 342 value belong to different classes. As a result, after training, the XOR representation data 464 represents features that are depicted in a particular modality but not in the other modality. For example, the XOR representation data 464 is associated with features that are absent from both the first and second modality images. The XOR representation data 464 may include data associated with one or more intermediate or final layers of the XOR processing module 460 or other portions of the processing module(s) 130.

[0081] In the illustrated implementation, there are two modalities used, resulting in XOR representation data 464 associated with a first modality and a second modality. For example, during training, if first intermediate representation data 406(1) is associated with an input image having a modality label 344 indicating a first modality, the XOR branch determines first modality XOR representation data 464(1). Continuing the example, if second intermediate representation data 406(2) is associated with an input image having a modality label 344 indicating a second modality, the XOR branch determines second modality XOR representation data 464(2).

[0082] During training, the loss values ​​determined by each loss function are used to determine trained model data, such as trained model data 452 for intersection processing module 450 and trained model data 462 for XOR processing module 460.

[0083] In some implementations, one or more of the backbone module(s) 404 may be omitted, and the input data may be processed by one or more of the processing modules 408. For example, general input image data 172 may be provided as input to a first processing module 408(1), an intersection processing module 450, and an XOR processing module 460. Continuing the example, biometric input image data 112 may be provided as input to a second processing module 408(2), an intersection processing module 450, and an XOR processing module 460.

[0084] 5 shows at 500 a block diagram of portions of training 440 and associated loss functions used during training of processing module(s) 130, according to some embodiments. As shown with respect to FIG. 4, training may be performed to train intersection processing module 450 and XOR processing module 460, while the remaining modules of processing module(s) 130 are not trained.

[0085] The intersection processing module 450 includes a first loss function module 514(1). The XOR processing module 460 includes another first loss function module 514(2). The first loss function modules 514(1) and 514(2) determine a loss value 574. The branches utilize label data 340 during training. The loss value 574 may also be provided to the second loss function module 550.

[0086] In one implementation, the first loss function module 514 may utilize a HyperSpheical loss function, as illustrated with respect to Equations 1 and 2. In other implementations, other loss functions may be used, such as Softmax, Cosine, AM-Softmax, Arcface, large margin cosine loss, etc.

[0087] The HyperSpheical loss (HSL) function minimizes L, which is the sum of a cross-entropy term and a regularization term to regularize the confidence scores (weighted by λ). j denotes the classifier weight for the jth class. C is the total number of training classes. M is the mini-batch size. m in these equations is a fixed angular margin.

number

number

number

[0088] During training, intersection processing module 450 may determine intersection representation data 454, one or more parameters of intersection representation data 454, etc. For example, the one or more parameters may include weights of one or more classes. The intersection representation data 454 and the associated one or more parameters may be stored as intersection data 520. During training, XOR processing module 460 may determine a plurality of first modality XOR representation data 464(1) and second modality XOR representation data 464(2), and one or more parameters of XOR representation data 464. The XOR representation data 464 and the associated one or more parameters may be stored as XOR data 522. Once training is complete, one or more of intersection data 520 or XOR data 522 may be deleted or otherwise discarded.

[0089] The probability distribution module 530 processes the data 520-522 to determine a set of probability distributions. The intersection data 520 is processed to determine an intersection probability distribution (Pi) 542(1). The XOR data 522 is processed to determine an XOR probability distribution (Pxp) 542(2).

[0090] The second loss function module 550 accepts the probability distributions 542 and determines a second loss value 576. In one embodiment, the second loss function module 550 may implement a Jensen-Shannon divergence (JSD) loss function. The JSD loss function measures the similarity between two probability distributions. For two probability distributions 542 P and Q, the JSD may be defined in one embodiment by the following equation:

number

[0091] KLD,

number

[0092] Given an image x of identity c, the above joint model training framework uses a first loss function, such as a hyperspherical loss, to determine a set of probability distributions 542: a cross probability distribution (Pi) 542(1) and an XOR probability distribution (Pxp) 542(2). For example, the loss values ​​574 determined by the first loss function module(s) 514 may be used as input to a JSD loss function. These probability distributions are N-dimensional, where N is the number of training identities 302 (N). This can be expressed as: Pi = [pi_1,pi_2,...,pi_N] (Equation 4) Pxp=[pxp_1,pxp_2,...,pxp_N] (Equation 5)

[0093] From each of these probability distributions 542, the entry corresponding to the correct identity c is removed and the vector is normalized to obtain an (N-1)-dimensional probability distribution of the incorrect classes Pi_n, Pxp_n, as shown in the following equation: Pi_n=[pi_1,pi_2,...,pi_c-1,pi_c+1,...,pi_N] / (1-pi_c) (Equation 6) Pxp_n=[pxp_1,pxp_2,...,pxp_c-1,pxp_c+1,...pxp_N] / (1-pxp_c) (Equation 7)

[0094] The JSD loss then minimizes the following equation:

number

[0095] A total loss value 560 is calculated based on the first loss value 572 and the second loss value 576. For example, the total loss value 560 may be calculated using the following formula: TotalLoss=Hyperspherical_loss+loss_weight*JSD_loss(Equation 9)

[0096] The total loss value 560 may then be provided to one or more of the intersection processing module 450 or the XOR processing module 460 for subsequent iterations during training. As a result of the training, trained model data 452 and 462, respectively, is determined. For example, the trained model data may include weight values, bias values, thresholds, etc. associated with particular nodes or functions within the processing module(s) 130. Once trained, the processing module(s) 130 may be used to determine representation data 132 for subsequent use.

[0097] FIG. 6 illustrates at 600 a block diagram of processing module(s) 130 during inference, according to some embodiments.

[0098] Once the training 440 portion has been trained as described above, a subset of the machine learning network may be used during inference. In the illustrated embodiment, the cross-processing module 402 during inference may include a first backbone module 404(1), a first processing module 408(1), a second backbone module 404(2), a second processing module 408(2), and a cross-processing module 450. In operation, input data 602, such as general input image data 172 including visible image data 176 and biometric input image data 112 including n-th modality image data 116, is provided to the trained processing module(s) 130.

[0099] First backbone module 404(1) can process general input image data 172 to determine first intermediate representation data 406(1). First intermediate representation data 406(1) is processed by first processing module 408(1) to determine first representation data 420. First intermediate representation data 406(1) is processed by intersection processing module 450 to determine first intersection representation data 454(1).

[0100] In some implementations, the first intermediate representation data 406(1) may be processed by the XOR processing module 460 to determine the first XOR representation data 464(1).

[0101] The second backbone module 404(2) may process the biometric input image data 112 to determine second intermediate representation data 406(2). The second intermediate representation data 406(2) is processed by the second processing module 408(2) to determine second representation data 422. The second intermediate representation data 406(2) is processed by the intersection processing module 450 to determine second intersection representation data 454(2).

[0102] In some implementations, the second intermediate representation data 406(2) may be processed by the XOR processing module 460 to determine the second XOR representation data 464(2).

[0103] The representation data 132 may include one or more of the first representation data 420, the second representation data 422, the first intersection representation data 454(1), the second intersection representation data 454(2), the first XOR representation data 464(1), or the second XOR representation data 464(2). The resulting representation data 132 may be used in subsequent processes, such as determining whether the general input image data 172 and the biometric input image data 112 match or mismatch, enrollment, identification, etc. The representation data 132 may include data related to one or more intermediate or final layers of the processing module(s) 130.

[0104] In some implementations, one or more of the backbone module(s) 404 may be omitted, and the input data 602 may be processed by one or more of the processing modules 408. For example, general input image data 172 may be provided as input to a first processing module 408(1), an intersection processing module 450, and an XOR processing module 460. Continuing the example, biometric input image data 112 may be provided as input to a second processing module 408(2), an intersection processing module 450, and an XOR processing module 460.

[0105] In addition to the above, once trained, a deployed implementation of processing module(s) 130 may omit one or more other modules that are used during training but not during inference. For example, processing module(s) 130 may omit first loss function module 514, probability distribution module 530, second loss function module 550, etc.

[0106] 7 is a block diagram 700 of a cross-comparison module that may be used for pre-registration, according to some embodiments. As noted above, the modalities associated with the general input image data 172 and the biometric input image data 112 may be different. In some embodiments, the modalities may be the same, but different processing modules 130 may be used to process the general input image data 172 and the biometric input image data 112 to determine the respective pre-registration representation data 182 and second representation data 132. In one embodiment, a trained cross-processing module 402 may be used to determine correspondence between the general input image data 172 and the biometric input image data 112.

[0107] In this example, the input data 602 includes general input image data 172 acquired at a first time and biometric input image data 112 acquired at a second time. Pre-enrollment identification data 184 may also be acquired at the first time and is associated with the general input image data 172.

[0108] In some implementations, multiple instances of image data may be acquired. For example, input data 602 may include ten pairs of general input image data 172 and biometric input image data 112 acquired from users who have opted in to use system 100.

[0109] Input data 602 is processed by intersection processing module 402 as described above with respect to FIG.

[0110] The intersection output data 702 can include one or more instances of expression data 132(1)-(N), each instance including first expression data 420 (stored as pre-registered expression data 182), first intersection expression data 454(1), second expression data 422, or second intersection expression data 454(2).

[0111] The intersection output data 702, or a portion thereof, may be processed by the pre-registration module 134. In the embodiment shown, the pre-registration module 134 includes a cross-comparison module 718. The cross-comparison module 718 is configured to accept as input first intersection representation data 454(1) and second intersection representation data 454(2) and determine comparison result data 724. The cross-comparison data 454 associated with one or more instances of representation data 132(1)-(N) in the intersection output data 702 may be provided to the cross-comparison module 718. In some embodiments, the cross-comparison module 718 may compare each instance of the first intersection representation data 454(1) with each instance of the second intersection representation data 454(2). Continuing with the previous example, if 10 pairs of biometric input image data 112(1)-(10) and general input image data 172(1)-(10) are processed, up to 10 x 10, or 100, comparisons may be performed. Comparison result data 724 may be determined for each comparison. Continuing with this example, comparison result data 724 may include a set of 100 values, each associated with a particular pair.

[0112] In one implementation illustrated herein, the intersection comparison module 718 may determine distance data 722 indicating a distance in intersection embedding space between a given combination of the first intersection representation data 454(1) and the second intersection representation data 454(2). For example, the distance data 722 may be calculated as the cosine distance between the first intersection representation data 454(1) and the second intersection representation data 454(2). The distance data 722 may be compared to a first threshold specified by the threshold data 720 to determine the comparison result data 724. For example, if the distance data 722 indicates a distance less than the first threshold, the general input image data 172 and the biometric input image data 112 may be considered to correspond to the same hand 102. In another example, if the distance data 722 indicates a distance equal to or greater than the threshold, the general input image data 172 and the biometric input image data 112 may be considered not to correspond or to correspond to different hands 102.

[0113] In some implementations, cross-comparison module 718 may include a classifier or other machine learning system that may be trained to accept first cross-representation data 454(1) and second cross-representation data 454(2) as inputs and provide comparison result data 724 indicating a classification of “{inputs_correspond}” or “{no_correspondence}” as output. In some implementations, the classifier may also be trained using one or more of first representation data 420, second representation data 422, XOR representation data 464(1) or 464(2), etc. In some implementations, threshold data 720 may specify one or more thresholds associated with the operation of the classifier. For example, threshold data 720 may specify a minimum confidence value to be used to provide an output.

[0114] At 740, the comparison result data 724 is evaluated. If at 740, the comparison result data 724 indicates that the general input image data 172 and the biometric input image data 112 correspond to one another, the process may proceed to complete enrollment by saving the second expression data 422 or other information based on the biometric input image data 112 as enrollment expression data 146 and saving the pre-enrollment identification data 184 associated with the general input image data 172 as identification data 144 associated with the enrollment expression data 146.

[0115] At 740, if the general input image data 172 and the biometric input image data 112 are deemed not to correspond to one another, the process may proceed to initiate a registration process. For example, the user may be prompted to opt in to use the system, may be presented with a user interface for providing identification data 144, or the like.

[0116] During operation of the system 100, the second expression data 422 based on the biometric input image data 112 may be processed by the identification module 150 to determine whether the user has been previously enrolled. For clarity of illustration, but not limitation, this comparison is not shown in FIG. 7 . If it is determined that the biometric input image data 112 does not correspond to a previously enrolled user, operations associated with the pre-enrollment module 134 may be performed to complete enrollment. If it is determined that the user has been previously enrolled, further discussion of this figure may be omitted.

[0117] 8 illustrates at 800 a method for processing training input data to determine transformer training data 850, according to some implementations. Preparation of the transformer training data 850 may be performed by one or more computing devices 118. The transformer training data 850 is obtained for use in training a transformer module 902, as described with respect to FIG.

[0118] As shown above, training data 120 is shown. The training data 120 may include one or more of general input image data 172 and biometric input image data 112 with associated label data 340, or synthetic input data with associated label data 340.

[0119] Training data 120 is processed by at least two processing models 808. General input image data 172, such as visible image data 176, is processed by a first processing module 808(1) to determine first training representation data 820(1) in a first embedding space 822(1). In some implementations, first training representation data 820(1) includes or is based on hidden layer data 810 and embedding layer data 812. Hidden layer data 810 may include values ​​associated with one or more layers of first processing module 808(1) while processing the input. Embedding layer data 812 includes representation data 132 provided by the output of first processing module 808(1). In one implementation, hidden layer data 810 may include values ​​of the penultimate layer of the neural network of first processing module 808(1). The penultimate layer may include the layer preceding the final output of embedding layer data 812. In one implementation, the hidden layer data 810 may include values ​​of a fully connected linear layer that precedes the output of the embedding layer data 812. For example, the embedding layer data 812 may have vectors of size 128, while the hidden layer data 810 has vectors of size 1280.

[0120] Continuing with the above embodiment, the first training representation data 820(1) may include a concatenation of the hidden layer data 810 and the embedding layer data 812. In other embodiments, the hidden layer data 810 and the embedding layer data 812 may be combined in other ways.

[0121] In some implementations, the use of intermediate tier data 810 significantly improves the overall performance of system 100.

[0122] The biometric input image data 112, such as the nth modality image data 116, is processed by a second processing module 808(2) to determine second training representation data 820(2) in a second embedding space 822(2). This pair of training representation data 820(1) and 820(2) may be associated with each other by a common value of sample ID 342. Thus, the pair represents the same input data from training data 120 as represented in two different embedding spaces. Each instance of training representation data 820 may be associated with label data 856. This associated label data 856 may include a modality label 344 and a model label 858 that indicates the processing module 808 used to generate the particular training representation data 820.

[0123] In some implementations, second training representation data 820(2) includes or is based on hidden layer data 814 and embedding layer data 816. Hidden layer data 814 may include values ​​associated with one or more layers of second processing module 808(2) while processing the input. Embedding layer data 816 includes representation data 132 provided by the output of second processing module 808(2). In one implementation, hidden layer data 814 may include values ​​of the penultimate layer of the neural network of second processing module 808(2). The penultimate layer may include the layer preceding the final output of embedding layer data 816. In one implementation, hidden layer data 814 may include values ​​of the fully connected linear layer preceding the output of embedding layer data 816. For example, embedding layer data 816 may have vectors of size 128, while hidden layer data 814 has vectors of size 1280.

[0124] Continuing with the above embodiment, the second training representation data 820(2) may include a concatenation of the hidden layer data 814 and the embedding layer data 816. In other embodiments, the hidden layer data 814 and the embedding layer data 816 may be combined in other ways.

[0125] In some implementations, the use of intermediate tier data 814 significantly improves the overall performance of the system 100 .

[0126] Transformer training data 850, including first training representation data 820(1), second training representation data 820(2), and associated or implied label data 856, can be used to train the transformer module 902, as described next.

[0127] 9 illustrates a transformer module 902 during training at 900, according to some implementations. The transformer module 902 may be implemented by one or more computing devices 118. The transformer module 902 includes a transformer network module 910, a classification module 916, a similarity loss module 918, and a divergence loss module 920.

[0128] The transformer network module 910 may include a neural network. During training, the transformer network module 910 accepts as input first training representation data 820(1) associated with a first embedding space 822 and produces as output transformed representation data 914. As training progresses, the quality of the resulting transformed representation data 914 may be expected to improve, as described below by the returned loss value 960.

[0129] In some implementations, the transformer network module 910 may include one or more multi-layer perceptrons (MLPs). Trained model data 912 associated with the operation of the transformer network module 910 is determined during training. For example, the trained model data 912 may include one or more of weight values, bias values, etc. associated with the operation of portions of a neural network.

[0130] The transformed representation data 914 is processed by a first classification module 916(1) to determine a first classification loss 942. In one implementation, the classification module 916 may utilize a HyperSpheical loss function, as illustrated with respect to Equations 1 and 2. In other implementations, other classification loss functions may be used. For example, other classification functions such as Softmax, Cosine, AM-Softmax, Arcface, large margin cosine loss, etc. may be used.

[0131] A HyperSpheical loss (HSL) function may also be used during training of the processing module(s) 130. The HSL loss minimizes L, which is the sum of a cross-entropy term and a regularization term for regularizing the confidence scores (weighted by λ). j denotes the classifier weight for the jth class. C is the total number of training classes. M is the mini-batch size. m in these equations is a fixed angular margin.

number

number

number

[0132] The second training representation data 820(2) is processed by a second classification module 916(2) to determine a second classification loss 948. The second classification module 916(2) may utilize the same loss function as the first classification module 916(1). For example, the second classification module 916(2) may utilize a HyperSpherical loss function.

[0133] The similarity loss module 918 accepts the transformed expression data 914 and the second training expression data 820(2) as input and determines a similarity loss 944.

[0134] In one implementation, the similarity loss module 918 may implement a mean squared error (MSE) and cosine distance loss function. In other implementations, other loss functions may be used. For example, an MSE loss may be used.

[0135] The divergence loss module 920 accepts the first classification loss 942 and the second classification loss 948 as input and determines the divergence loss 946. In one implementation, the divergence loss module 920 may implement a Kullback-Leibler divergence (KLD) function.

[0136] Loss value(s) 960, including one or more of the first classification loss 942, the second classification loss 948, the similarity loss 944, or the divergence loss 946, are then returned to the transformer network module 910 for subsequent iterations during training.

[0137] In some implementations, the Transformer Network module 910 may implement a cycle-consistency loss function. See, "Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks," Zhu, Jun-Yan, et al., arXiv:1703.01593v7, 24 August 2020.) The cycle-consistency loss function can be used to minimize the distance between the first training representation data 820(1) and the transformed representation data 914. In one implementation, this may result in a first portion of the Transformer Network module 910 being trained to transform the first representation data 420 into the transformed representation data 914, and a second portion of the Transformer Network module 910 being trained to transform the transformed representation data 914 into the third representation data 420. The first representation data 420 and the third representation data are represented in a common representation space, such as the same embedding space. For example, during training, the transformer network module 910 may be trained to accept an input printed image as input, generate a vein image from the printed image, and generate a second printed image from the vein image. The input printed image may then be compared to the second printed image, and the transformer network module 910 trained to minimize the distance between them.

[0138] FIG. 10 is a block diagram 1000 of using a transformer module and a comparison module for pre-registration, according to some embodiments.

[0139] As noted above, the modalities associated with the general input image data 172 and the biometric input image data 112 may be different. In some implementations, the modalities may be the same, but different processing modules 130 may be used to process the general input image data 172 and the biometric input image data 112 to determine the respective pre-enrolled expression data 182 and second expression data 132. In one implementation, correspondence between the general input image data 172 and the biometric input image data 112 may be determined by transforming one or more of the expression data 132 into a common representation space. For example, the pre-enrolled expression data 182 is transformed into transformed expression data 914 that is in the same representation space as the second expression data 132, allowing the two to be compared.

[0140] In this example, the input data 602 includes general input image data 172 acquired at a first time and biometric input image data 112 acquired at a second time. Pre-enrollment identification data 184 may also be acquired at the first time and is associated with the general input image data 172. As described with respect to FIG. 7 , in some implementations, multiple instances of image data may be acquired. For example, the input data 602 may include ten pairs of general input image data 172 and biometric input image data 112 acquired from users who have opted in to use the system 100.

[0141] The pre-registration module 134 may include a trained transformer network module 910. As mentioned above, in some implementations, the transformer network module 910 may implement a cycle consistency loss function during training.

[0142] In operation, pre-registration module 134 accepts as input first expression data 1010 and second expression data 1012. For example, general input image data 172 may be processed by first processing module 808(1) to determine first expression data 1010, and biometric input image data 112 may be processed by second processing module 808(2) to determine second expression data 1012. In other implementations, other processing modules 130 may be used to determine first expression data 1010 and second expression data 1012. For example, intersection processing module 402 may be used.

[0143] The transformer network module 910 accepts the first expression data 1010 as input and determines the transformed expression data 914 as output based on the trained model data 912. The transformed expression data 914 is in the same second embedding space 822(2) as the second expression data 1012. The transformed expression data 914 may be saved as pre-registered expression data 182 or may be used in other ways.

[0144] The comparison module 1018 accepts the transformed expression data 914 and the second expression data 1012 as inputs and determines comparison result data 1024 .

[0145] In the illustrated implementation, the comparison module 1018 may determine distance data 1022 indicating a distance in the second embedding space 822(2) between the transformed expression data 914 and the second expression data 1012. For example, the distance data 1022 may be calculated as the cosine distance between the transformed expression data 914 and the second expression data 1012. The distance data 1022 may be compared to a threshold specified by the threshold data 1020 to determine the comparison result data 1024. For example, if the distance data 1022 indicates a distance less than the threshold, the input data 602 may be deemed to correspond in that the biometric input image data 112 is associated with the general input image data 172. In another example, if the distance data 1022 indicates a distance equal to or greater than the threshold, the input data 602 may be deemed to not correspond in that the biometric input image data 112 and the general input image data 172 are deemed to relate to different hands 102.

[0146] In another embodiment, the comparison module 1018 may include a classifier or other machine learning system that may be trained to accept the transformed expression data 914 and the second expression data 1012 as inputs and provide comparison result data 1024 indicating a classification of "{inputs_correspond}" or "{no_correspondence}." In some embodiments, the classifier may be trained using one or more of the transformed expression data 914, the first expression data 1010, the second expression data 1012, the intersection expression data 454, the XOR expression data 464, etc.

[0147] In some implementations (not shown), the comparison module 1018 can alternatively or additionally perform a comparison between the first expression data 1010 and the second transformed expression data 914(2). The second transformed expression data 914(2) can be determined by using a second transformer network module 910(2). The second transformer network module 910(2) accepts the second expression data 1012 as input and determines the second transformed expression data 914(2) as output based on the trained model data 912(2). The second transformed expression data 914(2) is in the same first embedding space 822(1) as the first expression data 1010.

[0148] Implementations using additional modalities may include comparison of additional representations, a transformer network module 910, and the transformed representation data 914 produced thereby.

[0149] At 1040, the comparison result data 1024 is evaluated. If at 1040, the comparison result data 1024 indicates that the general input image data 172 and the biometric input image data 112 correspond to one another, the process may proceed to complete enrollment by saving the second expression data 1012 or other information based on the biometric input image data 112 as enrollment expression data 146 and saving the pre-enrollment identification data 184 associated with the general input image data 172 as identification data 144 associated with the enrollment expression data 146.

[0150] At 1040, if the general input image data 172 and the biometric input image data 112 are deemed not to correspond to one another, the process may proceed to initiate a registration process. For example, the user may be prompted to opt in to use the system, may be presented with a user interface for providing identification data 144, or the like.

[0151] During operation of the system 100, the second expression data 1012 based on the biometric input image data 112 may be processed by the identification module 150 to determine whether the user has been previously enrolled. For clarity of explanation, and not by way of limitation, this comparison is not shown in FIG. 10. If it is determined that the biometric input image data 112 does not correspond to a previously enrolled user, operations associated with the pre-enrollment module 134 may be performed to complete enrollment. If it is determined that the user has been previously enrolled, the description with respect to this figure may be omitted.

[0152] FIG. 11 is a block diagram 1100 of a computing device 118 for implementing system 100, according to some embodiments. Computing device 118 may be within biometric input device 104, may include a server, etc. Computing device 118 may be physically present at a facility, accessible over a network, or a combination of both. Computing device 118 does not require end-user knowledge of the physical location and configuration of the system delivering the service. Common terms associated with computing device 118 may include “embedded system,” “on-demand computing,” “software as a service (SaaS),” “platform computing,” “network-accessible platform,” “cloud services,” “data center,” etc. Services provided by computing device 118 may be distributed across one or more physical or virtual devices.

[0153] One or more power sources 1102 may be configured to provide adequate power to operate components within the computing device 118. The one or more power sources 1102 may include batteries, capacitors, fuel cells, solar cells, wireless power receivers, conductive couplings suitable for attachment to a power source such as that provided by a power utility, and the like. The computing device 118 may include one or more hardware processors 1104 configured to execute one or more stored instructions. The processors 1104 may include one or more cores. One or more clocks 1106 may provide information indicating dates, times, instants, and the like. For example, the processor 1104 may use data from the clock 1106 to associate a particular interaction with a particular point in time.

[0154] The computing device 118 may include one or more communication interfaces 1108, such as an input / output (I / O) interface 1110 and a network interface 1112. The communication interface 1108 allows the computing device 118 or components thereof to communicate with other devices or components. The communication interface 1108 may include one or more I / O interfaces 1110. The I / O interfaces 1110 may include an Integrated Circuit (I2C), a Serial Peripheral Interface Bus (SPI), a Universal Serial Bus (USB) developed by the USB Implementers Forum, RS-232, etc.

[0155] The I / O interface(s) 1110 may couple to one or more I / O devices 1114. The I / O devices 1114 may include input devices such as one or more of sensors 1116, a keyboard, a mouse, a scanner, etc. The I / O devices 1114 may also include output devices 1118, such as one or more of a display device, a printer, an audio speaker, etc. In some embodiments, the I / O devices 1114 may be physically integrated with the computing device 118 or may be external. The sensors 1116 may include a camera 108, a smart card reader, a touch sensor, a microphone, etc.

[0156] The network interface 1112 may be configured to facilitate communication between the computing device 118 and other devices, such as routers and access points. The network interface 1112 may include devices configured to couple to a personal area network (PAN), a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), etc. For example, the network interface 1112 may include devices compatible with Ethernet, Wi-Fi, Bluetooth, etc.

[0157] Computing device 118 may also include one or more buses or other internal communication hardware or software that allow data to be transferred between the various modules and components of computing device 118 .

[0158] 11, computing device 118 includes one or more memories 1120. Memory 1120 may include one or more non-transitory computer-readable storage media (CRSMs). CRSMs may be any one or more of electronic, magnetic, optical, quantum, and mechanical computer storage media. Memory 1120 provides storage of computer-readable instructions, data structures, program modules, and other data for operation of computing device 118. While some functional modules are stored in memory 1120, the same functionality may alternatively be implemented in hardware, firmware, or as a system-on-chip (SoC).

[0159] Memory 1120 may include at least one operating system (OS) module 1122. OS module 1122 is configured to manage hardware resource devices, such as I / O interface 1110, I / O devices 1114, and communication interface 1108, and to provide various services to applications or modules executing on processor 1104. OS module 1122 may implement a variant of the FreeBSD operating system as promulgated by the FreeBSD Project, other UNIX or UNIX-like variants, a variant of the Linux operating system as promulgated by Linus Torvalds, and the Windows operating system from Microsoft Corporation of Redmond, Washington, USA, etc.

[0160] The communications module 1126 may be configured to establish communications with the computing device 118, servers, other computing devices 106, or other devices. The communications may be authenticated, encrypted, or the like.

[0161] The processing module(s) 130 may be stored in memory 1120 .

[0162] The pre-registration module 134 may be stored in the memory 1120. In some implementations, the pre-registration module 134 may include one or more of the cross-comparison module 718, the transformer network module 910, the pre-registration module 134, or the comparison module 1018.

[0163] The registration module 140 may be stored in the memory 1120 .

[0164] The identification module 150 may be stored in the memory 1120 .

[0165] The facility management module 160 may be stored in the memory 1120 .

[0166] Also stored in memory 1120 may be a data store 1124 and one or more of the following modules, which may run as foreground applications, background tasks, daemons, etc. The data store 1124 may use flat files, databases, linked lists, trees, executable code, scripts, or other data structures for storing information. In some implementations, the data store 1124 or portions of the data store 1124 may be distributed across one or more other devices, including other computing devices 106, network-attached storage devices, etc.

[0167] The data store 1124 may store training data 120, transformer training data 850, etc. The data store 1124 may store trained model data 1134, such as trained model data 452, trained model data 462, and trained model data 912. The data store 1124 may store registered user data 142.

[0168] In some implementations, one or more of the general input image data 172 or the biometric input image data 112 may be temporarily stored during processing by the processing module(s) 130 or other modules. For example, the biometric input device 104 may obtain the biometric input image data 112, determine the expression data 132 based on the biometric input image data 112, and then clear the biometric input image data 112. The resulting expression data 132 may then be transmitted to a server or other computing device 118 to perform enrollment, comparison to assert identity, etc.

[0169] For example, facility management module 160 may perform various functions such as tracking items between different inventory locations or carts, generating restock orders, directing the operation of robots within the facility, and associating particular user identities with users within the facility using asserted identification data 152. During operation, facility management module 160 may access sensor data 1132, such as biometric input image data 112 or data from other sensors 1116.

[0170] Information used by facility management module 160 may be stored in data store 1124. For example, data store 1124 may be used to store physical layout data 1130, sensor data 1132, asserted identification data 152, user position data 1136, interaction data 1138, etc. For example, sensor data 1132 may include biometric input image data 112 obtained from biometric input devices 104 associated with the facility.

[0171] The physical layout data 1130 may provide information indicating where biometric input devices 104, cameras, weight sensors, antennas for wireless receivers, inventory locations, etc. are located relative to one another within a facility. For example, the physical layout data 1130 may include information representing a map or floor plan of a facility, with the relative locations of gates and inventory locations equipped with biometric input devices 104.

[0172] Facility management module 160 may generate user location data 1136 that indicates the user's location within the facility. For example, facility management module 160 may use image data captured by a camera to determine the user's location. In other implementations, other techniques may be used to determine user location data 1136. For example, data from a smart floor may be used to determine the user's location.

[0173] Identification data 144 may be associated with user location data 1136. For example, a user enters a facility and has their hand 102 scanned by a biometric input device 104, resulting in asserted identification data 152 that is associated with the time of entry and the location of the biometric input device 104. Tracking data indicating the user's path starting from the location of the biometric input device 104 at the time of entry may associate user location data 1136 with the user identifier in asserted identification data 152.

[0174] A particular interaction may be associated with a particular user's account based on the user location data 1136 and the interaction data 1138. For example, if the user location data 1136 indicates a user's presence in front of the inventory location 1192 at a time of 09:02:02, and the interaction data 1138 indicates a pick-up of a quantity of one item from an area on the inventory location 1192 at 09:04:13, the user may be charged a fee for the pick-up.

[0175] Facility management module 160 may use sensor data 1132 to generate interaction data 1138. Interaction data 1138 may include information regarding the type of items involved, the number involved, whether the interaction was a pick or place, and so on. Interactions may include a user picking an item from an inventory location, placing an item in an inventory location, touching an item at an inventory location, searching for an item at an inventory location, and so on. For example, facility management module 160 may generate interaction data 1138 indicating which item a user selected from a particular lane on a shelf and may then use this interaction data 1138 to adjust the total number of inventory contained in that lane. Interaction data 1138 may then be used to charge a fee to an account associated with a user identifier associated with the user who selected the item.

[0176] The facility management module 160 may process the sensor data 1132 and generate output data. For example, based on the interaction data 1138, the amount of a certain type of item at a particular inventory location may fall below a threshold restocking level. The system may generate output data including a restocking order indicating the inventory location, area, and quantity needed to replenish the inventory to a predetermined level. The restocking order may then be used to instruct a robot to restock that inventory location.

[0177] Other modules 1140 may also reside in other data 1142 in data store 1124 along with memory 1120. For example, a billing module may use interaction data 1138 and asserted identity data 152 to bill an account associated with a particular user.

[0178] 12A, 12B, and 12C show a process flow diagram 1200 for performing a pre-registration process according to some implementations. The process may be performed by one or more computing devices of system 100, such as a computing device of general input device 170, computing device 106 of biometric input device 104, computing device 118, or other devices.

[0179] The user opts in to use the system. The following process may then be initiated following the user opting in:

[0180] The operations shown with respect to 1202 through 1214 in FIG. 12A may be associated with a first time.

[0181] At 1202, first input image data is acquired using a first device, for example, general input device 170 may be used to acquire general input image data 172 at a first time.

[0182] At 1204, first expression data 132(1) is determined based on the first input image data. For example, general input image data 172 may be processed by one or more of processing modules 130 to determine first expression data 132(1). In some implementations, first expression data 132(1), or data based thereon, may be stored as pre-registered expression data 182.

[0183] At 1206, it is determined that first expression data 132(1) is not associated with stored enrolled user data 142. For example, if enrolled expression data 146 includes previously stored first expression data 132, identification module 150 may accept first expression data 132(1) as input and return data indicating that there is no corresponding enrolled expression data 146. In other implementations, first expression data 132(1) may be processed as described herein to determine whether there is an association between first expression data 132(1) and previously stored enrolled expression data 146.

[0184] At 1208, first identification data is determined. For example, the pre-registration identification data 184 may be received from the general input device 170 or may be retrieved from storage.

[0185] In some implementations, at 1210, the first identification data is deemed not associated with stored identification data 144 in registered user data 142. For example, the first identification data may include user data 204 such as a company name, a phone number, an account identifier, etc. The identification data 144 may be searched to determine whether the first identification data is already present in registered user data 142. If yes, the user may be presented with a prompt in the user interface indicating that they are already registered. If no, the process may proceed to 1212.

[0186] At 1212, the first identification data and the first expression data are saved, for example, pre-registered identification data 184 and pre-registered expression data 182 are saved for use at a second time.

[0187] In some implementations, first user interface data may be determined at 1214 indicating that the user may proceed to use the system. The first user interface data may then be used to generate output to the user. For example, the first user interface data may include text that may be presented to the user via an output device, such as a display device.

[0188] The operations shown with respect to 1216 through 1236 in FIG. 12B may be associated with a second time.

[0189] In some implementations, at 1216, first transaction data associated with the second device is determined. The transaction data may indicate a purchase, a request for physical access, authorization to perform a function, etc. For example, the transaction data may include a request to charge a payment account for a retail purchase. As described above, once the pre-registration process is completed a first time, an initial use of the biometric input device 104 may include processing the transaction data. As a result, the initial use may complete the registration process simultaneously with completing a transaction.

[0190] At 1218, second input image data is acquired using a second device. For example, biometric input device 104 may be used to acquire biometric input image data 112 at a second time.

[0191] At 1220, second expression data 132(2) is determined based on the second input image data. For example, biometric input image data 112 may be processed by one or more of processing modules 130 to determine second expression data 132(2).

[0192] At 1222, it is determined that the second expression data 132(2) is not associated with the stored enrolled user data 142. For example, the identification module 150 may accept the second expression data 132(2) as input and return data indicating that there is no corresponding enrolled expression data 146.

[0193] At 1224, a determination may be made as to whether the association between the first expression data 132(1) and the second expression data 132(2) is greater than a threshold. For example, the correspondence between the first expression data 132(1) and the second expression data 132(2) may be expressed as a confidence value that is compared to a threshold. If the association is less than or equal to the threshold, the process proceeds to 1226.

[0194] Additional data may be acquired at 1226. For example, the additional data may include user data 204. In some implementations, a second device may be used to acquire the additional data. For example, the biometric input device 104(1) may prompt the user to present an EMV card for reading by a reader of the biometric input device 104(1), may prompt the user to enter an authentication code, may prompt the user to enter a phone number, etc. In other implementations, the additional data may be acquired using another device, such as providing information about a device, such as a mobile phone, associated with the pre-registration identification data 184. In some implementations, the additional data may include audio data, such as the user's speech, image data, such as an image of the user's face, etc. The additional data may be evaluated, and if deemed to correspond to information associated with the pre-registration identification data 184, the process proceeds to 1228. Otherwise, an error message may be presented.

[0195] If the association between the first expression data 132(1) and the second expression data 132(2) is greater than the threshold, the process proceeds to 1228. At 1228, it is determined that the first expression data 132(1) is associated with the second expression data 132(2). For example, intersection output data 702 may be determined and processed to determine comparison result data 724 indicating that the first expression data 132(1) is associated with the second expression data 132(2).

[0196] At 1230, the second expression data 132(2) and the first identification data are stored. For example, the second expression data 132(2) may be stored as registered expression data 146, and the pre-registered identification data 184 or at least a portion of the data based thereon may be stored as identification data 144. An association between a particular registered expression data 146 and a particular identification data 144 is also stored.

[0197] Once the registration process is complete, the associated pre-registration data 136 may be deleted or otherwise discarded.

[0198] Once the registration process is complete, first transaction data associated with the user may also be processed.

[0199] At 1232, the first transaction data is associated with the first identification data. For example, the first transaction data may be considered to be associated with the identification data 144 in the registered user data 142.

[0200] At 1234, the first transaction data is processed using the first identification data. Continuing with this example, user identification information or payment account information associated with the first identification data may be determined and used to process the first transaction data and charge the payment account specified by the payment account information.

[0201] In some implementations, second user interface data indicating completion of the registration process may be determined at 136. The second user interface data may then be used to generate output to the user. For example, the second user interface data may include text that may be presented to the user via an output device, such as a display device.

[0202] 12C may be associated with a third time. In some implementations, different biometric input devices 104 may acquire biometric input image data 112 of different modalities. For example, biometric input device 104(1) may acquire biometric input image data 112(1) using a first and second modality, and biometric input device 104(2) may acquire biometric authentication input image data 112(2) using only the second modality.

[0203] During subsequent uses of the system 100, for example, during subsequent transactions by the user, use of the biometric input device 104 and the resulting biometric input image data 112 and corresponding expression data 132 may be associated with the previously registered user data 142 and stored as registered expression data 146, while also facilitating the completion of the transaction specified by the transaction data.

[0204] In some implementations, second transaction data associated with the third device is determined at 1238. The transaction data may indicate a purchase, a request for physical access, authorization to perform a function, etc.

[0205] At 1240, third input image data is acquired using a third device. For example, biometric input device 104 may be used to acquire biometric input image data 112(2) at a third time.

[0206] At 1242, third expression data 132(3) is determined based on the third input image data. For example, biometric input image data 112(2) may be processed by one or more of processing modules 130 to determine third expression data 132(3).

[0207] At 1244, it is determined that the third expression data 132(3) is associated with one or more of the first expression data 132(1) or the second expression data 132(2). Similar to the process described above with respect to 1224, if the association is greater than a threshold, the process may proceed to 1246. If not, additional data may be obtained and used for further comparison.

[0208] At 1246, the third expression data 132(3) is stored. The third expression data 132(3) is associated with the first identification data. For example, the third expression data 132(3) may be stored as registered expression data 146 and associated with the identification data 144.

[0209] At 1248, the second transaction data is determined to be associated with the first identification data. For example, the first transaction data may be considered to be associated with the identification data 144 in the registered user data 142.

[0210] The second transaction data is processed using the first identification data at 1250. Continuing with this example, user identification information or payment account information associated with the first identification data may be determined and used to process the second transaction data and charge the payment account specified by the payment account information.

[0211] 13 is a flow diagram 1300 of a process for performing a pre-registration process using cross-comparison, according to some implementations. The process may be performed by one or more computing devices of system 100, such as a computing device of general input device 170, computing device 106 of biometric input device 104, computing device 118, or other devices.

[0212] A user opts in to use the system, and subsequent processes may then be initiated following the user's opt-in.

[0213] At 1302, a first device is determined to be approved for use in the pre-registration process. For example, the first device may be a generic input device 170 having a specified make, model, manufacturer, operating system, specified applications, application versions, geographic location, etc. In one implementation, the first device may transmit information, such as first supplemental data 174(1), to computing device 118 as part of a request to initiate the pre-registration process. Computing device 118 may determine that first supplemental data 174(1) indicates an approved make and model of generic input device 170 located in an approved geographic location or range, such as a specified state or country. In some implementations, generic input device 170 approved for use may include a minimum resolution camera, flash capable of specified lighting, a secure computing environment (SCE), etc.

[0214] At 1304, first input image data is acquired using a first device. For example, general input device 170 may be used to acquire general input image data 172 at a first time.

[0215] At 1306, first supplemental data 174(1) associated with the first input image data is determined. The first supplemental data 174(1) may be determined by the first device or by another device or system associated with the first device. For example, the first supplemental data 174(1) may be determined by the general input device 170 concurrently with acquisition by the general input image data 172. In another example, an application server that communicates with an application executing on the general input device 170 to acquire the general input image data 172 may determine the first supplemental data 174(1).

[0216] At 1308, first intersection representation data in intersection space is determined based on the first input image data. For example, general input image data 172 may be processed as described with respect to FIG. 7 to determine first intersection representation data 454(1).

[0217] First identification data associated with the first input image data is determined at 1310. For example, the pre-enrollment identification data 184 may be received from the general input device 170 or may be retrieved from storage.

[0218] At 1312, one or more of first supplemental data 174(1), first intersection representation data, or first identification data are stored. For example, first intersection representation data may be stored as pre-registered representation data 182. Continuing with the example, pre-registered identification data 184 is stored. Explaining the example further, first supplemental data 174(1) may be stored. Also, an association between pre-registered representation data 182, pre-registered identification data 184, and first supplemental data 174(1) is stored.

[0219] At 1314, second input image data is acquired using a second device. For example, biometric input device 104 may be used to acquire biometric input image data 112 at a second time.

[0220] At 1316, second expression data 132(2) is determined based on the second input image data. For example, biometric input image data 112 may be processed by one or more of processing modules 130 to determine second expression data 132(2).

[0221] At 1318, it is determined that the second expression data 132(2) is not associated with the stored enrolled user data 142. For example, the identification module 150 may accept the second expression data 132(2) as input and return data indicating that there is no corresponding enrolled expression data 146.

[0222] At 1320, second supplemental data 174(2) associated with the second input image data is determined. The second supplemental data 174(2) may be determined by the biometric input device 104 or by another device or system associated with the biometric input device 104. For example, the second supplemental data 174(2) may be determined by the biometric input device 104 contemporaneously with the acquisition of the biometric input image data 112. Continuing with this example, the supplemental data 174(2) may include data indicative of operation of the biometric input device 104 while acquiring the second input image data, data acquired from electronically reading an EMV card, etc.

[0223] At 1322, it is determined that the second supplemental data 174(2) corresponds to one or more of the first identification data or the first supplemental data 174(1). For example, the supplemental data 174 may specify the geographic location information and time of use of each device. If the geographic location information is within a threshold distance and the time of use is within a threshold time, the two may be considered to correspond. In another example, if the first identification data indicates a user with the legal name "Abhi Patel" and the electronically retrieved EMV card data is associated with the legal name "Abhi Patel," the two may be considered to correspond.

[0224] At 1324, second intersection representation data in the intersection space is determined based on the second input image data. For example, the biometric input image data 112 may be processed as described with respect to FIG. 7 to determine the second intersection representation data 454(2).

[0225] At 1326, it is determined that the first intersection representation data is associated with the second intersection representation data. For example, intersection output data 702 may be determined and processed to determine comparison result data 724 indicating that the first intersection representation data 454(1) is associated with the second intersection representation data 454(2).

[0226] Because first cross-representation data 454(1) is based on processing of general input image data 172 and second cross-representation data 454(2) is based on processing of biometric input image data 112, general input image data 172 is considered to correspond to biometric input image data 112. Similarly, other data based thereon is also considered to correspond. For example, first representation data 420 is considered to correspond to second representation data 422.

[0227] At 1328, the second expression data 132(2) and the first identification data are stored. For example, the second expression data 132(2) may be stored as registered expression data 146, and the pre-registered identification data 184 or at least a portion of the data based thereon may be stored as identification data 144. An association between a particular registered expression data 146 and a particular identification data 144 is also stored.

[0228] Once the registration process is complete, the associated pre-registration data 136 may be deleted or otherwise discarded.

[0229] 14 is a flow diagram 1400 of a process for performing a pre-registration process using cross-comparison, according to some implementations. The process may be performed by one or more computing devices of system 100, such as a computing device of general input device 170, computing device 106 of biometric input device 104, computing device 118, or other devices.

[0230] The user opts in to use the system. The following process may then be initiated following the user opting in:

[0231] At 1402, a first device is determined to be approved for use in the pre-registration process. For example, the first device may be a generic input device 170 having a specified make, model, manufacturer, operating system, specified applications, application versions, geographic location, etc. In one implementation, the first device may transmit information, such as first supplemental data 174(1), to computing device 118 as part of a request to initiate the pre-registration process. Computing device 118 may determine that first supplemental data 174(1) indicates an approved make and model of generic input device 170 located in an approved geographic location or range, such as a specified state or country. In some implementations, generic input device 170 approved for use may include a minimum resolution camera, flash capable of specified lighting, SCE, etc.

[0232] At 1404, first input image data is acquired using a first device. For example, general input device 170 may be used to acquire general input image data 172 at a first time.

[0233] At 1406, first supplemental data 174(1) associated with the first input image data is determined. The first supplemental data 174(1) may be determined by the first device or by another device or system associated with the first device. For example, the first supplemental data 174(1) may be determined by the generic input device 170 contemporaneously with acquiring the generic input image data 172. In another example, an application server that communicates with an application executing on the generic input device 170 to acquire the generic input image data 172 may determine the first supplemental data 174(1).

[0234] At 1408, first representation data in a first representation space is determined based on the first input image data. For example, the general input image data 172 may be processed as described with respect to FIG. 10 to determine the first representation data 1010.

[0235] At 1410, first identification data is determined. For example, the pre-registration identification data 184 may be received from the general input device 170 or may be retrieved from storage.

[0236] At 1412, one or more of first supplemental data 174(1), first expression data, or first identification data are stored. For example, first expression data may be stored as pre-registered expression data 182. Continuing with the example, pre-registered identification data 184 is stored. Explaining the example further, first supplemental data 174(1) may be stored. Also, an association between pre-registered expression data 182, pre-registered identification data 184, and first supplemental data 174(1) is stored.

[0237] At 1414, first transformed representation data 914 in the second representation space is determined based on the first representation data. For example, the trained transformation network module 910 can accept the first representation data 1010 as input and provide the first transformed representation data 914 as output.

[0238] In some implementations, the operation at 1416 may be performed. At 1416, it is determined that the first transformed expression data 914 is not associated with stored enrolled user data 142. For example, the identification module 150 may accept the first transformed expression data 914 as input and return data indicating that there is no corresponding enrolled expression data 146.

[0239] At 1418, second input image data is acquired using a second device. For example, biometric input device 104 may be used to acquire biometric input image data 112 at a second time.

[0240] At 1420, second representation data 1012 in the second representation space is determined based on the second input image data. For example, the biometric input image data 112 may be processed by one or more of the processing modules 130 to determine the second representation data 1012.

[0241] At 1422, it is determined that the second expression data 1012 is not associated with the stored enrolled user data 142. For example, the identification module 150 may accept the second expression data 1012 as input and return data indicating that there is no corresponding enrolled expression data 146.

[0242] At 1424, second supplemental data 174(2) associated with the second input image data is determined. The second supplemental data 174(2) may be determined by the biometric input device 104 or by another device or system associated with the biometric input device 104. For example, the second supplemental data 174(2) may be determined by the biometric input device 104 contemporaneously with the acquisition of the biometric input image data 112. Continuing with this example, the second supplemental data 174(2) may include data indicative of operation of the biometric input device 104 while acquiring the second input image data, data acquired from electronically reading an EMV card, etc.

[0243] At 1426, it is determined that the second supplemental data 174(2) corresponds to one or more of the first identification data or the first supplemental data 174(1). For example, the supplemental data 174 may specify the geographic location information and time of use of each device. If the geographic location information is within a threshold distance and the time of use is within a threshold time, the two may be considered to correspond. In another example, if the first identification data indicates a user with the legal name "Abhi Patel" and the electronically retrieved EMV card data is associated with the legal name "Abhi Patel," the two may be considered to correspond.

[0244] At 1428, it is determined that the first transformed expression data 914 is associated with the second expression data 1012. For example, the transformed expression data 914 and the second expression data 1012 may be processed as described with respect to FIG. 10 to determine comparison result data 1024 indicating that the transformed expression data 914 is associated with the second expression data 1012. Because the transformed expression data 914 is based on the first expression data 1010, the first expression data 1010 is deemed to be associated with the second expression data 1012. As a result, the general input image data 172 is deemed to correspond to the biometric input image data 112.

[0245] At 1430, the second expression data 1012 and the first identification data are stored. For example, the second expression data 1012 may be stored as registered expression data 146, and the pre-registered identification data 184 or at least a portion of the data based thereon may be stored as identification data 144. Also, an association between a particular registered expression data 146 and a particular identification data 144 is stored.

[0246] Once the registration process is complete, the associated pre-registration data 136 may be deleted or otherwise discarded.

[0247] The devices and techniques described in this disclosure may be used in a variety of other settings. For example, system 100 may be used in conjunction with a point-of-sale (POS) device. A user may present hand 102 to biometric input device 104 to provide an indication of intent and authorization to pay via an account associated with asserted identification data 152. In another example, a robot may incorporate biometric input device 104. The robot may use asserted identification data 152 to determine whether to deliver a package to a user and, based on asserted identification data 152, to determine which package to deliver.

[0248] Although the input to system 100 is discussed with respect to image data, the system may be used with other types of input. For example, the input may include data obtained from one or more sensors 1116, data generated by another system, etc. For example, instead of image data generated by camera 108, the input to system 100 may include an array of data. Other modalities may be used. For example, the first modality may be visible light, the second modality may be sonar, etc.

[0249] Although system 100 is discussed with respect to processing biometric data, the system may be used with other types of data. For example, input may include satellite weather imagery, weather data, product images, data indicative of chemical composition, etc. For example, instead of image data generated by camera 108, input to system 100 may include an array of data.

[0250] The processes discussed herein may be implemented in hardware, software, or a combination thereof. In the software context, the described operations represent computer-executable instructions stored on one or more non-transitory computer-readable storage media that, when executed by one or more processors, perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, and data structures that perform particular functions or implement particular abstract data types. Those skilled in the art will readily recognize that particular steps or operations depicted in the above figures may be removed, combined, or performed in an alternative order. Any steps or operations may be performed serially or in parallel. Furthermore, the order in which operations are described is not intended to be construed as a limitation.

[0251] Embodiments may be provided as a software program or computer program product that includes a non-transitory computer-readable storage medium having stored thereon instructions (in compressed or uncompressed form), which can be used to program a computer (or other electronic device) to perform the processes or methods described herein. The computer-readable storage medium may be one or more of an electronic storage medium, a magnetic storage medium, an optical storage medium, a quantum storage medium, and the like. For example, the computer-readable storage medium may include, but is not limited to, a hard drive, an optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable ROM (EPROM), an electronically erasable programmable ROM (EEPROM), a flash memory, a magnetic or optical card, a solid-state memory device, or any other type of physical medium suitable for storing electronic instructions. Furthermore, embodiments may also be provided as a computer program product that includes a transitory machine-readable signal (in compressed or uncompressed form). Examples of transitory machine-readable signals include signals that can be configured to be accessed by a computer system or machine that hosts or executes a computer program, including, but not limited to, signals transmitted over one or more networks, whether modulated or unmodulated with a carrier. For example, a transitory machine-readable signal may include the transmission of software over the Internet.

[0252] The separate instances of those programs may be running on or distributed across any number of separate computer systems. Thus, while particular steps have been described as being performed by particular devices, software programs, processes, or entities, this need not be the case, and various alternative implementations will be appreciated by those skilled in the art.

[0253] Additionally, those skilled in the art will readily recognize that the above-described techniques may be utilized in a variety of devices, environments, and situations. Although the subject matter has been described in language specific to structural features or methodological acts, it will be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.

[0254] Embodiments of the present disclosure can be described in light of the following clauses.

[0255] (Clause 1) One or more memories storing first computer-executable instructions; one or more hardware processors; wherein the one or more hardware processors execute the first computer-executable instructions to: acquiring first input image data using a first device at a first time, the first input image data being associated with a first modality; determining first representation data using a first portion of a machine learning network at a second time to process the first input image data; determining, at a third time, first identification data associated with the first input image data; At a fourth time, storing the first identification data and the first expression data; At a fifth time, acquiring second input image data using a second device, the second input image data being associated with a second modality; and determining second representation data using a second portion of the machine learning network at a sixth time to process the second input image data; and at a seventh time, determining that the first expression data is associated with the second expression data; storing, at an eighth time, the first identification data and the second expression data, wherein the first identification data is associated with the second expression data; and The system.

[0256] (Clause 2) The one or more hardware processors execute the first computer-executable instructions to: determining transaction data associated with the second device after the fourth time period; and determining, after the eighth time, that the first identification data is associated with the transaction data; and processing the transaction data using the first identification data; 2. The system according to claim 1,

[0257] (Clause 3) The one or more hardware processors execute the first computer-executable instructions to: The system of clause 1 or clause 2, determining that use of one or more of the first device or applications running on the first device is authorized prior to the first time.

[0258] (Clause 4) The one or more hardware processors execute the first computer-executable instructions to: The system of any one of clauses 1 to 3, wherein, prior to the seventh time, it is determined that the second expression data is not associated with previously registered expression data.

[0259] (Clause 5) The one or more hardware processors execute the first computer-executable instructions to: The first expression data, the first representation data is associated with the second representation data; or the first expression data has been stored for longer than a specified time; The system according to any one of clauses 1 to 4, wherein the system is deleted after one of the above.

[0260] (Clause 6) The one or more hardware processors execute the first computer-executable instructions to: determining first supplemental data associated with the first device and the first time; determining second supplemental data associated with the second device and the second time; determining, prior to the seventh time, that at least a portion of the first supplemental data corresponds to the second supplemental data; The system according to any one of clauses 1 to 5,

[0261] (Clause 7) The one or more hardware processors execute the first computer-executable instructions to: determining supplemental data associated with the second device and the second time; determining, prior to the seventh time, that at least a portion of the supplemental data corresponds to the first identification data; The system according to any one of clauses 1 to 6,

[0262] (Clause 8) The one or more hardware processors execute the first computer-executable instructions to: processing the first input image data using a third portion of the machine learning network to determine a first intermediate representation; processing the first intermediate representation data using a fourth portion of the machine learning network to determine first cross representation data; processing the second input image data using a fifth portion of the machine learning network to determine second intermediate representation data; and processing the second intermediate representation data using a sixth portion of the machine learning network to determine second cross representation data; and determining, based on the first cross-representation data and the second cross-representation data, that the first representation data is associated with the second representation data; The system according to any one of clauses 1 to 7,

[0263] (Clause 9) The one or more hardware processors execute the first computer-executable instructions to: determining a first distance in an intersection embedding space between the first intersection representation data and the second intersection representation data; determining that the first representation data is associated with the second representation data based on the first distance being less than a threshold distance; The system according to any one of clauses 1 to 8,

[0264] (Clause 10) The one or more hardware processors execute the first computer-executable instructions to: determining transformed representation data using a second machine learning network to process the first representation data, wherein the transformed representation data and the second representation data are associated with an embedding space; and determining, based on the transformed expression data and the second expression data, that the first expression data is associated with the second expression data; The system according to any one of clauses 1 to 9,

[0265] Clause 11. A computer-implemented method comprising: acquiring first input image data using a first device; determining first expression data based on the first input image data; determining first identification data associated with the first input image data; storing the first identification data and the first expression data; acquiring second input image data using a second device; determining second expression data based on the second input image data; determining that the first expression data is associated with the second expression data; storing the first identification data and the second expression data, wherein the first identification data is associated with the second expression data; The computer-implemented method comprising:

[0266] (Clause 12) Obtaining third input image data using a third device; determining third expression data based on the third input image data; determining that the third expression data is associated with one or more of the first expression data or the second expression data; storing the third expression data, wherein the first identification data is associated with the third expression data; 12. The method of clause 11, further comprising:

[0267] (Clause 13) determining first supplemental data associated with the first input image data; determining second supplemental data associated with the second input image data; determining that at least a portion of the first supplemental data corresponds to the second supplemental data; 13. The method of clause 11 or clause 12, further comprising:

[0268] (Clause 14) processing the first input image data using a first portion of a machine learning network to determine first intermediate representation data; processing the first intermediate representation data using a second portion of the machine learning network to determine first cross representation data; processing the second input image data using a third portion of the machine learning network to determine a second intermediate representation; processing the second intermediate representation data using a fourth portion of the machine learning network to determine second cross representation data; determining, based on the first cross-representation data and the second cross-representation data, that the first representation data is associated with the second representation data; 14. The method of any one of clauses 11 to 13, further comprising:

[0269] (Clause 15) determining transformed representation data using a machine learning network to process the first representation data, wherein the transformed representation data and the second representation data are associated with an embedding space; and determining, based on the transformed expression data and the second expression data, that the first expression data is associated with the second expression data; 15. The method of any one of clauses 11 to 14, further comprising:

[0270] (Clause 16) One or more memories storing first computer-executable instructions; one or more hardware processors; wherein the one or more hardware processors execute the first computer-executable instructions to: acquiring first input image data using a first device; determining first expression data based on the first input image data; determining first identification data associated with the first input image data; storing the first identification data and the first expression data; acquiring second input image data using a second device; determining second expression data based on the second input image data; determining that the first expression data is associated with the second expression data; storing the first identification data and the second expression data, wherein the first identification data is associated with the second expression data; The system.

[0271] (Clause 17) The first device acquires the first input image data using a first modality; the second device acquires the second input image data using at least a second modality; The one or more hardware processors execute the first computer-executable instructions to: acquiring third input image data using a third device, wherein the third device acquires the third input image data utilizing at least a third modality; determining third expression data based on the third input image data; determining that the third expression data is associated with one or more of the first expression data or the second expression data; storing the third expression data, wherein the first identification data is associated with the third expression data; 17. The system of claim 16,

[0272] (Clause 18) The one or more hardware processors execute the first computer-executable instructions to: determining transaction data associated with the second device; determining that the first identification data is associated with the transaction data; processing the transaction data using the first identification data; 17. A system according to clause 16 or clause 17,

[0273] (Clause 19) The one or more hardware processors execute the first computer-executable instructions to: determining first supplemental data associated with the first input image data; determining second supplemental data associated with the second input image data; determining that at least a portion of the first supplemental data corresponds to the second supplemental data; The system according to any one of clauses 16 to 18,

[0274] (Clause 20) The system described in any one of clauses 16 to 19, wherein the first input image data is associated with a first modality and the second input image data is associated with a second modality.

Claims

1. one or more memories storing first computer-executable instructions; one or more hardware processors; wherein the one or more hardware processors execute the first computer-executable instructions to: acquiring first input image data using a first device; determining first representation data based on the first input image data; determining first identification data associated with the first input image data; storing the first identification data and the first expression data; acquiring second input image data using a second device; determining second representation data based on the second input image data; determining that the first expression data is associated with the second expression data; storing the first identification data and the second expression data, wherein the first identification data is associated with the second expression data; The system performs the above.

2. The system of claim 1 , wherein the first input image data is associated with a first modality and the second input image data is associated with a second modality.

3. The one or more hardware processors further execute the first computer-executable instructions to: determining transaction data associated with the second device; determining that the first identification data is associated with the transaction data; processing the transaction data using the first identification data; The system according to claim 1 or claim 2,

4. The one or more hardware processors further execute the first computer-executable instructions to: The system of any one of claims 1 to 3, further comprising determining that one or more subsequent uses of the first device or an application running on the first device are authorized before the first input image data is acquired.

5. The one or more hardware processors further execute the first computer-executable instructions to: A system as described in any one of claims 1 to 4, wherein before determining that the first expression data is associated with the second expression data, it is determined that the second expression data is not associated with previously registered expression data.

6. The one or more hardware processors further execute the first computer-executable instructions to: determining first supplemental data associated with the first input image data; determining second supplemental data associated with the second input image data; determining that at least a portion of the first supplemental data corresponds to the second supplemental data before determining that the first representation data is associated with the second representation data; The system according to any one of claims 1 to 5,

7. The one or more hardware processors further execute the first computer-executable instructions to: determining supplemental data associated with the second input image data; determining that at least a portion of the first supplemental data corresponds to the first identification data before determining that the first expression data is associated with the second expression data; The system according to any one of claims 1 to 6,

8. The one or more hardware processors further execute the first computer-executable instructions to: determining a first intermediate representation based on the first input image data; determining first intersection representation data based on the first intermediate representation data; determining second intermediate representation data based on the second input image data; determining second intersection representation data based on the second intermediate representation data; determining, based on the first cross-representation data and the second cross-representation data, that the first representation data is associated with the second representation data; The system according to any one of claims 1 to 7,

9. The one or more hardware processors further execute the first computer-executable instructions to: determining a first distance in an intersection embedding space between the first intersection representation data and the second intersection representation data; determining that the first expression data is associated with the second expression data based on the first distance being less than a threshold distance; The system according to any one of claims 1 to 8,

10. The one or more hardware processors further execute the first computer-executable instructions to: determining transformed representation data using a machine learning network to process the first representation data, wherein the transformed representation data and the second representation data are associated with an embedding space; determining, based on the transformed expression data and the second expression data, that the first expression data is associated with the second expression data; The system according to any one of claims 1 to 9,

11. 1. A computer-implemented method comprising: acquiring first input image data using a first device; determining first representation data based on the first input image data; determining first identification data associated with the first input image data; storing the first identification data and the first expression data; acquiring second input image data using a second device; determining second representation data based on the second input image data; determining that the first expression data is associated with the second expression data; storing the first identification data and the second expression data, wherein the first identification data is associated with the second expression data; The computer-implemented method comprising:

12. acquiring third input image data using a third device; determining third expression data based on the third input image data; determining that the third expression data is associated with one or more of the first expression data or the second expression data; storing the third expression data, wherein the first identification data is associated with the third expression data; The method of claim 11 further comprising:

13. determining first supplemental data associated with the first input image data; determining second supplemental data associated with the second input image data; determining that at least a portion of the first supplemental data corresponds to the second supplemental data; 13. The method of claim 11 or claim 12, further comprising:

14. processing the first input image data using a first portion of a machine learning network to determine a first intermediate representation; processing the first intermediate representation data using a second portion of the machine learning network to determine first intersecting representation data; processing the second input image data using a third portion of the machine learning network to determine a second intermediate representation; processing the second intermediate representation data using a fourth portion of the machine learning network to determine second cross representation data; and determining, based on the first cross-representation data and the second cross-representation data, that the first representation data is associated with the second representation data; The method of any one of claims 11 to 13, further comprising:

15. determining transformed representation data using a machine learning network to process the first representation data, wherein the transformed representation data and the second representation data are associated with an embedding space; determining, based on the transformed expression data and the second expression data, that the first expression data is associated with the second expression data; The method of any one of claims 11 to 14, further comprising:

Citation Information

Patent Citations

  • Method for associating multiple modalities, program therefor, and multi-modal system for associating multiple modalities

    JP2008129713A

  • Contactless biometric authentication system

    JP2021527872A

  • User tag generation method, device, computer program, and computer device

    JP2022508163A

  • Conditional and situational biometric authentication and enrollment

    US20140313007A1

  • Mobile enrollment using a known biometric

    US20200322328A1