Authentication system and authentication method
The authentication system enhances compatibility and accuracy by using multiple feature extraction models to handle different imaging methods and temporal variations, addressing the challenges of device incompatibility and biometric changes.
Patent Information
- Application Number
- JP2022060239
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-10-22
- Estimated Expiration
- 2042-03-31
AI Technical Summary
Existing biometric authentication systems face challenges in maintaining compatibility between different authentication devices and accommodating temporal changes in biometric features, leading to decreased accuracy and increased development costs.
An authentication system that employs a processor and memory to extract and compare biometric features using multiple imaging methods, including a first and second model to perform authentication based on captured images, ensuring compatibility and reducing the need for re-registration due to device changes or temporal variations.
Improves authentication accuracy while reducing development costs by maintaining compatibility across multiple systems and adapting to temporal changes in biometric features.
Smart Images

Figure 0007758627000001 
Figure 0007758627000002 
Figure 0007758627000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an authentication system and an authentication method. [Background technology]
[0002] In recent years, biometric authentication technology has begun to be used not only for access management such as entrance / exit control and PC login, but also as a means of identity authentication when making payments at point-of-sale (POS) registers in stores, with credit cards, and for online shopping, etc. In addition, the need for accurate identity authentication is increasing with the expansion of cashless payments using smartphones, and there is a demand for highly accurate biometric authentication technology that can be easily used on smartphones.
[0003] Currently widely used biometric authentication methods include face, iris, fingerprint, and finger and palm veins. In particular, finger authentication technology based on finger veins, fingerprints, skin prints, and knuckle prints is gaining popularity due to its high authentication accuracy and simple authentication operation, as it can utilize a wide range of biometric features.
[0004] In actual operation, authentication devices with various sensing methods are used depending on the application. For example, there are authentication devices that take high-resolution images of fingers inserted into holes in the device, small, open authentication devices that allow authentication by simply placing a finger on the device, contactless authentication devices that can simultaneously capture images of multiple fingers by irradiating them with reflected infrared and visible light when multiple fingers are held over the device, and authentication devices that take color images of fingers under visible light conditions using the rear camera of a typical smartphone or the front camera of a notebook PC (Pestronical Computer).
[0005] Since various types of authentication devices have different imaging principles and authentication algorithms, the clarity of the observed biometric features may differ, and the authentication algorithm may be customized for each authentication device, which may result in incompatibility between the enrollment templates registered on each authentication device.
[0006] Background art for realizing compatible biometric authentication is found in International Publication No. 2011 / 061862 (Patent Document 1), which describes the authentication system as comprising: "an input device for placing the biometric subject; an imaging device for photographing the biometric subject; a storage device for storing a plurality of first feature data including biometric features and pre-registered, and second feature data generated from each of the first feature data; and a matching processing unit for matching input data indicating the biometric features photographed by the imaging device with each of the first and second feature data, wherein each of the second feature data is smaller in size than each of the first feature data and includes at least a part of the biometric features" (see Abstract). [Prior art documents] [Patent documents]
[0007] [Patent Document 1] International Publication No. 2011 / 061862 Summary of the Invention [Problem to be solved by the invention]
[0008] Patent Document 1 discloses technology related to the automation of re-registration to maintain compatibility among finger vein authentication devices that differ in the parts of the body photographed, camera distortion, etc. The technology described in Patent Document 1 acquires image deformation information using a test chart or the like to calibrate camera distortion in order to generate patterns that are as similar as possible across multiple different devices, and performs geometric deformation correction to make it possible to determine similarity, and updates the registration data when a certain level of similarity is obtained.
[0009] However, Patent Document 1 does not describe a learning model that simultaneously eliminates compatibility between various authentication devices. Furthermore, the technology described in Patent Document 1 is premised on the assumption that the imaging principle is standardized to the infrared transmission method, and does not describe maintaining compatibility between different imaging methods, such as between visible light images and infrared images, or between reflected images and transmitted images.
[0010] The above-mentioned problems are not limited to finger veins, but also apply to various other biological elements such as the face, iris, auricle, facial veins, subconjunctival blood vessels, palmar veins, back of the hand veins, palm prints, inner and outer joint prints of the fingers, veins on the dorsum of the fingers, subcutaneous fat prints, and melanin prints, as well as to multimodal biometric authentication that combines these elements.
[0011] As described above, the technology described in Patent Document 1 has a problem in that it leads to a decrease in the compatibility accuracy between authentication devices when realizing a biometric authentication system using multiple different authentication devices. Therefore, one aspect of the present invention is to improve authentication accuracy while reducing development costs when realizing compatible authentication between multiple systems. [Means for solving the problem]
[0012] In order to solve the above-mentioned problems, one aspect of the present invention employs the following configuration. The authentication system includes a processor and a memory, and the memory stores: a first model that extracts a first feature corresponding to a first imaging method from a biometric image captured using a first imaging method corresponding to the authentication system; a second model that extracts a second feature extractable from a biometric image captured using a second imaging method corresponding to another authentication system, from the biometric image captured using the first imaging method; a biometric image of a person to be authenticated captured using the first imaging method, the first feature of a registered person, and the second feature of the registered person; the processor extracts the first feature of the person to be authenticated based on the biometric image of the person to be authenticated and the first model, and performs a first authentication of the person to be authenticated based on a comparison result between the extracted first feature of the person to be authenticated and the first feature of the registered person; and if the first authentication fails, extracts a second feature of the person to be authenticated based on the biometric image of the person to be authenticated and the second model, and performs a second authentication of the person to be authenticated based on a comparison result between the extracted second feature of the person to be authenticated and the second feature of the registered person. [Effects of the Invention]
[0013] According to one aspect of the present invention, when realizing compatible authentication between a plurality of systems, it is possible to improve authentication accuracy while reducing development costs.
[0014] Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief explanation of the drawings]
[0015] [Figure 1A] 1 is a block diagram showing an example of the overall configuration of a biometric authentication system using biometric features according to a first embodiment. [Figure 1B] 2 is a block diagram showing an example of a functional configuration of an authentication processing device according to the first embodiment; FIG. [Figure 2] 1 is a block diagram showing an example of the configuration of an integrated authentication system that maintains compatibility between authentication systems using different authentication devices according to a first embodiment. [Figure 3] 2 is a diagram illustrating an example of the internal configuration of a biometric information processing unit included in each of the authentication system A and the authentication system B according to the first embodiment. FIG. [Figure 4] 10 is a flowchart illustrating an example of a registration process for realizing compatibility between authentication systems according to the first embodiment. [Figure 5] 10 is a flowchart illustrating an example of an authentication process for realizing compatibility between authentication systems according to the first embodiment. [Figure 6A] FIG. 2 is a diagram illustrating an example of the configuration of a learning model during learning of an attribute information extraction unit and a common feature extraction unit in the first embodiment. [Figure 6B] 10 is a flowchart illustrating an example of a learning process for an attribute information extraction unit and a common feature extraction unit in the first embodiment. [Figure 7] FIG. 10 is an explanatory diagram illustrating an example of conformance determination by an attribute information extraction unit in the first embodiment. [Figure 8A] FIG. 2 is a diagram showing an example of pairs of training images as training data and non-conformance reason labels in the first embodiment. [Figure 8B] FIG. 10 is a diagram showing an example of a feature vector space onto which attribute information acquired from training images in Example 1 is mapped. [Figure 9A] 3 is a diagram illustrating an example of the configuration of a feature extraction unit that autonomously extracts any biometric feature that maximizes authentication accuracy in the first embodiment. FIG. [Figure 9B] 10A and 10B are diagrams illustrating an example of autonomous feature extraction processing when an image of a hand or fingers is captured by a visible light authentication terminal such as a smartphone in the first embodiment. [Figure 10A] 4 is a diagram illustrating an example of a learning model of a foreign object detection unit that detects the presence or absence of an abnormality in a biological image according to the first embodiment. FIG. [Figure 10B] FIG. 4 is a diagram illustrating an example of cooperation between an attribute information extraction unit and a foreign object detection unit in the first embodiment. [Figure 11A] FIG. 10 is a block diagram showing an example of the configuration of an integrated authentication system that maintains compatibility between authentication systems using different authentication devices according to a second embodiment. [Figure 11B] FIG. 11 is a diagram illustrating a configuration example of a registration database C of an authentication system C, which is an old authentication system in the second embodiment. [Figure 11C] FIG. 10 is a diagram illustrating an example of the configuration of a registration database A of an authentication system A, which is a new authentication system according to a second embodiment. [Figure 12A] FIG. 10 is a diagram illustrating an example of the internal processing configuration of a biometric information processing unit included in the authentication system A according to the second embodiment. [Figure 12B] FIG. 10 is a diagram illustrating an example of an internal processing configuration of a biometric information processing unit included in an authentication system C according to a second embodiment. [Figure 13] 10 is a flowchart showing an example of compatible authentication in which a user who has been registered only in the old authentication system is authenticated by the new authentication system in the second embodiment. [Figure 14A] FIG. 10 is a diagram illustrating an example of the configuration of a learning model of an imitation feature extraction unit B that generates imitation features of an old authentication system from an input image of a new authentication system in the second embodiment. [Figure 14B] 10 is a flowchart illustrating an example of an imitation feature generation process according to the second embodiment. [Figure 14C] FIG. 10 is a schematic diagram illustrating an example of dedicated features obtained from a biometric image of the old authentication system in the second embodiment. [Figure 14D]FIG. 10 is a schematic diagram showing an example of imitation features obtained from a biometric image of the new authentication system in the second embodiment. [Figure 15] FIG. 11 is a diagram illustrating a configuration example of a biometric information processing unit of an authentication system that performs compatible authentication for variations in biometric features that occur over time in a third embodiment. [Figure 16] 11 is a graph showing an example of changes in latent variables of a biometric image in response to short-term changes in daytime temperature and long-term changes in a user over time in Example 3. [Figure 17] 11A and 11B are schematic diagrams illustrating a change in a finger image over time and the effect of a process for suppressing the influence of the change over time in the third embodiment. [Figure 18] FIG. 10 is a block diagram showing a configuration example of an authentication system that achieves highly accurate authentication by autonomously learning an optimal control method for an authentication terminal in a fourth embodiment. [Figure 19] FIG. 10 is a diagram illustrating an example of the configuration of a feature extraction unit and a matching processing unit that autonomously learns optimal control of an authentication terminal A in a fourth embodiment. [Figure 20A] FIG. 13 is a diagram showing an example of a list of controllable device parameters of hardware included in an authentication terminal that is a control target in the fourth embodiment. [Figure 20B] FIG. 13 is a diagram showing an example of a user guidance list in the fourth embodiment. [Figure 20C] FIG. 13 is a diagram showing an example of an optimal behavior list corresponding to the attribute classification result in the fourth embodiment. [Figure 20D] FIG. 13 is a diagram illustrating a configuration example of a terminal control unit, which is a learning model that converts attribute information into optimal actions in the fourth embodiment. [Figure 21] 13 is a flowchart illustrating an example of a learning method of the terminal control unit in the fourth embodiment. [Figure 22] 13 is a flowchart illustrating an example of a learning process of a terminal control unit in the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0016] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The following description and drawings are examples for explaining the present invention, and some omissions and simplifications have been made as appropriate for clarity of explanation. The present invention can be implemented in various other forms. Unless otherwise specified, each component may be singular or plural.
[0017] In order to facilitate understanding of the invention, the position, size, shape, range, etc. of each component shown in the drawings may not represent the actual position, size, shape, range, etc. Therefore, the present invention is not necessarily limited to the position, size, shape, range, etc. disclosed in the drawings.
[0018] Furthermore, in the following description, processing performed by executing a program may be described, but the program is executed by a processor (e.g., a CPU (Central Processing Unit), a GPU (Graphics Processing Unit)) to perform the specified processing while appropriately using storage resources (e.g., memory) and / or interface devices (e.g., communication ports), etc., so the processor may be the subject of the processing. Similarly, the subject of the processing performed by executing a program may be a controller, device, system, computer, or node having a processor. The subject of the processing performed by executing a program may be any computing unit, and may include a dedicated circuit (e.g., an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit)) that performs specific processing.
[0019] In this embodiment, the biometric feature may be information indicating different anatomical biometric features such as finger veins, fingerprints, joint patterns, skin patterns, finger contour shapes, fat lobule patterns, finger length ratios, finger widths, finger areas, melanin patterns, palm veins, palm prints, back veins, facial veins, or ear veins, or information indicating multiple features extracted from information containing a mixture of multiple anatomically classified biometric parts and divided in any manner. Furthermore, the biometric feature may be information indicating a feature emitted from a living body, such as a voice. Hereinafter, in this embodiment, unless otherwise specified, an example will be described in which information indicating finger vein features is used as the biometric feature.
[0020] In the past, there were cases where the registration templates registered in each authentication device were incompatible, so even if registration had been completed in one authentication device, re-registration was required when a new authentication device was introduced or when the authentication algorithm of the authentication device was updated to improve authentication accuracy.
[0021] Even if the old and new authentication devices were linked as a system and the registration and authentication protocols and data formats for registration templates were unified, it would be difficult to correctly authenticate a person because the biometric features extracted would essentially be different.As a result, when using a new authentication device, the registration process had to be redone, which was a heavy burden for both users and system administrators.
[0022] In addition, when registering biometric information required to use biometric authentication, it is generally necessary to go to a dedicated registration counter to complete the registration process, but the location and time restrictions have made this a barrier to expanding the number of users.In contrast, if users could easily register at home using a separate device such as a personally owned smartphone, laptop, or other dedicated authentication device when using authentication devices installed in stores, the convenience of registration would be greatly improved.
[0023] To achieve this, it is necessary to link personal authentication terminals with authentication systems at stores, etc. However, since it is difficult for individuals to install dedicated terminals of the same format as those used in stores, it is necessary to enable registration using widely available devices such as smartphones. However, as mentioned above, achieving compatibility between different authentication devices is a challenge.
[0024] Furthermore, even if the authentication device and authentication system are the same, the biometrics used for authentication may change from the state at the time of enrollment due to the passage of time, physical condition, environmental influences, etc. Although most biometrics used in biometric authentication are selected to have little change over time, it is expected that they will change over time, and this change must be accommodated. Generally, re-enrollment is performed when authentication is no longer possible, but as with the issue of compatibility between devices mentioned above, this places a burden on users and system administrators, so it is desirable for the authentication system to cover this.
[0025] To solve these problems, it is necessary to develop an authentication algorithm that maintains compatibility between different devices and between biometric features that change over time.However, with conventional technology, achieving compatibility between devices increases the number of combinations as the number of device types increases, which increases development costs.In addition, since changes in biometric features over time require a large amount of development resources to analyze biometric features that are not affected by the passage of time, it is often difficult to maintain compatibility at low cost.
[0026] Thus, the challenge is to provide an authentication device or authentication system that achieves high authentication accuracy while being compatible with different authentication devices, is less susceptible to temporal biological variations, and has low management and development costs.
[0027] Furthermore, biometric authentication using fingers, in particular, involves capturing images of various biometric features, such as fingerprints, veins, joint patterns, fat patterns, and melanin patterns, using various sensing means for authentication. However, because the biometric features that can be captured clearly vary depending on the sensor used, the similarity of enrollment templates between different authentication devices often decreases even when the same finger is captured. This means that inter-device authentication is not possible, and re-enrollment is required when authenticating with an unregistered authentication device.
[0028] To ensure compatibility between different authentication devices, it is necessary to analyze and extract similar features between the two devices. This requires understanding detailed design differences, such as differences in how veins and joint patterns appear when captured by each device, and the degree of image distortion due to differences in lenses, and then standardizing these differences. However, achieving compatibility between the many authentication devices requires covering a wide range of combinations, which takes time and is not feasible at low cost. Furthermore, maintaining compatibility requires the use of legacy authentication algorithms, which hinders improvements in authentication accuracy through the application of cutting-edge technology, preventing the full potential of authentication accuracy. [Example]
[0029] FIG. 1A is a block diagram showing an example of the overall configuration of a biometric authentication system 1000 that uses biometric features. In this embodiment, the biometric authentication system 1000 includes multiple components. However, instead of the biometric authentication system 1000, a biometric authentication device that incorporates all or part of the multiple components in a single housing may be used. The biometric authentication device may be a personal authentication device that executes processes including authentication processing, or the authentication processing may be executed externally from the biometric authentication device, and the biometric authentication device may be a finger image acquisition device specialized for acquiring finger images or a finger feature image extraction device. Furthermore, the biometric authentication system 1000 may be implemented by a terminal such as a smartphone or a PC (Personal Computer).
[0030] A device including an imaging device that captures biometric images and an authentication processing unit that processes the captured image and performs biometric authentication is called a biometric authentication device, and a system including an imaging device that captures biometric images and a device that is connected to the imaging device via a network and performs authentication processing is called a biometric authentication system. Authentication systems include biometric authentication devices and biometric authentication systems.
[0031] 1A includes, for example, a biometric information acquisition device 2 that captures an image of a finger 1, an authentication processing device 10, an auxiliary storage device 14, a display device 15, an input device 16, a speaker 17, and an image input device 18. The biometric information acquisition device 2 includes an imaging device 9 installed inside a housing of the biometric information acquisition device 2, and may also include a light source 3 installed in the housing. The authentication processing device 10 has an image processing function.
[0032] The light source 3 is a light-emitting element such as an LED (Light Emitting Diode), and irradiates light onto a certain area of the user's living body (for example, the fingers 1 and face 4) presented to the biometric information acquisition device 2. Depending on the embodiment, the light source 3 may be capable of irradiating light of various wavelengths, or may be capable of irradiating light transmitted through the living body, or may be capable of irradiating light reflected from the living body. Furthermore, the light source 3 may be composed of multiple light-emitting elements, and may have a function that allows the authentication processing device 10 to adjust the light emission intensity of each of the multiple light-emitting elements.
[0033] The imaging device 9 captures images of the fingers 1 and face 4 presented to the biometric information acquisition device 2. The imaging device 9 may also capture images of biological components such as the iris, back of the hand, and palm at the same time as capturing images of the fingers 1 and face 4. The imaging device 9 is an optical sensor capable of capturing light of a single or multiple wavelengths, and may be a monochrome camera or a color camera, or may be a multispectral camera capable of simultaneously capturing ultraviolet light or infrared light in addition to visible light. The imaging device 9 may also be a distance camera capable of measuring the distance to a subject, or a stereo camera combining multiple identical cameras.
[0034] The biometric information acquisition device 2 may include a plurality of such imaging devices 9. Furthermore, the imaging device 9 may capture an image of one finger, or may capture an image of multiple fingers, or may capture an image of multiple fingers on both hands simultaneously.
[0035] The image input device 18 acquires an image captured by the imaging device 9 in the biometric image acquisition device 2, and outputs the acquired image to the authentication processing device 10. The image input device 18 is, for example, any of various reader devices (e.g., a video capture board) for reading an image.
[0036] In this embodiment, an example is described in which fingers 1 and a face 4 are used as biometric features, and therefore the biometric information acquisition device 2 includes a light source 3 and an imaging device 9, but if other biometric features are used, the biometric information acquisition device 2 is a sensor or the like for acquiring the other biometric features.
[0037] The authentication processing device 10 is configured by, for example, a computer including a CPU 11, a memory 12, and various interfaces (IF) 13. The CPU 11 is an arithmetic device that includes a processor and executes programs stored in the memory. The processor executes various programs to realize each functional unit provided by the authentication processing device.
[0038] The memory 12 includes a ROM, which is a non-volatile storage element, and a RAM, which is a volatile storage element. The ROM stores unchanging programs (e.g., BIOS) and the like. The RAM is a high-speed, volatile storage element such as a DRAM (Dynamic Random Access Memory), and temporarily stores programs executed by the processor and data used when the programs are executed. The memory 12 also temporarily stores images input from an image input device 18 and the like.
[0039] The program executed by the processor is provided to the authentication processing device 10 from a removable medium (such as a CD-ROM or flash memory) or via a network, and is stored in a non-volatile auxiliary storage device 14, which is a non-transitory storage medium. For this reason, the authentication processing device 10 should preferably have an interface for reading data from removable media.
[0040] Each interface 13 connects the authentication processing device 10 to an external device and functions as an input interface that receives input of information from the external device and / or an output interface that outputs information to the external device. Specifically, the interfaces 13 are devices having ports and the like for connecting the authentication processing device 10 to the biometric information acquisition device 2, the auxiliary storage device 14, the display device 15, the input device 16, the speaker 17, the image input device 18, and the like. Note that the authentication processing device 10 may include the biometric information acquisition device 2, the auxiliary storage device 14, the display device 15, the input device 16, the speaker 17, and the image input device 18.
[0041] The interface 13 also functions as a communication interface for the authentication processing device 10 to communicate with external devices via a communication network (not shown). The communication interface is, for example, a device that performs communication in accordance with the IEEE802.3 standard if the communication network 30 is a wired LAN (Local Area Network), and a device that performs communication in accordance with the IEEE802.11 standard if the communication network 30 is a wireless LAN.
[0042] The auxiliary storage device 14 is a large-capacity, non-volatile storage device such as a hard disk drive (HDD) or a solid state drive (SSD), and stores programs to be executed by the processor. That is, the programs are read from the auxiliary storage device 14, loaded into the memory 12, and executed by the processor, thereby realizing each function of the authentication processing device 10.
[0043] The auxiliary storage device 14 also stores user registration data, etc. The registration data is information obtained during the registration process for verifying users, and is stored in association with multiple biometric features for each user. The registration data is, for example, an image of facial features, finger features, or finger vein patterns, or a registration template in which biometric features are linked to a registration ID as user identification information.
[0044] The image of a finger vein pattern is, for example, an image of finger veins, which are blood vessels distributed under the skin of a finger, captured as a dark shadow pattern or a slightly bluish pattern.Furthermore, the feature data of a finger vein pattern is, for example, data obtained by converting an image of the vein portion into a binary or 8-bit image, data consisting of feature amounts generated from the coordinates of feature points such as the bends, branches, and endpoints of the vein or brightness information around the feature points, or encrypted data of these data.
[0045] The display device 15 is, for example, a liquid crystal display, and is an output device that displays information received from the authentication processing device 10, biometric posture guidance information, posture determination results, and other information.
[0046] The input device 16 is, for example, a keyboard or a touch panel, and receives information input from a user and transmits it to the authentication processing device 10. The speaker 17 is an output device that transmits information received from the authentication processing device 10 as an acoustic signal such as voice.
[0047] The authentication processing device 10 is a computer system configured on one physical computer or on multiple logically or physically configured computers, and may operate on a virtual computer constructed on multiple physical computer resources. For example, each functional unit may operate on a separate physical or logical computer, or multiple functional units may be combined to operate on a single physical or logical computer.
[0048] 1B is a block diagram showing an example of the functional configuration of the authentication processing device 10. Functional units that realize the respective functions are realized by programs stored (loaded) in the memory 12.
[0049] The memory 12 of the authentication processing device 10 includes a registration processing unit 20, an authentication processing unit 21, and a biometric information processing unit 27. The registration processing unit 20 associates an individual's biometric features with a personal ID (registered ID) and registers them in advance. The authentication processing unit 21 outputs an authentication result based on the biometric features extracted from an image captured during authentication and the registered biometric features.
[0050] The biometric information processing unit 27 performs image processing and signal processing necessary for performing biometric authentication. The biometric information processing unit 27 includes an image capture control unit 22, an attribute determination unit 23, a feature extraction unit 24, a matching processing unit 25, and an authentication determination unit 26.
[0051] The photographing control unit 22 photographs the presented biometric data under appropriate conditions. The attribute determination unit 23 analyzes and determines various attribute information such as the image quality of the photographed biometric image and the posture of the photographed biometric data. The feature extraction unit 24 extracts biometric features while appropriately correcting the posture of the biometric data during registration and authentication processes. The matching processing unit 25 calculates the similarity between biometric features. The authentication determination unit 26 determines the authentication result from one or more matching process results.
[0052] For example, the CPU 11 operates in accordance with the registration processing unit 20 (program) loaded into the memory 12 to execute the processing of the registration processing unit 20, and operates in accordance with the authentication processing unit 21 (program) loaded into the memory 12 to execute the processing of the authentication processing unit 21. The same relationship between the programs and the processing applies to the other functional units included in the memory 12.
[0053] 2 is a block diagram showing an example of the configuration of an integrated authentication system that maintains compatibility between authentication systems that use different authentication devices. The integrated authentication system includes multiple authentication systems, each of which includes a different authentication terminal, and achieves compatible authentication between the respective authentication systems.
[0054] In this embodiment, compatibility is maintained between two authentication systems each including a different finger authentication terminal, and the different authentication terminals are a visible light authentication terminal and an infrared reflective authentication terminal, respectively. Here, the authentication system including the visible light authentication terminal is called authentication system A41, and the authentication system including the infrared reflective authentication terminal 50 is called authentication system B42.
[0055] The authentication system A41 includes a registration database A43, an authentication server A45, and an authentication terminal A47 (visible light authentication terminal). The authentication system B42 includes a registration database B44, an authentication server B46, and an authentication terminal B48 (infrared reflective authentication terminal). The authentication server A45 and the authentication system B45 are connected to each other.
[0056] Each authentication server is realized by the authentication processing device 10 in FIG. 1, each registration database is realized by the auxiliary storage device 14 in FIG.
[0057] In both authentication systems, during enrollment, a biometric image captured by an authentication terminal of each authentication system is converted into an enrollment template on the authentication server, and the enrollment template is stored in a enrollment database. Also, in both authentication systems, during authentication, the enrollment template stored in the enrollment database is downloaded to the authentication server, and authentication processing is performed on the authentication server. Note that the authentication terminal may have the functions of the authentication server and enrollment database. Also, the enrollment database may be provided within the authentication server. Also, the enrollment database may be shared by both authentication systems.
[0058] It is also assumed that both authentication systems shown in Figure 2 are not yet in operation, and that the authentication system and authentication algorithm can be newly designed when the system is introduced.
[0059] 3 is a diagram showing an example of the internal configuration of the biometric information processing unit 27 provided in each of authentication systems A41 and B42. Unless otherwise specified in the following description, components with the same name will be distinguished from one another by adding an "a" to the end of the reference symbol for components included in authentication system A41 and adding a "b" to the end of the reference symbol for components included in authentication system B42. Furthermore, in the following description, when there is no need to distinguish between components with the same name that are included in authentication system A41 and those included in authentication system B41, or when it is necessary to specify which authentication system a component is included in, the "a" and "b" at the end of the reference symbol will be omitted.
[0060] The biometric information processing unit 27a includes an attribute information extraction unit 61a, an attribute determination unit 23a, a common feature extraction unit 63, a dedicated feature extraction unit 64a, and a matching processing unit 25a. The biometric information processing unit 27b includes an attribute information extraction unit 61b, an attribute determination unit 23b, a common feature extraction unit 63, a dedicated feature extraction unit 64b, and a matching processing unit 25b. Note that each processing unit is optimized for each authentication system, so the algorithms and associated parameters may differ for each authentication system. Also, although the imaging control unit 22 and authentication determination unit 26 are not shown in FIG. 3, they are included in the biometric information processing unit 27 as shown in FIG. 1B.
[0061] Below, we will explain each functional unit shown in Figure 3, but since similar processing is performed in the biometric information processing unit 27a and the biometric information processing unit 27b (authentication system A41, registration database A43, authentication server A45, authentication terminal A47, and dedicated feature 70a can be read as authentication system B42, registration database B44, authentication server B46, authentication terminal B48, and dedicated feature 70b, respectively), we will explain the functional units included in the biometric information processing unit 27a and omit explanation of the functional units included in the biometric information processing unit 27b.
[0062] The attribute information extraction unit 61a is a learning model that has been pre-trained to take as input a biometric image 66 taken by the authentication terminal A47 and a terminal control value 67 specific to the authentication terminal A47, and output attribute information 68 calculated from the characteristics of the biometric image 66 and the terminal control value 67.
[0063] A learning model is a machine learning framework that has, for example, hyperparameters such as the connection configuration of a deep neural network, the number of neurons in the input / output layer and hidden layer, the number of layers, and the type of activation function, and / or learning parameters such as the convolution kernel and the weights of the connections between neurons, and has the function of inferring and outputting the expected value for the input and the function of optimizing the learning parameters using training data.
[0064] The attribute information 68 is mainly used to determine whether to perform the subsequent feature extraction process and matching process. Examples of terminal control values specific to the authentication terminal A47 include values such as exposure time, white balance, gain, and focus of the camera of the biometric information acquisition device 2 included in the authentication terminal A47, as well as the illumination intensity of each of one or more light sources 3.
[0065] Additionally, if the authentication terminal A47 is a general-purpose device such as a smartphone, the attribute information 68 may include location information based on an illuminance sensor, an acceleration sensor, or a GPS (Global Positioning System) provided inside the authentication terminal A47. The attribute information 68 obtained by the attribute information extraction unit 61 includes feature amounts related to the currently input image and the state of the authentication terminal A47, and is used for attribute determination in a subsequent stage.
[0066] The attribute determination unit 62a receives the attribute information 68 output by the attribute information extraction unit 61a, determines the currently input image 66 and the state of the authentication terminal A47, and outputs an attribute determination result 71.
[0067] The attribute determination unit 62a, for example, determines whether the biometric data has been photographed in an appropriate environment, and determines whether to perform processing such as feature extraction or matching processing based on the determination result, thereby making it possible to perform an environmental determination as to whether the state in which the biometric data was photographed is an appropriate photographing state suitable for performing authentication processing.
[0068] Examples of photographic conditions that are inappropriate for performing authentication processing include when the biometrics is not presented, when the biometrics is not held in the appropriate location, when the image is too dark, when the image is too bright, when the finger has moved and the image is blurred, and when a foreign object or counterfeit is held.
[0069] If any of these events occur, the biometric image can be photographed repeatedly without performing the matching process. A list of reasons for non-compliance (reasons for the state being unsuitable for performing the authentication process) may be prepared in advance, and the attribute determination unit 62a may analyze the attribute information, obtain reasons for non-compliance from the list based on the analysis results, and display the reasons for non-compliance as guidance to the user. This allows the user to confirm the reasons for non-compliance, allowing them to hold their biometric image over the imaging device 9 in a state where authentication is more likely to be successful, thereby improving authentication accuracy and operability.
[0070] The common feature extraction unit 63 is a learning model that has been pre-trained to take a biometric image taken by the authentication terminal A47 as input and output a common feature 69 that can be commonly acquired by the authentication terminal A47 and the authentication terminal used in the target authentication system (authentication system B42) with which compatibility should be achieved.
[0071] The common feature extraction unit 63 is assumed to be trained to maximize authentication accuracy during compatible authentication with the authentication system with which compatibility is to be achieved. In other words, the common feature extraction unit 63 outputs common features that can be extracted from biometric images of both its own authentication system and the target authentication system with which compatibility is to be achieved.
[0072] In this embodiment, there are two authentication systems to be linked, and one authentication system only needs to be compatible with the other authentication system, so each authentication system only needs to include one common feature extraction unit 63. The common feature extraction unit 63 can use the same learning parameters in both authentication systems. When the authentication system is connected to two or more other systems to achieve compatible authentication, the system includes as many common feature extraction units 63 as the number of other connected authentication systems. Note that when an authentication system has multiple common feature extraction units 63, their learning models may be combined into a single common feature extraction unit.
[0073] The dedicated feature extraction unit 64a is a learning model that is pre-trained to input a biometric image captured by the authentication terminal A47a and output a dedicated feature 70a that maximizes authentication accuracy within the authentication system A41a. Authentication within each authentication system can be performed using the common feature 69 obtained by the common feature extraction unit 63, but by separately providing a dedicated feature extraction unit 64a that is trained to acquire features specialized for the characteristics of the authentication system itself, more accurate authentication can be achieved when compatible authentication is not required. While the common feature extraction unit 63 and the dedicated feature extraction unit 64a are described as separate modules in this embodiment, they may also be integrated into a single feature extraction unit 24.
[0074] The matching processing unit 25 is a learning model that is trained in advance to calculate the degree of difference between the common feature 69 or the dedicated feature 70a acquired during authentication and the common feature or the dedicated feature as a registered template that has been registered in advance.
[0075] The matching processing unit 25 compares the common features when performing compatible authentication between both authentication systems, or compares the dedicated features when performing authentication within its own authentication system, to calculate a matching score 72. The matching score 72 indicates the degree of difference (distance) between the enrolled template and the biometric features at the time of authentication. The matching processing unit 25 passes the matching score to the subsequent authentication assessment unit 26, which then compares the matching score with an authentication threshold.
[0076] If the authentication determination unit 26 determines that the matching score is smaller than the threshold, it determines that the registered template and the biometric feature at the time of authentication match, and determines that the authentication is successful. Specific examples of the matching score include Manhattan distance (L1 norm), Euclidean distance (L2 norm), and cosine similarity between features. Note that, since the calculation of these distance values is often an easy calculation to formulate, the matching processing unit 25 does not necessarily need to be a learning model.
[0077] All of the above-mentioned learning models can be constructed using deep neural networks, commonly known as deep learning. Specific baseline models include commonly known methods such as LuNet, AlexNet, VGG, GoogLeNet, ResNet, MobileNet, and SqueezeNet. Specific learning algorithms include style conversion based on generative adversarial networks (GANs) and feature extraction based on distance learning, such as ArcFace and AdaCos.
[0078] Fig. 4 is a flowchart showing an example of a registration process for achieving compatibility between authentication systems. Fig. 4 illustrates an example in which a user who has not completed registration with two authentication systems, A41 and B42, registers with the two authentication systems using authentication system A41 including a personally owned smartphone as authentication terminal A47.
[0079] The user launches the authentication application pre-installed on their authentication terminal A47, enters the necessary information, and selects to start registration. The authentication terminal A47 displays a preview image of their own rear camera, which is an example of the imaging device 9, and finger positioning guidance on the screen of the authentication terminal A47 (S401), and starts capturing a finger image while performing device-specific control such as the rear camera parameters and light source intensity (S402). At this time, the user follows the guidance and holds their finger over the screen so that the position of their finger aligns with the guidance on the screen. Details of the guidance screen will be described later.
[0080] The authentication terminal A47 transmits the captured finger image to the authentication server A45, the attribute information extraction unit 61a extracts attribute information from the finger image (S403), and the attribute determination unit 23a performs attribute determination (S404). In the attribute determination, the quality of the currently captured biometric image is determined based on, for example, the presence or absence of a finger, the finger's posture, the environmental brightness, and other conditions.
[0081] The attribute determination unit 23a determines whether the finger image is suitable for authentication based on the determination result in step S404 (S405), and if it is determined to be unsuitable (S405: N), the process returns to displaying the preview image and finger guide again (S401). At this time, if the attribute determination unit 23a can determine the reason why the finger image is determined to be unsuitable for authentication, guidance for resolving the reason may be displayed on the screen of the authentication terminal A47.
[0082] If the attribute determination unit 23a determines that the finger image is suitable for authentication (S405: Y), the dedicated feature extraction unit 64a performs dedicated feature extraction processing (S406). The dedicated feature extraction unit 64a, which is configured using the above-mentioned neural network or the like, extracts dedicated feature amounts 69a that maximize the authentication accuracy of the visible light image captured by the rear camera (image capture device 9) of the authentication terminal A47.
[0083] The common feature extraction unit 63 of the authentication server A45 extracts a common feature 70 that can be extracted in common between an image captured by the authentication terminal A47, which is a visible light authentication terminal, and an image captured by an infrared reflective authentication device included in the authentication system B42 that performs compatible authentication (S407). The registration processing unit 20 of the authentication server A45 stores the dedicated feature 69a extracted in step S406 and the common feature 70 extracted in step S407 in the memory 12 of the authentication server A45 as features to be registered (S408).
[0084] The registration processing unit 20 of the authentication server A45 determines whether three sets of finger registration candidates (sets of dedicated feature 69a and common feature 70) have been stored (S409), and if three sets have not been stored (S409: N), the process returns to displaying the preview image (S401). If three sets of finger registration candidates have been stored (S409: Y), the matching processing unit 25 of the authentication server A45 performs registration selection processing (S410).
[0085] This is a process to confirm that registration data has been acquired stably in the registration selection process. In this embodiment, the matching processing unit 25 of the authentication server A45 acquires three sets of registration candidates and confirms their stability.
[0086] In the registration selection process, the matching processing unit 25 of the authentication server A45 performs a round-robin comparison of the dedicated features 69a included in the three sets of registration candidates, for example, and verifies that the degrees of difference are low (for example, below a predetermined threshold) in all the comparison results, and selects the dedicated feature 69a with the smallest total degree of difference. The matching processing unit 25 of the authentication server A45 also performs a similar comparison process on the common features 70 included in the three sets of registration candidates, and selects the common feature 70 with the smallest total degree of difference.
[0087] Note that the matching processing unit 25 of the authentication server A45 may select multiple dedicated features 69a and multiple common features 70 in the registration selection process (for example, selecting a predetermined number of dedicated features 69a and a predetermined number of common features 70 in ascending order of the total degree of difference, or selecting all features in the registration selection process in step S410). In this case, if matching is successful using any of the multiple dedicated features 69a or the multiple common features 70 in the authentication process, it may be determined that authentication is successful.
[0088] In addition, in the registration selection process of step S410, the matching processing unit 25 of the authentication server A45 may select a predetermined statistical quantity (e.g., an average) of the dedicated features 69a included in the three sets of registration candidates and a predetermined statistical quantity (e.g., an average) of the common features 70 included in the three sets of registration candidates.
[0089] The enrollment processing unit 20 of the authentication server A45 determines whether the enrollment data has been correctly determined (S411). That is, for example, if the degree of difference between all the enrollment candidates in the enrollment selection process (S410) performed in a round-robin fashion is equal to or less than a predetermined threshold, the enrollment data has been correctly determined. If the enrollment processing unit 20 of the authentication server A45 determines that the enrollment data has been correctly determined (S411: Y), it issues an enrollment ID for identifying the enrollment template, generates an enrollment template in which the dedicated feature 69a selected in the enrollment selection process is linked to the enrollment ID, and an enrollment template in which the common feature 70 is linked to the dedicated feature 69a, and registers the enrollment template including the dedicated feature 69a in the enrollment database A43 of the authentication system A41 (because the dedicated feature 69a is used for authentication within the authentication system A41) (S412).
[0090] Furthermore, the registration processing unit 20 of the authentication server A45 transmits the registration template including the common feature 70 to the other system (authentication server B46), and the registration processing unit 20 of the authentication server B46 registers the received registration template in the registration database B44 (S413), thereby completing the registration process. Since the common feature will be used in the other system B42, there is no need to register it in the registration database A43 of the authentication system A41.
[0091] If the registration processing unit 20 of the authentication server A45 determines that the registration data was not correctly determined (S411: N), it determines whether a predetermined timeout period has elapsed (S414). If the registration processing unit 20 of the authentication server A45 determines that the timeout period has not elapsed (S414: N), it returns to displaying the preview image and finger guide (S401), and if the registration processing unit 20 determines that the timeout period has elapsed (S414: Y), it determines that the registration has failed (S415) and ends the registration process.
[0092] Fig. 5 is a flowchart showing an example of authentication processing for achieving compatibility between authentication systems. Fig. 5 illustrates an example in which a user who has completed registration with authentication system A41 through the above-described registration processing performs 1:N authentication using authentication system B42 without registering with authentication system B42.
[0093] 1:N authentication is an authentication method in which a matching registrant is selected from multiple registered templates, or the user is determined to match none of the registered templates, without the user having to enter their own registered ID. Note that if only 1:N authentication is performed in the authentication process, a registered ID may not be issued in the above-described registration process, and the registered template may not include a registered ID. In the example of Figure 5, it is assumed that a user is shopping at a store that supports biometric authentication payments and pays by holding their finger over an infrared reflective authentication terminal 50 installed in a POS (Point of Sales) register.
[0094] First, the authentication terminal B48 displays an indicator lamp as a cue to hold a finger over the authentication terminal B48, and provides guidance on the display or audio to indicate that the finger should be held over (S501). When the user (person to be authenticated) holds a finger over the camera (image capture device 9) of the authentication terminal B48, the authentication terminal B48 controls the camera, light source, and other device-specific controls and captures an image of the finger in a manner similar to that of step S402 (S502). The attribute information extraction unit 61b extracts attribute information from the finger image in a manner similar to that of step S403 (S503). The attribute determination unit 23b performs attribute determination (S504) and determines whether the finger image is suitable for authentication (S505) in a manner similar to that of steps S404 and S405. If the attribute determination unit 23b determines that the finger image is not suitable for authentication (S505: N), the attribute determination unit 23b returns to displaying the indicator and guidance (S501) and retakes an image of the biometrics.
[0095] If the attribute determination unit 23b determines that the finger image is suitable for authentication (S505: Y), the dedicated feature extraction unit 64b extracts dedicated features from the finger image (S506), and the matching processing unit 25b matches all dedicated features stored in the registration database B44 of the authentication system B42, which is its own system, with the dedicated feature 70b extracted in step S506 (S507).
[0096] The authentication judgment unit 26 of the authentication server B46 performs an authentication judgment based on the matching score obtained in the matching process of step S507 to determine whether there is a registered template (in the registration database B44) that includes a dedicated feature that matches the dedicated feature 70b obtained in step S506 (whether there is a registered person that matches the person to be authenticated) (S508).
[0097] If the authentication determination unit 26 of the authentication server B46 determines that there is a registrant that matches the person to be authenticated (S508: Y), authentication of the registrant is successful (S513), and the authentication process ends. On the other hand, if the authentication in step S508 fails (S508: N), compatible authentication is then performed in steps S509 to S511. In other words, if at least the dedicated feature 70b of the person to be authenticated is not registered in the authentication system B42, compatible authentication is performed.
[0098] The common feature extraction unit 63 of the authentication server B46 extracts a common feature 69 from the finger image (S509). The matching processing unit 25b matches all common features registered in the registration database B44 of the authentication system B42, which is its own system (features that can also be used in the authentication system A41), with the common feature 69 extracted in step S509 (S510).
[0099] The authentication judgment unit 26 of the authentication server B46 performs an authentication judgment based on the matching score obtained in the matching process of step S510 to determine whether there is a registered template (in the registration database B44) that includes a common feature that matches the common feature 69 obtained in step S509 (whether there is a registered person that matches the person to be authenticated) (S511).
[0100] If the authentication judgment unit 26 of the authentication server B46 determines that there is a registered person who matches the person to be authenticated (S511: Y), the registration processing unit 20 of the authentication server B46 includes the dedicated feature 69b extracted in step S506 in the registration template of the person to be authenticated and registers it in the registration database B44 of the authentication system B42, which is its own system (S512), and the authentication of the registered person is successful (S513), terminating the authentication process.
[0101] If the authentication determination unit 26 of the authentication server B46 determines that the authentication in step S511 has failed (S511: N), the authentication processing unit 21 of the authentication server B46 determines whether a predetermined timeout period has elapsed (S514). If the authentication processing unit 21 of the authentication server B46 determines that the timeout has not occurred (S514: N), it returns to displaying the indicator and guidance (S501), and if it determines that the timeout has occurred (S514: Y), it determines that the authentication has failed (S515) and ends the authentication process.
[0102] In the above example, the common feature registered in authentication system A41 exists in the registration database of authentication system B42, so authentication is successful in the compatible authentication, and the dedicated feature of authentication system B42 is automatically registered.
[0103] In the example of Figure 5, the authentication system used for registration generates common features that can be cross-matched with other authentication systems, and the generated common features are registered in the other systems. Therefore, the authentication system not used for registration becomes able to perform authentication using the common features, and a user can be authenticated using one authentication system by simply registering using the other authentication system.
[0104] In addition, in the example of Fig. 5, if the authentication system fails to authenticate using dedicated features and succeeds in authenticating using common features, the dedicated features are automatically registered, and there is an extremely high possibility that authentication using dedicated features will be successful in the next authentication process or thereafter. Since authentication using dedicated features has higher authentication accuracy than authentication using curved features, highly accurate authentication can be achieved.
[0105] Furthermore, in the example of Figure 5, the authentication system performs authentication within its own authentication system using dedicated features before compatible authentication using common features. This is because there is a high possibility that a registered template exists within the same authentication system (the same authentication system is used for registration and authentication), and there is a high probability that the authentication time can be shortened by performing authentication using dedicated features first.
[0106] 5, 1:N authentication is performed, but 1:1 authentication may also be performed in which the person to be authenticated inputs a registration ID to uniquely determine a registrant and then authenticate that the person is the registrant. When performing 1:1 authentication, the authentication system may determine whether to perform compatible authentication depending on whether a dedicated feature corresponding to the input registration ID is registered in the registration database of the authentication system. If the authentication system determines that a dedicated feature corresponding to the input registration ID is not registered in the lantern k database of the authentication system, it may omit the matching using the dedicated feature in steps S506 to S508.
[0107] The learning and inference methods of the attribute information extraction unit, common feature extraction unit, and dedicated feature extraction unit will be described below.
[0108] Fig. 6A is a diagram showing an example of a configuration of a learning model during learning of an attribute information extraction unit and a common feature extraction unit, Fig. 6B is a flowchart showing an example of a learning process for the attribute information extraction unit and the common feature extraction unit.
[0109] 6A, the learning data 121a and the learning data 121b are moving images of a living body captured by the authentication terminal A47 and the authentication terminal B48, respectively. The learning data 121a and the learning data 121b are data sets to which the type of device (the type of the imaging device 9 included in the authentication terminal) and an ID that distinguishes the subject are assigned, and the terminal control value 67 at the time of capturing is labeled for each frame of the moving image. The moving image is a series of images of fingers captured when expected operations are performed on each authentication terminal.
[0110] It is also assumed that a "suitable" label is assigned in advance to each frame of a video for which the designer has determined that the finger posture is ideal and the shooting conditions are good. The biometric information processing unit 27 includes a distance learning unit 122 and a loss calculation unit 123 for learning. In the example of FIG. 6A, the distance learning unit 122 and the loss calculation unit 123 are included in the biometric information processing units 27 of both the authentication server A45 and the authentication server B46, but for the sake of simplicity, only one distance learning unit 122 and one loss calculation unit 123 are depicted in FIG. 6A.
[0111] Furthermore, the distance learning unit 122 and the loss calculation unit 123 may be included in the biometric information processing unit 27 of either the authentication server A45 or the authentication server B46. Furthermore, a device independent of the authentication system A41 and the authentication system B42 may include the distance learning unit 122 and the loss calculation unit 123. The attribute information extraction unit, attribute determination unit, common feature extraction unit, and matching processing unit of the authentication server A45 and the authentication server B46 are optimally trained by the processing by the distance learning unit 122 and the loss calculation unit 123.
[0112] In the process of FIG. 6B, first, the attribute determination unit 23, the attribute information extraction unit 61a, and the attribute information extraction unit 61b of the authentication server A45 and the authentication server B46 perform pre-learning (S601) to improve the accuracy of attribute determination in advance.
[0113] The accuracy of feature extraction is improved in advance by executing pre-learning in the common feature extraction unit 63 of the authentication servers A45 and B46, the matching processing unit 25 of the authentication servers A45 and B46, and the distance learning unit 122 (S602). Learning is then executed in the attribute information extraction unit 61, attribute determination unit 62a, and attribute determination unit 62b of the authentication servers A45 and B46 using the learning results of the common feature extraction unit 63 of the authentication servers A45 and B46, the matching processing unit 25 of the authentication servers A45 and B46, and the distance learning unit 122, whose accuracy has been improved in advance (S603).
[0114] Using the learning results of the attribute information extraction unit 61a, the attribute information extraction unit 61b, and the attribute determination unit 62b of authentication server A45 and authentication server B46, learning is performed in the common feature extraction unit 63 of authentication server A45 and authentication server B46, the matching processing unit 25 of authentication server A45 and authentication server B46, and the distance learning unit 122 (S604).
[0115] The distance learning unit 122 determines whether the learning has converged by checking whether the classification accuracy has stably converged as a learning result (S605), and if it determines that the learning has not converged (S605: N), the process returns to step S603 again, and learning is repeatedly and alternately executed. If the distance learning unit 122 determines that the learning result has converged (S605: Y), the learning process ends.
[0116] First, the pre-learning (S601) of the attribute information extraction unit 61a, the attribute information extraction unit 61b, the attribute determination unit 62a, and the attribute determination unit 62b will be described in detail. In the pre-learning of the attribute information extraction unit 61a, the attribute information extraction unit 61b, the attribute determination unit 62a, and the attribute determination unit 62b, learning data 121 that has been labeled in advance with a subject ID and "suitable" is input, and learning of a binary classification problem is executed in which a result of "suitable" is output as accurately as possible for input images labeled with "suitable" and a result of "unsuitable" for other images.
[0117] At this time, the attribute information extraction unit 61a performs learning using learning data 121a including finger images captured by an authentication terminal A47, which is a visible light type authentication terminal such as a smartphone, and the attribute information extraction unit 61b performs learning using learning data 121b including finger images captured by an authentication terminal B48, which is an infrared reflection type authentication terminal.
[0118] As an example of a specific network configuration, the attribute information extraction unit 61a and the attribute information extraction unit 61b use ResNet, and can be configured to output a 512-dimensional feature vector as attribute information. The output layer of the attribute determination unit 62 can be configured to acquire output values of two nodes with softmax as an activation function as a one-hot vector, for example. This makes it possible to determine that if the first output node is activated, it is "suitable," and if the second output node is activated, it is "unsuitable." In this way, in pre-learning, supervised learning is performed using training data that has been assigned a "suitable" label by the designer in advance.
[0119] The pre-learning (S602) of the common feature extraction unit 63 will be described. For each frame of biometric images labeled as "suitable" in the training data, training is performed so that all subject IDs of images from the two devices used can be correctly classified. The training data includes training data A121 captured by authentication terminal A47 and training data B121 captured by authentication terminal B48, and the same subject ID is assigned to each image. Training proceeds by randomly selecting biometric images from training data 121a and training data 121b, and inputting one or more images that have been tensorized. As an example of the configuration of the common feature extraction unit 63, ResNet can be used as the baseline, and a 1024-dimensional feature vector can be output as the common feature 69.
[0120] The metric learning unit 122 is a learning model that optimizes the distance between features, such as the Euclidean distance (L2 distance) or cosine dissimilarity (cosine similarity), of common features (feature vectors) obtained through metric learning. Distance optimization means learning to obtain features such that the distance between features obtained from biometric images of the same subject is as close as possible, and the distance between features obtained from different subjects is as far as possible. Distance learning methods that can be used include Siamese Network, Triplet Network, ArcFace, and AdaCos.
[0121] In this embodiment, it is assumed that each of the authentication terminals A47 and B48 captures images of learning data for 1000 people whose subject IDs range from 001 to 999, and that AdaCos is adopted as the metric learning unit 122. In this case, the metric learning unit 122 is configured to receive the common feature 69 as input and output a 1000-dimensional one-hot vector, which is the total number of classes.
[0122] Furthermore, a loss function specific to AdaCos, which is a modified softmax loss, is used as the loss function for optimizing the learning, and the loss calculation unit 123 calculates the specific value of this loss. Then, as the learning of the common feature extraction unit 63 and the distance learning unit 122 proceeds to minimize this loss, the learning progresses so that the correct labels of the learning data and the inferred labels match more frequently, and the distance as the cosine dissimilarity between features is optimized. As a result, features that can classify 1,000 classes as accurately as possible can be acquired.
[0123] The advantage of introducing distance learning is that by solving a class classification problem, which is relatively easy to learn, learning proceeds so that the distance between the feature vectors of different subject IDs increases. Therefore, when an unknown biometric image is input, the distance becomes large relative to any registered data, which has the advantage of preventing the mistaken authentication of unregistered input. In general classification problems, input of an unknown class will be classified into one of the classes, so in biometric authentication, where it is essential to prevent input by unauthorized individuals, learning by simply solving the classification problem is insufficient, and distance learning is important.
[0124] The reason why common features can be obtained by the above-mentioned learning will be explained below. The learning data 121a and the learning data 121b include finger images captured by an authentication terminal A47, which is a visible light type authentication terminal such as a smartphone, and finger images captured by an authentication terminal B48, which is an infrared reflective type authentication terminal.
[0125] The fact that biometric images taken by any authentication device can be classified into the correct subject ID indicates that the biometric images of a subject with the same ID are composed of feature vectors that are highly similar (close to each other) regardless of which authentication device they are taken by. In other words, common features can be acquired from biometric images taken by different authentication devices.
[0126] Conversely, for biometrics of subjects with different IDs, the distance between feature vectors is learned to increase, so the degree of dissimilarity between the biometrics of subjects with different IDs increases, and outsiders will be correctly rejected regardless of which authentication terminal is used. Note that, since AdaCos is used in this embodiment, cosine similarity is used as the distance between vectors when comparing common features, but if other distance learning methods are used, distances such as L1 distance or L2 distance may also be used.
[0127] In addition, when the use of AdaCos is assumed as in this embodiment, AdaCos only needs to perform calculations using a formula to calculate cosine similarity during matching, so the matching processing unit 25 does not need to be a learning model, and learning of the matching processing unit 25 does not necessarily need to be performed.
[0128] Next, a description will be given of the process (S603, S604) of alternately and repeatedly learning and optimizing the attribute information extraction unit 61a, the attribute information extraction unit 61b, the attribute determination unit 62a, the attribute determination unit 62b, the common feature extraction unit 63 of authentication servers A45 and B46, the matching processing unit 25 of authentication servers A45 and B46, and the distance learning unit 122. The optimization procedure involves first learning the attribute information extraction unit 61a, the attribute information extraction unit 61b, the attribute determination unit 62a, and the attribute determination unit 62b, and then stopping the learning of the common feature extraction unit 63 and the matching processing unit 25 of authentication servers A45 and B46 to perform only inference in the current learning state.
[0129] The learning (S603) of attribute information extraction unit 61a, attribute information extraction unit 61b, attribute determination unit 62a, and attribute determination unit 62b will be described. First, a finger image captured by authentication terminal A47, which is a visible light authentication terminal such as a smartphone, is input to attribute information extraction unit 61a, and a finger image captured by authentication terminal B48, which is an infrared reflective authentication device, is input to attribute information extraction unit 61b, which then infers whether the image is "suitable" for use in authentication based on the current learning state.
[0130] In the above-mentioned pre-learning, supervised learning was performed by referring to the "suitable" label previously assigned to the dataset and proceeding with learning so as to match the correct answer, but in the learning (S603) of the attribute information extraction unit 61a, the attribute information extraction unit 61b, the attribute determination unit 62a, and the attribute determination unit 62b, unsupervised learning is performed to determine whether or not the data is "suitable" so as to maximize the authentication accuracy described below, without referring to the label.
[0131] If the unsupervised learning determines that the result is "suitable," subsequent feature extraction and matching are performed, and the actual error is obtained from the distribution of matching scores to calculate a loss function. Learning then proceeds in the attribute information extraction unit 61a, the attribute information extraction unit 61b, the attribute determination unit 62a, and the attribute determination unit 62b so as to minimize this loss function.
[0132] However, the attribute information extraction units 61a and 61b are required not only to improve the authentication accuracy of the matching processing unit 25 but also to determine as many inputs as possible as "matching." Therefore, in order to optimally train the attribute information extraction units 61a, 61b, attribute determination units 62a, and 62b, the following equation is defined as a loss function in this embodiment.
[0133] Loss function = α × classification error loss + β × mismatch loss per trial
[0134] Here, the first term of the loss function is the classification error loss, and for example, softmax loss or cross entropy loss is applied. The second term of the loss function is the proportion of all consecutive images in one trial that are rejected as non-matching. This loss function minimizes the classification error loss caused by the matching process in the subsequent stage, and also acts to reduce the number of "non-matching" judgments as much as possible in the match judgment. In other words, the loss function maximizes authentication accuracy while allowing for as much posture variation as possible. Allowing posture variation improves convenience by allowing users to be authenticated by simply holding their biometric device roughly over the sensor.
[0135] For example, if all input videos are judged to be "matching," the matching process will be performed even when no hands are visible or only a small part of the hand is visible, which will increase the matching error rate and result in a large loss. Conversely, if all input images are judged to be "mismatching," authentication will not fail at all, so the error rate will decrease, but the mismatch rate will be extremely high, so the loss will still be large in this case.
[0136] On the other hand, if images are judged as "mismatched" at an appropriate frequency, the system will train to carefully select images so that the mismatch loss is kept constant while minimizing the classification error loss. As the classification error rate decreases to a certain value, the mismatch loss gradually becomes larger, and the system will then train to reduce the frequency of mismatches. By repeating this process, the system will train to minimize classification errors while producing as many "matches" as possible, allowing for a certain degree of pose variation. This ultimately makes it possible to achieve both high authentication accuracy and ease of use. Note that α and β are weighting parameters for the classification error loss and the mismatch loss. Increasing α relative to β places more weight on the classification results, resulting in a "match" being judged only when the biometric image quality is high.
[0137] The learning (S604) of the common feature extraction unit 63 of authentication server A45 and authentication server B46, the matching processing unit 25 of authentication server A45 and authentication server B46, and the distance learning unit 122 will be described. First, the attribute information extraction unit 61a, the attribute information extraction unit 61b, the attribute determination unit 62a, and the attribute determination unit 62b, which have been trained in the processing of step S603, only execute inference, and in this state, only the common feature extraction unit 63 of authentication server A45 and authentication server B46, the matching processing unit 25 of authentication server A45 and authentication server B46, and the distance learning unit 122 are trained.
[0138] Pre-learning is performed in step S602, but further optimization is achieved by performing re-learning on inputs that the current attribute information extraction unit 61a, attribute information extraction unit 61b, attribute determination unit 62a, and attribute determination unit 62b have determined to be "suitable."
[0139] For example, a biometric image in which the hand is captured in a small area or in which the hand is held in a different position than expected may be judged as "matching" after learning, even though it was judged as "unsuitable" immediately after pre-learning. Therefore, biometric features are constantly re-learned to maximize authentication accuracy based on the most recent judgment results, and overall optimization is promoted. The basic learning method is the same as the pre-learning described above, so a detailed explanation is omitted. As a result, the features are customized to match the output results of the attribute information extraction unit 61a and the attribute information extraction unit 61b.
[0140] Subsequently, a phase of training only the attribute information extraction unit 61a, the attribute information extraction unit 61b, the attribute determination unit 62a, and the attribute determination unit 62b, and a phase of training only the common feature extraction unit 63 and the matching processing unit 25 of authentication server A45 and authentication server B46 are repeatedly carried out.
[0141] By alternately learning in this manner, the attribute information extraction units 61a and 61b reject input images that the subsequent common feature extraction unit 63 finds difficult to recognize as "unsuitable," while the common feature extraction unit 63 generates optimized features for images sent by the previous attribute information extraction unit as "suitable." Furthermore, when biometric images of the same user are input to both authentication servers A45 and B46, the common feature extraction units 63 of authentication servers A45 and B46 are trained so that the distance between the common features output by the respective common feature extraction units 63 is close (e.g., so that the distance between these common features is within a predetermined value). In other words, by training each learning model in a complementary manner, authentication accuracy can be optimized. Furthermore, by using such a learning method, it is not necessary to pre-evaluate, for example, the threshold value for the hand size in the image to be passed to the matching processing unit 25, and authentication accuracy can be optimized, thereby achieving high accuracy and reducing development costs.
[0142] 7 is an explanatory diagram showing an example of a conformance determination by the attribute information extraction unit 61a and the attribute determination unit 62a. In FIG. 7, an example is shown in which the conformance determination is performed by the attribute information extraction unit 61a and the attribute determination unit 62b on an image of a finger held over a rear camera (an example of the imaging device 9) of the smartphone 141 (an example of the authentication terminal A47).
[0143] On screen 142 of smartphone 141, the image captured by the rear camera is superimposed with finger guide 143, and the user can check how their fingers are being held while looking at screen 142. (A) shows a state where no fingers are being held over the rear camera. (B) shows a state where fingers 1 held over the rear camera are captured small. (C) shows a state where fingers 1 held over the rear camera are captured at a medium size. (D) shows a state where fingers 1 held over the rear camera are captured at approximately the same size as finger guide 143.
[0144] The attribute information extraction unit 61a receives the captured image and various device parameters as input and outputs attribute information. As described above, the attribute information represents attributes as a feature vector calculated based on the input image features and device control values. This feature vector is interpreted by the attribute determination unit 62a at a subsequent stage.
[0145] As described above, the attribute information extraction unit 61a and the attribute determination unit 62a are trained to improve authentication accuracy and increase the frequency of proceeding to the matching process as much as possible. If no fingers are visible, as in Figure 7A, authentication will clearly not be successful, and if this image is passed on to the matching process at a later stage, authentication accuracy will decrease. Therefore, the attribute information extraction unit 61a and the attribute determination unit 62a are trained to determine this image as "unsuitable."
[0146] Furthermore, when the fingers are generally small, there is a possibility that the accuracy of authentication may be degraded due to insufficient resolution. Therefore, even when the image of the fingers is insufficient as in (B), the attribute information extraction unit 61a and the attribute determination unit 62a are trained so that if there is a tendency for the loss (recognition incorrect rate) calculated during learning to increase, it will be determined to be "unsuitable."
[0147] Similarly, in (C), the hand photographed is thought to be slightly small relative to the finger guide 143, but if it is determined that this will not affect the authentication accuracy when training the attribute information extraction unit 61a and the attribute determination unit 62a, the attribute information extraction unit 61a and the attribute determination unit 62a can proceed to the subsequent matching process without any problems, and so the attribute information extraction unit 61a and the attribute determination unit 62a are trained to determine this as "suitable." Also, in (D), the photographing conditions are the best, and if it is determined that this will not actually affect the authentication accuracy during training, the hand will also be determined as "suitable."
[0148] In conventional technology, whether a hand is suitable or unsuitable is often determined based on an empirically determined hand size. It is particularly difficult to determine whether a hand is suitable or unsuitable for intermediate sizes like (C). Therefore, it is necessary to separately calculate the relationship between hand size and authentication accuracy, determine the hand size at which accuracy deteriorates, and determine the threshold for suitability or unsuitability. In contrast, this embodiment extracts image features and performs optimal determinations that increase matching frequency and improve authentication accuracy. Therefore, this embodiment extracts all image features, such as hand size, hand posture, and background lighting conditions, and performs flexible determinations to maximize authentication accuracy, thereby improving authentication accuracy and improving development efficiency.
[0149] Similarly, optimal matching determination is performed even when the finger extends beyond the screen 142, as in (E). Conventional technologies have often been designed to unconditionally output a non-matching determination based on experience when the finger extends beyond the screen 142, but in reality, it should be possible to determine a match if authentication is successfully performed based on the partially captured features. Therefore, by using a method for performing a matching determination from the perspective of maximizing authentication accuracy, as in this embodiment, it becomes possible to proceed to the matching process even in such cases, and the possibility of authentication being successful even when the user cannot properly place their finger over the screen is increased, which is expected to improve convenience.
[0150] Similarly, when the image is dark, as in (F), the match determination is performed to maximize authentication accuracy. For example, if this dark image is due to insufficient light from the light source 3, simply forcibly increasing the brightness by adjusting the contrast will emphasize the noise in the image, which will likely result in a deterioration in authentication accuracy. Therefore, in this case, the image should be determined as "non-matching." On the other hand, if the camera gain was simply smaller than necessary and the light from the light source 3 was sufficient, adjusting the contrast to brighten the image may not emphasize the noise, so in this case the image should be determined as "matching."
[0151] In conventional technology, a match or non-match was determined simply by using the average brightness of the image, but in this embodiment, various features on the image are analyzed in detail, and a match or non-match is determined based on whether the biometric features can be correctly extracted. As such, as in the above example, a "match" can be determined within a range that does not degrade accuracy, thereby improving authentication accuracy while increasing convenience.
[0152] In this way, in the past, compatibility or non-compatibility was determined empirically, and hand size judgment and overhang judgment were individually designed, but in this embodiment, it is possible to automatically make optimal judgments, which not only improves authentication accuracy and convenience but also reduces design costs.
[0153] An example of processing for automatically classifying the determination results of the attribute information extraction unit 61a and the attribute determination unit 62a will be described with reference to Figures 8A and 8B. Figure 8A is a diagram showing an example of a pair 161 of a training image and a non-conformance reason label as training data. Figure 8B is a diagram showing an example of a feature vector space onto which attribute information acquired from the training image is mapped.
[0154] 7A and 7B, the results of the captured images being determined to be "unsuitable" were due to "no hand being held over the subject" or "small hand," respectively. As described above, the learning of the attribute information extraction unit 61a, the attribute information extraction unit 61b, the attribute determination unit 62a, and the attribute determination unit 62b was performed to increase the matching frequency and improve authentication accuracy, and labels indicating which images should be determined to be suitable in advance were used in the pre-learning but were not used in the repeated learning. In other words, the determination results were optimized by unsupervised learning.
[0155] On the other hand, if the user knows the reason why the image was unsuitable, he or she can take steps to improve it, which improves convenience. Therefore, in this embodiment, typical reasons for unsuitability may be assigned as labels to the training images in advance, and the reason for unsuitability may be determined to be closest to the reason indicated by the label. An example of this determination method is shown below.
[0156] First, as shown in Fig. 8A, a pair 161 of a training image and a non-matching reason label is defined as training data. The pair 161 in Fig. 8A shows an example in which the images captured in the states shown in Fig. 7 ((A), (B), (E), and (F)) are given labels (manually classified into four classes) of "undetected," "small finger," "protruding," and "dark environment."
[0157] In addition, there are various other possible reasons why a finger may be judged as non-compliant, such as "the finger is fake," "the finger is bent," or "the hand is moving." However, if every possible condition is to be classified, the necessary evaluations would be required, making preparation difficult. In addition, it is possible that the situation may become so complex that it is impossible to categorize it manually, such as when multiple factors that result in non-compliance occur simultaneously. Therefore, it is desirable to prepare training data only for reasons that are clearly understood to be non-compliant and / or for reasons that are given high priority and should be fed back to the user as feedback.
[0158] As shown in Figures 6A and 6B, the attribute information acquired by the attribute information extraction unit 61a is, for example, a 512-dimensional feature vector, and the attribute determination unit 62a uses this feature to determine whether the captured image is a suitable image.
[0159] As shown in Figure 8B, when the attribute information acquired from the training images is mapped onto the space of feature vectors and observed along two feature axes, the feature vectors form several clusters, where feature vectors with the same label are distributed nearby.
[0160] 8B indicates the center of the cluster indicated by each label. The attribute determining unit 62a not only performs a binary determination of "suitable" or "unsuitable" as described above, but also performs unsupervised learning based on unsupervised learning techniques such as k-means and SOM (Self Organizing Maps) using the obtained attribute information vector as input, thereby learning to which cluster the input attribute information belongs.
[0161] When the attribute determination unit 62a uses the k-means method or the like, it can obtain not only the cluster center positions for each label but also several "unknown" clusters to which no labels are assigned. In the example of Figure 8B, two "unsuitable" results, "Unknown 1" and "Unknown 2," which cannot be determined in advance, are obtained.
[0162] Furthermore, in FIG. 8B, attribute information determined to be "suitable" is plotted in the upper right corner, but the "protruding" labels determined to be "unsuitable" are distributed relatively close together, suggesting that even among the "suitable" results, there are a few protruding images. Since it is assumed that multiple labels are assigned to a single image, it is desirable for the attribute determination unit 62a to not only determine whether an image is suitable or unsuitable, but also learn to output the probability of belonging to each cluster. Specifically, for example, a commonly known neural network such as ResNet can be used as the baseline, with output nodes whose number is equal to the number of clusters added to the two dimensions of suitability and unsuitability, and whose activation function for the output layer is an identity function.
[0163] If the attribute determination unit 62a determines that the attribute determination result for whether to send the image to the matching process is "incompatible," it provides the label of the cluster that is closest to the attribute information to the user, allowing the user to understand the reason for the failed photo shoot.
[0164] In addition, if the attribute determination unit 62a determines that the attribute determination result is "Unknown 1" or "Unknown 2" as shown in Figure 8B, it cannot explain the reason for the non-compliance to the user, so it can display an error message such as "Unexpected shooting conditions" or "Unknown environment detected" on the screen of the authentication terminal A47 to at least make it clear that there is a reason why authentication cannot be proceeded with.
[0165] If authentication still cannot proceed (cannot proceed to processing of step S505) after a certain period of time has elapsed, the attribute determination unit 62a may, for example, display a message on the screen of the authentication terminal A47 saying "Please enter your PIN as an alternative" and prompt the user to enter a PIN (Personal Identification Number) instead of performing biometric authentication, or may display a message on the screen of the authentication terminal A47 saying "Please notify an attendant" to prompt the user to notify an attendant at the site where the authentication system A41 is used.
[0166] It should be noted that for the unknown clusters "Unknown 1" and "Unknown 2" mentioned above, the designer can check the images classified into these clusters and later assign a label to the cluster to which the image belongs. For example, if the designer checks the images classified into these clusters and finds that "Unknown 1" has been judged to be non-compliant for reasons such as "high levels of image block noise" and "Unknown 2" is "backlit," then by assigning such labels to the "Unknown 1" and "Unknown 2" clusters and providing feedback, it will be possible to categorize unknown accuracy degradation conditions that the system designer did not initially anticipate in a way that users can understand later.
[0167] In this way, the attribute information extraction unit 61a, the attribute information extraction unit 61b, the attribute determination unit 62a, and the attribute determination unit 62b can automatically detect images that are unsuitable for authentication, and if the category name is known, it can be shown to the user to encourage improvement, thereby achieving highly accurate and convenient authentication and reducing development costs because it does not require the work of finely categorizing a large number of training images.
[0168] FIG. 9A shows an example of the configuration of a feature extraction unit 24 that autonomously extracts arbitrary biometric features that maximize authentication accuracy. In conventional technology, designers determine the biometric parts to be used for biometric authentication, and biometric feature extraction and matching technologies are also designed manually. However, faithful extraction of anatomical biometric parts does not necessarily maximize authentication accuracy. Rather, faithful extraction of anatomical biometric parts often results in insufficient accuracy from the perspective of maximizing authentication accuracy, such as when potentially useful biometric features are deemed noise and removed, or when patterns common to both the subject and others are extracted. Furthermore, trial and error often increases development costs.
[0169] Furthermore, with conventional technology, image segmentation, which separates the biological parts of interest from other unnecessary areas, required manual technical design, and in order to perform segmentation using machine learning, training data had to be created that paired the original image with the segmentation results, resulting in significant development costs.
[0170] Therefore, feature extraction technology based on unsupervised learning is effective in autonomously extracting the biometric features necessary to maximize authentication accuracy. In other words, when an input image and subject ID are given, it becomes possible to realize end-to-end biometric authentication processing learning that automatically extracts features that maximize authentication accuracy. This eliminates the need for preprocessing such as segmentation and biometric posture correction, which have previously been explicitly performed, or the need to separately design specialized feature extraction processing to extract biometric features used for authentication. Furthermore, it enables the realization of biometric authentication technology with higher accuracy and lower design costs than manual design.
[0171] 9A, the feature extraction unit 24 may be, for example, a multi-layer deep neural network. The feature extraction unit 24 may further include, for example, a landmark detection layer 181, a landmark normalization layer 183, a segmentation layer 185, and a feature selection layer 187.
[0172] The landmark detection layer 181 detects landmarks included in the biometric image 66. The landmark normalization layer 183 transforms and normalizes the image topology of the acquired landmark-information-added biometric image 182 based on the positions of the landmarks. The segmentation layer 185 segments biometric features from the transformed and corrected biometric image 184. The feature selection layer 187 selects and extracts biometric features from the obtained segmentation image 186.
[0173] The matching processor 25 then matches the obtained autonomously extracted feature 188 with the registered data. This autonomously extracted feature 188 corresponds to, for example, the dedicated feature 70a or the common feature 69 described above, and can be applied to generating a feature specific to a device as well as generating a common feature used in compatibility authentication.
[0174] An overview of the processing procedure when each of the above layers has been optimally trained will be described. First, the landmark detection layer 181 detects landmarks attributable to the living body from the input biometric image. For example, if the biometric image is an image of a hand, the tip, base, first joint, and second joint of each finger are automatically selected and detected as landmarks.
[0175] The landmark normalization layer 183 performs normalization such as scaling and deformation correction to standardize the position and size of the hand in the biometric image. The segmentation layer 185 acquires the area to be extracted as a biometric feature. The feature selection layer 187 extracts customized biometric features to maximize authentication accuracy.
[0176] Next, an example of the learning procedure of the feature extraction unit 24 in Fig. 9B will be described. It is assumed that learning data in which a subject ID is assigned to a biometric image captured by an arbitrary authentication terminal is input.
[0177] First, an example of learning in the landmark detection layer 181 will be described. First, partial features of images that are commonly observed in multiple biometric images of the same subject ID and multiple biometric images of different subject IDs are searched for. For example, a common feature, such as a semicircular curved outline like a fingertip, is found for each pixel.
[0178] For example, using SIFT (Scale-Invariant Feature Transform) features, corresponding points similar to this common feature can be found. These corresponding points are regarded as training data for provisional landmarks, and a fully convolutional neural network is adopted as the network configuration of the landmark detection layer 181, for example, and is trained to detect landmark parts pixel by pixel from the input image.
[0179] At this time, by minimizing the loss function using the cross-entropy error for the number of incorrectly detected pixels, the accuracy rate of accurately identifying landmarks improves and the feature amounts on the image for determining landmarks are internally learned. By repeatedly performing this process on various training images, eventually, when a biometric image 66 is input to the landmark detection layer 181, the landmark part will be automatically output.
[0180] In individual training images, for example, edges in the background may be accidentally judged as landmarks, leading to erroneous learning. However, if the part accidentally erroneously detected as a landmark is not consistently detected in any training image, it will gradually stop being detected. Therefore, if a large amount of training data is available, landmarks can be correctly detected based on features that can always be detected stably. Note that landmarks do not necessarily have to be represented by points, and may include edges, etc.
[0181] We will now describe an example of learning of the landmark normalization layer 183. When a biometric image 182 with landmark information is input to the landmark normalization layer 183, the layer outputs a biometric image in a state where the image has been deformed and corrected so that the position of each landmark overlaps with the average landmark position.
[0182] The average landmark position is obtained by learning to deform the image so that the average position of each landmark in all training data is given as the correct answer and the image approaches this correct answer. For example, by using GAN to determine in advance the distribution of latent variables when a biometric image is deformed to the average landmark position, and replacing the latent variables of the input biometric image with those of the biometric image with the average landmark position, a deformed and corrected biometric image 184 can be obtained in which the landmark positions have been corrected to the average positions.
[0183] At this time, inputs including erroneously detected landmarks may be passed to the landmark normalization layer 183, but by performing learning using a large amount of data, normalization can be performed without being influenced by exceptionally obtained landmarks. Note that the landmark normalization layer 183 may perform transformations using a formularizable mathematical model such as affine transformations, rather than machine learning.
[0184] Next, an example of learning in the segmentation layer 185 will be described. First, for input biometric images of the same subject ID, landmark extraction and normalization are performed in the preceding stages of the segmentation layer 185, and then the feature selection layer 187 extracts autonomously extracted features 188 from the image. Furthermore, for each pixel of the extracted autonomously extracted features 188 across the entire image, the similarity to the partial features around the pixel is measured, and the pixel is binarized into areas of high and low similarity. The binarized result is determined as provisional training data for segmentation. The partial features around the pixel can be, for example, autonomously extracted features of 11 pixels above, below, left, and right of the pixel of interest. For biometric images of the same subject, it is expected that the similarity of the biometric parts will be high and the similarity of the background parts will be low. Therefore, by obtaining pixels with high similarity, segmentation results for the biometric parts can be obtained.
[0185] The segmentation layer 185 is then trained to output this segmentation result from the input biometric image. By repeatedly performing this learning on various training data, only the parts that contribute to authentication are segmented. An image showing the segmentation result output by the segmentation layer 185 is called a segmentation image 186. Specifically, the segmentation image 186 is held as a black and white image in which, for example, pixels that contribute to authentication are set to 255 and other pixels are set to 0. Note that MobileNet-v3 or the like can be used as a specific network configuration of the learning model for the segmentation layer 185.
[0186] We will now describe an example of learning of the feature selection layer 187. After all of the processing before the feature selection layer 187 has been performed, the deformation-corrected biometric image 184 and the segmentation image 186 are input, and learning of the feature selection layer 187 is advanced so as to increase the accuracy rate of subject ID classification, similar to the learning of the dedicated feature extraction units 64a and 64b described above.
[0187] Once the learning has converged to a certain extent, learning of the feature selection layer 187 stops, and then learning is performed only on the segmentation layer 185. Once learning of the segmentation layer 185 has converged, learning is again performed only on the feature selection layer 187. By repeating this process, each layer is optimized, and when learning of all layers has converged, learning of all layers is completed.
[0188] The autonomously extracted features 188 extracted by the feature selection layer 187 may include the segmentation image 186 itself. Alternatively, the autonomously extracted features 188 may be abstract features of any dimension obtained from a biometric image and a segmentation image. Authentication accuracy is likely to improve when the autonomously extracted features 188 are extracted as an image plane having a two-dimensional spatial topology structure. For example, multiple different biometric features may be extracted for different image planes.
[0189] For example, when a finger image is used as the biometric image, finger veins can be extracted separately as plane 1 and knuckle prints as plane 2. The matching processor 25 at the subsequent stage can then perform matching independently of each other. By extracting multiple separate features in this way, for example, plane 1 can be used as an alignment feature, and plane 2 can be matched with the alignment result to calculate a matching score, and similarly plane 2 can be used as an alignment feature, and plane 1 can be matched with the alignment result to calculate a matching score, and the calculated scores can be finally combined. By performing such processing, it is possible to improve authentication accuracy.
[0190] It is also possible to improve the matching performance based on encryption by encrypting one plane and unencrypting the other. Of course, the multiple biometric features may be from different anatomical parts, or may be a mixture of multiple parts. It is also possible to extract multiple biometric features with low similarity, that is, to extract independent components as feature quantities. This increases accuracy because the features function complementary to each other, and also increases confidentiality because even if one biometric feature is leaked, the other biometric feature becomes more difficult to guess.
[0191] With the above method, by simply providing the input image and subject ID, each layer is autonomously optimized, and features that maximize authentication accuracy can be autonomously extracted. Note that the segmentation layer 185 may be omitted depending on the biometric capture conditions, which vary depending on the authentication terminal. For example, if an authentication terminal is used that captures only one biometric image and does not capture unnecessary background other than the user's biometrics, there is no need to perform segmentation optimization, and omitting the segmentation layer 185 can improve calculation speed.
[0192] 9B is a diagram showing an example of autonomous feature extraction processing when a photograph of fingers 1 is taken by a visible light authentication terminal such as a smartphone. Fingers 1 are photographed in biometric image 66, with five fingers and the upper part of the palm included in the field of view. An example of authentication when the above-mentioned autonomous feature extraction is performed on this biometric image 66 will be described.
[0193] First, landmarks 189 are obtained by the landmark detection layer 181. In the biometric image 66, landmarks 189 are detected at the fingertips, joints, and sides of the hand. In FIG. 9B, the landmarks 189 are indicated by small circles. In the example of FIG. 9B, erroneously detected landmarks 189 are also detected in the background outside the fingers 1.
[0194] The fingers in the biometric image 66 are deformed and corrected to a standard shape by the landmark normalization layer 183. In the deformation-corrected biometric image 184 in FIG. 9B, the size and angle of the hand have already been changed.
[0195] The segmentation layer 185 performs segmentation to determine regions effective for authentication on the deformation-corrected biometric image 184. In the example of FIG. 9B, the index finger region 190, middle finger region 191, ring finger region 192, and upper palm region 193 are selected as regions effective for authentication. On the other hand, the thumb and little finger are not selected. In this embodiment, only regions effective for authentication are extracted, so in the example of FIG. 9B, the little finger and thumb do not contribute to improving authentication accuracy. Since regions that do not contribute to improving authentication accuracy are not extracted, unnecessary processing time can be reduced and the data size of the enrollment template can also be suppressed.
[0196] The feature selection layer 187 extracts features from the segmented regions. In the example of FIG. 9B, a vein component 194 and a knuckle print component 195 are extracted from the fingers, and a palm print component 196 is extracted from the palm. When the obtained autonomously extracted features 188 are input to the matching processor 25, the matching processor 25 performs a matching process to determine the similarity between feature vectors that primarily include features from the valid region when matching with a registered template. As a result, the matching processor 25 can use only the parts that are valid for authentication for matching, thereby maximizing authentication accuracy.
[0197] Note that the segmentation image 186 may be displayed on the authentication terminal so that the designer can check which parts of the biometrics are being used by the current authentication system. When biometric features autonomously extracted through learning are used in actual operation, it is important for the user to be able to recognize which parts of the biometrics are being used for authentication so that the user can use the biometric authentication system with confidence. By checking the segmentation image, the designer can explain to the user which parts of the biometrics are being used.
[0198] Furthermore, depending on the environment in which the authentication terminal is used, there may be cases in which there are biometric parts that cannot be used, biometric parts that do not want to be used, or biometric parts that are forcibly used. In this case, the designer selects and masks the partial areas that do not want to be used or the partial areas that do want to be used in the segmentation image 186 obtained during the learning process. By removing or adding the masked areas and then re-learning the segmentation layer 185, it becomes possible to generate feature quantities for any part. As a result, in this embodiment, it is possible to realize customizable biometric authentication that can provide sufficient explanations to the user while performing autonomous feature extraction.
[0199] Fig. 10A is a diagram showing an example of a learning model of a foreign object detection unit that detects the presence or absence of abnormalities in a biometric image. Fig. 10B is a diagram showing an example of cooperation between an attribute information extraction unit and a foreign object detection unit. Since there are countless methods for creating abnormal biometric images such as images of forgeries, it is difficult to collect a large amount of training data corresponding to forgeries created using each method in advance.
[0200] Therefore, the memory 12 of the authentication processing device 10 includes a foreign object detection unit 201, which learns common features of a large number of correct biometric images and determines that an image that shows a certain difference from the learned features is an abnormal image. Learning models that can be used to achieve this type of anomaly detection include AnoGAN (Anomaly GAN), AdGAN, Efficient-GAN, and f-AnoGAN. Here, anomaly detection using f-AnoGAN will be described.
[0201] First, the generator G(x) 202 and the classifier D(x) 203 used in the GAN are trained in advance using a biometric image 66. Next, with the learning parameters of the generator G(x) 202 and the classifier D(x) 203 fixed, the encoder E(x) 204 is trained so that the generator G(x) 202 can restore the original biometric image.
[0202] Using a biometric image 66 as input, the image restored through the encoder E(x) 204 and generator G(x) 202 is the generated image G(E(x)) 205. When learning is performed using a genuine biometric image, the error between the genuine biometric image 66 and the generated image 205 is small. On the other hand, when an abnormal image is input as the biometric image 66, the original input image cannot be restored in the process of passing through the encoder E(x) 204 and the generator G(x) 202, that is, the error D1 between the biometric image 66 and the generated image 205 becomes large.
[0203] Similarly, in authenticity determination by discriminator D(x)203, the generated image G(E(x))205 is determined to be not the original input image, so the error D2 between the output of discriminator 203 for biometric image 66 and the output of discriminator 203 for generated image G(E(x))205 also becomes large.
[0204] Therefore, if D1 and D2 exceed a threshold value (a threshold value that defines the range of possible D1 and D2 for normal biometric images), the input biometric image 66 is rejected as an abnormal image. For example, 3σ (where σ is the standard deviation σ of the variations in D1 and D2) is an example of the threshold value for each of D1 and D2. The foreign object detection unit 201 compiles these errors D1 and D2 as a two-dimensional vector and outputs it as foreign object information 206.
[0205] 10B, the foreign object detection unit 201 can perform attribute determination in cooperation with the attribute information extraction unit 61 and the attribute determination unit 62. Because the foreign object information 206 is a two-dimensional vector as described above, this two-dimensional vector is input to the attribute determination unit 62, and the attribute determination unit 62 can separately determine the presence or absence of an abnormality based on predetermined thresholds for the errors D1 and D2 when performing attribute determination.
[0206] Alternatively, the attribute determination unit 62 may determine the attribute after learning the attribute information 68, to which the two-dimensional information of the foreign object information 206 is attached, using the method described above. The foreign object detection unit 201 may be included in the attribute information extraction unit 61. Similarly, the attribute information 68 may include information indicated by the foreign object information 206.
[0207] It is considered that the accuracy of identifying foreign objects increases when various input information is provided, so it is advisable to calculate, for example, the errors D1 and D2 for an image not illuminated by the light source 3 and an image illuminated by the light source 3, or to calculate the errors D1 and D2 for multiple images illuminated from different angles. This can improve the accuracy of anomaly detection. [Example]
[0208] In this embodiment, differences from embodiment 1 will be mainly described, and descriptions of similarities with embodiment 1 will be omitted where appropriate. Fig. 11A is a block diagram showing an example configuration of an integrated authentication system that maintains compatibility between authentication systems that use different authentication devices. In this embodiment, three types of authentication terminals (imaging devices 9) work together: authentication terminal A47, which is a visible light authentication terminal; authentication terminal B48, which is an infrared reflective authentication terminal; and authentication terminal C224, which is an infrared transmissive authentication terminal.
[0209] The authentication system C221 includes a registration database C222, an authentication server C223, and an authentication terminal C224 (visible light authentication terminal). The authentication server A45, the authentication server B46, and the authentication server C223 are connected to each other. The authentication server C223 is realized by the authentication processing device 10 in FIG. 1, the registration database C222 is realized by the auxiliary storage device 14 in FIG. 1, and the authentication terminal C224 includes the biometric information acquisition device 2.
[0210] In the example of Figure 11A, it is assumed that authentication system A41 including authentication terminal A47, which is a visible light authentication terminal, has been newly introduced, and authentication system B42 including authentication terminal B48, which is an infrared reflective authentication terminal, and authentication system C221 including authentication terminal C224, which is an infrared transmissive authentication terminal, have been in operation for some time.
[0211] Hereinafter, the newly introduced authentication system A41 will also be referred to as the new authentication system, and the authentication system B42 and authentication system C221 that have been in operation until now will also be referred to as the old authentication system B and old authentication system C.
[0212] Here, we assume that the two old authentication systems are already in use by many users, that the registration templates in the registration databases contained in each of the two old authentication systems cannot be re-registered for administrative reasons, and that the authentication algorithms of each of the two old authentication systems cannot be updated. Furthermore, we assume that the two old authentication systems were not originally capable of implementing compatible authentication. Therefore, in order for these three authentication systems to be compatible with each other, the new authentication system must absorb the differences between the two old authentication systems.
[0213] Two cases are assumed for compatible authentication in this embodiment. The first case is when a user of the old authentication system uses the new authentication system. The second case is when a user who has not used any of the old authentication systems registers with the new authentication system, and then uses the old authentication system without re-registering. Note that this embodiment will be described assuming 1:N authentication in which authentication is performed without entering a registration ID, but it goes without saying that this can also be applied to 1:1 authentication by entering a registration ID before registration and authentication to uniquely limit the registrant.
[0214] We will explain an example in which a user who has only used authentication system C221 now starts using authentication system A41. As shown in Figure 11A, both authentication systems include a registration database, an authentication server, and an authentication terminal. For example, authentication server C223 downloads registration data from registration database C222 and performs authentication.
[0215] When new authentication system A41 cooperates with old authentication system B or old authentication system C, if the new authentication system also adopts the protocol for downloading registration data from the registration database used by the old authentication system, the new authentication system can download the registration data of the old authentication system and perform authentication. That is, for example, authentication server A45 can perform authentication by downloading registration data from registration database C222. Similarly, if the new authentication system adopts the registration protocol of the old authentication system, it can also register a registration template generated by the new authentication system in the old authentication system.
[0216] 11B is a diagram showing an example of the configuration of the registration database C222 of the authentication system C221, which is the old authentication system. In this embodiment, the configuration of the registration database B44 is the same as the configuration of the registration database C222, so illustration and description thereof will be omitted.
[0217] Because the old authentication system does not take into consideration compatibility with different authentication systems, only the registration template for the user's own authentication system is registered together with the registration ID in the registration database C222. Note that for a user who has registered a registration template in both authentication system B42 and authentication system C221, the registration ID in authentication system B42 and authentication system C221 may not necessarily be the same.
[0218] 11C is a diagram showing an example of the configuration of registration database A43 of authentication system A41, which is a new authentication system. In addition to registration templates used by authentication system A41, registration database A43 can store registration IDs used by authentication systems B42 and C221. The registration IDs used by authentication systems B42 and C221 are used to link which registrants have been authenticated in the new authentication system when some authentication process is performed in the old authentication system after successful authentication in the new authentication system.
[0219] 12A and 12B are diagrams showing examples of the internal processing configuration of the biometric information processing unit 27 included in each of authentication systems A41 and C221. Unless otherwise specified in the following description, components with the same name will be distinguished from one another by adding an "a" to the end of the reference symbol for components included in authentication system A41, adding a "b" to the end of the reference symbol for components included in authentication system B42, and adding a "c" to the end of the reference symbol for components included in authentication system BC221. Furthermore, in the following description, when it is not necessary to distinguish between components included in authentication system A41, authentication system B41, and authentication system C221, or when it is necessary to specify which authentication system a component belongs to, the "a," "b," and "c" at the end of the reference symbol will be omitted.
[0220] In this embodiment, the internal configurations of the biometric information processing units 27b and 27c included in the old authentication systems are similar, and therefore illustration and description of the biometric information processing unit 27b will be omitted. Although the type of authentication terminal, the input biometric image 66, and the authentication algorithm used differ between the authentication system B42 and the authentication system C221, both systems have the functions of the photography determination unit, the dedicated feature extraction unit, and the matching processing unit.
[0221] The biometric information processing unit 27a further includes an imitation feature extraction unit B241 and an imitation feature extraction unit C242. The imitation feature extraction unit B241 is trained to output features that imitate the dedicated features used by the authentication system B42 when a finger image captured by a smartphone or the like that is the authentication terminal 49 of the authentication system A41 is input. Similarly, the imitation feature extraction unit C242 is trained to output features that imitate the dedicated features used by the authentication system C221 when a finger image captured by a smartphone or the like that is the authentication terminal 49 of the authentication system A41 is input.
[0222] The biometric information processing unit 27c includes, for example, an imaging determination unit 245, a dedicated feature extraction unit 64c, and a matching processing unit 25. The imaging determination unit 245 determines whether the input biometric image 66 is suitable for authentication based on the imaging state indicated by the terminal control value 67, and outputs the determination result as an imaging determination result 246. The imaging determination unit 245 corresponds to the attribute information extraction unit 61a and the attribute determination unit 62a in the authentication system A41. The dedicated feature extraction unit 64c and the matching processing unit 25 extract biometric features and perform matching based on the features, respectively. It is desirable that the dedicated feature extraction unit and the matching processing unit be optimally designed for each system.
[0223] 13 is a flowchart showing an example of compatible authentication in which a user who is registered only in the old authentication system is authenticated by the new authentication system. The user starts authentication through an authentication app installed on an authentication terminal A47 of the new authentication system (S1301).
[0224] The user holds their finger over a smartphone or other device that serves as an authentication terminal A47 of the new authentication system to capture a biometric image (S1302). The authentication server A45 performs 1:N authentication with all registered templates registered in the new authentication system (S1303). In step S1303, the attribute information extraction unit 61a and the attribute determination unit 62a perform a compatibility determination, the dedicated feature extraction unit 64a extracts features specialized for the new authentication system, and the matching processing unit 25 performs a matching process with all registered templates of the new authentication system (i.e., the processes of steps S503 to S507 are executed).
[0225] The authentication processing unit 21 of the authentication server A45 determines whether the 1:N authentication is successful (S1304). If the authentication is successful (S1304: Y), the user is already registered in the new authentication system, so compatible authentication is not performed, and the process for successful authentication, for example, payment processing, is performed (S1315). Furthermore, the imitation feature extraction unit B241 generates an imitation feature B243 and registers it in the authentication system B42, and the imitation feature extraction unit C242 generates an imitation feature C244 and registers it in the authentication system C221 (S1316), and the authentication process ends.
[0226] The process of registering the imitated features in the old authentication system in step S1316 is a process of registering the imitated features generated by the new authentication system in the old authentication system on behalf of only those authentication systems that do not have the features of the user registered in the linked authentication systems, and will be described in detail later.
[0227] In this embodiment, since it is assumed that the registered template for this user exists only in authentication system C221, 1:N authentication for new authentication system A41 is rejected and it is determined in the determination process of step S1304 that authentication has failed.
[0228] If the authentication fails (S1304: N), a compatible authentication process is performed on all registered data of the old authentication system B42 (S1305). In the compatible authentication process for the old authentication system B42, unlike the authentication process for the new authentication system (S1303) described above, the imitation feature extraction unit B241 extracts an imitation feature B243 from the biometric image and transmits it to the authentication server B46, and the authentication server B46 performs authentication using the imitation feature B243.
[0229] The imitation feature extraction unit B241 is configured by a learning model that has been trained in advance so that when a visible light biometric image 66 taken by an authentication terminal A47, which is a visible light authentication terminal such as a smartphone, is input, the imitation feature extraction unit B241 outputs a simulated feature amount generated by inputting an infrared reflection image taken by an authentication terminal B48, which is an infrared reflection authentication device of the authentication system B42. The method for configuring this learning model will be described later.
[0230] The imitation feature extraction unit B241 extracts the imitation feature B243, thereby realizing compatible authentication between the new authentication system A41 and the old authentication system B42. The authentication processing unit 21 of the authentication server B46 determines whether authentication for the old authentication system B42 is successful (S1306). If authentication for the old authentication system B42 fails (S1306: N), compatible authentication processing for the old authentication system C221 is performed in the same manner as in step S1305 (S1307).
[0231] The authentication processing unit 21 of the authentication server C223 determines whether authentication with the old authentication system C221 was successful (S1308), and if the authentication failed (S1308: N), determines whether a predetermined timeout occurred (S1309). If a timeout has not occurred (S1309: N), the process returns to step S1302. If a timeout has occurred (S1309: Y), it is determined that registration has failed (S1310), and the authentication process ends.
[0232] In this embodiment, since it is assumed that a user who has registered a registration template only in the old authentication system C221 is using the new authentication system A41 for the first time, authentication is successful in the authentication process for the old authentication system C221 (S1307, S1308).
[0233] In this way, if authentication with the new authentication system A41 fails but compatibility authentication with the old authentication system B42 or the old authentication system C221 succeeds (S1306: Y or S1309: N), it can be determined that registration with the new authentication system A41 has not been performed. Therefore, registration processing with the authentication system A41 is automatically performed.
[0234] The registration processing unit 20 of the authentication server A45 determines the registration ID in the new authentication system A41 (S1311), obtains the registration ID of the old authentication system that was successfully authenticated from the old authentication system, and writes the registration ID of the old authentication system that was successfully authenticated into the registration data of the new authentication system that corresponds to the determined registration ID (S1312).
[0235] To fully complete registration in the new authentication system, similar to the registration process in Fig. 4, registration data that are candidates for registration in the new authentication system are acquired three times to determine the registration data (S1313), and this is registered as a registration template for the new authentication system (S1314), and the process proceeds to step S1315. In the process in Fig. 13, the new authentication system executes access using the registration protocol of the old authentication system, thereby realizing compatible authentication without modifying the old authentication system.
[0236] Note that, due to the accuracy of authentication, it is possible that a user who has already been registered in an old authentication system may fail compatible authentication from the new authentication system. In this case, the automatic registration of mimicked features in step S1316 may result in the biometric information of the same person being registered multiple times in the same old authentication system with different IDs.
[0237] To prevent this, before or after automatic registration of imitation features, it is possible to confirm that the user is not registered in the old authentication system by having the user enter information such as date of birth and name, before registering the imitation features in the authentication system.
[0238] In addition, even if a user has already registered with the new authentication system, there may be cases where authentication with the new authentication system fails but compatibility matching with the old authentication system succeeds at a later stage. In such cases, the processing from step S1311 to step S1314 is performed, which causes new registration processing to be performed in the new authentication system, and the biometric information of the same person may be registered in multiple registration data in the new authentication system.
[0239] To prevent this, a full search may be performed to check whether the registration ID used when authentication was successful in the old authentication system is already registered in the registration database of the new authentication system, and whether a registration template for the new authentication system is already saved in the registration data of the new authentication system in which the registration ID was found. In this case, there is no need to perform the process of registering the new authentication system again; in other words, the automatic registration process to the new authentication system from step S1311 to step S1314 may be omitted.
[0240] In this way, by generating and registering templates that imitate the features of each old authentication system on the new authentication system side, authentication by all old authentication systems can be used. Therefore, even if it is difficult to re-register registered templates in the old authentication system that is still in operation for management reasons, compatibility between all authentication systems can be achieved by imitating the features and automatically adding them to the new authentication system.
[0241] An example of a learning method for the imitation feature extraction unit B241 will be described. Fig. 14A is a diagram showing an example of the configuration of a learning model for the imitation feature extraction unit B241 that generates imitation features of an old authentication system from an input image of a new authentication system. Fig. 14B is a flowchart showing an example of imitation feature generation processing. Note that the configuration example of the imitation feature extraction unit C242 and the processing by the imitation feature extraction unit C242 are similar to the configuration example of the imitation feature extraction unit B241 and the processing by the imitation feature extraction unit B241, respectively, and therefore will not be illustrated or described again.
[0242] In order for the imitation feature extraction unit B241 included in the authentication system A41 to extract the imitation feature B243, learning is performed using the dedicated feature extraction unit 64b and dedicated feature 70b included in the authentication system B42.
[0243] 14A and 14B differ from FIGS. 6A and 6B in that they do not include a distance learning unit 122, they include an imitation feature extraction unit B241 instead of a common feature extraction unit 63, and they perform pre-learning and iterative learning of the imitation feature extraction unit B241 and the matching processing unit 25. However, apart from these differences, they are the same as FIGS. 6A and 6B, and therefore the following description will mainly focus on these differences.
[0244] As a preliminary preparation, the attribute information extraction unit 61a, the attribute information extraction unit 61b, the attribute determination unit 62a, and the attribute determination unit 62b undergo preliminary learning (S1401) in the same manner as in step S601.
[0245] Pre-learning of the imitation feature extraction unit B241 is performed (S1402). A specific example of pre-learning will be described. The dedicated feature extraction unit 64b acquires dedicated features 70b from the biometric image included in the learning data 121b of the old authentication system B42. The attribute information extraction unit 61a and the attribute determination unit 62a determine whether the biometric image included in the learning data 121a of the new authentication system A is an image suitable for authentication.
[0246] The imitation feature extraction unit B241 generates an imitation feature B243 from a biometric image determined to be a match by the attribute determination unit 62a. Ideally, when biometric images of the same subject are input, the dedicated feature 70b generated by the authentication system B42 and the imitation feature B243 of the authentication system A41 should perfectly match.
[0247] Therefore, in the pre-learning of the imitation feature extraction unit B241, when all training data of the new and old authentication terminals for the same subject ID are input, the matching processing unit 25 of the authentication server A45 matches the imitation feature B243 with the dedicated feature 70b, the loss calculation unit 123 calculates the sum of the matching scores for the same IDs obtained in the matching process, and the learning proceeds so as to minimize the loss. Through these processes, the learning proceeds so as to obtain a pattern similar to the dedicated feature 70b as the imitation feature B243.
[0248] Learning of the attribute information extraction unit 61a and the attribute determination unit 62a (S1403) and learning of the imitation feature extraction unit B241 (S1404) are performed alternately, and it is determined that the learning has converged (S1405). The series of processes to end the learning when convergence is confirmed is the same as the series of processes from step S603 to step S605.
[0249] Through learning by the imitation feature extraction unit B241 and the imitation feature extraction unit C242, the new authentication system can acquire biometric features similar to the registered templates of the old authentication system, thereby realizing compatible authentication.
[0250] Fig. 14C is a schematic diagram showing an example of dedicated features obtained from a biometric image in the old authentication system. Fig. 14D is a schematic diagram showing an example of mimicked features obtained from a biometric image in the new authentication system. The subject ID of the features in Fig. 14C and Fig. 14D is both 07, and although the features in Fig. 14C and Fig. 14D were obtained from biometric images of the same subject, the biometric images 66 are different because the authentication terminals that captured the images were different.
[0251] However, the dedicated feature 70b obtained as described above is almost identical to the imitation feature B 243. As an intermediate step in the internal processing, a region 281 similar to the biometric image of authentication system B is automatically selected by the above-described learning in the biometric image of authentication system A41, and a biometric feature is extracted from region 281.
[0252] As another example of a learning model for the imitation feature extraction unit B241 and the imitation feature extraction unit B242, a type of GAN (Generative Adversarial Networks) called pix2pix may be used to generate a pattern similar to the dedicated feature 70b from the biometric image of the new authentication system. As another example of a learning model for the imitation feature extraction unit B241 and the imitation feature extraction unit B242, a style conversion technology called StyleGAN, which can make the style of an entire image similar to that of a different image, may be used to convert an image captured with the new authentication system into an image with a style that mimics an image captured with the old authentication system, and then the algorithm of the feature extraction unit of the old authentication system may be applied directly to the converted image to generate features. Because style-converted images are similar to each other, similar biometric features will inevitably be obtained if the feature extraction algorithm used in the subsequent stages is the same.
[0253] In this embodiment, the imitation feature extraction unit in the new authentication system is trained to imitate and output the enrolled templates of the old authentication system, but it may also be trained to extract common features of a different format from the enrolled templates of the old authentication system. In this case, it will no longer be possible to directly match the features generated by the new authentication system with the enrolled data of the old authentication system, but using common features may sometimes result in higher matching accuracy than completely imitating the enrolled templates of the old authentication system, and it will also be possible to generate common features with a small data size, which has the advantage of speeding up 1:N authentication.
[0254] This section explains a case where a user who has not registered a registration template in the old authentication system registers it for the first time using a visible light authentication terminal such as a smartphone in the new authentication system, and then uses the old authentication system in that state.
[0255] The processing in Figure 13 includes a process (S1316) for registering the imitation feature in the old authentication system. The processing in step S1316 is a process for performing full matching with all linked old authentication systems when the new authentication system's compatible authentication is successful, and automatically registering the imitation feature in the unregistered old authentication systems for which authentication was not successful. After the compatible authentication is successful through the processing in step S1316, the user can be successfully authenticated even if they perform authentication at a site where the old authentication system is in operation. Therefore, this case is also included in the processing in Figure 13.
[0256] If the user clearly knows whether they are registered or not for each authentication system, and if the user declares in advance that they are not registered or the authentication system they are registered with, they can select, as needed, compatible authentication processing for the old authentication system such as in steps S1305 and S1307, or authentication processing for their own authentication system. Furthermore, since registration processing for the new authentication system such as in step S1314 and registration processing for the old authentication system in step S1316 can be similarly selected, the processing in Figure 13 is simplified, processing speed is improved, and the risk of erroneous automatic registration is reduced. Furthermore, the registration processing can be designed so that the user is always required to declare an unregistered system when performing compatible authentication, which can also achieve the same effect. [Example]
[0257] In this embodiment, differences from the previous embodiment will be mainly described, and explanations of similarities with the previous embodiment will be omitted as appropriate. Fig. 15 is a diagram showing an example of the configuration of a biometric information processing unit 27a of an authentication system that performs compatible authentication for variations in biometric features that occur over time.
[0258] Biometric features often change over time due to factors such as environmental changes such as temperature and humidity due to seasonal and diurnal variations, changes in physical condition, and changes in the biometrics themselves over time. Therefore, in order to perform stable biometric authentication, it is also important to improve compatibility with respect to changes over time. Therefore, in this embodiment, a method for improving compatibility with changes in biometric features over time will be described. The configuration of the authentication system is similar to that shown in FIG. 2 or FIG. 11, for example.
[0259] The biometric information processing unit 27a of this embodiment further includes an aging conversion unit 301. The aging conversion unit 301 estimates and generates an aging image 302, and provides the generated aging image 302 to the common feature extraction unit 63 and / or the dedicated feature extraction unit 64a. This makes it possible to predict biometric features when aging occurs, improving authentication accuracy.
[0260] The functions and effects of the learning models of the attribute information extraction unit 61a, attribute determination unit 62a, common feature extraction unit 63, dedicated feature extraction unit 64a, and matching processing unit 25 are the same as those in the above-described embodiment, and therefore will not be described here. In Fig. 15, the aging conversion unit 301 is a separate module from the common feature extraction unit 63 and dedicated feature extraction unit 64a, but these may be integrated into one module. In Fig. 15, a visible light authentication terminal is used as the authentication terminal A47, but any type of authentication terminal may be used.
[0261] First, a description will be given of a specific example of the configuration of the neural network of the aging transformation unit 301. An example will be described in which StyleGAN, a representative method for performing image style transformation using GAN technology, is used.
[0262] StyleGAN has internal latent variables that can change various characteristics of the generated images. For example, if male and female facial images are given to StyleGAN as training data, the distribution of latent variables for the male training data and the female training data will tend to be different. By calculating the latent variables for the male image and making them closer to those for the female (replacing them), the male image can be converted into a female image.
[0263] Similarly, by taking biometric images of the same subject over a long period of time and observing changes in the distribution of latent variables, it is clear that the distribution of latent variables shifts in a specific direction over time. Therefore, by replacing the latent variables obtained from a biometric image at time t with the latent variables obtained from a biometric image at another time t' (where t' > t), it is possible to convert the biometric image at time t into the biometric image at time t', i.e., to estimate the biometric image at time t'. This is equivalent to estimating changes over time when t' - t has elapsed since time t. Similarly, it can be said that it is possible to retroactively estimate the biometric image at time t using the biometric image obtained at time t'.
[0264] An example of learning and inference by the aging transformation unit 301 will now be described. To learn trends in temporal changes, biometric images taken from a large number of subjects over as long a period as possible are collected as training data. Based on the aforementioned StyleGAN technology, latent variables are calculated for biometric images taken at various times, and changes in the latent space related to temporal changes in the distribution of the latent variables are obtained. For example, changes in the latent variables of biometric images taken in the morning and evening, or changes in the latent variables after one month, three months, six months, one year, and three years are obtained. At this time, changes over periods that cannot be obtained by actual measurement may be estimated by interpolation.
[0265] FIG. 16 is a graph showing an example of changes in the latent variables of a biometric image in response to short-term changes such as changes in daytime temperature and long-term changes in the user over time. The origin of the graph indicates the latent variable at the time of registration. The curve extending from the origin toward time t is a variation curve 321 of the latent variable. The variation curve 321 represents the transition of the latent variable over time.
[0266] The x-axis and y-axis represent two feature axes on which latent variables are located. The latent variables are assumed to have, for example, 512 dimensions x 18 items. The fluctuation curve 321 moves gradually upward and downward due to short-term fluctuations during the day. For example, if the biometric image is an image of finger veins, the blood flow may change between the cold morning hours and the warm afternoon hours, causing the biometric image to change slightly from the time of registration.
[0267] In the long term, for example, due to seasonal variations, the points on the variation curve 321 shift to the upper right. For example, in the case of a biometric finger vein image, the blood flow rate and the condition of rough skin may change between the cold winter and the hot summer. Such changes in conditions cause long-term variations.
[0268] Furthermore, after moving to the upper right on the variation curve 321, it turns back toward the origin again. This turning back toward the origin means, for example, that the seasons are changing and it is getting warmer again, and the biometric image is approaching the image at the time of enrollment. Furthermore, the points on the variation curve 321 do not necessarily return to the origin in the end, but rather there is a slight difference from the origin. This indicates essential changes over time in the living body (for example, aging, etc.).
[0269] If the time transition of such latent variables is investigated in advance, it becomes possible to remove the time-varying components. For example, if you want to convert the current biometric image to the biometric image at the time of enrollment, you can convert it to the image as it was taken at the time of enrollment by shifting the latent variable 322 of the current biometric image downward and to the left so that it overlaps with the latent variable 323 of the biometric image at the time of enrollment.
[0270] An example of aging transformation when actually performing authentication will be described. The aging transformation unit 301 calculates latent variables for the original biometric image input during authentication using a technique such as StyleGAN. The aging transformation unit 301 acquires the registration time 310 for each registration data to be matched from the registration database A43.
[0271] Based on the variation curve 321 of the latent variable shown in FIG. 16, the aging conversion unit 301 shifts the latent variable by the difference between the time when the registered data was registered and the current time before matching with each registered data, and generates a biometric image taken at the registration time, i.e., an aging image 302.
[0272] The aging image 302 is expected to resemble a biometric image taken back to the time of enrollment. The dedicated feature extraction unit 64a and / or the common feature extraction unit 63 of the authentication server A45 extract dedicated features and / or common features from the aging image 302, and the matching processing unit 25 of the authentication server A45 performs matching using the extracted features. By generating the aging image 302, the aging conversion unit 301 can estimate the biometric image at the time of enrollment, and the influence of aging can be suppressed, thereby improving authentication accuracy.
[0273] However, in the above-mentioned embodiment, since the aging transformation must be performed for each registered data to be compared, it is expected that the processing speed will decrease. In contrast to this, it is also possible to select in advance representative times when a biometric is typically prone to change, such as morning or afternoon, summer or winter, and check the changes in the latent variables of biometric images taken in each time period, and then convert the images to the time period closest to the registration time period for comparison.
[0274] If morning and afternoon, and summer and winter are selected, the biometric image at the time of authentication may be subjected to aging transformation a maximum of four times. Furthermore, the registration time may not be stored in the registration database A43a. If the registration time is unknown, aging images 302 may be generated for these four time periods, and the authentication result may be determined based on the matching result with the highest similarity.
[0275] In the above example, aging transformation is performed on the biometric image at the time of authentication, but aging transformation may also be performed at the time of registration. For example, the time period that tends to have the highest authentication accuracy among the combinations of morning or afternoon and summer or winter is evaluated in advance. Here, it is assumed that the combination of summer afternoon tends to have the highest authentication accuracy.
[0276] In this case, the aging conversion unit 301 always performs aging conversion to the summer daytime at the time of registration, regardless of the registration time zone, and the biometric features acquired from the aging image 302 obtained by this aging conversion are registered. The aging conversion unit 301 also performs aging conversion to the summer daytime at the time of authentication. Through these processes, the assumed time zones for both registration and authentication match, and temporal changes are always unified, thereby achieving authentication that is resistant to aging. Furthermore, since aging conversion is performed only once at the time of authentication, it is also advantageous in terms of processing speed.
[0277] 17A and 17B are schematic diagrams showing changes in a finger image over time and the effect of processing to suppress the effects of the changes over time. (A) is an example of a finger image 341 taken during the day in autumn, showing thick finger veins 344 and thin finger veins 345.
[0278] (B) is an example of a finger image 342 captured the morning after (A), i.e., on an autumn morning. In (B), about half a day has passed since (A), but because the finger image 342 was captured in the morning when the temperature was slightly lower, the thick finger veins 344 have become slightly thinner.
[0279] (C) is an example of a finger image 343 taken on a winter morning about three months after registration. In (C), thick finger veins 344 have become even thinner, and thin finger veins 345 are disappearing. Furthermore, due to the drop in temperature and humidity, rough skin 346 has been captured in the image, which is a significant change from finger image 341 at the time of registration in (A), and there is a possibility that authentication accuracy will deteriorate. In response to this, we consider performing aging conversion on finger image 341 in (C) to return it to the image quality of three months ago.
[0280] (D) is finger image 303 obtained by subjecting finger image 343 of (C) to aging transformation using the method described above, and converting it to a state that was captured during the day in autumn approximately three months ago. In finger image 303, slight roughness 346 on the epidermis remains and cannot be removed, but it can be seen that the blurring of thin finger veins 345 has improved and thick finger veins 344 have also become thicker. In other words, finger image 303 of (D) has a pattern that is close to the pattern at the time of enrollment shown by finger image 341 of (A). Therefore, even if a large amount of time has passed, changes due to the passage of time can be suppressed by aging transformation, and compatibility over time can be maintained without reducing authentication accuracy.
[0281] Although the above describes an example in which the biometric image is a finger image, the aging transformation unit 301 can also perform similar image transformation on a face image. For example, by acquiring in advance changes in the distribution of latent variables for all kinds of changes that can occur over time, such as whether or not a beard is grown, whether or not glasses are worn, whether or not a mask is worn, whether or not a hairstyle is changed, and the passage of 10 years, the style of the original image can be transformed by changing the latent variables to an arbitrary state, thereby converting the image into one that is close to the state at the time of enrollment, thereby improving authentication accuracy.
[0282] In particular, when the biometric image is a facial image, by standardizing factors such as whether or not a mask is worn, whether or not glasses are worn, and the type of hairstyle, and then converting the facial image into a consistent state at the time of registration and authentication before matching, authentication can be performed under consistent conditions, which is expected to improve accuracy.
[0283] Although the example in which the aging transformation unit 301 performs aging transformation on biometric images has been described, aging transformation including style transformation may also be performed on extracted features. In this case, the aging transformation unit 301 can simply perform aging transformation after extracting the features, and can examine changes in latent variables for the features during learning. Since biometric features generally have a smaller data volume than the original biometric images, improved conversion speed can be expected. [Example]
[0284] In this embodiment, differences from the previous embodiment will be mainly described, and explanations of similarities with the previous embodiment will be omitted as appropriate. Fig. 18 is a block diagram showing an example of the configuration of an authentication system that achieves high-precision authentication by autonomously learning an optimal control method for the authentication terminal.
[0285] As in the previous embodiments, the authentication system A41 in this embodiment includes a registration database A43, an authentication server A45, and an authentication terminal A47. The authentication terminal A47 also includes an authentication processing device 10, an imaging device 9, a light source 3, and a display device 15, and the authentication processing device 10 can transmit and receive data to and from each of the imaging device 9, the light source 3, and the display device 15.
[0286] The user 361 presents his / her biometric information to the imaging device 9 of the authentication terminal A47 while checking the guidance displayed on the display device 15. In this embodiment, the authentication systems B42 are linked to each other as in the previous embodiment, but only a single authentication system A41 may operate.
[0287] Fig. 19 is a diagram showing an example configuration of the feature extraction unit 24 and matching processing unit 25 that autonomously learn optimal control of the authentication terminal A47. The example configuration of Fig. 19 is almost the same as the example configuration of Fig. 3, but differs from the example configuration of Fig. 3 in that it includes a terminal control unit 381 and control information 382. The following mainly describes the terminal control unit 381 and control information 382, and as the other configurations are the same as those in the above-mentioned embodiments, explanations will be omitted as appropriate.
[0288] When a biometric image 66 and a terminal control value 67 are input, the attribute information extraction unit 61a outputs attribute information 68. As described above, the attribute information 68 is information acquired from image characteristics and the relationship between the image characteristics and the terminal control value, and indicates, for example, characteristics related to the hand posture, characteristics related to the hand shape, characteristics related to the image brightness, characteristics related to the image clarity, etc.
[0289] When the attribute information 68 is input, the terminal control unit 381 outputs control information 382 including information on optimal parameter control values for the authentication terminal A47 and optimal user guidance for the user. The control information 382 is fed back to the authentication terminal A47, and hardware included in the authentication terminal A47, such as the imaging device 9, light source 3, and display device 15, is controlled based on the control information 382. At this time, the user guidance is displayed on the display device 15, so that the user 361 is indirectly controlled optimally.
[0290] The terminal control unit 381 is configured with a learning model similar to the attribute information extraction unit 61a, and is trained to maximize authentication accuracy, as will be described later. Therefore, the terminal control unit 381 controls the imaging device 9 and light source 3 to maximize image quality, and can provide guidance to the user 361 on the biometric presentation method most suitable for authentication.
[0291] 20A is a diagram showing an example of a list of controllable device parameters of hardware included in an authentication terminal to be controlled. Parameter list 401 defines a list of controllable parameters of, for example, the image capture device 9 and the light source 3 included in the authentication terminal, along with their IDs.
[0292] FIG. 20B is a diagram showing an example of a user guidance list. User guidance list 402 is a list that predefines how to guide the user. In this embodiment, authentication accuracy is maximized by optimally controlling the device parameters listed in parameter list 401 and user guidance list 402 and optimally providing guidance to the user. Note that parameter list 401 and user guidance list 402 are stored in memory 12, for example, and FIGS. 20A and 20B are examples in which some of the items that should originally be defined are extracted.
[0293] 20C is a diagram showing an example of an optimal behavior list 403 corresponding to the attribute classification results. The left column of the optimal behavior list 403 stores an example of attribute information that may be included in the classification results using the methods of FIGS. 8A and 8B.
[0294] Examples of optimal actions corresponding to the classification results are stored in the right column of the optimal action list 403. The optimal actions defined in the optimal action list 403 are learned actions that maximize authentication accuracy.
[0295] In the optimal behavior list 403 of Figure 20C, for example, if the captured image is classified as being too dark, the optimal behavior is to increase the light intensity value of "light source 01" by 3, and if the image is classified as being one in which the hand is shifted to the left, the optimal behavior is to display the message "Move your hand to the right," which is "user guidance 04" defined in the user guidance list 402 of Figure 20A, on the display device 15.
[0296] However, because attribute classification is performed in a self-organizing manner through unsupervised learning as described above, it is often not possible to express it in a table such as optimal behavior list 403 shown in Fig. 20C. Therefore, a list equivalent to optimal behavior list 403 may be generated using a learning model.
[0297] 20D is a diagram showing an example of the configuration of terminal control unit 381, which is a learning model that converts attribute information into optimal behavior. When attribute information 68 is, for example, a 512-dimensional feature vector, the input layer of terminal control unit 381 is provided with a node consisting of a neural network with 512 elements so that the feature vector can be input.
[0298] The terminal control unit 381 passes through a multi-layer (e.g., 10-layer) neural network and finally outputs control information 382 from the output layer that outputs the behavior in response to the device parameters and the behavior in response to the user guidance. The control information 382 corresponds to the aforementioned optimal behavior list 403, and can acquire the current optimal behavior by analyzing complex attribute information.
[0299] The output layer of the neural network of the terminal control unit 381 is composed of the same number of nodes as the total number of all items in the device parameter list 401 and the user guidance list 402. In the example of Fig. 20D, information about device parameters is reflected above the control information 382, and information about user guidance is reflected below the control information 382.
[0300] The information about the device parameters includes information about which parameters should be controlled and how much they should be controlled. Therefore, the control value itself is output by using an identity function as the activation function of the output layer of the terminal control unit 381.
[0301] Similarly, for the output layer related to user guidance of the terminal control unit 381, an identity function can be used to obtain the level of guidance for each item. However, since the user 361 has difficulty understanding complex guidance, in this embodiment, the output values of all guidances are aggregated using a softmax function and normalized so that the sum is 1.0 so that only one most effective guidance is displayed. The output layer related to user guidance is trained so that the output value of the most effective guidance becomes the highest. As a result, the terminal control unit 381 can provide effective and easy-to-understand guidance to the user 361 by presenting the guidance with the highest output value from the output user guidance results to the user 361.
[0302] In this way, the terminal control unit 381 is trained to take optimal actions according to the attributes of the captured image, which makes it possible to improve authentication accuracy. In contrast, conventional technologies include, for example, technologies that control the light intensity of a light source to an optimal value, but these often control the light intensity to a predetermined value rather than controlling the light intensity to maximize authentication accuracy. Furthermore, control that takes into account the balance of multiple control values, such as exposure control and gain control in addition to light source control, is difficult to develop due to the complex design.
[0303] Therefore, in the conventional technology, the viewpoints of maximizing authentication accuracy and ease of development were not taken into consideration. However, in the optimal control technology according to the present embodiment, optimal parameter adjustment can be automatically performed based on a large amount of learning data, making it possible to obtain optimal control values.
[0304] For example, with regard to the dark image in Figure 7(F) in Example 1, it is assumed that the dark image is caused by a dark environment and by an exposure time that is too short. Since the attribute information extraction unit receives not only image features but also device control values including the exposure time, the attribute information extracted by the attribute information extraction unit reflects information indicating which factor is causing the image to be dark.
[0305] In other words, if the image is dark because the exposure time is too short, a brighter image can be captured by extending the exposure time, and if the light source is dark, a brighter image can be captured by increasing the light intensity. In this way, the terminal control unit 381 can capture a biometric image under optimal conditions by controlling the device based on the attribute information.
[0306] In this embodiment, the terminal control unit 381 also presents optimal guidance to the user 361. In conventional technology, a designer empirically defines the position and state of a biometric entity that results in a good image capture. In conventional technology, a technology for detecting whether a biometric entity is presented in the defined position and state is individually designed, and whether the image is suitable for authentication is determined based on the detection result of the presented biometric entity.
[0307] However, even if the image is determined to be suitable for authentication, this does not necessarily maximize authentication accuracy, and it is also difficult to appropriately determine the threshold for determining how much deviation from the ideal posture the user must have before receiving guidance.
[0308] On the other hand, in this embodiment, in order to guide the user 361 to maximize authentication accuracy, it is possible to automatically present appropriate guidance for the current photographing situation. Furthermore, if it is determined that authentication will be successful without any problems, there is no need to provide special guidance to the user 361, and guidance such as "Please wait as is" can be displayed. In other words, if it is determined that authentication accuracy will not deteriorate, unnecessary guidance will not be provided, thereby realizing highly accurate authentication and improving operability.
[0309] Fig. 21 is a flowchart showing an example of a learning method of the terminal control unit 381. In this embodiment, terminal control is performed based on reinforcement learning, in which trial and error is performed to determine what control should be performed from all the device parameters and user guidance shown in Figs. 20A and 20B. The biometric information processing unit 27a further includes the terminal control unit 381 and an authentication condition selection unit 423. In addition, authentication conditions 422 are stored in the memory 12 of the authentication server A45.
[0310] The terminal control unit 381 performs learning to increase the accuracy rate of classification of the input biometric image 66. Therefore, after feature extraction by the common feature extraction unit 63 and / or the dedicated feature extraction unit 64a and matching by the matching processing unit 25, reinforcement learning is performed to minimize loss values such as the cross entropy loss obtained by the loss calculation unit 123 as described above, that is, to maximize authentication accuracy.
[0311] In general, reinforcement learning involves taking actions to maximize rewards in the environment. Here, the authentication terminal A47 and the subject 421 correspond to the environment, the accuracy rate of authentication corresponds to the reward, and the device parameter settings and the output of user guidance correspond to actions. In other words, by checking the losses when the device parameters and user guidance are output in various ways and taking actions to minimize these losses, the system will gradually be able to take optimal actions.
[0312] The authentication condition 422 is information indicating whether to perform compatible authentication or dedicated authentication within the authentication system. The authentication condition selection unit 423 switches whether to input the biometric image 66 to the common feature extraction unit 63 or the dedicated feature extraction unit 64a according to the authentication condition 422.
[0313] Generally, the optimal control method may differ between performing authentication within one's own authentication system and performing compatible authentication between other authentication systems. For example, if the position at which the finger is photographed differs depending on the authentication terminal, it is expected that the guidance method will be changed so that the common part of the finger can be photographed. Therefore, in this embodiment, which authentication system will be used to perform matching is set in the authentication condition 422, and then the terminal control unit 381 is trained.
[0314] Furthermore, the authentication condition 422 is linked to the end of the attribute information 68 so that the authentication condition 422 is reflected in the learning of the terminal control unit 381. It should be noted that the authentication condition 422 is changed randomly and evenly during learning. For example, each time the subject makes a new input trial, it can be randomly determined whether to perform feature extraction using the dedicated feature extraction unit 64a or the common feature extraction unit 63.
[0315] 22 is a flowchart showing an example of the learning process of the terminal control unit 381. Learning of the attribute information extraction unit 61a, the attribute determination unit 62a, the common feature extraction unit 63, and the dedicated feature extraction unit 64a is performed in advance using the techniques of the above-mentioned embodiments or the like (S2201).
[0316] The terminal control unit 381 selects one unselected subject 421 from the reinforcement learning subjects 421 (S2202), determines the authentication conditions 422 using random numbers, and the authentication terminal A47 starts attempting to capture a biometric image (S2203). The selected subject 421 holds his / her biometric body while following instructions on the screen of the authentication terminal A47, and the authentication terminal A47 captures a biometric image 66 (S2204).
[0317] The attribute information extraction unit 61a extracts the attribute information 68 from the biometric image 66 captured in step S2204, and the attribute determination unit 62a performs attribute determination using, for example, the same method as step S404 (S2205).
[0318] The common feature extraction unit 63 or the dedicated feature extraction unit 64a extracts features from the biometric image 66 in accordance with an instruction from the authentication condition selection unit 423, and the matching processing unit 25 performs a round-robin match between the extracted features and the features of the subject 421 photographed so far, and obtains a matching score 72 (S2206). Note that, since there is no other person's data in the matching for the first subject 421, biometric images of multiple subjects may be photographed in advance in an ideal photographing environment, and the matching processing unit 25 may execute the matching process and obtain the matching score 72 using the features in these photographed images.
[0319] The loss calculation unit 123 calculates the loss (S2207). As in the above-described embodiment, the loss calculation unit 123 calculates the cross entropy loss based on, for example, the error rate of incorrect classification.
[0320] Furthermore, for example, the loss calculation unit 123 may determine the loss so that the loss is low when the distributions of the matching scores between biometrics of the same subject and the matching scores between biometrics of different subjects differ significantly, or conversely, so that the loss is low when the actual person and another person are almost indistinguishable. Specifically, the loss calculation unit 123 can, for example, actually measure or statistically estimate the intersection between the actual person's score distribution and the score distribution of another person, and calculate the equivalent error rate at which the actual person's errors and other person's errors match as such a loss. By proceeding with learning so as to minimize this equivalent error rate, the authentication error rate can be made as small as possible.
[0321] Learning of the terminal control unit 381 is carried out (S2208). In general reinforcement learning, a policy function can be defined that expresses the action to be taken in response to the current input as a probability, but in this embodiment, the terminal control unit 381 itself plays the role of the policy function. Then, the attribute information 68 is provided to the current terminal control unit 381 to infer and calculate current control information 382 (S2209), and the calculated control information 382 is fed back to the imaging device 9, light source 3, and / or display device 15 of the authentication terminal A47 (S2210).
[0322] In the early stages of learning, even if the biometric image 66 is dark, the terminal control unit 381 often takes meaningless actions, such as presenting the user with guidance such as "move your hand to the right" or further dimming the light.
[0323] However, if the terminal control unit 381 finds that these actions will degrade (or not improve) the authentication accuracy, it will take a different action. For example, if the biometric image 66 is dark, increasing the light intensity or lengthening the exposure time will result in a brighter image, which will improve the authentication accuracy.
[0324] The terminal control unit 381 can gradually improve its accuracy while learning this relationship between actions and rewards. An advantage of using a policy function in the terminal control unit 381 is that, because the policy function expresses the action to be taken in terms of probability, in the above-mentioned current control information inference (S2209), the terminal control unit 381 will take a suboptimal action with a small probability, but this will ultimately result in searching for a better action, making it possible to obtain a globally optimal solution without falling into a locally optimal solution.
[0325] When input attempts to the terminal control unit 381 are repeated for a certain period of time, various guidance messages are displayed one by one, and the subject 421 has to follow the guidance and hold the living body over the device again. In addition, the parameters of the light source 3 and the image capture device 9 change one by one, and control to improve authentication accuracy is achieved through trial and error.
[0326] The terminal control unit 381 determines whether the input attempt has timed out (S2211). If the input attempt has not timed out (S2211: N), the process returns to step S2209, and when the input attempt has timed out (S2211: Y), the terminal control unit 381 determines whether the input attempt for the same subject 421 has been repeated a predetermined number of times (S2212).
[0327] If the number of input attempts for the same subject 421 has not reached the predetermined number (S2212: N), the process returns to step S2203. Even for the same subject 421, repeated photographing can change the way the living body is held or the surrounding lighting environment, so the results of reinforcement learning can be made more stable by repeatedly photographing the subject.
[0328] When the input attempts for the same subject 421 reach a predetermined number (S2212: Y), the next subject 421 is assigned (S2213), and it is determined whether the input attempts for all subjects 421 have been completed (S2214). When there is a subject 421 for whom the input attempts have not been completed (S2214: N), the process returns to step S2202, and when the attempts for all subjects 421 have been completed (S2214: Y), the learning process of the terminal control unit 381 is terminated.
[0329] In this embodiment, the attribute information 68 is input to the terminal control unit 381, but the biometric image 66 may be input directly to the terminal control unit 381.
[0330] Furthermore, as described above, development costs increase when many subjects 421 are gathered and input trials to the terminal control unit 381 are repeated many times. Therefore, it is possible to collect, in advance, learning data such as changes in image quality that occur when camera parameters or light sources are changed, and changes in the way the user holds the device in response to the display of user guidance, and to use GAN technology such as StyleGAN to find the relationship between changes in each device parameter and the distribution of latent variables, and to simulate various changes in device parameters and ways of holding the device from a single base biometric image 66, thereby training the terminal control unit 318 in a simulated environment.
[0331] Furthermore, although biometric features generally differ depending on the subject 421, the biometric images 66 used as training data for the terminal control unit 318 may be simulated using GAN technology. By using this method, even if the number of subjects 421 is small, it is possible to efficiently generate simulated training data for the terminal control unit 318, thereby increasing the learning efficiency of various learning models included in the biometric information processing unit 27a and improving authentication accuracy.
[0332] The present invention is not limited to the above-described embodiments, but includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0333] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.
[0334] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]
[0335] 2 Biometric information acquisition device, 3 Light source, 9 Imaging device, 10 Authentication processing device, 11 CPU, 12 Memory, 13 Interface, 14 Auxiliary storage device, 15 Display device, 16 Input device, 17 Speaker, 18 Image input device, 20 Registration processing unit, 21 Authentication processing unit, 22 Photography control unit, 23 Attribute determination unit, 24 Feature extraction unit, 25 Matching processing unit, 26 Authentication determination unit, 27 Biometric information processing unit, 41 Authentication system A, 42 Authentication system B, 43 Registration database A, 44 Registration database B, 45 Authentication server A, 46 Authentication server B, 47 Authentication terminal A, 48 Authentication terminal B, 61a Attribute information extraction unit, 61b Attribute information extraction unit, 62a Attribute determination unit, 62b Attribute determination unit, 63 Common feature extraction unit, 64a Dedicated feature extraction unit, 64b Dedicated feature extraction unit, 64c Dedicated feature extraction unit, 122 Distance learning unit, 123 loss calculation unit, 181 landmark detection layer, 183 landmark normalization layer, 185 segmentation layer, 187 feature selection layer, 188 autonomously extracted feature, 201 foreign object detection unit, 221 authentication system C, 222 registration database C, 223 authentication server C, 224 authentication terminal C, 241 imitation feature extraction unit B, 242 imitation feature extraction unit C, 245 photography judgment unit, 1000 biometric authentication system
Claims
1. 1. An authentication system comprising: a processor and a memory, The memory includes: a first model that extracts a first feature corresponding to a first imaging method from a biometric image captured by the first imaging method corresponding to the authentication system; a second model that extracts, from a biometric image captured by the first imaging method, a second feature that can be extracted from a biometric image captured by a second imaging method corresponding to another authentication system; a biometric image of the person to be authenticated taken using the first imaging method; The first feature amount of the registrant; The second feature of the registered person is stored; the second model is trained using a first training biometric image of the subject captured by the first imaging method and a second training biometric image of the subject captured by the second imaging method; The processor: extracting a first feature of the person to be authenticated based on a biometric image of the person to be authenticated and the first model; performing a first authentication of the person to be authenticated based on a comparison result between the extracted first feature amount of the person to be authenticated and the first feature amount of a registered person; If the first authentication fails, extracting a second feature of the person to be authenticated based on a biometric image of the person to be authenticated and the second model; an authentication system that performs second authentication of the person to be authenticated based on a result of matching the extracted second feature amount of the person to be authenticated with the second feature amount of a registered person;
2. 2. The authentication system according to claim 1, When the second authentication is successful, the processor stores the extracted first feature of the person to be authenticated in the memory as the first feature of the registered person.
3. 2. The authentication system according to claim 1, a second feature amount that is output when a biometric image captured by the first imaging method is input to the second model; and An authentication system in which the second model is within a predetermined distance from a second feature output when a biometric image of a biometric organism identical to the biometric organism contained in the biometric image is input, the second feature being output when the biometric image is captured using the second capturing method.
4. 2. The authentication system according to claim 1, the other system holds a third model that extracts the second feature amount from a biometric image captured by the second imaging method; a second feature amount that is output when a biometric image captured by the first imaging method is input to the second model; and An authentication system in which the second feature output when a biometric image of a biometric organism identical to the biometric organism contained in the biometric image, photographed using the second photographing method, is input to the third model, and the second feature output is within a predetermined distance.
5. 2. The authentication system according to claim 1, connected to an imaging device that captures images using the first imaging method; the memory stores attribute information indicating at least one of image quality of a biometric image of the person to be authenticated and a biometric state captured in the biometric image; When the processor determines based on the attribute information that the biometric image of the person to be authenticated is not suitable for authentication, the processor outputs information instructing the imaging device to capture a biometric image of the person to be authenticated; An authentication system, wherein the first authentication and the second authentication are performed based on a biometric image of the person to be authenticated captured by the imaging device.
6. 6. The authentication system according to claim 5, the attribute information indicates the state of the living body captured in the living body image; the memory stores guidance information indicating a correspondence between the attribute information and guidance indicating an action to be taken by the subject to be authenticated in order to change the biometric state captured in the biometric image; and When the processor determines, based on the attribute information, that the biometric image of the person to be authenticated is not suitable for authentication, the processor generates data for outputting guidance corresponding to the attribute information in the guidance information.
7. 6. The authentication system according to claim 5, the attribute information indicates image quality of a biometric image of the person to be authenticated; the memory holds control value information indicating a correspondence between the attribute information and a control value of the imaging device for improving image quality of the living body image; When the processor determines that the biometric image of the person to be authenticated is not suitable for authentication based on the attribute information, the processor acquires a control value corresponding to the attribute information in the control value information; An authentication system that controls the imaging device based on the acquired control value.
8. 2. The authentication system according to claim 1, an authentication system in which the second model is trained based on a distance in feature space between the second features obtained by inputting the first training biometric image and the second training biometric image of the same subject into the second model, and a distance in feature space between the second features obtained by inputting the first training biometric image and the second training biometric image of a different subject into the second model.
9. An authentication system according to claim 8, An authentication system, wherein the second model is also capable of extracting the second feature from a biometric image captured using the second imaging method.
10. An authentication system according to claim 8, the memory stores attribute information indicating at least one of image quality of the first biometric learning image and the second biometric learning image and a state of the living body captured in the biometric learning image; An authentication system in which the second model is trained using the first training biometric image and the second training biometric image that are determined to be suitable for authentication based on the attribute information.
11. An authentication system according to claim 1, the second model is capable of extracting the second feature amount from a biological image captured by the second imaging method; the processor learns the second model based on a distance in a feature space between the second features obtained by inputting the first training biometric image and the second training biometric image of the same subject into the second model, and a distance in the feature space between the second features obtained by inputting the first training biometric image and the second training biometric image of a different subject into the second model; In training the second model, detecting biometric landmarks from the first training biometric image and the second training biometric image of the same subject; normalizing the biometric features in the first training biometric image and the second training biometric image of the same subject based on the detected landmarks; An authentication system that detects the second feature based on the normalized biometric information.
12. An authentication method using an authentication system, comprising: the authentication system includes a processor and a memory; The memory includes: a first model that extracts a first feature corresponding to a first imaging method from a biometric image captured by the first imaging method corresponding to the authentication system; a second model that extracts, from a biometric image captured by the first imaging method, a second feature that can be extracted from a biometric image captured by a second imaging method corresponding to another authentication system; a biometric image of the person to be authenticated taken using the first imaging method; The first feature amount of the registrant; The second feature of the registered person is stored; the second model is trained using a first training biometric image of the subject captured by the first imaging method and a second training biometric image of the subject captured by the second imaging method; The authentication method includes: the processor extracts a first feature of the person to be authenticated based on a biometric image of the person to be authenticated and the first model; the processor performs a first authentication of the person to be authenticated based on a comparison result between the extracted first feature amount of the person to be authenticated and the first feature amount of a registered person; If the processor fails the first authentication, the processor extracts a second feature of the person to be authenticated based on a biometric image of the person to be authenticated and the second model; The authentication method further comprises: the processor executing a second authentication of the person to be authenticated based on a result of matching the extracted second feature of the person to be authenticated with the second feature of a registered person.
Citation Information
Patent Citations
Fingerprint collating device
JP2002074364A
Biological information management method using two or more biometrics devices
JP2013186510A
Image processing apparatus, image processing method, program, and storage medium
JP2014164697A
Biometric authentication device and biometric authentication method
JP2022021537A
Authentication system using organism information, and authentication device
WO2011061862A1