Authentication system, authentication method, and program
Through an identity authentication system that accepts high-resolution processing parameters, high-resolution images are generated and matched, the problem that low-resolution images cannot accurately restore the original image after super-resolution processing is solved, and high-accurate identity recognition is achieved.
Patent Information
- Application Number
- JP2023188572
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-16
AI Technical Summary
When low-resolution facial images are processed with super-resolution, the original image cannot be accurately restored, resulting in incorrect results from the identity authentication system and the relevant personnel may be missed.
By accepting parameters of high resolution processing, high resolution images are generated and matched to determine whether the person represented by the input image has corresponding records in the registered image.
It realizes high-accuracy identification of human identity and reduces the phenomenon of misidentification and missed searches.
Smart Images

Figure 2025076753000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an authentication technology for identifying a person from a captured image. [Background technology]
[0002] In recent years, authentication systems have been put into practical use that identify people by comparing the facial image of a person captured by a camera with facial images registered in a database in advance. Using such authentication systems, it is possible to authenticate who the person is who is captured in the video of a surveillance camera installed on the street. For example, it can be used in criminal investigations by comparing the facial image captured by the surveillance camera with a database that registers facial images of criminals.
[0003] On the other hand, face images captured by surveillance cameras installed on the street are not always of high quality. Such images are often low resolution or noisy, making them unsuitable for authentication systems due to degradation in image quality. From this perspective, a method has been proposed for matching face images by first increasing the resolution of low-resolution face images using a process called super-resolution.
[0004] Patent Document 1 discloses a method for generating a high-resolution face image from multiple face images captured continuously by a camera and performing individual recognition. In addition, due to recent advances in neural network technology, Non-Patent Document 1 discloses a method for generating a high-resolution face image with high accuracy from a single low-resolution face image. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2010-55300 A [Non-patent literature]
[0006] [Non-Patent Document 1] Yu et al., Ultra-resolving face images by discriminative generative networks. In Proceedings of European Conference on Computer Vision (ECCV) [Non-Patent Document 2] Yu et al., Super-Resolving Very Low-Resolution Face Images with Supplementary Attributes. 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition [Non-Patent Document 3] Liu et al., SphereFace: Deep Hypersphere Embedding for Face Recognition. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) [Non-Patent Document 4] Ian J, Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio, “Generative Adversarial Networks”, (2014). arXiv:1406.2661 [Non-Patent Document 5] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer, “High-Resolution Image Synthesis with Latent Diffusion Models”, (2021) arXiv:2112.10752 [Non-Patent Document 6] Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, Dani Lischinski, "StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery", (2021), arXiv:2103.17249 Summary of the Invention [Problem to be solved by the invention]
[0007] However, when super-resolution processing is applied to a low-resolution face image, the face image of the person may not be restored, and the face image may resemble that of a different person. If an authentication system performs matching on such a face image, it may lead to incorrect results. In addition, there is a risk that the person corresponding to the image captured by the camera may be overlooked.
[0008] Therefore, the present disclosure provides an authentication technique that can identify a person with high accuracy. [Means for solving the problem]
[0009] An authentication system according to one embodiment of the present disclosure is characterized in having a receiving means for receiving input of parameters for high-resolution processing, a generating means for generating a high-resolution image from an input image based on the parameters received by the receiving means, and a determining means for determining whether an image presumed to represent the same person as the person represented by the input image is registered by performing a matching process based on the high-resolution image generated by the generating means and a registered image that has been registered in advance. Effect of the Invention
[0010] According to the present disclosure, people can be identified with high accuracy. [Brief description of the drawings]
[0011] [Figure 1] FIG. 2 is a block diagram showing a hardware configuration of the authentication system. [Diagram 2] FIG. 2 is a block diagram showing a functional configuration of the authentication system. [Diagram 3] 11 is a flowchart showing a process performed by the authentication system to compare a face image captured by a camera with a database. [Figure 4] 4 shows an example of a user interface of a parameter setting unit. [Diagram 5] FIG. 2 is a block diagram showing a configuration of a high-resolution processing unit. [Figure 6] 13 is a flowchart showing a process of a high-resolution processing unit. [Figure 7] 4 is a block diagram showing a configuration of a face image selection unit. FIG. [Figure 8] 13 is a flowchart showing a process of a face image selection unit. [Figure 9] FIG. 2 is a block diagram showing a configuration of a face authentication unit. [Figure 10] 13 is a flowchart showing a process of a face authentication unit. [Figure 11] 13 is a diagram showing an example of display of a matching result by a matching result output unit. FIG. [Figure 12] FIG. 2 is a block diagram showing a functional configuration of the authentication system. [Figure 13] FIG. 2 is a block diagram showing a configuration of a face authentication unit. [Figure 14] FIG. 2 is a block diagram showing a functional configuration of the authentication system. [Figure 15] 13 is a block diagram showing a configuration of a matching result statistics unit. [Figure 16] 13 is a diagram showing an example of display of a matching result by a matching result output unit. FIG. [Figure 17] FIG. 1A is a block diagram showing the functional configuration of an authentication system, and FIG. 1B is a diagram showing an example of the flow of a learning method in the block diagram of FIG. [Figure 18] 18 is a flowchart showing a processing method in the block diagram of FIG. 17(B). [Figure 19] FIG. 2 is a block diagram showing a functional configuration of the authentication system. [Figure 20] FIG. 2 is a block diagram showing a functional configuration of the authentication system. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings. In the drawings, elements having the same configuration or function are given the same reference numerals, and repeated description thereof will be omitted. The configurations shown in the following embodiments are merely examples, and the present disclosure is not limited to the configurations shown in the drawings.
[0013] 1 is a block diagram showing the hardware configuration of an authentication system according to this embodiment. The authentication system according to this embodiment includes a processor 1, a storage device 2, an input device 3, and an output device 4. Each device is configured to be able to communicate with each other, and is connected by a bus or the like.
[0014] The arithmetic processing device 1 executes programs stored in the storage device 2 and controls the operation of the entire authentication system. The arithmetic processing device 1 is a CPU (Central Processing Unit).
[0015] The storage device 2 is a ROM (Read Only Memory), a RAM (Random Access Memory), or other storage device. The storage device 2 stores programs loaded based on the operation of the arithmetic processing device 1, data that can be stored for a long time, and the like, and temporarily stores the programs and data for the execution of processing. As the storage device 2, a magnetic storage device, a semiconductor memory, or the like is used.
[0016] The arithmetic processing device 1 is not limited to a CPU, and may be a PLD (Programmable Logic Device) such as an FPGA (Field Programmable Gate Array). Alternatively, the arithmetic processing device 1 may be an ASIC (Application Specific Integrated Circuit) or a DSP (Digital Signal Processor). The arithmetic processing device 1 may be composed of a CPU and a GPU (Graphics Processing Unit).
[0017] In this embodiment, the functions of the authentication system and the processes related to the flowcharts described below are realized by the arithmetic processing device 1 performing processes according to the procedures of the programs stored in the storage device 2. The storage device 2 stores images to be processed by the authentication system and the processing results.
[0018] The input device 3 is a mouse, a keyboard, a touch panel device, and / or a button, etc., and is used to input various instructions. The input device 3 also includes an imaging device such as a camera. The camera is typically a surveillance camera, but may be a general-purpose camera that is widely used. The output device 4 is a liquid crystal panel, an external monitor, etc., and outputs various information.
[0019] The hardware configuration of the entire authentication system is not limited to the above configuration. For example, the authentication system may include an I / O device for communicating between various devices. For example, the I / O device may be an input / output unit such as a memory card or a USB cable, or a wired or wireless transmission / reception unit.
[0020] The authentication system may also include a network interface that can connect to an external device via an intranet or a public network. In this case, the input device 3, a part of the storage device 2 (e.g., a large-capacity storage device having a database function, etc.), and / or the output device 4 may be connected to the bus and the arithmetic processing device 1 as the external device via the network interface.
[0021] (First embodiment) 2 is a block diagram showing the functional configuration of the authentication system according to the first embodiment. As shown in the figure, the processes and functions of the authentication system include a face image acquisition unit 110, a high-resolution processing unit 120, a parameter setting unit 130, a face image selection unit 140, a face authentication unit 150, a registered face database 160, and a matching result output unit 170.
[0022] The face image acquisition unit 110 acquires one input image, an input face image here, from a video (multiple frames in a time series) captured by a camera. For example, the face image acquisition unit 110 has a user interface that plays back a video captured by a camera and specifies a rectangular area surrounding the face via, for example, a mouse operation of the input device 3 within a frame showing the face of a person to be matched.
[0023] The high-resolution processing unit 120 processes the input face image acquired by the face image acquisition unit 110 based on the parameters set by the parameter setting unit 130 to generate and output a high-resolution image. Hereinafter, an image (face image) captured by a camera or an image (face image) specified and acquired by the face image acquisition unit 110 is referred to as an input image (input face image).
[0024] The parameter setting unit 130 sets parameters relating to attributes representing a person for processing a face image in the high resolution processing unit 120. The parameter setting unit 130 has a user interface for designating attributes of the face of a person to be matched via, for example, a mouse operation of the input device 3. That is, (the user interface of) the parameter setting unit 130 accepts parameter values input by the user via the input device 3.
[0025] The face image selection unit 140 selects two or more face images so as to increase the variance of facial features (features that represent an individual face) from the face images output from the high-resolution processing unit 120. In other words, the face image selection unit 140 selects two or more face images with significantly different facial features from the multiple face images output from the high-resolution processing unit 120.
[0026] The face authentication unit 150 executes a matching process between the face image output from the high-resolution processing unit 120 and face images stored in the registered face database 160, and outputs the matching result.
[0027] The registered face database 160 stores face images of people to be authenticated by the authentication system. The face images stored in the registered face database 160 are hereinafter referred to as registered face images. The registered face database 160 also stores information about people, such as face feature values, person IDs, genders, and ages, associated with the registered face images.
[0028] The matching result output section 170 outputs the matching result (matching determination result) obtained by the face authentication section 150. The matching result output unit 170 having an "output" function is typically a functional block that performs control (e.g., display control) to output information to an output device 4 such as a display device, and does not include the output device 4 which is hardware. However, the matching result output unit 170 may be a concept that includes the output device 4. The same applies to the matching result output unit 370 described later.
[0029] The process of comparing a face image captured by a camera with a database by the authentication system will be described below. Fig. 3 is a flowchart showing the process. Details of steps S300, S400, and S500 in Fig. 3 will be described later in Figs. 6, 8, and 10. The processes shown in Figs. 3, 6, 8, and 10 are realized by the arithmetic processing device 1 reading and executing a program stored in the storage device 2 according to the processing contents.
[0030] In S100, the face image acquisition unit 110 acquires a face image from a video captured by a camera. When a user of the authentication system designates a face area by operating a mouse, the face image acquisition unit 110 acquires an image of the designated area as a face image.
[0031] In S200, the parameter setting unit 130 sets parameters related to facial attributes. Fig. 4 shows an example of a user interface that the parameter setting unit 130 displays on the output device 4. The setting panel 30 is an interface for specifying facial attributes. In the example shown in the figure, the parameter setting unit 130 includes a check box for specifying gender, age, and slider bars for specifying eye and nose shape as facial attributes. The figure shows an example in which gender: male, age: 40 years old, eyes and nose are specified as slightly narrow.
[0032] The face attributes, which are parameters, are not limited to these four attributes, and other attributes such as face outline, skin color, hairstyle, and hair color may be set. In addition to the face attributes, attributes such as height, body type, clothing, and movement may be set. Alternatively, the parameter setting unit 130 may have a function of adding one or more other attributes or deleting an attribute (or selecting "not set" as described below) according to a user's input operation.
[0033] Even if a person is captured in the surveillance camera video at a distance and the facial image has low resolution, if the person's attributes can be determined from their clothing, movements, and surrounding circumstances, the input by the parameter setting unit 130 will be useful information for narrowing down the people. The facial attribute information set by the parameter setting unit 130 is normalized to a value between 0 and 1 for each attribute, and is output from the parameter setting unit 130 as a facial attribute vector a having dimensions equal to the number of attributes. If all attributes cannot be guessed, the user does not necessarily need to set all attributes in the parameter setting unit 130 (for example, a function that allows the user to select "not set" for each attribute).
[0034] In S300, the high resolution processing unit 120 processes the input face image acquired by the face image acquisition unit 110 based on the parameters set by the parameter setting unit 130, and generates and outputs at least one high resolution image.
[0035] 5 is a block diagram showing the configuration of the high-resolution processing unit 120. The high-resolution processing unit 120 includes a super-resolution processing unit 21 and a noise addition unit 22.
[0036] The super-resolution processing unit 21 generates a high-resolution face image (high-resolution face image) from the low-resolution face image to be processed and the face attribute vector a. In this embodiment, the neural network of Non-Patent Document 2 may be used as the super-resolution processing unit 21. This neural network generates a high-resolution face image from the low-resolution face image via an Encoder-Decoder network and several deconvolution layers that expand the spatial size. This neural network can generate a high-resolution image (in this embodiment, a high-resolution face image) based on the face attribute vector a by combining the feature amount extracted by the Encoder network with the face attribute vector a and performing processing after the Decoder network.
[0037] The noise adding unit 22 adds a predetermined amount of random noise to each dimension of the face attribute vector a. By adding random noise, the super-resolution processing unit 21 can randomly generate a plurality of face images with a wide range of features.
[0038] FIG. 6 is a flowchart showing the process of S300 by the high-resolution processing unit 120 shown in FIG. In S310, the super-resolution processing unit 21 acquires a low-resolution face image (low-resolution face image). In S320, the noise adding unit 22 acquires a face attribute vector a. In S330, the noise adding unit 22 adds a predetermined amount of random noise to each dimension of the face attribute vector a.
[0039] In S340, the super-resolution processing unit 21 generates a high-resolution face image from the face attribute vector a to which noise has been added by the noise addition unit 22 and the low-resolution face image acquired in S310. In S350, the super-resolution processing unit 21 outputs the generated face image.
[0040] In S360, the high resolution processing unit 120 repeats the processes of S330 to S350 a predetermined number of times with different face attribute vectors a. Then, the high resolution processing unit 120 outputs a plurality of high resolution face images generated with the different face attribute vectors a.
[0041] The high-resolution processing unit 120 does not have to output multiple high-resolution images, and may output only one. This is because when there is a lot of information that contributes to the restoration of a face image, such as when the image quality of the face image acquired by the face image acquisition unit 110 is relatively good or when there are a large number of parameters (attributes) set by the parameter setting unit 130, there is a sufficient possibility of successful matching with just one high-resolution image. Alternatively, the high-resolution processing unit 120 may generate multiple high-resolution images based on one parameter.
[0042] Next, in S400 (FIG. 3), the facial image selection unit 140 selects two or more facial images with large variance of facial features from the multiple facial images output by the high-resolution processing unit 120. In other words, the facial image selection unit 140 selects two or more facial images with significantly different facial features from the multiple facial images output by the high-resolution processing unit 120.
[0043] 7 is a block diagram showing the configuration of the facial image selection unit 140. The facial image selection unit 140 includes a facial image storage unit 41, a facial feature extraction unit 42, a facial feature storage unit 43, a clustering unit 44, and a representative image selection unit 45.
[0044] The face image storage unit 41 stores the high-resolution face image generated by the high-resolution processing unit 120.
[0045] The facial feature extraction unit 42 extracts features (facial features) for matching facial images as feature vectors of a predetermined dimension. In this embodiment, the extraction method may be a feature extraction method using a neural network called SphereFace in Non-Patent Document 3. The facial feature storage unit 43 stores the facial feature amounts extracted by the facial feature extraction unit .
[0046] The clustering unit 44 performs clustering of a plurality of facial feature amounts stored in the facial feature storage unit 43, and extracts the facial feature amount closest to the cluster center of each cluster as a representative vector. In this embodiment, the K-means method is used for the clustering unit 44, but a hierarchical method such as a shortest distance method may also be used. The number of representative vectors to be extracted corresponds to the number of facial images to be selected by the facial image selection unit 140, and is set in advance.
[0047] The representative image selection unit 45 selects a face image corresponding to the face feature amount of the representative vector extracted by the clustering unit 44 from the face image storage unit 41, and outputs it.
[0048] FIG. 8 is a flowchart showing the process of S400 by the facial image selection unit 140 shown in FIG. In S410, the facial feature extraction unit 42 acquires high-resolution facial images generated by the high-resolution processing unit 120. In addition, these high-resolution facial images are stored in the facial image storage unit 41. In S420, the facial feature extraction unit 42 extracts facial feature amounts from the facial images by a neural network. The extracted facial feature amounts are stored in the facial feature storage unit 43.
[0049] In S430, the facial feature extraction unit 42 repeats the processes of S410 and S420 for all facial images generated by the high-resolution processing unit 120 to obtain facial feature amounts. When facial feature amounts have been obtained for all facial images, in S440, the clustering unit 44 performs clustering using all facial feature amounts stored in the facial feature storage unit 43, and extracts a representative vector for each cluster.
[0050] In S450, the representative image selection unit 45 selects and outputs one or more face images corresponding to the face feature amounts of each representative vector extracted in S440 from the face image storage unit 41. As described above, a preset number of face images are selected.
[0051] By the above-described processing of the face image selection unit 140, two or more face images can be selected from a plurality of face images so as to increase the variance of the facial features.
[0052] Next, in S500 (FIG. 3), face authentication unit 150 compares the face image selected by face image selection unit 140 with the registered face images stored in registered face database 160, and outputs the comparison result.
[0053] 9 is a block diagram showing the configuration of face authentication unit 150. Face authentication unit 150 includes a facial feature extraction unit 51 and a face matching unit 52.
[0054] The facial feature extraction unit 51 extracts facial features for matching facial images as a feature vector of a predetermined dimension. In this embodiment, the facial feature extraction unit 51 has the same function as the facial feature extraction unit 42 of the facial image selection unit 140.
[0055] The face matching unit 52 calculates the similarity between the facial feature amount extracted by the facial feature extraction unit 51 and the facial feature amount of the face image stored in the registered face database 160, and outputs the matching result. As a method for calculating the similarity between both facial feature amounts, cosine similarity, which is a similarity calculation method of Non-Patent Document 3, may be used.
[0056] FIG. 10 is a flowchart showing the process of S500 by the face authentication unit 150 shown in FIG. In S510, the facial feature extraction unit 51 acquires the high-resolution facial image selected by the facial image selection unit 140. In S520, the facial feature extraction unit 51 extracts facial features from the facial image acquired in S510 by using a neural network.
[0057] In S530, the face matching unit 52 selects one of the registered faces stored in the registered face database 160, and acquires the corresponding facial feature amount. In S540, the face matching unit 52 calculates the similarity between the facial feature amount extracted in S520 and the facial feature amount acquired in S530.
[0058] In S550, the face authentication unit 150 repeats the processes of S530 and S540 for all registered face images stored in the registered face database 160. In S560, the face authentication unit 150 determines that the registered face image having the highest similarity degree, which is higher than the threshold value, among the similarities obtained for all registered faces, is a face image that is estimated to represent the same person as the person represented by the input face image (i.e., it is determined that they match). Then, the face authentication unit 150 outputs the matching result. This matching result is a person ID linked to the registered face image that is determined to match.
[0059] The processes of S510 to S560 shown in FIG.
[0060] Next, in S600 (FIG. 3), the matching result output unit 170 outputs the matching result obtained by the face authentication unit 150. The matching result is displayed, for example, as screen information output by the output device 4. FIG. 11 is a diagram showing an example of the display of the matching result.
[0061] As a result of the matching, a low-resolution face image 71 captured by a camera is displayed. Also, high-resolution face images 72, 73, and 74 generated by high-resolution processing unit 120 and selected by face image selection unit 140 are displayed. Also, registered face images 75, 76, and 77 obtained as a result of matching by face authentication unit 150 and corresponding to high-resolution face images 72 to 74, respectively, are displayed.
[0062] In addition, information about the person, such as the person ID, gender, age, etc., is displayed together with the matched registered face image. In the example of Fig. 11, a plurality of high-resolution face images 72, 73, and 74 and their corresponding matching results are displayed for one original low-resolution face image 71. For example, even if the high-resolution face image 72 obtained by the super-resolution processing (high-resolution processing) of S300 cannot restore (reproduce) the face image of the person, the matching results (registered face images 76, 77) with other high-resolution face images 73, 74 are also output, thereby increasing the probability of identifying the person.
[0063] The matching result (screen information showing the correspondence between the high-resolution face images 72 to 74 and the registered face images 75 to 77) is not limited to the example shown in Fig. 11. For example, a frame surrounding the two images, high-resolution face image 72 and registered face image 75, and a frame surrounding the two images, high-resolution face image 73 and registered face image 76, may be provided, and the whole may be displayed like a table. Alternatively, for example, a line or a similar image may be displayed connecting the two images, the high-resolution face image 72 and the registered face image 75. Alternatively, for example, characters, symbols, or other images common to the two images, the high-resolution face image 72 and the registered face image 75, may be displayed. Furthermore, text or symbols (such as ◯ or ×) indicating the matching result (the result of determining whether or not the input face image matches the registered face image), a numerical value indicating the degree of similarity, etc. may also be displayed.
[0064] As described above, in the process of generating a high-resolution image from a low-resolution face image, the high-resolution processing unit 120 of this embodiment can improve the accuracy of restoration by providing information that contributes to the restoration of the face image from the user via the parameter setting unit 130. As a result, the face authentication unit 150 can identify a person with high accuracy in the matching process, and can suppress overlooking a person corresponding to the face image captured by the camera. For example, if the user sets the parameter "gender: male" in parameter setting unit 130, the high-resolution processing unit 120 does not generate a female facial image such as facial image 73. In this way, when information about attributes is known in advance from the camera image, it is possible to perform face authentication with higher accuracy by setting parameters related to the facial attributes.
[0065] Furthermore, in this embodiment, the face image selection unit 140 selects one or more face images from multiple high-resolution face images so that the variance of features representing an individual face is large, and therefore it is possible to appropriately narrow down the high-resolution face images that can be displayed on the matching result output unit 170. This makes it possible to reduce redundancy in the face authentication results or to improve viewability.
[0066] Second embodiment In the above first embodiment, the face image selection unit 140 selects a face image from a plurality of high-resolution face images, and the face authentication unit 150 extracts face features from the selected face image and compares it with face images stored in the registered face database 160. However, the face image selection unit 140 has already extracted the face features of the high-resolution face image, resulting in duplication of processing. In the second embodiment, unlike the first embodiment, the face authentication unit 150 reuses the face features extracted by the face image selection unit 140 to compare the face features.
[0067] Fig. 12 is a block diagram showing the functional configuration of an authentication system according to the second embodiment. The processes and functions of this authentication system are realized by a face image acquisition unit 110, a high-resolution processing unit 120, a parameter setting unit 130, a face image selection unit 240, a face authentication unit 250, a registered face database 160, and a matching result output unit 170. In the blocks shown in Fig. 3, elements having the same functions as those in Fig. 1 are given the same reference numerals, and their explanations will be omitted.
[0068] The facial image selection unit 240 is the same as the facial image selection unit 140 of the first embodiment in that it selects a facial image from a plurality of high-resolution facial images output by the high-resolution processing unit 120 so as to increase the variance of facial features. The representative image selection unit 45 of this embodiment (FIG. 7) is the same as the first embodiment in that it outputs a facial image corresponding to the facial feature amount of the representative vector extracted by the clustering unit 44, but differs from the first embodiment in that it further outputs the facial feature amount of the representative vector.
[0069] The face authentication unit 250 compares the facial feature amount output by the face image selection unit 240 with the facial feature amount of a registered face image stored in the registered face database 160, and outputs the comparison result. Fig. 13 is a block diagram showing the configuration of the face authentication unit 250. The face matching unit 52 calculates the similarity between the facial feature amount and the facial feature amount of a registered face image stored in the registered face database 160, and outputs the matching result.
[0070] According to this embodiment, the processing by the facial feature extraction section 51 of the face authentication section 150 can be omitted, so that the processing load can be reduced and the processing speed can be increased in the face authentication process.
[0071] Third embodiment In each of the above embodiments, the facial image selection unit 140 (240) selects a facial image from a plurality of high-resolution facial images, and the facial authentication unit 150 (250) compares the selected facial image with the facial images stored in the registered face database 160. However, the facial image selection unit 140 (240) may be omitted.
[0072] Fig. 14 is a block diagram showing the functional configuration of an authentication system according to a third embodiment. In this authentication system, a face image selection unit is not provided, and a matching result statistics unit 150 is newly provided. The matching result statistics unit 150 outputs statistical information of the matching result output from the face authentication unit 150. In the blocks shown in Fig. 14, the same reference numerals are used for those having the same functions as those in Fig. 1, and the description thereof will be omitted.
[0073] 15 is a block diagram showing the configuration of the matching result statistics unit 350. The matching result statistics unit 350 compiles the matching results from the face authentication unit 150, selects a predetermined number of person IDs that match frequently as a matching result when, for example, matching processing is performed multiple times, and outputs the selected person IDs together with the frequency. "Matching" means, for example, that the similarity of the registered face image to the input face image is equal to or greater than a threshold value.
[0074] The match result statistics section 350 includes a person ID count section 351 and a person ID sort section 352 .
[0075] The person ID counting unit 351 receives the person ID that is the matching result for each input face image matched by the face authentication unit 150, and counts the frequency of each person ID. When the face authentication unit 150 has finished matching all face images, the person ID sorting unit 352 sorts the person IDs counted by the person ID counting unit 351 in descending order of frequency.
[0076] The matching result output unit 370 outputs the matching result collected by the matching result statistics unit 350. The matching result is displayed on the output device 4. FIG. 16 is a diagram showing an example of the display of the matching result. As the matching result, a low-resolution face image 371 captured by a camera and a graph 372 displaying, for example, up to three person IDs matched with the low-resolution face image 371 in descending order of frequency are displayed. Furthermore, registered face images 373, 374, and 375 corresponding to the matched person IDs are displayed in order. Screen information including the graph 372 and the information of the registered face images 373, 374, and 375 arranged in order is an example of screen information including information on the order of frequency.
[0077] In this embodiment, multiple high-resolution face images are generated from a low-resolution face image, and the results of matching performed by the face authentication unit 150 are tallied by the matching result statistics unit 350, so that frequently occurring matching results are displayed. This makes it easier for the user to view the matching results.
[0078] (Fourth embodiment) In each of the above embodiments, the high-resolution processing unit 120 and the face authentication unit 150 are configured as independent modules. In the authentication system according to the fourth embodiment, the high-resolution processing unit 120 and the face authentication unit 150 in the first embodiment and the like are linked to learn. That is, learning is performed by linking the generation of a high-resolution image by the high-resolution processing unit 120 and the matching and judgment by the face authentication unit 150. Under conditions where there are sufficient learning data and computational resources, the authentication system can improve the accuracy of face authentication by learning by linking multiple modules together rather than individually. Learning by linking multiple modules means [1] linking multiple modules to learn simultaneously, or [2] integrating multiple modules into one large deep neural network to learn.
[0079] As a specific example of the learning method, end-to-end learning disclosed in Non-Patent Document 4 and the like can be used. Non-Patent Document 4 discloses a method of simultaneously learning each layer of two networks by connecting two neural networks, an image generator and an image authenticity discriminator, and returning error signals sequentially by the error backpropagation method. Note that the authentication system according to this embodiment does not have a face image selection unit, as in the third embodiment.
[0080] Fig. 17(A) is a block diagram showing the functional configuration of an authentication system according to the fourth embodiment. In the blocks shown in Fig. 17(A) and Fig. 17(B), the same reference numerals are used to denote the same functions as in Fig. 1, and the description thereof will be omitted.
[0081] The face authentication unit 155 has a high-resolution feature extraction unit 50. The high-resolution feature extraction unit 50 is a module that connects two neural networks, the high-resolution processing unit 120 and the facial feature extraction unit 51 in the first embodiment. As initial values of the high-resolution feature extraction unit 50, the weights of the high-resolution processing unit 120 and the facial feature extraction unit 51 described in the first embodiment and the like are used.
[0082] Fig. 17(B) is a diagram showing an example of the flow of a learning method in the block diagram shown in Fig. 17(A). Fig. 18 is a flowchart showing the processing of this embodiment. In Fig. 17(B), the present authentication system has a function of learning using one high-resolution image as a model. Therefore, in Fig. 17(B), the authentication system trains a neural network so that the input face image can be restored after the low-resolution conversion unit 100 once reduces the model high-resolution image to generate a low-resolution image.
[0083] 18, in S710, the low-resolution conversion unit 100 acquires matched images of persons X, Y, and Z. In S720, the low-resolution conversion unit 100 converts the matched images into low-resolution images, and the face image acquisition unit 110 acquires these low-resolution images as input face images.
[0084] In S730, the registered face database 160 reads out facial features of the registered face image of person X among persons X, Y, and Z, and attribute information of person X (e.g., Caucasian elderly male, etc.). In S740, the parameter setting unit 130 sets the attribute information of person X as a high-resolution parameter.
[0085] At S750, the high-resolution feature extraction unit 50 increases the resolution of the images of the persons X, Y, and Z based on the attribute information of the person X, which is a set parameter. At S760, the high-resolution feature extraction unit 50 extracts face features to be matched from the high-resolution face images. At S770, the face matching unit 52 calculates the similarity between the features of the registered face image of the person X and the face features to be matched of the persons X, Y, and Z.
[0086] In S780, the face authentication unit 155 performs weight adjustment on the matching result (calculation result of each similarity). The face authentication unit 155 performs the following process as the weight adjustment. First, the face authentication unit 155 compares each similarity calculated in S770 with a true value. Currently, the processing target is the similarity of the facial feature amount of person X, so the true value of person X is 1.0, and the true values of persons Y and Z are 0.0. The face authentication unit 155 adjusts the weight of each layer of the neural network of the high-resolution feature extraction unit 50 by the error backpropagation method so that each similarity approaches the corresponding true value. That is, the face authentication unit 155 executes the process of S760 to S780 using an error signal for the true value so that each similarity approaches the corresponding true value. Specifically, the weight adjustment is performed so that the similarity of the feature amount between the registered face image of person X and the matched image of person X becomes higher, and the similarity between the registered face image of person X and the matched images of persons Y and Z becomes lower.
[0087] In S790, the registered person is changed in sequence from person X to person Y, person Z, . . . and S730 to S780 are repeated.
[0088] It should be noted that the learning in this embodiment is not to improve the image quality (make the image clearer), but to concentrate the neural network resources on learning to restore features that are effective for identifying people from low-quality images. This is expected to improve the accuracy of the authentication system.
[0089] In the above description, the learning method [1] is taken as an example, but the method [2] may also be used. That is, it is possible to replace the two neural networks of the high-resolution processing unit 120 and the face authentication unit 150 in the first embodiment with one large deep neural network, and to provide true values as teaching values and learn from random number initial values. In this way, the implementation form is not limited to a specific form.
[0090] Fifth embodiment In each of the above embodiments, an example has been described in which candidates for facial attributes are presented in advance, and parameters (facial attributes) for high-resolution processing are set via a user interface that includes radio buttons and sliders for specifying values for those attributes, as shown in FIG. 4, for example.
[0091] Meanwhile, in recent years, there has been a technology for generating or modifying an image according to instructions in natural language, as disclosed in Non-Patent Documents 5 and 6. In the fifth embodiment, an example of a flexible face search system with a natural language interface will be described using a task of searching for a person in a video captured by a camera (e.g., a surveillance camera) as an example.
[0092] Fig. 19 is a block diagram showing the functional configuration of an authentication system according to the fifth embodiment. In the blocks shown in Fig. 19, elements having the same functions as those in Fig. 1 are given the same reference numerals, and the description thereof will be omitted.
[0093] The image acquisition unit 115 acquires, for example, a facial image of a target person, a whole-body image including a facial image, an image of a part of the body including a facial image, or an image of a part of the body not including a facial image from one or more frames of a video captured by a camera.
[0094] The parameter setting unit 135 has a language prompt input unit 1301 and an attribute feature conversion unit 1302. Note that the function of the noise addition unit 1303 can be used in another method that will be described later. For example, a prompt in a natural language is input as a parameter to the language prompt input unit 1301. The language prompt input unit 1301 receives linguistic information about a target person, such as “a boy in a lower grade on his way home from a soccer game,” from the user via the input device 3. The attribute feature conversion unit 1302 uses the methods of Non-Patent Document 5 and Non-Patent Document 6 to generate, from the sentence received by the language prompt input unit 1301, a feature vector in which the meaning of the sentence is embedded.
[0095] The high-resolution processing unit 125 converts the person image captured by the surveillance camera into a high-resolution image based on the feature vector generated by the attribute feature conversion unit 1302. The high-resolution image generated by the high-resolution processing unit 125 includes a face image of the target person, a whole-body image including a face image, an image of a part of the body including a face image, or an image of a part of the body not including a face image. Furthermore, the image generated by the high-resolution processing unit 125 is gently oriented to match the age indicated by <lower grades> in the text. The clothing is oriented to match the clothing and belongings one might wear on the way home from a soccer game. The high-resolution processing unit 125 can adjust the weights to determine how much importance is attached to these constraints, and the methods disclosed in Non-Patent Documents 5 and 6 can be used for this purpose.
[0096] A face image of a search target, a whole-body image including a face image, or a part of the face image is registered in the registered image database 165. In this case, a search target is assumed to be, for example, a lost child or a missing person.
[0097] The authentication unit 450 compares the features (facial features and / or features of the whole body (or part of the body)) of the generated high-resolution image with those of the person to be searched for who is registered in the registered image database 165, and outputs one or more images of the person who has a high similarity. Specifically, the feature extraction unit 451 extracts features from the high-resolution image (not limited to facial images). Then, the matching unit 452 executes matching processing by determining the similarity between the features and features of an image (not limited to facial images) stored in the registered database 165. The matching result output unit 170 (FIG. 1), not shown here, can display, for example, a ranking of the one or more images output from the authentication unit 450.
[0098] Alternatively, the high-resolution processing unit 125 may create multiple variations of the high-resolution image by adding perturbations to the attribute features in the noise addition unit 1303. The multiple high-resolution images thus obtained may be collated, and for example, a person image with a high average similarity may be displayed as a top candidate. As a method for determining the degree of consistency of the results in this way, multiple statistical and tabulation methods may be used. In addition, when multiple sentences are input to the language prompt input unit 1301, multiple high-resolution images may be generated according to which sentence is emphasized. For example, when two sentences, "a boy in a lower grade returning from a soccer game" and "a boy with a good physique and an active personality," are input, multiple high-resolution images may be generated, including a high-resolution image that is made high-resolution by emphasizing the first sentence and a high-resolution image that is made high-resolution by emphasizing the second sentence. In addition, when a prompt, "a boy in a lower grade returning from a soccer game," is input, multiple high-resolution images may be generated, including a high-resolution image that is made high-resolution by emphasizing the "boy in a lower grade returning from a soccer game" and a high-resolution image that is made high-resolution by emphasizing the "boy in a lower grade." As yet another method, multiple high-resolution images may be generated according to the strength of constraint based on the natural language input by the language prompt input unit 1301. For example, when a prompt such as "a boy in a lower grade on his way home from a soccer game" is input, multiple high-resolution images including a high-resolution image that strongly reflects the prompt and a high-resolution image that loosely reflects the prompt may be generated.
[0099] Other additional functions for displaying the results may also be considered. For example, the noise components and feature vectors of the most similar matched image may be used to increase the resolution of all other time-series images of the person in question and display them for confirmation. In this case, if the person obtained in the matching result is the correct person, the image is restored as a natural time-series image, but if the person is incorrect, the matching result may be inconsistent and the time-series image may appear unstable.
[0100] The processing procedure of this embodiment can be explained with reference to FIG. 3. In S100 of FIG. 3, the image acquisition unit 115 acquires a face image, a whole-body image including a face image, an image of a part of the body including a face image, or an image of a part of the body not including a face image of the target person from one or more frames of a moving image. In S200, the language prompt input unit 1301 accepts an input of a prompt in natural language as a parameter. In S300, the high-resolution processing unit 125 processes the input image based on a parameter corresponding to the natural language prompt to generate at least one high-resolution image. S400 may be skipped, but a functional block similar to the face image selection unit 140 may be provided between the high-resolution processing unit 125 and the authentication unit 450 so that two or more high-resolution images are selected from a plurality of high-resolution images. In S500, the authentication unit 450 compares the feature amount (face feature amount and / or whole-body (or part of body) feature amount) of the generated high-resolution image with that of the search target person registered in the registered image database 165, and outputs one or more images of the person with a high similarity. Then, in S600, the collation unit 452 outputs the collation result.
[0101] (Other embodiments) For example, the authentication systems shown in Fig. 14 and Fig. 17(A) show examples in which the facial image selection units 140, 240 shown in Fig. 2 etc. are not provided. Similarly, for example, the authentication systems according to the first and second embodiments described above may not have a facial image selection unit. Fig. 20 shows the functional configuration of such an authentication system. In an authentication system configured in this way, the high-resolution processing unit 120 may generate and output one high-resolution facial image, but it is not necessarily required to output only one. A plurality of high-resolution facial images may be output, and the facial authentication unit 150 may perform facial authentication on each of them.
[0102] In the above first to fourth embodiments, the high-resolution processing unit 120 mainly generates face images as high-resolution images. However, in the first to fourth embodiments, the image is not limited to a face image, and similar to the fifth embodiment, a face image of the target person, a whole-body image including a face image, an image of a part of the body including a face image, or an image of a part of the body not including a face image may be generated. In this case, similar to the fifth embodiment, a face image of the target person, a whole-body image including a face image, or a part thereof is registered in the registered image database, and the authentication unit performs a matching process on the face image, the whole-body image including a face image, etc.
[0103] In the above embodiments, an example has been described in which one face image is obtained from a video (multiple frames) captured by a camera and face recognition is performed. However, a configuration may be adopted in which a user specifies and obtains multiple face images of the same person from the multiple frames and performs face recognition. In this case, in the first and second embodiments, the high-resolution processing unit 120 may generate multiple high-resolution face images from each of the input face images acquired from multiple frames of the video, and the face image selection units 140 and 240 may perform clustering on all of the high-resolution face images. Furthermore, in the third embodiment, the high-resolution processing unit 120 generates multiple high-resolution face images for each of the face images acquired from multiple frames, and the matching result statistics unit 350 is configured to output statistics for the matching results of all the high-resolution face images. Alternatively, the user may specify a face image from one of the multiple frames and specify and acquire other input images of the person from the other frames, which may be a face image of the target person, a whole-body image including a face image, an image of a part of the body including a face image, or an image of a part of the body not including a face image. Alternatively, the user may specify a face image from one frame, and the face image acquisition unit may automatically specify face images from other frames by tracking processing.
[0104] As described above, when a user specifies a face image from one of the multiple frames and specifies and acquires other input images of the person from the other frames, the authentication system may perform the following process. For example, the high-resolution processing unit 120 or the like may further generate a high-resolution image from other input images of the person specified from other frames based on parameters used when generating a high-resolution image from an input image corresponding to the image specified from one frame. The other input images are as described above.
[0105] The matching result statistics section 350 (FIG. 14) in the third embodiment may be provided in the authentication systems in the first, second, fourth and fifth embodiments.
[0106] In the above embodiments, the image captured by the camera or the image acquired by the face image acquisition unit 110 has been described as being a low-resolution image. However, even if the image captured by the camera or the image acquired by the face image acquisition unit 110 is a high-resolution face image, there may be cases where the face image is noisy and has a lot of blur, as captured in a dark scene. Such cases are also included in the scope of the present disclosure. In such cases, the authentication system may temporarily reduce the resolution of the acquired face image by image reduction processing, and then perform subsequent processing. This allows the high-resolution processing unit 120 and the like to perform high-resolution processing on the face image with reduced noise.
[0107] In each of the above embodiments, an example has been given in which the user specifies a rectangular area surrounding the face and other body parts via a mouse operation, by the face image acquisition unit 110. However, the face image acquisition unit 110 is not limited to such a form, and may automatically specify a rectangular area surrounding the face and other body parts from an input image. In addition, in each of the above embodiments, the description has been given mainly on the example where the registered image (such as a facial image of the search target, a whole-body image including a facial image, or a part thereof) is registered in the registered face database 160 or the registered image database 165, but the present invention is not limited to this. For example, when searching for a lost child, a facial photo of the search target (lost child) may be input to the authentication system as a query image, and a matching process may be performed based on the facial photo and a high-resolution image generated from an image captured by a surveillance camera.
[0108] It is also possible to combine at least two of all the embodiments described above.
[0109] The present disclosure can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) for implementing one or more of the functions. It should be noted that the above-described embodiments are merely examples of the implementation of the present disclosure, and the technical scope of the present disclosure should not be interpreted as being limited by these embodiments. In other words, the present disclosure can be implemented in various forms without departing from its technical concept or main features.
[0110] The disclosure of this embodiment includes the following configuration, method, and program. (Configuration 1) A reception means for receiving an input of parameters for high resolution processing; a generating means for generating a high-resolution image from an input image based on the parameters received by the receiving means; a determination means for determining whether or not an image that is estimated to represent the same person as the person represented by the input image is registered by executing a matching process based on the high-resolution image generated by the generation means and a registered image that has been registered in advance; An authentication system comprising: (Configuration 2) The generating means generates the high-resolution image based on a plurality of the parameters. 2. An authentication system according to configuration 1. (Configuration 3) an output means for outputting the high-resolution image generated by the generating means and the determination result by the determining means; 3. The authentication system according to claim 1 or 2, further comprising: (Configuration 4) The output means outputs, as the determination result, screen information indicating that the high-resolution image corresponds to the registered image that is the image presumed to be of the same person. 4. An authentication system according to configuration 3. (Configuration 5) The determination means performs the determination based on the similarity of the feature amounts of the high-resolution image and the registered image. 5. The authentication system according to any one of configurations 1 to 4. (Configuration 6) The generating means generates a plurality of high resolution images; a selection means for selecting one or more high-resolution images from the plurality of high-resolution images generated by the generation means so that the variance of human features is large. 6. The authentication system according to any one of configurations 1 to 5, further comprising: (Configuration 7) the selection means extracts feature amounts of the plurality of high-resolution images generated by the generation means, and selects the one or more high-resolution images based on the extracted feature amounts; The determining means performs a determination based on each of the feature amounts extracted by the selecting means. 7. The authentication system according to configuration 6. (Configuration 8) A statistical means for outputting statistical information on the result of the judgment by the judging means. 8. The authentication system according to any one of claims 1 to 7, further comprising: (Configuration 9) The statistical means outputs screen information including information on the order of frequency at which the determining means determines that the images presumed to represent the same person are registered as the statistical information. 9. The authentication system according to configuration 8. (Configuration 10) The generation of the high-resolution image by the generating means and the comparison and judgment by the judging means are linked to carry out learning. 10. The authentication system according to any one of configurations 1 to 9. (Configuration 11) The parameter may be one or more attributes describing a person or a natural language. 11. The authentication system according to any one of configurations 1 to 10. (Configuration 12) A conversion means for converting the input image into a low resolution image, The generating means generates the high-resolution image from the low-resolution image converted by the converting means. 12. The authentication system according to any one of configurations 1 to 11. (Configuration 13) The generating means generates the high-resolution image from each of input images of the same person acquired from a plurality of frames constituting a moving image. 13. The authentication system according to any one of configurations 1 to 12. (Configuration 14) The generating means further generates a high-resolution image from another input image of the person based on the parameter used when the determining means determines that the image presumed to be of the same person is registered. 14. The authentication system according to any one of configurations 1 to 13. (Configuration 15) The generating means generates a high-resolution face image from an input face image as the input image, The determination means executes the matching process based on the generated high-resolution face image and a registered face image as the registered image. 15. An authentication system according to any one of configurations 1 to 14. (method) a receiving step of receiving an input of parameters for high resolution processing; a generating step of generating a high-resolution image from an input image based on the parameters received in the receiving step; a determination step of determining whether or not an image that is estimated to represent the same person as the person represented by the input image is registered by executing a matching process based on the high-resolution image generated by the generation step and a registered image that has been registered in advance; 13. An authentication method comprising: (program) A program for causing a computer to operate as the authentication system according to any one of configurations 1 to 15.
[0111] The present disclosure has been described above in detail based on preferred embodiments thereof, but the present disclosure is not limited to the above embodiments, and various modifications are possible based on the gist of the present disclosure, and these modifications are not excluded from the scope of the present disclosure. [Explanation of symbols]
[0112] 42, 51: Facial feature extraction section 50: High-resolution feature extraction unit 72, 73, 74: High-resolution facial images 75, 76, 77: Registered face image 100: Low resolution conversion section 110: Face image acquisition unit 120, 125: High resolution processing section 130, 135: Parameter setting section 140, 240: Face image selection section 150, 155, 250: Face recognition section 160: Registered face database 165:Registered image database 170, 370: Matching result output section 350: Matching result statistics section 450: Authentication section 1301: Language prompt input section
Claims
1. A reception means for receiving an input of parameters for high resolution processing; a generating means for generating a high-resolution image from an input image based on the parameters received by the receiving means; a determination means for determining whether or not an image that is estimated to represent the same person as the person represented by the input image is registered by executing a matching process based on the high-resolution image generated by the generation means and a registered image that has been registered in advance; An authentication system comprising:
2. The generating means generates the high-resolution image based on a plurality of the parameters.
2. The authentication system according to claim 1 .
3. an output means for outputting the high-resolution image generated by the generating means and the determination result by the determining means; 2. The authentication system of claim 1, further comprising:
4. The output means outputs, as the determination result, screen information indicating that the high-resolution image corresponds to the registered image that is the image presumed to be of the same person.
4. The authentication system according to claim 3.
5. The determination means performs the determination based on the similarity of the feature amounts of the high-resolution image and the registered image.
2. The authentication system according to claim 1 .
6. The generating means generates a plurality of high resolution images; a selection means for selecting one or more high-resolution images from the plurality of high-resolution images generated by the generation means so that the variance of human features is large.
2. The authentication system of claim 1, further comprising:
7. the selection means extracts feature amounts of the plurality of high-resolution images generated by the generation means, and selects the one or more high-resolution images based on the extracted feature amounts; The determining means performs a determination based on each of the feature amounts extracted by the selecting means.
7. The authentication system according to claim 6.
8. A statistical means for outputting statistical information on the result of the judgment by the judging means.
2. The authentication system of claim 1, further comprising:
9. The statistical means outputs screen information including information on an order of frequency at which the determining means determines that the images presumed to represent the same person are registered as the statistical information.
9. The authentication system according to claim 8.
10. The generation of the high-resolution image by the generating means and the comparison and judgment by the judging means are linked to carry out learning.
2. The authentication system according to claim 1 .
11. The parameters may be one or more attributes describing a person or a natural language.
2. The authentication system according to claim 1 .
12. A conversion means for converting the input image into a low resolution image, The generating means generates the high-resolution image from the low-resolution image converted by the converting means.
2. The authentication system according to claim 1 .
13. The generating means generates the high-resolution image from each of input images of the same person acquired from a plurality of frames constituting a moving image.
2. The authentication system according to claim 1 .
14. The generating means further generates a high-resolution image from another input image of the person based on the parameter used when the determining means determines that the image presumed to be of the same person is registered.
2. The authentication system according to claim 1 .
15. The generating means generates a high-resolution face image from an input face image as the input image, The determination means executes the matching process based on the generated high-resolution face image and a registered face image as the registered image.
2. The authentication system according to claim 1 .
16. a receiving step of receiving an input of parameters for high resolution processing; a generating step of generating a high-resolution image from an input image based on the parameters received in the receiving step; a determination step of determining whether or not an image that is estimated to represent the same person as the person represented by the input image is registered by executing a matching process based on the high-resolution image generated by the generation step and a registered image that has been registered in advance; 13. An authentication method comprising:
17. A program for causing a computer to operate as the authentication system according to any one of claims 1 to 15.
Citation Information
Patent Citations
Image processing apparatus and image processing method
JP2010055300A