Information processing device, information processing method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2023-08-22
- Publication Date
- 2026-05-11
Abstract
Description
Information processing device, information processing system, information processing method, and recording medium
[0001] The present disclosure relates to an information processing device, an information processing system, an information processing method, and a recording medium.
[0002] For example, Patent Document 1 discloses a normalization technique for deep neural networks.
[0003] In the deep neural network normalization method described in Patent Literature 1, an input data set is input to a deep neural network, and the normalization method includes normalizing a feature map set output by a network layer in the deep neural network in at least one dimension to obtain a variance of at least one dimension and a mean value of at least one dimension.
[0004] The feature map set includes at least one feature map corresponding to at least one channel, and each channel corresponds to at least one feature map. The normalized target feature map set is identified based on the variance of at least one dimension and the mean value of at least one dimension.
[0005] According to the description in Patent Document 1, normalization along at least one dimension incorporates statistical information of each dimension through the normalization operation, ensuring high robustness to the statistics of each dimension while not being overly dependent on batch size.
[0006] Special table 2020-537204 publication
[0007] The present disclosure aims to improve upon the techniques described in the prior art documents mentioned above.
[0008] The information processing device of the present disclosure includes: a target information acquisition means for acquiring target information regarding a target region in an input image; a calculation means for calculating, based on the target information, correction parameters for correcting parameters included in a feature extraction model that extracts features of the target region; an extraction means for extracting features of the target region using the feature extraction model corrected with the correction parameters; and a comparison means for outputting a result of comparing the features of the target region with pre-registered registration information.
[0009] The information processing system of the present disclosure includes a target information acquisition means for acquiring target information related to a target region in an input image; a calculation means for calculating, based on the target information, correction parameters for correcting parameters included in a feature extraction model that extracts features of the target region; an extraction means for extracting features of the target region using the feature extraction model corrected with the correction parameters; and a comparison means for outputting a result of comparing the features of the target region with pre-registered registration information.
[0010] The information processing method of the present disclosure includes one or more computers acquiring target information regarding a target region in an input image, calculating correction parameters for correcting parameters included in a feature extraction model that extracts features of the target region based on the target information, extracting features of the target region using the feature extraction model corrected with the correction parameters, and outputting a result of comparing the features of the target region with pre-registered registration information.
[0011] The recording medium in the present disclosure has recorded thereon a program for causing one or more computers to acquire target information regarding a target region in an input image, calculate correction parameters for correcting parameters included in a feature extraction model that extracts features of the target region based on the target information, extract features of the target region using the feature extraction model corrected with the correction parameters, and output a result of comparing the features of the target region with pre-registered registration information.
[0012] 1 is a block diagram showing a configuration of a first information processing device according to the present disclosure. FIG. 1 is a block diagram showing a configuration of a first information processing system according to the present disclosure. FIG. 2 is a flowchart showing a processing operation of the first information processing device according to the present disclosure. FIG. 3 is a diagram showing an example of a binocular image as an input image according to the present disclosure. FIG. 4 is a diagram showing an example of an iris region as a target region according to the present disclosure. FIG. 5 is a diagram showing an example of a physical configuration of a first information processing device according to the present disclosure. FIG. 6 is a block diagram showing a configuration of a second information processing system and a second information processing device according to the present disclosure. FIG. 7 is a diagram showing an example of a rectangular iris region as a target region converted into a predetermined shape according to the present disclosure. FIG. 8 is a flowchart showing a processing operation of the second information processing device according to the present disclosure. FIG. 9 is a block diagram showing a configuration of a third information processing system and a third information processing device according to the present disclosure. FIG. 10 is a flowchart showing a processing operation of the third information processing device according to the present disclosure. FIG. 11 is a block diagram showing a configuration of a fourth information processing system and a fourth information processing device according to the present disclosure. FIG. 12 is a block diagram showing an example of a configuration of a fourth calculation unit according to the present disclosure. FIG. 13 is a flowchart showing a processing operation of the fourth information processing device according to the present disclosure. FIG. 14 is a flowchart showing an example of a correction parameter calculation process included in the processing operation of the fourth information processing device according to the present disclosure. FIG. 15 is a block diagram showing a configuration of a fifth information processing system and a fifth information processing device according to the present disclosure. FIG. 16 is a block diagram showing an example of a configuration of a fifth calculation unit according to the present disclosure. FIG. 17 is a block diagram showing a configuration of a sixth information processing system and a sixth information processing device according to the present disclosure. FIG. 10 is a block diagram showing the configuration of a learning unit included in a sixth information processing device according to the present disclosure. FIG. 11 is a flowchart showing an example of a learning process performed by the sixth information processing device according to the present disclosure. FIG. 12 is a block diagram showing the configurations of a seventh information processing system and a seventh information processing device according to the present disclosure. FIG. 13 is a block diagram showing the configuration of a first learning unit included in the seventh information processing device according to the present disclosure. FIG. 14 is a block diagram showing the configuration of a second learning unit included in the seventh information processing device according to the present disclosure. FIG. 15 is a flowchart showing an example of a first learning process performed by the seventh information processing device according to the present disclosure. FIG. 16 is a flowchart showing an example of a second learning process performed by the seventh information processing device according to the present disclosure. FIG. 17 is a block diagram showing the configurations of an eighth information processing system and an eighth information processing device according to the present disclosure.10 is a block diagram showing the configuration of a first learning unit included in an eighth information processing device according to the present disclosure. FIG. 11 is a block diagram showing the configuration of a second learning unit included in the eighth information processing device according to the present disclosure. FIG. 12 is a flowchart showing an example of a first learning process executed by the eighth information processing device according to the present disclosure. FIG. 13 is a flowchart showing an example of a second learning process executed by the eighth information processing device according to the present disclosure. FIG. 14 is a block diagram showing the configuration of a ninth information processing system and a ninth information processing device according to the present disclosure. FIG. 15 is a block diagram showing the configuration of a learning unit included in the ninth information processing device according to the present disclosure. FIG. 16 is a flowchart showing an example of a learning process executed by the ninth information processing device according to the present disclosure. FIG. 17 is a block diagram showing the configuration of a tenth information processing system and a tenth information processing device according to the present disclosure. FIG. 18 is a block diagram showing the configuration of a first learning unit included in the tenth information processing device according to the present disclosure. FIG. 19 is a block diagram showing the configuration of a second learning unit included in the tenth information processing device according to the present disclosure. FIG. 19 is a flowchart showing an example of a first learning process executed by the tenth information processing device according to the present disclosure. FIG. 19 is a flowchart showing an example of a second learning process executed by the tenth information processing device according to the present disclosure.
[0013] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that in this disclosure, the drawings relate to one or more embodiments. In all drawings, similar components are designated by similar reference numerals, and descriptions thereof will be omitted as appropriate.
[0014] [First Embodiment] (Summary) For example, in iris authentication, generally, if the quality of the image used for iris authentication, particularly the iris region, is low, such as if the image resolution is low, authentication accuracy may decrease. Therefore, in iris authentication, it is desirable to photograph the authentication target using an imaging device capable of capturing high-quality images, such as high resolution, but it may be difficult to use an imaging device capable of capturing high-quality images due to installation space, cost, etc. of the imaging device.
[0015] Thus, when performing processing using images, high-quality images such as high resolution are required, but it can be difficult to obtain high-quality images.
[0016] The normalization technique described in Patent Document 1 may be able to obtain high robustness against statistics of each dimension without excessive dependency on batch size. However, even if the normalization technique described in Patent Document 1 is used, it is difficult to obtain image features that improve authentication accuracy from low-quality input images.
[0017] Furthermore, as a technique for improving image quality, there is, for example, an image super-resolution technique for improving image resolution (see, for example, Japanese Patent Application Laid-Open No. 2009-282925). However, image super-resolution techniques generally often require large processing costs, such as the time required for processing and memory consumption.
[0018] One of the problems that the invention according to the present disclosure aims to solve is to perform authentication with high accuracy while suppressing increases in processing costs.
[0019] (Configuration example of information processing device 100 and information processing system SYS1) As shown in Fig. 1, the information processing device 100 includes a target information acquisition unit 115, a calculation unit 116, an extraction unit 117, and a collation unit 118. As shown in Fig. 2, the information processing system SYS1 includes a target information acquisition unit 115, a calculation unit 116, an extraction unit 117, and a collation unit 118.
[0020] The target information acquisition unit 115 acquires target information relating to a target region in an input image.
[0021] The calculation unit 116 calculates, based on the target information, correction parameters for correcting parameters included in a feature amount extraction model that extracts feature amounts of the target region.
[0022] The extraction unit 117 extracts features of the target region using the feature extraction model corrected with the correction parameters.
[0023] The matching unit 118 outputs the result of matching the feature amount of the target region with pre-registered registration information.
[0024] (Operations and Effects) According to the information processing device 100, it is possible to extract features of a target region using a feature extraction model corrected with correction parameters based on target information relating to the target region.
[0025] Therefore, even if the input image is low quality, it is possible to extract the feature values of the target region with high accuracy and obtain matching results using the feature values. Furthermore, the processing cost for extracting such accurate feature values of the target region is lower than when applying general image super-resolution technology.
[0026] Therefore, it is possible to suppress an increase in processing costs and perform authentication with high accuracy.
[0027] According to this information processing system SYS1, similar to the information processing device 100, it is possible to suppress an increase in processing costs and perform authentication with high accuracy.
[0028] (Example of Processing Operation of Information Processing Device 100) The information processing device 100 executes information processing as shown in FIG.
[0029] The target information acquisition unit 115 acquires target information relating to the target region in the input image (step S115).
[0030] The calculation unit 116 calculates correction parameters for correcting parameters included in a feature amount extraction model that extracts feature amounts of the target region, based on the target information (step S116).
[0031] The extraction unit 117 extracts features of the target region using the feature extraction model corrected with the correction parameters (step S117).
[0032] The matching unit 118 outputs the result of matching the feature amount of the target region with the registered information that has been registered in advance (step S118).
[0033] According to this information processing, it is possible to extract features of a target region using a feature extraction model corrected with correction parameters based on target information relating to the target region.
[0034] Therefore, even if the input image is low quality, it is possible to extract the feature values of the target region with high accuracy and obtain matching results using the feature values. Furthermore, the processing cost for extracting such accurate feature values of the target region is lower than when applying general image super-resolution technology.
[0035] Therefore, it is possible to suppress an increase in processing costs and perform authentication with high accuracy.
[0036] (Detailed Example) Hereinafter, detailed examples of the information processing device 100, the information processing system SYS1, and information processing will be described. Note that the present disclosure may be realized by a program for causing one or more computers to execute information processing, a recording medium on which the program is recorded, or the like.
[0037] (Input Image and Target Region) The input image is an image that includes a target region to be used for matching. In other words, the target region is an image of the region of the input image that is to be used for matching.
[0038] The input image is, for example, a binocular image including both eyes, a monocular image including one predetermined eye of the left and right eyes, a facial image including a face, or a biometric image such as a vein image including veins.
[0039] The target region is, for example, an iris region (i.e., an image of an area showing an iris), a face region (i.e., an image of an area showing a face), a fingerprint region (i.e., an image of an area showing a fingerprint), a vein region (i.e., an image of an area showing veins), etc. When the target region is an iris region, a face region, a fingerprint region, or a vein region, iris authentication, face authentication, fingerprint authentication, or vein authentication can be performed, respectively.
[0040] Note that iris authentication, face authentication, and vein authentication are examples of biometric authentication. Authentication according to the present disclosure is not limited to biometric authentication. Therefore, the input image is not limited to a biometric image. Furthermore, the iris, face, etc. are typically those of a human, but may also be those of an animal such as a dog or snake.
[0041] In the embodiment according to the present disclosure, an example will be described in which the input image is a binocular image including both eyes of a person (see FIG. 4 ), and the target region is an iris region (see FIG. 5 ).
[0042] (Regarding Target Information) The target information is information relating to a target region. The target information is information including at least one value according to the quality of the target region in the input image. The value included in the target information is, for example, a continuous value.
[0043] In detail, the target information may include, for example, at least one of quality information of the target region, intermediate feature amounts of the target region, and statistics related to the target region.
[0044] The quality information of the target region is information indicating the quality of the target region. For example, the quality information of the target region may include one or more of the following: resolution of the target region, focus blur when the input image is captured by a photographing device such as a camera, motion blur, brightness of the target region, information on eyeglass reflection, etc. The resolution may be, for example, the iris diameter when the input image includes an iris.
[0045] The iris diameter may be, for example, the diameter of the iris or the radius of the iris. The iris diameter may be expressed, for example, by the number of pixels. Furthermore, for example, if the iris image is elliptical, the iris diameter may be either the minor axis or the major axis of the iris, or both. Note that the method of expressing the iris diameter is not limited to the number of pixels, and may be, for example, a value according to the size of the iris region.
[0046] The information about eyeglasses reflection may include, for example, at least one of the intensity, area, and proportion of the area of eyeglasses reflection. The intensity of eyeglasses reflection may be, but is not limited to, the average value, maximum value, etc. of the intensity of eyeglasses reflection included in the target region. The area or proportion of eyeglasses reflection may be, but is not limited to, the area or proportion of the region of the iris region that includes eyeglasses reflection. Note that the quality information is not limited to the examples given here.
[0047] When the target information includes quality information of the target region, the target information acquisition unit 115 may acquire the target information using, for example, a quality estimation model for estimating the quality information of the target region. When the target region is input, the quality estimation model outputs the quality information of the target region.
[0048] In an embodiment according to the present disclosure, when the target information includes quality information, an example in which the quality information is the iris diameter will be described.
[0049] When the target information includes intermediate features of the target region, the target information acquisition unit 115 acquires the target information using an intermediate feature extraction model for extracting intermediate features (image features) of the target region. When a target region is input, the intermediate feature extraction model outputs intermediate features of the target region. The intermediate features are, for example, features extracted in the feature extraction model at a stage before the final features described below are extracted. Image features such as intermediate features may be represented, for example, as numerical vectors.
[0050] The intermediate feature extraction model is configured using, for example, a multi-layer neural network. In detail, for example, the intermediate feature extraction model includes at least one convolutional layer, at least one normalization layer, etc. The intermediate feature extraction model may be, for example, a learning model that constitutes a part of the feature extraction model.
[0051] The statistics regarding the target region are, for example, statistics based on intermediate features of the target region. In detail, for example, the statistics regarding the target region are the average μ *C , variance σ *C etc.
[0052] (Feature Extraction Model) The feature extraction model is a learning model for extracting features (image features) of a target region. When a target region or intermediate features of the target region are input, the feature extraction model extracts and outputs the features of the target region.
[0053] The features extracted by the feature extraction model are features (image features) used for matching. The features extracted by such a feature extraction model can also be considered as features (final features) that are finally extracted from the input image. Like the intermediate features, the final features are also image features and may be represented, for example, by a numerical vector.
[0054] The feature extraction model is configured using, for example, a multi-layer neural network. In detail, for example, the feature extraction model includes at least one convolution layer, at least one normalization layer, etc.
[0055] For example, when the target information acquisition unit 115 uses an intermediate feature extraction model, the input to the feature extraction model is, for example, intermediate features of the target region. In this case, the feature extraction model and the intermediate feature extraction model may be configured, for example, as a series of neural networks that, when a target region is input, output features of the target region (i.e., features used for matching). In detail, for example, the intermediate feature extraction model and the feature extraction model may be the front and back stages, respectively, when the series of neural networks is divided into two.
[0056] Note that the series of neural networks is not limited to being divided into two parts, an earlier stage and a later stage, but may be divided into three or more parts. When the series of neural networks is divided into three or more parts, for example, the parts other than the final stage may be used as intermediate feature extraction models, and the final stage may be used as a feature extraction model.
[0057] (Regarding Correction Parameters) The correction parameters are values for correcting parameters included in the feature extraction model. The calculation unit 116 may be configured to calculate correction parameters that improve authentication accuracy compared to default parameters included in the feature extraction model.
[0058] The parameters included in the feature extraction model are, for example, parameters for each channel C in the neural network that constitutes the feature extraction model. A channel C is a neuron set in a convolution layer that corresponds to each filter.
[0059] For example, when the feature extraction model is configured from a neural network including at least one normalization layer, the parameters corrected using the correction parameters are parameters used in at least one normalization layer.
[0060] In detail, for example, the parameters corrected using the correction parameters are parameters for each channel C used in at least one normalization layer. For example, the parameters for each channel C used in at least one normalization layer may be replaced with the correction parameters.
[0061] More specifically, for example, the correction parameters are μ for each channel C as follows: c , σ c , γ c , β c In the following equations (1) and (2), xic represents the feature map of channel C of sample i. μ c , σ c and γ represent the mean and variance of the batch norm channel C, respectively. c , β c represent the shift and scale parameters of channel C, respectively.
[0062]
[0063]
[0064] Note that correction using the correction parameters may be performed when the resolution of the target region is low, equal to or less than a predetermined threshold. In this case, the calculation unit 116 may determine whether the target region has low resolution, and calculate the correction parameters if the target region has low resolution. If the target region does not have low resolution, the calculation unit 116 does not need to calculate the correction parameters. In this case, the extraction unit 117 may extract features from the target region using a feature extraction model that has not been corrected with the correction parameters (i.e., a feature extraction model whose parameters are default values).
[0065] (Regarding the Matching Result) The matching result is the result of matching the feature amounts (final feature amounts) of the target region with the registered information that has been registered in advance. That is, the matching unit 118 matches the feature amounts of the target region with the registered information that has been registered in advance, and outputs the matching result. The matching result may include at least one of information including the similarity (e.g., cosine similarity) between the final feature amounts and the feature amounts included in the registered information, and information indicating whether the similarity satisfies a predetermined matching condition.
[0066] The matching condition may be, for example, that the similarity is equal to or greater than a predetermined matching threshold. If the similarity satisfies the matching condition, the matching result may be, for example, information indicating successful authentication. If the similarity does not satisfy the matching condition, the matching result may be, for example, information indicating unsuccessful authentication.
[0067] The matching result may include, for example, information indicating at least one likelihood of the similarity, the determination result of whether or not the matching condition is satisfied, etc. Note that the matching result is not limited to the examples given here.
[0068] The matching result may be output by displaying it on a display unit (not shown) or by transmitting it to another device (not shown). In detail, for example, the matching unit 118 may display the matching result on a display unit provided in the information processing device 100 or provided in another device (not shown) (for example, a mobile terminal used by the user). Also, for example, the matching unit 118 may transmit the matching result to another device (not shown) (for example, an information processing device that uses the matching result) via the network NT.
[0069] (Regarding Information Processing) The information processing device 100 may start information processing, for example, when it acquires an input image.
[0070] The trigger for the information processing device 100 to start information processing is not limited to this.
[0071] (Example of physical configuration of information processing device 100) The information processing device 100 physically includes a bus 2010, a processor 2020, a memory 2030, a storage device 2040, a network interface 2050, an input interface 2060, and an output interface 2070, as shown in FIG. 6, for example.
[0072] The bus 2010 is a data transmission path for transmitting and receiving data among the processor 2020, memory 2030, storage device 2040, network interface 2050, input interface 2060, and output interface 2070. However, the method of connecting the processor 2020 and the like to each other is not limited to bus connection.
[0073] The processor 2020 is implemented by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like.
[0074] The memory 2030 is a main storage device realized by a RAM (Random Access Memory) or the like.
[0075] The storage device 2040 is an auxiliary storage device realized by a hard disk drive (HDD), a solid state drive (SSD), a memory card, a read only memory (ROM), or the like. The storage device 2040 stores program modules for realizing the functions of the information processing device 100. The processor 2020 reads each of these program modules into the memory 2030 and executes them, thereby realizing the function corresponding to the program module.
[0076] The network interface 2050 is an interface for connecting the information processing device 100 to a network NT. The network NT is a communication network for transmitting and receiving information to and from other devices (not shown), and may be configured as a wired or wireless network or a combination of these.
[0077] The input interface 2060 is an interface for the user to input information, and is composed of, for example, a touch panel, a keyboard, a mouse, and the like.
[0078] The output interface 2070 is an interface for presenting information to the user, and is configured, for example, by a liquid crystal panel, an organic EL (Electro-Luminescence) panel, or the like.
[0079] Although the information processing device 100 has been described as being physically composed of one device (e.g., a computer), the information processing device 100 may be composed of multiple devices (e.g., computers) that transmit and receive information to and from each other via a network NT. In this case, the multiple devices may cooperate to execute information processing.
[0080] (Operations and Effects) As described above, according to this embodiment, the input image is an image that includes an iris region, which is a target region. The feature extraction model is configured from a neural network that includes at least one normalization layer. The parameters corrected using the correction parameters are the parameters used in the normalization layer.
[0081] This allows the normalization layer to suppress variations in the distribution of intermediate features that can arise due to variations in the quality of the input image, and extract features of the target region. Therefore, it is possible to extract features of the target region with high accuracy regardless of the quality of the input image, and obtain matching results using these features. Furthermore, the processing cost for extracting such accurate features of the target region is lower than when applying general image super-resolution technology.
[0082] Therefore, it is possible to suppress an increase in processing costs and perform authentication with high accuracy.
[0083] According to this embodiment, the target information includes at least one of quality information of the target region, intermediate feature amounts of the target region, and statistics related to the target region.
[0084] This makes it possible to extract features of the target region using a feature extraction model corrected with correction parameters based on at least one of quality information of the target region, intermediate features of the target region, and statistics related to the target region.
[0085] Therefore, even if the input image is low quality, it is possible to extract the feature values of the target region with high accuracy and obtain matching results using the feature values. Furthermore, the processing cost for extracting such accurate feature values of the target region is lower than when applying general image super-resolution technology.
[0086] Therefore, it is possible to suppress an increase in processing costs and perform authentication with high accuracy.
[0087] Second Embodiment In a second embodiment, an example will be described in which target information includes quality information of a target region estimated by inputting the target region into a quality estimation model for estimating quality information of the target region in an input image, and a correction parameter is calculated by inputting the target information into a first correction parameter estimation model for estimating the correction parameter.
[0088] 7 , the information processing system SYS2 includes an imaging device 201 and an information processing device 200. The information processing device 200 includes a target position estimation unit 211, a target area generation unit 212, a target information acquisition unit 215, a calculation unit 216, an extraction unit 117, and a matching unit 118.
[0089] The image capturing device 201 is a device that captures an image of a subject and generates an input image. The image capturing device 201 is, for example, a visible light camera, an infrared camera, a near-infrared camera, or the like.
[0090] In detail, for example, when the image capturing device 201 receives a signal indicating that a person is at a predetermined position from a sensor such as a motion sensor (not shown), the image capturing device 201 captures an image of the vicinity of both eyes of the person and generates a binocular image. The image capturing device 201 may output the generated binocular image.
[0091] When the image capturing device 201 acquires an input image, the target position estimation unit 211 estimates the position of the target area in the input image.
[0092] In more detail, for example, when the target region is an iris region, the target position estimation unit 211 estimates the eye position including the position of the pupil center, pupil diameter, and iris diameter as the position of the target region. The pupil diameter may be the diameter of the pupil or the radius of the pupil. The pupil diameter may be expressed, for example, by the number of pixels. Note that the method of expressing the pupil diameter is not limited to the number of pixels.
[0093] The technology for estimating the position of the target area may be a general technology such as pattern matching or a machine learning model. When a machine learning model is used, the target position estimation unit 211 may use an input image as input and estimate the position of the target area in the input image using a learning model that has been trained to estimate the position of the target area from a learning input image. In learning this learning model, supervised learning may be performed using the learning input image and training data including the position of the target area in the learning input image.
[0094] It should be noted that the eye position and the technique for estimating the same are not limited to those exemplified here.
[0095] The target region generating unit 212 generates a target region (an image showing only the target region) based on the input image and the estimated position of the target region.
[0096] In detail, for example, the target region generation unit 212 may generate a target region (for example, a roughly annular iris region as shown in FIG. 5 ) by cutting out the target region from the input image. Alternatively, the target region generation unit 212 may generate a target region in a predetermined shape (for example, a roughly rectangular iris region as shown in FIG. 8 ) by cutting out the target region from the input image and further performing a predetermined transformation.
[0097] For example, when an eyelid or the like covers a part of the iris region in a binocular image, the iris region may be changed from a circular or rectangular shape to a shape in which the part covered by the eyelid or the like is missing.
[0098] The target information acquisition unit 215 acquires quality information relating to a target region in an input image as target information.
[0099] In detail, for example, the target information acquisition unit 215 acquires estimated quality information of the target region by inputting the target region into a quality estimation model. The quality estimation model is a learning model for estimating quality information of the target region in the input image. For example, when the target region is input, the quality estimation model may output quality information of the target region. The quality information may be, for example, the iris diameter.
[0100] In training the quality estimation model, supervised learning may be performed using training input images and training data including quality information on the training input images.
[0101] Note that the technique for estimating the quality information is not limited to the example given here. Furthermore, when the quality information is obtained by the target position estimation unit 211, such as when the quality information is the iris diameter, the target information acquisition unit 215 may acquire the quality information from the target position estimation unit 211.
[0102] The calculation unit 216 calculates correction parameters for correcting parameters included in the feature extraction model based on the quality information.
[0103] In detail, for example, the calculation unit 216 calculates the correction parameter by inputting quality information into a first correction parameter estimation model. The first correction parameter estimation model is a learning model for estimating the correction parameter. For example, when quality information of the target region is input, the first correction parameter estimation model outputs the correction parameter. This correction parameter is, for example, the average μ for each channel as described above. c , variance σ c , shift γ c , scale parameter β c is.
[0104] The first correction parameter estimation model may be configured using, for example, a neural network such as a convolutional neural network, an attention mechanism, or the like.
[0105] A method for learning the first correction parameter estimation model will be described in another embodiment.
[0106] (Example of processing operation of information processing device 200) The information processing device 200 executes information processing as shown in Fig. 9. The information processing device 200 may start the information processing when it acquires an input image from the image capturing device 201, for example.
[0107] The target position estimation unit 211 estimates the position of the target area in the input image based on the input image generated by the imaging device 201 (step S211).
[0108] The target area generating unit 212 generates a target area (an image showing only the target area) based on the input image generated by the imaging device 201 and the position of the target area estimated in step S211 (step S212).
[0109] The target information acquisition unit 215 inputs the target region generated in step S212 into a quality estimation model, thereby acquiring quality information of the target region, which is target information (step S215).
[0110] The calculation unit 216 inputs the quality information acquired in step S215 into the first correction parameter estimation model to calculate the correction parameters (step S216).
[0111] The extraction unit 117 extracts features of the target region using the feature extraction model corrected with the correction parameters (step S117).
[0112] In detail, for example, the extraction unit 117 extracts features of the target region by inputting the target region generated in step S212 into a feature extraction model corrected with the correction parameters calculated in step S216.
[0113] The matching unit 118 outputs the result of matching the feature amount of the target region extracted in step S117 with the registered information that has been registered in advance (step S118), and ends the information processing.
[0114] (Actions and Effects) As described above, according to this embodiment, the target information includes quality information estimated by inputting the target area into a quality estimation model for estimating quality information of the target area. The correction parameters are calculated by inputting the target information into a first correction parameter estimation model for estimating the correction parameters.
[0115] This makes it possible to extract features of the target region using a feature extraction model corrected with correction parameters based on quality information of the target region.
[0116] Therefore, even if the input image is low quality, it is possible to extract the feature values of the target region with high accuracy and obtain matching results using the feature values. Furthermore, the processing cost for extracting such accurate feature values of the target region is lower than when applying general image super-resolution technology.
[0117] Therefore, it is possible to suppress an increase in processing costs and perform authentication with high accuracy.
[0118] Third Embodiment In a third embodiment, an example will be described in which target information includes intermediate features extracted by inputting a foreground region into an intermediate feature extraction model for extracting intermediate features of a target region in an input image. Also, an example will be described in which correction parameters are calculated by inputting the target information into a second correction parameter estimation model for estimating the correction parameters.
[0119] 10 , the information processing system SYS3 includes an imaging device 201 and an information processing device 300. The information processing device 300 includes a target position estimation unit 211, a target area generation unit 212, a target information acquisition unit 315, a calculation unit 316, an extraction unit 117, and a matching unit 118.
[0120] The target information acquisition unit 315 acquires intermediate feature amounts of a target region in an input image as target information.
[0121] In detail, for example, the target information acquisition unit 315 acquires intermediate features of the target region by inputting the target region into the intermediate feature extraction model described above.
[0122] A method for learning the intermediate feature extraction model will be described in another embodiment.
[0123] The calculation unit 316 calculates correction parameters for correcting parameters included in the feature extraction model based on the intermediate feature amounts of the target region.
[0124] In detail, for example, the calculation unit 316 calculates the correction parameters by inputting the intermediate feature values of the target region into the second correction parameter estimation model. The second correction parameter estimation model is a learning model for estimating the correction parameters. For example, when the intermediate feature values of the target region are input, the second correction parameter estimation model outputs the correction parameters. The correction parameters are, for example, the average μ for each channel as described above. c , variance σ c , shift γ c , scale parameter β c is.
[0125] The second correction parameter estimation model may be configured using, for example, a neural network such as a convolutional neural network, an attention mechanism, or the like.
[0126] A method for learning the second correction parameter estimation model will be described in another embodiment. Note that, instead of the intermediate feature values, norms of the intermediate feature values may be used.
[0127] (Example of processing operation of information processing device 300) The information processing device 300 executes information processing as shown in Fig. 11. The information processing device 300 may start information processing when it acquires an input image from the image capturing device 201, for example.
[0128] The above-described steps S211 and S212 are executed.
[0129] The target information acquisition unit 315 inputs the target region generated in step S212 into an intermediate feature extraction model, thereby acquiring target information including intermediate features of the target region (step S315).
[0130] The calculation unit 316 calculates correction parameters by inputting the intermediate feature amount of the target region acquired in step S315 into the second correction parameter estimation model (step S316).
[0131] The extraction unit 117 extracts features of the target region using the feature extraction model corrected with the correction parameters (step S117).
[0132] In detail, for example, the extraction unit 117 extracts features of the target region by inputting the intermediate features acquired in step S315 into a feature extraction model corrected with the correction parameters calculated in step S316.
[0133] The matching unit 118 outputs the result of matching the feature amount of the target region extracted in step S117 with the registered information that has been registered in advance (step S118), and ends the information processing.
[0134] As described above, according to this embodiment, the target information includes intermediate features extracted by inputting the target region into an intermediate feature extraction model for extracting intermediate features of the target region. The correction parameters are calculated by inputting the target information into a second correction parameter estimation model for estimating the correction parameters.
[0135] This makes it possible to extract features of the target region using a feature extraction model corrected with correction parameters based on the intermediate features of the target region.
[0136] Generally, there is a correlation between the quality of an image and the norm of a feature vector extracted from that image. For example, the higher the quality of an image, the larger the norm of the feature vector extracted from that image. Intermediate features also hold quality information that indicates the quality of the image from which the intermediate features were extracted. For example, the quality information can be estimated by calculating the norm.
[0137] Therefore, the correction parameters can be estimated by using a correction parameter estimation model with the norm of the intermediate feature or the intermediate feature itself as input. By extracting the feature of the target region using the correction parameter based on such intermediate feature, it is possible to extract the feature of the target region with high accuracy even if the input image is of low quality, and obtain matching results using the feature. Furthermore, the processing cost for extracting the feature of the target region with high accuracy is lower than that when applying general image super-resolution technology.
[0138] Therefore, it is possible to suppress an increase in processing costs and perform authentication with high accuracy.
[0139] In the fourth embodiment, an example will be described in which the target information includes quality information of the target region or intermediate feature amounts of the target region. The correction parameters are obtained by inputting the target information into a correction coefficient estimation model for estimating the correction coefficients S and B, and a dictionary parameter μ d c , σ d c , γ d c , β d c An example in which the value of the first relationship is calculated by applying the above formula to a predetermined first relationship will be described.
[0140] 12 , the information processing system SYS4 includes an imaging device 201 and an information processing device 400. The information processing device 400 includes a target position estimation unit 211, a target area generation unit 212, a target information acquisition unit 415, a calculation unit 416, an extraction unit 117, and a matching unit 118.
[0141] The target information acquisition unit 415 acquires target information in the input image. The target information according to this embodiment is quality information or intermediate feature amounts of the target region.
[0142] When acquiring quality information, the target information acquisition unit 415 may have the same functions as the target information acquisition unit 215 described in embodiment 2. When acquiring intermediate features, the target information acquisition unit 415 may have the same functions as the target information acquisition unit 315 described in embodiment 3.
[0143] The calculation unit 416 calculates correction parameters for correcting parameters included in the feature extraction model based on the target information acquired by the target information acquisition unit 415 .
[0144] As shown in FIG. 13, the calculation unit 416 includes a dictionary parameter storage unit 416a, a correction coefficient acquisition unit 416b, and a correction parameter calculation unit 416c.
[0145] The dictionary parameter storage unit 416a stores dictionary parameters in advance. The dictionary parameters include, for example, the average μ d c, variance σ d c , shift γ d c , scale parameter β d c The dictionary parameter, the average μ d c , variance σ d c , shift γ d c , scale parameter β d c may be a standard value that is set appropriately.
[0146] The correction coefficient acquisition unit 416b acquires correction coefficients S and B used to calculate correction parameters based on the target information (quality information or intermediate feature amount).
[0147] For example, the correction coefficient acquisition unit 416b acquires the correction coefficients S and B by inputting target information (quality information or intermediate feature amounts) into a correction coefficient estimation model. The correction coefficient estimation model is a learning model for estimating the correction coefficients S and B. When the quality information or intermediate feature amounts are input, the correction coefficient estimation model outputs the correction coefficients S and B.
[0148] The correction parameter calculation unit 416c calculates a correction parameter by applying a dictionary parameter and a correction coefficient stored in advance to a first relationship that is determined in advance. The correction parameter is, for example, the average μ c , variance σ c , shift γ c , scale parameter β c is.
[0149] The first relationship is defined, for example, by the following formula (3): X in formula (3) represents each of μ, σ, γ, and β. Note that the first relationship is not limited to the example given here, and may be defined, for example, by another relational expression or a multidimensional matrix.
[0150]
[0151] A method for learning the correction coefficient estimation model will be described in another embodiment.
[0152] (Example of processing operation of information processing device 400) The information processing device 400 executes information processing as shown in Fig. 14. The information processing device 400 may start information processing when it acquires an input image from the image capturing device 201, for example.
[0153] The above-described steps S211 and S212 are executed.
[0154] The target information acquisition unit 415 acquires target information (quality information or intermediate features of the target region) using the target region generated in step S212 and the quality estimation model or intermediate feature extraction model (step S415).
[0155] The correction coefficient acquisition unit 416b calculates correction parameters based on the target information (quality information or intermediate feature amount) acquired in step S415 (step S416).
[0156] For example, as shown in FIG. 15, the correction coefficient acquisition unit 416b acquires the correction coefficients S and B by inputting target information (quality information or intermediate feature amounts) into a correction coefficient estimation model (step S416a).
[0157] The correction parameter calculation unit 416c calculates the correction parameters using the pre-stored dictionary parameters, the correction coefficients S and B obtained in step S416a, and the predetermined first relationship (step S416b), and returns to information processing.
[0158] For example, the correction parameter calculation unit 416c substitutes the dictionary parameters and correction coefficients S and B into the formula (3) in which X is replaced with μ, σ, γ, and β, respectively, to calculate the average μ for each channel. c , variance σ c , shift γ c , scale parameter β c Calculate.
[0159] Referring again to Fig. 14, the extraction unit 117 extracts features of the target region using the feature extraction model corrected with the correction parameters as described above (step S117).
[0160] In detail, for example, when the target information includes quality information, as described in the second embodiment, the extraction unit 117 may extract features of the target region by inputting the target region generated in step S212 into the feature extraction model corrected by the correction parameters.
[0161] Furthermore, for example, when the target information includes intermediate features, as described in the third embodiment, the extraction unit 117 may extract features of the target region by inputting the intermediate features acquired in step S315 into the feature extraction model corrected with the correction parameters.
[0162] The matching unit 118 outputs the result of matching the feature amount of the target region extracted in step S117 with the registered information that has been registered in advance (step S118), and ends the information processing.
[0163] According to the present embodiment, the target information includes quality information of the target region or intermediate feature amounts of the target region. The correction parameters are calculated by applying the correction coefficients, which are obtained by inputting the target information into a correction coefficient estimation model, and the dictionary parameters, which are stored in advance, to a predetermined first relationship.
[0164] According to this, since the correction parameters are calculated using the correction coefficients, the number of coefficients estimated based on the target information can be made smaller than the number of correction parameters, thereby reducing the amount of processing required for estimation based on the target information.
[0165] Furthermore, since the feature extraction model corrected by the correction parameters can be used to extract the features of the target region, even if the input image is of low quality, the feature of the target region can be extracted with high accuracy, and matching results can be obtained using the features.
[0166] Therefore, it is possible to further suppress increases in processing costs and perform authentication with high accuracy.
[0167] Fifth Embodiment In a fifth embodiment, an example will be described in which the target information includes statistics related to the target region calculated using intermediate features of the target region, and the correction parameters are calculated by applying the statistics related to the target region and pre-stored dictionary parameters to a predetermined second relationship.
[0168] 16 , the information processing system SYS5 includes an imaging device 201 and an information processing device 500. The information processing device 500 includes a target position estimation unit 211, a target area generation unit 212, a target information acquisition unit 515, a calculation unit 516, an extraction unit 117, and a matching unit 118.
[0169] The target information acquisition unit 515 acquires statistics relating to the target region in the input image as target information. The statistics relating to the target region may be, for example, the average μ c , variance σ c is.
[0170] In detail, for example, the target information acquisition unit 515 has the same function as the target information acquisition unit 315 described in the third embodiment, and thereby acquires intermediate features of the target region, i.e., intermediate features for each channel related to the target region. The target information acquisition unit 515 performs statistical processing on the intermediate features to obtain the average μ *C , variance σ *C Calculate.
[0171] The calculation unit 516 calculates correction parameters for correcting parameters included in the feature extraction model based on the target information acquired by the target information acquisition unit 415 .
[0172] As shown in FIG. 17, the calculation unit 516 includes a dictionary parameter storage unit 416a and a correction parameter calculation unit 516b.
[0173] The correction parameter calculation unit 516b calculates the correction parameter by applying the statistics related to the target region and the dictionary parameters stored in advance to a predetermined second relationship.
[0174] The second relationship is defined, for example, by the following formula (4). In formula (4), Y represents each of μ and σ. Note that the second relationship is not limited to the example given here, and may be defined, for example, by another relational expression or a multidimensional matrix. m is a weight that is determined appropriately, and is, for example, 0<m<1.
[0175]
[0176] According to the present embodiment, the target information includes statistics related to the target region calculated using intermediate features of the target region. The correction parameters are calculated by applying the statistics related to the target region and pre-stored dictionary parameters to a predetermined second relationship.
[0177] This makes it possible to extract features of the target region using a feature extraction model corrected with correction parameters based on statistics related to the target region.
[0178] Therefore, even if the input image is low quality, it is possible to extract the feature values of the target region with high accuracy and obtain matching results using the feature values. Furthermore, the processing cost for extracting such accurate feature values of the target region is lower than when applying general image super-resolution technology.
[0179] Therefore, it is possible to suppress an increase in processing costs and perform authentication with high accuracy.
[0180] Sixth Embodiment In the fifth embodiment, the target information acquisition unit 515 acquires intermediate feature values for each channel related to the target region and the average μ *C , variance σ *C In this case, the calculation unit 516 may further include functions similar to those of the calculation unit 316 described in the third embodiment, in addition to the functions described in the fifth embodiment. This allows the calculation unit 516 to calculate correction parameters based on the intermediate feature amounts of the target region and statistics related to the target region.
[0181] The extraction unit 117 may extract the feature of the target region by inputting the intermediate feature to a feature extraction model in which, for example, some of the parameters of the feature extraction model are corrected by correction parameters based on statistics and the remaining parameters are corrected by correction parameters based on intermediate feature. The parameters corrected by correction parameters based on statistics include, for example, the average μ c , variance σ c At least one of the above may be used.
[0182] (Operations and Effects) As described above, according to this embodiment, the target information includes intermediate feature amounts of the target region and statistics related to the target region. The correction parameters are calculated based on the intermediate feature amounts of the target region and statistics related to the target region.
[0183] This makes it possible to extract features of the target region using a feature extraction model that has been more appropriately corrected with correction parameters based on intermediate features of the target region and statistics related to the target region.
[0184] Therefore, even if the input image is of low quality, it is possible to extract the feature values of the target region with high accuracy and obtain matching results using the feature values. Furthermore, the processing cost for extracting such accurate feature values of the target region is lower than when applying general image super-resolution technology.
[0185] Therefore, it is possible to suppress an increase in processing costs and perform authentication with higher accuracy.
[0186] Seventh Embodiment In this embodiment, a first example of the learning method for the first correction parameter estimation model and the feature quantity extraction model described in the second embodiment will be described.
[0187] (Configuration example of information processing system SYS6 and information processing device 600) As shown in Fig. 18 , the information processing system SYS6 includes an imaging device 201 and an information processing device 600. The information processing device 600 includes a learning unit 619 in addition to the functions of the information processing device 200.
[0188] The learning unit 619 uses training data prepared in advance to train the feature extraction model and the first correction parameter estimation model. That is, the learning unit 619 uses training data common to the feature extraction model and the first correction parameter estimation model to simultaneously train the feature extraction model and the first correction parameter estimation model. The training data may include, for example, training input images and correct labels.
[0189] In detail, for example, as shown in FIG. 19, the learning unit 619 includes a learning calculation unit 619a, a learning feature extraction unit 619b, a loss calculation unit 619c, and an update unit 619d.
[0190] The learning calculation unit 619a inputs quality information of the learning input image included in the training data into the first correction parameter estimation model to calculate the learning correction parameters.
[0191] The training feature extraction unit 619b inputs the training input image included in the training data to a feature extraction model corrected with the training correction parameters, and extracts training features. The training features are features of the target region included in the training input image.
[0192] The loss calculation unit 619c calculates the loss based on the correct labels and learning features included in the training data. As the loss function, a general loss function such as the sum of squares error or the cross entropy error may be used.
[0193] The update unit 619 d updates the parameters included in the feature extraction model and the first correction parameter estimation model based on the loss. This updating may be performed using a general optimization method such as gradient descent or stochastic gradient descent.
[0194] (Example of processing operation of information processing device 600) The information processing performed by the information processing device 600 may further include a learning process as shown in FIG. 20 in addition to the information processing described in the second embodiment. The learning process may be started, for example, in response to an instruction from a user. Training data may be prepared before the learning process is started. Note that the trigger for starting the learning process is not limited to this.
[0195] The learning calculation unit 619a inputs quality information relating to a target region in a learning input image to a first correction parameter estimation model, and calculates learning correction parameters (step S601).
[0196] Here, the quality information regarding the learning input image may be acquired using, for example, the functions of the target region generation unit 212 and the target information acquisition unit 215, or may be included in the training data. Furthermore, the learning unit 619 may further include an image degradation unit (not shown) that degrades the learning input image to an appropriate quality specified by the user. In this case, the quality information regarding the learning input image is the quality specified by the user and may be acquired from the image degradation unit.
[0197] The learning feature extraction unit 619b inputs the learning input image to the feature extraction model corrected with the learning correction parameters calculated in step S601, and extracts learning features (step S602).
[0198] The loss calculation unit 619c calculates the loss based on the correct label and the learning features (step S603).
[0199] The update unit 619d updates the parameters included in the feature extraction model and the first correction parameter estimation model based on the loss calculated in step S603 (step S604), and ends the learning process.
[0200] (Operations and Effects) As described above, according to this embodiment, the information processing device 600 includes the learning unit 619 that uses training data to learn the feature extraction model and the first correction parameter estimation model.
[0201] This allows the feature extraction model and the first correction parameter estimation model to be trained simultaneously using common training data. Therefore, the effort required for preparing training data can be reduced compared to preparing training data for each model. Furthermore, the effort required for training can be reduced compared to training each of the feature extraction model and the first correction parameter estimation model individually. Therefore, the first correction parameter estimation model and the feature extraction model can be easily trained.
[0202] Eighth Embodiment In this embodiment, a second example of the learning method for the first correction parameter estimation model and the feature quantity extraction model described in the second embodiment will be described.
[0203] (Configuration example of information processing system SYS7 and information processing device 700) As shown in Fig. 21 , the information processing system SYS7 includes an imaging device 201 and an information processing device 700. In addition to the functions of the information processing device 200, the information processing device 700 includes a first learning unit 720 and a second learning unit 721.
[0204] The first learning unit 720 uses first training data prepared in advance to learn a feature extraction model, and the second learning unit 721 uses second training data prepared in advance to learn a first correction parameter estimation model.
[0205] The first training data includes first input images that are input images for learning, and may further include first correct labels that are correct labels associated with the first input images.
[0206] The second training data includes second input images that are input images for learning. The second training data may further include second correct labels that are correct labels associated with the second input images.
[0207] The target region in the first input image has higher quality than the target region in the second input image. Such a second input image may be prepared by an appropriate method. For example, the information processing device 700 may further include an image degradation unit (not shown) that degrades the image to an appropriate quality specified by the user. In this case, the second input image may be obtained by degrading the image quality of the first input image using the image degradation unit. Note that the method for preparing the second input image is not limited to this.
[0208] The first learning unit 720 and the second learning unit 721 use first training data and second training data including learning images of different image quality to individually learn the feature extraction model and the first correction parameter estimation model, respectively.
[0209] (Configuration Example of First Learning Unit 720) As shown in FIG. 22, the first learning unit 720 includes a first feature amount extracting unit 720a, a first loss calculating unit 720b, and a first updating unit 720c.
[0210] The first feature extraction unit 720a inputs the first input image to a feature extraction model to extract a first feature, which is a feature of a target region included in the first input image, and the same applies hereinafter.
[0211] The first loss calculation unit 720b calculates the loss based on the first correct label and the first feature amount. As the loss function, a general loss function such as a sum of squares error or a cross entropy error may be used.
[0212] The first update unit 720c updates the parameters included in the feature extraction model based on the first loss. This updating may be performed using a general optimization method such as gradient descent or stochastic gradient descent.
[0213] (Configuration Example of Second Learning Unit 721) As shown in FIG. 23, the second learning unit 721 includes a learning calculation unit 721a, a second feature amount extraction unit 721b, a second loss calculation unit 721c, and a second update unit 721d.
[0214] The learning calculation unit 721a inputs the quality information of the second input image into the first correction parameter estimation model to calculate learning correction parameters.
[0215] Here, the quality information regarding the second input image may be acquired using, for example, the functions of the target region generation unit 212 and the target information acquisition unit 215, or may be included in the second training data. Furthermore, if the second input image is created using an image degradation unit (not shown), the quality information regarding the second input image is a quality specified by the user and may be acquired from the image degradation unit.
[0216] The second feature extraction unit 721b inputs the second input image to a trained feature extraction model corrected with the training correction parameters, and extracts the second feature, which is a feature of a target region included in the second input image, and the same applies hereinafter.
[0217] The second loss calculation unit 721c calculates the second loss based on the second correct label and the second feature amount. As the loss function, a general loss function such as a sum of squares error or a cross entropy error may be used.
[0218] The second update unit 721d updates the parameters included in the first correction parameter estimation model based on the second loss. This update may be performed using a general optimization method such as gradient descent or stochastic gradient descent.
[0219] (Example of processing operation of information processing device 700) The information processing performed by the information processing device 700 may further include, in addition to the information processing described in embodiment 2, a first learning process and a second learning process as shown in Figures 24 and 25, respectively.
[0220] The first learning process is a process for training a feature extraction model. The second learning process is a process for training a first correction parameter estimation model. The second learning process is performed using the feature extraction model that has been trained by executing the first learning process. Therefore, it is preferable that the second learning process be performed after the first learning process.
[0221] The first learning process may be started, for example, in response to an instruction from a user. The second learning process may be automatically executed following the first learning process, or may be started in response to an instruction from a user. The first training data and the second training data may be prepared before the learning process is started. Note that the trigger for starting the learning process is not limited to this.
[0222] (First Learning Process) See Fig. 24. The first feature extraction unit 720a inputs the first input image to the feature extraction model and extracts the first feature (step S701).
[0223] The first loss calculation unit 720b calculates a loss based on the first correct label and the first feature amount (step S702).
[0224] The first update unit 720c updates the parameters included in the feature extraction model based on the first loss (step S703), and ends the first learning process, thereby completing learning of the feature extraction model.
[0225] (Second Learning Process) In the second learning process, a feature extraction model that has been trained by executing the first learning process, that is, a trained feature extraction model, is used.
[0226] See Fig. 25. The learning calculation unit 721a inputs the quality information of the second input image to the first correction parameter estimation model to calculate learning correction parameters (step S711).
[0227] The second feature extraction unit 721b inputs the second input image to the trained feature extraction model corrected with the learning correction parameters, and extracts the second feature (step S712).
[0228] The second loss calculation unit 721c calculates the second loss based on the second correct label and the second feature amount (step S713).
[0229] The second update unit 721d updates the parameters included in the first correction parameter estimation model based on the second loss (step S714), and ends the second learning process, thereby completing the learning of the first correction parameter estimation model.
[0230] (Actions and Effects) As described above, according to this embodiment, the information processing device 700 includes a first learning unit 720 that uses first training data to learn a feature extraction model, and a second learning unit 721 that uses second training data to learn a first correction parameter estimation model.
[0231] The first training data includes a first input image that is an input image for learning, and the second training data includes a second input image that is an input image for learning, and the target region in the first input image has higher quality than the target region in the second input image.
[0232] This allows the feature extraction model to be trained using high-quality training images (first input images), and the first correction parameter estimation model to be trained using low-quality training images (second input images). Therefore, the feature extraction model and the first correction parameter estimation model can be trained so that the features of the target region can be extracted with high accuracy. This allows for more accurate authentication.
[0233] [Embodiment 8] This embodiment describes a third example of the method for learning the first correction parameter estimation model and the feature quantity extraction model described in Embodiment 2. This embodiment describes an example in which the feature quantity extraction model and the first correction parameter estimation model are individually learned using a feature quantity norm (norm of a feature quantity vector).
[0234] (Configuration example of information processing system SYS8 and information processing device 800) As shown in Fig. 26 , the information processing system SYS8 includes an imaging device 201 and an information processing device 800. In addition to the functions of the information processing device 200, the information processing device 800 includes a first learning unit 820 and a second learning unit 821.
[0235] The first learning unit 820 uses first training data prepared in advance to learn the feature extraction model, and the second learning unit 821 uses second training data prepared in advance to learn the first correction parameter estimation model.
[0236] The first training data includes first input images that are input images for learning. The first training data may further include first correct labels that are correct labels associated with the first input images, and a first correct norm.
[0237] The first correct norm is, for example, a correct value of the norm of a feature extracted from a target region in the first input image. The feature norm is the norm of a vector representing the feature (feature vector), and the same applies hereinafter.
[0238] The second training data includes second input images that are input images for learning. The second training data may further include second correct labels that are correct labels associated with the second input images, and a second correct norm.
[0239] The second correct norm is, for example, a correct value of the norm of the feature amount extracted from the target region in the second input image.
[0240] The target region in the first input image has higher quality than the target region in the second input image. Such a second input image may be prepared by an appropriate method, as described in the seventh embodiment. That is, for example, the second input image may be obtained by degrading the image quality of the first input image using an image degradation unit (not shown) included in the information processing device 800. Note that the method for preparing the second input image is not limited to this.
[0241] The first learning unit 820 and the second learning unit 821 use first training data and second training data containing learning images of different image quality to individually learn the feature extraction model and the first correction parameter estimation model, respectively.
[0242] (Configuration example of first learning unit 820) As shown in FIG. 27 , the first learning unit 820 includes a first feature quantity extraction unit 720a, a first loss calculation unit 720b, a first norm calculation unit 820c, a first norm loss calculation unit 820d, a loss integrating unit 820e, and a first update unit 820f.
[0243] The first norm calculation unit 820c calculates the first norm, which is the norm of the first feature amount (feature amount vector).
[0244] The first norm loss calculation unit 820d calculates the first norm loss based on the first ground truth norm included in the first training data and the first norm. As the loss function, a general loss function such as a sum of squares error or a cross entropy error may be used.
[0245] The loss integrator 820e calculates an integrated loss by integrating the first loss and the first norm loss. For example, the loss integrator 820e calculates the integrated loss by multiplying each of the first loss and the first norm loss by a predetermined weight and then adding the results together.
[0246] The first update unit 820f updates the parameters included in the feature extraction model. This updating may be performed using a general optimization method such as gradient descent or stochastic gradient descent.
[0247] (Example configuration of second learning unit 821) As shown in FIG. 28 , the second learning unit 821 includes a learning calculation unit 721a, a second feature extraction unit 721b, a second norm calculation unit 821c, a second norm loss calculation unit 821d, and a second update unit 821e.
[0248] The second norm calculation unit 821c calculates the second norm, which is the norm of the second feature amount.
[0249] The second norm loss calculation unit 821d calculates the second norm loss based on the ground truth norm included in the second training data and the second norm. As the loss function, a general loss function such as a sum of squares error or a cross entropy error may be used.
[0250] The second update unit 821e updates the parameters included in the first correction parameter estimation model based on the second norm loss. This update may be performed using a general optimization method such as gradient descent or stochastic gradient descent.
[0251] (Example of processing operation of information processing device 800) The information processing performed by the information processing device 800 may further include, in addition to the information processing described in embodiment 2, a first learning process and a second learning process as shown in Figures 29 and 30, respectively.
[0252] The first learning process is a process for training a feature extraction model. The second learning process is a process for training a first correction parameter estimation model. The second learning process is performed using the feature extraction model that has been trained by executing the first learning process. Therefore, it is preferable that the second learning process be performed after the first learning process.
[0253] The first learning process may be started, for example, in response to an instruction from a user. The second learning process may be automatically executed following the first learning process, or may be started in response to an instruction from a user. The first training data and the second training data may be prepared before the learning process is started. Note that the trigger for starting the learning process is not limited to this.
[0254] (First Learning Process) See Fig. 29. Steps S701 and S702 described in the seventh embodiment are executed.
[0255] The first norm calculation unit 820c calculates the first norm, which is the norm of the first feature amount (step S803).
[0256] The first norm loss calculation unit 820d calculates the first norm loss based on the first ground truth norm included in the first training data and the first norm (step S804).
[0257] The loss integrating unit 820e calculates an integrated loss by integrating the first loss and the first norm loss (step S805).
[0258] The first update unit 820f updates the parameters included in the feature extraction model (step S806), and ends the first learning process, thereby completing learning of the feature extraction model.
[0259] (Second Learning Process) See Fig. 30. Steps S711 to S712 described in the seventh embodiment are executed.
[0260] The learning unit 721 includes a learning calculation unit 721a, a second feature amount extraction unit 721b, a second norm calculation unit 821c, a second norm loss calculation unit 821d, and a second update unit 821e.
[0261] The second norm calculation unit 821c calculates the second norm, which is the norm of the second feature amount (step S813).
[0262] The second norm loss calculation unit 821d calculates the second norm loss based on the ground truth norm included in the second training data and the second norm calculated in step S813 (step S814).
[0263] The second update unit 821e updates the parameters included in the first correction parameter estimation model based on the second norm loss (step S815), and ends the second learning process, thereby completing the learning of the first correction parameter estimation model.
[0264] (Operations and Effects) As described above, according to this embodiment, the information processing device 800 includes the first learning unit 820 and the second learning unit 821 .
[0265] The first learning unit 820 includes a first feature extractor 720a, a first loss calculator 720b, a first norm calculator 820c, a first norm loss calculator 820d, a loss integrator 820e, and a first updater 820f.
[0266] The first feature extraction unit 720a inputs a first input image to the feature extraction model and extracts a first feature that is a feature of the first input image. The first loss calculation unit 720b calculates a first loss based on the correct label included in the first training data and the first feature.
[0267] The first norm calculation unit 820c calculates a first norm, which is the norm of the first feature. The first norm loss calculation unit 820d calculates a first norm loss based on the ground truth norm included in the first training data and the first norm. The loss integration unit 820e calculates an integrated loss by integrating the first loss and the first norm loss. The first update unit 820f updates parameters included in the feature extraction model based on the integrated loss.
[0268] The second learning unit 821 includes a learning calculation unit 721a, a second feature amount extraction unit 721b, a second norm calculation unit 821c, a second norm loss calculation unit 821d, and a second update unit 821e.
[0269] The learning calculation unit 721a inputs quality information of the second input image to the first correction parameter estimation model to calculate learning correction parameters. The second feature extraction unit 721b inputs the second input image to a trained feature extraction model corrected with the learning correction parameters to extract second feature values that are feature values of the second input image.
[0270] The second norm calculation unit 821c calculates a second norm, which is the norm of the second feature amount. The second norm loss calculation unit 821d calculates a second norm loss based on the ground truth norm included in the second training data and the second norm. The second update unit 821e updates the parameters included in the first correction parameter estimation model based on the second norm loss.
[0271] This allows the feature extraction model to be trained using a high-quality training image (first input image), and the first correction parameter estimation model to be trained using a low-quality training image (second input image). Furthermore, as described above, there is a correlation between image quality and the norm of the feature vector extracted from the image. Therefore, the feature extraction model and the first correction parameter estimation model can be trained to extract features of the target region with even greater accuracy. This allows for even more accurate authentication.
[0272] [Embodiment 10] In the embodiment 9, an example was described in which the correct norm included in the second training data used in step S814 is the correct value of the norm of the feature amount extracted from the target region in the second input image. However, the correct norm included in the second training data is not limited to this.
[0273] Generally, the higher the quality of an image, the longer the norm of the feature extracted from the image, and the lower the error rate of the authentication result using the image. Therefore, by training the first correction parameter estimation model to obtain correction parameters that increase the norm length of the feature extracted from the image using the feature extraction model, it is possible to reduce the error rate of the authentication result using the image.
[0274] Therefore, for example, the average value of the correct norms included in the first training data (hereinafter referred to as the "first correct norm") may be used as the correct norm included in the second training data (hereinafter referred to as the "second correct norm"). As described above, the first input image has higher quality than the second input image. Therefore, by using such second training data, the first correction parameter estimation model can be trained to obtain correction parameters that bring the length of the norm of the feature extracted from the low-quality image closer to the same length as that extracted from the high-quality image. This enables accurate authentication.
[0275] Eleventh Embodiment In this embodiment, a first example of the learning method for the intermediate feature extraction model, the second correction parameter estimation model, and the feature extraction model described in the third embodiment will be described.
[0276] (Configuration example of information processing system SYS9 and information processing device 900) As shown in Fig. 31 , the information processing system SYS9 includes an imaging device 201 and an information processing device 900. The information processing device 900 includes a learning unit 919 in addition to the functions of the information processing device 300.
[0277] The learning unit 919 uses training data prepared in advance to train the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model. That is, the learning unit 919 uses training data common to the feature extraction model and the first correction parameter estimation model to simultaneously train the feature extraction model and the first correction parameter estimation model in an end-to-end manner. The training data may include, for example, training input images and correct labels.
[0278] In detail, for example, the learning unit 919 includes a learning target information acquisition unit 919a, a learning calculation unit 919b, a learning feature extraction unit 919c, a loss calculation unit 919d, and an update unit 919e, as shown in FIG. 32.
[0279] The learning target information acquisition unit 919a inputs the learning input image to an intermediate feature extraction model to extract learning intermediate features. The learning intermediate features are intermediate features of a target region in the learning input image.
[0280] The learning calculation unit 919b inputs the learning intermediate feature amounts to the second correction parameter estimation model to calculate learning correction parameters.
[0281] The training feature extraction unit 919c inputs the training intermediate features to a feature extraction model corrected with the training correction parameters, and extracts training features. The training features are features of a target region in the training input image.
[0282] The loss calculation unit 919d calculates the loss based on the correct labels and learning features included in the training data. As the loss function, a general loss function such as the sum of squares error or the cross entropy error may be used.
[0283] The update unit 919e updates the parameters included in each of the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model based on the loss. This updating may be performed using a general optimization method such as gradient descent or stochastic gradient descent.
[0284] (Example of processing operation of information processing device 900) The information processing performed by the information processing device 600 may further include a learning process as shown in FIG. 33 in addition to the information processing described in the third embodiment. The learning process may be started, for example, in response to an instruction from a user. Training data may be prepared before the learning process is started. Note that the trigger for starting the learning process is not limited to this.
[0285] The learning object information acquisition unit 919a inputs a learning input image to an intermediate feature extraction model to extract learning intermediate features (step S901).
[0286] The learning calculation unit 919b inputs the learning intermediate feature amounts to the second correction parameter estimation model to calculate learning correction parameters (step S902).
[0287] The learning feature extraction unit 919c inputs the learning intermediate feature to the feature extraction model corrected with the learning correction parameters, and extracts learning feature (step S903).
[0288] The loss calculation unit 919d calculates the loss based on the correct labels and learning features included in the training data (step S904).
[0289] The update unit 919e updates the parameters included in the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model based on the loss (step S905), and ends the learning process.
[0290] (Operations and Effects) As described above, according to this embodiment, the information processing device 900 includes the learning unit 919 that uses training data to learn the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model.
[0291] This allows the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model to be trained simultaneously using common training data. This reduces the effort required to prepare training data compared to preparing training data for each model. Furthermore, this reduces the effort required for training compared to training each of the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model individually. This makes it possible to easily train the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model.
[0292] Eleventh Embodiment In this embodiment, a second example of the learning method for the intermediate feature extraction model, the second correction parameter estimation model, and the feature extraction model described in the third embodiment will be described.
[0293] (Configuration example of information processing system SYS10 and information processing device 1000) As shown in Fig. 34 , the information processing system SYS10 includes an imaging device 201 and an information processing device 1000. In addition to the functions of the information processing device 300, the information processing device 1000 includes a first learning unit 1020 and a second learning unit 1021.
[0294] The first learning unit 1020 uses first training data prepared in advance to learn a feature extraction model, and the second learning unit 1021 uses second training data prepared in advance to learn a second correction parameter estimation model.
[0295] The first training data and the second training data may be similar to the first training data and the second training data described in embodiment 8. That is, for example, the first training data may include first input images that are input images for learning and first correct labels. Also, for example, the second training data may include second input images that are input images for learning and second correct labels.
[0296] The target region in the first input image has higher quality than the target region in the second input image. The method for preparing such a second input image may be the same as the method described in the eighth embodiment.
[0297] The first learning unit 1020 and the second learning unit 1021 use first training data and second training data including learning images of different image quality to individually learn the feature extraction model and the first correction parameter estimation model.
[0298] (Example configuration of first learning unit 1020) As shown in FIG. 35, the first learning unit 1020 includes a first learning object information acquisition unit 1020a, a first feature extraction unit 1020b, a first loss calculation unit 1020c, and a first update unit 1020d.
[0299] The first learning target information acquisition unit 1020a inputs the first input image to an intermediate feature extraction model to extract first intermediate features, which are intermediate features of a target region included in the first input image.
[0300] The first feature extraction unit 1020b inputs the first intermediate feature to a feature extraction model to extract the first feature.
[0301] The first loss calculation unit 1020c calculates the first loss based on the first correct label and the first feature amount. As the loss function, a general loss function such as a sum of squares error or a cross entropy error may be used.
[0302] The first update unit 1020d updates the parameters included in each of the intermediate feature extraction model and the feature extraction model based on the first loss. This updating may be performed using a general optimization method such as gradient descent or stochastic gradient descent.
[0303] (Example configuration of second learning unit 1021) As shown in FIG. 36, the second learning unit 1021 includes a second learning target information acquisition unit 1021a, a learning calculation unit 1021b, a second feature extraction unit 1021c, a second loss calculation unit 1021d, and a second update unit 1021e.
[0304] The second learning target information acquisition unit 1021a inputs the second input image to a trained intermediate feature extraction model to extract second intermediate features, which are intermediate features of a target region included in the second input image.
[0305] The learning calculation unit 1021b inputs the second intermediate feature amount to the second correction parameter estimation model to calculate learning correction parameters.
[0306] The second feature extraction unit 1021c inputs the second input image to a trained feature extraction model corrected with the learning correction parameters, and extracts the second feature.
[0307] The second loss calculation unit 1021d calculates the second loss based on the second correct label and the second feature amount. As the loss function, a general loss function such as a sum of squares error or a cross entropy error may be used.
[0308] The second update unit 1021e updates the parameters included in the second correction parameter estimation model based on the second loss. This update may be performed using a general optimization method such as gradient descent or stochastic gradient descent.
[0309] (Example of processing operation of information processing device 1000) The information processing performed by the information processing device 1000 may further include, in addition to the information processing described in embodiment 3, a first learning process and a second learning process as shown in Figures 37 and 38, respectively.
[0310] The first learning process is a process for training an intermediate feature extraction model and a feature extraction model. The second learning process is a process for training a second correction parameter estimation model. The second learning process is performed using the intermediate feature extraction model and the feature extraction model that have been trained by executing the first learning process. Therefore, it is preferable that the second learning process be performed after the first learning process.
[0311] The first learning process may be started, for example, in response to an instruction from a user. The second learning process may be automatically executed following the first learning process, or may be started in response to an instruction from a user. The first training data and the second training data may be prepared before the learning process is started. Note that the trigger for starting the learning process is not limited to this.
[0312] See Fig. 37. The first learning object information acquisition unit 1020a inputs a first input image to an intermediate feature extraction model and extracts first intermediate features (step S1001).
[0313] The first feature extraction unit 1020b inputs the first intermediate feature into a feature extraction model to extract the first feature (step S1002).
[0314] The first loss calculation unit 1020c calculates the first loss based on the first correct label and the first feature amount (step S1003).
[0315] The first updating unit 1020d updates the parameters included in each of the intermediate feature extraction model and the feature extraction model based on the first loss (step S1004), and ends the first learning process. This completes the learning of the intermediate feature extraction model and the feature extraction model.
[0316] (Regarding the Second Learning Process) In the second learning process, the intermediate feature extraction model and feature extraction model learned by executing the first learning process, i.e., the trained intermediate feature extraction model and feature extraction model, are used.
[0317] See Fig. 38. The second learning target information acquisition unit 1021a inputs the second input image to the trained intermediate feature extraction model to extract second intermediate features (step S1011).
[0318] The learning calculation unit 1021b inputs the second intermediate feature amount to the second correction parameter estimation model to calculate learning correction parameters (step S1012).
[0319] The second feature extraction unit 1021c inputs the second input image to the trained feature extraction model corrected with the learning correction parameters, and extracts the second feature (step S1013).
[0320] The second loss calculation unit 1021d calculates the second loss based on the second correct label and the second feature amount (step S1014).
[0321] The second update unit 1021e updates the parameters included in the second correction parameter estimation model based on the second loss (step S1015), and ends the second learning process, thereby completing the learning of the second correction parameter estimation model.
[0322] (Operations and Effects) As described above, according to this embodiment, the information processing device 700 includes a first learning unit 1020 and a second learning unit 1021. The first learning unit 1020 uses first training data to learn a feature extraction model and an intermediate feature extraction model. The second learning unit 1021 uses second training data to learn a second correction parameter estimation model.
[0323] The first training data includes a first input image that is the input image for training, and the second training data includes a second input image that is the input image for training, and the target region in the first input image has higher quality than the target region in the second input image.
[0324] This allows the feature extraction model to be trained using high-quality training images (first input images), and the second correction parameter estimation model to be trained using low-quality training images (second input images). Therefore, the feature extraction model and the second correction parameter estimation model can be trained separately so that the features of the target region can be extracted with high accuracy. This allows for more accurate authentication.
[0325] [Embodiment 12] This embodiment describes a third example of the method for learning the intermediate feature extraction model, the second correction parameter estimation model, and the feature extraction model described in Embodiment 3. As described in Embodiment 8, the intermediate feature extraction model, the feature extraction model, and the first correction parameter estimation model may be learned individually using a feature norm (norm of a feature vector).
[0326] In this case, the “feature extraction model” in the eighth embodiment may be replaced with the “intermediate feature extraction model and feature extraction model.” Also, the “first correction parameter estimation model” in the eighth embodiment may be replaced with the “second correction parameter estimation model.”
[0327] That is, in this embodiment, the first feature extraction unit 720a inputs a first input image to the intermediate feature extraction model and the feature extraction model, and extracts first features that are features of the first input image. The first loss calculation unit 720b calculates the first loss based on the correct labels included in the first training data and the first features.
[0328] The first norm calculation unit 820c calculates a first norm, which is the norm of the first feature. The first norm loss calculation unit 820d calculates a first norm loss based on the ground truth norm included in the first training data and the first norm. The loss integration unit 820e calculates an integrated loss by integrating the first loss and the first norm loss. The first update unit 820f updates parameters included in the feature extraction model based on the integrated loss.
[0329] The learning calculation unit 721a inputs quality information of the second input image to a second correction parameter estimation model to calculate learning correction parameters. The second feature extraction unit 721b inputs the second input image to a trained intermediate feature extraction model and a feature extraction model corrected with the learning correction parameters, and extracts second features that are features of the second input image.
[0330] The second norm calculation unit 821c calculates a second norm, which is the norm of the second feature amount. The second norm loss calculation unit 821d calculates a second norm loss based on the ground truth norm included in the second training data and the second norm. The second update unit 821e updates the parameters included in the second correction parameter estimation model based on the second norm loss.
[0331] This allows the intermediate feature extraction model and the feature extraction model to be trained using high-quality training images (first input images), and the second correction parameter estimation model to be trained using low-quality training images (second input images). Furthermore, as described above, there is a correlation between image quality and the norm of the feature vector extracted from the image. Therefore, the intermediate feature extraction model and the feature extraction model, and the second correction parameter estimation model can be trained separately to enable more accurate extraction of features of the target region. This allows for more accurate authentication.
[0332] [Embodiment 13] In the learning of the correction coefficient estimation model described in embodiment 4, the methods described in embodiments 7 to 12 above can be applied.
[0333] For example, when quality information of the target region is used as target information in the fourth embodiment, the methods described in the seventh to ninth embodiments can be applied to training the correction coefficient estimation model. In detail, for example, the "first correction parameter estimation model" in the seventh to ninth embodiments may be replaced with a "correction coefficient estimation model." During training, the dictionary parameters and second relationship set in advance may also be used as necessary.
[0334] More specifically, for example, when the method described in the seventh embodiment is applied to learning a correction coefficient estimation model, the learning unit 619 may use training data to learn the feature extraction model and the correction coefficient estimation model.
[0335] When the method described in the eighth embodiment is applied to learning the correction coefficient estimation model, the second learning unit 721 may use the second training data to learn the second correction parameter estimation model.
[0336] When the method described in the ninth embodiment is applied to training of a correction coefficient estimation model, the training calculation unit 721a may input quality information of the second input image to the correction coefficient estimation model to calculate training correction parameters. The second update unit 821e may update the parameters included in the correction coefficient estimation model based on the second norm loss.
[0337] This provides substantially the same effects as those of the seventh to ninth embodiments.
[0338] For example, when intermediate feature amounts of a target region are used as target information in the fourth embodiment, the methods described in the tenth to twelfth embodiments can be applied to training of a correction coefficient estimation model. In detail, for example, the "second correction parameter estimation model" in the tenth to twelfth embodiments may be replaced with a "correction coefficient estimation model." During training, the dictionary parameters and second relationship set in advance may also be used as necessary.
[0339] More specifically, for example, when the method described in the tenth embodiment is applied to learning a correction coefficient estimation model, the learning unit 919 may use training data to learn a feature extraction model, a correction coefficient estimation model, and an intermediate feature extraction model.
[0340] When the method described in the eleventh embodiment is applied to learning the correction coefficient estimation model, the second learning unit 1021 may use the second training data to learn the correction coefficient estimation model.
[0341] When the method described in the twelfth embodiment is applied to training of a correction coefficient estimation model, the training calculation unit 721a may input quality information of the second input image to the correction coefficient estimation model to calculate training correction parameters. The second update unit 821e may update the parameters included in the correction coefficient estimation model based on the second norm loss.
[0342] This provides substantially the same effects as those of the tenth to twelfth embodiments.
[0343] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0344] In addition, although the flowcharts used in the above description show a sequence of steps (processes), the order of steps executed in each embodiment is not limited to the sequence shown in the flowcharts. In each embodiment, the order of steps shown in the diagrams can be changed as long as it does not cause any problems in terms of the content.
[0345] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes. 1. An information processing device comprising: object information acquisition means for acquiring object information related to a target region in an input image; calculation means for calculating, based on the object information, correction parameters for correcting parameters included in a feature extraction model that extracts features of the target region; extraction means for extracting features of the target region using the feature extraction model corrected with the correction parameters; and comparison means for outputting a result of comparing the features of the target region with pre-registered information. 2. The information processing device described in 1., wherein the input image is an image including an iris region that is the target region; the feature extraction model is composed of a neural network including at least one normalization layer; and the parameters corrected using the correction parameters are parameters used in the normalization layer. 3. The information processing device described in 1. or 2., wherein the object information includes at least one of quality information of the target region, intermediate features of the target region, and statistics related to the target region. 4. The information processing device described in 3., wherein the target information includes quality information estimated by inputting the target region into a quality estimation model for estimating the quality information of the target region, and the correction parameter is calculated by inputting the target information into a first correction parameter estimation model for estimating the correction parameter. 5. The information processing device described in 3., wherein the target information includes intermediate features extracted by inputting the target region into an intermediate feature extraction model for extracting the intermediate features of the target region (in the input image), and the correction parameter is calculated by inputting the target information into a second correction parameter estimation model for estimating the correction parameter. 6. The information processing device described in 3., wherein the target information includes the quality information of the target region or the intermediate features of the target region, and the correction parameter is calculated by applying a correction coefficient acquired by inputting the target information into a correction coefficient estimation model and a dictionary parameter stored in advance to a predetermined first relationship.7. The information processing device described in 3., wherein the target information includes the statistics related to the target region calculated using the intermediate features of the target region, and the correction parameters are calculated by applying the statistics related to the target region and pre-stored dictionary parameters to a predetermined second relationship. 8. The information processing device described in 3., wherein the target information includes the intermediate features of the target region and the statistics related to the target region, and the correction parameters are calculated based on the intermediate features of the target region and the statistics related to the target region. 9. The information processing device described in 4., further comprising learning means for learning the feature extraction model and the first correction parameter estimation model using training data. 10. 4. The information processing device according to claim 4, comprising: a first learning means for learning the feature extraction model using first training data; and a second learning means for learning the first correction parameter estimation model using second training data, wherein the first training data includes a first input image that is the input image for learning; the second training data includes a second input image that is the input image for learning; and the target region in the first input image is of higher quality than the target region in the second input image.11. the first learning means includes: a first feature extraction means for inputting the first input image into the feature extraction model and extracting a first feature that is a feature of the first input image; a first loss calculation means for calculating a first loss based on a correct label included in the first training data and the first feature; a first norm calculation means for calculating a first norm that is a norm of the first feature; a first norm loss calculation means for calculating a first norm loss based on a correct norm included in the first training data and the first norm; a loss integration means for calculating an integrated loss by integrating the first loss and the first norm loss; and a first update means for updating parameters included in the feature extraction model based on the integrated loss; and the second learning means includes: a learning calculation means for inputting quality information of a second input image into the first correction parameter estimation model and calculating learning correction parameters; a second feature extraction means for inputting the second input image into the trained feature extraction model corrected with the learning correction parameters and extracting a second feature that is a feature of the second input image; 10. The information processing device according to 10., comprising: second norm calculation means for calculating a second norm which is a norm of the second feature amount; second norm loss calculation means for calculating a second norm loss based on the second norm and a ground truth norm included in the second training data; and second update means for updating parameters included in the first correction parameter estimation model based on the second norm loss. 12. The information processing device according to 5., further comprising learning means for training the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model using training data. 13. The information processing device according to 5., comprising: first learning means for training the feature extraction model and the intermediate feature extraction model using first training data; and second learning means for training the second correction parameter estimation model using second training data, wherein the first training data includes a first input image which is the input image for training, the second training data includes a second input image which is the input image for training, and the target region in the first input image is of higher quality than the target region in the second input image. The information processing device described in14. An information processing system comprising: a target information acquisition means for acquiring target information related to a target region in an input image; a calculation means for calculating, based on the target information, correction parameters for correcting parameters included in a feature extraction model that extracts features of the target region; an extraction means for extracting features of the target region using the feature extraction model corrected with the correction parameters; and a matching means for outputting a result of matching the features of the target region with pre-registered information. 15. The information processing system described in 14., wherein the input image is an image including an iris region that is the target region; the feature extraction model is composed of a neural network including at least one normalization layer; and the parameters corrected using the correction parameters are parameters used in the normalization layer. 16. The information processing system described in 14. or 15., wherein the target information includes at least one of quality information of the target region, intermediate features of the target region, and statistics related to the target region. 17. The information processing system described in 16., wherein the target information includes the quality information estimated by inputting the target region into a quality estimation model for estimating the quality information of the target region, and the correction parameter is calculated by inputting the target information into a first correction parameter estimation model for estimating the correction parameter. 18. The information processing system described in 16., wherein the target information includes the intermediate feature extracted by inputting the target region into an intermediate feature extraction model for extracting the intermediate feature of the target region (in the input image), and the correction parameter is calculated by inputting the target information into a second correction parameter estimation model for estimating the correction parameter. 19. The information processing system described in 16., wherein the target information includes the quality information of the target region or the intermediate feature of the target region, and the correction parameter is calculated by applying a correction coefficient acquired by inputting the target information into a correction coefficient estimation model and a dictionary parameter stored in advance to a predetermined first relationship.20. The information processing system described in 16., wherein the target information includes the statistics related to the target region calculated using the intermediate features of the target region, and the correction parameters are calculated by applying the statistics related to the target region and pre-stored dictionary parameters to a predetermined second relationship. 21. The information processing system described in 16., wherein the target information includes the intermediate features of the target region and the statistics related to the target region, and the correction parameters are calculated based on the intermediate features of the target region and the statistics related to the target region. 22. The information processing system described in 17., further comprising learning means for learning the feature extraction model and the first correction parameter estimation model using training data. 23. 17. An information processing system according to claim 17, comprising: a first learning means for learning the feature extraction model using first training data; and a second learning means for learning the first correction parameter estimation model using second training data, wherein the first training data includes first input images that are the input images for learning, the second training data includes second input images that are the input images for learning, and the target region in the first input image is of higher quality than the target region in the second input image.24. the first learning means includes: a first feature extraction means for inputting the first input image into the feature extraction model and extracting a first feature that is a feature of the first input image; a first loss calculation means for calculating a first loss based on a correct label included in the first training data and the first feature; a first norm calculation means for calculating a first norm that is a norm of the first feature; a first norm loss calculation means for calculating a first norm loss based on a correct norm included in the first training data and the first norm; a loss integration means for calculating an integrated loss by integrating the first loss and the first norm loss; and a first update means for updating parameters included in the feature extraction model based on the integrated loss; and the second learning means includes: a learning calculation means for inputting quality information of a second input image into the first correction parameter estimation model and calculating learning correction parameters; a second feature extraction means for inputting the second input image into the trained feature extraction model corrected with the learning correction parameters and extracting a second feature that is a feature of the second input image; 23. The information processing system according to 23, comprising: second norm calculation means for calculating a second norm which is the norm of the second feature; second norm loss calculation means for calculating a second norm loss based on the second norm and a ground truth norm included in the second training data; and second update means for updating parameters included in the first correction parameter estimation model based on the second norm loss. 25. The information processing system according to 18, further comprising learning means for training the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model using training data. 26. The information processing system according to 18, comprising: first learning means for training the feature extraction model and the intermediate feature extraction model using first training data; and second learning means for training the second correction parameter estimation model using second training data, wherein the first training data includes a first input image which is the input image for training, the second training data includes a second input image which is the input image for training, and the target region in the first input image is of higher quality than the target region in the second input image. An information processing system according to claim 1.27. An information processing method in which one or more computers acquire object information regarding a target region in an input image, calculate correction parameters for correcting parameters included in a feature extraction model that extracts features of the target region based on the object information, extract features of the target region using the feature extraction model corrected with the correction parameters, and output a result of matching the features of the target region with pre-registered registration information. 28. The information processing method described in 27., in which the input image is an image including an iris region that is the target region, the feature extraction model is composed of a neural network including at least one normalization layer, and the parameters corrected using the correction parameters are parameters used in the normalization layer. 29. The information processing method described in 27. or 28., in which the object information includes at least one of quality information of the target region, intermediate features of the target region, and statistics related to the target region. 30. The information processing method according to 29., wherein the target information includes quality information estimated by inputting the target region into a quality estimation model for estimating the quality information of the target region, and the correction parameters are calculated by inputting the target information into a first correction parameter estimation model for estimating the correction parameters. 31. The information processing method according to 29., wherein the target information includes intermediate features extracted by inputting the target region into an intermediate feature extraction model for extracting the intermediate features of the target region (in the input image), and the correction parameters are calculated by inputting the target information into a second correction parameter estimation model for estimating the correction parameters. 32. The information processing method according to 29., wherein the target information includes the quality information of the target region or the intermediate features of the target region, and the correction parameters are calculated by applying correction coefficients acquired by inputting the target information into a correction coefficient estimation model and pre-stored dictionary parameters to a predetermined first relationship.33. The information processing method described in 29., wherein the target information includes the statistics related to the target region calculated using the intermediate features of the target region, and the correction parameters are calculated by applying the statistics related to the target region and pre-stored dictionary parameters to a predetermined second relationship. 34. The information processing method described in 29., wherein the target information includes the intermediate features of the target region and the statistics related to the target region, and the correction parameters are calculated based on the intermediate features of the target region and the statistics related to the target region. 35. The information processing method described in 30., further comprising training the feature extraction model and the first correction parameter estimation model using training data. 36. 30. The information processing method described in Item 30, further comprising: training the feature extraction model using first training data; and training the first correction parameter estimation model using second training data, wherein the first training data includes first input images that are the input images for training; the second training data includes second input images that are the input images for training; and the target region in the first input image is of higher quality than the target region in the second input image.37. The training of the feature extraction model includes: inputting the first input image into the feature extraction model to extract first features that are features of the first input image; calculating a first loss based on a correct label included in the first training data and the first features; calculating a first norm that is a norm of the first features; calculating a first norm loss based on the correct norm included in the first training data and the first norm; loss integration means that calculates an integrated loss by integrating the first loss and the first norm loss; and updating parameters included in the feature extraction model based on the integrated loss; and the training of the first correction parameter estimation model includes: inputting quality information of a second input image into the first correction parameter estimation model to calculate learning correction parameters; inputting the second input image into the trained feature extraction model that has been corrected with the learning correction parameters to extract second features that are features of the second input image; and calculating a second norm that is the norm of the second features. The information processing method according to 36., which includes calculating a second norm loss based on a ground truth norm included in the second training data and the second norm, and updating parameters included in the first correction parameter estimation model based on the second norm loss. 38. The information processing method according to 31., which further includes training the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model using training data. 39. The information processing method according to 31., which further includes training the feature extraction model and the intermediate feature extraction model using first training data, and training the second correction parameter estimation model using second training data, wherein the first training data includes first input images that are the input images for training, the second training data includes second input images that are the input images for training, and the target region in the first input image is of higher quality than the target region in the second input image.40. A program causing one or more computers to execute the following steps: acquire object information regarding a target region in an input image; calculate, based on the object information, correction parameters for correcting parameters included in a feature extraction model that extracts features of the target region; extracting features of the target region using the feature extraction model corrected with the correction parameters; and outputting a result of matching the features of the target region with pre-registered registration information. 41. The program described in 40., wherein the input image is an image including an iris region that is the target region; the feature extraction model is composed of a neural network including at least one normalization layer; and the parameters corrected using the correction parameters are parameters used in the normalization layer. 42. The program described in 40. or 41., wherein the object information includes at least one of quality information of the target region, intermediate features of the target region, and statistics related to the target region. 43. The program described in 42., wherein the target information includes quality information estimated by inputting the target region into a quality estimation model for estimating the quality information of the target region, and the correction parameters are calculated by inputting the target information into a first correction parameter estimation model for estimating the correction parameters. 44. The program described in 43., wherein the target information includes intermediate features extracted by inputting the target region into an intermediate feature extraction model for extracting the intermediate features of the target region (in the input image), and the correction parameters are calculated by inputting the target information into a second correction parameter estimation model for estimating the correction parameters. 45. The program described in 44., wherein the target information includes the quality information of the target region or the intermediate features of the target region, and the correction parameters are calculated by applying correction coefficients acquired by inputting the target information into a correction coefficient estimation model and pre-stored dictionary parameters to a predetermined first relationship.46. The program according to 42., wherein the target information includes the statistics related to the target region calculated using the intermediate features of the target region, and the correction parameters are calculated by applying the statistics related to the target region and pre-stored dictionary parameters to a predetermined second relationship. 47. The program according to 42., wherein the target information includes the intermediate features of the target region and the statistics related to the target region, and the correction parameters are calculated based on the intermediate features of the target region and the statistics related to the target region. 48. The program according to 43., further causing the program to execute learning of the feature extraction model and the first correction parameter estimation model using training data. 49. 43. The program described in 43. further executes training of the feature extraction model using first training data, and training of the first correction parameter estimation model using second training data, wherein the first training data includes first input images that are the input images for training, the second training data includes second input images that are the input images for training, and the target region in the first input image is of higher quality than the target region in the second input image.50. The training of the feature extraction model includes: inputting the first input image into the feature extraction model to extract first features that are features of the first input image; calculating a first loss based on a correct label included in the first training data and the first features; calculating a first norm that is a norm of the first features; calculating a first norm loss based on the correct norm included in the first training data and the first norm; loss integration means that calculates an integrated loss by integrating the first loss and the first norm loss; and updating parameters included in the feature extraction model based on the integrated loss; and the training of the first correction parameter estimation model includes: inputting quality information of a second input image into the first correction parameter estimation model to calculate learning correction parameters; inputting the second input image into the trained feature extraction model that has been corrected with the learning correction parameters to extract second features that are features of the second input image; and calculating a second norm that is a norm of the second features. The program described in 49., which includes calculating a second norm loss based on a ground truth norm included in the second training data and the second norm, and updating parameters included in the first correction parameter estimation model based on the second norm loss. 51. The program described in 44., which is further executed to train the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model using training data. 52. The program described in 44., which is further executed to train the feature extraction model and the intermediate feature extraction model using first training data, and to train the second correction parameter estimation model using second training data, wherein the first training data includes first input images that are the input images for training, the second training data includes second input images that are the input images for training, and the target region in the first input image is of higher quality than the target region in the second input image.53. A recording medium having recorded thereon a program for causing one or more computers to acquire object information regarding a target region in an input image, calculate correction parameters for correcting parameters included in a feature extraction model that extracts features of the target region based on the object information, extract features of the target region using the feature extraction model corrected with the correction parameters, and output a result of matching the features of the target region with pre-registered registration information. 54. The recording medium described in 53., having recorded thereon a program in which the input image is an image including an iris region that is the target region, the feature extraction model is composed of a neural network including at least one normalization layer, and the parameters corrected using the correction parameters are parameters used in the normalization layer. 55. The recording medium described in 53. or 54., having recorded thereon a program in which the object information includes at least one of quality information of the target region, intermediate features of the target region, and statistics related to the target region. 56. The recording medium described in 55., on which a program is recorded in which the target information includes quality information estimated by inputting the target region into a quality estimation model for estimating the quality information of the target region, and the correction parameters are calculated by inputting the target information into a first correction parameter estimation model for estimating the correction parameters. 57. The recording medium described in 56., on which a program is recorded in which the target information includes intermediate features extracted by inputting the target region into an intermediate feature extraction model for extracting the intermediate features of the target region (in the input image), and the correction parameters are calculated by inputting the target information into a second correction parameter estimation model for estimating the correction parameters. 58. The recording medium described in 57., on which a program is recorded in which the target information includes the quality information of the target region or the intermediate features of the target region, and the correction parameters are calculated by applying correction coefficients obtained by inputting the target information into a correction coefficient estimation model and pre-stored dictionary parameters to a predetermined first relationship.59. The recording medium described in 55., on which a program is recorded, in which the target information includes the statistics related to the target region calculated using the intermediate features of the target region, and the correction parameters are calculated by applying the statistics related to the target region and pre-stored dictionary parameters to a predetermined second relationship. 60. The recording medium described in 55., on which a program is recorded, in which the target information includes the intermediate features of the target region and the statistics related to the target region, and the correction parameters are calculated based on the intermediate features of the target region and the statistics related to the target region. 61. The recording medium described in 56., on which a program is recorded for further executing learning of the feature extraction model and the first correction parameter estimation model using training data. 62. 56. The recording medium described in Item 56, on which a program is recorded that further executes training of the feature extraction model using first training data and training of the first correction parameter estimation model using second training data, wherein the first training data includes first input images that are the input images for training, the second training data includes second input images that are the input images for training, and the target region in the first input image has higher quality than the target region in the second input image.63. The training of the feature extraction model includes: inputting the first input image into the feature extraction model to extract first features that are features of the first input image; calculating a first loss based on a correct label included in the first training data and the first features; calculating a first norm that is a norm of the first features; calculating a first norm loss based on the correct norm included in the first training data and the first norm; loss integration means that calculates an integrated loss by integrating the first loss and the first norm loss; and updating parameters included in the feature extraction model based on the integrated loss; and the training of the first correction parameter estimation model includes: inputting quality information of a second input image into the first correction parameter estimation model to calculate learning correction parameters; inputting the second input image into the trained feature extraction model that has been corrected with the learning correction parameters to extract second features that are features of the second input image; and calculating a second norm that is a norm of the second features. The recording medium described in 62. has recorded thereon a program including calculating a second norm loss based on a ground truth norm included in the second training data and the second norm, and updating parameters included in the first correction parameter estimation model based on the second norm loss. 64. The recording medium described in 57. has recorded thereon a program for further executing training of the feature extraction model, the second correction parameter estimation model, and the intermediate feature extraction model using training data. 65. The recording medium described in 57. has recorded thereon a program for further executing training of the feature extraction model and the intermediate feature extraction model using first training data, and training of the second correction parameter estimation model using second training data, wherein the first training data includes first input images that are the input images for training, the second training data includes second input images that are the input images for training, and the target region in the first input image is of higher quality than the target region in the second input image.
[0346] SYS1 to SYS10 Information processing systems 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 Information processing device 115 Object information acquisition unit 116 Calculation unit 117 Extraction unit 118 Collation unit 201 Imaging device 211 Object position estimation unit 212 Object area generation unit 215, 315, 415, 515 Object information acquisition unit 216, 316, 416, 516, 616 Calculation unit 416a Dictionary parameter storage unit 416b Correction coefficient acquisition unit 416c Correction parameter calculation unit 516b Correction parameter calculation unit 619, 919 Learning unit 720, 820, 1020 First learning unit 721, 821, 1021 Second learning unit
Claims
1. A target information acquisition means for acquiring target information relating to a target region in an input image, A calculation means for calculating correction parameters for correcting the parameters included in the feature extraction model that extracts features of the target region based on the aforementioned target information, An extraction means for extracting features of the target region using the feature extraction model corrected with the correction parameters, The system includes a matching means that outputs the result of comparing the feature quantities of the target region with pre-registered registration information. Information processing device.
2. The input image is an image that includes the iris region, which is the target region. The feature extraction model consists of a neural network including at least one normalization layer. The parameter corrected using the aforementioned correction parameter is the parameter used in the normalization layer. The information processing apparatus according to claim 1.
3. The aforementioned target information includes at least one of the quality information of the target region, the intermediate features of the target region, and statistics relating to the target region. The information processing apparatus according to claim 1 or 2.
4. The aforementioned target information includes the quality information estimated by inputting the target area into a quality estimation model for estimating the quality information of the target area, The correction parameter is calculated by inputting the target information into a first correction parameter estimation model for estimating the correction parameter. The information processing apparatus according to claim 3.
5. The aforementioned target information includes the intermediate features extracted by inputting the target region into an intermediate feature extraction model for extracting the intermediate features of the target region (in the input image), The aforementioned correction parameter is calculated by inputting the target information into a second correction parameter estimation model for estimating the aforementioned correction parameter. The information processing apparatus according to claim 3.
6. The aforementioned target information includes the quality information of the target region or the intermediate feature quantities of the target region, The aforementioned correction parameter is calculated by applying the correction coefficient obtained by inputting the target information into the correction coefficient estimation model, and the pre-stored dictionary parameters, to a predetermined first relationship. The information processing apparatus according to claim 3.
7. The aforementioned target information includes the aforementioned statistics relating to the target region calculated using the aforementioned intermediate features of the target region, The correction parameter is calculated by applying the statistic relating to the target region and the pre-stored dictionary parameter to a predetermined second relationship. The information processing apparatus according to claim 3.
8. The aforementioned target information includes the intermediate feature quantities of the target region and the statistical quantities relating to the target region, The correction parameter is calculated based on the intermediate features of the target region and the statistics relating to the target region. The information processing apparatus according to claim 3.
9. One or more computers, Obtain target information regarding the target region in the input image. Based on the aforementioned target information, correction parameters are calculated to correct the parameters included in the feature extraction model that extracts the features of the target region. Using the feature extraction model corrected with the aforementioned correction parameters, the features of the target region are extracted. The results of comparing the feature quantities of the target region with pre-registered registration information are output. Information processing methods.
10. On one or more computers, Obtain target information regarding the target region in the input image. Based on the aforementioned target information, correction parameters are calculated to correct the parameters included in the feature extraction model that extracts the features of the target region. Using the feature extraction model corrected with the aforementioned correction parameters, the features of the target region are extracted. A program for outputting the results of comparing the feature quantities of the target region with pre-registered registration information.